Skip to content

encoding/hex: vectorize Encode/Decode - #82089

Open
AskAlexSharov wants to merge 5 commits into
golang:masterfrom
AskAlexSharov:hex-simd
Open

AskAlexSharov wants to merge 5 commits into
golang:masterfrom
AskAlexSharov:hex-simd

Conversation

@AskAlexSharov

@AskAlexSharov AskAlexSharov commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

simd and simdarch (arm64 and amd64) implementations of hex encode/decode


Implementation details

On amd64 Encode widens each byte 0xHL to the uint16 0x00HL and
multiplies it by 0x1001, which gives 0xL0HL; shifted right by 4 it is
0x0L0H, the two nibbles in output order, and one VPSHUFB turns them into
digits. Decode computes the nibbles and checks the characters in one
pass, by algorithm 3 of
http://0x80.pl/notesen/2022-01-17-validating-hex-parse.html, and packs
them with VPMADDUBSW and VPSHUFB.

On arm64 Encode looks up both nibbles with TBL and interleaves the
digits with ZIP1 and ZIP2. Decode computes the nibbles as on amd64 and
packs them with UZP1 and UZP2.


GiB/s (Decode counted in input characters):

amd64, AMD EPYC 4344P (Zen 4)
                  scalar  archsimd  portable 256-bit  portable 512-bit
Encode 16B          1.46      3.72                 -                 -
Encode 64B          1.53     10.66              2.62              3.07
Encode 256B         1.51     20.94              3.43              4.22
Encode 4KB          1.55     29.18              3.80              4.75
Encode 128KB        1.54     29.74              3.81              4.79
Decode 64B          3.27      8.10              3.69              4.25
Decode 256B         3.45     16.43              5.62              7.02
Decode 4KB          3.51     24.14              6.65              8.77
Decode 128KB        3.52     24.90              6.78              8.95

arm64, Apple M4 Max
                  scalar  archsimd  portable 128-bit
Encode 16B          1.61      4.98              1.41
Encode 64B          1.79     13.00              2.14
Encode 256B         1.77     19.68              2.47
Encode 4KB          1.88     22.91              2.59
Encode 128KB        1.88     23.80              2.60
Decode 64B          3.24     12.24              3.57
Decode 256B         3.35     21.04              4.58
Decode 4KB          3.71     21.74              4.93
Decode 128KB        3.72     21.64              5.00

The portable kernels are slower because every 64-bit lane works alone
and scalar stores join the lanes. What portable simd package lacks:

  • interleaving the bytes of two vectors, or zero-extending half of a
    Uint8s to Uint16s (VPUNPCKLBW or VPMOVZXBW; ZIP1 or UXTL), for Encode
  • narrowing Uint16s to Uint8s, or taking the even bytes of two vectors
    (VPACKUSWB; UZP1 or XTN), for Decode
  • a byte table lookup (VPSHUFB; TBL), which saves a few operations in
    Encode
  • comparisons and shifts on Uint8s, which only Int8s and wider types
    have
  • clearing the upper bits of the AVX registers before returning:
    archsimd.ClearAVXUpperBits has no portable counterpart, and the
    compiler does not insert VZEROUPPER (cmd/compile: legacy SSE encodings emitted in functions using simd/archsimd intrinsics cause AVX-SSE transition penalties #80835)

All benchmarks, scalar against archsimd:

goos: linux
goarch: amd64
pkg: encoding/hex
cpu: AMD EPYC 4344P 8-Core Processor
                   │    scalar     │                archsimd                 │
                   │      B/s      │      B/s       vs base                  │
Encode/16B           1.464Gi ±  0%    4.639Gi ± 0%   +216.75% (p=0.000 n=10)
Encode/32B           1.506Gi ±  1%    7.457Gi ± 0%   +395.02% (p=0.000 n=10)
Encode/64B           1.533Gi ±  1%   12.405Gi ± 0%   +709.25% (p=0.000 n=10)
Encode/256B          1.507Gi ±  0%   22.174Gi ± 0%  +1371.23% (p=0.000 n=10)
Encode/1KB           1.535Gi ±  1%   27.576Gi ± 1%  +1696.38% (p=0.000 n=10)
Encode/4KB           1.546Gi ±  1%   29.313Gi ± 0%  +1795.86% (p=0.000 n=10)
Encode/16KB          1.538Gi ±  1%   29.410Gi ± 0%  +1812.22% (p=0.000 n=10)
Encode/128KB         1.538Gi ±  0%   29.786Gi ± 0%  +1837.30% (p=0.000 n=10)
Decode/32B           3.045Gi ±  5%    5.987Gi ± 0%    +96.64% (p=0.000 n=10)
Decode/64B           3.268Gi ±  4%    9.675Gi ± 0%   +196.03% (p=0.000 n=10)
Decode/128B          3.416Gi ±  3%   13.950Gi ± 0%   +308.39% (p=0.000 n=10)
Decode/256B          3.451Gi ±  2%   17.958Gi ± 1%   +420.45% (p=0.000 n=10)
Decode/1KB           3.462Gi ±  4%   22.729Gi ± 0%   +556.49% (p=0.000 n=10)
Decode/4KB           3.512Gi ±  2%   24.390Gi ± 1%   +594.48% (p=0.000 n=10)
Decode/16KB          3.533Gi ±  4%   24.587Gi ± 1%   +595.99% (p=0.000 n=10)
Decode/128KB         3.521Gi ±  2%   24.881Gi ± 0%   +606.61% (p=0.000 n=10)
DecodeString/256B    2.696Gi ± 13%    4.024Gi ± 1%    +49.26% (p=0.000 n=10)
DecodeString/1KB     2.725Gi ± 10%    6.519Gi ± 1%   +139.24% (p=0.000 n=10)
DecodeString/4KB     2.816Gi ± 13%    7.559Gi ± 2%   +168.39% (p=0.000 n=10)
DecodeString/16KB    2.825Gi ±  2%    8.229Gi ± 1%   +191.25% (p=0.000 n=10)
DecodeString/128KB   2.910Gi ±  1%   10.135Gi ± 1%   +248.32% (p=0.000 n=10)
geomean              2.387Gi          13.50Gi        +465.62%

goos: darwin
goarch: arm64
pkg: encoding/hex
cpu: Apple M4 Max
                      │    scalar    │                archsimd                 │
                      │     B/s      │      B/s       vs base                  │
Encode/16B              1.607Gi ± 2%    5.166Gi ± 3%   +221.38% (p=0.000 n=10)
Encode/32B              1.712Gi ± 1%    8.785Gi ± 4%   +413.09% (p=0.000 n=10)
Encode/64B              1.785Gi ± 1%   13.576Gi ± 3%   +660.41% (p=0.000 n=10)
Encode/256B             1.774Gi ± 1%   19.933Gi ± 1%  +1023.41% (p=0.000 n=10)
Encode/1KB              1.855Gi ± 3%   22.690Gi ± 2%  +1122.98% (p=0.000 n=10)
Encode/4KB              1.876Gi ± 2%   22.998Gi ± 2%  +1126.17% (p=0.000 n=10)
Encode/16KB             1.883Gi ± 1%   23.197Gi ± 3%  +1131.93% (p=0.000 n=10)
Encode/128KB            1.879Gi ± 2%   23.822Gi ± 0%  +1168.08% (p=0.000 n=10)
Decode/32B              2.905Gi ± 2%    9.607Gi ± 4%   +230.68% (p=0.000 n=10)
Decode/64B              3.239Gi ± 1%   14.035Gi ± 1%   +333.35% (p=0.000 n=10)
Decode/128B             3.471Gi ± 1%   18.724Gi ± 3%   +439.40% (p=0.000 n=10)
Decode/256B             3.346Gi ± 2%   21.018Gi ± 1%   +528.18% (p=0.000 n=10)
Decode/1KB              3.625Gi ± 2%   21.533Gi ± 1%   +493.94% (p=0.000 n=10)
Decode/4KB              3.712Gi ± 4%   21.813Gi ± 2%   +487.68% (p=0.000 n=10)
Decode/16KB             3.716Gi ± 2%   21.872Gi ± 3%   +488.54% (p=0.000 n=10)
Decode/128KB            3.722Gi ± 2%   21.763Gi ± 1%   +484.75% (p=0.000 n=10)
DecodeString/256B       2.917Gi ± 3%    5.000Gi ± 4%    +71.43% (p=0.000 n=10)
DecodeString/1KB        3.322Gi ± 5%    6.455Gi ± 3%    +94.32% (p=0.000 n=10)
DecodeString/4KB        3.458Gi ± 4%    6.754Gi ± 5%    +95.34% (p=0.000 n=10)
DecodeString/16KB       3.505Gi ± 2%    7.618Gi ± 7%   +117.33% (p=0.000 n=10)
DecodeString/128KB      3.599Gi ± 3%    7.965Gi ± 4%   +121.34% (p=0.000 n=10)
geomean                 2.672Gi         13.51Gi        +405.60%

With GOEXPERIMENT=simd on a CPU with AVX2, Encode handles whole
16-byte blocks and Decode whole 32-character blocks with
simd/archsimd. Shorter inputs and the tails use the scalar loops.

Encode widens each byte 0xHL to the uint16 0x00HL and multiplies it
by 0x1001, which gives 0xL0HL; shifted right by 4 it is 0x0L0H, the
two nibbles in output order, and one VPSHUFB turns them into digits.

Decode computes the nibbles and checks the characters in one pass,
by algorithm 3 of
http://0x80.pl/notesen/2022-01-17-validating-hex-parse.html, and
packs them with VPMADDUBSW and VPSHUFB. A block with an invalid
character is left to the scalar loop, so the error and the partial
output do not change.

goos: linux
goarch: amd64
pkg: encoding/hex
cpu: AMD EPYC 4344P 8-Core Processor
                   │      old      │                   new                   │
                   │      B/s      │      B/s       vs base                  │
Encode/16            1.464Gi ±  0%    4.648Gi ± 1%   +217.53% (p=0.000 n=10)
Encode/32            1.508Gi ±  0%    7.850Gi ± 0%   +420.42% (p=0.000 n=10)
Encode/64            1.535Gi ±  0%   12.445Gi ± 0%   +710.69% (p=0.000 n=10)
Encode/256           1.508Gi ±  0%   22.212Gi ± 0%  +1372.80% (p=0.000 n=10)
Encode/1024          1.539Gi ±  0%   27.676Gi ± 0%  +1698.23% (p=0.000 n=10)
Encode/4096          1.547Gi ±  0%   29.456Gi ± 0%  +1803.71% (p=0.000 n=10)
Encode/16384         1.539Gi ±  0%   29.504Gi ± 1%  +1817.25% (p=0.000 n=10)
Decode/32            3.089Gi ±  4%    6.009Gi ± 0%    +94.51% (p=0.000 n=10)
Decode/64            3.314Gi ±  3%    9.699Gi ± 0%   +192.72% (p=0.000 n=10)
Decode/128           3.384Gi ±  3%   13.978Gi ± 0%   +313.03% (p=0.000 n=10)
Decode/256           3.499Gi ±  2%   17.957Gi ± 0%   +413.25% (p=0.000 n=10)
Decode/1024          3.484Gi ±  1%   22.818Gi ± 0%   +554.90% (p=0.000 n=10)
Decode/4096          3.529Gi ±  2%   24.472Gi ± 0%   +593.40% (p=0.000 n=10)
Decode/16384         3.538Gi ±  5%   24.763Gi ± 0%   +599.96% (p=0.000 n=10)
DecodeString/256     2.702Gi ± 13%    4.125Gi ± 1%    +52.63% (p=0.000 n=10)
DecodeString/1024    2.834Gi ±  6%    6.652Gi ± 1%   +134.69% (p=0.000 n=10)
DecodeString/4096    2.813Gi ±  1%    7.732Gi ± 1%   +174.85% (p=0.000 n=10)
DecodeString/16384   2.842Gi ±  6%    8.366Gi ± 1%   +194.40% (p=0.000 n=10)
geomean              2.381Gi          12.82Gi        +438.29%

The portable simd package cannot express these kernels yet: it cannot
move bytes between lanes. Running the same nibble arithmetic in each
64-bit lane and joining the lanes with scalar stores reaches, for
16 KiB on the same machine:

           Encode     Decode
  512-bit  4.79GiB/s  4.34GiB/s
  256-bit  3.83GiB/s  3.28GiB/s
  128-bit  2.68GiB/s  2.20GiB/s

(Decode counted in output bytes), that is 1.2x-3.1x the scalar loops
but 9-38% of these kernels, which reach 29.5GiB/s and 11.3GiB/s
counted the same way. The simd package would need:

  - interleaving the bytes of two vectors, or zero-extending half of
    a Uint8s to Uint16s (VPUNPCKLBW or VPMOVZXBW, ZIP1 or UXTL), for
    Encode;
  - narrowing Uint16s to Uint8s, or taking the even bytes of two
    vectors (VPACKUSWB, UZP1 or XTN), for Decode;
  - a byte table lookup (VPSHUFB, TBL), which saves a few operations
    in Encode;
  - comparisons and shifts on Uint8s, which only Int8s and wider
    types have;
  - clearing the upper bits of the AVX registers before returning:
    archsimd.ClearAVXUpperBits has no portable counterpart, and the
    compiler does not insert VZEROUPPER (golang#80835).
@AskAlexSharov AskAlexSharov changed the title encoding/hex: vectorize Encode and Decode on amd64 encoding/hex: vectorize Encode and Decode on amd64 and arm64 Oct 9, 2026
@gopherbot

Copy link
Copy Markdown
Contributor

This PR (HEAD: cbdb9c7) has been imported to Gerrit for code review.

Please visit Gerrit at https://go-review.googlesource.com/c/go/+/847806.

Important tips:

  • Don't comment on this PR. All discussion takes place in Gerrit.
  • You need a Gmail or other Google account to log in to Gerrit.
  • To change your code in response to feedback:
    • Push a new commit to the branch used by your GitHub PR.
    • A new "patch set" will then appear in Gerrit.
    • Respond to each comment by marking as Done in Gerrit if implemented as suggested. You can alternatively write a reply.
    • Critical: you must click the blue Reply button near the top to publish your Gerrit responses.
    • Multiple commits in the PR will be squashed by GerritBot.
  • The title and description of the GitHub PR are used to construct the final commit message.
    • Edit these as needed via the GitHub web interface (not via Gerrit or git).
    • You should word wrap the PR description at ~76 characters unless you need longer lines (e.g., for tables or URLs).
  • See the Sending a change via GitHub and Reviews sections of the Contribution Guide as well as the FAQ for details.

@AskAlexSharov AskAlexSharov changed the title encoding/hex: vectorize Encode and Decode on amd64 and arm64 encoding/hex: vectorize Encode/Decode Oct 9, 2026
@gopherbot

Copy link
Copy Markdown
Contributor

Message from Gopher Robot:

Patch Set 1:

(1 comment)


Please don’t reply on this GitHub thread. Visit golang.org/cl/847806.
After addressing review feedback, remember to publish your drafts!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants