Repository navigation
encoding/hex: vectorize Encode/Decode - #82089
Open
AskAlexSharov wants to merge 5 commits into
Open
AskAlexSharov wants to merge 5 commits into
AskAlexSharov wants to merge 5 commits into
Conversation
With GOEXPERIMENT=simd on a CPU with AVX2, Encode handles whole 16-byte blocks and Decode whole 32-character blocks with simd/archsimd. Shorter inputs and the tails use the scalar loops. Encode widens each byte 0xHL to the uint16 0x00HL and multiplies it by 0x1001, which gives 0xL0HL; shifted right by 4 it is 0x0L0H, the two nibbles in output order, and one VPSHUFB turns them into digits. Decode computes the nibbles and checks the characters in one pass, by algorithm 3 of http://0x80.pl/notesen/2022-01-17-validating-hex-parse.html, and packs them with VPMADDUBSW and VPSHUFB. A block with an invalid character is left to the scalar loop, so the error and the partial output do not change. goos: linux goarch: amd64 pkg: encoding/hex cpu: AMD EPYC 4344P 8-Core Processor │ old │ new │ │ B/s │ B/s vs base │ Encode/16 1.464Gi ± 0% 4.648Gi ± 1% +217.53% (p=0.000 n=10) Encode/32 1.508Gi ± 0% 7.850Gi ± 0% +420.42% (p=0.000 n=10) Encode/64 1.535Gi ± 0% 12.445Gi ± 0% +710.69% (p=0.000 n=10) Encode/256 1.508Gi ± 0% 22.212Gi ± 0% +1372.80% (p=0.000 n=10) Encode/1024 1.539Gi ± 0% 27.676Gi ± 0% +1698.23% (p=0.000 n=10) Encode/4096 1.547Gi ± 0% 29.456Gi ± 0% +1803.71% (p=0.000 n=10) Encode/16384 1.539Gi ± 0% 29.504Gi ± 1% +1817.25% (p=0.000 n=10) Decode/32 3.089Gi ± 4% 6.009Gi ± 0% +94.51% (p=0.000 n=10) Decode/64 3.314Gi ± 3% 9.699Gi ± 0% +192.72% (p=0.000 n=10) Decode/128 3.384Gi ± 3% 13.978Gi ± 0% +313.03% (p=0.000 n=10) Decode/256 3.499Gi ± 2% 17.957Gi ± 0% +413.25% (p=0.000 n=10) Decode/1024 3.484Gi ± 1% 22.818Gi ± 0% +554.90% (p=0.000 n=10) Decode/4096 3.529Gi ± 2% 24.472Gi ± 0% +593.40% (p=0.000 n=10) Decode/16384 3.538Gi ± 5% 24.763Gi ± 0% +599.96% (p=0.000 n=10) DecodeString/256 2.702Gi ± 13% 4.125Gi ± 1% +52.63% (p=0.000 n=10) DecodeString/1024 2.834Gi ± 6% 6.652Gi ± 1% +134.69% (p=0.000 n=10) DecodeString/4096 2.813Gi ± 1% 7.732Gi ± 1% +174.85% (p=0.000 n=10) DecodeString/16384 2.842Gi ± 6% 8.366Gi ± 1% +194.40% (p=0.000 n=10) geomean 2.381Gi 12.82Gi +438.29% The portable simd package cannot express these kernels yet: it cannot move bytes between lanes. Running the same nibble arithmetic in each 64-bit lane and joining the lanes with scalar stores reaches, for 16 KiB on the same machine: Encode Decode 512-bit 4.79GiB/s 4.34GiB/s 256-bit 3.83GiB/s 3.28GiB/s 128-bit 2.68GiB/s 2.20GiB/s (Decode counted in output bytes), that is 1.2x-3.1x the scalar loops but 9-38% of these kernels, which reach 29.5GiB/s and 11.3GiB/s counted the same way. The simd package would need: - interleaving the bytes of two vectors, or zero-extending half of a Uint8s to Uint16s (VPUNPCKLBW or VPMOVZXBW, ZIP1 or UXTL), for Encode; - narrowing Uint16s to Uint8s, or taking the even bytes of two vectors (VPACKUSWB, UZP1 or XTN), for Decode; - a byte table lookup (VPSHUFB, TBL), which saves a few operations in Encode; - comparisons and shifts on Uint8s, which only Int8s and wider types have; - clearing the upper bits of the AVX registers before returning: archsimd.ClearAVXUpperBits has no portable counterpart, and the compiler does not insert VZEROUPPER (golang#80835).
Contributor
|
This PR (HEAD: cbdb9c7) has been imported to Gerrit for code review. Please visit Gerrit at https://go-review.googlesource.com/c/go/+/847806. Important tips:
|
Contributor
|
Message from Gopher Robot: Patch Set 1: (1 comment) Please don’t reply on this GitHub thread. Visit golang.org/cl/847806. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
simdandsimdarch(arm64 and amd64) implementations of hex encode/decodeImplementation details
On amd64 Encode widens each byte 0xHL to the uint16 0x00HL and
multiplies it by 0x1001, which gives 0xL0HL; shifted right by 4 it is
0x0L0H, the two nibbles in output order, and one VPSHUFB turns them into
digits. Decode computes the nibbles and checks the characters in one
pass, by algorithm 3 of
http://0x80.pl/notesen/2022-01-17-validating-hex-parse.html, and packs
them with VPMADDUBSW and VPSHUFB.
On arm64 Encode looks up both nibbles with TBL and interleaves the
digits with ZIP1 and ZIP2. Decode computes the nibbles as on amd64 and
packs them with UZP1 and UZP2.
GiB/s (Decode counted in input characters):
The portable kernels are slower because every 64-bit lane works alone
and scalar stores join the lanes. What portable
simdpackage lacks:Uint8s to Uint16s (VPUNPCKLBW or VPMOVZXBW; ZIP1 or UXTL), for Encode
(VPACKUSWB; UZP1 or XTN), for Decode
Encode
have
archsimd.ClearAVXUpperBits has no portable counterpart, and the
compiler does not insert VZEROUPPER (cmd/compile: legacy SSE encodings emitted in functions using simd/archsimd intrinsics cause AVX-SSE transition penalties #80835)
All benchmarks, scalar against archsimd: