Repository navigation
db/recsplit: AVX-512 findBijection - #24418
AskAlexSharov wants to merge 5 commits into
Conversation
The eight salt candidates the scalar loop unrolls by hand fit one 512-bit register, so the splitmix64 finaliser and remap16 run once per key instead of eight times. AVX2 cannot host it: both need a 64x64 multiply, which is VPMULLQ (AVX512DQ), and emulating that from VPMULUDQ partial products costs more than the unrolled scalar form. findSplit is left alone — it histograms into a per-fanout count array, which is a scatter. Experimental: builds only under go1.27 with GOEXPERIMENT=simd on amd64, and falls back to the scalar path when AVX-512 is absent.
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
The hardware-specific path depends on experimental Go SIMD support and requires final validation on AVX-512 hardware.
Review effort: Balanced
Findings: None
What changed in this PR
Adds an AVX-512 implementation of RecSplit’s bijection search while preserving scalar fallback behavior.
Changes:
- Renames the scalar implementation for explicit fallback use.
- Adds SIMD and generic build-tag dispatchers.
- Tests SIMD/scalar result consistency across bucket sizes.
| File | Description |
|---|---|
db/recsplit/recsplit.go |
Exposes the scalar implementation as findBijectionGeneric. |
db/recsplit/recsplit_test.go |
Adds implementation-equivalence coverage. |
db/recsplit/bijection_simd.go |
Implements AVX-512 vectorized salt searching. |
db/recsplit/bijection_generic.go |
Provides the non-SIMD dispatcher. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
Review of this PR, plus two follow-ups. Both applied. Everything else I checked and left: OptimisationThe first version collapsed the scalar's eight independent salt chains into one register, which is latency-bound: splitmix64 is serial. Two explicit chains restore the ILP. Four chains in a n5, AMD EPYC 4344P, go1.27.1, 3 runs each. Portable simd packageNot possible today. The portable |
|
Correction to my note above: the portable Both gaps are synthesizable:
Parity against the scalar form passes on arm64 (m up to 12). The portable package fixes one vector width per execution and reports On amd64 it does not compile — the 512-bit specialization emits an invalid instruction. Minimal reproducer on go1.27.1 linux/amd64: func BitAt(r, one simd.Uint64s) simd.Uint64s {
zero := simd.BroadcastUint64s(0)
return one.ShiftAllLeft(1).IfElse(r.And(simd.BroadcastUint64s(1)).NotEqual(zero), one)
}So |
Experimental, alongside #24410's SIMD build setup.
findBijection's eight hand-unrolled salt candidates fit one 512-bit register, so splitmix64 and remap16 run once per key instead of eight times.n5, AMD EPYC 4344P, go1.27.1, 5 and 3 runs.
TestFindBijectionMatchesGenericchecks the compiled implementation against the scalar one, so it runs in both build modes.AVX2 cannot host this: splitmix64 and remap16 both need a 64x64 multiply, which is VPMULLQ (AVX512DQ). Emulating it from VPMULUDQ partial products costs more than the unrolled scalar form saves.
findSplitis left alone — it histograms into a per-fanout count array, which is a scatter.Falls back to scalar when AVX-512 is absent, and builds only under go1.27 with GOEXPERIMENT=simd on amd64.