Vectorize the bitmap-container word operations, fusing the cardinality pass
The bitmap-container boolean operations wrote their result with a plain Go
loop over 1024 words and then made a second pass to count it. Add six word
helpers -- orSlice, andSlice, xorSlice, andNotSlice, and the orCardSlice and
andCardSlice variants that also return the population count via VPOPCNTQ --
and use them to replace eight hand-written loops in bitmapcontainer.go.
orBitmap, iorBitmap and iandBitmap now make a single fused pass. andBitmap,
xorBitmap, ixorBitmap, andNotBitmap and iandNotBitmapSurely need the
cardinality before choosing a bitmap or array result, so they keep the
count-first shape and only their write loop is vectorized. lazyIORBitmap and
lazyORBitmap lose their manual four-way unroll.
dst may alias either source: sources are loaded before the destination is
stored within an iteration, which is what the in-place operations rely on.
One 8 KB container on a Xeon Gold 6548N:
write only 599 ns -> 52 ns (11.5x)
write + cardinality 631 ns -> 115 ns (5.5x)
Two bitmaps of 64 dense bitmap containers:
Or 152985 ns -> 90806 ns (1.68x)
AndNot 153366 ns -> 105304 ns (1.46x)
And 152147 ns -> 105404 ns (1.44x)
Xor 151482 ns -> 105158 ns (1.44x)
The end-to-end gap is allocation: each operation allocates 64 fresh 8 KB
containers.
Claude-Session: https://claude.ai/code/session_0123ePBefqPjrCvhxdwWFakj D
Daniel Lemire committed
ff94de30087e244f171ed377769e7ece11c0501b
Parent: 92bd7e5