Repository navigation
Speed up scans with a roaring row selection - #10359
Conversation
Merging this PR will regress 2 benchmarks
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | bitpack_blocked_compress_avx2 |
6.7 µs | 7.6 µs | -11.69% |
| ❌ | WallTime | scalar_subtract_neon |
10.9 µs | 12.1 µs | -10.29% |
| ⚡ |
Simulation | density_sweep_dense_runs[0.001] |
47.8 µs | 29.7 µs | +60.73% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ji/roaring-selection-scan (1939f12) with develop (731a231)
Footnotes
-
359 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
Index-driven scans pass the matching rows as a `Selection`. With `IncludeRoaring`, every split built a one-range treemap, intersected it with a clone of the whole selection, collected the result into a `Vec<usize>` and only then built the mask. Sparse roaring selections also never got the exact split ranges that `IncludeByIndex` gets, so they were planned over natural splits. - `Selection::row_mask` now counts the selected rows in the split with `range_cardinality`, returning all-false/all-true masks directly, and otherwise seeks the treemap iterator to the split and sets bits in place. - `attempt_split_ranges` plans exact ranges for `IncludeRoaring` the same way as for `IncludeByIndex`. On a 14M-row file with 840K selected rows, a count over the selection drops from 29.5 ms to 20.6 ms of CPU. Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JRTKSkM9xUFT8WHRyS97aR
The minimal-versions and WASM CI checks resolve roaring 0.11.0, which does not have range_cardinality. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Robert Kruszewski <github@robertk.io>
a40d1c2 to
1939f12
Compare
Summary
Scans that take their rows from an external index pass those rows as a
Selection. WithIncludeRoaring, every split did the following:Vec<usize>,Sparse roaring selections also never got the exact split ranges that
IncludeByIndexselections get, so they were planned over natural splits.Changes
Selection::row_maskforIncludeRoaring/ExcludeRoaringcounts the selected rows in the split withrange_cardinality. It returns an all-false or all-true mask directly when it can. Otherwise it seeks the treemap iterator to the split and sets the bits in place.attempt_split_rangesplans exact sparse ranges forIncludeRoaringthe same way it does forIncludeByIndex. The range-building loop is shared between the two.New tests check that roaring and index selections plan the same split ranges, and that dense roaring selections fall back to natural splits.
On a 14M-row file with 840K rows selected, a
count(*)over the selection drops from 29.5 ms to 20.6 ms of CPU. With 58K rows selected it drops from 7.0 ms to 5.9 ms.🤖 Generated with Claude Code
https://claude.ai/code/session_01JRTKSkM9xUFT8WHRyS97aR
Generated by Claude Code