Skip to content

Speed up scans with a roaring row selection - #10359

Merged
robert3005 merged 2 commits into
developfrom
ji/roaring-selection-scan
Oct 10, 2026
Merged

robert3005 merged 2 commits into
developfrom
ji/roaring-selection-scan

Conversation

@joseph-isaacs

Copy link
Copy Markdown
Contributor

Summary

Scans that take their rows from an external index pass those rows as a Selection. With IncludeRoaring, every split did the following:

  • built a one-range treemap,
  • intersected it with a clone of the whole selection,
  • collected the result into a Vec<usize>,
  • and only then built the mask.

Sparse roaring selections also never got the exact split ranges that IncludeByIndex selections get, so they were planned over natural splits.

Changes

  • Selection::row_mask for IncludeRoaring/ExcludeRoaring counts the selected rows in the split with range_cardinality. It returns an all-false or all-true mask directly when it can. Otherwise it seeks the treemap iterator to the split and sets the bits in place.
  • attempt_split_ranges plans exact sparse ranges for IncludeRoaring the same way it does for IncludeByIndex. The range-building loop is shared between the two.

New tests check that roaring and index selections plan the same split ranges, and that dense roaring selections fall back to natural splits.

On a 14M-row file with 840K rows selected, a count(*) over the selection drops from 29.5 ms to 20.6 ms of CPU. With 58K rows selected it drops from 7.0 ms to 5.9 ms.

🤖 Generated with Claude Code

https://claude.ai/code/session_01JRTKSkM9xUFT8WHRyS97aR


Generated by Claude Code

@codspeed

codspeed Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Merging this PR will regress 2 benchmarks

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ 19 benchmarks spent significant time in system calls

System calls cannot be consistently instrumented, so they are not included in the measure, which understates the real cost. Please switch to the Walltime instrument to accurately measure system calls.

Measurement and system calls

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 1 improved benchmark
❌ 2 regressed benchmarks
✅ 2180 untouched benchmarks
⏩ 359 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ WallTime bitpack_blocked_compress_avx2 6.7 µs 7.6 µs -11.69%
❌ WallTime scalar_subtract_neon 10.9 µs 12.1 µs -10.29%
⚡ ⚠️ Simulation density_sweep_dense_runs[0.001] 47.8 µs 29.7 µs +60.73%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing ji/roaring-selection-scan (1939f12) with develop (731a231)

Open in CodSpeed

Footnotes

  1. 359 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@robert3005 robert3005 added the changelog/performance A performance improvement label Oct 10, 2026
joseph-isaacs and others added 2 commits October 10, 2026 16:18
Index-driven scans pass the matching rows as a `Selection`. With
`IncludeRoaring`, every split built a one-range treemap, intersected it
with a clone of the whole selection, collected the result into a
`Vec<usize>` and only then built the mask. Sparse roaring selections also
never got the exact split ranges that `IncludeByIndex` gets, so they were
planned over natural splits.

- `Selection::row_mask` now counts the selected rows in the split with
  `range_cardinality`, returning all-false/all-true masks directly, and
  otherwise seeks the treemap iterator to the split and sets bits in place.
- `attempt_split_ranges` plans exact ranges for `IncludeRoaring` the same
  way as for `IncludeByIndex`.

On a 14M-row file with 840K selected rows, a count over the selection drops
from 29.5 ms to 20.6 ms of CPU.

Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JRTKSkM9xUFT8WHRyS97aR
The minimal-versions and WASM CI checks resolve roaring 0.11.0, which does
not have range_cardinality.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Robert Kruszewski <github@robertk.io>
@robert3005
robert3005 force-pushed the ji/roaring-selection-scan branch from a40d1c2 to 1939f12 Compare October 10, 2026 15:18
@robert3005
robert3005 merged commit 3d12093 into develop Oct 10, 2026
91 of 92 checks passed
@robert3005
robert3005 deleted the ji/roaring-selection-scan branch October 10, 2026 17:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/performance A performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants