Repository navigation
Add fused compare and filter kernels for Delta - #10247
Draft
joseph-isaacs wants to merge 6 commits into
Draft
joseph-isaacs wants to merge 6 commits into
joseph-isaacs wants to merge 6 commits into
Conversation
Delta had no compute kernels besides cast, so comparing or filtering a Delta array decoded it in full first: an 8 MB write and read per million 64-bit values, repeated for every predicate on the column. Decode one 1024-value chunk at a time into stack buffers instead. Compare against a constant writes only the result bits, flipping the sign bit so signed order compares as unsigned. Filter decodes only the chunks that hold a selected value and gathers from them; masks denser than one half keep the decode-then-filter path. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Copying one selected index at a time made the filter kernel slower than decode-then-filter for dense run-shaped masks over 32-bit values. Walk the mask's slices instead and copy each run's part of a decoded chunk in one go. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Gather by 64-bit mask words: skip empty words, bulk-copy full words, and walk set bits otherwise. When the mask touches at least 90% of chunks at 15%+ density, decline so the canonical decode-then-filter path runs. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
A contiguous mask executes as a zero-copy slice that decodes only the selected range, which the chunk gather was preempting. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Check cached slices or indices first, and otherwise scan only the candidate run or the tail after it, so declining a contiguous mask stays cheap. Also format and satisfy clippy in the Delta kernels. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Merging this PR will not alter performance
|
Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on #10233, which makes the compressor choose Delta for metrics timestamps and counters. The westermo and lo2 queries then compare and filter Delta columns, and both operations currently decode the whole array first. These kernels work one 1024-value chunk at a time in stack buffers and never materialize the decoded array.
Changes
decode_chunkindelta_decompress.rsdecodes a single chunk: undelta, then untranspose.CompareKernel for Deltacompares against a constant of the same ptype.FilterKernel for Deltadecodes only the chunks the mask touches.delta/vtable/kernels.rs.Kernel benchmarks
The benchmark uses 2^20 values in five shapes: i64 series-major timestamps, i64 jittered timestamps, a u32 counter, an i32 negative sawtooth, and i64 with nulls. Each shape is run full and sliced. The baseline decodes the whole array, then runs the same compare or filter. All results were checked equal to the baseline. Figures are geomeans, with minimum–maximum in brackets.
In the ≥ 15% rows the kernel declines and the existing path runs. Running the same benchmark with the kernel switched off gives the same timings, so the gap to the baseline is not caused by this PR. Some of these rows show single-run variance of 0.5–1.4x.
Query benchmarks
These use the westermo and lo2 suites from #10214 and #10216, run on the Delta-recompressed files from #10233. Each cell is the Vortex time with this PR over the Vortex time on the same files without it: 5 iterations per query, on a 4-core cloud VM.
Queries that changed by more than 30% were rerun with 15 iterations:
The apparent slowdowns in the first run, lo2 Q02, Q03 and Q15 on DataFusion and westermo Q01 on DuckDB, returned to their earlier times on the rerun. lo2 Q22 on DuckDB has a median of 101 ms against 86 ms before, over 30 iterations. Its p10 is 87 ms, so the change is within this VM's run-to-run variance and not established either way.
Checks run:
cargo test -p vortex-fastlanes --lib -- delta: 89 passed.cargo clippy -p vortex-fastlanes -p vortex-btrblocks -p vortex-bench --all-targets --all-features -- -D warnings: clean, withclippy::manual_mapallowed for an existing lint invortex-bufferthat this PR does not touch.cargo +nightly-2026-09-10 fmt --all --check: clean.🤖 Generated with Claude Code
https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ