Skip to content

Add fused compare and filter kernels for Delta - #10247

Draft
joseph-isaacs wants to merge 6 commits into
ji/ts-compressionfrom
ji/delta-kernels
Draft

joseph-isaacs wants to merge 6 commits into
ji/ts-compressionfrom
ji/delta-kernels

Conversation

@joseph-isaacs

@joseph-isaacs joseph-isaacs commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Stacked on #10233, which makes the compressor choose Delta for metrics timestamps and counters. The westermo and lo2 queries then compare and filter Delta columns, and both operations currently decode the whole array first. These kernels work one 1024-value chunk at a time in stack buffers and never materialize the decoded array.

Changes

  • decode_chunk in delta_decompress.rs decodes a single chunk: undelta, then untranspose.
  • CompareKernel for Delta compares against a constant of the same ptype.
    • Each chunk is decoded and packed straight into the result bit words.
    • Signed types compare in the unsigned domain by flipping the sign bit.
  • FilterKernel for Delta decodes only the chunks the mask touches.
    • It walks the mask 64 bits at a time: empty words are skipped, full words are bulk-copied, and partial words gather their set bits.
    • It declines, so the existing paths run, in two cases:
      • the mask is contiguous, which the filter executor already turns into a zero-copy slice;
      • the mask touches ≥ 90% of chunks at ≥ 15% density, where decoding everything and then bulk filtering wins.
  • Both kernels are registered as execute-parent kernels in delta/vtable/kernels.rs.
  • Tests cover:
    • all six operators against negative thresholds;
    • sliced arrays with nulls;
    • filter masks that are scattered, contiguous, in the last chunk, made of full words crossing a chunk boundary, or dense enough to fall back;
    • filtering a slice.

Kernel benchmarks

The benchmark uses 2^20 values in five shapes: i64 series-major timestamps, i64 jittered timestamps, a u32 counter, an i32 negative sawtooth, and i64 with nulls. Each shape is run full and sliced. The baseline decodes the whole array, then runs the same compare or filter. All results were checked equal to the baseline. Figures are geomeans, with minimum–maximum in brackets.

Operation Mask density Speedup
compare (Gt / Lte / Eq vs median) – 1.69x (1.10–2.72x)
filter, one run < 1% 496x
filter, one run 1–15% 74x
filter, clusters of 64 < 1% 70x
filter, clusters of 64 1–15% 3.8x
filter, uniform random < 1% 5.0x
filter, uniform random 1–15% 1.6x
filter, strided < 1% 5.3x
filter, strided 1–15% 1.8x
filter, any pattern ≥ 15% ~0.9x

In the ≥ 15% rows the kernel declines and the existing path runs. Running the same benchmark with the kernel switched off gives the same timings, so the gap to the baseline is not caused by this PR. Some of these rows show single-run variance of 0.5–1.4x.

Query benchmarks

These use the westermo and lo2 suites from #10214 and #10216, run on the Delta-recompressed files from #10233. Each cell is the Vortex time with this PR over the Vortex time on the same files without it: 5 iterations per query, on a 4-core cloud VM.

Suite DataFusion DuckDB
Westermo, geomean of 15 queries 1.09x faster 1.03x faster
LO2, geomean of 28 queries 1.21x faster 1.03x faster

Queries that changed by more than 30% were rerun with 15 iterations:

Query Before After
westermo Q01, DataFusion 87.8 ms 29.4 ms
westermo Q14, DuckDB 32.6 ms 16.6 ms
lo2 Q12, DataFusion 62.3 ms 24.4 ms
lo2 Q13, DataFusion 61.7 ms 19.9 ms
lo2 Q16, DataFusion 84.2 ms 35.4 ms
lo2 Q06, DuckDB 107.5 ms 82.3 ms

The apparent slowdowns in the first run, lo2 Q02, Q03 and Q15 on DataFusion and westermo Q01 on DuckDB, returned to their earlier times on the rerun. lo2 Q22 on DuckDB has a median of 101 ms against 86 ms before, over 30 iterations. Its p10 is 87 ms, so the change is within this VM's run-to-run variance and not established either way.

Checks run:

  • cargo test -p vortex-fastlanes --lib -- delta: 89 passed.
  • cargo clippy -p vortex-fastlanes -p vortex-btrblocks -p vortex-bench --all-targets --all-features -- -D warnings: clean, with clippy::manual_map allowed for an existing lint in vortex-buffer that this PR does not touch.
  • cargo +nightly-2026-09-10 fmt --all --check: clean.

🤖 Generated with Claude Code

https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ

claude added 5 commits October 2, 2026 20:39
Delta had no compute kernels besides cast, so comparing or filtering a
Delta array decoded it in full first: an 8 MB write and read per million
64-bit values, repeated for every predicate on the column.

Decode one 1024-value chunk at a time into stack buffers instead.
Compare against a constant writes only the result bits, flipping the
sign bit so signed order compares as unsigned. Filter decodes only the
chunks that hold a selected value and gathers from them; masks denser
than one half keep the decode-then-filter path.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Copying one selected index at a time made the filter kernel slower than
decode-then-filter for dense run-shaped masks over 32-bit values. Walk
the mask's slices instead and copy each run's part of a decoded chunk in
one go.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Gather by 64-bit mask words: skip empty words, bulk-copy full words, and
walk set bits otherwise. When the mask touches at least 90% of chunks at
15%+ density, decline so the canonical decode-then-filter path runs.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
A contiguous mask executes as a zero-copy slice that decodes only the
selected range, which the chunk gather was preempting.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Check cached slices or indices first, and otherwise scan only the
candidate run or the tail after it, so declining a contiguous mask
stays cheap. Also format and satisfy clippy in the Delta kernels.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
@joseph-isaacs joseph-isaacs added the changelog/performance A performance improvement label Oct 2, 2026 — with Claude
@codspeed

codspeed Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

✅ 2097 untouched benchmarks
⏩ 503 skipped benchmarks1


Comparing ji/delta-kernels (e2f97a9) with ji/ts-compression (b2edd3c)

Open in CodSpeed

Footnotes

  1. 503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/performance A performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants