Skip to content

refactor: make bitpacked CPU kernels consume stored chunk layouts - #9938

Closed
mhk197 wants to merge 1 commit into
mk/bitpacked-stack-04-offset-childfrom
mk/bitpacked-stack-02-cpu-layout
Closed

mhk197 wants to merge 1 commit into
mk/bitpacked-stack-04-offset-childfrom
mk/bitpacked-stack-02-cpu-layout

Conversation

@mhk197

@mhk197 mhk197 commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Make CPU readers locate each 1024-value chunk through the offsets child and derive its width as (end - start) / 128. This replaces the assumption that every chunk has one scalar width and a fixed byte stride.

Bulk decoding, take, filter, comparisons, streaming predicates, constant detection, and fused frame-of-reference decoding prepare and validate ChunkLayout once per operation. They borrow materialized offsets or execute the compressed child once, then use direct buffer access in their loops. Scalar decoding reads only the selected chunk's boundaries, using execute_scalar when necessary. Slicing preserves the relevant boundaries and their origin, including partial and zero-width chunks.

Remove the temporary scalar bit_width and allow compressed offsets. Validate boundary differences and packed-buffer bounds before unpacking, with regression tests for malformed materialized and compressed children. Encoders still choose a uniform width, and serialization still accepts only the original v1 format. CUDA retains its uniform-width restriction.

Validation: 396 FastLanes/BtrBlocks tests passed (1 skipped). Focused Clippy passed with all targets, all features, and warnings denied.

@codspeed

codspeed Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Merging this PR will regress 1 benchmark

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ 4 benchmarks measured no execution time

Nothing ran under measurement, usually because the compiler removed the code under test. These results are not comparable, so they count as unchanged.

Preventing compiler optimizations

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 2 improved benchmarks
❌ 1 regressed benchmark
✅ 2229 untouched benchmarks
⏩ 329 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ Simulation take_fsl_random[128, 10] 32.7 µs 58.5 µs -44.13%
⚡ Simulation take_fsl_u8_random[256, 100] 98.7 µs 45.8 µs ×2.2
⚡ Simulation take_fsl_nullable_random[16, 100] 96.9 µs 50.5 µs +92.04%
⚠️ Simulation take_fsl_f16_random[16, 100] 61.5 µs < 1 ns N/A
⚠️ Simulation fixed_16_advancing_ptr_safe[100] < 1 ns < 1 ns N/A
⚠️ Simulation preverify_advancing_ptr_unchecked[1000] < 1 ns < 1 ns N/A
⚠️ Simulation preverify_advancing_ptr_unchecked[10000] < 1 ns < 1 ns N/A

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing mk/bitpacked-stack-02-cpu-layout (797717f) with mk/bitpacked-stack-04-offset-child (06905e1)

Open in CodSpeed

Footnotes

  1. 329 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@mhk197
mhk197 added this pull request to stack #9946 September 18, 2026 19:27
@mhk197
mhk197 force-pushed the mk/bitpacked-stack-01-wire-boundary branch from 950de57 to 9681199 Compare September 18, 2026 19:42
@mhk197
mhk197 force-pushed the mk/bitpacked-stack-02-cpu-layout branch from 9737c00 to 7390ffe Compare September 18, 2026 19:42
@mhk197
mhk197 deployed to duckdb-build September 18, 2026 19:43 — with GitHub Actions Active
@mhk197
mhk197 force-pushed the mk/bitpacked-stack-02-cpu-layout branch 3 times, most recently from bd67461 to d0d31cc Compare September 23, 2026 18:17
@mhk197
mhk197 removed this pull request from stack #9946 September 23, 2026 18:19
@mhk197 mhk197 changed the title refactor: make bitpacked CPU kernels consume chunk layouts refactor: make bitpacked CPU kernels consume stored chunk layouts Sep 23, 2026
@mhk197
mhk197 changed the base branch from mk/bitpacked-stack-01-wire-boundary to mk/bitpacked-stack-04-offset-child September 23, 2026 18:22
@mhk197
mhk197 added this pull request to stack #10006 September 23, 2026 18:22
Signed-off-by: "Matt Katz" <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
@mhk197
mhk197 removed this pull request from stack #10006 September 23, 2026 19:43
@mhk197
mhk197 added this pull request to stack #10012 September 23, 2026 19:43
@mhk197
mhk197 force-pushed the mk/bitpacked-stack-02-cpu-layout branch from d0d31cc to 797717f Compare September 23, 2026 19:43
@mhk197
mhk197 force-pushed the mk/bitpacked-stack-04-offset-child branch from ecd1bd3 to 06905e1 Compare September 23, 2026 19:43
@mhk197 mhk197 closed this Oct 2, 2026
@mhk197
mhk197 removed this pull request from stack #10012 October 2, 2026 02:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant