Repository navigation
Conversation
Merging this PR will regress 1 benchmark
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | take_fsl_random[128, 10] |
32.7 µs | 58.5 µs | -44.13% |
| ⚡ | Simulation | take_fsl_u8_random[256, 100] |
98.7 µs | 45.8 µs | ×2.2 |
| ⚡ | Simulation | take_fsl_nullable_random[16, 100] |
96.9 µs | 50.5 µs | +92.04% |
| Simulation | take_fsl_f16_random[16, 100] |
61.5 µs | < 1 ns | N/A | |
| Simulation | fixed_16_advancing_ptr_safe[100] |
< 1 ns | < 1 ns | N/A | |
| Simulation | preverify_advancing_ptr_unchecked[1000] |
< 1 ns | < 1 ns | N/A | |
| Simulation | preverify_advancing_ptr_unchecked[10000] |
< 1 ns | < 1 ns | N/A |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing mk/bitpacked-stack-02-cpu-layout (797717f) with mk/bitpacked-stack-04-offset-child (06905e1)
Footnotes
-
329 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
950de57 to
9681199
Compare
9737c00 to
7390ffe
Compare
bd67461 to
d0d31cc
Compare
Signed-off-by: "Matt Katz" <mhkatz97@gmail.com> Signed-off-by: Matt Katz <mhkatz97@gmail.com>
d0d31cc to
797717f
Compare
ecd1bd3 to
06905e1
Compare
Make CPU readers locate each 1024-value chunk through the offsets child and derive its width as
(end - start) / 128. This replaces the assumption that every chunk has one scalar width and a fixed byte stride.Bulk decoding,
take,filter, comparisons, streaming predicates, constant detection, and fused frame-of-reference decoding prepare and validateChunkLayoutonce per operation. They borrow materialized offsets or execute the compressed child once, then use direct buffer access in their loops. Scalar decoding reads only the selected chunk's boundaries, usingexecute_scalarwhen necessary. Slicing preserves the relevant boundaries and their origin, including partial and zero-width chunks.Remove the temporary scalar
bit_widthand allow compressed offsets. Validate boundary differences and packed-buffer bounds before unpacking, with regression tests for malformed materialized and compressed children. Encoders still choose a uniform width, and serialization still accepts only the original v1 format. CUDA retains its uniform-width restriction.Validation: 396 FastLanes/BtrBlocks tests passed (1 skipped). Focused Clippy passed with all targets, all features, and warnings denied.