Repository navigation
Conversation
Merging this PR will improve performance by 79.37%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | Simulation | take_fsl_random[128, 10] |
58.6 µs | 32.7 µs | +79.37% |
| Simulation | take_fsl_u32_random[256, 10] |
< 1 ns | < 1 ns | N/A | |
| Simulation | fixed_16_advancing_ptr_safe[100] |
< 1 ns | < 1 ns | N/A | |
| Simulation | preverify_advancing_ptr_unchecked[1000] |
< 1 ns | < 1 ns | N/A | |
| Simulation | preverify_advancing_ptr_unchecked[10000] |
< 1 ns | < 1 ns | N/A |
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing mk/bitpacked-stack-06-v2-wire (60ce538) with mk/bitpacked-stack-05-explicit-packing (0e28beb)2
Footnotes
-
329 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
No successful run was found on
mk/bitpacked-stack-05-explicit-packing(82c7619) during the generation of this report, so a3ce2bd was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩
899c596 to
db2c61b
Compare
377f70c to
e265aeb
Compare
b055b6b to
a90e6fe
Compare
e265aeb to
4b029d8
Compare
a90e6fe to
0e28beb
Compare
Signed-off-by: "Matt Katz" <mhkatz97@gmail.com> Signed-off-by: Matt Katz <mhkatz97@gmail.com>
0e28beb to
82c7619
Compare
4b029d8 to
60ce538
Compare
Serialize arrays with differing chunk widths under
fastlanes.bitpacked_v2. The format stores one offsets child containingnum_chunks + 1byte boundaries; widths are derived from adjacent differences. Metadata contains only bounded array and patch information.Uniform arrays retain the v1 wire format, which omits the offsets child and stores one width in metadata. Empty arrays use width zero because there are no chunks from which to derive a width. Deserialization still accepts legacy empty arrays with a nonzero width.
Add recursive round trips, execution with compressed offsets, validation of malformed compressed boundaries, and rejection of mismatched format IDs and child layouts. Extract shared wire-format helpers while preserving the v1 validation order. CUDA support remains limited to uniform widths.
Validation: 403 FastLanes/BtrBlocks tests passed (1 skipped).