Repository navigation
Conversation
Merging this PR will regress 3 benchmarks
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | take_fsl_f16_random[256, 100] |
85.4 µs | 141.1 µs | -39.52% |
| ❌ | Simulation | take_chunked_fsl_sorted[32, 64] |
196 µs | 252.4 µs | -22.34% |
| ❌ | Simulation | take_fsl_random[64, 100] |
124.9 µs | 143.8 µs | -13.16% |
| ⚡ | Simulation | take_fsl_u8_random[256, 100] |
98.6 µs | 45.8 µs | ×2.2 |
| ⚡ | WallTime | filtered_sink_i64_neon[NineNullsInTen] |
19.6 µs | 17.6 µs | +11.8% |
| ⚡ | WallTime | filtered_owned_i64_neon[NineNullsInTen] |
19.5 µs | 17.6 µs | +10.92% |
| ⚡ | WallTime | scalar_subtract_neon |
13.4 µs | 12.1 µs | +10.62% |
| 🆕 | Simulation | compress_v1[u16, drift] |
N/A | 291.2 µs | N/A |
| 🆕 | Simulation | compress_v1[u16, drift+exc1%] |
N/A | 829.2 µs | N/A |
| 🆕 | Simulation | compress_v1[u16, drift+null10%] |
N/A | 519.3 µs | N/A |
| 🆕 | Simulation | compress_v1[u16, random] |
N/A | 297.4 µs | N/A |
| 🆕 | Simulation | compress_v1[u16, spiky] |
N/A | 884.5 µs | N/A |
| 🆕 | Simulation | compress_v1[u16, spiky+exc1%] |
N/A | 891.4 µs | N/A |
| 🆕 | Simulation | compress_v1[u16, uniform] |
N/A | 267.5 µs | N/A |
| 🆕 | Simulation | compress_v1[u16, uniform+exc1%] |
N/A | 810.9 µs | N/A |
| 🆕 | Simulation | compress_v1[u16, uniform+null10%] |
N/A | 496.9 µs | N/A |
| 🆕 | Simulation | compress_v1[u16, zero_heavy] |
N/A | 267.9 µs | N/A |
| 🆕 | Simulation | compress_v1[u32, drift] |
N/A | 427.4 µs | N/A |
| 🆕 | Simulation | compress_v1[u32, drift+exc1%] |
N/A | 988.9 µs | N/A |
| 🆕 | Simulation | compress_v1[u32, drift+null10%] |
N/A | 1.5 ms | N/A |
| ... | ... | ... | ... | ... | ... |
ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing mk/bitpacked-stack-09-benchmarks (5c86d38) with mk/bitpacked-stack-08-fused-encoder (34b2162)2
Footnotes
-
323 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
No successful run was found on
mk/bitpacked-stack-08-fused-encoder(ab1d8e1) during the generation of this report, so 86f78ee was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩
d40e0fd to
910328c
Compare
28385b3 to
4cc3a9d
Compare
910328c to
f2104c5
Compare
545cb8a to
c8fc0a4
Compare
f2104c5 to
5fc8320
Compare
c8fc0a4 to
20df1a9
Compare
Signed-off-by: "Matt Katz" <mhkatz97@gmail.com> Signed-off-by: Matt Katz <mhkatz97@gmail.com>
34b2162 to
ab1d8e1
Compare
20df1a9 to
5c86d38
Compare
Add a benchmark sweep comparing global-width and per-chunk encoding and decoding across uniform, alternating, drifting, and outlier-heavy distributions. Report compressed sizes alongside timing to make compression and runtime tradeoffs visible.
Size accounting includes the offsets child for v2 arrays and excludes it for uniform arrays, whose v1 wire format stores only the scalar width.
Validation: 435 FastLanes/BtrBlocks tests passed (1 skipped). Focused Clippy passed with all targets, all features, and warnings denied. FastLanes doctests and workspace formatting checks passed. CUDA runtime tests were not run. The focused u32 decoding benchmark completed across uniform and varying-width cases (40 samples, 64 iterations per sample).