Repository navigation
feat(compress): retain Narrow integer children through compression - #10333
connortsui20 wants to merge 4 commits into
Conversation
Merging this PR will regress 3 benchmarks
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | take_map[(0.1, 1.0)] |
225.3 µs | 289.3 µs | -22.12% |
| ❌ | WallTime | bitpack_blocked_compress_avx2 |
6.7 µs | 7.6 µs | -11.67% |
| ❌ | Simulation | take_map[(0.1, 0.5)] |
147.7 µs | 166.8 µs | -11.44% |
| ⚡ | Simulation | encode_primitives[u8, (1000, 2)] |
117.3 µs | 74 µs | +58.52% |
| ⚡ | Simulation | encode_primitives[u8, (1000, 4)] |
117.5 µs | 74.2 µs | +58.36% |
| ⚡ | Simulation | encode_primitives[u8, (1000, 8)] |
118.7 µs | 75.3 µs | +57.72% |
| ⚡ | Simulation | encode_primitives[u8, (1000, 32)] |
120.1 µs | 77 µs | +55.97% |
| ⚡ | Simulation | encode_primitives[u8, (1000, 512)] |
136.1 µs | 92.8 µs | +46.68% |
| ⚡ | Simulation | encode_primitives[u8, (2000, 4)] |
169.4 µs | 125.9 µs | +34.54% |
| ⚡ | Simulation | encode_primitives[u8, (2000, 2)] |
169.2 µs | 125.8 µs | +34.48% |
| ⚡ | Simulation | encode_primitives[u8, (2000, 8)] |
170.1 µs | 126.9 µs | +34.09% |
| ⚡ | Simulation | encode_primitives[u8, (2000, 32)] |
172.1 µs | 128.8 µs | +33.6% |
| ⚡ | Simulation | encode_primitives[u8, (2000, 512)] |
188.4 µs | 145 µs | +29.91% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ct/narrow-compression (71f9458) with ct/narrow-pushdown (557a613)
Footnotes
-
518 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
76da12f to
756d7f0
Compare
756d7f0 to
0310216
Compare
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
b641848 to
71f9458
Compare
Summary
Progress towards #10020. Stacked on #10332.
Compressing a narrow integer child lets codecs decode at the smaller width while a Narrow wrapper retains the column's logical type. Session presets now use this representation when
preview2026.10.0is enabled. Core-only and CUDA presets keep it disabled.Changes
The old primitive
narrowmethod is replaced by shared width selection inNarrowArray. Internal buffers whose dtypes are chosen at construction useencode_values; arrays with a logical dtype contract useencode. Both retain signedness, so nonnegative signed internal buffers no longer silently become unsigned.API Changes
Replaces
PrimitiveArrayExt::narrowwithNarrowArray::encode_valuesfor internal buffers. UseNarrowArray::encodewhen the logical dtype must be retained.