Convert trivial filters into slices during reduction, MaskValues::last uses BitBuffer::last_set_index - #9831
Conversation
Merging this PR will regress 6 benchmarks
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | filter_powerlaw_by_random[250000] |
130.4 µs | 168.2 µs | -22.48% |
| ❌ | Simulation | take_filter_primitive_nullable_slice_mask_random_indices[4096, 1000] |
96.2 µs | 121.7 µs | -20.93% |
| ❌ | Simulation | take_filter_primitive_nullable_slice_mask_random_indices[16384, 1000] |
114.7 µs | 144.1 µs | -20.42% |
| ❌ | WallTime | filtered_sink_i64_avx512[OneNullInEight] |
22.6 µs | 26.6 µs | -15.3% |
| ❌ | Simulation | take_filter_list_slice_mask_sequential_indices[768, 50] |
128.9 µs | 148.5 µs | -13.19% |
| ❌ | Simulation | take_filter_list_slice_mask_sequential_indices[256, 50] |
130.1 µs | 148.6 µs | -12.4% |
| ⚡ | Simulation | take_fsl_f16_force_manual_range_copy[2048, 10] |
62.8 µs | 9.9 µs | ×6.4 |
| ⚡ | Simulation | density_sweep_dense_runs[0.9] |
59.4 µs | 15.3 µs | ×3.9 |
| ⚡ | Simulation | density_sweep_single_slice[0.9999] |
84.5 µs | 23.2 µs | ×3.6 |
| ⚡ | Simulation | density_sweep_single_slice[0.999] |
84.2 µs | 23.2 µs | ×3.6 |
| ⚡ | Simulation | density_sweep_single_slice[0.95] |
82.3 µs | 23.2 µs | ×3.5 |
| ⚡ | Simulation | density_sweep_single_slice[0.99] |
82 µs | 23.2 µs | ×3.5 |
| ⚡ | Simulation | density_sweep_single_slice[0.5] |
64 µs | 23.3 µs | ×2.8 |
| ⚡ | Simulation | density_sweep_single_slice[0.9] |
48.1 µs | 23.3 µs | ×2.1 |
| ⚡ | Simulation | density_sweep_single_slice[0.1] |
48 µs | 23.3 µs | ×2.1 |
| ⚡ | Simulation | density_sweep_single_slice[0.05] |
46.3 µs | 23.3 µs | +98.88% |
| ⚡ | Simulation | density_sweep_single_slice[0.01] |
40.8 µs | 23.2 µs | +75.78% |
| ⚡ | Simulation | density_sweep_random[0.02] |
89 µs | 63.4 µs | +40.41% |
| ⚡ | Simulation | density_sweep_single_slice[0.005] |
32.2 µs | 23.2 µs | +39.16% |
| ⚡ | Simulation | patterns_i128[Contiguous] |
29.3 µs | 22.1 µs | +32.86% |
| ... | ... | ... | ... | ... | ... |
ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing rk/trivialfilter (a556e0c) with develop (38fa7e3)
Footnotes
-
385 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
Polar Signals Profiling ResultsLatest Run
Previous Runs (10)
Powered by Polar Signals Cloud |
Benchmarks: PolarSignals Profiling 📖Commits: PR datafusion / vortex-file-compressed / ns (1.016x ➖, 0↑ 1↓)
File Size Changes (1 files changed, +8.6% overall, 1↑ 0↓)
Totals:
|
Benchmarks: TPC-H SF=1 on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.994x ➖, 0↑ 0↓)
datafusion / parquet / ns (1.000x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (1.013x ➖, 0↑ 2↓)
duckdb / parquet / ns (1.002x ➖, 0↑ 0↓)
File Size Changes (8 files changed, +16.7% overall, 6↑ 2↓)
Totals:
|
Benchmarks: FineWeb NVMe 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.984x ➖, 3↑ 2↓)
datafusion / parquet / ns (1.008x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (0.926x ➖, 3↑ 3↓)
duckdb / parquet / ns (1.035x ➖, 0↑ 1↓)
File Size Changes (1 files changed, +25.7% overall, 1↑ 0↓)
Totals:
|
Benchmarks: Clickbench Sorted on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.736x ✅, 10↑ 0↓)
datafusion / parquet / ns (1.006x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (0.720x ✅, 9↑ 0↓)
duckdb / parquet / ns (0.989x ➖, 0↑ 0↓)
File Size Changes (100 files changed, +26.4% overall, 100↑ 0↓)
Totals:
|
Benchmarks: TPC-DS SF=1 on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.005x ➖, 0↑ 1↓)
datafusion / parquet / ns (1.005x ➖, 0↑ 1↓)
duckdb / vortex-file-compressed / ns (0.978x ➖, 8↑ 3↓)
duckdb / parquet / ns (0.998x ➖, 2↑ 3↓)
File Size Changes (24 files changed, +3.0% overall, 16↑ 8↓)
Totals:
|
Benchmarks: TPC-H SF=10 on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.992x ➖, 0↑ 0↓)
datafusion / parquet / ns (1.002x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (0.983x ➖, 0↑ 1↓)
duckdb / parquet / ns (0.991x ➖, 0↑ 0↓)
File Size Changes (8 files changed, +17.1% overall, 6↑ 2↓)
Totals:
|
Benchmarks: FineWeb S3 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.247x ➖, 0↑ 3↓)
datafusion / parquet / ns (0.963x ➖, 1↑ 0↓)
duckdb / vortex-file-compressed / ns (1.153x ➖, 0↑ 1↓)
duckdb / parquet / ns (1.007x ➖, 0↑ 0↓)
|
Benchmarks: Statistical and Population Genetics 📖Commits: PR How to read Verdict and Engines
duckdb / vortex-file-compressed / ns (0.860x ✅, 6↑ 0↓)
duckdb / parquet / ns (1.022x ➖, 0↑ 0↓)
File Size Changes (1 files changed, +16.6% overall, 1↑ 0↓)
Totals:
|
Benchmarks: Clickbench on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.893x ✅, 13↑ 1↓)
datafusion / parquet / ns (0.997x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (0.963x ➖, 10↑ 4↓)
duckdb / parquet / ns (1.006x ➖, 2↑ 0↓)
File Size Changes (100 files changed, +29.3% overall, 100↑ 0↓)
Totals:
|
Benchmarks: TPC-H SF=1 on S3 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.272x ➖, 0↑ 7↓)
datafusion / parquet / ns (1.155x ➖, 0↑ 5↓)
duckdb / vortex-file-compressed / ns (1.173x ➖, 0↑ 6↓)
duckdb / parquet / ns (1.105x ➖, 0↑ 0↓)
|
156426e to
d997574
Compare
…t uses BitBuffer:last_set_index Signed-off-by: Robert Kruszewski <github@robertk.io>
d997574 to
a556e0c
Compare
| Mask::Values(_) => contiguous_filter_range(array.filter_mask()) | ||
| .map(|range| array.child().slice(range)) | ||
| .transpose(), |
There was a problem hiding this comment.
This looks at buffers so I would think of this as a execute rule
There was a problem hiding this comment.
Can you make an issue to fix this usage of buffer in reduce
Optimise more filters away