Convert trivial filters into slices during reduction, MaskValues::last uses BitBuffer::last_set_index - #9831
6 benchmarks regressed
⚠️ Unknown Walltime execution environment detected
Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.
For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.
⚠️ 3 benchmarks measured no execution time
Nothing ran under measurement, usually because the compiler removed the code under test. These results are not comparable, so they count as unchanged.
⚡ 22 improved benchmarks
❌ 6 regressed benchmarks
✅ 2152 untouched benchmarks
⏩ 385 skipped benchmarks1
Warning
Please fix the performance issues or acknowledge them on CodSpeed.
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | filter_powerlaw_by_random[250000] |
130.4 µs | 168.2 µs | -22.48% |
| ❌ | Simulation | take_filter_primitive_nullable_slice_mask_random_indices[4096, 1000] |
96.2 µs | 121.7 µs | -20.93% |
| ❌ | Simulation | take_filter_primitive_nullable_slice_mask_random_indices[16384, 1000] |
114.7 µs | 144.1 µs | -20.42% |
| ❌ | WallTime | filtered_sink_i64_avx512[OneNullInEight] |
22.6 µs | 26.6 µs | -15.3% |
| ❌ | Simulation | take_filter_list_slice_mask_sequential_indices[768, 50] |
128.9 µs | 148.5 µs | -13.19% |
| ❌ | Simulation | take_filter_list_slice_mask_sequential_indices[256, 50] |
130.1 µs | 148.6 µs | -12.4% |
| ⚡ | Simulation | take_fsl_f16_force_manual_range_copy[2048, 10] |
62.8 µs | 9.9 µs | ×6.4 |
| ⚡ | Simulation | density_sweep_dense_runs[0.9] |
59.4 µs | 15.3 µs | ×3.9 |
| ⚡ | Simulation | density_sweep_single_slice[0.9999] |
84.5 µs | 23.2 µs | ×3.6 |
| ⚡ | Simulation | density_sweep_single_slice[0.999] |
84.2 µs | 23.2 µs | ×3.6 |
| ⚡ | Simulation | density_sweep_single_slice[0.95] |
82.3 µs | 23.2 µs | ×3.5 |
| ⚡ | Simulation | density_sweep_single_slice[0.99] |
82 µs | 23.2 µs | ×3.5 |
| ⚡ | Simulation | density_sweep_single_slice[0.5] |
64 µs | 23.3 µs | ×2.8 |
| ⚡ | Simulation | density_sweep_single_slice[0.9] |
48.1 µs | 23.3 µs | ×2.1 |
| ⚡ | Simulation | density_sweep_single_slice[0.1] |
48 µs | 23.3 µs | ×2.1 |
| ⚡ | Simulation | density_sweep_single_slice[0.05] |
46.3 µs | 23.3 µs | +98.88% |
| ⚡ | Simulation | density_sweep_single_slice[0.01] |
40.8 µs | 23.2 µs | +75.78% |
| ⚡ | Simulation | density_sweep_random[0.02] |
89 µs | 63.4 µs | +40.41% |
| ⚡ | Simulation | density_sweep_single_slice[0.005] |
32.2 µs | 23.2 µs | +39.16% |
| ⚡ | Simulation | patterns_i128[Contiguous] |
29.3 µs | 22.1 µs | +32.86% |
| ... | ... | ... | ... | ... | ... |
ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing rk/trivialfilter (a556e0c) with develop (38fa7e3)
Footnotes
-
385 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩