Repository navigation
Blocked FoR with one reference per 1024-element chunk - #10105
Conversation
Merging this PR will degrade performance by 4.54%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | dict_canonicalize_gt_u8_avx512[1000000] |
426.9 µs | 515.8 µs | -17.23% |
| ⚡ | WallTime | words_gather_dispatch_neon[65536] |
2.3 µs | 2.1 µs | +10.1% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing mk/for-chunked-02-references (38d5787) with develop (709eaa1)
Footnotes
-
461 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
1 benchmark was run, but is now archived. If it was deleted in another branch, consider rebasing to remove it from the report. Instead if it was added back, click here to restore it. ↩
FoR reference per 1024-element chunkFoR with one reference per 1024-element chunk
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
dc56b84 to
38d5787
Compare
|
can you add a follow up issue |
Stacked on #10105. ## Summary Adds an encoder that picks one reference per 1024-element chunk, and a fused decode for chunked FoR over BitPacked. No wire format changes: arrays with varying references still can't be serialized until the next PR. ## Encoder `FoR::encode_chunked(array, ctx)` uses the same rule as `FoR::encode`, applied per chunk: - Each chunk's reference is the minimum of its valid values, so every encoded value is small and non-negative. Null slots don't count towards the minimum. - Null rows encode as 0, as in `FoR::encode`. - The minimum and the subtraction run per chunk while it is in cache, without branches. Each value is masked with all ones or all zeros from its validity bit, over fixed 64-value blocks, so both loops vectorize at every integer width. I checked the aarch64 assembly: `umin`/`smin` and `sub` at `.8b`/`.8h`/`.4s`/`.2d`, with no per-value branches. - A chunk with no valid values reuses the previous chunk's reference (or the first valid reference at the start), so the references compress into runs. - The references are a plain primitive child and the offset is 0. Compressing the references is left to the compressor scheme. ```text chunk 0 values: 1_000_000 .. 1_000_900 → reference 1_000_000, encoded 0..900 chunk 1 values: all null → reference 1_000_000, encoded 0 chunk 2 values: 9_000_000 .. 9_000_950 → reference 9_000_000, encoded 0..950 ``` Encoding, `cargo bench -p vortex-fastlanes --bench for_encode`, median per iteration on an M5 Max with 512 KiB of input (128Ki `u32` or 64Ki `i64` values): | | `FoR::encode` | `FoR::encode_chunked` | |---|---|---| | u32, non-null | 13.4 µs | 10.3 µs | | i64, non-null | 18.0 µs | 12.5 µs | | u32, 10% null | 112 µs | 27.6 µs | | i64, 10% null | 61.5 µs | 25.0 µs | ## Decode `decompress_many_refs` now dispatches like `decompress_one_ref`: - **Fused**: unsigned arrays whose `encoded` child is `BitPacked` with the same offset as the FoR. For each packed chunk it calls `unchecked_unfor_pack` with that chunk's reference, writing full chunks straight into the output and partial first/last chunks through a scratch buffer. Patch values are added to the reference of the chunk they fall in. - **Unfused** (`add_references`): everything else. It decodes `encoded`, then adds each chunk's reference. A uniquely owned buffer is updated in place. A shared one, as when `encoded` is a plain primitive child, gets the references added while it is copied, in one pass. Copying first and then adding made this path 2.2× slower than single-reference decode. A BitPacked child with patches slices lazily into `Slice(BitPacked)`, so sliced arrays with patches take the unfused path. Single-reference FoR behaves the same way today. Decoding, `--bench for_decode`, same machine and input size. Every chunk spans the same range, so both encodings pack at 7 bits and the difference is the cost of the per-chunk references: | | one reference | per-chunk references | |---|---|---| | u32, primitive child | 8.46 µs | 8.71 µs | | i64, primitive child | 8.46 µs | 8.50 µs | | u32, BitPacked child (fused) | 9.58 µs | 9.67 µs | | i64, BitPacked child (unfused) | 19.1 µs | 19.2 µs | Single-reference decode is as fast as it was before #10105. On aarch64, the per-chunk reference adds compile to four 128-bit `add.4s`/`add.2d` per iteration, and fused decode calls the same `unchecked_unfor_pack` kernel as single-reference decode. ## Benchmarks `for_encode` and `for_decode` cover `u32` and `i64` inputs of 256 KiB and 512 KiB. They carry `#[cpu_features]`, so CodSpeed measures them on the walltime legs instead of in simulation. Locally, every case takes 4–112 µs per iteration. ## CUDA CUDA FoR decoding now returns an error for per-chunk references instead of falling back to the CPU. --------- Signed-off-by: Matt Katz <mhkatz97@gmail.com>
## Summary FoR's `CastReduce` now pushes down only casts that change nullability (`eq_ignore_nullability`). Every cast that changes the integer type declines, so the array decodes and the decoded values are cast. This fixes three correctness bugs in casting FoR arrays. Two of them have been there since before the new FoR array. ## Background FoR stores each value as an offset from a reference: `encoded = value.wrapping_sub(reference)` in the source type, and decoding is `encoded.wrapping_add(reference)` in whatever type the array has. The encoded values are therefore offsets modulo 2ⁿ, stored in the source type, and not ordinary values. The old kernel cast the encoded child and the references to the target type separately, then rebuilt the FoR array. That treats the offsets as plain values, which is only correct when every offset fits the source type as a non-negative number and every decoded value fits the target type. ## Bugs Reproduced on `develop` and, for the first two, on the commit before the blocked FoR stack (`ba5fcfcde0^`): | Cast | Expected | Got | |---|---|---| | `[-128i8, 127]` → `i16` | `[-128, 127]` | `[-128, -129]` | | `[100i16, 200]` → `i8` | `Err(value exceeds target range)` | `[100, -56]` | | `5..10` (reference `-5`) → `u32` | `[5, 6, 7, 8, 9]` | `Err(No CastReduce to cast constant array from i32 to u32)` on decode | 1. **Widening a signed type whose range exceeds `T::MAX` gives wrong values.** `[-128i8, 127]` has reference `-128`, so `127` is stored as `255`, which wraps to `-1` in `i8`. Casting the encoded child to `i16` sign-extends it to `-1`, which decodes as `-128 + -1 = -129`. Over `i8` values `-128..=127`, 128 of 256 values come out wrong, with no error. `FoR::encode` produces this whenever a signed column's range exceeds the type's maximum. 2. **Narrowing wraps instead of erroring.** In `[100i16, 200]` cast to `i8`, the reference `100` and the offsets `0` and `100` each fit `i8`, but `100 + 100` doesn't, and the wrapping decode hides the overflow. Sign changes such as `i8` → `u8` and `u8` → `i8` fail the same way. The opposite also happens: `-100i32..100` fits `i8`, but its offsets `0..=199` don't, so the cast fails on decode. 3. **An out-of-range reference leaves an undecodable array.** Values `5..10` with reference `-5` can remain after filtering away the negatives. Cast to `u32`, the values fit but the reference doesn't. `Constant`'s `CastReduce` declines, so the references slot is left holding a lazy `vortex.cast`, which `try_new_chunked` accepts because it only checks dtype and length. The cast appears to succeed, and decoding fails. Before #10105, the scalar reference cast failed eagerly, so the whole cast errored instead. Per-chunk references hit the same issue, because a slice keeps the references of every chunk it overlaps. The existing tests didn't catch these: every conformance case has a small range and widens. ## Why decline instead of proving the cast safe - **The pushdown doesn't save work.** BitPacked's `CastReduce` only handles nullability, and its widening `CastKernel` unpacks into a full-width buffer. A pushed-down widening cast therefore executes as `FoR(Cast(BitPacked))`: it unpacks into a wide buffer, then adds the reference in a second pass, because the child is no longer BitPacked and FoR's fused unpack-and-add path doesn't apply. Declining gives one fused decode in the narrow type, then a widening primitive cast. The pass count is the same, with the fused pass over narrower data. - **Nothing downstream benefits.** For example, a compare on `FoR(Cast(BitPacked))` can't use BitPacked's compare kernel. - **It wouldn't apply to per-chunk references.** Since #10136, `FoRScheme` writes per-chunk references when `fastlanes.for.v2` is allowed. The proof needs the minimum and maximum of the references child, which may be lazy inside a reduce rule, so every type change would decline there anyway. A nullability-only cast is always safe: it changes neither the values nor the non-nullable references, so it is still pushed down, for single and per-chunk references alike. --------- Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Stacked on #10104.
Summary
The in-memory
FoRarray now allows one reference per 1024-element chunk instead of a single global reference per array.This is supported by adding a
referenceschild that records each reference. Elementidecodes asencoded[i] + references[(offset + i) / 1024]with wrapping arithmetic.This also requires adding an
offsetpointer. This is the position of the first element within the first chunk and only matters for sliced arrays.The wire format is unchanged.
fastlanes.forreads its single reference as aConstantArraychild, andFoRPluginwritesfastlanes.foronly when the references are constant.Note that FoR encoding always uses a constant reference per array still and serialization refuses arrays that have non-constant references, so this change will not lead to breaks.
Example
References
A 3000-row array spans three chunks, each with its own reference.
encodedholds each value minus its chunk's reference.Offset
Slicing is zero-copy, so a slice can start partway through a chunk.
offsetrecords how far into its first chunk the slice starts.The new offset is always below 1024. Slicing a slice adds the offsets and slices the references again.
Constant references (existing files)
API
FoRSlotsgainsreferences.FoRDataholds onlyoffset: u16. The stored reference scalar is gone.FoRArrayExt::reference_scalar() -> &Scalaris replaced byconstant_reference() -> Option<Scalar>, which isSomewhenreferencesis aConstantArray.offset()andptype()move onto the extension trait.FoR::try_new(encoded, reference)keeps its signature and builds the constant child.FoR::try_new_chunked(encoded, references, offset)accepts any references.Kernels
With constant references, every kernel takes its existing path. That includes Eq/NotEq compare pushdown, is_constant, is_sorted, take, filter, the fused BitPacked decode, and CUDA FFOR and dynamic dispatch.
With varying references:
scalar_atlooks up the element's chunk reference.sliceslices the references by chunk and carries the new offset.castcasts the references.Noneand fall back to decoding.Callers that rewrap a FoR's encoded child (the btrblocks FoR scheme, benches, CUDA tests) now pass through
references()andoffset()viatry_new_chunked.