Skip to content

Blocked FoR with one reference per 1024-element chunk - #10105

Merged
mhk197 merged 8 commits into
developfrom
mk/for-chunked-02-references
Sep 29, 2026
Merged

mhk197 merged 8 commits into
developfrom
mk/for-chunked-02-references

Conversation

@mhk197

@mhk197 mhk197 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #10104.

Summary

The in-memory FoR array now allows one reference per 1024-element chunk instead of a single global reference per array.

This is supported by adding a references child that records each reference. Element i decodes as encoded[i] + references[(offset + i) / 1024] with wrapping arithmetic.

This also requires adding an offset pointer. This is the position of the first element within the first chunk and only matters for sliced arrays.

The wire format is unchanged. fastlanes.for reads its single reference as a ConstantArray child, and FoRPlugin writes fastlanes.for only when the references are constant.

Note that FoR encoding always uses a constant reference per array still and serialization refuses arrays that have non-constant references, so this change will not lead to breaks.

Example

References

A 3000-row array spans three chunks, each with its own reference. encoded holds each value minus its chunk's reference.

rows         0 ─────────── 1023 │ 1024 ────────── 2047 │ 2048 ──────── 2999
references   [    1_000_000     │      5_000_000       │     9_000_000     ]
encoded      [  small values    │    small values      │   small values    ]

value[1500] = encoded[1500] + references[1500 / 1024]
            = encoded[1500] + references[1]
            = encoded[1500] + 5_000_000

Offset

Slicing is zero-copy, so a slice can start partway through a chunk. offset records how far into its first chunk the slice starts.

             │◄──── chunk 0 ────►│◄──── chunk 1 ────►│◄──── chunk 2 ────►│
rows         0                 1024      1500      2048      2500      3000
slice                                     [═══════════════════)
                                   │◄ 476 ►│

slice(1500..2500):
  encoded    = encoded.slice(1500..2500)          1000 rows
  references = references.slice(1..3)             [5_000_000, 9_000_000]
  offset     = 1500 % 1024                        476

slice row i uses references[(476 + i) / 1024]:
  i = 0..=547    →  references[0] = 5_000_000     (original chunk 1)
  i = 548..=999  →  references[1] = 9_000_000     (original chunk 2)

The new offset is always below 1024. Slicing a slice adds the offsets and slices the references again.

Constant references (existing files)

fastlanes.for on disk      metadata = reference 42, children = [encoded]
        │ read                                  ▲ write (only when references are constant)
        ▼                                       │
in memory                  references = ConstantArray(42, len = num_chunks), offset = 0

API

  • FoRSlots gains references.
  • FoRData holds only offset: u16. The stored reference scalar is gone.
  • FoRArrayExt::reference_scalar() -> &Scalar is replaced by constant_reference() -> Option<Scalar>, which is Some when references is a ConstantArray. offset() and ptype() move onto the extension trait.
  • FoR::try_new(encoded, reference) keeps its signature and builds the constant child. FoR::try_new_chunked(encoded, references, offset) accepts any references.

Kernels

With constant references, every kernel takes its existing path. That includes Eq/NotEq compare pushdown, is_constant, is_sorted, take, filter, the fused BitPacked decode, and CUDA FFOR and dynamic dispatch.

With varying references:

  • Decoding adds each chunk's reference in place. scalar_at looks up the element's chunk reference. slice slices the references by chunk and carries the new offset. cast casts the references.
  • Compare, is_constant, is_sorted, take and filter return None and fall back to decoding.
  • CUDA decodes these arrays on the CPU and never fuses them into a dispatch plan.

Callers that rewrap a FoR's encoded child (the btrblocks FoR scheme, benches, CUDA tests) now pass through references() and offset() via try_new_chunked.

@codspeed

codspeed Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Merging this PR will degrade performance by 4.54%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚡ 1 improved benchmark
❌ 1 regressed benchmark
✅ 2102 untouched benchmarks
⏩ 461 skipped benchmarks1
🗄️ 1 archived benchmark run2

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ WallTime dict_canonicalize_gt_u8_avx512[1000000] 426.9 µs 515.8 µs -17.23%
⚡ WallTime words_gather_dispatch_neon[65536] 2.3 µs 2.1 µs +10.1%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing mk/for-chunked-02-references (38d5787) with develop (709eaa1)

Open in CodSpeed

Footnotes

  1. 461 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

  2. 1 benchmark was run, but is now archived. If it was deleted in another branch, consider rebasing to remove it from the report. Instead if it was added back, click here to restore it. ↩

@mhk197 mhk197 added the changelog/feature A new feature label Sep 28, 2026
@mhk197 mhk197 changed the title Store one FoR reference per 1024-element chunk Blocked FoR with one reference per 1024-element chunk Sep 28, 2026
@mhk197
mhk197 added this pull request to stack #10109 September 28, 2026 16:50
@mhk197
mhk197 marked this pull request as ready for review September 28, 2026 19:57
Base automatically changed from mk/for-chunked-01-plugin to develop September 28, 2026 20:53
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
@mhk197
mhk197 force-pushed the mk/for-chunked-02-references branch from dc56b84 to 38d5787 Compare September 28, 2026 20:53
@joseph-isaacs

Copy link
Copy Markdown
Contributor

can you add a follow up issue

@joseph-isaacs joseph-isaacs left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

z

@mhk197
mhk197 merged commit ba5fcfc into develop Sep 29, 2026
114 of 115 checks passed
@mhk197
mhk197 deleted the mk/for-chunked-02-references branch September 29, 2026 14:58
@mhk197 mhk197 linked an issue Sep 29, 2026 that may be closed by this pull request
mhk197 added a commit that referenced this pull request Sep 30, 2026
Stacked on #10105.

## Summary

Adds an encoder that picks one reference per 1024-element chunk, and a
fused decode for chunked FoR over BitPacked. No wire format changes:
arrays with varying references still can't be serialized until the next
PR.

## Encoder

`FoR::encode_chunked(array, ctx)` uses the same rule as `FoR::encode`,
applied per chunk:

- Each chunk's reference is the minimum of its valid values, so every
encoded value is small and non-negative. Null slots don't count towards
the minimum.
- Null rows encode as 0, as in `FoR::encode`.
- The minimum and the subtraction run per chunk while it is in cache,
without branches. Each value is masked with all ones or all zeros from
its validity bit, over fixed 64-value blocks, so both loops vectorize at
every integer width. I checked the aarch64 assembly: `umin`/`smin` and
`sub` at `.8b`/`.8h`/`.4s`/`.2d`, with no per-value branches.
- A chunk with no valid values reuses the previous chunk's reference (or
the first valid reference at the start), so the references compress into
runs.
- The references are a plain primitive child and the offset is 0.
Compressing the references is left to the compressor scheme.

```text
chunk 0 values: 1_000_000 .. 1_000_900   → reference 1_000_000, encoded 0..900
chunk 1 values: all null                 → reference 1_000_000, encoded 0
chunk 2 values: 9_000_000 .. 9_000_950   → reference 9_000_000, encoded 0..950
```

Encoding, `cargo bench -p vortex-fastlanes --bench for_encode`, median
per iteration on an M5 Max with 512 KiB of input (128Ki `u32` or 64Ki
`i64` values):

| | `FoR::encode` | `FoR::encode_chunked` |
|---|---|---|
| u32, non-null | 13.4 µs | 10.3 µs |
| i64, non-null | 18.0 µs | 12.5 µs |
| u32, 10% null | 112 µs | 27.6 µs |
| i64, 10% null | 61.5 µs | 25.0 µs |

## Decode

`decompress_many_refs` now dispatches like `decompress_one_ref`:

- **Fused**: unsigned arrays whose `encoded` child is `BitPacked` with
the same offset as the FoR. For each packed chunk it calls
`unchecked_unfor_pack` with that chunk's reference, writing full chunks
straight into the output and partial first/last chunks through a scratch
buffer. Patch values are added to the reference of the chunk they fall
in.
- **Unfused** (`add_references`): everything else. It decodes `encoded`,
then adds each chunk's reference. A uniquely owned buffer is updated in
place. A shared one, as when `encoded` is a plain primitive child, gets
the references added while it is copied, in one pass. Copying first and
then adding made this path 2.2× slower than single-reference decode.

A BitPacked child with patches slices lazily into `Slice(BitPacked)`, so
sliced arrays with patches take the unfused path. Single-reference FoR
behaves the same way today.

Decoding, `--bench for_decode`, same machine and input size. Every chunk
spans the same range, so both encodings pack at 7 bits and the
difference is the cost of the per-chunk references:

| | one reference | per-chunk references |
|---|---|---|
| u32, primitive child | 8.46 µs | 8.71 µs |
| i64, primitive child | 8.46 µs | 8.50 µs |
| u32, BitPacked child (fused) | 9.58 µs | 9.67 µs |
| i64, BitPacked child (unfused) | 19.1 µs | 19.2 µs |

Single-reference decode is as fast as it was before #10105. On aarch64,
the per-chunk reference adds compile to four 128-bit `add.4s`/`add.2d`
per iteration, and fused decode calls the same `unchecked_unfor_pack`
kernel as single-reference decode.

## Benchmarks

`for_encode` and `for_decode` cover `u32` and `i64` inputs of 256 KiB
and 512 KiB. They carry `#[cpu_features]`, so CodSpeed measures them on
the walltime legs instead of in simulation. Locally, every case takes
4–112 µs per iteration.

## CUDA

CUDA FoR decoding now returns an error for per-chunk references instead
of falling back to the CPU.

---------

Signed-off-by: Matt Katz <mhkatz97@gmail.com>
mhk197 added a commit that referenced this pull request Oct 1, 2026
## Summary

FoR's `CastReduce` now pushes down only casts that change nullability
(`eq_ignore_nullability`). Every cast that changes the integer type
declines, so the array decodes and the decoded values are cast. This
fixes three correctness bugs in casting FoR arrays. Two of them have
been there since before the new FoR array.

## Background

FoR stores each value as an offset from a reference: `encoded =
value.wrapping_sub(reference)` in the source type, and decoding is
`encoded.wrapping_add(reference)` in whatever type the array has. The
encoded values are therefore offsets modulo 2ⁿ, stored in the source
type, and not ordinary values.

The old kernel cast the encoded child and the references to the target
type separately, then rebuilt the FoR array. That treats the offsets as
plain values, which is only correct when every offset fits the source
type as a non-negative number and every decoded value fits the target
type.

## Bugs

Reproduced on `develop` and, for the first two, on the commit before the
blocked FoR stack (`ba5fcfcde0^`):

| Cast | Expected | Got |
|---|---|---|
| `[-128i8, 127]` → `i16` | `[-128, 127]` | `[-128, -129]` |
| `[100i16, 200]` → `i8` | `Err(value exceeds target range)` | `[100,
-56]` |
| `5..10` (reference `-5`) → `u32` | `[5, 6, 7, 8, 9]` | `Err(No
CastReduce to cast constant array from i32 to u32)` on decode |

1. **Widening a signed type whose range exceeds `T::MAX` gives wrong
values.** `[-128i8, 127]` has reference `-128`, so `127` is stored as
`255`, which wraps to `-1` in `i8`. Casting the encoded child to `i16`
sign-extends it to `-1`, which decodes as `-128 + -1 = -129`. Over `i8`
values `-128..=127`, 128 of 256 values come out wrong, with no error.
`FoR::encode` produces this whenever a signed column's range exceeds the
type's maximum.
2. **Narrowing wraps instead of erroring.** In `[100i16, 200]` cast to
`i8`, the reference `100` and the offsets `0` and `100` each fit `i8`,
but `100 + 100` doesn't, and the wrapping decode hides the overflow.
Sign changes such as `i8` → `u8` and `u8` → `i8` fail the same way. The
opposite also happens: `-100i32..100` fits `i8`, but its offsets
`0..=199` don't, so the cast fails on decode.
3. **An out-of-range reference leaves an undecodable array.** Values
`5..10` with reference `-5` can remain after filtering away the
negatives. Cast to `u32`, the values fit but the reference doesn't.
`Constant`'s `CastReduce` declines, so the references slot is left
holding a lazy `vortex.cast`, which `try_new_chunked` accepts because it
only checks dtype and length. The cast appears to succeed, and decoding
fails. Before #10105, the scalar reference cast failed eagerly, so the
whole cast errored instead. Per-chunk references hit the same issue,
because a slice keeps the references of every chunk it overlaps.

The existing tests didn't catch these: every conformance case has a
small range and widens.

## Why decline instead of proving the cast safe

- **The pushdown doesn't save work.** BitPacked's `CastReduce` only
handles nullability, and its widening `CastKernel` unpacks into a
full-width buffer. A pushed-down widening cast therefore executes as
`FoR(Cast(BitPacked))`: it unpacks into a wide buffer, then adds the
reference in a second pass, because the child is no longer BitPacked and
FoR's fused unpack-and-add path doesn't apply. Declining gives one fused
decode in the narrow type, then a widening primitive cast. The pass
count is the same, with the fused pass over narrower data.
- **Nothing downstream benefits.** For example, a compare on
`FoR(Cast(BitPacked))` can't use BitPacked's compare kernel.
- **It wouldn't apply to per-chunk references.** Since #10136,
`FoRScheme` writes per-chunk references when `fastlanes.for.v2` is
allowed. The proof needs the minimum and maximum of the references
child, which may be lazy inside a reduce rule, so every type change
would decline there anyway.

A nullability-only cast is always safe: it changes neither the values
nor the non-nullable references, so it is still pushed down, for single
and per-chunk references alike.

---------

Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/feature A new feature

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Blocked FoR: Support references child in FoR Array

3 participants