Skip to content

Encode/Decode Kernels for BlockedFoR - #10118

Merged
mhk197 merged 13 commits into
developfrom
mk/for-chunked-03-encoder
Sep 30, 2026
Merged

mhk197 merged 13 commits into
developfrom
mk/for-chunked-03-encoder

Conversation

@mhk197

@mhk197 mhk197 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #10105.

Summary

Adds an encoder that picks one reference per 1024-element chunk, and a fused decode for chunked FoR over BitPacked. No wire format changes: arrays with varying references still can't be serialized until the next PR.

Encoder

FoR::encode_chunked(array, ctx) uses the same rule as FoR::encode, applied per chunk:

  • Each chunk's reference is the minimum of its valid values, so every encoded value is small and non-negative. Null slots don't count towards the minimum.
  • Null rows encode as 0, as in FoR::encode.
  • The minimum and the subtraction run per chunk while it is in cache, without branches. Each value is masked with all ones or all zeros from its validity bit, over fixed 64-value blocks, so both loops vectorize at every integer width. I checked the aarch64 assembly: umin/smin and sub at .8b/.8h/.4s/.2d, with no per-value branches.
  • A chunk with no valid values reuses the previous chunk's reference (or the first valid reference at the start), so the references compress into runs.
  • The references are a plain primitive child and the offset is 0. Compressing the references is left to the compressor scheme.
chunk 0 values: 1_000_000 .. 1_000_900   → reference 1_000_000, encoded 0..900
chunk 1 values: all null                 → reference 1_000_000, encoded 0
chunk 2 values: 9_000_000 .. 9_000_950   → reference 9_000_000, encoded 0..950

Encoding, cargo bench -p vortex-fastlanes --bench for_encode, median per iteration on an M5 Max with 512 KiB of input (128Ki u32 or 64Ki i64 values):

FoR::encode FoR::encode_chunked
u32, non-null 13.4 µs 10.3 µs
i64, non-null 18.0 µs 12.5 µs
u32, 10% null 112 µs 27.6 µs
i64, 10% null 61.5 µs 25.0 µs

Decode

decompress_many_refs now dispatches like decompress_one_ref:

  • Fused: unsigned arrays whose encoded child is BitPacked with the same offset as the FoR. For each packed chunk it calls unchecked_unfor_pack with that chunk's reference, writing full chunks straight into the output and partial first/last chunks through a scratch buffer. Patch values are added to the reference of the chunk they fall in.
  • Unfused (add_references): everything else. It decodes encoded, then adds each chunk's reference. A uniquely owned buffer is updated in place. A shared one, as when encoded is a plain primitive child, gets the references added while it is copied, in one pass. Copying first and then adding made this path 2.2× slower than single-reference decode.

A BitPacked child with patches slices lazily into Slice(BitPacked), so sliced arrays with patches take the unfused path. Single-reference FoR behaves the same way today.

Decoding, --bench for_decode, same machine and input size. Every chunk spans the same range, so both encodings pack at 7 bits and the difference is the cost of the per-chunk references:

one reference per-chunk references
u32, primitive child 8.46 µs 8.71 µs
i64, primitive child 8.46 µs 8.50 µs
u32, BitPacked child (fused) 9.58 µs 9.67 µs
i64, BitPacked child (unfused) 19.1 µs 19.2 µs

Single-reference decode is as fast as it was before #10105. On aarch64, the per-chunk reference adds compile to four 128-bit add.4s/add.2d per iteration, and fused decode calls the same unchecked_unfor_pack kernel as single-reference decode.

Benchmarks

for_encode and for_decode cover u32 and i64 inputs of 256 KiB and 512 KiB. They carry #[cpu_features], so CodSpeed measures them on the walltime legs instead of in simulation. Locally, every case takes 4–112 µs per iteration.

CUDA

CUDA FoR decoding now returns an error for per-chunk references instead of falling back to the CPU.

@mhk197
mhk197 added this pull request to stack #10109 September 28, 2026 20:07
@codspeed

codspeed Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

✅ 2112 untouched benchmarks
🆕 15 new benchmarks
⏩ 461 skipped benchmarks1
🗄️ 1 archived benchmark run2

Performance Changes

Mode Benchmark BASE HEAD Efficiency
🆕 WallTime decode_bitpacked_chunked_avx2[i64, 524288] N/A 34.2 µs N/A
🆕 WallTime decode_bitpacked_chunked_avx2[u32, 524288] N/A 19.8 µs N/A
🆕 WallTime decode_chunked_avx2[i64, 524288] N/A 19.1 µs N/A
🆕 WallTime encode_chunked_avx2[i64, 524288] N/A 23.8 µs N/A
🆕 WallTime encode_chunked_nullable_avx2[i64, 524288] N/A 135.6 µs N/A
🆕 WallTime decode_bitpacked_chunked_avx512[i64, 524288] N/A 31.8 µs N/A
🆕 WallTime decode_bitpacked_chunked_avx512[u32, 524288] N/A 17.2 µs N/A
🆕 WallTime decode_chunked_avx512[i64, 524288] N/A 18.5 µs N/A
🆕 WallTime encode_chunked_avx512[i64, 524288] N/A 18.7 µs N/A
🆕 WallTime encode_chunked_nullable_avx512[i64, 524288] N/A 59.4 µs N/A
🆕 WallTime decode_bitpacked_chunked_neon[i64, 524288] N/A 40.3 µs N/A
🆕 WallTime decode_bitpacked_chunked_neon[u32, 524288] N/A 26 µs N/A
🆕 WallTime decode_chunked_neon[i64, 524288] N/A 13.8 µs N/A
🆕 WallTime encode_chunked_neon[i64, 524288] N/A 35.6 µs N/A
🆕 WallTime encode_chunked_nullable_neon[i64, 524288] N/A 64.1 µs N/A

Comparing mk/for-chunked-03-encoder (28649fb) with develop (e868676)

Open in CodSpeed

Footnotes

  1. 461 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

  2. 1 benchmark was run, but is now archived. If it was deleted in another branch, consider rebasing to remove it from the report. Instead if it was added back, click here to restore it. ↩

@mhk197 mhk197 added the changelog/feature A new feature label Sep 28, 2026
@mhk197
mhk197 force-pushed the mk/for-chunked-03-encoder branch from e5f9ac0 to 66ba06b Compare September 28, 2026 20:53
@mhk197 mhk197 changed the title Encode FoR with one reference per 1024-element chunk Encode/Decode Kernels for BlockedFoR Sep 29, 2026
Base automatically changed from mk/for-chunked-02-references to develop September 29, 2026 14:58
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
@mhk197
mhk197 force-pushed the mk/for-chunked-03-encoder branch from b851595 to 03bb58f Compare September 29, 2026 14:58
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
@mhk197 mhk197 linked an issue Sep 29, 2026 that may be closed by this pull request
@joseph-isaacs

Copy link
Copy Markdown
Contributor

Do we need 96 new benchmarks?

Comment on lines +48 to +50
let (encoded, references) = match_each_integer_ptype!(array.ptype(), |T| {
let (encoded, references) = compress_chunked::<T>(array.as_slice::<T>(), &mask);
(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you put the match each in its own func?

Comment on lines +265 to +305
let reference = references[range.start / FL_CHUNK_SIZE];
// `range` counts from the start of the first chunk, and the output starts at `offset`.
let skip = offset.saturating_sub(range.start);
let dst = &mut output[range.start + skip - offset..range.end - offset];
if dst.len() == FL_CHUNK_SIZE {
// SAFETY: `packed` holds one chunk at `bit_width` and `dst` has room for a chunk.
unsafe {
FoR::unchecked_unfor_pack(
bit_width,
packed,
reference,
mem::transmute::<&mut [MaybeUninit<T>], &mut [T]>(dst),
);
}
} else {
// SAFETY: as above, with `scratch` as the destination.
unsafe {
FoR::unchecked_unfor_pack(
bit_width,
packed,
reference,
mem::transmute::<&mut [MaybeUninit<T>], &mut [T]>(&mut scratch[..]),
);
}
dst.copy_from_slice(&scratch[skip..range.len()]);
}
},
)?;

if let Some(patches) = bp.patches() {
let indices = patches.indices().clone().execute::<PrimitiveArray>(ctx)?;
let values = patches.values().clone().execute::<PrimitiveArray>(ctx)?;
let values = values.as_slice::<T>();
match_each_unsigned_integer_ptype!(indices.ptype(), |P| {
for (&index, &value) in indices.as_slice::<P>().iter().zip_eq(values) {
let index = <P as AsPrimitive<usize>>::as_(index) - patches.offset();
let reference = references[(offset + index) / FL_CHUNK_SIZE];
uninit_range.set_value(index, value.wrapping_add(&reference));
}
});
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you share logic with fused_decompress or at least funcs inside

}

/// Encode a primitive array with one Frame of Reference per 1024-element chunk.
pub fn encode_chunked(array: PrimitiveArray, ctx: &mut ExecutionCtx) -> VortexResult<FoRArray> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

just a note: it would be nice to avoid the ctx

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we need validity though?

Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
@mhk197
mhk197 marked this pull request as ready for review September 30, 2026 14:12
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Comment on lines +84 to +106
#[vortex_bench_support::cpu_features]
#[divan::bench(types = [i64], args = INPUT_BYTES)]
fn encode<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) {
run::<T>(bencher, bytes, false, FoR::encode);
}

#[vortex_bench_support::cpu_features]
#[divan::bench(types = [i64], args = INPUT_BYTES)]
fn encode_chunked<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) {
run::<T>(bencher, bytes, false, FoR::encode_chunked);
}

#[vortex_bench_support::cpu_features]
#[divan::bench(types = [i64], args = INPUT_BYTES)]
fn encode_nullable<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) {
run::<T>(bencher, bytes, true, FoR::encode);
}

#[vortex_bench_support::cpu_features]
#[divan::bench(types = [i64], args = INPUT_BYTES)]
fn encode_chunked_nullable<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) {
run::<T>(bencher, bytes, true, FoR::encode_chunked);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we just have one here. Does the CPU actually matter?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for nullable yes it matters:

Screenshot 2026-09-30 at 10 34 17 AM

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can get rid of encode non-chunked ones tho? Although it may be good to see perf diff bw global and blocked references

Comment on lines +96 to +111
#[divan::bench(types = [i64], args = INPUT_BYTES)]
fn decode<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) {
run::<T>(bencher, bytes, false, false);
}

#[vortex_bench_support::cpu_features]
#[divan::bench(types = [i64], args = INPUT_BYTES)]
fn decode_chunked<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) {
run::<T>(bencher, bytes, true, false);
}

#[vortex_bench_support::cpu_features]
#[divan::bench(types = [i64], args = INPUT_BYTES)]
fn decode_bitpacked<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) {
run::<T>(bencher, bytes, false, true);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do you use custom CPU instrs here?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should autovectorize...

Signed-off-by: Matt Katz <mhkatz97@gmail.com>
@mhk197
mhk197 merged commit 812de38 into develop Sep 30, 2026
134 of 144 checks passed
@mhk197
mhk197 deleted the mk/for-chunked-03-encoder branch September 30, 2026 17:30
mhk197 added a commit that referenced this pull request Sep 30, 2026
Stacked on #10118.

## Summary

Adds a wire format for `FoR` arrays whose chunks have different
references, and makes it the in-memory ID, following `DecimalByteParts`.

| ID | Role |
|---|---|
| `fastlanes.for` | Frozen wire format for single-reference arrays.
Unchanged. |
| `fastlanes.for.v2` | In-memory ID, and the wire format for varying
references. |

`FoRPlugin` declares both serialized IDs and picks one per array:

- **Constant references** write `fastlanes.for` exactly as before: the
reference in the metadata and one `encoded` child. Existing files read
as before, and default writes stay byte-identical.
- **Varying references** write `fastlanes.for.v2`.

## `fastlanes.for.v2` layout

- **Children**: `[encoded, references]`. The references are serialized
like any other child, so the compressor can compress them.
- **Metadata**: a prost message holding only `offset`, the position of
the first element within the first chunk. The references' dtype (the
array's dtype, non-nullable) and length (`ceil((offset + len) / 1024)`)
are derived rather than stored.

Deserialization checks there are no buffers and exactly two children,
and validates the offset before deriving the references length.

## Editions

`fastlanes.for.v2` is in no edition. Writers that enforce editions
reject it, and readers that predate it report an unknown encoding. The
compressor never produces varying references yet, so default writes
don't change. `FoRScheme::produced_encodings` now declares
`for_v1_id()`, since `FoR.id()` is the in-memory ID.

---------

Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/feature A new feature

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Blocked FoR: Support blocked references in encode and decode kernels

2 participants