Repository navigation
Encode/Decode Kernels for BlockedFoR - #10118
Conversation
Merging this PR will not alter performance
|
e5f9ac0 to
66ba06b
Compare
FoR with one reference per 1024-element chunkFoR
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
b851595 to
03bb58f
Compare
|
Do we need 96 new benchmarks? |
| let (encoded, references) = match_each_integer_ptype!(array.ptype(), |T| { | ||
| let (encoded, references) = compress_chunked::<T>(array.as_slice::<T>(), &mask); | ||
| ( |
There was a problem hiding this comment.
can you put the match each in its own func?
| let reference = references[range.start / FL_CHUNK_SIZE]; | ||
| // `range` counts from the start of the first chunk, and the output starts at `offset`. | ||
| let skip = offset.saturating_sub(range.start); | ||
| let dst = &mut output[range.start + skip - offset..range.end - offset]; | ||
| if dst.len() == FL_CHUNK_SIZE { | ||
| // SAFETY: `packed` holds one chunk at `bit_width` and `dst` has room for a chunk. | ||
| unsafe { | ||
| FoR::unchecked_unfor_pack( | ||
| bit_width, | ||
| packed, | ||
| reference, | ||
| mem::transmute::<&mut [MaybeUninit<T>], &mut [T]>(dst), | ||
| ); | ||
| } | ||
| } else { | ||
| // SAFETY: as above, with `scratch` as the destination. | ||
| unsafe { | ||
| FoR::unchecked_unfor_pack( | ||
| bit_width, | ||
| packed, | ||
| reference, | ||
| mem::transmute::<&mut [MaybeUninit<T>], &mut [T]>(&mut scratch[..]), | ||
| ); | ||
| } | ||
| dst.copy_from_slice(&scratch[skip..range.len()]); | ||
| } | ||
| }, | ||
| )?; | ||
|
|
||
| if let Some(patches) = bp.patches() { | ||
| let indices = patches.indices().clone().execute::<PrimitiveArray>(ctx)?; | ||
| let values = patches.values().clone().execute::<PrimitiveArray>(ctx)?; | ||
| let values = values.as_slice::<T>(); | ||
| match_each_unsigned_integer_ptype!(indices.ptype(), |P| { | ||
| for (&index, &value) in indices.as_slice::<P>().iter().zip_eq(values) { | ||
| let index = <P as AsPrimitive<usize>>::as_(index) - patches.offset(); | ||
| let reference = references[(offset + index) / FL_CHUNK_SIZE]; | ||
| uninit_range.set_value(index, value.wrapping_add(&reference)); | ||
| } | ||
| }); | ||
| } |
There was a problem hiding this comment.
can you share logic with fused_decompress or at least funcs inside
| } | ||
|
|
||
| /// Encode a primitive array with one Frame of Reference per 1024-element chunk. | ||
| pub fn encode_chunked(array: PrimitiveArray, ctx: &mut ExecutionCtx) -> VortexResult<FoRArray> { |
There was a problem hiding this comment.
just a note: it would be nice to avoid the ctx
There was a problem hiding this comment.
we need validity though?
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
| #[vortex_bench_support::cpu_features] | ||
| #[divan::bench(types = [i64], args = INPUT_BYTES)] | ||
| fn encode<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) { | ||
| run::<T>(bencher, bytes, false, FoR::encode); | ||
| } | ||
|
|
||
| #[vortex_bench_support::cpu_features] | ||
| #[divan::bench(types = [i64], args = INPUT_BYTES)] | ||
| fn encode_chunked<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) { | ||
| run::<T>(bencher, bytes, false, FoR::encode_chunked); | ||
| } | ||
|
|
||
| #[vortex_bench_support::cpu_features] | ||
| #[divan::bench(types = [i64], args = INPUT_BYTES)] | ||
| fn encode_nullable<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) { | ||
| run::<T>(bencher, bytes, true, FoR::encode); | ||
| } | ||
|
|
||
| #[vortex_bench_support::cpu_features] | ||
| #[divan::bench(types = [i64], args = INPUT_BYTES)] | ||
| fn encode_chunked_nullable<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) { | ||
| run::<T>(bencher, bytes, true, FoR::encode_chunked); | ||
| } |
There was a problem hiding this comment.
can we just have one here. Does the CPU actually matter?
There was a problem hiding this comment.
I can get rid of encode non-chunked ones tho? Although it may be good to see perf diff bw global and blocked references
| #[divan::bench(types = [i64], args = INPUT_BYTES)] | ||
| fn decode<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) { | ||
| run::<T>(bencher, bytes, false, false); | ||
| } | ||
|
|
||
| #[vortex_bench_support::cpu_features] | ||
| #[divan::bench(types = [i64], args = INPUT_BYTES)] | ||
| fn decode_chunked<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) { | ||
| run::<T>(bencher, bytes, true, false); | ||
| } | ||
|
|
||
| #[vortex_bench_support::cpu_features] | ||
| #[divan::bench(types = [i64], args = INPUT_BYTES)] | ||
| fn decode_bitpacked<T: NativePType + TryFrom<usize>>(bencher: Bencher, bytes: usize) { | ||
| run::<T>(bencher, bytes, false, true); | ||
| } |
There was a problem hiding this comment.
do you use custom CPU instrs here?
There was a problem hiding this comment.
Should autovectorize...
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Stacked on #10118. ## Summary Adds a wire format for `FoR` arrays whose chunks have different references, and makes it the in-memory ID, following `DecimalByteParts`. | ID | Role | |---|---| | `fastlanes.for` | Frozen wire format for single-reference arrays. Unchanged. | | `fastlanes.for.v2` | In-memory ID, and the wire format for varying references. | `FoRPlugin` declares both serialized IDs and picks one per array: - **Constant references** write `fastlanes.for` exactly as before: the reference in the metadata and one `encoded` child. Existing files read as before, and default writes stay byte-identical. - **Varying references** write `fastlanes.for.v2`. ## `fastlanes.for.v2` layout - **Children**: `[encoded, references]`. The references are serialized like any other child, so the compressor can compress them. - **Metadata**: a prost message holding only `offset`, the position of the first element within the first chunk. The references' dtype (the array's dtype, non-nullable) and length (`ceil((offset + len) / 1024)`) are derived rather than stored. Deserialization checks there are no buffers and exactly two children, and validates the offset before deriving the references length. ## Editions `fastlanes.for.v2` is in no edition. Writers that enforce editions reject it, and readers that predate it report an unknown encoding. The compressor never produces varying references yet, so default writes don't change. `FoRScheme::produced_encodings` now declares `for_v1_id()`, since `FoR.id()` is the in-memory ID. --------- Signed-off-by: Matt Katz <mhkatz97@gmail.com>

Stacked on #10105.
Summary
Adds an encoder that picks one reference per 1024-element chunk, and a fused decode for chunked FoR over BitPacked. No wire format changes: arrays with varying references still can't be serialized until the next PR.
Encoder
FoR::encode_chunked(array, ctx)uses the same rule asFoR::encode, applied per chunk:FoR::encode.umin/sminandsubat.8b/.8h/.4s/.2d, with no per-value branches.Encoding,
cargo bench -p vortex-fastlanes --bench for_encode, median per iteration on an M5 Max with 512 KiB of input (128Kiu32or 64Kii64values):FoR::encodeFoR::encode_chunkedDecode
decompress_many_refsnow dispatches likedecompress_one_ref:encodedchild isBitPackedwith the same offset as the FoR. For each packed chunk it callsunchecked_unfor_packwith that chunk's reference, writing full chunks straight into the output and partial first/last chunks through a scratch buffer. Patch values are added to the reference of the chunk they fall in.add_references): everything else. It decodesencoded, then adds each chunk's reference. A uniquely owned buffer is updated in place. A shared one, as whenencodedis a plain primitive child, gets the references added while it is copied, in one pass. Copying first and then adding made this path 2.2× slower than single-reference decode.A BitPacked child with patches slices lazily into
Slice(BitPacked), so sliced arrays with patches take the unfused path. Single-reference FoR behaves the same way today.Decoding,
--bench for_decode, same machine and input size. Every chunk spans the same range, so both encodings pack at 7 bits and the difference is the cost of the per-chunk references:Single-reference decode is as fast as it was before #10105. On aarch64, the per-chunk reference adds compile to four 128-bit
add.4s/add.2dper iteration, and fused decode calls the sameunchecked_unfor_packkernel as single-reference decode.Benchmarks
for_encodeandfor_decodecoveru32andi64inputs of 256 KiB and 512 KiB. They carry#[cpu_features], so CodSpeed measures them on the walltime legs instead of in simulation. Locally, every case takes 4–112 µs per iteration.CUDA
CUDA FoR decoding now returns an error for per-chunk references instead of falling back to the CPU.