Repository navigation
[python] Return search results through distributed Ray row lookup - #10429
Conversation
|
[P2] Preserve block order when dropping columns after the global score sort In I reproduced this using a real table with 24 rows across four files, with row ds = (table.search([0.0, 0.0]).select(["id"])
.with_score().order_by_score().limit(24)
.to_ray(execution="local", override_num_blocks=4))Adding a one-second delay to the first sorted partition's column-dropping task, without changing its data or projection, produced IDs Please explicitly preserve order for this Dataset when score ordering is requested, without changing the global DataContext, and add a regression test with multiple blocks and a slow first partition. The 11 new tests pass with the default configuration, so their current cases do not expose this scheduling-dependent failure. |
|
+1 |
Purpose
Add
to_ray()to vector search and batch vector search queries. Candidate search resolves one snapshot eagerly; final row lookup uses the existing Paimon Ray datasource lazily on workers, avoiding driver materialization of selected columns. Batch queries return one Dataset per input vector.Preserve pre/post filters, deletion vectors, score output, deterministic global score ordering, typed empty results, and snapshot consistency across later updates/deletes. Return BLOB descriptors with the metadata needed by
map_with_blobs. Document snapshot retention and the independent batch lookup pipelines.Tests
11 real-Ray integration tests passed for local/Ray candidate execution, distributed lookup, score ordering, snapshot pinning, filters/empty schema (including never-written tables), metadata-only results, batch queries, and resolving BLOBs on workers. Another 25 existing multimodal search/Ray regression tests passed. Flake8 and
git diff --checkpassed. No benchmark code included.Integration validation with the other four proposed index/Ray changes also passed 46 targeted tests, including Ray full-text/Hybrid queries returning Datasets and Ray-built indexes followed by atomic maintenance and distributed lookup.
Review follow-up: explicitly preserve block order on the returned Dataset when score ordering is requested, including projection stages after Sort, without modifying the global DataContext or unrelated Datasets. A slow-first-partition regression fails against the previous implementation on Ray 2.59.0. The fixed implementation passes all 13 search Dataset/order tests on Ray 2.59.0, and both new ordering regressions also pass on Ray 2.54.0.