Skip to content

[python] Return search results through distributed Ray row lookup - #10429

Merged
JingsongLi merged 3 commits into
apache:masterfrom
TheR1sing3un:contribute/ray-search-dataset
Oct 9, 2026
Merged

JingsongLi merged 3 commits into
apache:masterfrom
TheR1sing3un:contribute/ray-search-dataset

Conversation

@TheR1sing3un

@TheR1sing3un TheR1sing3un commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Purpose

Add to_ray() to vector search and batch vector search queries. Candidate search resolves one snapshot eagerly; final row lookup uses the existing Paimon Ray datasource lazily on workers, avoiding driver materialization of selected columns. Batch queries return one Dataset per input vector.

Preserve pre/post filters, deletion vectors, score output, deterministic global score ordering, typed empty results, and snapshot consistency across later updates/deletes. Return BLOB descriptors with the metadata needed by map_with_blobs. Document snapshot retention and the independent batch lookup pipelines.

Tests

11 real-Ray integration tests passed for local/Ray candidate execution, distributed lookup, score ordering, snapshot pinning, filters/empty schema (including never-written tables), metadata-only results, batch queries, and resolving BLOBs on workers. Another 25 existing multimodal search/Ray regression tests passed. Flake8 and git diff --check passed. No benchmark code included.

Integration validation with the other four proposed index/Ray changes also passed 46 targeted tests, including Ray full-text/Hybrid queries returning Datasets and Ray-built indexes followed by atomic maintenance and distributed lookup.

Review follow-up: explicitly preserve block order on the returned Dataset when score ordering is requested, including projection stages after Sort, without modifying the global DataContext or unrelated Datasets. A slow-first-partition regression fails against the previous implementation on Ray 2.59.0. The fixed implementation passes all 13 search Dataset/order tests on Ray 2.59.0, and both new ordering regressions also pass on Ray 2.54.0.

@JingsongLi

Copy link
Copy Markdown
Contributor

[P2] Preserve block order when dropping columns after the global score sort

In search_result.py:74–77, the sorted Dataset is followed by drop_columns calls. With Ray 2.59.0, preserve_order defaults to False, and Sort does not automatically enable it. Since finish_batch leaves the inferred output schema unknown, these calls run as parallel MapBatches stages and emit blocks in task completion order. A slow earlier partition can therefore place lower-scoring rows before higher-scoring rows, violating order_by_score() and potentially causing take(1) to return a candidate other than the highest-scoring one. This also affects batch queries using the same helper.

I reproduced this using a real table with 24 rows across four files, with row i containing vector [float(i), 0.0]:

ds = (table.search([0.0, 0.0]).select(["id"])
      .with_score().order_by_score().limit(24)
      .to_ray(execution="local", override_num_blocks=4))

Adding a one-second delay to the first sorted partition's column-dropping task, without changing its data or projection, produced IDs [5, ..., 23, 0, ..., 4] instead of [0, ..., 23]. The same reproduction with preserve_order=True returned the expected order. The unmodified execution plan also confirms that both column-dropping stages occur after Sort.

Please explicitly preserve order for this Dataset when score ordering is requested, without changing the global DataContext, and add a regression test with multiple blocks and a slow first partition. The 11 new tests pass with the default configuration, so their current cases do not expose this scheduling-dependent failure.

@JingsongLi

Copy link
Copy Markdown
Contributor

+1

@JingsongLi
JingsongLi merged commit 787f16c into apache:master Oct 9, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants