Skip to content

[Python] Delegate local full-text search to Rust core - #10494

Merged
JingsongLi merged 5 commits into
apache:masterfrom
JingsongLi:codex/native-local-full-text-search
Oct 10, 2026
Merged

JingsongLi merged 5 commits into
apache:masterfrom
JingsongLi:codex/native-local-full-text-search

Conversation

@JingsongLi

@JingsongLi JingsongLi commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Delegate existing local full-text searches to Rust core for eligible REST Data Evolution tables with row tracking and Parquet files. Query parsing, snapshot/index planning, scalar filtering, raw scans and ranking run in core; Python forwards builder configuration and wraps scored results.

The full-text implementation in apache/paimon-rust#1095 and the predicate normalization fix in apache/paimon-rust#1096 have both merged. CI keeps installing Rust main; no fork/branch pin is introduced.

  • Reuse the existing Python builder/result API and full-text analyzer option normalization.
  • Qualify support before searching. Tables outside the route or predicates not representable by the typed FFI use the Python reader. Once native execution starts, errors propagate without a second search.
  • Align the classic Python reader with current Java scalar FAST/FULL/DETAIL coverage and configurable candidate refinement, instead of forcing every scalar predicate through FULL reads.
  • Pin classic read plans, attach scalar files to indexed splits, and require an existing full-text definition before raw search.
  • Share Java's predicate exactness rules between full-text and vector searches. Contains, endsWith and residual LIKE require refinement even when a bitmap reader supplies exact rows; AND/OR cannot hide such leaves. Reuse one LIKE optimizer for equality/prefix rewrites and preserve escaped patterns.
  • Preserve Java's empty/all-wildcard search semantics through the Rust core change. Compare LIKE empty/percent/double-percent and empty contains/endsWith/startsWith on BTree/bitmap, including NULL labels, refinement on/off and complete/partial coverage.
  • Implement Java's exact BTree prefix interval query and share SST entry traversal with ordinary range queries. Query and metadata selection compare UTF-8 byte boundaries directly, including upper bounds that cannot decode to Unicode. Prefix matches now remain available with refinement disabled in both full-text and vector searches.

Tests

  • 242 passed with native plan/read/write/update/commit CI gates enabled against an extension rebuilt from Rust [doc] Add trino time travel doc #1096: 234 full-text REST cases plus 8 gate regressions.
  • Gate-batch native counters: 245 plans, 238 reads, 9 writes, 1 commit; row-ID=5, grouped=2, predicate=1, upsert=3, incremental=5, merge=1.
  • 234 REST tests compare native/classic result IDs and scores, including all full-text/scalar modes, BTree/bitmap and partial coverage, Boolean/boost queries, pre-Top-K filters, candidate refinement, typed stop words, partition predicates, pinned/empty/error REST snapshots, schema-only renames, deletion vectors, text update invalidation, invalid limits and query authorization. Candidate tests include LIKE/contains/endsWith with refinement on/off and AND/OR, plus equality/prefix and escaped LIKE controls.
  • The original empty/all-wildcard comparison matrix reproduced 20 failing refinement-disabled cases and 20 passing controls before the Rust fix; all 40 now pass. It also covers empty StartsWith as an exact predicate.
  • Prefix regressions reproduced seven ordinary/empty BTree mismatches and eight Unicode-bound errors before the Python fix. All 52 focused prefix cases now pass, covering StartsWith/optimized LIKE, BTree/bitmap, matching/missing/empty literals, refinement off/on and complete/partial coverage.
  • Eligible-route tests prohibit Python full-text scanning. Missing-index and REST execution failures verify that native errors are not retried through Python.
  • Existing classic full-text/scalar/vector regressions: 202 passed and 19 subtests passed. Vector prefix expectations follow Java exactness and verify that exact BTree prefixes do not re-read filter columns.
  • BTree file, bloom, thread and metadata regressions: 25 passed and 59 subtests passed, including duplicate keys, small data blocks, NULL-only files, empty/missing prefixes and UTF-8 successor boundaries under both pread and seek/read inputs.
  • Flake8, Python 3.6 source syntax and diff whitespace checks: passed.
  • Independent review found no remaining actionable regressions and verified concurrent prefix/range queries under both IO paths.

PK and distributed full-text searches retain their existing path. No new index type or storage format is added.

@leaves12138 leaves12138 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed head 0d903a5 against Rust head e1069c5a9f4c424144d2a201283daa94e90bea7c and the current Java full-text/scalar implementation.

The native qualification boundary, typed analyzer/predicate forwarding, authoritative snapshot selection, candidate filtering before Top-K, and fail-closed execution behavior look sound in the reviewed paths. I built the paired Rust wheel locally rather than testing against an older installed extension. The 98 new full-text cases plus 8 native-gate regression cases passed (106 total), with all five gates enabled and nonzero counters for every update category. Changed-file flake8 and Python 3.6 syntax checks passed.

However, the existing full-text regression suite was not migrated to the changed Java-aligned fallback semantics. This is a confirmed CI blocker, not a request to restore the old non-Java behavior. Details and a base/head comparison are inline. Java parity was checked at source level, not via a Java engine interoperability run.

Comment thread paimon-python/pypaimon/table/source/full_text_read.py

@leaves12138 leaves12138 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed head ab44d71 with Rust head 3badea107c3a8f434c82aec239a0b89cc3ad8049.

The four old regression failures are addressed by the updated fixtures and explicit Java-aligned coverage/refinement expectations. The selected classic suites now pass 83 cases. The existing native full-text suite plus native-gate regressions pass 106 cases, with all five gates and every update category exercised. Changed-file flake8 and Python 3.6 syntax checks passed. Production sources are unchanged since the prior head; the paired extension was reused only after verifying identical Rust/binding production code.

Additional cross-route coverage reveals one remaining mismatch for bitmap contains/endsWith/LIKE with global-index.filter.refine-from-data=false. The current Python result-level exactness exemption does not match the current Java/Rust predicate-level policy. A 24-case REST comparison matrix has 6 failures (bitmap, refinement disabled, each operator with and without an unindexed tail) and 18 passes (BTree controls and refinement-enabled cases). Details and concrete row IDs are inline.

Please keep Rust/Java behavior intact and align the classic Python route and its updated test expectation. Hosted CI is still running; Java was compared at source level, not via a Java engine run.

Comment thread paimon-python/pypaimon/table/source/full_text_read.py Outdated

@leaves12138 leaves12138 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rechecked Python head a604c96fe4b5603133f9eb358eaa06a2ab8837ce against Rust head 3badea107c3a8f434c82aec239a0b89cc3ad8049 and the Java predicate/exactness contract.

The previous bitmap candidate-policy finding is fixed. The shared predicate-level helper now handles ordinary contains/endsWith/residual LIKE and the earlier independent 24-case BTree/bitmap parity matrix passes. I will resolve that previous thread.

Validation completed with the matching Rust wheel, Python 3.10 and PyArrow 19:

  • Classic full-text/scalar/vector/sorted-index selection: 217 passed, 2 skipped, 19 subtests passed.
  • Native full-text plus planner/reader/writer/update/merge/commit gate regressions: 152 passed; every enabled native gate was exercised.
  • Native vector search regression selection: 94 passed.
  • Previous independent candidate-policy matrix: 24 passed.
  • Changed-file flake8 and Python 3.6 syntax parsing: passed for 12 Python files.

One remaining Java-consistency issue is reproduced: native predicate normalization folds empty/all-wildcard leaves into exact Eq/IsNotNull predicates, bypassing the candidate refinement contract that the corrected classic route retains. The focused 40-case matrix gives 20 failures with refinement disabled and 20 passing controls with refinement enabled. Details are in the inline comment. Java comparison here is source-contract verification, not a Java engine interoperability run.

Comment thread paimon-python/pypaimon/table/source/native_full_text_search.py

@leaves12138 leaves12138 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rechecked Python head 5d2c640095b7d88f8b6641cb424007aa784f2076 with a rebuilt local Rust wheel from PR #1096 head 1ccf5baa1ad29e35f6220c3595cc24fdf4488b02.

The previous native predicate-folding issue is fixed: all 40 independent empty/all-wildcard candidate-policy comparisons pass. The earlier 24 ordinary contains/endsWith/LIKE candidate-policy comparisons also pass. I will resolve the previous folding thread.

One remaining Java-consistency issue is reproduced: Python's BTree reader has only an inexact all-non-null fallback for StartsWith, whereas Java/Rust use exact prefix intervals. The new full-text refinement gate therefore excludes ordinary matching BTree prefixes on classic with the default refinement=false, while native retains them. The inline comment includes complete/partial coverage reproduction and a four-case base-branch confirmation showing this PR's classic behavior regression.

Validation with Python 3.10, PyArrow 19 and the rebuilt local wheel:

  • Classic full-text/scalar/vector/sorted-index selection: 217 passed, 2 skipped, 19 subtests passed.
  • Native full-text and all enabled planner/reader/writer/update/merge/commit gates: 192 passed.
  • Native vector search: 94 passed.
  • Independent review matrix: 98 passed, 6 failed, all six failures from BTree prefix exactness with refinement disabled (four ordinary-prefix and two empty-prefix cases).
  • Base-branch classic ordinary-prefix controls: 4 passed.
  • Flake8 and Python 3.6 syntax parsing: passed for all 12 changed Python files.

Java comparison is source-contract verification, not a Java-engine interoperability run. Please align the Python BTree prefix implementation rather than relaxing the correct Rust/Java contract.

Comment thread paimon-python/pypaimon/table/source/full_text_read.py

@leaves12138 leaves12138 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rechecked head e1c9bd6bd6b8f2e9f1e9fbf240412ab0aa789678 against the Java source contract and the unchanged Rust #1096 head 1ccf5baa1ad29e35f6220c3595cc24fdf4488b02 (now merged).

The previous BTree-prefix finding is fixed. The classic reader now performs a real prefix interval query before returning an exact result, rather than marking an all-non-null fallback exact. The reader and metadata selector compare serialized UTF-8 byte boundaries directly, avoiding attempts to decode an exclusive upper bound that is not a valid Unicode string. Ordinary and empty StartsWith, optimized prefix LIKE, nonmatching prefixes, null-only files, Unicode boundaries, partial coverage and both refinement settings are covered.

Validation on this head with Python 3.10, PyArrow 19 and the matching local Rust wheel:

  • Independent review matrix: 104 passed. All six failures from the previous review are gone; the earlier 40 empty/all-wildcard cases and 24 ordinary candidate-policy cases still pass.
  • Classic BTree/full-text/scalar/vector/sorted-index selection: 231 passed, 2 skipped, 78 subtests passed.
  • Adjacent BTree thread-safety/bloom/composite-fallback and primary-key sorted-index regressions: 25 passed, 129 subtests passed.
  • Native full-text plus all enabled planner/reader/writer/update/merge/commit gates: 242 passed. Every enabled gate was exercised.
  • Native vector search: 94 passed.
  • Flake8 and Python 3.6 syntax parsing: passed for all 15 changed Python files.

No additional blocking issue found in the reviewed changes. I am resolving the previous BTree-prefix thread. Java comparison is source-contract verification, not a Java-engine interoperability run; Rust core tests are the prior run on the unchanged Rust head, not a new Rust test run in this recheck.

@JingsongLi
JingsongLi marked this pull request as ready for review October 10, 2026 15:09
@JingsongLi
JingsongLi merged commit af4b262 into apache:master Oct 10, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants