Repository navigation
[Python] Delegate local full-text search to Rust core - #10494
Conversation
leaves12138
left a comment
There was a problem hiding this comment.
Reviewed head 0d903a5 against Rust head e1069c5a9f4c424144d2a201283daa94e90bea7c and the current Java full-text/scalar implementation.
The native qualification boundary, typed analyzer/predicate forwarding, authoritative snapshot selection, candidate filtering before Top-K, and fail-closed execution behavior look sound in the reviewed paths. I built the paired Rust wheel locally rather than testing against an older installed extension. The 98 new full-text cases plus 8 native-gate regression cases passed (106 total), with all five gates enabled and nonzero counters for every update category. Changed-file flake8 and Python 3.6 syntax checks passed.
However, the existing full-text regression suite was not migrated to the changed Java-aligned fallback semantics. This is a confirmed CI blocker, not a request to restore the old non-Java behavior. Details and a base/head comparison are inline. Java parity was checked at source level, not via a Java engine interoperability run.
leaves12138
left a comment
There was a problem hiding this comment.
Re-reviewed head ab44d71 with Rust head 3badea107c3a8f434c82aec239a0b89cc3ad8049.
The four old regression failures are addressed by the updated fixtures and explicit Java-aligned coverage/refinement expectations. The selected classic suites now pass 83 cases. The existing native full-text suite plus native-gate regressions pass 106 cases, with all five gates and every update category exercised. Changed-file flake8 and Python 3.6 syntax checks passed. Production sources are unchanged since the prior head; the paired extension was reused only after verifying identical Rust/binding production code.
Additional cross-route coverage reveals one remaining mismatch for bitmap contains/endsWith/LIKE with global-index.filter.refine-from-data=false. The current Python result-level exactness exemption does not match the current Java/Rust predicate-level policy. A 24-case REST comparison matrix has 6 failures (bitmap, refinement disabled, each operator with and without an unindexed tail) and 18 passes (BTree controls and refinement-enabled cases). Details and concrete row IDs are inline.
Please keep Rust/Java behavior intact and align the classic Python route and its updated test expectation. Hosted CI is still running; Java was compared at source level, not via a Java engine run.
leaves12138
left a comment
There was a problem hiding this comment.
Rechecked Python head a604c96fe4b5603133f9eb358eaa06a2ab8837ce against Rust head 3badea107c3a8f434c82aec239a0b89cc3ad8049 and the Java predicate/exactness contract.
The previous bitmap candidate-policy finding is fixed. The shared predicate-level helper now handles ordinary contains/endsWith/residual LIKE and the earlier independent 24-case BTree/bitmap parity matrix passes. I will resolve that previous thread.
Validation completed with the matching Rust wheel, Python 3.10 and PyArrow 19:
- Classic full-text/scalar/vector/sorted-index selection: 217 passed, 2 skipped, 19 subtests passed.
- Native full-text plus planner/reader/writer/update/merge/commit gate regressions: 152 passed; every enabled native gate was exercised.
- Native vector search regression selection: 94 passed.
- Previous independent candidate-policy matrix: 24 passed.
- Changed-file flake8 and Python 3.6 syntax parsing: passed for 12 Python files.
One remaining Java-consistency issue is reproduced: native predicate normalization folds empty/all-wildcard leaves into exact Eq/IsNotNull predicates, bypassing the candidate refinement contract that the corrected classic route retains. The focused 40-case matrix gives 20 failures with refinement disabled and 20 passing controls with refinement enabled. Details are in the inline comment. Java comparison here is source-contract verification, not a Java engine interoperability run.
leaves12138
left a comment
There was a problem hiding this comment.
Rechecked Python head 5d2c640095b7d88f8b6641cb424007aa784f2076 with a rebuilt local Rust wheel from PR #1096 head 1ccf5baa1ad29e35f6220c3595cc24fdf4488b02.
The previous native predicate-folding issue is fixed: all 40 independent empty/all-wildcard candidate-policy comparisons pass. The earlier 24 ordinary contains/endsWith/LIKE candidate-policy comparisons also pass. I will resolve the previous folding thread.
One remaining Java-consistency issue is reproduced: Python's BTree reader has only an inexact all-non-null fallback for StartsWith, whereas Java/Rust use exact prefix intervals. The new full-text refinement gate therefore excludes ordinary matching BTree prefixes on classic with the default refinement=false, while native retains them. The inline comment includes complete/partial coverage reproduction and a four-case base-branch confirmation showing this PR's classic behavior regression.
Validation with Python 3.10, PyArrow 19 and the rebuilt local wheel:
- Classic full-text/scalar/vector/sorted-index selection: 217 passed, 2 skipped, 19 subtests passed.
- Native full-text and all enabled planner/reader/writer/update/merge/commit gates: 192 passed.
- Native vector search: 94 passed.
- Independent review matrix: 98 passed, 6 failed, all six failures from BTree prefix exactness with refinement disabled (four ordinary-prefix and two empty-prefix cases).
- Base-branch classic ordinary-prefix controls: 4 passed.
- Flake8 and Python 3.6 syntax parsing: passed for all 12 changed Python files.
Java comparison is source-contract verification, not a Java-engine interoperability run. Please align the Python BTree prefix implementation rather than relaxing the correct Rust/Java contract.
leaves12138
left a comment
There was a problem hiding this comment.
Rechecked head e1c9bd6bd6b8f2e9f1e9fbf240412ab0aa789678 against the Java source contract and the unchanged Rust #1096 head 1ccf5baa1ad29e35f6220c3595cc24fdf4488b02 (now merged).
The previous BTree-prefix finding is fixed. The classic reader now performs a real prefix interval query before returning an exact result, rather than marking an all-non-null fallback exact. The reader and metadata selector compare serialized UTF-8 byte boundaries directly, avoiding attempts to decode an exclusive upper bound that is not a valid Unicode string. Ordinary and empty StartsWith, optimized prefix LIKE, nonmatching prefixes, null-only files, Unicode boundaries, partial coverage and both refinement settings are covered.
Validation on this head with Python 3.10, PyArrow 19 and the matching local Rust wheel:
- Independent review matrix: 104 passed. All six failures from the previous review are gone; the earlier 40 empty/all-wildcard cases and 24 ordinary candidate-policy cases still pass.
- Classic BTree/full-text/scalar/vector/sorted-index selection: 231 passed, 2 skipped, 78 subtests passed.
- Adjacent BTree thread-safety/bloom/composite-fallback and primary-key sorted-index regressions: 25 passed, 129 subtests passed.
- Native full-text plus all enabled planner/reader/writer/update/merge/commit gates: 242 passed. Every enabled gate was exercised.
- Native vector search: 94 passed.
- Flake8 and Python 3.6 syntax parsing: passed for all 15 changed Python files.
No additional blocking issue found in the reviewed changes. I am resolving the previous BTree-prefix thread. Java comparison is source-contract verification, not a Java-engine interoperability run; Rust core tests are the prior run on the unchanged Rust head, not a new Rust test run in this recheck.
Purpose
Delegate existing local full-text searches to Rust core for eligible REST Data Evolution tables with row tracking and Parquet files. Query parsing, snapshot/index planning, scalar filtering, raw scans and ranking run in core; Python forwards builder configuration and wraps scored results.
The full-text implementation in apache/paimon-rust#1095 and the predicate normalization fix in apache/paimon-rust#1096 have both merged. CI keeps installing Rust main; no fork/branch pin is introduced.
Tests
PK and distributed full-text searches retain their existing path. No new index type or storage format is added.