Skip to content

[python] Add atomic global index maintenance and replacement - #10431

Open
TheR1sing3un wants to merge 7 commits into
apache:masterfrom
TheR1sing3un:contribute/global-index-maintenance
Open

TheR1sing3un wants to merge 7 commits into
apache:masterfrom
TheR1sing3un:contribute/global-index-maintenance

Conversation

@TheR1sing3un

Copy link
Copy Markdown
Member

Purpose

Add opt-in maintain_global_indexes / multimodal maintain_indexes for configured global indexes. Each pass plans one snapshot, fills uncovered ranges, and publishes all requested definitions in one commit. Repeated catch-up is a no-op; keeping explicit definitions lets scheduled ingestion/backfill jobs rebuild coverage removed by DROP_PARTITION_INDEX updates.

rebuild=True atomically replaces matching index entries, optionally for selected partitions, without removing files retained by older snapshots. A strict Python-committer snapshot guard rejects publication after any concurrent commit. Failed builds preserve existing coverage and clean completed uncommitted outputs; ambiguous commit failures leave files for orphan maintenance. The API does not create a background service: callers own scheduling, backoff and retries. IGNORE updates require explicit rebuild.

Tests

83 targeted maintenance/build/conflict tests passed across the initial 81-test run and the final expanded 10-test maintenance run. Coverage includes idempotence, append catch-up, rebuilding update-invalidated partitions, one-commit publication of multiple definitions, partition-scoped replacement, historical index retention, concurrent-commit rejection, failed-build cleanup, ignored-update repair and validation. Flake8 and git diff --check passed. No benchmark code included.

@JingsongLi

Copy link
Copy Markdown
Contributor

[P2] Reject snapshot changes without rolling back concurrent compaction

The new strict snapshot check in ConflictDetection.check_conflicts returns a normal commit conflict. Maintenance sets check_from_snapshot on every message, which also sets allow_rollback=True in FileStoreCommit.commit. When a rollback-capable catalog has published a concurrent COMPACT, _try_commit_once consequently calls CommitRollback.try_to_rollback, rolls the table back to the preceding snapshot, and retries the maintenance commit. It can then publish successfully instead of rejecting the stale plan.

This contradicts the documented strict guard and lets a scheduled index maintenance pass undo a completed compaction. Please make this snapshot-change conflict non-rollbackable, or disable compaction rollback for maintenance/index-only messages. The existing concurrent-append test uses a filesystem catalog without catalog rollback and does not exercise this branch.

Validation: 90 targeted tests and 22 subtests passed locally, plus project Flake8. I also reproduced the rollback using real Parquet data, a real BTree rebuild and the production committer; only the catalog rollback endpoint was injected to delegate to local rollback. Planned snapshot 2 was followed by a physical-file-replacing COMPACT 3. Current code called rollback 3→2 and published maintenance as APPEND 3, restoring the original data file. Using the pre-PR conflict-check implementation instead published APPEND 4 with zero rollback calls and preserved the compacted file. Both variants retained the two logical rows; this finding does not claim data loss.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants