Skip to content

Vortex spark integration is a fileformat not a tableprovider - #9658

Draft
robert3005 wants to merge 1 commit into
developfrom
rk/sparkfileformat
Draft

robert3005 wants to merge 1 commit into
developfrom
rk/sparkfileformat

Conversation

@robert3005

Copy link
Copy Markdown
Contributor

Following in the footsteps of duckdb integration we now claim to be a fileformat
instead of tableprovider thus better aligning with existing functionality

We also add shared java io layer that can you JvmWrite/JvmReadAt

Signed-off-by: Robert Kruszewski github@robertk.io

@robert3005
robert3005 marked this pull request as draft August 27, 2026 11:34
@codspeed

codspeed Bot commented Aug 27, 2026 •

Copy link
Copy Markdown

Merging this PR will degrade performance by 11.56%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ 12 benchmarks spent significant time in system calls

System calls cannot be consistently instrumented, so they are not included in the measure, which understates the real cost. Please switch to the Walltime instrument to accurately measure system calls.

Measurement and system calls

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

❌ 1 regressed benchmark
✅ 2182 untouched benchmarks
⏩ 359 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ WallTime bitpack_blocked_compress_avx2 6.8 µs 7.6 µs -11.56%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing rk/sparkfileformat (4c1310f) with develop (3d12093)

Open in CodSpeed

Footnotes

  1. 359 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@github-actions

Copy link
Copy Markdown
Contributor

This PR has been marked as stale because it has been open for 14 days with no activity. Please comment or remove the stale label if you wish to keep it active, otherwise it will be closed in 7 days

@github-actions github-actions Bot added the stale This PR is stale and will be auto-closed soon label Sep 23, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

This PR was closed because it has been inactive for 7 days since being marked as stale.

@github-actions github-actions Bot closed this Oct 2, 2026
@robert3005 robert3005 reopened this Oct 6, 2026
@robert3005 robert3005 removed the stale This PR is stale and will be auto-closed soon label Oct 6, 2026
@github-actions
github-actions Bot temporarily deployed to docs-preview/pr-9658 October 6, 2026 15:51 Inactive
@robert3005 robert3005 added the changelog/break A breaking API change label Oct 6, 2026
Add Hadoop-backed Java IO layer for Spark

Signed-off-by: Robert Kruszewski <github@robertk.io>

This branch was successfully deployed

1 active deployment
docs-preview/pr-9658 — 4c1310f9 Deployed Oct 10, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/break A breaking API change

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant