Repository navigation
Add LO2 Prometheus metrics benchmark with real PromQL queries - #10216
Draft
joseph-isaacs wants to merge 11 commits into
Draft
joseph-isaacs wants to merge 11 commits into
joseph-isaacs wants to merge 11 commits into
Conversation
`action/bench-sql` keeps the quicker `pr` preset, now without Clickbench Sorted. The new `action/bench-sql-extended` label runs the `pr-full` preset, which covers every regular SQL benchmark. Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFeuTV21uQCX8GcNyZbKwf
Move statpopgen and FineWeb on S3 to the extended label only, and test that `pr-full` covers every `pr` target. Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFeuTV21uQCX8GcNyZbKwf
Add a `westermo` vx-bench suite over the Westermo test system performance data set (CC BY 4.0): real Prometheus node_exporter scrapes from 19 servers every 30 seconds for a month. The harness downloads the CSVs pinned to an upstream commit and converts them to Prometheus layout, one row per sample with a labels struct, sorted by series then time. The 15 queries are PromQL expressions translated to SQL, covering label matchers including regex, aggregation by label at a step, vector matching, newest-point queries, and the label and series metadata APIs. The suite runs on DataFusion only and is not added to the CI matrix. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Store `ts` as int64 milliseconds, the type Prometheus uses for sample timestamps, so time bounds are integer literals and step buckets are `ts - ts % step`. Rewrite the newest-point query as a join on each series' max timestamp. Every query now runs unchanged on DataFusion and DuckDB, so drop the DataFusion-only engine restriction. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Add an `lo2` vx-bench suite over the sample run of the LO2v2 data set (CC BY 4.0): real node_exporter and cAdvisor scrapes from the light-oauth2 microservice under load, 3,833 series and about 19 million samples with 86 label keys. The harness downloads the archive from Zenodo and converts it to Prometheus layout, one row per sample with a labels struct holding every label key as a nullable dictionary field, sorted by series then time. The 28 queries are real PromQL taken from node-mixin recording rules and alerts, awesome-prometheus-alerts host rules, the Node Exporter Full Grafana dashboard and prombench, translated to SQL that runs unchanged on DataFusion and DuckDB. Each query quotes its source expression. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Zenodo answers requests without a User-Agent with 403 Forbidden, which broke the LO2 data download. Set one on the shared reqwest client. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Add a CI catalog entry that schedules Westermo in the pr-full preset only, on DataFusion and DuckDB over Parquet and Vortex, and document it. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Add a CI catalog entry that schedules LO2 in the pr-full preset only, on DataFusion and DuckDB over Parquet and Vortex, and document it. Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Merging this PR will not alter performance
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| Simulation | bench_compare_sliced_dict_primitive[(3333, 10000)] |
79.1 µs | < 1 ns | N/A |
Comparing ji/lo2-bench (04dc516) with ji/westermo-bench (b4c33cf)
Footnotes
-
503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Signed-off-by: Claude <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
joseph-isaacs
force-pushed
the
ji/westermo-bench
branch
from
October 5, 2026 10:31
b4c33cf to
c38bf58
Compare
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on #10214, which is stacked on #10213. This PR adds an
lo2vx-bench suite that runs real PromQL queries, translated to SQL, over real Prometheus data. In CI it runs only underaction/bench-sql-extended.The data is the sample run from LO2v2, licensed CC BY 4.0. It contains node_exporter and cAdvisor scrapes from the light-oauth2 microservice under Locust load: 54 test cases over about 90 minutes, exported by the authors at a one-second step. That gives 3,833 series, about 19 million samples and 86 label keys. cAdvisor series carry up to 18 labels, including long Docker Compose hash values.
The 28 queries come from four sources that people run in production or use to benchmark Prometheus. Each query quotes its source expression, and every source is pinned to a commit or revision:
Changes
vortex-bench/src/lo2/downloads the 78 MB archive from Zenodo and converts it to Prometheus layout:labels,tsandvalue;labelsstruct holds every label key in the data as a nullable dictionary string field;tsis int64 milliseconds;vortex-bench/sql/lo2.sqlholds the queries andlo2.mddocuments the suite. The translation rules are in the SQL file's header:rate()is the increase over the window divided by the sampled time span. It omits Prometheus' extrapolation and counter-reset handling.predict_linearusesregr_slopeandregr_intercept.zip, already a workspace dependency, is added tovortex-bench.pr-fullpreset only, on DataFusion and DuckDB over Parquet and Vortex. The matrix test and the benchmarking docs are updated to match.Local results on a 4-core cloud VM, 5 iterations per query, with files in the page cache:
Checks run:
uv run --project bench-orchestrator --with pytest pytest bench-orchestrator/tests/test_matrix.py: 13 passed.uvx ruff format --checkanduvx ruff checkonbench-orchestrator: clean.vx-bench matrix pr-fulllistslo2; the other presets do not.data-gen,datafusion-benchandduckdb-benchbuilt inrelease_debugwith no warnings, anddata-gen lo2 --formats parquet,vortexsucceeded.vx-bench run lo2 -e datafusion,duckdb -f parquet,vortex -i 5completed all 112 runs.Not run:
cargo clippyandcargo nextestforvortex-bench, including the new unit tests.🤖 Generated with Claude Code
https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Generated by Claude Code