Skip to content

Add LO2 Prometheus metrics benchmark with real PromQL queries - #10216

Draft
joseph-isaacs wants to merge 11 commits into
ji/westermo-benchfrom
ji/lo2-bench
Draft

joseph-isaacs wants to merge 11 commits into
ji/westermo-benchfrom
ji/lo2-bench

Conversation

@joseph-isaacs

Copy link
Copy Markdown
Contributor

Summary

Stacked on #10214, which is stacked on #10213. This PR adds an lo2 vx-bench suite that runs real PromQL queries, translated to SQL, over real Prometheus data. In CI it runs only under action/bench-sql-extended.

The data is the sample run from LO2v2, licensed CC BY 4.0. It contains node_exporter and cAdvisor scrapes from the light-oauth2 microservice under Locust load: 54 test cases over about 90 minutes, exported by the authors at a one-second step. That gives 3,833 series, about 19 million samples and 86 label keys. cAdvisor series carry up to 18 labels, including long Docker Compose hash values.

The 28 queries come from four sources that people run in production or use to benchmark Prometheus. Each query quotes its source expression, and every source is pinned to a commit or revision:

Source Queries
node-mixin recording rules and one alert (prometheus/node_exporter) Q0 to Q9
awesome-prometheus-alerts host rules, as run by VictoriaMetrics' prometheus-benchmark Q10 to Q16
Node Exporter Full, Grafana dashboard 1860 Q17 to Q22
prombench load generator (prometheus/test-infra) Q23 to Q27

Changes

  • Data conversion. vortex-bench/src/lo2/ downloads the 78 MB archive from Zenodo and converts it to Prometheus layout:
    • one row per sample, with columns labels, ts and value;
    • the labels struct holds every label key in the data as a nullable dictionary string field;
    • ts is int64 milliseconds;
    • rows are sorted by series, then time.
  • Queries and docs. vortex-bench/sql/lo2.sql holds the queries and lo2.md documents the suite. The translation rules are in the SQL file's header:
    • Instant queries are evaluated at a fixed time. Dashboard panels run as a one-hour range query at a one-minute step.
    • rate() is the increase over the window divided by the sampled time span. It omits Prometheus' extrapolation and counter-reset handling.
    • predict_linear uses regr_slope and regr_intercept.
    • Vector matching becomes joins or pivots.
  • Portable SQL. Every query runs unchanged on DataFusion and DuckDB.
  • Expected empty results. The host was healthy during the run, so the alert queries return no rows. They still read and filter the same data a Prometheus server would.
  • Dependency. zip, already a workspace dependency, is added to vortex-bench.
  • Download fix. The shared download client now sends a User-Agent, because Zenodo answers requests without one with 403 Forbidden.
  • CI. A catalog entry schedules the suite in the pr-full preset only, on DataFusion and DuckDB over Parquet and Vortex. The matrix test and the benchmarking docs are updated to match.
  • Registration. The suite is registered alongside Westermo in the CLI, the dataset enum, the v3 emitter, the orchestrator and the bench-performance skill.

Local results on a 4-core cloud VM, 5 iterations per query, with files in the page cache:

DataFusion DuckDB
Vortex time relative to Parquet, geometric mean 0.20 0.80
Queries where Vortex is faster 28 of 28 22 of 28
File Size
Parquet, zstd level 3 50 MB
Vortex 79 MB

Checks run:

  • uv run --project bench-orchestrator --with pytest pytest bench-orchestrator/tests/test_matrix.py: 13 passed.
  • uvx ruff format --check and uvx ruff check on bench-orchestrator: clean.
  • vx-bench matrix pr-full lists lo2; the other presets do not.
  • Before rebasing onto Split SQL PR benchmarks into bench-sql and bench-sql-extended #10213: data-gen, datafusion-bench and duckdb-bench built in release_debug with no warnings, and data-gen lo2 --formats parquet,vortex succeeded.
  • Before rebasing onto Split SQL PR benchmarks into bench-sql and bench-sql-extended #10213: vx-bench run lo2 -e datafusion,duckdb -f parquet,vortex -i 5 completed all 112 runs.
  • All 28 queries return identical results on DataFusion and DuckDB.
  • The Parquet file written by the Rust converter gives identical results on all 28 queries to an independent Python conversion of the same archive.
  • Every metric's series is uniquely identified by the labels the queries group on.

Not run:

  • cargo clippy and cargo nextest for vortex-bench, including the new unit tests.
  • A rebuild after the rebase.

🤖 Generated with Claude Code

https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ


Generated by Claude Code

claude added 8 commits October 2, 2026 08:40
`action/bench-sql` keeps the quicker `pr` preset, now without Clickbench
Sorted. The new `action/bench-sql-extended` label runs the `pr-full`
preset, which covers every regular SQL benchmark.

Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFeuTV21uQCX8GcNyZbKwf
Move statpopgen and FineWeb on S3 to the extended label only, and test
that `pr-full` covers every `pr` target.

Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFeuTV21uQCX8GcNyZbKwf
Add a `westermo` vx-bench suite over the Westermo test system performance
data set (CC BY 4.0): real Prometheus node_exporter scrapes from 19 servers
every 30 seconds for a month. The harness downloads the CSVs pinned to an
upstream commit and converts them to Prometheus layout, one row per sample
with a labels struct, sorted by series then time.

The 15 queries are PromQL expressions translated to SQL, covering label
matchers including regex, aggregation by label at a step, vector matching,
newest-point queries, and the label and series metadata APIs. The suite
runs on DataFusion only and is not added to the CI matrix.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Store `ts` as int64 milliseconds, the type Prometheus uses for sample
timestamps, so time bounds are integer literals and step buckets are
`ts - ts % step`. Rewrite the newest-point query as a join on each
series' max timestamp. Every query now runs unchanged on DataFusion and
DuckDB, so drop the DataFusion-only engine restriction.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Add an `lo2` vx-bench suite over the sample run of the LO2v2 data set
(CC BY 4.0): real node_exporter and cAdvisor scrapes from the
light-oauth2 microservice under load, 3,833 series and about 19 million
samples with 86 label keys. The harness downloads the archive from
Zenodo and converts it to Prometheus layout, one row per sample with a
labels struct holding every label key as a nullable dictionary field,
sorted by series then time.

The 28 queries are real PromQL taken from node-mixin recording rules and
alerts, awesome-prometheus-alerts host rules, the Node Exporter Full
Grafana dashboard and prombench, translated to SQL that runs unchanged
on DataFusion and DuckDB. Each query quotes its source expression.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Zenodo answers requests without a User-Agent with 403 Forbidden, which
broke the LO2 data download. Set one on the shared reqwest client.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Add a CI catalog entry that schedules Westermo in the pr-full preset
only, on DataFusion and DuckDB over Parquet and Vortex, and document it.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Add a CI catalog entry that schedules LO2 in the pr-full preset only,
on DataFusion and DuckDB over Parquet and Vortex, and document it.

Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
@joseph-isaacs joseph-isaacs added the changelog/chore A trivial change label Oct 2, 2026 — with Claude
@codspeed

codspeed Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ 1 benchmark measured no execution time

Nothing ran under measurement, usually because the compiler removed the code under test. This result is not comparable, so it counts as unchanged.

Preventing compiler optimizations

✅ 2097 untouched benchmarks
⏩ 503 skipped benchmarks1

Performance Changes

Mode Benchmark BASE HEAD Efficiency
⚠️ Simulation bench_compare_sliced_dict_primitive[(3333, 10000)] 79.1 µs < 1 ns N/A

Comparing ji/lo2-bench (04dc516) with ji/westermo-bench (b4c33cf)

Open in CodSpeed

Footnotes

  1. 503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

claude added 3 commits October 3, 2026 08:35
Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ
Signed-off-by: Claude <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXCBFkSfZNsuNNRowQfwdQ

This branch was successfully deployed

1 active deployment
docs-preview/pr-10216 — 04dc516b Deployed Oct 3, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/chore A trivial change

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants