Skip to content

Official Task Variations (3/3): Add Official AlpacaEval, Arena-Hard, and MT-Bench Tasks - #115

Merged
kargibora merged 11 commits into
refactor/composable-pairwise-metricsfrom
feat/official-task-variants-v2
Sep 15, 2026
Merged

kargibora merged 11 commits into
refactor/composable-pairwise-metricsfrom
feat/official-task-variants-v2

Conversation

@kargibora

@kargibora kargibora commented Sep 3, 2026 •

Copy link
Copy Markdown
Member

This PR is stacked on #113. It adds the official AlpacaEval, Arena-Hard, and MT-Bench task protocols.

Problem

The existing task names used JudgeArena defaults rather than the original benchmark protocols. This made it difficult to reproduce official results or compare a JudgeArena run with the upstream benchmark.

Canonical tasks

The canonical task names now use the official defaults:

Task Default protocol
alpaca-eval Pinned AlpacaEval 2.0 data, gpt4_1106_preview baseline, official prompt, random answer order, first-token logprobs, and official length-controlled scoring
arena-hard-v0.1 Official v0.1 data and baseline, official verdict prompt, both answer orders, and official weighted scoring
arena-hard-v2.0 Official v2.0 data, category-specific baselines and prompts, both answer orders, and judge-specific style-controlled scoring
mt-bench Pinned FastChat questions and references, gpt-3.5-turbo baseline, FastChat pairwise prompt, and both answer orders

Normal runtime options can still override model and generation settings.

JudgeArena variants

The existing JudgeArena protocols remain available under explicit names:

Task Protocol
alpaca-eval-ja JudgeArena prompt, fixed answer order, pairwise win rate, and JudgeArena length-controlled win rate
arena-hard-v0.1-ja JudgeArena prompt, fixed answer order, and pairwise win rate
arena-hard-v2.0-ja Official v2.0 data and category baselines with the JudgeArena prompt, fixed answer order, and pairwise win rate

These are separate protocols rather than aliases for the canonical tasks.

Implementation

  • Add the official AlpacaEval length-controlled metric.
  • Add the official Arena-Hard v0.1 weighted metric.
  • Add the official Arena-Hard v2 style-controlled metric and packaged calibration data.
  • Put data sources, revisions, baselines, prompts, generation settings, and metric parameters in task YAML.
  • Reuse the existing JudgeArena table adapter for AlpacaEval.
  • Keep the native Arena-Hard adapter for its versioned JSONL and nested answer formats.
  • Use the shared battle-dataframe and metric interfaces from Official Task Variations (2/3): Make Scoring more Composable #113.
  • Route MT-Bench through its FastChat-compatible prompt and judging path.
  • Require an OpenRouter provider that supports top logprobs when a task requests them.

The Arena-Hard sources contain:

  • v0.1: 500 instructions, 37,510 outputs, and 76 models
  • v2.0: 750 instructions, 22,496 outputs, and 30 models

Parsed judge output

The official parsers use the structured parser result introduced in #108:

  • AlpacaEval stores the normalized preference, emitted m or M label, and the relevant raw token logprobs.
  • Arena-Hard stores the graded preference and final normalized verdict, such as A>B or A>>B.

Metrics continue to receive the canonical numeric pref column.

Protocol constraints

AlpacaEval requires first-token top logprobs. Responses without them remain unparsed; textual labels are not used as a fallback. If every preference is missing, the metric is saved as unavailable.

Arena-Hard v2 style control is judge-specific. The packaged calibration supports gpt-4.1 and gemini-2.5 with their required settings. Other judges need their own calibration data. arena-hard-v2.0-ja remains available for pairwise evaluation with arbitrary judges.

Tests

The tests cover task configuration, dataset normalization, prompt and parser behavior, official numerical results, Arena-Hard calibration, MT-Bench integration, OpenRouter capability routing, and package contents.

uv run ruff check judgearena tests
uv run pytest -q tests

Result: 359 passed locally.

Raw diff

Relative to #113 at 6a42472. Parent changes are excluded. Pre-existing tests are restored; remaining test reductions apply only to stack-added coverage.

File type Added Deleted Changed
Production Python 1,284 128 1,412
Tests 804 125 929
Task YAML, prompts, docs, packaging 237 45 282
Total 2,325 298 2,623

Estimated production changes

 ┌─────────────────────────┬───────────────┐
 │ Type of change          │ Estimated LoC │
 ├─────────────────────────┼───────────────┤
 │ New implementation      │ 1,095–1,120   │
 ├─────────────────────────┼───────────────┤
 │ Moved or extracted code │ 45–70         │
 ├─────────────────────────┼───────────────┤
 │ Small refactors         │ 5–10          │
 ├─────────────────────────┼───────────────┤
 │ Removed implementation  │ 55–65         │
 └─────────────────────────┴───────────────┘

from judgearena.prompts.registry import DEFAULT_JUDGE_PROMPT_PRESET
from judgearena.utils import strip_thinking_tags

FASTCHAT_TEMPERATURE_CONFIG: dict[str, float] = {

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Moved it into the task definition

self,
model: str,
*,
instruction_ids: pd.Index | None = None,

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need instruction_ids because some official protocol asks for filtered instructions and groups to judge different prompts with different configurations

Comment thread judgearena/evaluate.py Outdated
n_no_logprobs,
len(ann_list),
label,
raise ValueError(

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we skip or give warning or raise a ValueError here?

Comment thread judgearena/models.py
# OpenRouter may otherwise route individual requests through
# providers that silently omit requested logprobs.
provider = dict(extra_body.get("provider", {}) or {})
provider["require_parameters"] = True

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

without this, some provider does not return logprobs and hence, these backends can be used during API routing which fails the pipeline

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this intentional?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, arena hard uses some pre-defined CSV file to configure their scoring. Perhaps we can avoid having this .csv but should put these data somehow if we want true replication

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can this be regenerated or is it fixed?

Comment thread judgearena/config.py
)

protocol = resolved_task.spec.protocol
task_generation = getattr(protocol, "generation", None)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The task defaults added here leave generation.truncate_all_input_chars at the generic 8192. The pinned Arena-Hard data has 2 v0.1 prompts and 15 v2.0 prompts over that limit, up to 27,586 chars, while upstream passes the full question to generation. Could official tasks set this default to None? Otherwise newly generated candidate answers use truncated prompts on those rows.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes this is fair. I will fix it

Comment thread judgearena/benchmarks/pairwise/scoring/arena_hard.py
@kargibora
kargibora force-pushed the feat/official-task-variants-v2 branch from c1bf395 to 7f148d7 Compare September 10, 2026 09:55
@kargibora

Copy link
Copy Markdown
Member Author

Upstream–JudgeArena comparisons and repeatability

Check Result
Same-annotation parser and metrics Match to numerical precision; Arena-Hard v2 hard scoring remains partial.

1. AlpacaEval

AlpacaEval 2 — Upstream vs JA; baseline: gpt4_1106_preview; length-controlled win rate ± SE

Candidate Judge: upstream → JA Upstream (%) JA (%) Δ (pp) Agreement (%)
GPT-4 GPT-4-1106 → GPT-4 42.90 ± 1.69 31.28 ± 1.37 −11.62 84.50
GPT-4 Gemma, run 1 41.93 ± 1.03 41.79 ± 1.04 −0.14 99.45
GPT-4 Gemma, run 2 45.30 ± 1.00 40.94 ± 1.02 −4.36 98.36
Llama 3.1-70B Instruct (Turbo) Gemma 38.85 ± 0.99 40.09 ± 0.92 +1.24 93.50

AlpacaEval 2 — Gemma repeatability; candidate: GPT-4; baseline: gpt4_1106_preview; winner/tie agreement

Pipeline Run 1 (%) Run 2 (%) Δ₂₋₁ (pp) Agreement (%)
JA 41.79 ± 1.04 40.94 ± 1.02 −0.85 98.36
Upstream 41.93 ± 1.03 45.30 ± 1.00 +3.37 99.45

2. MT-Bench

Upstream vs JA — candidate: GPT-4; baseline: gpt-3.5-turbo; paired win rate; per-game verdict agreement

Judge: upstream → JA Upstream (%) JA (%) Δ (pp) Agreement (%)
GPT-4 (released → new) 82.50 84.06 +1.56 91.25
Gemma, run 1 78.75 75.63 −3.13 92.50
Gemma, run 2 77.19 77.81 +0.63 93.75

Additional GPT-4.1 pilot — annotation agreement only; score not computed

Candidate Baseline Judge (both pipelines) Panel Verdict agreement (%)
gpt-4 gpt-3.5-turbo GPT-4.1 10 questions 90.00

Gemma repeatability — candidate: GPT-4; baseline: gpt-3.5-turbo; paired win rate; per-game verdict agreement

Pipeline Run 1 (%) Run 2 (%) Δ₂₋₁ (pp) Agreement (%)
JA 75.63 77.81 +2.19 94.38
Upstream 78.75 77.19 −1.56 93.13

3. Arena-Hard v0.1

Upstream vs JA — baseline: gpt-4-0314; decisive-weighted win rate [95% CI]

Candidate Judge: upstream → JA Upstream (%) JA (%) Δ (pp) Exact / winner-tie (%)
GPT-4-0613 GPT-4-1106 → GPT-4 39.18 [35.50, 42.55] 32.28 [28.80, 35.01] −6.90 51.60 / 63.83
GPT-4-0613 Gemma, run 1 40.00 [36.89, 43.61] 37.59 [33.79, 41.46] −2.41 75.88 / 89.16
GPT-4-0613 Gemma, run 2 40.19 [37.36, 43.95] 38.32 [35.16, 41.82] −1.88 80.22 / 89.97
GPT-3.5-turbo-0125 Gemma 21.13 [18.86, 23.47] 20.58 [18.24, 24.08] −0.55 77.64 / 89.20

Additional GPT-4.1 pilot — annotation agreement only; score not computed

Candidate Baseline Judge (both pipelines) Panel Exact (%) Winner/tie (%)
gpt-4-0613 gpt-4-0314 GPT-4.1 10 questions 70.00 85.00

Gemma repeatability — candidate: GPT-4-0613; baseline: gpt-4-0314; decisive-weighted win rate [95% CI]

Pipeline Run 1 (%) Run 2 (%) Δ₂₋₁ (pp) Exact (%) Winner/tie (%)
JA 37.59 [33.79, 41.46] 38.32 [35.16, 41.82] +0.72 76.42 88.62
Upstream 40.00 [36.89, 43.61] 40.19 [37.36, 43.95] +0.19 76.96 87.80

4. Arena-Hard v2

Upstream vs JA — hard baseline: o3-mini-2025-01-31; creative baseline: gemini-2.0-flash-001; scores [90% CI]

Variant Candidate Judge (both pipelines) Upstream (%) JA (%) Δ (pp)
Hard DeepSeek-R1 GPT-4.1 (100-question pilot) 52.45 [47.54, 56.20] 51.87 [47.60, 56.30] −0.58
Creative DeepSeek-R1 GPT-4.1 (10-question pilot) 97.20 [92.50, 100.00] 97.20 [92.50, 100.00] 0.00
Creative DeepSeek-R1 Gemma 73.06 [70.51, 75.38] 72.12 [69.63, 74.27] −0.94
Creative Qwen3-32B Gemma, run 1 39.20 [36.59, 41.74] 39.99 [37.55, 42.68] +0.79
Creative Qwen3-32B Gemma, run 2 41.15 [38.54, 43.44] 39.62 [37.52, 41.78] −1.53

Upstream vs JA — annotation agreement; same baselines as the score table

Variant Candidate Judge (both pipelines) Exact (%) Winner/tie (%)
Hard DeepSeek-R1 GPT-4.1 (100-question pilot) 69.50 78.00
Hard DeepSeek-R1 GPT-4.1 (10-question pilot) 70.00 70.00
Hard DeepSeek-R1 Gemma 73.25 87.00
Hard Qwen3-32B Gemma, run 1 78.93 88.83
Hard Qwen3-32B Gemma, run 2 75.13 85.53
Hard GPT-4.1-mini Gemma 77.25 88.50
Creative DeepSeek-R1 GPT-4.1 (10-question pilot) 60.00 100.00
Creative DeepSeek-R1 Gemma 78.89 91.71
Creative Qwen3-32B Gemma, run 1 81.31 91.41
Creative Qwen3-32B Gemma, run 2 78.79 88.89

Creative Gemma repeatability — candidate: Qwen3-32B; baseline: gemini-2.0-flash-001; bootstrap-mean win rate [90% CI]

Pipeline Run 1 (%) Run 2 (%) Δ₂₋₁ (pp) Exact (%) Winner/tie (%)
JA 39.99 [37.55, 42.68] 39.62 [37.52, 41.78] −0.36 80.05 90.66
Upstream 39.20 [36.59, 41.74] 41.15 [38.54, 43.44] +1.95 82.58 90.66

Hard Gemma repeatability — candidate: Qwen3-32B; baseline: o3-mini-2025-01-31; annotation agreement only

Comparison Exact agreement (%) Winner/tie agreement (%)
JA 1 vs JA 2 77.16 90.36
Upstream 1 vs Upstream 2 78.17 87.56

Use the native fixed-step L-BFGS procedure without a PyTorch dependency. Preserve source attribution and acknowledge remaining float32 numerical differences. Return unavailable preferences per row when AlpacaEval logprobs are missing, rather than aborting the batch.
@kargibora
kargibora merged commit 677212a into main Sep 15, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants