Official Task Variations (3/3): Add Official AlpacaEval, Arena-Hard, and MT-Bench Tasks - #115
Conversation
| from judgearena.prompts.registry import DEFAULT_JUDGE_PROMPT_PRESET | ||
| from judgearena.utils import strip_thinking_tags | ||
|
|
||
| FASTCHAT_TEMPERATURE_CONFIG: dict[str, float] = { |
There was a problem hiding this comment.
Moved it into the task definition
| self, | ||
| model: str, | ||
| *, | ||
| instruction_ids: pd.Index | None = None, |
There was a problem hiding this comment.
We need instruction_ids because some official protocol asks for filtered instructions and groups to judge different prompts with different configurations
| n_no_logprobs, | ||
| len(ann_list), | ||
| label, | ||
| raise ValueError( |
There was a problem hiding this comment.
Should we skip or give warning or raise a ValueError here?
| # OpenRouter may otherwise route individual requests through | ||
| # providers that silently omit requested logprobs. | ||
| provider = dict(extra_body.get("provider", {}) or {}) | ||
| provider["require_parameters"] = True |
There was a problem hiding this comment.
without this, some provider does not return logprobs and hence, these backends can be used during API routing which fails the pipeline
7cc7660 to
c7a1374
Compare
c7a1374 to
dc8107b
Compare
There was a problem hiding this comment.
Is this intentional?
There was a problem hiding this comment.
Yes, arena hard uses some pre-defined CSV file to configure their scoring. Perhaps we can avoid having this .csv but should put these data somehow if we want true replication
There was a problem hiding this comment.
Can this be regenerated or is it fixed?
| ) | ||
|
|
||
| protocol = resolved_task.spec.protocol | ||
| task_generation = getattr(protocol, "generation", None) |
There was a problem hiding this comment.
The task defaults added here leave generation.truncate_all_input_chars at the generic 8192. The pinned Arena-Hard data has 2 v0.1 prompts and 15 v2.0 prompts over that limit, up to 27,586 chars, while upstream passes the full question to generation. Could official tasks set this default to None? Otherwise newly generated candidate answers use truncated prompts on those rows.
There was a problem hiding this comment.
Yes this is fair. I will fix it
c1bf395 to
7f148d7
Compare
Upstream–JudgeArena comparisons and repeatability
1. AlpacaEvalAlpacaEval 2 — Upstream vs JA; baseline: gpt4_1106_preview; length-controlled win rate ± SE
AlpacaEval 2 — Gemma repeatability; candidate: GPT-4; baseline: gpt4_1106_preview; winner/tie agreement
2. MT-BenchUpstream vs JA — candidate: GPT-4; baseline: gpt-3.5-turbo; paired win rate; per-game verdict agreement
Additional GPT-4.1 pilot — annotation agreement only; score not computed
Gemma repeatability — candidate: GPT-4; baseline: gpt-3.5-turbo; paired win rate; per-game verdict agreement
3. Arena-Hard v0.1Upstream vs JA — baseline: gpt-4-0314; decisive-weighted win rate [95% CI]
Additional GPT-4.1 pilot — annotation agreement only; score not computed
Gemma repeatability — candidate: GPT-4-0613; baseline: gpt-4-0314; decisive-weighted win rate [95% CI]
4. Arena-Hard v2Upstream vs JA — hard baseline: o3-mini-2025-01-31; creative baseline: gemini-2.0-flash-001; scores [90% CI]
Upstream vs JA — annotation agreement; same baselines as the score table
Creative Gemma repeatability — candidate: Qwen3-32B; baseline: gemini-2.0-flash-001; bootstrap-mean win rate [90% CI]
Hard Gemma repeatability — candidate: Qwen3-32B; baseline: o3-mini-2025-01-31; annotation agreement only
|
Use the native fixed-step L-BFGS procedure without a PyTorch dependency. Preserve source attribution and acknowledge remaining float32 numerical differences. Return unavailable preferences per row when AlpacaEval logprobs are missing, rather than aborting the batch.
# Conflicts: # tests/test_config.py
This PR is stacked on #113. It adds the official AlpacaEval, Arena-Hard, and MT-Bench task protocols.
Problem
The existing task names used JudgeArena defaults rather than the original benchmark protocols. This made it difficult to reproduce official results or compare a JudgeArena run with the upstream benchmark.
Canonical tasks
The canonical task names now use the official defaults:
alpaca-evalgpt4_1106_previewbaseline, official prompt, random answer order, first-token logprobs, and official length-controlled scoringarena-hard-v0.1arena-hard-v2.0mt-benchgpt-3.5-turbobaseline, FastChat pairwise prompt, and both answer ordersNormal runtime options can still override model and generation settings.
JudgeArena variants
The existing JudgeArena protocols remain available under explicit names:
alpaca-eval-jaarena-hard-v0.1-jaarena-hard-v2.0-jaThese are separate protocols rather than aliases for the canonical tasks.
Implementation
The Arena-Hard sources contain:
Parsed judge output
The official parsers use the structured parser result introduced in #108:
morMlabel, and the relevant raw token logprobs.A>BorA>>B.Metrics continue to receive the canonical numeric
prefcolumn.Protocol constraints
AlpacaEval requires first-token top logprobs. Responses without them remain unparsed; textual labels are not used as a fallback. If every preference is missing, the metric is saved as unavailable.
Arena-Hard v2 style control is judge-specific. The packaged calibration supports
gpt-4.1andgemini-2.5with their required settings. Other judges need their own calibration data.arena-hard-v2.0-jaremains available for pairwise evaluation with arbitrary judges.Tests
The tests cover task configuration, dataset normalization, prompt and parser behavior, official numerical results, Arena-Hard calibration, MT-Bench integration, OpenRouter capability routing, and package contents.
Result:
359 passedlocally.Raw diff
Relative to #113 at
6a42472. Parent changes are excluded. Pre-existing tests are restored; remaining test reductions apply only to stack-added coverage.Estimated production changes