Conversation
Adds the claude/claude-sonnet engine identity, model registry entries,
and a claude-code-json result parser for the `claude -p --output-format
json` envelope.
The parser distinguishes transport/API failures (rate limits, timeouts —
excluded from the scoreboard as non-scoreable "infrastructure", no
verify/retry) from agent-level failures (max turns, tool/permission
rejection — still scoreable, so a real quality loss isn't hidden from
the scoreboard). It reads the JSON result structurally (whole-payload
decode, one JSON object per line) instead of scanning for a bare `{`,
so text an agent quotes inside its own result can't be mistaken for the
envelope. Model attribution from `modelUsage` only returns a family
match when every candidate of that family agrees on the same canonical
model, rather than guessing the first one when a director and a
same-family subagent both appear.
Adds a schema v3->v4 migration (scoreable, failure_kind columns) with a
backward-compatible ALTER TABLE.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
claude/claude-sonnetengine identity and model registry entries (claude-sonnet-5,claude-opus-5,claude-fable-5-1,claude-haiku-4-5, aliassonnet) so Claude Code CLI can run as a worker engine alongside Codex, Grok Build, and OpenCode/OpenRouter.result_parser = "claude-code-json"andparse_claude_code_json_result()to parse theclaude -p --output-format jsonresult envelope: real token usage fromusage, and the effective model frommodelUsage(Claude Code can resolve an alias likesonnetto a dated model id).scoreable=False,failure_kind="infrastructure", excluded from the scoreboard, no verify/retry) from agent-level failures (max turns, tool/permission rejection — stillscoreable=True, so a genuine quality loss isn't hidden from the scoreboard). An earlier version of this PR classified everyis_erroras infrastructure; fixed after a Codex read-only audit caught it.{. The earlier bracket-scan could match a{an agent happened to quote inside its ownresulttext and misattribute tokens/model/error to that embedded fragment. Also fixed after the same audit._claude_reported_modelonly returns a family match (sonnet/opus/fable/haiku) when every same-family candidate inmodelUsageagrees on the same canonical model; otherwise it returnsNoneinstead of guessing the first one. Guards against misattributing identity when a director and a same-family subagent both appear (e.g. two different Sonnet versions).scoreable,failure_kindcolumns) with a backward-compatibleALTER TABLE.Motivation
I wanted to run real production tasks through Ringer using Claude Code CLI as the worker, and needed the scoreboard to stay honest about why a run failed — an API rate limit isn't the model's fault, but Ringer's existing token/model regex parsers had no way to tell an infra failure apart from a real one for a JSON-emitting engine.
Test plan
python3 -m pytest -q— 278/278 passed (was 275 before this branch's own fixes)tests/test_claude_engine.py: JSON result selection (including the embedded-JSON-in-string case), infra vs. agent error classification, ambiguous same-family model attribution, schema migration🤖 Generated with Claude Code