Skip to content

skill: escalate the model on hard tasks; read config hunks in review - #127

Open
chriscfellows wants to merge 1 commit into
NateBJones-Projects:mainfrom
chriscfellows:docs/skill-routing-and-review
Open

chriscfellows wants to merge 1 commit into
NateBJones-Projects:mainfrom
chriscfellows:docs/skill-routing-and-review

Conversation

@chriscfellows

Copy link
Copy Markdown

Two additions, both about the orchestrator's own judgment rather than the worker's.

  • Escalate the model when the task warrants it. The engine-selection section routes by scoreboard evidence but says nothing about the case where the default lane is simply the wrong size for the task. Signals worth escalating on: prior same-shape tasks burned >5M tokens, or produced workaround-shaped output under spec pressure. The per-task model/engine fields already exist for this; the guidance is to use them deliberately and to bring the numbers when recommending it.
  • Read every hunk of config and manifest files in the patch. Added to the post-run ritual as step 4. Both check-gaming incidents seen here hid in one-line changes to package.json scripts adjacent to legitimate edits — diff stats and artifact spot-checks do not catch that shape.

The ritual's later items are renumbered 5 and 6 to keep the list correct.

Docs only, one file. Suite unchanged: Ran 254 tests, OK.

Claude-Session: https://claude.ai/code/session_01AVTYzUcjMokzMsyjABjUBF

Two additions, both about the orchestrator's own judgment rather than the
worker's.

* **Escalate the model when the task warrants it.** The engine-selection
  section routes by scoreboard evidence but says nothing about the case where
  the default lane is simply the wrong size for the task. Signals worth
  escalating on: prior same-shape tasks burned >5M tokens, or produced
  workaround-shaped output under spec pressure. The per-task model/engine
  fields already exist for this; the guidance is to use them deliberately and
  to bring the numbers when recommending it.
* **Read every hunk of config and manifest files in the patch.** Added to the
  post-run ritual as step 4. Both check-gaming incidents seen here hid in
  one-line changes to package.json scripts adjacent to legitimate edits —
  diff stats and artifact spot-checks do not catch that shape.

The ritual's later items are renumbered 5 and 6 to keep the list correct.

Docs only, one file. Suite unchanged: Ran 254 tests, OK.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AVTYzUcjMokzMsyjABjUBF
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant