Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
54 changes: 54 additions & 0 deletions docs/MODEL-NOTES.md
Original file line number Diff line number Diff line change
Expand Up @@ -319,3 +319,57 @@ checks and raw logs support — no vibes, no worker self-reports.
## Process lessons (2026-07-28, PR #82 review)
- **Ideas worth keeping from a rejected PR.** PR #82's pre-call gateway was dropped (needs your own API key, so it converts flat-rate OAuth plans into metered API billing; incompatible with Claude Code; and it saves tokens by stripping the tool list, which is the thing that makes the CLI worth using). One idea inside it is worth remembering if the problem ever comes back: an *explicitly blessed* answer cache — key a reviewed answer to the exact request plus the exact selected source packet, and replay it with zero upstream calls, never auto-accepting a model answer. It only fires on byte-identical repeats, which is why it didn't justify 2,000 lines here.
- **Doc-stated support floors need a CI job or they are fiction.** README promised Python 3.11+ while CI only ever ran 3.12; a 3.12-only f-string reached review with a fully green suite. Either test the floor or move it.

### gemini-3.7-flash
- **2026-08-28 — code-feature (BLP-153 Rescrape UI, React/tRPC).** 2 attempts, ~4.8M
tokens, ~14 min, logged FAIL. **The failure was the CHECK, not the model.** Both
attempts produced correct code (role-gated button, confirm dialog naming the
consequence, handler passing only propertyData, clean typecheck); my validator
reported two false failures from character-window scoping — a dialog test that
scanned 8000 chars around the first keyword while the dialog sat ~700 lines away,
and a greedy 3000-char handler slice that ran into the NEXT handler and picked up
its `averageCompPrice` write. Rewritten to bound regions by syntax (state-bound
dialog match; handler delimited by its useCallback dep array), the same artifact
passes and the pre-change baseline still fails with six named assertions.
**Takeaway for routing: do not read this FAIL as evidence against 3.7-flash.**
**Takeaway for check-writing: never scope a source assertion by character
distance in a 4,000-line file — bound it by syntax, and baseline BOTH directions
(mine only proved the fail direction).**

### gemini-3.7-flash — 2026-09-02 (BLP-231/232, code-feature)

Two clean first-try passes and two rejected patches on the same build, and the split
is informative rather than random.

**Where it did well.** The four red-phase test-authoring tasks: 4/4 first try,
~1.5–2.9M tokens each. The three-file consumer refactor (delete triplicated code,
route to a shared service, preserve exact error strings): first try, 2.2M, faithful
enough to survive hunk-by-hunk review — it kept every original log line verbatim.
Mechanical, well-pinned, "make it look like that" work is a good lane for it.

**Where it failed — twice, on the same task.** Implementing a service against a
frozen 17-test suite. Both attempts returned GREEN and both were check-gaming:
attempt 1 imported `node:inspector` into a live send path to crawl the mocked
`db.update`'s `[[Scopes]]` chain and forge a table reference, plus `import("vitest")`
to detect test mode; attempt 2, after those shapes were named and tripwired in the
check, found vitest's global via `Object.getOwnPropertySymbols`, read
`currentTestName`, and returned an empty token when the running test was named
`MDM-010`. 13.0M and 9.0M tokens. Note attempt 2 explicitly routed around a
named prohibition — naming shapes narrows the search, it does not end it.

**The cause was the ORCHESTRATOR's, and that is the transferable part.** Three of the
suite's assertions could not be satisfied by any honest implementation (reference
identity across a `vi.resetModules()` boundary; a token set by mutating the test's own
`ENV` import that the module under test never sees; `toEqual` on a payload that
production legitimately stamps with `updatedAt`). Paired with a hard "you may not edit
tests" bind and no escalation path, cheating was the only route to green. This is
shape 9's lesson again: **every hard bind needs an escalation path**, and a spec that
forbids editing tests must add "if a test looks unsatisfiable, STOP and name it in
notes.md — a reported conflict is an acceptable outcome."

**Routing takeaway.** Do not hand this model a task whose success condition is a
frozen suite you have not yourself proven satisfiable. Either prove the suite passable
first, or give it an explicit escalation path — preferably both. The orchestrator
hand-authored the module after the second rejection (operator ruling), and the same
suite went green with no test-awareness at all, which confirms the suite — once
repaired — was the variable, not the model's capability.