Prove a task can pass before buying it - #1
Conversation
Every control that existed fired after the spend: a meter, a budget, an abort. This is the half that comes before. Baseline stops being a flag. `run --baseline` was already shipped and had never once been used here, which is what a flag people are supposed to remember is worth. Every `run` now baselines itself first, and refuses the dispatch on the two shapes that make a task unbuyable: a check that is already green (green now, green at the end, so it can never tell you the work happened) and a check that could not be executed at all. A check that FAILS baseline is the wanted result and dispatches normally -- that is the whole point of the phase, and refusing on it would refuse every honest manifest. The canary is a stop, not a smaller batch. A multi-task run releases its first task alone, judges it by its own executed check, and only then releases the rest; a bad verdict marks every remaining task SKIPPED without spawning. Measured both ways on the same night: a 3-ticket run that did this caught a design fault on its first task, and a 32-task run that did not lost a whole round to a fault its first task had already demonstrated. --canary-confirm adds a human on top, deliberately not by default -- a pause met on every run becomes a keypress people learn to hit. Both gates take a REASON to skip, not a bare flag, and print WAIVED (not proved, not verified). A blank reason is rejected rather than quietly re-enabling the gate the operator believed they had turned off. The whole verdict lands in the run record as a `preflight` block, so "was this checked?" is answerable later instead of remembered. Two existing tests now waive baseline explicitly: the budget e2e checks `exit 0` on purpose, and the workdir-escape test needs to reach the runtime guard the gate would otherwise pre-empt. Both keep testing what they are named for, and both refusals are pinned independently in the new file. Every new test was run against the unmodified tree first: all nine fail there. Demo keeps its parallel fan-out -- the canary is skipped for it with the reason recorded, since serialising the first of three workers would hide the thing the demo exists to show. Its baseline passes cleanly (3 fail, 0 pass, 0 error). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KX9rgZWnE5zAmJmhaBz35S
…deliverable a real refusal Two findings from the lens review, both about this change. The record answered "was a canary configured?" but not "was it judged, and what did it say?" -- and only the second question tells you whether the batch was released on evidence. The verdict now lands in the run record at all four points where the canary is actually decided: released, held, waived, and the single-task auto-skip. It lives on the state writer rather than in Preflight, which stays frozen and describes only what was decided BEFORE dispatch. The unreachable deliverable had no demonstrated failure case. The test provoked the `error` outcome through an escaping task key, which is a different fault -- the ticket's named case is a deliverable no sandboxed worker can write, and nothing detected that at all. Baseline now probes declared absolute `expect_files` against the nearest existing ancestor and refuses when nothing could be created there. Only absolute paths are probed, on purpose: a relative deliverable lands in the scratch dir the harness makes and is always writable, so probing those would refuse nearly every honest manifest. That complement is pinned too. The third finding, a failure-counter docstring contradicting its implementation, is declined as out of scope: `failure_counts` is untouched by this change and belongs to the cost-control work. It arrived because the review kit staged the wrong diff -- see the PR thread. All three new assertions were run against the previous commit first and fail there. Suite: 327 pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KX9rgZWnE5zAmJmhaBz35S
Review pass 1 — and a defect in the review kit itselfPass 1 returned three findings. Two were real and are fixed in The staging defect
git diff "$(git merge-base origin/main "$HEAD_SHA")"..."$HEAD_SHA"It hard-codes Two of the three "Lens" items were consequently about code in NateBJones-Projects#129, not this PR ( It fails silently. This is not specific to a fork — any stacked PR in any repo (base = a feature branch rather than main) gets the same contamination. The three findings
review-accepted: Failure-counter docstring contradicts its implementation — the symbol does not appear in this PR's commits; it entered the review through the mis-staged diff described above and belongs to NateBJones-Projects#129. All three new assertions were run against the pre-fix commit first and fail there. Suite: 327 pass, 1 pre-existing skip. |
The lens is right that `unwritable_deliverables` cannot establish that a WORKER can write a path: it probes as the dispatcher, and a sandbox can deny what `os.access` here calls writable. Widening the code is not available -- proving it needs a spawned worker, and spawning nothing is exactly what makes baseline free enough to run before every dispatch. So the claim is narrowed to what the code does. Naming the layers while correcting it, because the boundary is the useful part: lint's `worker_unwritable_paths` reasons about sandbox SCOPE and is the one that matches the measured incident; baseline catches what no process could write at all; the canary buys whatever neither could know statically, once rather than once per task. The docstring, the error text and the README now each say which of the three they are. Suite: 327 pass, unchanged -- this narrows claims, not behavior. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KX9rgZWnE5zAmJmhaBz35S
…pawning nothing The docstring promised "Spawn nothing" and the function does spawn things -- every check is a subprocess, and the worktree path launches git helpers. The guarantee that actually matters, and the one the refusal depends on, is that no worker starts: no model, no billable token. Claiming more invites a maintainer to assume there are no side effects at all, when a check can legitimately export files, which the fix-swarm pattern relies on. Inherited wording, but this function is rewritten here, so it is corrected here. Suite: 327 pass, unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KX9rgZWnE5zAmJmhaBz35S
The demo disabled the canary by command name. Nobody asked for it, and unlike a real waiver it printed nothing -- it was recorded in the run record and silent on the terminal, which is the difference between a decision and a default nobody can see. A gate with a third path that the tool takes on your behalf is the shape this whole change exists to remove. The justification was that the demo shows parallel fan-out and the canary serialises the first of its three workers. That is true and it is not worth a special case: a demo of a path real runs never take teaches the wrong behaviour, and the canary is now what a real run does. The demo still passes end to end -- its baseline is clean (3 fail, 0 pass, 0 error) and its first task writes the file its own check demands. If the three-at-once visual is wanted back it is one flag away, and that flag announces itself. Which is the design. Suite: 327 pass, unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KX9rgZWnE5zAmJmhaBz35S
Refs bryfa/mulhacenlabs-work#1031
Every expensive thing measured on 2026-09-10 had the same shape: work was dispatched before anyone had proved it could succeed. Not five problems — one problem, five times. The controls that already exist (meter, budget, abort) all fire after the spend. This is the half that comes before.
Base is
fix/configurable-check-timeout, the checkout every factory runs, so this stacks on the cost controls rather than conflicting with them. An upstream PR gets cut separately once NateBJones-Projects#74 and NateBJones-Projects#129 land.1. Baseline stops being a flag
run --baselineshipped already and had never once been used here — which is what a flag people are supposed to remember is worth. Everyrunnow baselines itself first.A run is refused before any worker spawns (exit 2) on the two shapes that make a task unbuyable:
Refusing on a failing baseline would refuse every honest manifest, so it does not.
expect_filesas the dispatcher sees them, so it catches paths the filesystem itself refuses; a sandbox can still deny a path that looks writable from here. Three layers divide that ground:worker_unwritable_paths, already shipped in NateBJones-Projects#129)2. The canary is a stop, not a smaller batch
A multi-task run releases its first task alone, judges it by its own executed check, and only then releases the rest. A bad verdict marks every remaining task
SKIPPEDwithout spawning.Measured both ways on the same night: a 3-ticket run that did this caught a design fault on its first task; a 32-task run that did not lost a whole round to a fault its first task had already demonstrated.
--canary-confirmadds a human on top — deliberately not default. A pause met on every run becomes a keypress people learn to hit, which buys the appearance of a gate and none of the substance. On a non-TTY it stops rather than fabricating approval.3. Escapes that announce themselves
Each takes a reason, prints
WAIVED (not proved, not verified), and records it. A blank reason is rejected rather than silently re-enabling the gate the operator believed they had turned off.Every run record carries a
preflightblock with the per-task baseline verdict and the canary's verdict — released, held, waived, or auto-skipped — so "was this checked, and what did it say?" is answerable later rather than remembered.Tests
tests/test_prove_before_buy.py— 12 tests, each refusal provoked on purpose: a check that cannot fail, a check that cannot be executed, a deliverable the filesystem refuses, a bad canary verdict holding the batch, a good canary releasing it, the single-task auto-skip, both waivers, a blank reason, and the record.Every new test was run against the tree without the fix first, and fails there. A test that has never failed proves nothing.
Full suite: 327 pass (was 315), 1 pre-existing skip.
Two existing tests now waive baseline explicitly rather than being weakened — the budget e2e checks
exit 0on purpose, and the workdir-escape test needs to reach the runtime guard the gate would otherwise pre-empt.demokeeps its parallel fan-out: the canary is skipped for it with the reason recorded, since serialising the first of three workers would hide the thing the demo exists to show. Its baseline passes cleanly (3 fail, 0 pass, 0 error), verified against the shipped manifest.Review
Three lens passes. Fixed: the canary verdict was missing from the run record; the unreachable deliverable had no detection at all, not merely no test;
unwritable_deliverablesoverstated what it proves;execute_baseline's docstring said "spawn nothing" when it spawns checks and git helpers (it spawns no workers).review-accepted: Baseline deliverable probing does not prove worker-sandbox writeability — correct, and the claim is now narrowed everywhere to what the code does. The suggested fix, probing inside the worker sandbox, would require baseline to spawn a worker; spawning none is what makes it cheap enough to run before every dispatch, and is the guarantee the refusal itself depends on. Sandbox scope is already reasoned about statically by lint's
worker_unwritable_paths(NateBJones-Projects#129), and what neither can know statically is what the canary is for — one task's spend instead of the manifest's.review-accepted: Failure-counter docstring contradicts its implementation —
failure_countshas zero occurrences in this PR's commits; it belongs to NateBJones-Projects#129 and reached the reviewer through a staging defect in the review kit, fixed separately in bryfa/nexo#259 (work#1034).Deliberately not built
No baseline caching across runs. The ticket raised skipping the baseline on a re-run of an unchanged manifest. The tree moves under the manifest, so a cached baseline can be green about code that no longer exists — and the baseline spawns no workers, so it is nearly free. Keeping it cheap beats caching a lie.
🤖 Generated with Claude Code
https://claude.ai/code/session_01KX9rgZWnE5zAmJmhaBz35S