Offline research assets for Loomarr-specific model experiments. This repository is deliberately separate from the Go application: Loomarr consumes released model bytes through its existing provider boundary and never imports this toolchain.
The active custom-planner lane is being rebuilt against the current production contract under loomarr-models#22. Historical v3/v4 corpora are preserved but excluded; the replacement no-spend artifacts, disjointness proof, stock-baseline preregistration, and stop conditions are documented in docs/planner-current-contract-v1.md.
The follow-on no-spend baseline harness is tracked in loomarr-models#25 and documented in docs/planner-current-stock-baseline.md. It adds current-contract scoring, preflight, runtime, and publication replay while all execution authority remains disabled.
Independent review of the current-contract training drafts is tracked in loomarr-models#27 and documented in docs/planner-current-review.md. Pending, rejected, stale, or self-reviewed rows cannot enter the active training corpus.
The current-contract QLoRA v2 run is terminal and settled under
loomarr-models#20. It stopped before optimizer
step 1 because the preregistered renderer sorted JSON keys while the pinned Transformers runtime
preserved insertion order. All 24 token counts matched, with a 5,416-token maximum, but the exact
rendered-byte identity correctly failed closed. No adapter was produced and evaluation did not run;
the exact Runpod charge was $0.06595210370142013. A corrected retry requires a new immutable
capacity report, experiment identity, and explicit authorization.
The first milestone is loomarr/loomarr#937: a validated 50-trace planner smoke corpus. Current files establish the fail-closed trace contract, the pinned Qwen 3.8 / Unsloth candidate environment, and the reproducible evidence from the first bounded GPU smoke. Model weights, adapters, caches, and raw logs remain local ignored artifacts; only compact hash-bound evidence is committed.
make checkvalidate-corpus accepts reviewed artifacts only. Draft validation is available explicitly for the
independent-review workflow and never promotes a draft into training data.
Every corrected trace requires separate Gemini 3.1 Pro and GPT-5.4 attestations through pinned OpenRouter
provider routes. The two model families remain blind to each other's output and outside the Qwen
candidate family. Once both pass all six criteria for all 50 traces, make finalize-corpus creates the
immutable artifact, manifest, and validation report. Any disagreement, rejection, invalid response,
route drift, partial run, or unsettled charge remains non-approved and produces a targeted escalation.
The completed v7 review established a 50/50 valid rate for Gemini and an 18/50 valid rate for Sonnet,
with 13 unanimous approvals and 37 targeted escalations. Its exact $2.212744 cost and all replayable
evidence are staged under reviews/planner-smoke-v1/planner-model-review-v7/; partial results do not alter
the canonical pending corpus.
Review v8 re-ran the complete corrected corpus with Gemini 3.1 Pro and GPT-5.4. All 100 calls settled
for exactly $2.4979775: Gemini produced 50 valid approvals, while GPT-5.4 produced 42 valid reviews
and eight quarantined length completions. The paired evidence independently approves 36 traces and
leaves 14 pending. Five ambiguous-mood traces and one conflicting-intent trace exposed two remaining
generator defects; the other eight pending traces require replacement attestations for invalid GPT
completions. Replayable evidence is staged under
reviews/planner-smoke-v1/planner-model-review-v8/, and the canonical corpus remains unchanged.
Review v9 corrects the two remaining generator families and re-reviews the complete hash-bound corpus.
Ambiguous-mood candidates now contain explicit tone evidence; conflicting-intent traces use a named
title that the same request both requires and excludes, avoiding fixture-only search language. The
GPT-5.4 completion ceiling is raised to 4,000 tokens to reduce invalid reasoning-only completions while
preserving one call per trace and no automatic inference retry. Its first launch stopped before inference
when the pinned OpenAI route became unavailable; v9 now pins the same model and upstream revision through
the single healthy openai/flex route. That route then rate-limited its first GPT call after all 50
Gemini reviews had settled, so the partial v9 run stopped and charged only the exact $0.967292 Gemini
cost. V10 uses the healthy openai/fast route and adds explicit provider-error-envelope validation.
V10 completed 100/100 valid reviews for exactly $3.999040, independently approving 46 traces and
leaving four disagreements. Two rejections misread the deliberately synthetic fixture provenance as
assistant-invented content. Two recovery traces exposed real contract gaps: one reused a bare genre as
a title query, and one omitted the requested genre from final policy. The canonical corpus remains
pending while those reviewer and generator defects are corrected.
V11 applies the recovery correction across all five variants and makes the auditor's fixture semantics explicit. A tool-returned reserved-ID fixture stands in for real catalog content and is not an invented title. The complete corpus will be reviewed again through the same healthy Gemini and GPT-5.4 fast routes before any trace is promoted.
V11 completed with 100 valid attestations, 50 unanimous approvals, zero escalations, and an exact
$3.785636 cost. Its replayable publication is the sole input to the fail-closed canonical promotion
step; no earlier partial decisions are combined with it.
The candidate NVIDIA environment is resolved with uv 0.12.9 for Linux x86_64, Python 3.12, CUDA
12.8, and PyTorch 2.8. make lock-qwen38-a40 reproduces the hash-bound lock; inside the pinned
container, make sync-qwen38-a40 installs it using uv's cu128 package backend. Neither command
downloads model weights or starts training.
Issue loomarr/loomarr#938 owns the bounded QLoRA smoke
runner. The first pinned A40 run completed all 20 steps against the reviewed-frozen 50-trace corpus;
its compact publication is under runs/planner-qwen38-smoke-v1/. The result proves the environment,
memory envelope, and adapter-only save path, but does not certify or authorize the adapter for release.
See docs/qwen38-qlora-smoke.md for the result, the NVIDIA training lane,
and the 64 GB Mac development/evaluation lane.
The leakage-free development comparison in
loomarr-models#5 is complete. Across 50 frozen
synthetic cases, the adapter improved weighted quality and policy accuracy, but retained 16 hard
failures, missed the absolute quality gates, and regressed recovery. Its hash-bound publication is
under runs/planner-adapter-eval-v1/; the decision is adapter-rejected-no-release, so no
certification run, packaging, serving, or release is authorized. See
docs/planner-adapter-eval.md. The exact Runpod charge was
$0.7827729525743052.
The exhaustive no-spend follow-up in loomarr-models#7 classifies all 16 hard failures by first divergence. The current 50-trace corpus does not authorize another QLoRA configuration; targeted reviewed traces and a newly frozen disjoint development set must exist first. See docs/planner-adapter-failure-analysis.md.
The no-spend first stage of
loomarr-models#9 generates 120 targeted pending
training drafts and a separate 60-case development gate across the six observed corrective behaviors.
All four planner splits pass pairwise identity and normalized-content leakage checks. The review plan
is hash-bound and its compact 240-call envelope passes a $13.906540 worst-case preflight. Paid review
is authorized only by its separate reviewed plan; the execution wrapper reconstructs every committed
request and refuses route, price, budget, or source drift. Publication, exact settlement, unanimous-only
promotion, and corpus freezing are implemented and hash-bound before any paid call. This does not
authorize retraining. See
docs/planner-behavior-corpus-v2.md.
The completed review settled all 240 calls for exactly $4.080632. Both reviewers approved 118
traces; two disagreements remain pending because the compact reviewer packet did not make the
contract's title-query and empty-picks confidence semantics explicit enough. The completed plan is
disabled, and the disputed traces cannot enter training until a corrected independent review resolves
them.
The no-spend planner-behavior-review-v3 plan corrects the packet and repeats the full 120-trace review
with both reviewers; it does not selectively retry the two disagreements. Each request now includes
the exact tool declaration and explicit targeted audit semantics, including query title search and
per-existing-pick confidence. Its conservative worst case is $15.617740, within a $16.50
reservation and the maintainer-authorized $40.00 aggregate cap. A separately reviewed authorization
commit may enable only this exact plan after a fresh route snapshot.
See CONTRIBUTING.md before opening a change. Report security issues privately as described in SECURITY.md. Product bugs and feature requests belong in the main Loomarr repository.