Skip to content

About

Reproducible, privacy-safe model research and certification tooling for Loomarr

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

loomarr-models

CI

Offline research assets for Loomarr-specific model experiments. This repository is deliberately separate from the Go application: Loomarr consumes released model bytes through its existing provider boundary and never imports this toolchain.

The active custom-planner lane is being rebuilt against the current production contract under loomarr-models#22. Historical v3/v4 corpora are preserved but excluded; the replacement no-spend artifacts, disjointness proof, stock-baseline preregistration, and stop conditions are documented in docs/planner-current-contract-v1.md.

The follow-on no-spend baseline harness is tracked in loomarr-models#25 and documented in docs/planner-current-stock-baseline.md. It adds current-contract scoring, preflight, runtime, and publication replay while all execution authority remains disabled.

Independent review of the current-contract training drafts is tracked in loomarr-models#27 and documented in docs/planner-current-review.md. Pending, rejected, stale, or self-reviewed rows cannot enter the active training corpus.

The current-contract QLoRA v2 run is terminal and settled under loomarr-models#20. It stopped before optimizer step 1 because the preregistered renderer sorted JSON keys while the pinned Transformers runtime preserved insertion order. All 24 token counts matched, with a 5,416-token maximum, but the exact rendered-byte identity correctly failed closed. No adapter was produced and evaluation did not run; the exact Runpod charge was $0.06595210370142013. A corrected retry requires a new immutable capacity report, experiment identity, and explicit authorization.

The first milestone is loomarr/loomarr#937: a validated 50-trace planner smoke corpus. Current files establish the fail-closed trace contract, the pinned Qwen 3.8 / Unsloth candidate environment, and the reproducible evidence from the first bounded GPU smoke. Model weights, adapters, caches, and raw logs remain local ignored artifacts; only compact hash-bound evidence is committed.

Checks

make check

validate-corpus accepts reviewed artifacts only. Draft validation is available explicitly for the independent-review workflow and never promotes a draft into training data.

Every corrected trace requires separate Gemini 3.1 Pro and GPT-5.4 attestations through pinned OpenRouter provider routes. The two model families remain blind to each other's output and outside the Qwen candidate family. Once both pass all six criteria for all 50 traces, make finalize-corpus creates the immutable artifact, manifest, and validation report. Any disagreement, rejection, invalid response, route drift, partial run, or unsettled charge remains non-approved and produces a targeted escalation.

The completed v7 review established a 50/50 valid rate for Gemini and an 18/50 valid rate for Sonnet, with 13 unanimous approvals and 37 targeted escalations. Its exact $2.212744 cost and all replayable evidence are staged under reviews/planner-smoke-v1/planner-model-review-v7/; partial results do not alter the canonical pending corpus.

Review v8 re-ran the complete corrected corpus with Gemini 3.1 Pro and GPT-5.4. All 100 calls settled for exactly $2.4979775: Gemini produced 50 valid approvals, while GPT-5.4 produced 42 valid reviews and eight quarantined length completions. The paired evidence independently approves 36 traces and leaves 14 pending. Five ambiguous-mood traces and one conflicting-intent trace exposed two remaining generator defects; the other eight pending traces require replacement attestations for invalid GPT completions. Replayable evidence is staged under reviews/planner-smoke-v1/planner-model-review-v8/, and the canonical corpus remains unchanged.

Review v9 corrects the two remaining generator families and re-reviews the complete hash-bound corpus. Ambiguous-mood candidates now contain explicit tone evidence; conflicting-intent traces use a named title that the same request both requires and excludes, avoiding fixture-only search language. The GPT-5.4 completion ceiling is raised to 4,000 tokens to reduce invalid reasoning-only completions while preserving one call per trace and no automatic inference retry. Its first launch stopped before inference when the pinned OpenAI route became unavailable; v9 now pins the same model and upstream revision through the single healthy openai/flex route. That route then rate-limited its first GPT call after all 50 Gemini reviews had settled, so the partial v9 run stopped and charged only the exact $0.967292 Gemini cost. V10 uses the healthy openai/fast route and adds explicit provider-error-envelope validation.

V10 completed 100/100 valid reviews for exactly $3.999040, independently approving 46 traces and leaving four disagreements. Two rejections misread the deliberately synthetic fixture provenance as assistant-invented content. Two recovery traces exposed real contract gaps: one reused a bare genre as a title query, and one omitted the requested genre from final policy. The canonical corpus remains pending while those reviewer and generator defects are corrected.

V11 applies the recovery correction across all five variants and makes the auditor's fixture semantics explicit. A tool-returned reserved-ID fixture stands in for real catalog content and is not an invented title. The complete corpus will be reviewed again through the same healthy Gemini and GPT-5.4 fast routes before any trace is promoted.

V11 completed with 100 valid attestations, 50 unanimous approvals, zero escalations, and an exact $3.785636 cost. Its replayable publication is the sole input to the fail-closed canonical promotion step; no earlier partial decisions are combined with it.

The candidate NVIDIA environment is resolved with uv 0.12.9 for Linux x86_64, Python 3.12, CUDA 12.8, and PyTorch 2.8. make lock-qwen38-a40 reproduces the hash-bound lock; inside the pinned container, make sync-qwen38-a40 installs it using uv's cu128 package backend. Neither command downloads model weights or starts training.

Issue loomarr/loomarr#938 owns the bounded QLoRA smoke runner. The first pinned A40 run completed all 20 steps against the reviewed-frozen 50-trace corpus; its compact publication is under runs/planner-qwen38-smoke-v1/. The result proves the environment, memory envelope, and adapter-only save path, but does not certify or authorize the adapter for release. See docs/qwen38-qlora-smoke.md for the result, the NVIDIA training lane, and the 64 GB Mac development/evaluation lane.

The leakage-free development comparison in loomarr-models#5 is complete. Across 50 frozen synthetic cases, the adapter improved weighted quality and policy accuracy, but retained 16 hard failures, missed the absolute quality gates, and regressed recovery. Its hash-bound publication is under runs/planner-adapter-eval-v1/; the decision is adapter-rejected-no-release, so no certification run, packaging, serving, or release is authorized. See docs/planner-adapter-eval.md. The exact Runpod charge was $0.7827729525743052.

The exhaustive no-spend follow-up in loomarr-models#7 classifies all 16 hard failures by first divergence. The current 50-trace corpus does not authorize another QLoRA configuration; targeted reviewed traces and a newly frozen disjoint development set must exist first. See docs/planner-adapter-failure-analysis.md.

The no-spend first stage of loomarr-models#9 generates 120 targeted pending training drafts and a separate 60-case development gate across the six observed corrective behaviors. All four planner splits pass pairwise identity and normalized-content leakage checks. The review plan is hash-bound and its compact 240-call envelope passes a $13.906540 worst-case preflight. Paid review is authorized only by its separate reviewed plan; the execution wrapper reconstructs every committed request and refuses route, price, budget, or source drift. Publication, exact settlement, unanimous-only promotion, and corpus freezing are implemented and hash-bound before any paid call. This does not authorize retraining. See docs/planner-behavior-corpus-v2.md.

The completed review settled all 240 calls for exactly $4.080632. Both reviewers approved 118 traces; two disagreements remain pending because the compact reviewer packet did not make the contract's title-query and empty-picks confidence semantics explicit enough. The completed plan is disabled, and the disputed traces cannot enter training until a corrected independent review resolves them.

The no-spend planner-behavior-review-v3 plan corrects the packet and repeats the full 120-trace review with both reviewers; it does not selectively retry the two disagreements. Each request now includes the exact tool declaration and explicit targeted audit semantics, including query title search and per-existing-pick confidence. Its conservative worst case is $15.617740, within a $16.50 reservation and the maintainer-authorized $40.00 aggregate cap. A separately reviewed authorization commit may enable only this exact plan after a fresh route snapshot.

Contributing and security

See CONTRIBUTING.md before opening a change. Report security issues privately as described in SECURITY.md. Product bugs and feature requests belong in the main Loomarr repository.

About

Reproducible, privacy-safe model research and certification tooling for Loomarr

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages