Repository navigation
Conditionally train and evaluate one current-contract planner adapter #20
Description
Activity
The no-spend QLoRA v2 preregistration is now published in draft PR #21: #21
Exact-head CI passed: https://github.com/loomarr/loomarr-models/actions/runs/33820386753/job/100861595512
The plan binds Qwen3.8-27B, the frozen 120-trace corpus, the A40 environment, 45-step adapter-only recipe, and the
$1.50reservation. Paid training and paid evaluation remain disabled. The next state change requires a separate exact authorization commit.Read-only Runpod readiness check at 2026-09-04T00:00Z (no pod created):
NVIDIA A40/ poolAMPERE_48: availabilityHIGH- 48 GB VRAM; CUDA 12.8 currently available
- community
$0.35/hour; secure$0.49/hour - maximum secure GPU charge under the 9,000-second alarm:
$1.225 - account GPU pods before launch:
0
The
$1.50training reservation still covers the live secure-compute ceiling with$0.275headroom for pod disk. Availability and price are temporal and must be rechecked immediately before provisioning. No paid training or evaluation was authorized or started by this check.No-spend Runpod execution-readiness audit is published in PR #21 at commit
1120d57f80d412d4143506517a285f662a999658.The milestone closes the operational gaps between a valid experiment manifest and a safely recoverable GPU run:
- checked-in Runpod lifecycle for the exact authorization commit, secure A40/image digest, same-DC 40 GB network volume, detached execution, artifact retrieval, and teardown;
- post-save reload of the persisted PEFT adapter into a fresh pinned base-model instance plus a deterministic generation probe before a run manifest can exist;
- fail-closed artifact verifier for all 45 steps, source/config/corpus/environment identities, package/runtime envelope, complete file hashes, adapter-only output, and the reload probe;
- six sabotage tests covering digest drift, wrong source commit, incomplete steps, unauthorized state, and merged/base-weight output;
- the verifier is now a clean tracked preflight-critical input.
Current provider audit (no resources created):
- Runpod GPU pods:
0; network volumes:0; - secure A40 availability:
HIGH; price:$0.49/hour; CUDA 12.8 available; runpodctl 2.12.0installed locally;- important provider caveat: v2.12 removed
--stop-after/--terminate-afterbecause the backend accepted but never enforced them (fix(pod): remove --stop-after and --terminate-after runpod/runpodctl#330). The runbook therefore requires uninterrupted supervision and explicit deletion by 9,000 seconds from provider creation. At the live rate, GPU maximum is$1.225; estimated 40 GB container + network storage remains below$0.03, inside the$1.50reservation.
Verification at this commit:
make check— 117 tests pass;make preflight-qlora-v2-plan— exact 120/120 reviewed corpus, projected aggregate$29.4826615675672820 / $40.00;- generated plan check and
git diff --checkpass.
Paid training and evaluation remain disabled. No pod, volume, model download, inference, or spend was started by this audit. The next state change still requires explicit maintainer authorization for one
$1.50training run from this exact reviewed commit, with no automatic paid retry.Exact-head CI is green for
1120d57f80d412d4143506517a285f662a999658: https://github.com/loomarr/loomarr-models/actions/runs/33824479285/job/100874084985PR #21 remains draft, open, mergeable, and unmerged. The paid gates are unchanged.
Paid execution is now blocked on #22.
Production contract drift is material: loomarr/loomarr#967 advances the planner to suggester-prompt-v4 / catalog-search-v4 and adds independently scored network, cast, and creator routes. This PR's exact commit 1120d57 remains valid no-spend evidence for the older v3 contract, but it is no longer an authorization-ready training source because its 120 traces and 60-case gate cannot teach or detect regressions in the new fields.
Do not provision from 1120d57. #22 owns the no-spend v6 import, leakage check, new reviewed route contrasts, disjoint development cases, stock-v6 baseline, and replacement exact authorization commit. The existing $1.50 ceiling and no-automatic-retry boundary remain unchanged; no reservation is spent or newly committed by this comment.
Acceptance addition from the latest live Loomarr TGIF diagnosis: the targeted adapter should improve the first retrieval turn, not rely on an app-side second chance. For a terse named programming block such as
TGIF, target one valid collection tool call with at least three accurate, diverse exact members inside the normal latency bound. Score real-but-wrong members as failures and a smaller honest set above padding. Include non-TGIF collections/networks/franchises in the held-out gate so this is retrieval-strategy learning, not acronym memorization. See corpus issue #9 for the full observed evidence and split requirements.- changed the title
[-]Train and evaluate one targeted planner QLoRA v2 adapter[/-][+]Conditionally train and evaluate one current-contract planner adapter[/+]on Sep 16, 2026 Current-contract prerequisite update:
- draft PR Rebaseline planner model lane on current production contract #24 freezes Loomarr
e4193f1c385128ccc01837315beeb0996f724c04and replaces obsolete v3/v4 active data with 24 pending current-contract training drafts plus a disjoint 24-case development gate; - draft PR Add no-spend current planner stock baseline harness #26 adds the fail-closed current-contract stock baseline harness with a
$0reservation and all paid execution disabled; - draft PR Land planner research stack (#3–#28) #28 adds the independent-review ledger, packet, and immutable promotion gate. All 24 review decisions remain pending.
All three PRs have green CI. Training remains blocked on genuine independent draft review, a completed authorized stock-baseline quality miss, the application recovery prerequisite, and a separate aggregate-spend authorization. No paid action occurred.
- draft PR Rebaseline planner model lane on current production contract #24 freezes Loomarr
Current-contract stock prerequisite is now a genuine model-quality miss (issue #29).
Validated evidence:
- exact source commit:
aa1fa71e908240a3eae43f50e9b763332ec8cea5 - exact model revision:
8aa5f05d26b7205477066e1449e0af13f762a299 - 24/24 current-contract development cases completed on one secure A40
- run-manifest SHA-256:
d2e496999a8f97b88b3d821aeadcdeb20cf39895813262ec7b5d129add2ffc7a - complete model-quality result: 32 hard failures; grounded completion / correct tool operation / schema validity
0.375; policy accuracy / recovery0.0 - decision:
qloraJustified=true,trainingAuthorized=false - pod deleted; zero active Runpod pods verified
This satisfies the stock-result quality predicate for preparing a replacement current-contract QLoRA authorization package. It does not authorize training. The package must use only the independently approved 24-row
corpus/planner-current-v1/traces.jsonl, bind the current v5 contract and hash-only holdout denylist, keep the 24 development cases out of training, preserve one configuration/no retry/adapter-only output, and bind the immutable settled v3 publication once exact provider billing appears.The old
planner-qwen38-qlora-v2plan remains obsolete v3-contract evidence and must not be provisioned. Exact v3 billing/publication is still pending provider settlement.- exact source commit:
The replacement current-contract QLoRA authorization package is published at exact commit
c75e9f00700d94e54ab91dc62089b85686832f45in PR #28.Resolved prerequisites and exact plan:
- settled stock prerequisite:
runs/planner-current-qwen-stock-baseline-v3/publication.json, statusqlora-justified-settled - training corpus: 24 independently approved synthetic traces, SHA-256
a4ce7ac3e3365148ee4bd057928cfe3f72847637fe9c8ae7ebbadc29b608dcfa - development gate: 24 disjoint cases, excluded from training
- base artifact revision:
8aa5f05d26b7205477066e1449e0af13f762a299 - exact rendered-token proof: 24/24 traces, maximum 5,416 tokens; training context 8,192; no truncation
- one response-only adapter run: batch 1, accumulation 4, 9 optimizer steps, 36 example exposures / 1.5 passes
- no sweep, no automatic paid retry, adapter-only save, fresh-base reload probe required
- published stock results are reused for evaluation; stock is not rerun
- exact ceiling:
$1.50training +$1.50adapter evaluation =$3.00combined - current committed spend:
$29.1962968898326090175; projected maximum:$32.1962968898326090175 / $40.00 make check: 213 tests plus all generated-artifact, capacity, reconciliation, and environment checks pass
Paid flags remain false. The earlier in-chat authorization preceded this exact commit and did not name the
$3.00ceiling, so it has been treated as authorization to prepare the phase, not as spend authorization. The next state change is one explicit maintainer authorization naming commitc75e9f00700d94e54ab91dc62089b85686832f45and the$3.00combined ceiling. No new content review or licensing gate is required.- settled stock prerequisite:
Exact-head CI passed for
c75e9f00700d94e54ab91dc62089b85686832f45: https://github.com/loomarr/loomarr-models/actions/runs/35236676148/job/105254249717. The plan is clean and authorization-ready; all paid flags remain false.Authorization recorded from the repository owner on 2026-09-17:
- Authorized plan commit:
c75e9f00700d94e54ab91dc62089b85686832f45 - Combined maximum spend: $3.00
- Training slice activated now: up to $1.50, exactly one QLoRA training run, no sweep and no automatic paid retry
- Adapter-evaluation slice reserved for the later artifact-bound evaluation: up to $1.50
- Projected aggregate ledger ceiling after both slices:
32.1962968898326090175 / 40.00
This authorization does not introduce a licensing or additional human-review requirement. The current reviewed 24-trace corpus and deterministic Go authority boundaries remain unchanged.
- Authorized plan commit:
Training pod launched for the authorized current-contract QLoRA run.
- Execution source commit:
4900124cbd052a19d87230f0d0e309da7efb8571 - Exact-head CI: https://github.com/loomarr/loomarr-models/actions/runs/35298009779
- Pod:
xjnydoqn8p7cdy - Shape: one secure NVIDIA A40, CUDA 12.8,
CA-MTL-1 - Live GPU price at launch: $0.49/hour
- Storage: 40 GB container disk + 40 GB pod-persistent
/workspace - Provider creation time:
2026-09-18T02:08:08.843Z - Hard deletion deadline:
2026-09-18T04:38:08.843Z(9,000 seconds) - Training reservation: $1.50; maximum GPU component at the deadline: $1.225
- Run policy: exactly one training launch, no sweep, no automatic paid retry
The pod will be deleted after adapter verification and retrieval, before the deadline. Exact provider cost will be recorded only after billing settles.
- Execution source commit:
The single authorized QLoRA training launch ended in a terminal infrastructure/storage failure before optimizer step 1. No retry will be attempted.
- Execution source:
4900124cbd052a19d87230f0d0e309da7efb8571 - Pod:
xjnydoqn8p7cdy - Exit code:
1 - Failure: Hugging Face/Xet model reconstruction exhausted the 40 GB pod-persistent quota while acquiring
unsloth/Qwen3.8-27B-unsloth-bnb-4bitat revision8aa5f05d26b7205477066e1449e0af13f762a299 - Training progress: 0 optimizer steps; no adapter files produced
- Retrieved evidence archive SHA-256:
8858d36143a4b7c64dbb3cea5ef32ec1c3be696814735b1d995d0f3183b4fc38 - Preserved evidence: dependency sync, 213-test repository gate, paid preflight, compatibility probe, training log, exit code, GPU/storage snapshots, and empty adapter inventory
The pod is being deleted now. Exact provider cost will be posted after billing settles; no estimate will be substituted.
- Execution source:
Teardown confirmed for terminal training failure:
- Deleted pod:
xjnydoqn8p7cdy - Runpod read-back: pod
404; account pod list empty - Deleted before hard deadline
2026-09-18T04:38:08.843Z - Local evidence archive SHA-256 reverified as
8858d36143a4b7c64dbb3cea5ef32ec1c3be696814735b1d995d0f3183b4fc38 - Initial provider billing read has not settled yet (zero records), so no cost has been posted or estimated
Next action is exact provider settlement and a terminal, hash-bound failure publication. The reserved adapter evaluation cannot run because no adapter exists.
- Deleted pod:
Terminal failure publisher is now committed and CI-green while provider settlement is pending:
- Publisher commit:
c40603ffc172bf674c20415e97197d0b9e626252 - Exact-head CI: https://github.com/loomarr/loomarr-models/actions/runs/35299953187
- Repository gate: 217 tests plus all generators, budget reconciliation, environment verification, and corpus validation
- Publisher fails closed on evidence archive digest, source commit, clean checkout, exit code 1, zero adapter inventory, storage-exhaustion class, deleted resource, exact provider component sum, and the $1.50 training ceiling
No resource is running. Publication waits only for Runpod to emit the exact pod billing row.
- Publisher commit:
Terminal settlement is published for the failed current-contract QLoRA v1 run.
- Settlement commit:
14d5483e008bef7b2e56e0428d162ae08ce335e9 - Exact-head CI: https://github.com/loomarr/loomarr-models/actions/runs/35345174948
- Provider billing row: GPU
$0.15634634345769882, disk$0.003703703638166189, CPU$0, exact total$0.160050047095865 - Failure: model-download storage exhaustion before optimizer step 1; no adapter produced
- Failure archive SHA-256:
8858d36143a4b7c64dbb3cea5ef32ec1c3be696814735b1d995d0f3183b4fc38 - Publication:
runs/planner-current-qwen38-qlora-v1/publication.json, SHA-2564727eddfa2f6fed3ff0e9f6e84fb221ecf3a4a32ab4802b24147ea05b1ae3c81 - Canonical ledger after settlement:
29.3563469369284740175 / 40.00, with zero outstanding reservations - Runpod pod inventory: empty
- All training and paid-evaluation authority for v1 is revoked; the old runner now fails closed as terminal
make checkpasses 218 tests plus generated-artifact, billing-reconciliation, environment, and corpus checks. The next phase is a separate no-spend corrected training plan with a larger storage envelope; no paid retry is authorized by this settlement.- Settlement commit:
The corrected no-spend current-contract QLoRA v2 plan is published and exact-head CI is green.
- Plan commit:
adbde9e6d9e42f0eaf8402fb0aef5325d2fe4d9b - Exact-head CI: https://github.com/loomarr/loomarr-models/actions/runs/35345845118
- Experiment:
planner-current-qwen38-qlora-v2 - Correction: 40 GB container disk for
/opt/loomarr-venv; 80 GB pod-persistent/workspacefor Hugging Face cache, logs, and artifacts - Model acquisition:
HF_HOME=/workspace/hf-cache;HF_HUB_DISABLE_XET=1is applied before Unsloth/Hugging Face imports; launch requires at least 70 GB free on/workspace - Unchanged behavior inputs: exact Qwen3.8 revision, 24 independently approved training traces, disjoint 24-case development gate, max rendered length 5,416, context 8,192, nine optimizer steps, response-only adapter training
- Run policy: one configuration, no sweep, no automatic paid retry, adapter-only output, fresh-base reload probe
- Current committed spend:
$29.3563469369284740175 / $40.00 - Proposed ceiling:
$1.50training +$1.50adapter evaluation =$3.00combined - Maximum projected aggregate after both slices:
$32.3563469369284740175 / $40.00 - Exact clean plan preflight passes;
make checkpasses 222 tests plus generated-artifact, billing-reconciliation, environment, and corpus checks
All paid flags remain false. No pod, model download, training, evaluation, or new spend occurred. The next state change requires explicit maintainer authorization naming commit
adbde9e6d9e42f0eaf8402fb0aef5325d2fe4d9band the$3.00combined ceiling.- Plan commit:
Maintainer authorization recorded at 2026-09-18T12:58:05Z.
Authorized plan commit:
adbde9e6d9e42f0eaf8402fb0aef5325d2fe4d9b
Aggregate ceiling for this phase: $3.00 combinedThis authorization activates only the training slice now: at most $1.50 for exactly one
planner-current-qwen38-qlora-v2run, with no sweep and no automatic paid retry. The remaining $1.50 stays reserved for the preregistered adapter evaluation and may be activated only after training produces a verified, hash-bound adapter. Published stock-v3 results are reused.The program ledger before this phase is
$29.3563469369284740175; the maximum combined commitment is therefore$32.3563469369284740175 / $40.00. This authorization adds no separate review gate and grants no deployment, certification, or release authority.Training resource created at
2026-09-18T13:01:05.84Z.- Source authorization commit:
6c1a43d8af777be0d475060c3a3ff22df7b312db - Runpod pod:
dlrdzo9tq4hqif - Placement: secure cloud,
CA-MTL-1, one NVIDIA A40 - Live GPU rate at creation:
$0.49/hour - Image:
runpod/pytorch@sha256:4d1721e62b56d345c83b4fd6090664be6daf9312caab5b2e76f23d8231941851 - Storage: 40 GB container disk; 80 GB host-local persistent volume mounted at
/workspace - Supervisor termination deadline:
2026-09-18T15:31:05.101Z(9,000 seconds after create) - Training ceiling:
$1.50; combined training/evaluation ceiling:$3.00
Exactly one training launch is authorized. No sweep and no automatic or manual paid retry.
- Source authorization commit:
The single authorized training process launched at
2026-09-18T13:07:56Zon poddlrdzo9tq4hqiffrom source commit6c1a43d8af777be0d475060c3a3ff22df7b312db.Remote validation before launch passed: full repository gate, exact paid preflight, storage floor, pinned package import, CUDA availability, and A40 identity. The process is downloading the exact pinned Hugging Face revision into
/workspace/hf-cache;HF_HUB_DISABLE_XET=1is active. This is the one permitted training launch. Any failure is terminal and will not be retried.The single authorized v2 training launch failed closed at
2026-09-18T13:21:26Z, before optimizer step 1.- Pod:
dlrdzo9tq4hqif - Source:
6c1a43d8af777be0d475060c3a3ff22df7b312db - Exit code:
2 - Terminal error:
live rendered training capacity differs from the preregistered report - Model acquisition and A40 loading completed successfully before the capacity identity check failed.
- No adapter was produced.
Per the authorization, this run will not be retried. Evidence retrieval, pod deletion, and exact provider settlement are in progress. Paid adapter evaluation remains disabled.
- Pod:
Failure evidence was retrieved before teardown.
- Local evidence archive SHA-256:
958fa29db90c4cae4c85719c25b1472851f6f50cccb30f515769b3b80b0984ad - Training log SHA-256:
30a6717a5784adc529f515c05bff51ac3dddcdced285560c13b7f75beb5f0a03 - Exit record SHA-256:
53c234e5e8472b6ac51c1ae1cab3fe06fad053beb8ebfd8977b010655bfdd3c3 - Output inventory: empty experiment directory; no adapter or run manifest
- Pinned model cache snapshot:
8aa5f05d26b7205477066e1449e0af13f762a299
Pod
dlrdzo9tq4hqifwas deleted at2026-09-18T13:24:34.424Z. Provider read-back returns 404 and the active pod list is empty. Exact billing settlement is pending; no estimated cost will be posted.- Local evidence archive SHA-256:
Root-cause analysis for the terminal
planner-current-qwen38-qlora-v2run (no additional spend):- Reproduced the exact pinned
transformers==5.15.1,tokenizers==0.22.2, andjinja2==3.1.6rendering path locally against the hash-bound tokenizer files and all 24 training traces. - Live
AutoTokenizer.apply_chat_template(...)token counts match the preregistered report for every trace; the maximum is exactly 5,416 tokens, below the 8,192 limit. - The rendered-byte SHA differs: preregistered
35b00e0cffd1ac50e254148ceeaf13ba25f3dee7113c5685b74aa0e0472dd6cf; exact Transformers path0b121663f00262fb744080ece3e1318d3b0a2bfe0ed56dc6ba7ad3c165fc04f0. - Cause: the offline capacity checker used Jinja's default
tojsonpolicy (sort_keys=True), while the Transformers renderer preserves mapping insertion order. Example: offline tool JSON begins withfunction; live tool JSON begins withtype. This changes bytes without changing any token count. - Setting the offline Jinja policy to unsorted UTF-8 JSON reproduces the live Transformers output byte-for-byte across all 24 traces.
Conclusion: this was a harness parity defect, not a model-capacity failure and not evidence for or against fine-tuning quality. The run stopped before optimizer step 1 and produced no adapter. A corrected report/plan will require a new committed identity and new explicit authorization before any retry. Provider billing settlement is still pending; the pod is deleted and the account has zero active pods.
- Reproduced the exact pinned
planner-current-qwen38-qlora-v2is now terminal, published, and fully settled.- Terminal publication commit:
2b1d53b4e4ad164c81063d16db9c9831d1445210 - Exact-head CI: https://github.com/loomarr/loomarr-models/actions/runs/35354873900 (passed)
- Provider settlement: GPU
$0.06456321477890015+ disk$0.0013888889225199819+ CPU$0= exact total$0.06595210370142013 - Canonical committed spend:
$29.4222990406298941475of$40.00; outstanding reservations$0 - Pod and pod-persistent storage deleted; zero active pods
- Source commit:
6c1a43d8af777be0d475060c3a3ff22df7b312db - Source config SHA-256:
c5d38651b1c2a17540d9b69de26e1b3f605d8d50f565409906ce4add660a7595 - Failure archive SHA-256:
958fa29db90c4cae4c85719c25b1472851f6f50cccb30f515769b3b80b0984ad - Outcome: stopped before optimizer step 1; no adapter; evaluation not run; every authority revoked
- Root cause: offline Jinja
tojsonsorted object keys while pinned Transformers preserved insertion order. All 24 token counts matched exactly (maximum 5,416), but the rendered-byte SHA correctly failed closed.
The repository now records the immutable run evidence, provider settlement, terminal authorization/config, canonical budget reconciliation, and renderer-parity diagnosis.
make checkpasses all 226 tests. V2 cannot be rerun. The next experiment must have a new identity, a byte-exact capacity artifact generated with Transformers-compatible JSON serialization, and a new explicit authorization before any paid retry.- Terminal publication commit:
No-spend follow-up landed in
068fe77b56ff1965c9b190eccdf3fecd35b825d4: the capacity generator now explicitly configures Jinjatojsonfor unsorted UTF-8 JSON, matching pinned Transformers mapping-order semantics, with a regression assertion that the policy is installed before template compilation. Exact-head CI passed: https://github.com/loomarr/loomarr-models/actions/runs/35355029428This fixes the identified harness defect but does not create or authorize a retry. The next step is a new v3 experiment identity and newly generated capacity report whose rendered SHA is expected to be
0b121663f00262fb744080ece3e1318d3b0a2bfe0ed56dc6ba7ad3c165fc04f0; that report and plan must be committed and reviewed before any new paid authorization.Superseded by #32. The conditional 27B adapter rested on the v1 gate, which failed prompt-following models on policy, exact tool arguments, and date spans (PR #31,
src/loomarr_models/current_gate_v2.py). The target was also wrong: Loomarr serves Flash-Next, which runs a full 24-case screen about 23× faster than 27B bf16 on the local appliance. Any adapter is now decided per role after the stock screen in #35.
Outcome
Only if a preregistered stock Qwen baseline fails the current production planner contract for a repeatable model-quality reason, train one bounded adapter and decide whether it materially improves the optional local/offline planner lane.
This issue owns one conditional training configuration and one blind stock-versus-adapter evaluation. It does not authorize corpus construction, packaging, serving, release, application integration, production learning, or any paid action by itself.
Current-contract prerequisite: #22
Historical corpus/evidence: #9, #5, and #7
Application programs: loomarr/loomarr#828 and loomarr/loomarr#856
Recovery-evidence prerequisite: loomarr/loomarr#1195
Entry gates
Training remains blocked until all of the following are true:
tmdbId/tvdbIdoutput or omit current mandatory fields.Passing the stock thresholds ends this issue with a no-training decision.
Product hypothesis
A current-contract adapter may improve bounded evidence use while preserving useful creativity:
Controlled serendipity remains subordinate to grounding and hard constraints. Deterministic Loomarr code retains identity, source, policy, approval, authorization, scheduling, and playback authority.
Training contract
Evaluation contract
Compare on the same untouched current-contract holdout and runtime envelope:
Report hard failures and per-capability results for schema validity, exact tool arguments, exact-key grounding, date meaning, unsupported identities, authority violations, constraint preservation, source-evidence use, recovery, abstention, relevance, latency, memory, and settled cost. Novelty/diversity is scored only inside the grounded eligible set and never offsets a hard failure.
Promotion requires a preregistered meaningful gain over stock, improvement in the targeted capability families, zero unsupported-identity/authority/executable-envelope regressions, and acceptable latency for the declared local lane. Otherwise publish rejection and stop.
Two-speed learning boundary
This adapter is the slow shared-model loop. Immediate household personalization remains in explicit feedback, exposure memory, and deterministic local reranking. No online weight updates, hidden telemetry, unreviewed production ingestion, or cross-household training data is authorized.
Budget and authority
The earlier
$1.50A40 proposal is expired planning evidence, not a standing reservation. Before any paid action, publish:Opening or editing this issue authorizes
$0of provider, model-download, inference, GPU, storage, training, or evaluation spend.Acceptance criteria
make checkremains hermetic and reproduces every committed no-spend artifact.Stop point
Stop after the one stock-versus-adapter decision and its immutable evidence. Merging, quantizing, packaging, serving, deployment, release, and application activation remain separately tracked and unauthorized.