Skip to content

Conditionally train and evaluate one current-contract planner adapter #20

Description

@mantonx

Outcome

Only if a preregistered stock Qwen baseline fails the current production planner contract for a repeatable model-quality reason, train one bounded adapter and decide whether it materially improves the optional local/offline planner lane.

This issue owns one conditional training configuration and one blind stock-versus-adapter evaluation. It does not authorize corpus construction, packaging, serving, release, application integration, production learning, or any paid action by itself.

Current-contract prerequisite: #22
Historical corpus/evidence: #9, #5, and #7
Application programs: loomarr/loomarr#828 and loomarr/loomarr#856
Recovery-evidence prerequisite: loomarr/loomarr#1195

Entry gates

Training remains blocked until all of the following are true:

  1. Rebaseline the planner custom-model lane on the current production contract #22 binds one exact reviewed Loomarr production contract and publishes a current-contract training/development plan with no application holdout leakage.
  2. Historical v3/v4 final-response rows are excluded or regenerated and freshly reviewed; they may not teach obsolete tmdbId/tvdbId output or omit current mandatory fields.
  3. One exact stock Qwen artifact completes the current-contract baseline on the intended inference host.
  4. The stock result misses a preregistered quality threshold for a stable learned-behavior gap. Runtime, chat-template, tool-adapter, deterministic grounding, source retrieval, accounting, timeout, or fixture defects must be repaired outside training first.
  5. Recovery is measured with an actually injected and observed failure rather than the non-exercised historical fixture described in Make planner recovery certification inject and observe an actual tool failure loomarr#1195.
  6. The aggregate external-spend ledger is reconciled across the model and application programs, and a new exact reservation/authorization commit is reviewed.

Passing the stock thresholds ends this issue with a no-training decision.

Product hypothesis

A current-contract adapter may improve bounded evidence use while preserving useful creativity:

  • choose the correct first retrieval operation without invented narrowing;
  • copy exact catalog keys and emit complete executable Proposal JSON;
  • represent date meaning and multiple date axes correctly;
  • consume source-grounded collection evidence without memorizing membership;
  • preserve franchise, person/network, format, audience, region/language, ownership, inclusion/exclusion, refinement, and season constraints;
  • abstain honestly on conflict, ambiguity, thin evidence, or empty results;
  • recover from a fault only after actually observing one;
  • prefer a smaller accurate result over real-but-wrong padding.

Controlled serendipity remains subordinate to grounding and hard constraints. Deterministic Loomarr code retains identity, source, policy, approval, authorization, scheduling, and playback authority.

Training contract

  • Use one revision-pinned base/training artifact and one preregistered configuration; no sweep and no automatic paid retry.
  • Train only on newly approved current-contract rows from Rebaseline the planner custom-model lane on the current production contract #22. Historical v3/v4 corpora are evidence, not active training data, unless regenerated as new reviewed artifacts.
  • Never train on Loomarr certification/release/development cases, failed candidate generations, reviewer rationales, household data, histories, catalogs, decisions, paths, credentials, or raw production traces.
  • Preserve response-only SFT, deterministic seed, pinned environment and source commit, clean preflight, hard wall-clock alarm, and adapter-only output.
  • Post-save verification must reload the persisted adapter into a fresh pinned base and run a deterministic generation probe before publication.

Evaluation contract

Compare on the same untouched current-contract holdout and runtime envelope:

  1. the exact stock base;
  2. the new adapter;
  3. the prior rejected adapter only where historical results are explicitly non-comparative;
  4. the current hosted production bar as a separate product comparison after it is requalified on the same current contract.

Report hard failures and per-capability results for schema validity, exact tool arguments, exact-key grounding, date meaning, unsupported identities, authority violations, constraint preservation, source-evidence use, recovery, abstention, relevance, latency, memory, and settled cost. Novelty/diversity is scored only inside the grounded eligible set and never offsets a hard failure.

Promotion requires a preregistered meaningful gain over stock, improvement in the targeted capability families, zero unsupported-identity/authority/executable-envelope regressions, and acceptable latency for the declared local lane. Otherwise publish rejection and stop.

Two-speed learning boundary

This adapter is the slow shared-model loop. Immediate household personalization remains in explicit feedback, exposure memory, and deterministic local reranking. No online weight updates, hidden telemetry, unreviewed production ingestion, or cross-household training data is authorized.

Budget and authority

The earlier $1.50 A40 proposal is expired planning evidence, not a standing reservation. Before any paid action, publish:

  • the reconciled aggregate ledger and remaining authorization;
  • live provider/GPU availability and price;
  • one exact training-plus-evaluation ceiling;
  • a clean source commit with paid flags still false;
  • a separate maintainer authorization naming that commit and ceiling.

Opening or editing this issue authorizes $0 of provider, model-download, inference, GPU, storage, training, or evaluation spend.

Acceptance criteria

  • Every entry gate is satisfied or the issue closes with a no-training decision.
  • One exact configuration, source commit, corpus digest, base revision, environment lock, seed, and budget reservation are preregistered.
  • Preflight rejects split leakage, stale response schemas, dirty inputs, revision drift, extra configurations, output escape, and budget overflow before importing the training stack.
  • One authorized run saves adapter-only artifacts and a hash-bound manifest, or a terminal failure is published without automatic paid retry.
  • Stock and adapter inference consume the same untouched cases in the same order and runtime envelope.
  • The publication reports every hard gate, capability family, latency/memory measure, exact settled cost, and promotion decision.
  • The final decision names the measured Track the three-pillar Loomarr AI model program loomarr#856 pillar(s) and explicitly leaves unmeasured pillars unclaimed.
  • make check remains hermetic and reproduces every committed no-spend artifact.

Stop point

Stop after the one stock-versus-adapter decision and its immutable evidence. Merging, quantizing, packaging, serving, deployment, release, and application activation remain separately tracked and unauthorized.

Activity

  1. mantonx commented on Sep 4, 2026

    @mantonx
    ContributorAuthor

    The no-spend QLoRA v2 preregistration is now published in draft PR #21: #21

    Exact-head CI passed: https://github.com/loomarr/loomarr-models/actions/runs/33820386753/job/100861595512

    The plan binds Qwen3.8-27B, the frozen 120-trace corpus, the A40 environment, 45-step adapter-only recipe, and the $1.50 reservation. Paid training and paid evaluation remain disabled. The next state change requires a separate exact authorization commit.

  2. mantonx commented on Sep 4, 2026

    @mantonx
    ContributorAuthor

    Read-only Runpod readiness check at 2026-09-04T00:00Z (no pod created):

    • NVIDIA A40 / pool AMPERE_48: availability HIGH
    • 48 GB VRAM; CUDA 12.8 currently available
    • community $0.35/hour; secure $0.49/hour
    • maximum secure GPU charge under the 9,000-second alarm: $1.225
    • account GPU pods before launch: 0

    The $1.50 training reservation still covers the live secure-compute ceiling with $0.275 headroom for pod disk. Availability and price are temporal and must be rechecked immediately before provisioning. No paid training or evaluation was authorized or started by this check.

  3. mantonx commented on Sep 4, 2026

    @mantonx
    ContributorAuthor

    No-spend Runpod execution-readiness audit is published in PR #21 at commit 1120d57f80d412d4143506517a285f662a999658.

    The milestone closes the operational gaps between a valid experiment manifest and a safely recoverable GPU run:

    • checked-in Runpod lifecycle for the exact authorization commit, secure A40/image digest, same-DC 40 GB network volume, detached execution, artifact retrieval, and teardown;
    • post-save reload of the persisted PEFT adapter into a fresh pinned base-model instance plus a deterministic generation probe before a run manifest can exist;
    • fail-closed artifact verifier for all 45 steps, source/config/corpus/environment identities, package/runtime envelope, complete file hashes, adapter-only output, and the reload probe;
    • six sabotage tests covering digest drift, wrong source commit, incomplete steps, unauthorized state, and merged/base-weight output;
    • the verifier is now a clean tracked preflight-critical input.

    Current provider audit (no resources created):

    • Runpod GPU pods: 0; network volumes: 0;
    • secure A40 availability: HIGH; price: $0.49/hour; CUDA 12.8 available;
    • runpodctl 2.12.0 installed locally;
    • important provider caveat: v2.12 removed --stop-after / --terminate-after because the backend accepted but never enforced them (fix(pod): remove --stop-after and --terminate-after runpod/runpodctl#330). The runbook therefore requires uninterrupted supervision and explicit deletion by 9,000 seconds from provider creation. At the live rate, GPU maximum is $1.225; estimated 40 GB container + network storage remains below $0.03, inside the $1.50 reservation.

    Verification at this commit:

    • make check — 117 tests pass;
    • make preflight-qlora-v2-plan — exact 120/120 reviewed corpus, projected aggregate $29.4826615675672820 / $40.00;
    • generated plan check and git diff --check pass.

    Paid training and evaluation remain disabled. No pod, volume, model download, inference, or spend was started by this audit. The next state change still requires explicit maintainer authorization for one $1.50 training run from this exact reviewed commit, with no automatic paid retry.

  4. mantonx commented on Sep 4, 2026

    @mantonx
    ContributorAuthor

    Exact-head CI is green for 1120d57f80d412d4143506517a285f662a999658: https://github.com/loomarr/loomarr-models/actions/runs/33824479285/job/100874084985

    PR #21 remains draft, open, mergeable, and unmerged. The paid gates are unchanged.

  5. mantonx commented on Sep 4, 2026

    @mantonx
    ContributorAuthor

    Paid execution is now blocked on #22.

    Production contract drift is material: loomarr/loomarr#967 advances the planner to suggester-prompt-v4 / catalog-search-v4 and adds independently scored network, cast, and creator routes. This PR's exact commit 1120d57 remains valid no-spend evidence for the older v3 contract, but it is no longer an authorization-ready training source because its 120 traces and 60-case gate cannot teach or detect regressions in the new fields.

    Do not provision from 1120d57. #22 owns the no-spend v6 import, leakage check, new reviewed route contrasts, disjoint development cases, stock-v6 baseline, and replacement exact authorization commit. The existing $1.50 ceiling and no-automatic-retry boundary remain unchanged; no reservation is spent or newly committed by this comment.

  6. mantonx commented on Sep 4, 2026

    @mantonx
    ContributorAuthor

    Acceptance addition from the latest live Loomarr TGIF diagnosis: the targeted adapter should improve the first retrieval turn, not rely on an app-side second chance. For a terse named programming block such as TGIF, target one valid collection tool call with at least three accurate, diverse exact members inside the normal latency bound. Score real-but-wrong members as failures and a smaller honest set above padding. Include non-TGIF collections/networks/franchises in the held-out gate so this is retrieval-strategy learning, not acronym memorization. See corpus issue #9 for the full observed evidence and split requirements.

  7. changed the title [-]Train and evaluate one targeted planner QLoRA v2 adapter[/-] [+]Conditionally train and evaluate one current-contract planner adapter[/+] on Sep 16, 2026
  8. mantonx commented on Sep 17, 2026

    @mantonx
    ContributorAuthor

    Current-contract prerequisite update:

    All three PRs have green CI. Training remains blocked on genuine independent draft review, a completed authorized stock-baseline quality miss, the application recovery prerequisite, and a separate aggregate-spend authorization. No paid action occurred.

  9. mantonx commented on Sep 17, 2026

    @mantonx
    ContributorAuthor

    Current-contract stock prerequisite is now a genuine model-quality miss (issue #29).

    Validated evidence:

    • exact source commit: aa1fa71e908240a3eae43f50e9b763332ec8cea5
    • exact model revision: 8aa5f05d26b7205477066e1449e0af13f762a299
    • 24/24 current-contract development cases completed on one secure A40
    • run-manifest SHA-256: d2e496999a8f97b88b3d821aeadcdeb20cf39895813262ec7b5d129add2ffc7a
    • complete model-quality result: 32 hard failures; grounded completion / correct tool operation / schema validity 0.375; policy accuracy / recovery 0.0
    • decision: qloraJustified=true, trainingAuthorized=false
    • pod deleted; zero active Runpod pods verified

    This satisfies the stock-result quality predicate for preparing a replacement current-contract QLoRA authorization package. It does not authorize training. The package must use only the independently approved 24-row corpus/planner-current-v1/traces.jsonl, bind the current v5 contract and hash-only holdout denylist, keep the 24 development cases out of training, preserve one configuration/no retry/adapter-only output, and bind the immutable settled v3 publication once exact provider billing appears.

    The old planner-qwen38-qlora-v2 plan remains obsolete v3-contract evidence and must not be provisioned. Exact v3 billing/publication is still pending provider settlement.

  10. mantonx commented on Sep 17, 2026

    @mantonx
    ContributorAuthor

    The replacement current-contract QLoRA authorization package is published at exact commit c75e9f00700d94e54ab91dc62089b85686832f45 in PR #28.

    Resolved prerequisites and exact plan:

    • settled stock prerequisite: runs/planner-current-qwen-stock-baseline-v3/publication.json, status qlora-justified-settled
    • training corpus: 24 independently approved synthetic traces, SHA-256 a4ce7ac3e3365148ee4bd057928cfe3f72847637fe9c8ae7ebbadc29b608dcfa
    • development gate: 24 disjoint cases, excluded from training
    • base artifact revision: 8aa5f05d26b7205477066e1449e0af13f762a299
    • exact rendered-token proof: 24/24 traces, maximum 5,416 tokens; training context 8,192; no truncation
    • one response-only adapter run: batch 1, accumulation 4, 9 optimizer steps, 36 example exposures / 1.5 passes
    • no sweep, no automatic paid retry, adapter-only save, fresh-base reload probe required
    • published stock results are reused for evaluation; stock is not rerun
    • exact ceiling: $1.50 training + $1.50 adapter evaluation = $3.00 combined
    • current committed spend: $29.1962968898326090175; projected maximum: $32.1962968898326090175 / $40.00
    • make check: 213 tests plus all generated-artifact, capacity, reconciliation, and environment checks pass

    Paid flags remain false. The earlier in-chat authorization preceded this exact commit and did not name the $3.00 ceiling, so it has been treated as authorization to prepare the phase, not as spend authorization. The next state change is one explicit maintainer authorization naming commit c75e9f00700d94e54ab91dc62089b85686832f45 and the $3.00 combined ceiling. No new content review or licensing gate is required.

  11. mantonx commented on Sep 17, 2026

    @mantonx
    ContributorAuthor

    Exact-head CI passed for c75e9f00700d94e54ab91dc62089b85686832f45: https://github.com/loomarr/loomarr-models/actions/runs/35236676148/job/105254249717. The plan is clean and authorization-ready; all paid flags remain false.

  12. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    Authorization recorded from the repository owner on 2026-09-17:

    • Authorized plan commit: c75e9f00700d94e54ab91dc62089b85686832f45
    • Combined maximum spend: $3.00
    • Training slice activated now: up to $1.50, exactly one QLoRA training run, no sweep and no automatic paid retry
    • Adapter-evaluation slice reserved for the later artifact-bound evaluation: up to $1.50
    • Projected aggregate ledger ceiling after both slices: 32.1962968898326090175 / 40.00

    This authorization does not introduce a licensing or additional human-review requirement. The current reviewed 24-trace corpus and deterministic Go authority boundaries remain unchanged.

  13. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    Training pod launched for the authorized current-contract QLoRA run.

    • Execution source commit: 4900124cbd052a19d87230f0d0e309da7efb8571
    • Exact-head CI: https://github.com/loomarr/loomarr-models/actions/runs/35298009779
    • Pod: xjnydoqn8p7cdy
    • Shape: one secure NVIDIA A40, CUDA 12.8, CA-MTL-1
    • Live GPU price at launch: $0.49/hour
    • Storage: 40 GB container disk + 40 GB pod-persistent /workspace
    • Provider creation time: 2026-09-18T02:08:08.843Z
    • Hard deletion deadline: 2026-09-18T04:38:08.843Z (9,000 seconds)
    • Training reservation: $1.50; maximum GPU component at the deadline: $1.225
    • Run policy: exactly one training launch, no sweep, no automatic paid retry

    The pod will be deleted after adapter verification and retrieval, before the deadline. Exact provider cost will be recorded only after billing settles.

  14. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    The single authorized QLoRA training launch ended in a terminal infrastructure/storage failure before optimizer step 1. No retry will be attempted.

    • Execution source: 4900124cbd052a19d87230f0d0e309da7efb8571
    • Pod: xjnydoqn8p7cdy
    • Exit code: 1
    • Failure: Hugging Face/Xet model reconstruction exhausted the 40 GB pod-persistent quota while acquiring unsloth/Qwen3.8-27B-unsloth-bnb-4bit at revision 8aa5f05d26b7205477066e1449e0af13f762a299
    • Training progress: 0 optimizer steps; no adapter files produced
    • Retrieved evidence archive SHA-256: 8858d36143a4b7c64dbb3cea5ef32ec1c3be696814735b1d995d0f3183b4fc38
    • Preserved evidence: dependency sync, 213-test repository gate, paid preflight, compatibility probe, training log, exit code, GPU/storage snapshots, and empty adapter inventory

    The pod is being deleted now. Exact provider cost will be posted after billing settles; no estimate will be substituted.

  15. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    Teardown confirmed for terminal training failure:

    • Deleted pod: xjnydoqn8p7cdy
    • Runpod read-back: pod 404; account pod list empty
    • Deleted before hard deadline 2026-09-18T04:38:08.843Z
    • Local evidence archive SHA-256 reverified as 8858d36143a4b7c64dbb3cea5ef32ec1c3be696814735b1d995d0f3183b4fc38
    • Initial provider billing read has not settled yet (zero records), so no cost has been posted or estimated

    Next action is exact provider settlement and a terminal, hash-bound failure publication. The reserved adapter evaluation cannot run because no adapter exists.

  16. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    Terminal failure publisher is now committed and CI-green while provider settlement is pending:

    • Publisher commit: c40603ffc172bf674c20415e97197d0b9e626252
    • Exact-head CI: https://github.com/loomarr/loomarr-models/actions/runs/35299953187
    • Repository gate: 217 tests plus all generators, budget reconciliation, environment verification, and corpus validation
    • Publisher fails closed on evidence archive digest, source commit, clean checkout, exit code 1, zero adapter inventory, storage-exhaustion class, deleted resource, exact provider component sum, and the $1.50 training ceiling

    No resource is running. Publication waits only for Runpod to emit the exact pod billing row.

  17. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    Terminal settlement is published for the failed current-contract QLoRA v1 run.

    • Settlement commit: 14d5483e008bef7b2e56e0428d162ae08ce335e9
    • Exact-head CI: https://github.com/loomarr/loomarr-models/actions/runs/35345174948
    • Provider billing row: GPU $0.15634634345769882, disk $0.003703703638166189, CPU $0, exact total $0.160050047095865
    • Failure: model-download storage exhaustion before optimizer step 1; no adapter produced
    • Failure archive SHA-256: 8858d36143a4b7c64dbb3cea5ef32ec1c3be696814735b1d995d0f3183b4fc38
    • Publication: runs/planner-current-qwen38-qlora-v1/publication.json, SHA-256 4727eddfa2f6fed3ff0e9f6e84fb221ecf3a4a32ab4802b24147ea05b1ae3c81
    • Canonical ledger after settlement: 29.3563469369284740175 / 40.00, with zero outstanding reservations
    • Runpod pod inventory: empty
    • All training and paid-evaluation authority for v1 is revoked; the old runner now fails closed as terminal

    make check passes 218 tests plus generated-artifact, billing-reconciliation, environment, and corpus checks. The next phase is a separate no-spend corrected training plan with a larger storage envelope; no paid retry is authorized by this settlement.

  18. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    The corrected no-spend current-contract QLoRA v2 plan is published and exact-head CI is green.

    • Plan commit: adbde9e6d9e42f0eaf8402fb0aef5325d2fe4d9b
    • Exact-head CI: https://github.com/loomarr/loomarr-models/actions/runs/35345845118
    • Experiment: planner-current-qwen38-qlora-v2
    • Correction: 40 GB container disk for /opt/loomarr-venv; 80 GB pod-persistent /workspace for Hugging Face cache, logs, and artifacts
    • Model acquisition: HF_HOME=/workspace/hf-cache; HF_HUB_DISABLE_XET=1 is applied before Unsloth/Hugging Face imports; launch requires at least 70 GB free on /workspace
    • Unchanged behavior inputs: exact Qwen3.8 revision, 24 independently approved training traces, disjoint 24-case development gate, max rendered length 5,416, context 8,192, nine optimizer steps, response-only adapter training
    • Run policy: one configuration, no sweep, no automatic paid retry, adapter-only output, fresh-base reload probe
    • Current committed spend: $29.3563469369284740175 / $40.00
    • Proposed ceiling: $1.50 training + $1.50 adapter evaluation = $3.00 combined
    • Maximum projected aggregate after both slices: $32.3563469369284740175 / $40.00
    • Exact clean plan preflight passes; make check passes 222 tests plus generated-artifact, billing-reconciliation, environment, and corpus checks

    All paid flags remain false. No pod, model download, training, evaluation, or new spend occurred. The next state change requires explicit maintainer authorization naming commit adbde9e6d9e42f0eaf8402fb0aef5325d2fe4d9b and the $3.00 combined ceiling.

  19. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    Maintainer authorization recorded at 2026-09-18T12:58:05Z.

    Authorized plan commit: adbde9e6d9e42f0eaf8402fb0aef5325d2fe4d9b
    Aggregate ceiling for this phase: $3.00 combined

    This authorization activates only the training slice now: at most $1.50 for exactly one planner-current-qwen38-qlora-v2 run, with no sweep and no automatic paid retry. The remaining $1.50 stays reserved for the preregistered adapter evaluation and may be activated only after training produces a verified, hash-bound adapter. Published stock-v3 results are reused.

    The program ledger before this phase is $29.3563469369284740175; the maximum combined commitment is therefore $32.3563469369284740175 / $40.00. This authorization adds no separate review gate and grants no deployment, certification, or release authority.

  20. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    Training resource created at 2026-09-18T13:01:05.84Z.

    • Source authorization commit: 6c1a43d8af777be0d475060c3a3ff22df7b312db
    • Runpod pod: dlrdzo9tq4hqif
    • Placement: secure cloud, CA-MTL-1, one NVIDIA A40
    • Live GPU rate at creation: $0.49/hour
    • Image: runpod/pytorch@sha256:4d1721e62b56d345c83b4fd6090664be6daf9312caab5b2e76f23d8231941851
    • Storage: 40 GB container disk; 80 GB host-local persistent volume mounted at /workspace
    • Supervisor termination deadline: 2026-09-18T15:31:05.101Z (9,000 seconds after create)
    • Training ceiling: $1.50; combined training/evaluation ceiling: $3.00

    Exactly one training launch is authorized. No sweep and no automatic or manual paid retry.

  21. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    The single authorized training process launched at 2026-09-18T13:07:56Z on pod dlrdzo9tq4hqif from source commit 6c1a43d8af777be0d475060c3a3ff22df7b312db.

    Remote validation before launch passed: full repository gate, exact paid preflight, storage floor, pinned package import, CUDA availability, and A40 identity. The process is downloading the exact pinned Hugging Face revision into /workspace/hf-cache; HF_HUB_DISABLE_XET=1 is active. This is the one permitted training launch. Any failure is terminal and will not be retried.

  22. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    The single authorized v2 training launch failed closed at 2026-09-18T13:21:26Z, before optimizer step 1.

    • Pod: dlrdzo9tq4hqif
    • Source: 6c1a43d8af777be0d475060c3a3ff22df7b312db
    • Exit code: 2
    • Terminal error: live rendered training capacity differs from the preregistered report
    • Model acquisition and A40 loading completed successfully before the capacity identity check failed.
    • No adapter was produced.

    Per the authorization, this run will not be retried. Evidence retrieval, pod deletion, and exact provider settlement are in progress. Paid adapter evaluation remains disabled.

  23. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    Failure evidence was retrieved before teardown.

    • Local evidence archive SHA-256: 958fa29db90c4cae4c85719c25b1472851f6f50cccb30f515769b3b80b0984ad
    • Training log SHA-256: 30a6717a5784adc529f515c05bff51ac3dddcdced285560c13b7f75beb5f0a03
    • Exit record SHA-256: 53c234e5e8472b6ac51c1ae1cab3fe06fad053beb8ebfd8977b010655bfdd3c3
    • Output inventory: empty experiment directory; no adapter or run manifest
    • Pinned model cache snapshot: 8aa5f05d26b7205477066e1449e0af13f762a299

    Pod dlrdzo9tq4hqif was deleted at 2026-09-18T13:24:34.424Z. Provider read-back returns 404 and the active pod list is empty. Exact billing settlement is pending; no estimated cost will be posted.

  24. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    Root-cause analysis for the terminal planner-current-qwen38-qlora-v2 run (no additional spend):

    • Reproduced the exact pinned transformers==5.15.1, tokenizers==0.22.2, and jinja2==3.1.6 rendering path locally against the hash-bound tokenizer files and all 24 training traces.
    • Live AutoTokenizer.apply_chat_template(...) token counts match the preregistered report for every trace; the maximum is exactly 5,416 tokens, below the 8,192 limit.
    • The rendered-byte SHA differs: preregistered 35b00e0cffd1ac50e254148ceeaf13ba25f3dee7113c5685b74aa0e0472dd6cf; exact Transformers path 0b121663f00262fb744080ece3e1318d3b0a2bfe0ed56dc6ba7ad3c165fc04f0.
    • Cause: the offline capacity checker used Jinja's default tojson policy (sort_keys=True), while the Transformers renderer preserves mapping insertion order. Example: offline tool JSON begins with function; live tool JSON begins with type. This changes bytes without changing any token count.
    • Setting the offline Jinja policy to unsorted UTF-8 JSON reproduces the live Transformers output byte-for-byte across all 24 traces.

    Conclusion: this was a harness parity defect, not a model-capacity failure and not evidence for or against fine-tuning quality. The run stopped before optimizer step 1 and produced no adapter. A corrected report/plan will require a new committed identity and new explicit authorization before any retry. Provider billing settlement is still pending; the pod is deleted and the account has zero active pods.

  25. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    planner-current-qwen38-qlora-v2 is now terminal, published, and fully settled.

    • Terminal publication commit: 2b1d53b4e4ad164c81063d16db9c9831d1445210
    • Exact-head CI: https://github.com/loomarr/loomarr-models/actions/runs/35354873900 (passed)
    • Provider settlement: GPU $0.06456321477890015 + disk $0.0013888889225199819 + CPU $0 = exact total $0.06595210370142013
    • Canonical committed spend: $29.4222990406298941475 of $40.00; outstanding reservations $0
    • Pod and pod-persistent storage deleted; zero active pods
    • Source commit: 6c1a43d8af777be0d475060c3a3ff22df7b312db
    • Source config SHA-256: c5d38651b1c2a17540d9b69de26e1b3f605d8d50f565409906ce4add660a7595
    • Failure archive SHA-256: 958fa29db90c4cae4c85719c25b1472851f6f50cccb30f515769b3b80b0984ad
    • Outcome: stopped before optimizer step 1; no adapter; evaluation not run; every authority revoked
    • Root cause: offline Jinja tojson sorted object keys while pinned Transformers preserved insertion order. All 24 token counts matched exactly (maximum 5,416), but the rendered-byte SHA correctly failed closed.

    The repository now records the immutable run evidence, provider settlement, terminal authorization/config, canonical budget reconciliation, and renderer-parity diagnosis. make check passes all 226 tests. V2 cannot be rerun. The next experiment must have a new identity, a byte-exact capacity artifact generated with Transformers-compatible JSON serialization, and a new explicit authorization before any paid retry.

  26. mantonx commented on Sep 18, 2026

    @mantonx
    ContributorAuthor

    No-spend follow-up landed in 068fe77b56ff1965c9b190eccdf3fecd35b825d4: the capacity generator now explicitly configures Jinja tojson for unsorted UTF-8 JSON, matching pinned Transformers mapping-order semantics, with a regression assertion that the policy is installed before template compilation. Exact-head CI passed: https://github.com/loomarr/loomarr-models/actions/runs/35355029428

    This fixes the identified harness defect but does not create or authorize a retry. The next step is a new v3 experiment identity and newly generated capacity report whose rendered SHA is expected to be 0b121663f00262fb744080ece3e1318d3b0a2bfe0ed56dc6ba7ad3c165fc04f0; that report and plan must be committed and reviewed before any new paid authorization.

  27. mantonx commented on Sep 27, 2026

    @mantonx
    ContributorAuthor

    Superseded by #32. The conditional 27B adapter rested on the v1 gate, which failed prompt-following models on policy, exact tool arguments, and date spans (PR #31, src/loomarr_models/current_gate_v2.py). The target was also wrong: Loomarr serves Flash-Next, which runs a full 24-case screen about 23× faster than 27B bf16 on the local appliance. Any adapter is now decided per role after the stock screen in #35.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions