Fix manufactured confidence on ambiguous inputs with one-vs-all distillation - #6
Merged
Merged
Conversation
…llation (#5) Issue #5: `# Heading` + a single-word dash list scored Yaml 0.91 / Markdown 0.07 even though the Magika teacher reads it as markdown 0.96. Root cause: the teacher predicts over 214 labels, and the training cache renormalized its probabilities over the 48 exported labels. That conditions on "the input is one of the head labels" and manufactures confident targets for inputs the teacher mostly places on txt/unknown - a bare `- item` list (valid YAML and valid Markdown) trained toward yaml 0.9+ while the teacher kept ~70% of its mass out-of-head. The fix retrains the shipped checkpoint on a rebuilt public corpus with a target scheme that never renormalizes away out-of-head mass (cf. "Revisiting One-vs-All Classifiers for Predictive Uncertainty and OOD Detection", Padhy et al. 2020): - Cache the teacher's raw per-class head marginals (--head-marginal-targets) and distill them with per-class sigmoid cross-entropy (--soft-loss-mode bce). Out-of-scope mass simply lowers every class target instead of being renormalized into false confidence. - Discount the hard argmax target by the teacher's in-head mass (--mass-discounted-hard-labels) and drop rows whose in-head mass is noise (--min-teacher-head-mass 0.1). - Self-distill from a ~2x larger one-vs-all parent (wordseq-b1536-k3-m2048-med-3conv-hidden, 0.9492 test teacher parity) whose sigmoid marginals are cached via the new scripts/cache_self_distill.py. - The runtime is unchanged: softmax over BCE-trained per-class log-odds is near-uniform when every class is unlikely and confident when one dominates. The corpus is rebuilt from public sources with the original layout (scripts/build_finetune_corpus.py): bigcode/the-stack-smol-xl and bigcode/the-stack per-language samples, GitHub repos for labels absent from The Stack (objectivec, gradle, gemfile), and a small synthetic set covering the issue #5 ambiguity, including train-only bare dash lists whose teacher targets are genuinely split. Files from one repository always land in the same split. Results on the exact issue reproduction: Markdown 0.86 / Yaml 0.06 (was Yaml 0.91 / Markdown 0.07). Genuinely balanced inputs now report split probabilities: a bare `- first\n- second` list reads Yaml 0.54 / Markdown 0.28 (was 0.99), name lists 0.43/0.39, and the median top-1 over 200 sampled bare lists drops 0.92 -> 0.51. Aggregates improve across the board on the rebuilt held-out split versus the previous artifact: fs accuracy 0.9262 -> 0.9424, teacher parity 0.9221 -> 0.9441, macro recall 0.9299 -> 0.9397; tree-mode accuracy on local repos: dioxus 95.6% -> 97.1%, wasm-bindgen 90.2% -> 95.4%, docsite 91.2% -> 93.9%. Also included: - Regression tests for the issue snippet, capitalized name lists, bare-list uncertainty (top-2 = {yaml, markdown}, top-1 < 0.9), and YAML guards. - scripts/build_fs_labels.py and scripts/cache_self_distill.py, both previously referenced by the tooling but missing from the repo, and scripts/make_pruned48_config.py to regenerate the pruned teacher config from the magika pip package. - Teacher-vetted synthetic hard-boundary generators (hard_gen_*.py) for the remaining student-error confusion cells, with negative A/B results documented: at this capacity the c/cpp and js/ts sets only relocate errors inside genuinely arbitrary boundaries, so the shipped recipe does not sample them. - Trainer support for all of the above plus --checkpoint-weights (npz save/resume for non-exportable parent architectures), with the legacy renormalizing recipe preserved as the default. - Docs: README/MODEL_CARD metrics and SHA-256 for the new artifact, an analysis of the remaining ambiguous pairs in the confusion matrix, the full reproducible fine-tune recipe in TRAINING.md (including A/B findings: FP-then-QAT collapses on short schedules; hard-loss 0.5 re-sharpens ambiguous rows on long ones), and regenerated confusion reports/images. README confusion images are pinned to this branch name; re-pin to the merge commit after landing, matching the existing convention. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <166158716+staging-devin-ai-integration[bot]@users.noreply.github.com>
The model inspects just the first and last 4096 bytes of a file, but the example read every file in full, so tree walks over directories with large artifacts spent their time streaming bytes the model never looks at. Read the head and tail blocks and report the real file size instead; output is byte-identical on full-tree runs. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <166158716+staging-devin-ai-integration[bot]@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The fix retrains the shipped checkpoint on a rebuilt public corpus with a target scheme that never renormalizes away out-of-head mass (cf. "Revisiting One-vs-All Classifiers for Predictive Uncertainty and OOD Detection", Padhy et al. 2020):
The corpus is rebuilt from public sources with the original layout (scripts/build_finetune_corpus.py): bigcode/the-stack-smol-xl and bigcode/the-stack per-language samples, GitHub repos for labels absent from The Stack (objectivec, gradle, gemfile), and a small synthetic set covering the issue #5 ambiguity, including train-only bare dash lists whose teacher targets are genuinely split. Files from one repository always land in the same split.
Results on the exact issue reproduction: Markdown 0.86 / Yaml 0.06 (was Yaml 0.91 / Markdown 0.07). Genuinely balanced inputs now report split probabilities: a bare
- first\n- secondlist reads Yaml 0.54 / Markdown 0.28 (was 0.99), name lists 0.43/0.39, and the median top-1 over 200 sampled bare lists drops 0.92 -> 0.51. Aggregates improve across the board on the rebuilt held-out split versus the previousartifact: fs accuracy 0.9262 -> 0.9424, teacher parity 0.9221 ->
0.9441, macro recall 0.9299 -> 0.9397; tree-mode accuracy on local repos: dioxus 95.6% -> 97.1%, wasm-bindgen 90.2% -> 95.4%, docsite 91.2% -> 93.9%.
Also included:
README confusion images are pinned to this branch name; re-pin to the merge commit after landing, matching the existing convention.
Fixes #5
Devin Review