Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
63 changes: 54 additions & 9 deletions MODEL_CARD.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
- File: `assets/magika/source-student-q4.bin`
- Format: weights-only MSQ1 quantized tensor payload
- Size: 47,840 bytes
- SHA-256: `59ef24167bddd1364eb9c1650add8a67e1a542b5155fac67f5e1cda07df0c0f0`
- SHA-256: `8493d2d3757572c8661141e414b1c0755aa08d4c4e5382dfbbc6b73b02d89083`
- Architecture: `wordseq-b1024-k3-m2048-tiny-3conv-hidden`
- Tokenizer: word-unit tokenizer version 3
- Output head: 48 model labels exposed one-to-one as public `Language` variants
Expand All @@ -19,21 +19,63 @@ decisions, malware classification, or legal identification of file provenance.

## Training Source

The student was trained from Google's Magika v3.3 teacher predictions over a
source-language corpus assembled from extension-suffixed files. The corpus used
for the shipped model was the `bigorig` split, extracted from a GitHub
partial-clone blob index of roughly 6,000 popular code repositories.
The student was originally trained from Google's Magika v3.3 teacher
predictions over a source-language corpus assembled from extension-suffixed
files (the `bigorig` split, extracted from a GitHub partial-clone blob index of
roughly 6,000 popular code repositories).

The shipped artifact is that model fine-tuned from its exported checkpoint on a
rebuilt public corpus (see `scripts/build_finetune_corpus.py`): per-language
samples from `bigcode/the-stack-smol-xl` and `bigcode/the-stack`, GitHub repo
files for labels absent from The Stack, and a small synthetic set targeting the
Markdown/YAML bare-list ambiguity reported in issue #5, including train-only
bare `- item` lists.

The fine-tune distills the same Magika v3.3 teacher one-vs-all: per-class
binary cross-entropy against the teacher's raw head-label marginals, plus a
hard term against the teacher argmax discounted by the teacher's in-head
probability mass. Unlike the original softmax distillation, this never
renormalizes away the probability mass the teacher assigns to labels outside
the 48-label head (such as `txt`), so inputs the teacher considers ambiguous
or out-of-scope train toward uniformly low logits and keep low softmax
confidence at inference.

The shipped run additionally self-distills from a ~2x larger intermediate
parent (`wordseq-b1536-k3-m2048-med-3conv-hidden`, 0.9492 test teacher parity)
trained on the same corpus with the same one-vs-all scheme; the parent's
per-class sigmoid marginals are cached with `scripts/cache_self_distill.py`
and added as a second BCE term.

The model distills teacher probabilities and filesystem-extension labels. It
does not contain original source files, but its labels and soft targets are
derived from the training corpus and Magika teacher.

## Evaluation

Manifest-aligned held-out filesystem-label test split:
Held-out filesystem-label test split of the rebuilt corpus (34,087 files,
train/valid/test repositories are disjoint, rows where the teacher keeps at
most 10% of its probability mass on the head labels are excluded):

- `test_fs_accuracy=0.965238`
- `macro_recall=0.965411`
- `test_fs_accuracy=0.942353`
- `macro_recall=0.939690`
- `test_teacher_parity=0.944055`

For comparison, the pre-fine-tune artifact scores `test_fs_accuracy=0.926160`,
`macro_recall=0.929904`, and `test_teacher_parity=0.922111` on the same split.
The rebuilt split is balanced across all 48 labels (including rare classes),
so these numbers are not comparable to the `bigorig` metrics reported for
earlier artifacts.

On ambiguous bare `- item` lists (valid YAML and valid Markdown), the median
top-1 probability drops from 0.92 (pre-fine-tune) to 0.51, while YAML or
Markdown remains the top prediction.

Most remaining confusion sits on genuinely ambiguous pairs where the teacher
also splits its probability on the confused files: `c`/`cpp`,
`javascript`/`typescript`, `markdown`/`yaml`, `ini`/`toml`, `batch`/`shell`,
and `php`/`html`. Several other cells are corpus extension-label noise (for
example `.vb` files containing SQL dumps) where the teacher agrees with the
model on 80%+ of the confused files.

The README confusion matrix groups the same held-out split by file-size bucket.

Expand All @@ -42,7 +84,10 @@ The README confusion matrix groups the same held-out split by file-size bucket.
- Very short inputs are intentionally rejected when fewer than eight
non-whitespace bytes are available.
- Ambiguous snippets can put several languages close together even when a human
can infer the language from file naming context.
can infer the language from file naming context. A bare `- item` list with no
heading is valid YAML and valid Markdown; the model reports a split
YAML/Markdown distribution for such inputs rather than picking one with
certainty.
- The classifier uses content only. It does not inspect file names, extensions,
shebangs outside the model window, repository metadata, or build-system
context.
Expand Down
19 changes: 10 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,20 +40,21 @@ one-to-one with no label aggregation.

The confusion matrix uses the same labels:

![Betlang wordseq confusion](https://raw.githubusercontent.com/ealmloff/betlang/92da743c8c97fd11bb57645ec394371ee7cf836f/assets/confusion-overall.png)
![Betlang wordseq confusion](https://raw.githubusercontent.com/ealmloff/betlang/issue5-ova-calibration/assets/confusion-overall.png)

## Model

The embedded model is `assets/magika/source-student-q4.bin`, a 47,840-byte
weights-only MSQ1 payload with SHA-256:

```text
59ef24167bddd1364eb9c1650add8a67e1a542b5155fac67f5e1cda07df0c0f0
8493d2d3757572c8661141e414b1c0755aa08d4c4e5382dfbbc6b73b02d89083
```

Architecture: `wordseq-b1024-k3-m2048-tiny-3conv-hidden`, tokenizer version 3.
On the manifest-aligned held-out filesystem-label test split it reaches
`test_fs_accuracy=0.965238` with `macro_recall=0.965411`.
On the held-out filesystem-label test split it reaches
`test_fs_accuracy=0.942353` with `macro_recall=0.939690`. Probabilities are
calibrated: ambiguous inputs report split scores instead of a confident label.

See [MODEL_CARD.md](MODEL_CARD.md) for the training and evaluation summary.

Expand All @@ -75,9 +76,9 @@ Apache-2.0. Keep this attribution with redistributed model artifacts.

## Confusion By File Size

The shipped wordseq model is evaluated below on the held-out `bigorig` test
split. Each panel is a row-normalized confusion matrix for one file-size
bucket: actual labels are rows, predicted labels are columns, and the diagonal
is correct classification.
The shipped wordseq model is evaluated below on the held-out test split. Each
panel is a row-normalized confusion matrix for one file-size bucket: actual
labels are rows, predicted labels are columns, and the diagonal is correct
classification.

![Betlang wordseq confusion by file size](https://raw.githubusercontent.com/ealmloff/betlang/92da743c8c97fd11bb57645ec394371ee7cf836f/assets/confusion-by-size.png)
![Betlang wordseq confusion by file size](https://raw.githubusercontent.com/ealmloff/betlang/issue5-ova-calibration/assets/confusion-by-size.png)
2 changes: 2 additions & 0 deletions _typos.toml
Original file line number Diff line number Diff line change
Expand Up @@ -9,3 +9,5 @@ extend-exclude = [
[default.extend-words]
cpy = "cpy"
edn = "edn"
# Verilog data-out signal name used by the synthetic hard-pair generator.
dout = "dout"
Binary file modified assets/confusion-by-size.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file modified assets/confusion-overall.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file modified assets/magika/source-student-q4.bin
Binary file not shown.
32 changes: 27 additions & 5 deletions examples/detect.rs
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@

use std::collections::HashMap;
use std::fs;
use std::io::{self, Read};
use std::io::{self, Read, Seek, SeekFrom};
use std::path::{Path, PathBuf};
use std::process::ExitCode;

Expand Down Expand Up @@ -58,8 +58,8 @@ fn detect_stdin() -> ExitCode {
}

fn detect_file(path: &Path) -> ExitCode {
let bytes = match fs::read(path) {
Ok(bytes) => bytes,
let (bytes, _) = match read_model_window(path) {
Ok(window) => window,
Err(err) => {
eprintln!("betlang: failed to read {}: {err}", path.display());
return ExitCode::from(2);
Expand All @@ -68,6 +68,29 @@ fn detect_file(path: &Path) -> ExitCode {
report_single(betlang::detect(bytes))
}

/// The model only inspects the first and last 4096 bytes of a file, so read
/// just those two windows instead of the whole file. `build_window` derives
/// its begin/end blocks from `source[..4096]` and `source[len - 4096..]`,
/// which the concatenated head + tail preserves exactly. Returns the window
/// bytes and the file's real size.
fn read_model_window(path: &Path) -> io::Result<(Vec<u8>, u64)> {
const BLOCK: u64 = 4096;

let mut file = fs::File::open(path)?;
let size = file.metadata()?.len();
if size <= 2 * BLOCK {
let mut bytes = Vec::with_capacity(size as usize);
file.read_to_end(&mut bytes)?;
return Ok((bytes, size));
}

let mut bytes = vec![0u8; (2 * BLOCK) as usize];
file.read_exact(&mut bytes[..BLOCK as usize])?;
file.seek(SeekFrom::End(-(BLOCK as i64)))?;
file.read_exact(&mut bytes[BLOCK as usize..])?;
Ok((bytes, size))
}

fn report_single(detection: betlang::Detection) -> ExitCode {
match detection.language() {
Some(language) => {
Expand Down Expand Up @@ -247,8 +270,7 @@ fn display_name(path: &Path, root: &Path, depth: usize) -> String {
}

fn classify_file(path: &Path) -> Option<(Option<Language>, u64)> {
let bytes = fs::read(path).ok()?;
let size = bytes.len() as u64;
let (bytes, size) = read_model_window(path).ok()?;
let language = betlang::detect(bytes).language();
Some((language, size))
}
Expand Down
Loading
Loading