Skip to content

About

Layer atlas of an open video diffusion transformer (Wan2.1-1.3B)

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

vdit-atlas

A layer atlas of an open video diffusion transformer: semantics live mid-depth, the arrow of time lives late.

Linear probes over the internal activations of a frozen, fully open video DiT (Wan2.1-T2V-1.3B) — all 30 blocks x 9 noise levels — plus an arrow-of-time probe and a zero-shot diffusion classifier. Total compute: ~8 GPU-hours on one RTX 3090 Ti (~2 for the UCF run, ~6 for the SSv2 replication).

Layer dissociation

TL;DR

  1. Semantics peak mid-depth. Action-classification accuracy of a linear probe over frozen DiT features follows an inverted U across depth, peaking at blocks 10-15 of 30 (sigma 0.2-0.7). Best cell: 96.2% on 10-way UCF101-subset — above VideoMAE-base (94.6%) under the identical probe protocol.
  2. The arrow of time lives late. A probe distinguishing forward from reversed videos (paired noise, video-grouped folds) reaches AUROC 0.917 with the peak at blocks 22-25 — while VideoMAE-base features sit at 0.603, barely above chance. A model trained only to generate video knows which way time flows; a strong masked-autoencoding baseline barely does.
  3. The conditional likelihood is already a classifier. Scoring the flow-matching velocity loss under 10 class prompts (shared noise) gives 46.7% top-1 / 81.3% top-3 zero-shot (chance 10%) with zero training and a naive prompt template.

Update (Aug 2026): everything replicates on Something-Something v2. Same protocol on 30 temporally-loaded SSv2 classes: the semantic peak lands in the same mid-depth band (per-block profile r = 0.945 vs UCF, now far from ceiling), the arrow-of-time peak lands in the exact same cell (block 22, sigma 0.3), and the frozen DiT again edges out VideoMAE under the identical probe. Details below.

The dissociation in (1)+(2) — what happens is most readable at ~40% depth, which way time flows at ~80% — is, to our knowledge, the first such map for an open video DiT.

Why

Video DiTs are trained only to denoise; any understanding inside them is a byproduct of generation. Recent work found strong readable semantics (Gen4U) and internal physics (The Invisible Hand of Physics) in proprietary models (Veo3), with open-weight models reported as weaker. This repo maps where understanding lives in a fully open 1.3B model, with a protocol cheap enough for a single consumer GPU — and finds the open model is stronger than that reading suggests, at least at this scale of evaluation.

Method (short version)

  • Features: forward hooks on every transformer.blocks[i]; one forward per (video, sigma). Flow-matching noising x_s = (1-s) z0 + s eps, timestep s*1000, empty-prompt conditioning (UMT5 embedding computed once and cached). Latents normalized with the VAE's latents_mean/std. Tokens are frame-major, so per-latent-frame pooling is exact.
  • Probes: logistic regression + standardization, stratified 6-fold CV, identical protocol for the DiT and the baseline (VideoMAE-base, mean-pooled last hidden state).
  • Arrow of time: same video forward vs reversed, same noise for both members of a pair (seeds are stable per video id), StratifiedGroupKFold so a video never appears in both train and test.
  • Zero-shot: velocity-loss MSE under K class prompts with shared noise (paired comparison), argmin over classes; sigmas {0.3, 0.5, 0.7}.

Results

Semantic atlas (392 videos, 10 classes, chance 0.10)

best cell acc AUROC
Wan2.1-1.3B (frozen) block 10, sigma 0.4 0.962 0.998
VideoMAE-base — 0.946 ± 0.019 0.998

28 of 270 (block, sigma) cells beat the baseline; all of them sit in the mid-depth band. Accuracy is flat across sigma 0.2-0.7 and collapses only at 0.9.

Semantic atlas

Arrow of time (200 videos x2, paired noise)

features best cell AUROC acc
DiT per-frame sequence block 22, sigma 0.3 0.917 ± 0.033 0.840
DiT global mean-pool block 22, sigma 0.3 0.893 ± 0.022 0.817
VideoMAE-base mean-pool — 0.603 0.590

Signal is distributed (every block ≥ 0.86 at sigma 0.3) but peaks late, and degrades quickly with noise (0.917 → 0.762 from sigma 0.3 → 0.7). The global mean-pool result shows time direction is encoded inside per-frame features, not just in token order.

AoT heatmap

Zero-shot diffusion classifier (75 held-out videos)

top-1 0.467, top-3 0.813 (chance 0.100), template "a video of a person {}.". Confusions are semantically structured (Basketball vs BasketballDunk accounts for a large share; merged they reach ~0.73 recall). Correct predictions have ~2x the top1-top2 loss margin of incorrect ones.

Update: replication on Something-Something v2 (Aug 2026)

Same harness, harder benchmark: 30 temporally-loaded SSv2 classes (directional pairs like pushing left→right / right→left, cover/uncover, open/close, plus pretending-to classes), 120 videos each — 3,600 videos, 30-way, chance 0.033. Appearance alone does not solve these classes; frame order does. Nothing was re-tuned.

UCF101-subset (10-way) SSv2 temporal-30 (30-way)
Semantic peak block 10, σ0.4 — acc 0.962 block 11, σ0.5 — acc 0.494 (AUROC 0.918)
VideoMAE-base, same protocol 0.946 ± 0.019 0.482 ± 0.009 (AUROC 0.911)
Cells beating the baseline 28/270 — all in blocks 10-15 18/270 — all in blocks 8-14
Arrow-of-time peak block 22, σ0.3 — AUROC 0.917 block 22, σ0.3 — AUROC 0.861 ± 0.015
VideoMAE-base arrow of time 0.603 0.585
  • The semantic map is the same map. Per-block accuracy profiles correlate at r = 0.945 across datasets, despite different classes, chance levels and video sources. And the benchmark is no longer saturated (AUROC 0.918, not 0.998), so matches-or-beats-VideoMAE now holds off-ceiling: ~15× chance with a linear probe on frozen generative features.
  • The arrow of time peaks in the exact same cell (block 22, σ0.3); the top-8 SSv2 cells all sit at σ0.3 in late blocks, with the same monotone decay in noise (0.861 → 0.820 → 0.765 for σ 0.3/0.5/0.7). VideoMAE stays near chance on both datasets (0.603 / 0.585) — a +0.28 AUROC gap.
  • Two SSv2-specific observations: the AoT per-block profile is less flat than UCF's (profile r = 0.60 — what replicates precisely is the peak cell, the dominant late band, and the sigma ordering, not the full curve), and global mean-pooling costs more here (0.861 → 0.773): on SSv2, time direction lives more in the frame sequence and less inside single-frame features — consistent with classes defined by order.

SSv2 replication

Reproduce

Windows / PowerShell (the whole pipeline, unattended, resumable):

powershell -ExecutionPolicy Bypass -File setup_env.ps1
python expA_atlas\download_data.py
powershell -ExecutionPolicy Bypass -File run_overnight.ps1   # heavy caches go to F:\ by default; override with -Work/-HfHome/-PipCache

SSv2 replication (needs the official Qualcomm download: 2 file parts + a labels zip, ~19 GB):

python expA_atlas\download_ssv2.py --extract <raw_download_dir> --to <ssv2_dir>
python expA_atlas\organize_ssv2.py --raw <ssv2_dir> --out <ssv2_dir>_temporal30 --max-per-class 120
python expA_atlas\run_extract.py --data <ssv2_dir>_temporal30 --out results\features_ssv2 --max-per-class 120
python expA_atlas\run_probes.py --features results\features_ssv2 --baseline videomae
python expA_atlas\run_arrow_of_time.py --data <ssv2_dir>_temporal30 --out results\aot_ssv2 --baseline videomae

Linux/macOS: create a venv, pip install torch (CUDA build) + pip install -r requirements.txt, then run the same expA_atlas/*.py and expC_diffusion_classifier/*.py scripts directly.

There is a no-GPU smoke mode that validates the full pipeline in ~2 minutes: python expA_atlas/make_smoke_data.py && python expA_atlas/run_extract.py --smoke --data data_smoke --out results/features_smoke.

Timings on one RTX 3090 Ti — UCF run: atlas extraction 25 min (3.9 s/video, 9 forwards each), probes 12 min, arrow of time ~35 min, zero-shot ~25 min. SSv2 run: extraction 3 h 15 (3.26 s/video, resumable — it survived an interruption at 457/3600), probes 1 h 15, arrow of time ~2 h, baseline 13 min. First run downloads ~15 GB (the UMT5-XXL text encoder is loaded once; its prompt embeddings are cached to disk).

Caveats

  1. UCF101-subset is small (10 classes) and appearance-biased, and its semantic benchmark is near saturation (AUROC 0.998 for both models). The SSv2 replication above addresses both concerns (30-way, off-ceiling). What remains true on both datasets: the margin over VideoMAE is small (+1.2 to +1.6 pts) — the shape of the atlas, not the margin, is the finding.
  2. The arrow-of-time baseline uses mean-pooled VideoMAE features; a stronger temporal readout would raise its number (though closing +0.26-0.31 AUROC gaps entirely seems unlikely).
  3. One model (Wan2.1-1.3B), one noise seed per video (stable and paired), one prompt template for zero-shot. The SSv2 subset is a hand-picked temporal-30; the full 174-class run is future work.

Related work

Gen4U · The Invisible Hand of Physics · UniVideo · V-JEPA 2 · Video models are zero-shot learners and reasoners

Cite

@misc{neuregex2026vditatlas,
  author = {neuregex},
  title  = {vdit-atlas: A layer atlas of an open video diffusion transformer},
  year   = {2026},
  url    = {https://github.com/neuregex/vdit-atlas}
}

MIT license.

About

Layer atlas of an open video diffusion transformer (Wan2.1-1.3B)

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages