A layer atlas of an open video diffusion transformer: semantics live mid-depth, the arrow of time lives late.
Linear probes over the internal activations of a frozen, fully open video DiT (Wan2.1-T2V-1.3B) — all 30 blocks x 9 noise levels — plus an arrow-of-time probe and a zero-shot diffusion classifier. Total compute: ~8 GPU-hours on one RTX 3090 Ti (~2 for the UCF run, ~6 for the SSv2 replication).
- Semantics peak mid-depth. Action-classification accuracy of a linear probe over frozen DiT features follows an inverted U across depth, peaking at blocks 10-15 of 30 (sigma 0.2-0.7). Best cell: 96.2% on 10-way UCF101-subset — above VideoMAE-base (94.6%) under the identical probe protocol.
- The arrow of time lives late. A probe distinguishing forward from reversed videos (paired noise, video-grouped folds) reaches AUROC 0.917 with the peak at blocks 22-25 — while VideoMAE-base features sit at 0.603, barely above chance. A model trained only to generate video knows which way time flows; a strong masked-autoencoding baseline barely does.
- The conditional likelihood is already a classifier. Scoring the flow-matching velocity loss under 10 class prompts (shared noise) gives 46.7% top-1 / 81.3% top-3 zero-shot (chance 10%) with zero training and a naive prompt template.
Update (Aug 2026): everything replicates on Something-Something v2. Same protocol on 30 temporally-loaded SSv2 classes: the semantic peak lands in the same mid-depth band (per-block profile r = 0.945 vs UCF, now far from ceiling), the arrow-of-time peak lands in the exact same cell (block 22, sigma 0.3), and the frozen DiT again edges out VideoMAE under the identical probe. Details below.
The dissociation in (1)+(2) — what happens is most readable at ~40% depth, which way time flows at ~80% — is, to our knowledge, the first such map for an open video DiT.
Video DiTs are trained only to denoise; any understanding inside them is a byproduct of generation. Recent work found strong readable semantics (Gen4U) and internal physics (The Invisible Hand of Physics) in proprietary models (Veo3), with open-weight models reported as weaker. This repo maps where understanding lives in a fully open 1.3B model, with a protocol cheap enough for a single consumer GPU — and finds the open model is stronger than that reading suggests, at least at this scale of evaluation.
- Features: forward hooks on every
transformer.blocks[i]; one forward per (video, sigma). Flow-matching noisingx_s = (1-s) z0 + s eps, timesteps*1000, empty-prompt conditioning (UMT5 embedding computed once and cached). Latents normalized with the VAE'slatents_mean/std. Tokens are frame-major, so per-latent-frame pooling is exact. - Probes: logistic regression + standardization, stratified 6-fold CV, identical protocol for the DiT and the baseline (VideoMAE-base, mean-pooled last hidden state).
- Arrow of time: same video forward vs reversed, same noise for both members of a pair (seeds are stable per video id), StratifiedGroupKFold so a video never appears in both train and test.
- Zero-shot: velocity-loss MSE under K class prompts with shared noise (paired comparison), argmin over classes; sigmas {0.3, 0.5, 0.7}.
| best cell | acc | AUROC | |
|---|---|---|---|
| Wan2.1-1.3B (frozen) | block 10, sigma 0.4 | 0.962 | 0.998 |
| VideoMAE-base | — | 0.946 ± 0.019 | 0.998 |
28 of 270 (block, sigma) cells beat the baseline; all of them sit in the mid-depth band. Accuracy is flat across sigma 0.2-0.7 and collapses only at 0.9.
| features | best cell | AUROC | acc |
|---|---|---|---|
| DiT per-frame sequence | block 22, sigma 0.3 | 0.917 ± 0.033 | 0.840 |
| DiT global mean-pool | block 22, sigma 0.3 | 0.893 ± 0.022 | 0.817 |
| VideoMAE-base mean-pool | — | 0.603 | 0.590 |
Signal is distributed (every block ≥ 0.86 at sigma 0.3) but peaks late, and degrades quickly with noise (0.917 → 0.762 from sigma 0.3 → 0.7). The global mean-pool result shows time direction is encoded inside per-frame features, not just in token order.
top-1 0.467, top-3 0.813 (chance 0.100), template "a video of a person {}.".
Confusions are semantically structured (Basketball vs BasketballDunk accounts
for a large share; merged they reach ~0.73 recall). Correct predictions have ~2x
the top1-top2 loss margin of incorrect ones.
Same harness, harder benchmark: 30 temporally-loaded SSv2 classes (directional pairs like pushing left→right / right→left, cover/uncover, open/close, plus pretending-to classes), 120 videos each — 3,600 videos, 30-way, chance 0.033. Appearance alone does not solve these classes; frame order does. Nothing was re-tuned.
| UCF101-subset (10-way) | SSv2 temporal-30 (30-way) | |
|---|---|---|
| Semantic peak | block 10, σ0.4 — acc 0.962 | block 11, σ0.5 — acc 0.494 (AUROC 0.918) |
| VideoMAE-base, same protocol | 0.946 ± 0.019 | 0.482 ± 0.009 (AUROC 0.911) |
| Cells beating the baseline | 28/270 — all in blocks 10-15 | 18/270 — all in blocks 8-14 |
| Arrow-of-time peak | block 22, σ0.3 — AUROC 0.917 | block 22, σ0.3 — AUROC 0.861 ± 0.015 |
| VideoMAE-base arrow of time | 0.603 | 0.585 |
- The semantic map is the same map. Per-block accuracy profiles correlate at r = 0.945 across datasets, despite different classes, chance levels and video sources. And the benchmark is no longer saturated (AUROC 0.918, not 0.998), so matches-or-beats-VideoMAE now holds off-ceiling: ~15× chance with a linear probe on frozen generative features.
- The arrow of time peaks in the exact same cell (block 22, σ0.3); the top-8 SSv2 cells all sit at σ0.3 in late blocks, with the same monotone decay in noise (0.861 → 0.820 → 0.765 for σ 0.3/0.5/0.7). VideoMAE stays near chance on both datasets (0.603 / 0.585) — a +0.28 AUROC gap.
- Two SSv2-specific observations: the AoT per-block profile is less flat than UCF's (profile r = 0.60 — what replicates precisely is the peak cell, the dominant late band, and the sigma ordering, not the full curve), and global mean-pooling costs more here (0.861 → 0.773): on SSv2, time direction lives more in the frame sequence and less inside single-frame features — consistent with classes defined by order.
Windows / PowerShell (the whole pipeline, unattended, resumable):
powershell -ExecutionPolicy Bypass -File setup_env.ps1
python expA_atlas\download_data.py
powershell -ExecutionPolicy Bypass -File run_overnight.ps1 # heavy caches go to F:\ by default; override with -Work/-HfHome/-PipCacheSSv2 replication (needs the official Qualcomm download: 2 file parts + a labels zip, ~19 GB):
python expA_atlas\download_ssv2.py --extract <raw_download_dir> --to <ssv2_dir>
python expA_atlas\organize_ssv2.py --raw <ssv2_dir> --out <ssv2_dir>_temporal30 --max-per-class 120
python expA_atlas\run_extract.py --data <ssv2_dir>_temporal30 --out results\features_ssv2 --max-per-class 120
python expA_atlas\run_probes.py --features results\features_ssv2 --baseline videomae
python expA_atlas\run_arrow_of_time.py --data <ssv2_dir>_temporal30 --out results\aot_ssv2 --baseline videomaeLinux/macOS: create a venv, pip install torch (CUDA build) +
pip install -r requirements.txt, then run the same expA_atlas/*.py and
expC_diffusion_classifier/*.py scripts directly.
There is a no-GPU smoke mode that validates the full pipeline in ~2 minutes:
python expA_atlas/make_smoke_data.py && python expA_atlas/run_extract.py --smoke --data data_smoke --out results/features_smoke.
Timings on one RTX 3090 Ti — UCF run: atlas extraction 25 min (3.9 s/video, 9 forwards each), probes 12 min, arrow of time ~35 min, zero-shot ~25 min. SSv2 run: extraction 3 h 15 (3.26 s/video, resumable — it survived an interruption at 457/3600), probes 1 h 15, arrow of time ~2 h, baseline 13 min. First run downloads ~15 GB (the UMT5-XXL text encoder is loaded once; its prompt embeddings are cached to disk).
- UCF101-subset is small (10 classes) and appearance-biased, and its semantic benchmark is near saturation (AUROC 0.998 for both models). The SSv2 replication above addresses both concerns (30-way, off-ceiling). What remains true on both datasets: the margin over VideoMAE is small (+1.2 to +1.6 pts) — the shape of the atlas, not the margin, is the finding.
- The arrow-of-time baseline uses mean-pooled VideoMAE features; a stronger temporal readout would raise its number (though closing +0.26-0.31 AUROC gaps entirely seems unlikely).
- One model (Wan2.1-1.3B), one noise seed per video (stable and paired), one prompt template for zero-shot. The SSv2 subset is a hand-picked temporal-30; the full 174-class run is future work.
Gen4U · The Invisible Hand of Physics · UniVideo · V-JEPA 2 · Video models are zero-shot learners and reasoners
@misc{neuregex2026vditatlas,
author = {neuregex},
title = {vdit-atlas: A layer atlas of an open video diffusion transformer},
year = {2026},
url = {https://github.com/neuregex/vdit-atlas}
}MIT license.



