Three base models are the product:
| Base model | Size | Primary capability |
|---|---|---|
| Cosmos3-Edge | 4B | Compact omnimodal world model for on-device and low-latency deployment |
| Cosmos3-Nano | 16B | Omnimodal world model — understanding, simulation, prediction, action reasoning on a single GPU |
| Cosmos3-Super | 64B | Frontier-scale omnimodal world model |
Post-trained example checkpoints (cosmos3-examples) demonstrate fine-tuning capabilities and are not part of the product line:
| Example checkpoint | Base | Demonstrates |
|---|---|---|
| Cosmos3-Super-Text2Image | Super | Elite quality text-to-image |
| Cosmos3-Super-Text2Image-4Step | Super | Elite quality text-to-image, 17-25x faster |
| Cosmos3-Super-Image2Video | Super | Elite quality image-to-video |
| Cosmos3-Super-Image2Video-4Step | Super | Elite quality image-to-video, 17-25x faster |
| Cosmos3-Nano-Policy-DROID | Nano | Open SOTA DROID robot policy, runs on RTX Pro 6000 |
| Cosmos3-Edge-Policy-DROID | Edge | DROID robot policy at edge-deployable scale |
Cosmos 3 is built on a unified Mixture-of-Transformers (MoT) architecture combining an autoregressive transformer (reasoning; causal self-attention, next-token prediction) with a diffusion transformer (generation; full attention over noisy image/video/audio/action tokens). Both modes share transformer weights, multimodal attention layers, and a unified 3D mRoPE that encodes spatial and temporal structure across modalities.
| Workflow | Input | Output Cosmos3-Super |
Output Cosmos3-Nano |
Output Cosmos3-Edge |
|---|---|---|---|---|
| t2v | Text | Video (256p, 480p, 720p) | Video (256p, 480p, 720p) | Video (256p, 480p) |
| t2sv | Text | Video with sound (256p, 480p, 720p) | Video with sound (256p, 480p, 720p) | Not supported |
| i2v — 1st frame | Text + image (JPG/PNG/JPEG/WEBP) | Video (256p, 480p, 720p) | Video (256p, 480p, 720p) | Video (256p, 480p) |
| i2sv | Text + image (JPG/PNG/JPEG/WEBP) | Video with sound (256p, 480p, 720p) | Video with sound (256p, 480p, 720p) | Not supported |
| v2v — transfer | Text + MP4 video (edge, blur, depth, segmentation map) | Video (256p, 480p, 720p) | Video (256p, 480p, 720p) | Not supported |
| v2v — predict | Text + MP4 video (first 5 frames used, up to ~3 s of conditioning) | Video (256p, 480p, 720p) | Video (256p, 480p, 720p) | Not supported |
| t2i | Text string | Image (256p, 480p, 720p) | Image (256p, 480p, 720p) | Not supported |
| Visual reasoning | Text, MP4 video, image (JPG/PNG/JPEG/WEBP) | Text | Text | Text |
| Forward dynamics | Text, image, video, action | Video (256p, 480p, 720p) | Video (256p, 480p, 720p) | Video (256p, 480p) |
| Inverse dynamics | Text, video | Action | Action | Action |
| Action policy | Text, image, video | Action, video | Action, video | Action, video |
* 720p uses 1280×720, 480p uses 832×480, and 256p uses 320×192.
* Video conditioning (v2v — predict) uses 5 frames at the matching resolution.
* Action conditioning: camera motion (9D) · AV (9D) · egocentric (57D) · single-arm (10D: DROID/UR/Fractal/Bridge/UMI) · dual-arm (20D) · humanoid (29D: AgiBot)
* Output format: JPG, MP4, AAC-in-MP4 (stereo 48 kHz), JSON actions, text
* Prompt length: < 300 words recommended for world generation
Every workflow has a runnable notebook — see cookbooks.
| Setting | Supported values |
|---|---|
| Resolution tiers | 256p, 480p, 720p (default 480p) |
| Aspect ratios | 16:9, 4:3, 1:1, 3:4, 9:16 (default 16:9) |
| Frame rates | 10, 16, 24, 30 FPS (default 24) |
| Frame count | 5–300 (default 189) |
| Precision | BF16 tested |
| OS / GPU | Linux; Ampere, Hopper, Blackwell |
* Cosmos3-Edge supports only 256p and 480p, frame count is 50-150.
Generator prompt upsampling: max_tokens=20000, temperature=0.7, top_p=0.8, top_k=20, repetition_penalty=1.0, presence_penalty=1.5, seed=3407.
Reasoner:
| Parameter | Without reasoning | With reasoning |
|---|---|---|
temperature |
0.7 | 0.6 |
top_p |
0.8 | 0.95 |
top_k |
20 | 20 |
presence_penalty |
1.5 | 0.0 |
repetition_penalty |
1.0 | 1.0 |
Cosmos 3 can produce artifacts in long, high-resolution, or physically complex outputs: temporal inconsistency, unstable camera/object motion, sound-video misalignment, action-state inconsistency, object morphing, inaccurate 3D structure, implausible dynamics. Physically grounded simulation, safety-critical control, and complex multi-agent behavior require additional validation, guardrails, and system-level safety analysis before deployment.
