Skip to content

Latest commit

 

History

History
94 lines (73 loc) · 6.91 KB

File metadata and controls

94 lines (73 loc) · 6.91 KB

Model Reference

Family

Three base models are the product:

Base model Size Primary capability
Cosmos3-Edge 4B Compact omnimodal world model for on-device and low-latency deployment
Cosmos3-Nano 16B Omnimodal world model — understanding, simulation, prediction, action reasoning on a single GPU
Cosmos3-Super 64B Frontier-scale omnimodal world model

Post-trained example checkpoints (cosmos3-examples) demonstrate fine-tuning capabilities and are not part of the product line:

Example checkpoint Base Demonstrates
Cosmos3-Super-Text2Image Super Elite quality text-to-image
Cosmos3-Super-Text2Image-4Step Super Elite quality text-to-image, 17-25x faster
Cosmos3-Super-Image2Video Super Elite quality image-to-video
Cosmos3-Super-Image2Video-4Step Super Elite quality image-to-video, 17-25x faster
Cosmos3-Nano-Policy-DROID Nano Open SOTA DROID robot policy, runs on RTX Pro 6000
Cosmos3-Edge-Policy-DROID Edge DROID robot policy at edge-deployable scale

Architecture

Cosmos 3 is built on a unified Mixture-of-Transformers (MoT) architecture combining an autoregressive transformer (reasoning; causal self-attention, next-token prediction) with a diffusion transformer (generation; full attention over noisy image/video/audio/action tokens). Both modes share transformer weights, multimodal attention layers, and a unified 3D mRoPE that encodes spatial and temporal structure across modalities.

Cosmos 3 model architecture

Input / output and workflows

Workflow Input Output
Cosmos3-Super
Output
Cosmos3-Nano
Output
Cosmos3-Edge
t2vTextVideo
(256p, 480p, 720p)
Video
(256p, 480p, 720p)
Video
(256p, 480p)
t2svTextVideo with sound
(256p, 480p, 720p)
Video with sound
(256p, 480p, 720p)
Not supported
i2v — 1st frameText + image
(JPG/PNG/JPEG/WEBP)
Video
(256p, 480p, 720p)
Video
(256p, 480p, 720p)
Video
(256p, 480p)
i2svText + image
(JPG/PNG/JPEG/WEBP)
Video with sound
(256p, 480p, 720p)
Video with sound
(256p, 480p, 720p)
Not supported
v2v — transferText + MP4 video
(edge, blur, depth, segmentation map)
Video
(256p, 480p, 720p)
Video
(256p, 480p, 720p)
Not supported
v2v — predictText + MP4 video
(first 5 frames used, up to ~3 s of conditioning)
Video
(256p, 480p, 720p)
Video
(256p, 480p, 720p)
Not supported
t2iText stringImage
(256p, 480p, 720p)
Image
(256p, 480p, 720p)
Not supported
Visual reasoningText, MP4 video, image
(JPG/PNG/JPEG/WEBP)
TextTextText
Forward dynamicsText, image, video, actionVideo
(256p, 480p, 720p)
Video
(256p, 480p, 720p)
Video
(256p, 480p)
Inverse dynamicsText, videoActionActionAction
Action policyText, image, videoAction, videoAction, videoAction, video

* 720p uses 1280×720, 480p uses 832×480, and 256p uses 320×192.
* Video conditioning (v2v — predict) uses 5 frames at the matching resolution.
* Action conditioning: camera motion (9D) · AV (9D) · egocentric (57D) · single-arm (10D: DROID/UR/Fractal/Bridge/UMI) · dual-arm (20D) · humanoid (29D: AgiBot)
* Output format: JPG, MP4, AAC-in-MP4 (stereo 48 kHz), JSON actions, text
* Prompt length: < 300 words recommended for world generation

Every workflow has a runnable notebook — see cookbooks.

Generation settings

Setting Supported values
Resolution tiers 256p, 480p, 720p (default 480p)
Aspect ratios 16:9, 4:3, 1:1, 3:4, 9:16 (default 16:9)
Frame rates 10, 16, 24, 30 FPS (default 24)
Frame count 5–300 (default 189)
Precision BF16 tested
OS / GPU Linux; Ampere, Hopper, Blackwell

* Cosmos3-Edge supports only 256p and 480p, frame count is 50-150.

Sampling defaults

Generator prompt upsampling: max_tokens=20000, temperature=0.7, top_p=0.8, top_k=20, repetition_penalty=1.0, presence_penalty=1.5, seed=3407.

Reasoner:

Parameter Without reasoning With reasoning
temperature 0.7 0.6
top_p 0.8 0.95
top_k 20 20
presence_penalty 1.5 0.0
repetition_penalty 1.0 1.0

Limitations

Cosmos 3 can produce artifacts in long, high-resolution, or physically complex outputs: temporal inconsistency, unstable camera/object motion, sound-video misalignment, action-state inconsistency, object morphing, inaccurate 3D structure, implausible dynamics. Physically grounded simulation, safety-critical control, and complex multi-agent behavior require additional validation, guardrails, and system-level safety analysis before deployment.