Hi, thank you for releasing Cosmos 3.
The paper mentions that the video captioner is a LoRA-finetuned Qwen3-VL-8B model used to generate structured training captions.
Would it be possible to release the fine-tuned video captioner checkpoint or LoRA weights? It would be very helpful for reproducing the Cosmos 3 data annotation pipeline.
Thank you!
Hi, thank you for releasing Cosmos 3.
The paper mentions that the video captioner is a LoRA-finetuned Qwen3-VL-8B model used to generate structured training captions.
Would it be possible to release the fine-tuned video captioner checkpoint or LoRA weights? It would be very helpful for reproducing the Cosmos 3 data annotation pipeline.
Thank you!