I am Wang Chengyao. You may find more information here
🔭 Interested in anything towards human-like Multimodal Intelligence.
☕️ Feel free to buy me a coffee and talk about anything you want.
📧 Contact me via wcy1122@link.cuhk.edu.hk
I am Wang Chengyao. You may find more information here
🔭 Interested in anything towards human-like Multimodal Intelligence.
☕️ Feel free to buy me a coffee and talk about anything you want.
📧 Contact me via wcy1122@link.cuhk.edu.hk
Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models (ECCV 2024)
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
[ICCV 2025] Official Implementation for "Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition"
Pointcept: Perceive the world with sparse points, a codebase for point cloud perception research. Latest works: Utonia (ICML'26), Concerto (NeurIPS'25), Sonata (CVPR'25 Highlight), PTv3 (CVPR'24 Oral)
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing