Currently, the text and vision capabilities support >30 languages, but the speech conversation and full-duplex features are strictly limited to English and Chinese. I am writing to request/suggest expanding the speech modality to include Italian (and potentially other major European languages).
Technical Context & Bottlenecks
From my understanding of the architecture, the foundation for multilingual support is already there, but the speech pipeline is bottlenecked by a couple of specific modules:
- Audio Encoder (Whisper Medium - 307M) & LLM Backbone (Qwen): Both already possess strong native multilingual capabilities, including Italian.
- Audio Projector (2-layer MLP - 21M): It seems this projector was trained/aligned primarily on EN/ZH audio-text pairs during the Omni Pretraining and SFT phases.
- Speech Decoder/Talker (ChatTTS / CosyVoice2): These modules are currently highly optimized for EN/ZH, lacking expressive voice clones or native support for Italian.
Proposal / Questions for the Devs
To bring Italian/multilingual speech to MiniCPM-o, what would be the most efficient path forward for the community?
- Are there any plans from OpenBMB to release an updated 21M Audio Projector aligned with a broader multilingual speech dataset (e.g., Common Voice)?
- Is there a recommended pipeline/script to fine-tune the 21M Projector locally for our own languages?
- Would swapping the current Talker/TTS backend with a more multilingual-friendly model (like XTTS or out-of-the-box multilingual CosyVoice) break the Omni-Flow end-to-end architecture?
Use Case
Expanding speech to Italian would massively increase the adoption of MiniCPM-o in European edge-device projects, robotics, and local AI assistants where local language Nuance is critical.
Thank you for your time and for open-sourcing such a masterpiece!
Currently, the text and vision capabilities support >30 languages, but the speech conversation and full-duplex features are strictly limited to English and Chinese. I am writing to request/suggest expanding the speech modality to include Italian (and potentially other major European languages).
Technical Context & Bottlenecks
From my understanding of the architecture, the foundation for multilingual support is already there, but the speech pipeline is bottlenecked by a couple of specific modules:
Proposal / Questions for the Devs
To bring Italian/multilingual speech to MiniCPM-o, what would be the most efficient path forward for the community?
Use Case
Expanding speech to Italian would massively increase the adoption of MiniCPM-o in European edge-device projects, robotics, and local AI assistants where local language Nuance is critical.
Thank you for your time and for open-sourcing such a masterpiece!