Qwen3-Omni is a native omni-modal foundation model that can process diverse inputs including text, images, audio, and video, and deliver real-time streaming responses in both text and natural speech. The model supports 119 text languages, 19 speech input languages, and 10 speech output languages. It has a novel architecture with a MoE-based Thinker–Talker design and a multi-codebook design for low latency. The model is available for download and can be used with Hugging Face Transformers or vLLM for inference.