AI News (2026/1/23): Qwen3-TTS Series Open-Sourced!
Executive Summary:
The Qwen3-TTS series of speech synthesis models, developed by the Tongyi Qwen team, has been officially open-sourced. The series includes models with 1.7B and 0.6B parameters, supporting voice cloning, voice creation, and human-like speech synthesis. It uses an innovative 12Hz multi-codebook speech encoder and a dual-track modeling architecture to achieve efficient speech compression and high-fidelity restoration, with the first audio packet delay as low as 97 milliseconds. The models cover 10 major languages and dialects, including Chinese, English, Japanese, and Korean, and support precise control of voice, emotion, and prosody through natural language commands.
Qwen3-TTS Series Open-Sourced!
January 23, Friday
Details
Tongyi Qwen team announced on January 23 that they have officially open-sourced the Qwen3-TTS series of speech synthesis models. The series includes two versions: 1.7B parameters and 0.6B parameters, designed to meet the needs of different application scenarios. Qwen3-TTS not only supports voice cloning and voice creation but can also generate human-like speech, providing users with a more natural and realistic auditory experience.
Key Points
Multi-Codebook Speech Encoder: Qwen3-TTS uses an innovative 12Hz multi-codebook speech encoder, which can efficiently compress speech signals. This encoder maintains high fidelity while significantly reducing the computational resource requirements, enabling the model to run efficiently on various devices.
Dual-Track Modeling Architecture: The model employs a dual-track modeling architecture to handle acoustic and prosodic features separately. This architecture not only enhances the naturalness and fluency of the generated speech but also allows for seamless switching between different languages and dialects. The first audio packet delay is as low as 97 milliseconds, greatly improving user experience.
Multilingual Support: Qwen3-TTS covers 10 major languages and dialects, including Chinese, English, Japanese, and Korean, and supports precise control of voice, emotion, and prosody through natural language commands. Users can adjust the style of the generated speech with simple text commands to better fit specific scenarios.
AI-ALL In-Depth Commentary
The open-sourcing of Qwen3-TTS marks another significant advancement by the Tongyi Qwen team in the field of speech synthesis. The combination of the multi-codebook speech encoder and the dual-track modeling architecture not only improves the model's performance and efficiency but also provides developers with more flexibility and control. The release of this series of models will further promote the application of AI speech technology in multilingual environments, especially in scenarios requiring high customization and natural interaction. For AI developers, the open-sourcing of Qwen3-TTS means they can more easily integrate high-quality speech synthesis capabilities into their projects, thereby enhancing product competitiveness and user experience.
