Recently, the Alibaba Qwen team officially launched the new real-time speech synthesis model Qwen-Audio-3.0-TTS, pushing speech synthesis from "being able to speak" to "being able to express." The Plus version has ranked first globally in the Speech Arena, a prestigious global list by Artificial Analysis, outperforming mainstream models such as Gemini 3.1 TTS and ElevenLabs v3.

This release includes two versions: the Flash version is designed for real-time interaction with an initial latency of about 300ms, suitable for low-latency scenarios like smart assistants; the Plus version focuses on high-quality generation, offering better naturalness and voice similarity.
Four major breakthroughs in core capabilities:
First, multilingual and dialect coverage has been upgraded, supporting 16 languages, with the Plus version achieving an average speaker similarity of 82.75% across all 16 languages, ranking first in the industry. It also supports 20 Chinese dialects, avoiding the issue of "weakened dialect characteristics," making it more in line with native speakers' expressions.
Second, support for free-style natural language instructions allows users to generate corresponding speech accurately using natural language instructions like "gentle customer service tone" or "live-streaming host style," without requiring professional annotations.
Third, fine-grained tag control supports structured tags such as [gasp] and [angry], enabling precise control over non-verbal details like breathing and laughter, suitable for scenarios such as games and audiobooks.
Fourth, strong acoustic robustness, even if the reference audio contains high noise or reverberation, it can automatically filter out background noise while preserving voice quality, ensuring stable synthesis results.
The accompanying premium voice library covers various types of voices, including instructions, dialects, and small languages, supports 48K high-definition audio output (expected to be open on July 24th), and can synthesize up to 3 minutes of long text in a single session. The model is now fully available on the Alibaba Cloud BaiLian platform, allowing developers to access and experience it, jointly exploring new possibilities in voice interaction.

