The arms race in the voice field by large model manufacturers has added another heavyweight player. On July 20, Qwen officially launched the audio synthesis large model Qwen-Audio-3.0-TTS, returning the control of speaking to a natural language command, instead of tedious parameter settings.
The most distinctive feature of this new model is Free-style natural language instruction control. In the past, to make AI speak with a certain tone, rhythm, or style, it often required adjusting a series of engineering parameters in the background; now, users just need to describe the desired effect in everyday language, and the model can synthesize the voice according to the instructions, greatly lowering the barrier.
More importantly, Qwen has prepared two versions for different scenarios. The Flash version for real-time interaction reduces the first packet delay to the 300-millisecond level, which means that voice responses in conversations are almost imperceptible, sufficient to support real-time Q&A, voice assistants, and other scenarios that are extremely sensitive to latency. The Plus version for high-quality generation focuses on sound quality and expressiveness, targeting tasks such as audiobooks, dubbing, and content creation that require subtle voice texture. A single product line simultaneously covers both fast and high-quality demand curves that were previously conflicting.
When natural language becomes the remote control for voice, Qwen has taken another step forward, pushing voice synthesis from an engineering task into the toolbox of ordinary people.
