On September 23, the Qwen large model launched the Qwen-Audio-3.1 series of audio large models, forming a complete audio capability stack covering understanding, generation, interaction, and creation. The series includes speech recognition Qwen-Audio-3.1-ASR, audio understanding ASR-Next, text-to-speech TTS, audio creation TTS-Next, and real-time interaction Realtime. The API has been listed on the Qwen AI platform, and the ASR-Next API will be launched soon. At the same time, all products are being discounted, with TTS dropping by about 70%, Realtime by about 85%, and ASR by up to 95%.

QQ20260923-152735.jpg

In terms of capabilities, ASR enhances multilingual and dialect recognition as well as context understanding. It now supports native transcription polishing, automatically removing filler words and repeated expressions, and supports end-to-end role-based transcription, jointly outputting speaker tags, timestamps, and text; one model supports 30 languages and 16 Chinese dialects, with a first-word response time of about 160 milliseconds for streaming recognition.

ASR-Next is based on the next-generation architecture, expanding audio understanding from text in speech to emotions, ambient sounds, and mechanical sounds, supporting voice description, event localization, and audio question-answering reasoning. TTS supports natural cross-language voice color migration and controls emotion and speaking speed through instructions. TTS-Next is aimed at audio creation, unifying the generation of human voice, sound effects, and ambient sounds, supporting multi-role voice replication, fine-grained timestamps, and 48kHz output, serving scenarios such as podcasts, audiobooks, and film and game applications. Realtime realizes full-duplex interaction based on a multi-teacher distillation architecture, allowing interruptions at any time, supporting real-time language switching, emotion understanding, and tool calls during dialogue.

In testing, ASR achieved an average CER of 4.55% on open-source dialect test sets, an average CER of 10.38% on self-built Chinese dialect sets, and an average semantic sentence accuracy of 82.10% for converting dialects to Mandarin; ASR-Flash-Next and ASR-Flash ranked first in five and three out of eight metrics across four test sets, comprehensively outperforming the previous Fun-ASR cascading system.

Currently, Realtime is integrated with agents like Qoder and Qwen Office, as well as smart hardware like the Qwen AI glasses, continuing the trend of voice becoming a natural interface for human-computer interaction.

  • Qwen-Audio-3.1-ASR:

https://www.qianwenai.com/models/qwen-audio-3.1-asr-flash

  • Qwen-Audio-3.1-TTS:

https://www.qianwenai.com/models/qwen-audio-3.1-tts-flash

  • Qwen-Audio-3.1-TTS-Next:

https://www.qianwenai.com/models/qwen-audio-3.1-tts-next

  • Qwen-Audio-3.1-Realtime:

https://www.qianwenai.com/models/qwen-audio-3.1-realtime-plus

(The Qwen-Audio-3.1-ASR-Next API will be launched soon)