Microsoft officially released the new speech recognition model MAI-Transcribe-2 on September 3. As the most powerful and efficient transcription product to date, this model achieved an average word error rate (WER) of 5.2% in the FLEURS benchmark test (covering 60 languages), and ranked second in the Artificial Analysis word error rate ranking.

QQ20260904-092216.jpg

While maintaining high accuracy, its long audio processing speed can be up to 10 times faster than that of competitors, with a processing time of about 10 seconds for one hour of audio. The model is now available on Microsoft Foundry, MAI Playground, and Open Router, and offers a limited-time discounted price of just $0.10 per hour of audio.

MAI-Transcribe-2 has been comprehensively upgraded for real-world complex scenarios, integrating features such as speaker segmentation, word-level timestamps, keyword preferences, and automatic language detection. It supports code-switching phenomena such as Hinglish (Indian English) and Spanglish (Spanish-English mix), and has strong noise resistance. Developers can also freely switch between "word-for-word transcription" and "concise transcription" styles to accurately adapt to diverse production environments such as legal, clinical, subtitles, and high-compliance analysis.

QQ20260904-092235.jpg

As AI applications evolve from single-modal interaction to multi-modal and real-time speech intelligence, high-throughput, low-latency, and low-cost speech infrastructure is becoming a key competitive barrier. MAI-Transcribe-2 redefines the accuracy and latency Pareto frontier of speech transcription, significantly lowering the deployment threshold for multilingual audio, and injecting new momentum into the development of end-to-end real-time speech ecosystems.