Amazon Launches New ASR System Supporting Over 100 Languages


Alibaba has fully launched the Qwen-Audio-3.0 series of speech models on the Tongyi AI platform, including three types: speech recognition, speech synthesis, and real-time voice interaction. Officially, the series performed exceptionally well in the July 2024 speech ranking list by the authoritative evaluation platform Artificial Analysis, securing first place in all three categories: speech recognition, real-time interaction, and speech synthesis.
Qwen introduces the real-time speech recognition model Fun-ASR-Realtime, reducing the first-word latency to the millisecond level and achieving a smooth "speak and feedback immediately" interaction. Its recognition accuracy is close to that of offline models, achieving high precision while breaking through the real-time performance bottleneck, marking a new height in voice interaction experience.
iFlytek launched its AI hardware and software integrated solution at the 2025 1024 Developer Festival. By deeply integrating algorithms and hardware, it solves recognition challenges in complex environments such as high noise and far-field conditions, improving the accuracy of voice and visual intelligence, marking a significant breakthrough in this field.
Recently, Tongyi Lab of Alibaba officially released its latest end-to-end speech recognition large model - FunAudio-ASR. The biggest highlight of this model is its innovative "Context Module," which significantly improves the accuracy of speech recognition in high-noise environments. The hallucination rate has been reduced from 78.5% to 10.7%, a decrease of nearly 70%. This technological breakthrough has set a new benchmark for the speech recognition industry, especially suitable for noisy environments such as meetings and public places. FunAudio-AS
Recently, the OpenAI Evals tool received a significant and exciting update, adding native audio input and evaluation features. This innovation means that developers can now evaluate speech recognition and generation models directly using audio files, without going through the cumbersome process of text transcription. This change greatly simplifies the evaluation process, making the development of audio applications more efficient. In previous evaluation processes, developers often needed to first convert audio content into text, which was time-consuming and labor-intensive, and could also affect