Tencent Hunyuan has officially released the new generation speech recognition model Hy ASR3.0 preview. It leverages the latest Hy3 large language model's language understanding capabilities, combining high-precision speech recognition with deep semantic understanding. It has significantly improved in core dimensions such as general recognition, context awareness, multi-scenario robustness, and dialect coverage, providing more accurate, coherent, and user-intent-aligned transcription results in more complex real-world inputs — as the official description says, it has evolved from "word-by-word transcription and single-point optimization" to "context understanding, scenario compatibility, and one-click output."

image.png

The performance is impressive. In overseas open-source evaluation sets, Hy ASR3.0 preview has reduced the word error rate (WER) of multiple languages to around 3%: 3.34% for Mandarin Chinese, 2.62% for English, and 3.12% for Cantonese. On Tencent's own built evaluation set, WER remains low in various scenarios including general recognition, dialect recognition, context understanding, specialized term understanding, and complex acoustic situations like high noise or whispering. It achieved the lowest error rate on the comprehensive evaluation set.

image.png