Meta has launched Muse Voice Transcribe, its first real-time audio-aware model. Developers can call this capability through the Meta Model API, with a pricing of $3 (approximately 20.2 RMB) per 1000 minutes of audio, equivalent to $0.18 per hour.

Speak and Type, Automatically Separate 20+ Speakers
The model integrates streaming automatic speech recognition, speaker separation, and endpoint detection in the same process. "Streaming" means that it transcribes as people speak, without waiting for the entire audio to finish before producing results. Instead, it continuously outputs text in small segments, reducing latency and making it suitable for real-time transcription, meeting notes, and call subtitles. When users speak, the system can output text in real time, without needing to wait until the end of the recording to process it separately.
Muse Voice Transcribe can separate different speakers in recordings with more than 20 speakers, and it can also identify when a speaker stops talking, thus automatically determining the boundaries of an input. In terms of language support, the model covers over 70 languages, with 25 languages verified at launch; it supports audio longer than one hour and natural language switching within sentences or between sentences. Developers can also improve recognition accuracy for specific languages, terminology, or scenarios by using language, keyword, and context bias.

Adaptive Delay Mechanism Balances Speed and Accuracy
To balance real-time response and recognition accuracy, Meta introduced a dynamic decision-making mechanism called "adaptive delay." The system does not make a fixed trade-off between speed and accuracy for all words, but instead determines how much additional audio is needed before outputting the recognition result for each word: for easier-to-recognize speech, the system will submit the transcription results faster; for words that are unclear, have many technical terms, or are in complex contexts, it will use more audio information from the surrounding context to improve accuracy.
According to Meta, as of September 1, 2026, Muse Voice Transcribe ranked first on the Artificial Analysis Streaming Speech-to-Text Leaderboard.
Join Now