On September 27, NVIDIA officially released the free artificial intelligence model Nemotron3Diarization with approximately 100 million parameters. This model focuses on the speech segmentation clustering (Diarization) task, and it has opened up weight data, allowing accurate identification of speakers at any moment in a conversation. It supports real-time and recorded audio processing for up to eight people speaking simultaneously.
In terms of technical breakthroughs and performance, Nemotron3Diarization topped the VoiceArena speech segmentation benchmark test (v1 version) with an error rate (DER) of 14.72%, significantly outperforming the second-ranked system (19.3%). Compared to its predecessor, Streaming Sortformer, the new model achieved an average error rate reduction of 41% across eight test scenarios using a 1.04-second audio buffer. To balance system latency and accuracy, the model supports four levels of dynamic audio buffer settings ranging from 0.32 seconds to 30.4 seconds.

In terms of functional applications, the new model can seamlessly collaborate with speech recognition systems such as Parakeet to automatically generate text with anonymous speaker tags (e.g., "speaker_2"). Although the error rate may increase when there are more participants, excessive background noise, or severe reverberation, its powerful underlying architecture can effectively address traditional challenges such as overlapping speech.
This release of a lightweight and free high-precision speech segmentation model by NVIDIA significantly lowers the technical barriers for complex speech analysis. This move not only provides a highly cost-effective underlying support for applications such as real-time meeting records, smart customer service, and multi-speaker voice interaction, but also indicates that the deployment of multi-speaker identification technology in edge devices and real-time commercial scenarios will accelerate further.