Picture published by the creator on Hugging Face.
Label who spoke when in any audio recording
Made by NVIDIA
View profile on Hugging Face (opens in a new tab)Nemotron-3-Diarization is an open-weight model from NVIDIA that does speaker diarization — working out "who spoke when" in a recording — for up to eight speakers.
It handles complete files and live streaming audio, splitting sound into independently processed chunks so recordings of any length work. Speaker labels follow arrival order: whoever speaks first becomes speaker one, following the Sortformer approach. At about 100 million parameters it is small by modern standards, and it runs through NVIDIA NeMo, NeMo-Speech.cpp, or the transformers library. NVIDIA states the model is ready for commercial and non-commercial use, and reports streaming latency down to 0.32 seconds of input buffering (compute time extra).
Can I use this?
You'll need Python with NeMo or the transformers library, and audio to analyze · Setup needed
Worth knowing The card states the model is ready for commercial and non-commercial use under its openmdw-1.1 licence; diarization accuracy depends on audio quality, and overlapping speech remains difficult for any such system. BuildTube has not run or verified this model.
Why it's here
- 579 people have liked it on Hugging Face.
- It was downloaded 40,936 times in the last 30 days.
Numbers from the snapshot taken 1 October 2026; not refreshed since.
Behind it
See the code on Hugging FaceBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?