BuildTube Preview About
Back to the feed

Picture published by the creator on Hugging Face.

Label who spoke when in any audio recording

Nemotron-3-Diarization is an open-weight model from NVIDIA that does speaker diarization — working out "who spoke when" in a recording — for up to eight speakers.

It handles complete files and live streaming audio, splitting sound into independently processed chunks so recordings of any length work. Speaker labels follow arrival order: whoever speaks first becomes speaker one, following the Sortformer approach. At about 100 million parameters it is small by modern standards, and it runs through NVIDIA NeMo, NeMo-Speech.cpp, or the transformers library. NVIDIA states the model is ready for commercial and non-commercial use, and reports streaming latency down to 0.32 seconds of input buffering (compute time extra).

Can I use this?

Needs a model runtime on your computer

You'll need Python with NeMo or the transformers library, and audio to analyze · Setup needed

Worth knowing The card states the model is ready for commercial and non-commercial use under its openmdw-1.1 licence; diarization accuracy depends on audio quality, and overlapping speech remains difficult for any such system. BuildTube has not run or verified this model.

Why it's here

  • 579 people have liked it on Hugging Face.
  • It was downloaded 40,936 times in the last 30 days.

Numbers from the snapshot taken 1 October 2026; not refreshed since.

Behind it

See the code on Hugging Face

BuildTube has not run or verified this project. Everything above is written from what the creator published.

Made this?

Back to the feed