NVIDIA Nemotron 3 Diarization in Glimpse
Glimpse is an open-source desktop app for macOS and Windows that combines voice dictation, meeting recording, transcription, and NVIDIA Nemotron 3 Diarization. The model runs on your computer and labels up to 8 speakers, live while you record and again in full when you stop.
- speakers per track
- 8speakers per track
- fewer errors than the previous model
- 42%fewer errors than the previous model
- faster than real time on an M2 Pro
- 339×faster than real time on an M2 Pro
- on disk, at full accuracy
- 106 MBon disk, at full accuracy
At a glance
- Model
- NVIDIA Nemotron 3 Diarization, built on Streaming Sortformer v3. It answers “who spoke when” at 10 ms resolution. It doesn't produce text, so Glimpse pairs it with a speech recognition model.
- Where it runs
- On your computer, through transcribe.cpp, Glimpse's on-device speech engine. No upload, no account, and it works with the wifi off.
- Platforms
- macOS 14 or later, and Windows 10 or later.
- Speakers
- Up to 8 per track.
- Hardware
- The graphics chip through Metal on Mac and Vulkan on Windows, so it doesn't need an NVIDIA card. It falls back to the CPU when no GPU is available.
- Model file
- An 8-bit GGUF (Q8_0), 106 MB, which scores the same as full precision. F16 and F32 versions are published on Hugging Face.
- Speech recognition
- Whichever model you pick: Parakeet TDT V3, Parakeet Unified, Whisper, Nemotron Streaming or Qwen3-ASR, all local. Cloud models you add with your own key work too.
How Glimpse uses it
Speaker labels for meetings, calls, lectures and interviews, from the first word to the export.
- 1
Two tracks, labeled separately
A recording keeps your microphone and your computer's audio apart, so your own lines are always yours. Nemotron 3 Diarization runs on each track on its own, so the people in the room with you never get mixed up with the people on the call.
- 2
Live, while you record
The live transcript runs the model in streaming mode with about one second of lookahead, so each line shows up already labeled by speaker.
- 3
A full pass when you stop
Once the recording ends, Glimpse runs it again over each whole track with 30 seconds of lookahead, NVIDIA's published setting for the best accuracy. Speaker labels then line up with the words from your speech model.
- 4
Easy to fix
Rename, recolor, merge or hide speakers in the Library, or move a single line to someone else. Detecting speakers again keeps the names you gave them. Files you import get the same treatment.
Benchmarks
On the AMI meeting corpus: 16 real meetings, 9.06 hours, with a 0 second collar and overlapped speech included. Every row went through the same pipeline. The port matches NVIDIA's reference at every stage. Methodology and the full table are on the model card.
Diarization error rate
Lower is better. Live mode, with one second of lookahead, still beats the previous model's best.
- Glimpse, Q8_030.4 s lookahead9.23%
- Glimpse, Q8_0, live1.04 s lookahead9.48%
- NVIDIA reference, fp3230.4 s lookahead9.22%
- Sortformer v2.1, previous model30.4 s lookahead15.99%
Speed
Times faster than real time on an Apple M2 Pro. A 10 minute meeting takes under two seconds on the GPU.
- Nemotron 3 DiarizationGPU (Metal)339×
- Nemotron 3 DiarizationCPU59×
- Sortformer v2.1GPU (Metal)105×
- Sortformer v2.1CPU44×
Common questions
Does Glimpse use NVIDIA Nemotron 3 Diarization?
Yes. Since version 1.2.7, Glimpse labels speakers with NVIDIA Nemotron 3 Diarization, running locally on Mac and Windows. It's used for live transcripts while you record, for the full pass after a recording, and for audio and video files you import.
Is it the same as Sortformer?
Nemotron 3 Diarization is NVIDIA's Streaming Sortformer v3. It replaced Sortformer v2.1 in Glimpse, with 42% fewer errors overall, about four times less speaker confusion, and twice as many speakers.
Do I need an NVIDIA graphics card?
No. Glimpse runs the model through transcribe.cpp, using Metal on Mac and Vulkan on Windows, so it works on Apple Silicon, AMD, Intel and NVIDIA graphics, and on the CPU when there's no GPU.
Does speaker detection work offline?
Yes. The model is downloaded once and runs on your computer. Your recordings are never uploaded for speaker detection.
How many speakers can it tell apart?
Up to 8 per track. Because your microphone and your computer's audio are separate tracks, a call with a few people in the room and more on the line still comes out clean.
Is speaker detection free?
Speaker detection is part of the Library and Recording Mode, which come with every Glimpse license and the two-week free trial. Dictation is free and unlimited.
Can I use the converted model in my own project?
Yes. The GGUF files are on Hugging Face under NVIDIA's OpenMDW 1.1 license, which allows commercial use, and they run with transcribe.cpp. Glimpse itself is open source under AGPL-3.0.
Try it on your next meeting
Dictation is free and unlimited. Recording and speaker detection come with a two-week free trial.