NVIDIA Nemotron 3 Diarization in Glimpse

Glimpse is an open-source desktop app for macOS and Windows that combines voice dictation, meeting recording, transcription, and NVIDIA Nemotron 3 Diarization. The model runs on your computer and labels up to 8 speakers, live while you record and again in full when you stop.

Who spoke whenNemotron 3 Diarization
LeslieBenAnnAndy
0:000:150:300:451:00
speakers per track
8speakers per track
fewer errors than the previous model
42%fewer errors than the previous model
faster than real time on an M2 Pro
339×faster than real time on an M2 Pro
on disk, at full accuracy
106 MBon disk, at full accuracy

At a glance

Model
NVIDIA Nemotron 3 Diarization, built on Streaming Sortformer v3. It answers “who spoke when” at 10 ms resolution. It doesn't produce text, so Glimpse pairs it with a speech recognition model.
Where it runs
On your computer, through transcribe.cpp, Glimpse's on-device speech engine. No upload, no account, and it works with the wifi off.
Platforms
macOS 14 or later, and Windows 10 or later.
Speakers
Up to 8 per track.
Hardware
The graphics chip through Metal on Mac and Vulkan on Windows, so it doesn't need an NVIDIA card. It falls back to the CPU when no GPU is available.
Model file
An 8-bit GGUF (Q8_0), 106 MB, which scores the same as full precision. F16 and F32 versions are published on Hugging Face.
Speech recognition
Whichever model you pick: Parakeet TDT V3, Parakeet Unified, Whisper, Nemotron Streaming or Qwen3-ASR, all local. Cloud models you add with your own key work too.

How Glimpse uses it

Speaker labels for meetings, calls, lectures and interviews, from the first word to the export.

  1. 1

    Two tracks, labeled separately

    A recording keeps your microphone and your computer's audio apart, so your own lines are always yours. Nemotron 3 Diarization runs on each track on its own, so the people in the room with you never get mixed up with the people on the call.

  2. 2

    Live, while you record

    The live transcript runs the model in streaming mode with about one second of lookahead, so each line shows up already labeled by speaker.

  3. 3

    A full pass when you stop

    Once the recording ends, Glimpse runs it again over each whole track with 30 seconds of lookahead, NVIDIA's published setting for the best accuracy. Speaker labels then line up with the words from your speech model.

  4. 4

    Easy to fix

    Rename, recolor, merge or hide speakers in the Library, or move a single line to someone else. Detecting speakers again keeps the names you gave them. Files you import get the same treatment.

Benchmarks

On the AMI meeting corpus: 16 real meetings, 9.06 hours, with a 0 second collar and overlapped speech included. Every row went through the same pipeline. The port matches NVIDIA's reference at every stage. Methodology and the full table are on the model card.

Diarization error rate

Lower is better. Live mode, with one second of lookahead, still beats the previous model's best.

  • Glimpse, Q8_030.4 s lookahead9.23%
  • Glimpse, Q8_0, live1.04 s lookahead9.48%
  • NVIDIA reference, fp3230.4 s lookahead9.22%
  • Sortformer v2.1, previous model30.4 s lookahead15.99%

Speed

Times faster than real time on an Apple M2 Pro. A 10 minute meeting takes under two seconds on the GPU.

  • Nemotron 3 DiarizationGPU (Metal)339×
  • Nemotron 3 DiarizationCPU59×
  • Sortformer v2.1GPU (Metal)105×
  • Sortformer v2.1CPU44×

Common questions

Does Glimpse use NVIDIA Nemotron 3 Diarization?

Yes. Since version 1.2.7, Glimpse labels speakers with NVIDIA Nemotron 3 Diarization, running locally on Mac and Windows. It's used for live transcripts while you record, for the full pass after a recording, and for audio and video files you import.

Is it the same as Sortformer?

Nemotron 3 Diarization is NVIDIA's Streaming Sortformer v3. It replaced Sortformer v2.1 in Glimpse, with 42% fewer errors overall, about four times less speaker confusion, and twice as many speakers.

Do I need an NVIDIA graphics card?

No. Glimpse runs the model through transcribe.cpp, using Metal on Mac and Vulkan on Windows, so it works on Apple Silicon, AMD, Intel and NVIDIA graphics, and on the CPU when there's no GPU.

Does speaker detection work offline?

Yes. The model is downloaded once and runs on your computer. Your recordings are never uploaded for speaker detection.

How many speakers can it tell apart?

Up to 8 per track. Because your microphone and your computer's audio are separate tracks, a call with a few people in the room and more on the line still comes out clean.

Is speaker detection free?

Speaker detection is part of the Library and Recording Mode, which come with every Glimpse license and the two-week free trial. Dictation is free and unlimited.

Can I use the converted model in my own project?

Yes. The GGUF files are on Hugging Face under NVIDIA's OpenMDW 1.1 license, which allows commercial use, and they run with transcribe.cpp. Glimpse itself is open source under AGPL-3.0.

Try it on your next meeting

Dictation is free and unlimited. Recording and speaker detection come with a two-week free trial.