Back home

Local speech to text benchmarks for Mac and Windows

How fast and how accurate are the speech models you can run on your own computer? We measured the speech models in Glimpse on an Apple M2 Pro and a Windows PC, on the Neural Engine, the GPU and the CPU, with public test sets and a 9 minute recording. Updated October 5, 2026.

faster than real time: Parakeet TDT V3 on the M2 Pro Neural Engine
313×faster than real time: Parakeet TDT V3 on the M2 Pro Neural Engine
median for dictation clips under 5 seconds
16 msmedian for dictation clips under 5 seconds
clips behind the Parakeet accuracy check
14,058clips behind the Parakeet accuracy check
Neural Engine, Metal, Vulkan and CPU
2 machinesNeural Engine, Metal, Vulkan and CPU

Mac: Apple M2 Pro

Glimpse 1.3.1 as it ships, on MacBook with 32 GB memory, macOS 27 beta (26A428). Parakeet and Nemotron rows come from the release benchmark, with the 1.3.0 time for the same file alongside. More on what changed in the 1.3.1 post.

Time to transcribe the 9 minute file on an M2 Pro

Seconds, shorter is better. Each model on its fastest Mac backend. The exact numbers, every backend and the method are in the tables below.

  • Parakeet Unified ENEnglish only, Neural Engine (Core ML)1.51 s
  • Parakeet TDT V3int8 encoder, macOS 15 or later, Neural Engine (Core ML)1.77 s
  • Parakeet TDT V3int8 encoder, macOS 14, Neural Engine (Core ML)1.92 s
  • Whisper TinyQ8_0, Neural Engine (Core ML)3.55 s
  • Nemotron Streaming ENEnglish only, files on the GPU, GPU (Metal)6.30 s
  • Whisper Large V3 TurboQ8_0, Neural Engine (Core ML)13.05 s
  • Whisper SmallQ8_0, Neural Engine (Core ML)13.51 s
  • Qwen3-ASR 0.6BQ8_0, Neural Engine (Core ML)15.41 s
  • Nemotron 3.5 Streamingmultilingual, CPU25.59 s
  • Whisper Large V3Q5, Neural Engine (Core ML)39.36 s

Parakeet and Nemotron in Glimpse 1.3.1

Parakeet and Nemotron speed on an Apple M2 Pro, Glimpse 1.3.1
ModelRuns onShort clip, medianClip under 5 s, median9 min file× real time9 min file in 1.3.0Peak memory, 9 min fileDownload
Parakeet TDT V3int8 encoder, macOS 15 or laterNeural Engine (Core ML)26.7 ms16.0 ms1.77 s313×4.21 s617 MB549 MB
Parakeet TDT V3int8 encoder, macOS 14Neural Engine (Core ML)42.1 ms36.3 ms1.92 s288×4.21 s617 MB543 MB
Parakeet Unified ENEnglish onlyNeural Engine (Core ML)38.9 ms36.4 ms1.51 s367×2.18 s1,309 MB1,822 MB
Nemotron Streaming ENEnglish only, files on the GPUGPU (Metal)86.5 ms54.5 ms6.30 s88×47.39 s2,202 MB730 MB
Nemotron 3.5 StreamingmultilingualCPU315.9 ms160.1 ms25.59 s22×51.80 s2,149 MB751 MB

Short clip: 320 clips of 1 to 15 seconds. 9 min file: 553.9 seconds of mixed English, Spanish, German and French. Memory is the peak for the process transcribing the 9 minute file.

The macOS 15 Parakeet encoder carries 15, 10 and 5 second versions, which is why short clips are faster; its first load after download takes 26.9 seconds (once) and later loads 0.74 seconds, against 12.0 and 0.20 seconds for the macOS 14 version.

Nemotron Streaming EN transcribes files on the GPU; live dictation stays on the CPU. Nemotron 3.5 Streaming stays on the CPU on the Mac, because the GPU version was measurably less accurate in Hungarian.

Whisper and Qwen3-ASR

Whisper and Qwen3-ASR speed on an Apple M2 Pro
ModelRuns onShort clip, medianShort clip, p90Clip under 5 s, median9 min file× real time
Whisper Large V3 TurboQ8_0Neural Engine (Core ML)470.0 ms519.8 ms440.6 ms13.05 s42×
Whisper Large V3 TurboQ8_0GPU (Metal)671.3 ms720.5 ms643.2 ms17.13 s32×
Whisper Large V3Q5Neural Engine (Core ML)799.1 ms987.3 ms663.1 ms39.36 s*14×
Whisper Large V3Q5GPU (Metal)1,075.1 ms1,266.3 ms932.4 ms45.25 s*12×
Whisper SmallQ8_0Neural Engine (Core ML)185.4 ms253.2 ms137.4 ms13.51 s*41×
Whisper SmallQ8_0GPU (Metal)222.7 ms288.4 ms177.0 ms9.84 s56×
Whisper TinyQ8_0Neural Engine (Core ML)62.6 ms97.9 ms39.5 ms3.55 s*156×
Whisper TinyQ8_0GPU (Metal)66.9 ms96.6 ms47.2 ms4.15 s*133×
Qwen3-ASR 0.6BQ8_0Neural Engine (Core ML)236.0 ms378.8 ms124.9 ms15.41 s36×
Qwen3-ASR 0.6BQ8_0GPU (Metal)261.5 ms425.5 ms131.3 ms17.33 s32×

Measured October 1, 2026 with each clip run once in a single process, so these are slightly noisier than the table above. Glimpse 1.3.1 doesn't change these models: the same runs on the 1.3.0 engine gave timings within noise and no change in accuracy.

* Whisper gave different transcripts of the 9 minute file from run to run on these backends (up to 4 different texts in 4 runs), and its time varies with the text.

Windows: Xeon and RTX 4000 Ada

Glimpse 1.3.1 on an Intel Xeon w7-2495X + NVIDIA RTX 4000 Ada (24 cores, 20 GB graphics card, Windows 11 (build 26200)). Vulkan is what Glimpse picks when a supported graphics card is present; the CPU rows are the same machine with no GPU, using all 24 cores. See also dictation for Windows.

Local speech model speed on Windows, Glimpse 1.3.1
ModelRuns onShort clip, medianClip under 5 s, median9 min file× real time9 min file in 1.3.0Peak memory, 9 min fileDownload
Parakeet TDT V3GPU (Vulkan)142.5 ms105.8 ms5.18 s107×7.35 s323 MB740 MB
Parakeet TDT V3CPU226.8 ms125.3 ms23.56 s24×118.62 s1,400 MB740 MB
Parakeet Unified ENGPU (Vulkan)150.4 ms98.7 ms2.62 s211×no text294 MB731 MB
Parakeet Unified ENCPU229.7 ms129.4 ms18.33 s30×no text1,367 MB731 MB
Nemotron Streaming ENfiles on the GPUGPU (Vulkan)204.9 ms136.7 ms2.79 s199×91.80 s1,243 MB730 MB
Nemotron Streaming ENCPU241.2 ms133.8 ms17.61 s31×91.80 s1,755 MB730 MB
Nemotron 3.5 Streamingfiles on the GPUGPU (Vulkan)225.2 ms143.9 ms5.15 s108×94.92 s1,385 MB751 MB
Nemotron 3.5 StreamingCPU277.4 ms146.1 ms19.84 s28×94.92 s1,836 MB751 MB
Whisper Large V3 TurboQ8_0GPU (Vulkan)259.7 ms236.7 ms7.65 s72×8.02 s299 MB886 MB
Qwen3-ASR 0.6BGPU (Vulkan)165.6 ms82.7 ms10.77 s51×10.75 s463 MB850 MB

“No text”: Parakeet Unified EN in 1.3.0 returned an empty transcript for the 9 minute file without Core ML, which 1.3.1 fixes. Nemotron in 1.3.0 ran everything on the CPU, so both of its rows share one 1.3.0 time.

Whisper Large V3 Turbo and Qwen3-ASR were run as a quicker check (4 runs each) to confirm 1.3.1 left them unchanged; they were.

Accuracy

Word error rate: the share of words a model got wrong, lower is better. First, every model on the same 320 clips, so the rows compare directly.

Word error rate on the same 320 clips, Apple M2 Pro
ModelRuns onEnglish (200)Spanish (40)German (40)French (40)
Parakeet TDT V3Neural Engine (Core ML)2.01%3.51%4.57%7.15%
Parakeet Unified ENNeural Engine (Core ML)1.76%n/an/an/a
Nemotron Streaming ENCPU1.80%n/an/an/a
Nemotron 3.5 StreamingCPU2.20%5.42%7.22%12.83%
Qwen3-ASR 0.6BNeural Engine (Core ML)1.80%4.78%7.34%8.62%
Whisper Large V3 TurboNeural Engine (Core ML)2.10%2.98%2.65%7.25%
Whisper Large V3 (Q5)Neural Engine (Core ML)2.32%2.66%2.77%7.25%
Whisper SmallNeural Engine (Core ML)3.28%5.84%9.03%16.26%
Whisper TinyNeural Engine (Core ML)7.27%21.68%29.36%46.43%

English: LibriSpeech test-clean. Spanish, German, French: FLEURS. n/a: English-only model. Measured on the 1.3.0 engine, October 1, 2026. Across all 14,058 clips of the large sets, 1.3.1 left no language significantly worse for Parakeet TDT V3 (including the int8 encoder) and Parakeet Unified EN. 40 clips per language is a small sample: treat differences under a point or two as noise.

On the large public test sets

Word error rate on LibriSpeech and FLEURS, Glimpse 1.3.1
ModelRuns onLibriSpeech test-cleanLibriSpeech test-otherFLEURS
Parakeet TDT V3 (int8, macOS 15+)Neural Engine (Core ML)2.19%2,620 clips4.03%2,939 clips13.00%8,499 clips, 25 languages
Parakeet Unified ENNeural Engine (Core ML)1.83%2,620 clips3.42%2,939 clips4.34%350 clips, English
Nemotron Streaming ENGPU (Metal)2.38%500 clips5.00%500 clips6.74%100 clips, English
Nemotron 3.5 StreamingCPU3.14%500 clips6.92%500 clips17.12%2,000 clips, 20 languages

Parakeet rows use every clip in each set. Nemotron rows use a fixed, seeded sample (500 clips per LibriSpeech set, 100 FLEURS clips per language), and the FLEURS columns cover different languages per model, so compare FLEURS only within a row's own language mix. Apple M2 Pro, Glimpse 1.3.1.

Next to Apple's built-in recognizer

macOS has its own on-device speech recognition (SpeechAnalyzer). We ran it through the same harness on the 47-clip dictation set (English, Spanish, German and French, plus silences and a 65 second passage) and the 9 minute file. Compare with Apple Dictation.

Apple Speech and local models on 47 dictation clips, Apple M2 Pro
ModelRuns onClip, medianClip, p909 min fileEnglish WER
Parakeet TDT V3Neural Engine (Core ML)50.3 ms86.3 ms2.22 s2.40%
Qwen3-ASR 0.6BNeural Engine (Core ML)265.1 ms524.3 ms15.41 s2.90%
Apple SpeechSystem295.5 ms1,200.1 ms42.99 s9.21%
Whisper Large V3 TurboNeural Engine (Core ML)476.8 ms533.5 ms13.05 s2.90%

Engine harness, October 1, 2026, on an early build of the 1.3.1 engine (Glimpse-Speech 112aa43) with Parakeet's original encoder; the shipping int8 encoder is faster (see the Mac table). Apple Speech needs no download. Its 9 minute time is the median of 2 runs after a warm-up.

Next to Superwhisper

Superwhisper also offers Parakeet V3 on the Mac. Both apps transcribed the 9 minute file on the same M2 Pro and were timed from the outside. We never loaded Superwhisper's model files ourselves. More in Glimpse vs Superwhisper.

Superwhisper and Glimpse on the 9 minute file, Apple M2 Pro
AppModel9 min file
SuperwhisperParakeet V3 (6-bit palettized Argmax encoder)3.9 to 4.2 s
Glimpse 1.3.0Parakeet TDT V34.21 s
Glimpse 1.3.1Parakeet TDT V3, int8 encoder1.73 to 1.77 s

Ranges across the runs taken. The Superwhisper version and the number of runs were not recorded, so treat this as a spot check, not a controlled benchmark.

Methodology

Machines

Mac: Apple M2 Pro, MacBook with 32 GB memory, macOS 27 beta (26A428). Windows: Intel Xeon w7-2495X + NVIDIA RTX 4000 Ada, 24 cores, 20 GB graphics card, Windows 11 (build 26200). One benchmark process at a time; timed parts waited for a load average under 4. Every Mac run logged load, thermal state and any other busy process; Windows runs logged CPU and GPU use.

Test audio

Short clips: 320 clips of 1 to 15 seconds (200 LibriSpeech test-clean in English, 40 FLEURS each in Spanish, German and French), 2,321 seconds in total. Under 5 s is the 105 of them shorter than 5 seconds. The 9 minute file is 553.9 seconds: 47 clips in four languages joined with half a second of silence, including a 65 second passage, synthesized meeting, email and numbers clips, and silences. All audio is 16 kHz mono.

Timing

Wall-clock time inside Glimpse-Speech and the transcribe.cpp fork Glimpse runs, with the model already loaded; load times are measured separately. 1.3.1 rows: each clip timed in 10 separate processes (5 for Nemotron, 3 on the Windows CPU, 4 for Whisper and Qwen3-ASR on Windows) and the median taken; the 9 minute file once to warm up, then the median of 10 runs (5 for Nemotron, 3 on the Windows CPU, 4 for Whisper and Qwen3-ASR on Windows). Mac Whisper and Qwen3-ASR rows: each clip once in one process, the 9 minute file as the median of 3 runs after a warm-up. “× real time” is 553.9 seconds divided by the measured time.

Accuracy

Word error rate after lowercasing, removing punctuation and spelling out numbers, so “42” and “forty-two” count the same. Sets: LibriSpeech test-clean and test-other and FLEURS test (both CC BY 4.0). Changes between engine versions are judged per language with a paired bootstrap 95% confidence interval; a language counts as worse only when the interval excludes zero.

Engine versus app

The engine numbers skip work the app adds (its own chunking, silence gating, passing text to the window). Checked against the real app on the Mac: the 9 minute file took 1.75 seconds in Glimpse with the int8 Parakeet encoder (10 runs) against 1.77 in the engine.

Versions

Glimpse 1.3.0 uses Glimpse-Speech 2.0.2; Glimpse 1.3.1 uses Glimpse-Speech 2.0.3. Model files are the ones Glimpse downloads, and each run records their sha256. Measured October 1 to 3, 2026.

Not measured

  • Energy use and battery drain. Reading the power counters on macOS needs root, so they weren't recorded.
  • Other Macs and PCs. Every number comes from the two machines above; a base M-series chip, an Intel Mac or a laptop GPU will be slower or faster in ways these tables don't predict.
  • Live streaming latency for Nemotron. Its live dictation runs on the CPU in both versions; the short clip numbers for Nemotron time whole clips, the way files and recordings are transcribed.
  • Whisper and Qwen3-ASR word error rates on the large sets, and Mac speed for the other Whisper sizes Glimpse offers (Base, Medium, Distil).
  • The Superwhisper version, and how many runs were timed for it. Only the range was recorded.

Common questions

What is the fastest local speech to text model on a Mac?

In these tests, Parakeet TDT V3 on the Neural Engine: a 553.9 second recording in 1.77 seconds on an M2 Pro, about 313 times faster than real time, and 16 ms for a median dictation under 5 seconds. Parakeet Unified EN was faster on the long file (1.51 seconds) but only does English.

Is Parakeet faster than Whisper?

On the same M2 Pro and the same 320 clips, Parakeet TDT V3 took 26.7 ms per clip and Whisper Large V3 Turbo 470 ms, both on the Neural Engine. On the 9 minute file it was 1.77 seconds against 13.05. Whisper Large V3 Turbo was more accurate in German and Spanish on those clips; Parakeet was slightly better in English.

How does Apple's built-in speech recognition compare?

On the 47 dictation clips, Apple's on-device recognizer had a median of 295.5 ms and an English word error rate of 9.21%, against 50.3 ms and 2.4% for Parakeet TDT V3. On the 9 minute file it took 42.99 seconds.

Do I need a graphics card on Windows?

No. Without one, Parakeet TDT V3 transcribed the 9 minute file in 23.56 seconds on a 24-core Xeon, and short clips in 226.8 ms. With an RTX 4000 Ada through Vulkan it took 5.18 seconds. Fewer cores will be slower.

Can I reproduce these numbers?

The models are the public files linked in the tables, and the test audio comes from LibriSpeech and FLEURS. Glimpse, Glimpse-Speech and transcribe.cpp are open source, so the same engine runs on your machine.

Run them on your own machine

Every Glimpse model on this page is a free download, with free, unlimited dictation on Mac and Windows.