Local speech to text benchmarks for Mac and Windows
How fast and how accurate are the speech models you can run on your own computer? We measured the speech models in Glimpse on an Apple M2 Pro and a Windows PC, on the Neural Engine, the GPU and the CPU, with public test sets and a 9 minute recording. Updated October 5, 2026.
- faster than real time: Parakeet TDT V3 on the M2 Pro Neural Engine
- 313×faster than real time: Parakeet TDT V3 on the M2 Pro Neural Engine
- median for dictation clips under 5 seconds
- 16 msmedian for dictation clips under 5 seconds
- clips behind the Parakeet accuracy check
- 14,058clips behind the Parakeet accuracy check
- Neural Engine, Metal, Vulkan and CPU
- 2 machinesNeural Engine, Metal, Vulkan and CPU
Mac: Apple M2 Pro
Glimpse 1.3.1 as it ships, on MacBook with 32 GB memory, macOS 27 beta (26A428). Parakeet and Nemotron rows come from the release benchmark, with the 1.3.0 time for the same file alongside. More on what changed in the 1.3.1 post.
Time to transcribe the 9 minute file on an M2 Pro
Seconds, shorter is better. Each model on its fastest Mac backend. The exact numbers, every backend and the method are in the tables below.
- Parakeet Unified ENEnglish only, Neural Engine (Core ML)1.51 s
- Parakeet TDT V3int8 encoder, macOS 15 or later, Neural Engine (Core ML)1.77 s
- Parakeet TDT V3int8 encoder, macOS 14, Neural Engine (Core ML)1.92 s
- Whisper TinyQ8_0, Neural Engine (Core ML)3.55 s
- Nemotron Streaming ENEnglish only, files on the GPU, GPU (Metal)6.30 s
- Whisper Large V3 TurboQ8_0, Neural Engine (Core ML)13.05 s
- Whisper SmallQ8_0, Neural Engine (Core ML)13.51 s
- Qwen3-ASR 0.6BQ8_0, Neural Engine (Core ML)15.41 s
- Nemotron 3.5 Streamingmultilingual, CPU25.59 s
- Whisper Large V3Q5, Neural Engine (Core ML)39.36 s
Parakeet and Nemotron in Glimpse 1.3.1
| Model | Runs on | Short clip, median | Clip under 5 s, median | 9 min file | × real time | 9 min file in 1.3.0 | Peak memory, 9 min file | Download |
|---|---|---|---|---|---|---|---|---|
| Parakeet TDT V3int8 encoder, macOS 15 or later | Neural Engine (Core ML) | 26.7 ms | 16.0 ms | 1.77 s | 313× | 4.21 s | 617 MB | 549 MB |
| Parakeet TDT V3int8 encoder, macOS 14 | Neural Engine (Core ML) | 42.1 ms | 36.3 ms | 1.92 s | 288× | 4.21 s | 617 MB | 543 MB |
| Parakeet Unified ENEnglish only | Neural Engine (Core ML) | 38.9 ms | 36.4 ms | 1.51 s | 367× | 2.18 s | 1,309 MB | 1,822 MB |
| Nemotron Streaming ENEnglish only, files on the GPU | GPU (Metal) | 86.5 ms | 54.5 ms | 6.30 s | 88× | 47.39 s | 2,202 MB | 730 MB |
| Nemotron 3.5 Streamingmultilingual | CPU | 315.9 ms | 160.1 ms | 25.59 s | 22× | 51.80 s | 2,149 MB | 751 MB |
Short clip: 320 clips of 1 to 15 seconds. 9 min file: 553.9 seconds of mixed English, Spanish, German and French. Memory is the peak for the process transcribing the 9 minute file.
The macOS 15 Parakeet encoder carries 15, 10 and 5 second versions, which is why short clips are faster; its first load after download takes 26.9 seconds (once) and later loads 0.74 seconds, against 12.0 and 0.20 seconds for the macOS 14 version.
Nemotron Streaming EN transcribes files on the GPU; live dictation stays on the CPU. Nemotron 3.5 Streaming stays on the CPU on the Mac, because the GPU version was measurably less accurate in Hungarian.
Whisper and Qwen3-ASR
| Model | Runs on | Short clip, median | Short clip, p90 | Clip under 5 s, median | 9 min file | × real time |
|---|---|---|---|---|---|---|
| Whisper Large V3 TurboQ8_0 | Neural Engine (Core ML) | 470.0 ms | 519.8 ms | 440.6 ms | 13.05 s | 42× |
| Whisper Large V3 TurboQ8_0 | GPU (Metal) | 671.3 ms | 720.5 ms | 643.2 ms | 17.13 s | 32× |
| Whisper Large V3Q5 | Neural Engine (Core ML) | 799.1 ms | 987.3 ms | 663.1 ms | 39.36 s* | 14× |
| Whisper Large V3Q5 | GPU (Metal) | 1,075.1 ms | 1,266.3 ms | 932.4 ms | 45.25 s* | 12× |
| Whisper SmallQ8_0 | Neural Engine (Core ML) | 185.4 ms | 253.2 ms | 137.4 ms | 13.51 s* | 41× |
| Whisper SmallQ8_0 | GPU (Metal) | 222.7 ms | 288.4 ms | 177.0 ms | 9.84 s | 56× |
| Whisper TinyQ8_0 | Neural Engine (Core ML) | 62.6 ms | 97.9 ms | 39.5 ms | 3.55 s* | 156× |
| Whisper TinyQ8_0 | GPU (Metal) | 66.9 ms | 96.6 ms | 47.2 ms | 4.15 s* | 133× |
| Qwen3-ASR 0.6BQ8_0 | Neural Engine (Core ML) | 236.0 ms | 378.8 ms | 124.9 ms | 15.41 s | 36× |
| Qwen3-ASR 0.6BQ8_0 | GPU (Metal) | 261.5 ms | 425.5 ms | 131.3 ms | 17.33 s | 32× |
Measured October 1, 2026 with each clip run once in a single process, so these are slightly noisier than the table above. Glimpse 1.3.1 doesn't change these models: the same runs on the 1.3.0 engine gave timings within noise and no change in accuracy.
* Whisper gave different transcripts of the 9 minute file from run to run on these backends (up to 4 different texts in 4 runs), and its time varies with the text.
Windows: Xeon and RTX 4000 Ada
Glimpse 1.3.1 on an Intel Xeon w7-2495X + NVIDIA RTX 4000 Ada (24 cores, 20 GB graphics card, Windows 11 (build 26200)). Vulkan is what Glimpse picks when a supported graphics card is present; the CPU rows are the same machine with no GPU, using all 24 cores. See also dictation for Windows.
| Model | Runs on | Short clip, median | Clip under 5 s, median | 9 min file | × real time | 9 min file in 1.3.0 | Peak memory, 9 min file | Download |
|---|---|---|---|---|---|---|---|---|
| Parakeet TDT V3 | GPU (Vulkan) | 142.5 ms | 105.8 ms | 5.18 s | 107× | 7.35 s | 323 MB | 740 MB |
| Parakeet TDT V3 | CPU | 226.8 ms | 125.3 ms | 23.56 s | 24× | 118.62 s | 1,400 MB | 740 MB |
| Parakeet Unified EN | GPU (Vulkan) | 150.4 ms | 98.7 ms | 2.62 s | 211× | no text | 294 MB | 731 MB |
| Parakeet Unified EN | CPU | 229.7 ms | 129.4 ms | 18.33 s | 30× | no text | 1,367 MB | 731 MB |
| Nemotron Streaming ENfiles on the GPU | GPU (Vulkan) | 204.9 ms | 136.7 ms | 2.79 s | 199× | 91.80 s | 1,243 MB | 730 MB |
| Nemotron Streaming EN | CPU | 241.2 ms | 133.8 ms | 17.61 s | 31× | 91.80 s | 1,755 MB | 730 MB |
| Nemotron 3.5 Streamingfiles on the GPU | GPU (Vulkan) | 225.2 ms | 143.9 ms | 5.15 s | 108× | 94.92 s | 1,385 MB | 751 MB |
| Nemotron 3.5 Streaming | CPU | 277.4 ms | 146.1 ms | 19.84 s | 28× | 94.92 s | 1,836 MB | 751 MB |
| Whisper Large V3 TurboQ8_0 | GPU (Vulkan) | 259.7 ms | 236.7 ms | 7.65 s | 72× | 8.02 s | 299 MB | 886 MB |
| Qwen3-ASR 0.6B | GPU (Vulkan) | 165.6 ms | 82.7 ms | 10.77 s | 51× | 10.75 s | 463 MB | 850 MB |
“No text”: Parakeet Unified EN in 1.3.0 returned an empty transcript for the 9 minute file without Core ML, which 1.3.1 fixes. Nemotron in 1.3.0 ran everything on the CPU, so both of its rows share one 1.3.0 time.
Whisper Large V3 Turbo and Qwen3-ASR were run as a quicker check (4 runs each) to confirm 1.3.1 left them unchanged; they were.
Accuracy
Word error rate: the share of words a model got wrong, lower is better. First, every model on the same 320 clips, so the rows compare directly.
| Model | Runs on | English (200) | Spanish (40) | German (40) | French (40) |
|---|---|---|---|---|---|
| Parakeet TDT V3 | Neural Engine (Core ML) | 2.01% | 3.51% | 4.57% | 7.15% |
| Parakeet Unified EN | Neural Engine (Core ML) | 1.76% | n/a | n/a | n/a |
| Nemotron Streaming EN | CPU | 1.80% | n/a | n/a | n/a |
| Nemotron 3.5 Streaming | CPU | 2.20% | 5.42% | 7.22% | 12.83% |
| Qwen3-ASR 0.6B | Neural Engine (Core ML) | 1.80% | 4.78% | 7.34% | 8.62% |
| Whisper Large V3 Turbo | Neural Engine (Core ML) | 2.10% | 2.98% | 2.65% | 7.25% |
| Whisper Large V3 (Q5) | Neural Engine (Core ML) | 2.32% | 2.66% | 2.77% | 7.25% |
| Whisper Small | Neural Engine (Core ML) | 3.28% | 5.84% | 9.03% | 16.26% |
| Whisper Tiny | Neural Engine (Core ML) | 7.27% | 21.68% | 29.36% | 46.43% |
English: LibriSpeech test-clean. Spanish, German, French: FLEURS. n/a: English-only model. Measured on the 1.3.0 engine, October 1, 2026. Across all 14,058 clips of the large sets, 1.3.1 left no language significantly worse for Parakeet TDT V3 (including the int8 encoder) and Parakeet Unified EN. 40 clips per language is a small sample: treat differences under a point or two as noise.
On the large public test sets
| Model | Runs on | LibriSpeech test-clean | LibriSpeech test-other | FLEURS |
|---|---|---|---|---|
| Parakeet TDT V3 (int8, macOS 15+) | Neural Engine (Core ML) | 2.19%2,620 clips | 4.03%2,939 clips | 13.00%8,499 clips, 25 languages |
| Parakeet Unified EN | Neural Engine (Core ML) | 1.83%2,620 clips | 3.42%2,939 clips | 4.34%350 clips, English |
| Nemotron Streaming EN | GPU (Metal) | 2.38%500 clips | 5.00%500 clips | 6.74%100 clips, English |
| Nemotron 3.5 Streaming | CPU | 3.14%500 clips | 6.92%500 clips | 17.12%2,000 clips, 20 languages |
Parakeet rows use every clip in each set. Nemotron rows use a fixed, seeded sample (500 clips per LibriSpeech set, 100 FLEURS clips per language), and the FLEURS columns cover different languages per model, so compare FLEURS only within a row's own language mix. Apple M2 Pro, Glimpse 1.3.1.
Next to Apple's built-in recognizer
macOS has its own on-device speech recognition (SpeechAnalyzer). We ran it through the same harness on the 47-clip dictation set (English, Spanish, German and French, plus silences and a 65 second passage) and the 9 minute file. Compare with Apple Dictation.
| Model | Runs on | Clip, median | Clip, p90 | 9 min file | English WER |
|---|---|---|---|---|---|
| Parakeet TDT V3 | Neural Engine (Core ML) | 50.3 ms | 86.3 ms | 2.22 s | 2.40% |
| Qwen3-ASR 0.6B | Neural Engine (Core ML) | 265.1 ms | 524.3 ms | 15.41 s | 2.90% |
| Apple Speech | System | 295.5 ms | 1,200.1 ms | 42.99 s | 9.21% |
| Whisper Large V3 Turbo | Neural Engine (Core ML) | 476.8 ms | 533.5 ms | 13.05 s | 2.90% |
Engine harness, October 1, 2026, on an early build of the 1.3.1 engine (Glimpse-Speech 112aa43) with Parakeet's original encoder; the shipping int8 encoder is faster (see the Mac table). Apple Speech needs no download. Its 9 minute time is the median of 2 runs after a warm-up.
Next to Superwhisper
Superwhisper also offers Parakeet V3 on the Mac. Both apps transcribed the 9 minute file on the same M2 Pro and were timed from the outside. We never loaded Superwhisper's model files ourselves. More in Glimpse vs Superwhisper.
| App | Model | 9 min file |
|---|---|---|
| Superwhisper | Parakeet V3 (6-bit palettized Argmax encoder) | 3.9 to 4.2 s |
| Glimpse 1.3.0 | Parakeet TDT V3 | 4.21 s |
| Glimpse 1.3.1 | Parakeet TDT V3, int8 encoder | 1.73 to 1.77 s |
Ranges across the runs taken. The Superwhisper version and the number of runs were not recorded, so treat this as a spot check, not a controlled benchmark.
Methodology
Machines
Test audio
Timing
Accuracy
Engine versus app
Versions
Not measured
- Energy use and battery drain. Reading the power counters on macOS needs root, so they weren't recorded.
- Other Macs and PCs. Every number comes from the two machines above; a base M-series chip, an Intel Mac or a laptop GPU will be slower or faster in ways these tables don't predict.
- Live streaming latency for Nemotron. Its live dictation runs on the CPU in both versions; the short clip numbers for Nemotron time whole clips, the way files and recordings are transcribed.
- Whisper and Qwen3-ASR word error rates on the large sets, and Mac speed for the other Whisper sizes Glimpse offers (Base, Medium, Distil).
- The Superwhisper version, and how many runs were timed for it. Only the range was recorded.
Common questions
What is the fastest local speech to text model on a Mac?
In these tests, Parakeet TDT V3 on the Neural Engine: a 553.9 second recording in 1.77 seconds on an M2 Pro, about 313 times faster than real time, and 16 ms for a median dictation under 5 seconds. Parakeet Unified EN was faster on the long file (1.51 seconds) but only does English.
Is Parakeet faster than Whisper?
On the same M2 Pro and the same 320 clips, Parakeet TDT V3 took 26.7 ms per clip and Whisper Large V3 Turbo 470 ms, both on the Neural Engine. On the 9 minute file it was 1.77 seconds against 13.05. Whisper Large V3 Turbo was more accurate in German and Spanish on those clips; Parakeet was slightly better in English.
How does Apple's built-in speech recognition compare?
On the 47 dictation clips, Apple's on-device recognizer had a median of 295.5 ms and an English word error rate of 9.21%, against 50.3 ms and 2.4% for Parakeet TDT V3. On the 9 minute file it took 42.99 seconds.
Do I need a graphics card on Windows?
No. Without one, Parakeet TDT V3 transcribed the 9 minute file in 23.56 seconds on a 24-core Xeon, and short clips in 226.8 ms. With an RTX 4000 Ada through Vulkan it took 5.18 seconds. Fewer cores will be slower.
Can I reproduce these numbers?
The models are the public files linked in the tables, and the test audio comes from LibriSpeech and FLEURS. Glimpse, Glimpse-Speech and transcribe.cpp are open source, so the same engine runs on your machine.
Run them on your own machine
Every Glimpse model on this page is a free download, with free, unlimited dictation on Mac and Windows.