Benchmarks
Every number, in full.
Frozen, public test sets run through the same code path the app uses. Lower is better for error rates and latency.
Speech to text
- Common Voice Spontaneous 4.0 · 262 clips · 51 min
- MacBook Pro M5
- Measured 2026-08-18
Accuracy vs speed
Common Voice Spontaneous 4.0 · 262 clips · 51 min · MacBook Pro M5. Shorter error bars and longer speed bars are better.
- Parakeet V3Default6.9%73×
- Parakeet V27.3%73×
- Whisper Large V3 Turbo8.2%11×
- Cohere Transcribe10.8%5×
- Nemotron 3.513.8%45×
p50 decode: Parakeet V3 0.13 s · Parakeet V2 0.12 s · Whisper Large V3 Turbo 1.00 s · Cohere Transcribe 1.93 s · Nemotron 3.5 0.20 s
Speech engines
Lower is better.
- Parakeet TDT 0.6B V3Default6.9%
- Parakeet TDT 0.6B V27.3%
- Whisper Large V3 Turbo8.2%
- Cohere Transcribe10.8%
- Nemotron 3.5 Multilingual 0.6B13.8%
Full table
| Engine | WER | Segmented WER | p50 decode | p95 decode | Real time |
|---|---|---|---|---|---|
| NVIDIA Parakeet TDT 0.6B V3Lowest WER | 6.9% | 8.3% | 0.13 s | 0.37 s | 73× |
| NVIDIA Parakeet TDT 0.6B V2 | 7.3% | 9.3% | 0.12 s | 0.37 s | 73× |
| Whisper Large V3 Turbo | 8.2% | 9.8% | 1.00 s | 2.19 s | 11× |
| Cohere Transcribe | 10.8% | 12.5% | 1.93 s | 5.75 s | 5× |
| NVIDIA Nemotron 3.5 Multilingual 0.6B | 13.8% | 14.0% | 0.20 s | 0.60 s | 45× |
Not yet measured on this set: Apple Speech Analyzer, Qwen3-ASR 0.6B, Whisper Small, Whisper Tiny, Hinglish Prime. Apple Speech needs an authorized manual-QA run. WER is whole-clip; segmented WER re-decodes through live segmentation. Real time = audio duration ÷ decode time.
Smart formatting (ZenPolish)
- 500 held-out dictations · 7 categories · leakage-checked
- Apple Silicon, 24 GB, MLX Metal
- Measured 2026-09-25
Smart formatting
500 held-out dictations · 7 categories · leakage-checked. Lower is better.
- ZenPolish 1.7B v2.2ZenVoice6.7%
- Stock Qwen3-4B (no fine-tune)21.6%
- Stock Qwen3-1.7B (no fine-tune)33.7%
Full table
| Model | WER ↓ | Exact match ↑ | Punctuation F1 ↑ | Capitalization ↑ | Sentence starts ↑ | Added content ↓ |
|---|---|---|---|---|---|---|
| Stock Qwen3-1.7B (no fine-tune) | 33.7% | 13.8% | 35.2% | 85.1% | 38.7% | 22.7% |
| Stock Qwen3-4B (no fine-tune) | 21.6% | 39.2% | 75.5% | 95.3% | 74.1% | 16.6% |
| ZenPolish 1.7B v2.2Ships in ZenVoice | 6.7% | 53.6% | 75.9% | 97.8% | 91.6% | 2.2% |
Added content = words in the output that were never spoken. ZenPolish 1.7B v2.2 makes a fifth of the word errors of Qwen3-1.7B, the stock model it is trained from, and adds 90% fewer words that were never said.
Formatting WER by category
- Same 500 dictations
- Best value per row in bold
Formatting WER by category
Word error rate on each kind of speech, same 500 dictations. Lower is better.
- ZenPolish 1.7B v2.2
- Stock Qwen3-1.7B (its base)
- Stock Qwen3-4B
All dictations
500 dictations
−26.9 pts6.7%
1.7B v2.2
6.7%33.7%21.6%1.7B v2.2Qwen3-1.7BQwen3-4BHeavy fillers and restarts
94 dictations
−72.5 pts3.9%
1.7B v2.2
3.9%76.4%46.4%1.7B v2.2Qwen3-1.7BQwen3-4BFillers
43 dictations
−63.4 pts7.2%
1.7B v2.2
7.2%70.6%22.2%1.7B v2.2Qwen3-1.7BQwen3-4BFacts: numbers, money, dates
105 dictations
−29.4 pts15.2%
1.7B v2.2
15.2%44.6%24.8%1.7B v2.2Qwen3-1.7BQwen3-4BQuestions
72 dictations
−11.0 pts0.3%
1.7B v2.2
0.3%11.3%26.1%1.7B v2.2Qwen3-1.7BQwen3-4BFragments
62 dictations
−0.9 pts3.9%
1.7B v2.2
3.9%4.8%4.7%1.7B v2.2Qwen3-1.7BQwen3-4BExclamations
60 dictations
level4.5%
1.7B v2.2
4.5%4.8%4.0%1.7B v2.2Qwen3-1.7BQwen3-4BRun-ons
64 dictations
level8.9%
1.7B v2.2
8.9%8.4%7.7%1.7B v2.2Qwen3-1.7BQwen3-4B
Full table
| Category | Rows | Stock Qwen3-1.7B | Stock Qwen3-4B | ZenPolish 1.7B v2.2 |
|---|---|---|---|---|
| Heavy fillers and restarts | 94 | 76.4% | 46.4% | 3.9% |
| Fillers | 43 | 70.6% | 22.2% | 7.2% |
| Facts: numbers, money, dates | 105 | 44.6% | 24.8% | 15.2% |
| Questions | 72 | 11.3% | 26.1% | 0.3% |
| Fragments | 62 | 4.8% | 4.7% | 3.9% |
| Exclamations | 60 | 4.8% | 4.0% | 4.5% |
| Run-ons | 64 | 8.4% | 7.7% | 8.9% |
The weakest category is numbers, money and dates (15.2%); in the app, fixed rules write those. Run-ons are level with the stock models. On questions, fragments and exclamations every spoken word is kept, so word error rate only measures words a model damaged.