Skip to content
ZenVoice

Benchmarks

Every number, in full.

Frozen, public test sets run through the same code path the app uses. Lower is better for error rates and latency.

Speech to text

Common Voice Spontaneous 4.0 · 262 clips · 51 min
MacBook Pro M5
Measured 2026-08-18

Accuracy vs speed

Common Voice Spontaneous 4.0 · 262 clips · 51 min · MacBook Pro M5. Shorter error bars and longer speed bars are better.

  • Parakeet V3Default6.9%73×
  • Parakeet V27.3%73×
  • Whisper Large V3 Turbo8.2%11×
  • Cohere Transcribe10.8%5×
  • Nemotron 3.513.8%45×

p50 decode: Parakeet V3 0.13 s · Parakeet V2 0.12 s · Whisper Large V3 Turbo 1.00 s · Cohere Transcribe 1.93 s · Nemotron 3.5 0.20 s

Speech engines

Lower is better.

  • Parakeet TDT 0.6B V3Default6.9%
  • Parakeet TDT 0.6B V27.3%
  • Whisper Large V3 Turbo8.2%
  • Cohere Transcribe10.8%
  • Nemotron 3.5 Multilingual 0.6B13.8%

Full table

Speech engine benchmark on Common Voice Spontaneous 4.0 · 262 clips · 51 min, MacBook Pro M5
EngineWERSegmented WERp50 decodep95 decodeReal time
NVIDIA Parakeet TDT 0.6B V3Lowest WER6.9%8.3%0.13 s0.37 s73×
NVIDIA Parakeet TDT 0.6B V27.3%9.3%0.12 s0.37 s73×
Whisper Large V3 Turbo8.2%9.8%1.00 s2.19 s11×
Cohere Transcribe10.8%12.5%1.93 s5.75 s5×
NVIDIA Nemotron 3.5 Multilingual 0.6B13.8%14.0%0.20 s0.60 s45×

Not yet measured on this set: Apple Speech Analyzer, Qwen3-ASR 0.6B, Whisper Small, Whisper Tiny, Hinglish Prime. Apple Speech needs an authorized manual-QA run. WER is whole-clip; segmented WER re-decodes through live segmentation. Real time = audio duration ÷ decode time.

Smart formatting (ZenPolish)

500 held-out dictations · 7 categories · leakage-checked
Apple Silicon, 24 GB, MLX Metal
Measured 2026-09-25

Smart formatting

500 held-out dictations · 7 categories · leakage-checked. Lower is better.

  • ZenPolish 1.7B v2.2ZenVoice6.7%
  • Stock Qwen3-4B (no fine-tune)21.6%
  • Stock Qwen3-1.7B (no fine-tune)33.7%

Full table

ZenPolish formatting benchmark
ModelWER ↓Exact match ↑Punctuation F1 ↑Capitalization ↑Sentence starts ↑Added content ↓
Stock Qwen3-1.7B (no fine-tune)33.7%13.8%35.2%85.1%38.7%22.7%
Stock Qwen3-4B (no fine-tune)21.6%39.2%75.5%95.3%74.1%16.6%
ZenPolish 1.7B v2.2Ships in ZenVoice6.7%53.6%75.9%97.8%91.6%2.2%

Added content = words in the output that were never spoken. ZenPolish 1.7B v2.2 makes a fifth of the word errors of Qwen3-1.7B, the stock model it is trained from, and adds 90% fewer words that were never said.

Formatting WER by category

Same 500 dictations
Best value per row in bold

Formatting WER by category

Word error rate on each kind of speech, same 500 dictations. Lower is better.

  • ZenPolish 1.7B v2.2
  • Stock Qwen3-1.7B (its base)
  • Stock Qwen3-4B
  • All dictations

    500 dictations

    6.7%

    1.7B v2.2

    −26.9 pts
    6.7%
    33.7%
    21.6%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Heavy fillers and restarts

    94 dictations

    3.9%

    1.7B v2.2

    −72.5 pts
    3.9%
    76.4%
    46.4%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Fillers

    43 dictations

    7.2%

    1.7B v2.2

    −63.4 pts
    7.2%
    70.6%
    22.2%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Facts: numbers, money, dates

    105 dictations

    15.2%

    1.7B v2.2

    −29.4 pts
    15.2%
    44.6%
    24.8%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Questions

    72 dictations

    0.3%

    1.7B v2.2

    −11.0 pts
    0.3%
    11.3%
    26.1%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Fragments

    62 dictations

    3.9%

    1.7B v2.2

    −0.9 pts
    3.9%
    4.8%
    4.7%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Exclamations

    60 dictations

    4.5%

    1.7B v2.2

    level
    4.5%
    4.8%
    4.0%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Run-ons

    64 dictations

    8.9%

    1.7B v2.2

    level
    8.9%
    8.4%
    7.7%
    1.7B v2.2Qwen3-1.7BQwen3-4B

Full table

Formatting word error rate per category
CategoryRowsStock Qwen3-1.7BStock Qwen3-4BZenPolish 1.7B v2.2
Heavy fillers and restarts9476.4%46.4%3.9%
Fillers4370.6%22.2%7.2%
Facts: numbers, money, dates10544.6%24.8%15.2%
Questions7211.3%26.1%0.3%
Fragments624.8%4.7%3.9%
Exclamations604.8%4.0%4.5%
Run-ons648.4%7.7%8.9%

The weakest category is numbers, money and dates (15.2%); in the app, fixed rules write those. Run-ons are level with the stock models. On questions, fragments and exclamations every spoken word is kept, so word error rate only measures words a model damaged.

Ask AI about this pageChatGPT ↗Claude ↗Perplexity ↗Grok ↗Markdown