# Benchmarks · ZenVoice

Source: https://getzenvoice.com/benchmarks

> Word error rate, latency and formatting quality for ZenVoice's on-device models.

Benchmarks

# Every number, in full.

Frozen, public test sets run through the same code path the app uses. Lower is better for error rates and latency.

## Speech to text

Common Voice Spontaneous 4.0 · 262 clips · 51 min

MacBook Pro M5

Measured 2026-08-18

### Accuracy vs speed

Common Voice Spontaneous 4.0 · 262 clips · 51 min · MacBook Pro M5. Shorter error bars and longer speed bars are better.

EngineWord error rate · lower is betterSpeed · times faster than real time (log)

-   Parakeet V3Default6.9%73×
-   Parakeet V27.3%73×
-   Whisper Large V3 Turbo8.2%11×
-   Cohere Transcribe10.8%5×
-   Nemotron 3.513.8%45×

p50 decode: Parakeet V3 0.13 s · Parakeet V2 0.12 s · Whisper Large V3 Turbo 1.00 s · Cohere Transcribe 1.93 s · Nemotron 3.5 0.20 s

### Speech engines

Lower is better.

-   Parakeet TDT 0.6B V3Default6.9%
-   Parakeet TDT 0.6B V27.3%
-   Whisper Large V3 Turbo8.2%
-   Cohere Transcribe10.8%
-   Nemotron 3.5 Multilingual 0.6B13.8%

### Full table

Engine

WER

Segmented WER

p50 decode

p95 decode

Real time

NVIDIA Parakeet TDT 0.6B V3Lowest WER

6.9%

8.3%

0.13 s

0.37 s

73×

NVIDIA Parakeet TDT 0.6B V2

7.3%

9.3%

0.12 s

0.37 s

73×

Whisper Large V3 Turbo

8.2%

9.8%

1.00 s

2.19 s

11×

Cohere Transcribe

10.8%

12.5%

1.93 s

5.75 s

5×

NVIDIA Nemotron 3.5 Multilingual 0.6B

13.8%

14.0%

0.20 s

0.60 s

45×

Not yet measured on this set: Apple Speech Analyzer, Qwen3-ASR 0.6B, Whisper Small, Whisper Tiny, Hinglish Prime. Apple Speech needs an authorized manual-QA run. WER is whole-clip; segmented WER re-decodes through live segmentation. Real time = audio duration ÷ decode time.

## Smart formatting (ZenPolish)

500 held-out dictations · 7 categories · leakage-checked

Apple Silicon, 24 GB, MLX Metal

Measured 2026-09-25

### Smart formatting

500 held-out dictations · 7 categories · leakage-checked. Lower is better.

-   ZenPolish 1.7B v2.2ZenVoice6.7%
-   Stock Qwen3-4B (no fine-tune)21.6%
-   Stock Qwen3-1.7B (no fine-tune)33.7%

### Full table

Model

WER ↓

Exact match ↑

Punctuation F1 ↑

Capitalization ↑

Sentence starts ↑

Added content ↓

Stock Qwen3-1.7B (no fine-tune)

33.7%

13.8%

35.2%

85.1%

38.7%

22.7%

Stock Qwen3-4B (no fine-tune)

21.6%

39.2%

75.5%

95.3%

74.1%

16.6%

ZenPolish 1.7B v2.2Ships in ZenVoice

6.7%

53.6%

75.9%

97.8%

91.6%

2.2%

Added content = words in the output that were never spoken. ZenPolish 1.7B v2.2 makes a fifth of the word errors of Qwen3-1.7B, the stock model it is trained from, and adds 90% fewer words that were never said.

## Formatting WER by category

Same 500 dictations

Best value per row in bold

### Formatting WER by category

Word error rate on each kind of speech, same 500 dictations. Lower is better.

-   ZenPolish 1.7B v2.2
-   Stock Qwen3-1.7B (its base)
-   Stock Qwen3-4B

-   All dictations
    
    500 dictations
    
    6.7%
    
    1.7B v2.2
    
    −26.9 pts
    
    6.7%
    
    33.7%
    
    21.6%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Heavy fillers and restarts
    
    94 dictations
    
    3.9%
    
    1.7B v2.2
    
    −72.5 pts
    
    3.9%
    
    76.4%
    
    46.4%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Fillers
    
    43 dictations
    
    7.2%
    
    1.7B v2.2
    
    −63.4 pts
    
    7.2%
    
    70.6%
    
    22.2%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Facts: numbers, money, dates
    
    105 dictations
    
    15.2%
    
    1.7B v2.2
    
    −29.4 pts
    
    15.2%
    
    44.6%
    
    24.8%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Questions
    
    72 dictations
    
    0.3%
    
    1.7B v2.2
    
    −11.0 pts
    
    0.3%
    
    11.3%
    
    26.1%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Fragments
    
    62 dictations
    
    3.9%
    
    1.7B v2.2
    
    −0.9 pts
    
    3.9%
    
    4.8%
    
    4.7%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Exclamations
    
    60 dictations
    
    4.5%
    
    1.7B v2.2
    
    level
    
    4.5%
    
    4.8%
    
    4.0%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Run-ons
    
    64 dictations
    
    8.9%
    
    1.7B v2.2
    
    level
    
    8.9%
    
    8.4%
    
    7.7%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    

### Full table

Category

Rows

Stock Qwen3-1.7B

Stock Qwen3-4B

ZenPolish 1.7B v2.2

Heavy fillers and restarts

94

76.4%

46.4%

3.9%

Fillers

43

70.6%

22.2%

7.2%

Facts: numbers, money, dates

105

44.6%

24.8%

15.2%

Questions

72

11.3%

26.1%

0.3%

Fragments

62

4.8%

4.7%

3.9%

Exclamations

60

4.8%

4.0%

4.5%

Run-ons

64

8.4%

7.7%

8.9%

The weakest category is numbers, money and dates (15.2%); in the app, fixed rules write those. Run-ons are level with the stock models. On questions, fragments and exclamations every spoken word is kept, so word error rate only measures words a model damaged.
