September 25, 2026
ZenPolish 1.7B v2.2
ZenVoice’s own model for turning spoken drafts into clean writing. On your Mac.
(1)
Introduction
Speech doesn’t come out as writing. We say “um”, repeat words, restart sentences, and say “two thousand dollars” where we’d type $2,000. ZenPolish is ZenVoice’s own model for the step between what you said and what you meant to write.
Today we’re introducing ZenPolish 1.7B v2.2. It removes fillers and false starts, fixes punctuation and capitalization, and writes numbers, dates, times and money the way you’d type them, while keeping your words. It runs entirely on your Mac, inside ZenVoice, with no network at inference time.
On 500 held-out dictations it makes a fifth of the word errors of the stock model it’s trained from, and adds 90% fewer words you never said. The biggest gains are in the hardest speech: fillers, restarts and thinking out loud.
- −80%
- word errors, against its own stock base model
- −90%
- words you never said, against the stock model
- 54%
- outputs exactly right, against 14% for the stock model
- 1.8 GB
- memory while it runs, on your Mac
(2)
Performance
We measure ZenPolish on 500 held-out dictations across seven kinds of speech: our original 234-row suite and 266 new rows built the same way. Every row was checked against the training data, and the numbers in it don’t overlap the training templates, so the model can’t have memorised the answers.
Fine-tuning is what does the work, not size: ZenPolish 1.7B v2.2 cuts its own stock base model’s word error rate from 33.7% to 6.7%, a fifth of it, and beats Qwen3-4B, a stock model more than twice its size, on every measure.
| Metric | ZenPolish 1.7B v2.2ours | Qwen3-1.7Bits stock base | Qwen3-4Ba larger stock model |
|---|---|---|---|
| Word error ratelower is better | 6.7% | 33.7% | 21.6% |
| Added contentwords never spoken · lower is better | 2.2% | 22.7% | 16.6% |
| Exact matchoutput identical to reference | 53.6% | 13.8% | 39.2% |
| Punctuation F1token-level | 75.9% | 35.2% | 75.5% |
| Capitalizationaccuracy | 97.8% | 85.1% | 95.3% |
| Sentence startsaccuracy | 91.6% | 38.7% | 74.1% |
Best value per row in bold. Word error rate is word-only alignment against a hand-written reference; punctuation and capitalization are token-level; all metrics macro-averaged across rows. Every model runs exactly as ZenVoice loads it, with the same instructions, Apple Silicon (24 GB), MLX Metal, greedy decoding. Measured September 25, 2026.
Smart formatting
500 held-out dictations · 7 categories · leakage-checked. Lower is better.
- ZenPolish 1.7B v2.2ZenVoice6.7%
- Stock Qwen3-4B (no fine-tune)21.6%
- Stock Qwen3-1.7B (no fine-tune)33.7%
Formatting WER by category
Word error rate on each kind of speech, same 500 dictations. Lower is better.
- ZenPolish 1.7B v2.2
- Stock Qwen3-1.7B (its base)
- Stock Qwen3-4B
All dictations
500 dictations
−26.9 pts6.7%
1.7B v2.2
6.7%33.7%21.6%1.7B v2.2Qwen3-1.7BQwen3-4BHeavy fillers and restarts
94 dictations
−72.5 pts3.9%
1.7B v2.2
3.9%76.4%46.4%1.7B v2.2Qwen3-1.7BQwen3-4BFillers
43 dictations
−63.4 pts7.2%
1.7B v2.2
7.2%70.6%22.2%1.7B v2.2Qwen3-1.7BQwen3-4BFacts: numbers, money, dates
105 dictations
−29.4 pts15.2%
1.7B v2.2
15.2%44.6%24.8%1.7B v2.2Qwen3-1.7BQwen3-4BQuestions
72 dictations
−11.0 pts0.3%
1.7B v2.2
0.3%11.3%26.1%1.7B v2.2Qwen3-1.7BQwen3-4BFragments
62 dictations
−0.9 pts3.9%
1.7B v2.2
3.9%4.8%4.7%1.7B v2.2Qwen3-1.7BQwen3-4BExclamations
60 dictations
level4.5%
1.7B v2.2
4.5%4.8%4.0%1.7B v2.2Qwen3-1.7BQwen3-4BRun-ons
64 dictations
level8.9%
1.7B v2.2
8.9%8.4%7.7%1.7B v2.2Qwen3-1.7BQwen3-4B
Where the gains are
| Category | Qwen3-1.7B (stock base) | ZenPolish 1.7B v2.2 | Change |
|---|---|---|---|
| Heavy fillers and restarts | 76.4% | 3.9% | −72.5 pts |
| Fillers | 70.6% | 7.2% | −63.4 pts |
| Questions | 11.3% | 0.3% | −11.0 pts |
| Fragments (1–5 words) | 4.8% | 3.9% | −0.9 pts |
| Exclamations | 4.8% | 4.5% | level |
| Run-ons | 8.4% | 8.9% | level |
| Facts: numbers, money, dates | 44.6% | 15.2% | −29.4 pts |
Word error rate per category, lower is better; the last column is the change in percentage points against the stock model v2.2 is trained from. Speech full of fillers and restarts is where it pulls furthest ahead: 3.9% against 76.4%. Run-ons are level. On questions, fragments and exclamations every spoken word is kept, so word error rate there only counts words a model damaged.
(3)
Before and after
Here is what that looks like on real rows from the evaluation, next to the stock model it’s trained from. The difference shows most when speech is messy: the stock model keeps the “um” and “you know”; v2.2 finds the sentence underneath.
What you said · Slack
“so the thing is um yeah like the the partner leads are you know warm”
Qwen3-1.7B, stock
so the thing is um yeah like the partner leads are you know warm
ZenPolish 1.7B v2.2
The partner leads are warm.
Real outputs from the 500-dictation evaluation. ZenPolish and the stock model it’s trained from got the same input and instructions.
Where it still falls short
v2.2 is weakest on numbers and dates, where its word error rate is 15.2%; in ZenVoice, fixed rules write money, times and dates before and after the model, so they come out right. It also sometimes leaves the full stop or question mark off a short sentence.
You said “my bonus was two thousand dollars”
Expected
My bonus was $2,000.
ZenPolish 1.7B v2.2
My bonus was two thousand dollars.
Leaves money in words. Numbers and dates are its weakest area; in the app, fixed rules write them.
You said “can you send me the link”
Expected
Can you send me the link?
ZenPolish 1.7B v2.2
can you send me the link
Sometimes drops the capital and the question mark on a short question.
You said “standup moved to ten fifteen”
Expected
Standup moved to 10:15.
ZenPolish 1.7B v2.2
Standup moved to 15:15.
Rarely, it writes the wrong time. We saw this once in 500 rows.
It leaves clean text alone
We also test on 120 real transcripts that are already readable prose, where the best a model can do is leave them alone. There, v2.2 changes no more than the stock models do, and adds almost nothing that wasn’t said.
| Model | Word error rate ↓ | Added content ↓ |
|---|---|---|
| ZenPolish 1.7B v2.2 | 11.0% | 0.9% |
| ZenPolish 1.7B v2 (previous) | 7.6% | 0.9% |
| Qwen3-1.7B (stock) | 11.0% | 0.9% |
| Qwen3-4B (stock) | 14.6% | 1.5% |
120 real dictation transcripts that are already readable prose, scored against the original sentences. Lower means the model changed less of text that didn’t need changing.
(4)
Speed and size
A typical dictation is a sentence or two. ZenPolish 1.7B v2.2 writes it in about 0.4 seconds on Apple Silicon, and needs 1.8 GB of memory while it runs.
| Model | Per dictation | First token | Output speed | Peak memory | Download |
|---|---|---|---|---|---|
| ZenPolish 1.7B v2.2 | 0.40 s | 0.20 s | 42 tok/s | 1.77 GB | 1.55 GB |
bench_speed_v500.py · 50 dictations · two passes, averaged · other GPU work was running, so absolute times are pessimistic. Apple Silicon (24 GB), MLX Metal, greedy decoding.
(5)
How it's built
ZenPolish 1.7B v2.2 is a LoRA fine-tune of Alibaba’s Qwen3-1.7B, trained at rank 32 with a cosine learning-rate schedule. It ships as a 6-bit MLX base with the adapter applied live when it loads. We don’t merge it into the weights: merging measurably hurt money formatting and filler removal.
The training set has about 90,000 examples from ten sources, including parliamentary transcripts, earnings calls, podcasts, conversational question answering and dialogue from books, alongside speech-disfluency data. For v2.2 we also made the labels consistent where earlier data disagreed with itself: fillers in the middle of a sentence, money, exclamations and casing.
| Spec | ZenPolish 1.7B v2.2 |
|---|---|
| Base model | Qwen3-1.7B |
| Parameters | 1.7 B |
| Fine-tuning | LoRA rank 32, applied live on the base |
| Quantization | 6-bit MLX |
| Download | 1.5 GB |
| Memory while running | 1.8 GB |
| Training examples | ~90,000 from 10 sources |
| Licence | Apache-2.0 |
Private by construction
ZenPolish runs in-process inside ZenVoice. Your dictation never leaves your Mac to be formatted, and the model needs no network once it’s downloaded. The download itself is pinned to a reviewed revision and checked by SHA-256 before it loads.
(6)
Availability
ZenPolish 1.7B v2.2 arrives in ZenVoice 0.5.0. Choose it in Models → Dictation enhancement; it’s a 1.5 GB download, alongside Apple Intelligence and Qwen 3.5 2B.
The weights are open under Apache-2.0 on Hugging Face.
