Skip to content
ZenVoice

September 25, 2026

ZenPolish 1.7B v2.2

ZenVoice’s own model for turning spoken drafts into clean writing. On your Mac.

(1)

Introduction

Speech doesn’t come out as writing. We say “um”, repeat words, restart sentences, and say “two thousand dollars” where we’d type $2,000. ZenPolish is ZenVoice’s own model for the step between what you said and what you meant to write.

Today we’re introducing ZenPolish 1.7B v2.2. It removes fillers and false starts, fixes punctuation and capitalization, and writes numbers, dates, times and money the way you’d type them, while keeping your words. It runs entirely on your Mac, inside ZenVoice, with no network at inference time.

On 500 held-out dictations it makes a fifth of the word errors of the stock model it’s trained from, and adds 90% fewer words you never said. The biggest gains are in the hardest speech: fillers, restarts and thinking out loud.

−80%
word errors, against its own stock base model
−90%
words you never said, against the stock model
54%
outputs exactly right, against 14% for the stock model
1.8 GB
memory while it runs, on your Mac

(2)

Performance

We measure ZenPolish on 500 held-out dictations across seven kinds of speech: our original 234-row suite and 266 new rows built the same way. Every row was checked against the training data, and the numbers in it don’t overlap the training templates, so the model can’t have memorised the answers.

Fine-tuning is what does the work, not size: ZenPolish 1.7B v2.2 cuts its own stock base model’s word error rate from 33.7% to 6.7%, a fifth of it, and beats Qwen3-4B, a stock model more than twice its size, on every measure.

ZenPolish 1.7B v2.2 compared with two stock models
MetricZenPolish 1.7B v2.2oursQwen3-1.7Bits stock baseQwen3-4Ba larger stock model
Word error ratelower is better6.7%33.7%21.6%
Added contentwords never spoken · lower is better2.2%22.7%16.6%
Exact matchoutput identical to reference53.6%13.8%39.2%
Punctuation F1token-level75.9%35.2%75.5%
Capitalizationaccuracy97.8%85.1%95.3%
Sentence startsaccuracy91.6%38.7%74.1%

Best value per row in bold. Word error rate is word-only alignment against a hand-written reference; punctuation and capitalization are token-level; all metrics macro-averaged across rows. Every model runs exactly as ZenVoice loads it, with the same instructions, Apple Silicon (24 GB), MLX Metal, greedy decoding. Measured September 25, 2026.

Smart formatting

500 held-out dictations · 7 categories · leakage-checked. Lower is better.

  • ZenPolish 1.7B v2.2ZenVoice6.7%
  • Stock Qwen3-4B (no fine-tune)21.6%
  • Stock Qwen3-1.7B (no fine-tune)33.7%

Formatting WER by category

Word error rate on each kind of speech, same 500 dictations. Lower is better.

  • ZenPolish 1.7B v2.2
  • Stock Qwen3-1.7B (its base)
  • Stock Qwen3-4B
  • All dictations

    500 dictations

    6.7%

    1.7B v2.2

    −26.9 pts
    6.7%
    33.7%
    21.6%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Heavy fillers and restarts

    94 dictations

    3.9%

    1.7B v2.2

    −72.5 pts
    3.9%
    76.4%
    46.4%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Fillers

    43 dictations

    7.2%

    1.7B v2.2

    −63.4 pts
    7.2%
    70.6%
    22.2%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Facts: numbers, money, dates

    105 dictations

    15.2%

    1.7B v2.2

    −29.4 pts
    15.2%
    44.6%
    24.8%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Questions

    72 dictations

    0.3%

    1.7B v2.2

    −11.0 pts
    0.3%
    11.3%
    26.1%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Fragments

    62 dictations

    3.9%

    1.7B v2.2

    −0.9 pts
    3.9%
    4.8%
    4.7%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Exclamations

    60 dictations

    4.5%

    1.7B v2.2

    level
    4.5%
    4.8%
    4.0%
    1.7B v2.2Qwen3-1.7BQwen3-4B
  • Run-ons

    64 dictations

    8.9%

    1.7B v2.2

    level
    8.9%
    8.4%
    7.7%
    1.7B v2.2Qwen3-1.7BQwen3-4B

Where the gains are

Word error rate per speech category
CategoryQwen3-1.7B (stock base)ZenPolish 1.7B v2.2Change
Heavy fillers and restarts76.4%3.9%−72.5 pts
Fillers70.6%7.2%−63.4 pts
Questions11.3%0.3%−11.0 pts
Fragments (1–5 words)4.8%3.9%−0.9 pts
Exclamations4.8%4.5%level
Run-ons8.4%8.9%level
Facts: numbers, money, dates44.6%15.2%−29.4 pts

Word error rate per category, lower is better; the last column is the change in percentage points against the stock model v2.2 is trained from. Speech full of fillers and restarts is where it pulls furthest ahead: 3.9% against 76.4%. Run-ons are level. On questions, fragments and exclamations every spoken word is kept, so word error rate there only counts words a model damaged.

(3)

Before and after

Here is what that looks like on real rows from the evaluation, next to the stock model it’s trained from. The difference shows most when speech is messy: the stock model keeps the “um” and “you know”; v2.2 finds the sentence underneath.

What you said · Slack

“so the thing is um yeah like the the partner leads are you know warm”

Qwen3-1.7B, stock

so the thing is um yeah like the partner leads are you know warm

ZenPolish 1.7B v2.2

The partner leads are warm.

Real outputs from the 500-dictation evaluation. ZenPolish and the stock model it’s trained from got the same input and instructions.

Where it still falls short

v2.2 is weakest on numbers and dates, where its word error rate is 15.2%; in ZenVoice, fixed rules write money, times and dates before and after the model, so they come out right. It also sometimes leaves the full stop or question mark off a short sentence.

You said “my bonus was two thousand dollars”

Expected

My bonus was $2,000.

ZenPolish 1.7B v2.2

My bonus was two thousand dollars.

Leaves money in words. Numbers and dates are its weakest area; in the app, fixed rules write them.

You said “can you send me the link”

Expected

Can you send me the link?

ZenPolish 1.7B v2.2

can you send me the link

Sometimes drops the capital and the question mark on a short question.

You said “standup moved to ten fifteen”

Expected

Standup moved to 10:15.

ZenPolish 1.7B v2.2

Standup moved to 15:15.

Rarely, it writes the wrong time. We saw this once in 500 rows.

It leaves clean text alone

We also test on 120 real transcripts that are already readable prose, where the best a model can do is leave them alone. There, v2.2 changes no more than the stock models do, and adds almost nothing that wasn’t said.

Clean-input test: how much each model changes text that is already clean
ModelWord error rate ↓Added content ↓
ZenPolish 1.7B v2.211.0%0.9%
ZenPolish 1.7B v2 (previous)7.6%0.9%
Qwen3-1.7B (stock)11.0%0.9%
Qwen3-4B (stock)14.6%1.5%

120 real dictation transcripts that are already readable prose, scored against the original sentences. Lower means the model changed less of text that didn’t need changing.

(4)

Speed and size

A typical dictation is a sentence or two. ZenPolish 1.7B v2.2 writes it in about 0.4 seconds on Apple Silicon, and needs 1.8 GB of memory while it runs.

Latency, throughput and memory
ModelPer dictationFirst tokenOutput speedPeak memoryDownload
ZenPolish 1.7B v2.20.40 s0.20 s42 tok/s1.77 GB1.55 GB

bench_speed_v500.py · 50 dictations · two passes, averaged · other GPU work was running, so absolute times are pessimistic. Apple Silicon (24 GB), MLX Metal, greedy decoding.

(5)

How it's built

ZenPolish 1.7B v2.2 is a LoRA fine-tune of Alibaba’s Qwen3-1.7B, trained at rank 32 with a cosine learning-rate schedule. It ships as a 6-bit MLX base with the adapter applied live when it loads. We don’t merge it into the weights: merging measurably hurt money formatting and filler removal.

The training set has about 90,000 examples from ten sources, including parliamentary transcripts, earnings calls, podcasts, conversational question answering and dialogue from books, alongside speech-disfluency data. For v2.2 we also made the labels consistent where earlier data disagreed with itself: fillers in the middle of a sentence, money, exclamations and casing.

Model specifications
SpecZenPolish 1.7B v2.2
Base modelQwen3-1.7B
Parameters1.7 B
Fine-tuningLoRA rank 32, applied live on the base
Quantization6-bit MLX
Download1.5 GB
Memory while running1.8 GB
Training examples~90,000 from 10 sources
LicenceApache-2.0

Private by construction

ZenPolish runs in-process inside ZenVoice. Your dictation never leaves your Mac to be formatted, and the model needs no network once it’s downloaded. The download itself is pinned to a reviewed revision and checked by SHA-256 before it loads.

(6)

Availability

ZenPolish 1.7B v2.2 arrives in ZenVoice 0.5.0. Choose it in Models → Dictation enhancement; it’s a 1.5 GB download, alongside Apple Intelligence and Qwen 3.5 2B.

The weights are open under Apache-2.0 on Hugging Face.

Ask AI about this pageChatGPT ↗Claude ↗Perplexity ↗Grok ↗Markdown