# ZenPolish 1.7B v2.2 · ZenVoice

Source: https://getzenvoice.com/zenpolish

> ZenVoice's own on-device model for turning spoken drafts into clean writing: a fifth of the word errors of its stock base model and 90% fewer words you never said, in 1.8 GB of memory.

![](/films/zenpolish-end.jpg)

Skip ↓

September 25, 2026

# ZenPolish 1.7B v2.2

ZenVoice’s own model for turning spoken drafts into clean writing. On your Mac.

1.  (1)Introduction
2.  (2)Performance
3.  (3)Before and after
4.  (4)Speed and size
5.  (5)How it's built
6.  (6)Availability

September 25, 2026

# ZenPolish 1.7B v2.2

ZenVoice’s own model for turning spoken drafts into clean writing. On your Mac.

1.  (1)Introduction
2.  (2)Performance
3.  (3)Before and after
4.  (4)Speed and size
5.  (5)How it's built
6.  (6)Availability

(1)

## Introduction

Speech doesn’t come out as writing. We say “um”, repeat words, restart sentences, and say “two thousand dollars” where we’d type **$2,000**. ZenPolish is ZenVoice’s own model for the step between what you said and what you meant to write.

Today we’re introducing **ZenPolish 1.7B v2.2**. It removes fillers and false starts, fixes punctuation and capitalization, and writes numbers, dates, times and money the way you’d type them, while keeping your words. It runs entirely on your Mac, inside ZenVoice, with no network at inference time.

On 500 held-out dictations it makes **a fifth of the word errors** of the stock model it’s trained from, and adds **90% fewer** words you never said. The biggest gains are in the hardest speech: fillers, restarts and thinking out loud.

−80%

word errors, against its own stock base model

−90%

words you never said, against the stock model

54%

outputs exactly right, against 14% for the stock model

1.8 GB

memory while it runs, on your Mac

(2)

## Performance

We measure ZenPolish on 500 held-out dictations across seven kinds of speech: our original 234-row suite and 266 new rows built the same way. Every row was checked against the training data, and the numbers in it don’t overlap the training templates, so the model can’t have memorised the answers.

Fine-tuning is what does the work, not size: ZenPolish 1.7B v2.2 cuts its own stock base model’s word error rate from 33.7% to 6.7%, a fifth of it, and beats Qwen3-4B, a stock model more than twice its size, on every measure.

ZenPolish 1.7B v2.2ours

Qwen3-1.7Bits stock base

Qwen3-4Ba larger stock model

Word error ratelower is better

6.7%

33.7%

21.6%

Added contentwords never spoken · lower is better

2.2%

22.7%

16.6%

Exact matchoutput identical to reference

53.6%

13.8%

39.2%

Punctuation F1token-level

75.9%

35.2%

75.5%

Capitalizationaccuracy

97.8%

85.1%

95.3%

Sentence startsaccuracy

91.6%

38.7%

74.1%

Best value per row in bold. Word error rate is word-only alignment against a hand-written reference; punctuation and capitalization are token-level; all metrics macro-averaged across rows. Every model runs exactly as ZenVoice loads it, with the same instructions, Apple Silicon (24 GB), MLX Metal, greedy decoding. Measured September 25, 2026.

### Smart formatting

500 held-out dictations · 7 categories · leakage-checked. Lower is better.

-   ZenPolish 1.7B v2.2ZenVoice6.7%
-   Stock Qwen3-4B (no fine-tune)21.6%
-   Stock Qwen3-1.7B (no fine-tune)33.7%

### Formatting WER by category

Word error rate on each kind of speech, same 500 dictations. Lower is better.

-   ZenPolish 1.7B v2.2
-   Stock Qwen3-1.7B (its base)
-   Stock Qwen3-4B

-   All dictations
    
    500 dictations
    
    6.7%
    
    1.7B v2.2
    
    −26.9 pts
    
    6.7%
    
    33.7%
    
    21.6%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Heavy fillers and restarts
    
    94 dictations
    
    3.9%
    
    1.7B v2.2
    
    −72.5 pts
    
    3.9%
    
    76.4%
    
    46.4%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Fillers
    
    43 dictations
    
    7.2%
    
    1.7B v2.2
    
    −63.4 pts
    
    7.2%
    
    70.6%
    
    22.2%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Facts: numbers, money, dates
    
    105 dictations
    
    15.2%
    
    1.7B v2.2
    
    −29.4 pts
    
    15.2%
    
    44.6%
    
    24.8%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Questions
    
    72 dictations
    
    0.3%
    
    1.7B v2.2
    
    −11.0 pts
    
    0.3%
    
    11.3%
    
    26.1%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Fragments
    
    62 dictations
    
    3.9%
    
    1.7B v2.2
    
    −0.9 pts
    
    3.9%
    
    4.8%
    
    4.7%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Exclamations
    
    60 dictations
    
    4.5%
    
    1.7B v2.2
    
    level
    
    4.5%
    
    4.8%
    
    4.0%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    
-   Run-ons
    
    64 dictations
    
    8.9%
    
    1.7B v2.2
    
    level
    
    8.9%
    
    8.4%
    
    7.7%
    
    1.7B v2.2Qwen3-1.7BQwen3-4B
    

### Where the gains are

Category

Qwen3-1.7B (stock base)

ZenPolish 1.7B v2.2

Change

Heavy fillers and restarts

76.4%

3.9%

−72.5 pts

Fillers

70.6%

7.2%

−63.4 pts

Questions

11.3%

0.3%

−11.0 pts

Fragments (1–5 words)

4.8%

3.9%

−0.9 pts

Exclamations

4.8%

4.5%

level

Run-ons

8.4%

8.9%

level

Facts: numbers, money, dates

44.6%

15.2%

−29.4 pts

Word error rate per category, lower is better; the last column is the change in percentage points against the stock model v2.2 is trained from. Speech full of fillers and restarts is where it pulls furthest ahead: 3.9% against 76.4%. Run-ons are level. On questions, fragments and exclamations every spoken word is kept, so word error rate there only counts words a model damaged.

(3)

## Before and after

Here is what that looks like on real rows from the evaluation, next to the stock model it’s trained from. The difference shows most when speech is messy: the stock model keeps the “um” and “you know”; v2.2 finds the sentence underneath.

What you said · Slack

“so the thing is um yeah like the the partner leads are you know warm”

Qwen3-1.7B, stock

so the thing is um yeah like the partner leads are you know warm

ZenPolish 1.7B v2.2

The partner leads are warm.

Real outputs from the 500-dictation evaluation. ZenPolish and the stock model it’s trained from got the same input and instructions.

### Where it still falls short

v2.2 is weakest on numbers and dates, where its word error rate is 15.2%; in ZenVoice, fixed rules write money, times and dates before and after the model, so they come out right. It also sometimes leaves the full stop or question mark off a short sentence.

You said “my bonus was two thousand dollars”

Expected

My bonus was $2,000.

ZenPolish 1.7B v2.2

My bonus was two thousand dollars.

Leaves money in words. Numbers and dates are its weakest area; in the app, fixed rules write them.

You said “can you send me the link”

Expected

Can you send me the link?

ZenPolish 1.7B v2.2

can you send me the link

Sometimes drops the capital and the question mark on a short question.

You said “standup moved to ten fifteen”

Expected

Standup moved to 10:15.

ZenPolish 1.7B v2.2

Standup moved to 15:15.

Rarely, it writes the wrong time. We saw this once in 500 rows.

### It leaves clean text alone

We also test on 120 real transcripts that are already readable prose, where the best a model can do is leave them alone. There, v2.2 changes no more than the stock models do, and adds almost nothing that wasn’t said.

Model

Word error rate ↓

Added content ↓

ZenPolish 1.7B v2.2

11.0%

0.9%

ZenPolish 1.7B v2 (previous)

7.6%

0.9%

Qwen3-1.7B (stock)

11.0%

0.9%

Qwen3-4B (stock)

14.6%

1.5%

120 real dictation transcripts that are already readable prose, scored against the original sentences. Lower means the model changed less of text that didn’t need changing.

(4)

## Speed and size

A typical dictation is a sentence or two. ZenPolish 1.7B v2.2 writes it in about 0.4 seconds on Apple Silicon, and needs 1.8 GB of memory while it runs.

Model

Per dictation

First token

Output speed

Peak memory

Download

ZenPolish 1.7B v2.2

0.40 s

0.20 s

42 tok/s

1.77 GB

1.55 GB

bench\_speed\_v500.py · 50 dictations · two passes, averaged · other GPU work was running, so absolute times are pessimistic. Apple Silicon (24 GB), MLX Metal, greedy decoding.

(5)

## How it's built

ZenPolish 1.7B v2.2 is a LoRA fine-tune of Alibaba’s **Qwen3-1.7B**, trained at rank 32 with a cosine learning-rate schedule. It ships as a 6-bit MLX base with the adapter applied live when it loads. We don’t merge it into the weights: merging measurably hurt money formatting and filler removal.

The training set has about **90,000** examples from ten sources, including parliamentary transcripts, earnings calls, podcasts, conversational question answering and dialogue from books, alongside speech-disfluency data. For v2.2 we also made the labels consistent where earlier data disagreed with itself: fillers in the middle of a sentence, money, exclamations and casing.

ZenPolish 1.7B v2.2

Base model

Qwen3-1.7B

Parameters

1.7 B

Fine-tuning

LoRA rank 32, applied live on the base

Quantization

6-bit MLX

Download

1.5 GB

Memory while running

1.8 GB

Training examples

~90,000 from 10 sources

Licence

Apache-2.0

### Private by construction

ZenPolish runs in-process inside ZenVoice. Your dictation never leaves your Mac to be formatted, and the model needs no network once it’s downloaded. The download itself is pinned to a reviewed revision and checked by SHA-256 before it loads.

(6)

## Availability

ZenPolish 1.7B v2.2 arrives in **ZenVoice 0.5.0**. Choose it in **Models → Dictation enhancement**; it’s a 1.5 GB download, alongside Apple Intelligence and Qwen 3.5 2B.

The weights are open under Apache-2.0 on Hugging Face.

[Get ZenVoice](https://getzenvoice.com/#download)[ZenPolish 1.7B v2.2 on Hugging Face ↗](https://huggingface.co/imYChaudhary22/zen-polish-v22-4bit)
