Zerm

34 releases

Changelog

Every published release, in full. Newest first.

v2.8.6 Latest Download .dmg 32.7 MB

2.8.6

The biggest Zerm update yet: dictation always uses the model you chose, AI enhancement is rebuilt, Meetings is replaced by file transcription with speaker identification, and the model catalog and the whole app are refreshed, in English and Hebrew.

Fixed

  • The right model, every time. Power Modes no longer lock in the dictation model, language and prompt that happened to be selected when they were created, so dictation stops switching models depending on which app is in front. Every Power Mode setting now shows "Use global setting" unless you choose otherwise, and a Power Mode never changes your global settings.
  • Parakeet keeps your last words. Stopping right after you finish speaking no longer cuts the end off. Switching between Whisper models loads the new one once instead of on every dictation.
  • AI enhancement, rebuilt. One Output setting decides what happens: Instant, Instant + Refine, or Enhanced. Each dictation is enhanced with exactly the prompt, provider and model it started with, the timeout setting works, and when enhancement is skipped or fails Zerm tells you why instead of silently pasting the raw text. Prompts can now use their own provider and model.
  • Dashboard time saved and totals follow the range you select.
  • Escape double-press to cancel is reliable again.

New

  • Transcribe File. Drop in audio or video files and get a full transcript with speaker identification, timestamps and renameable speakers. Export to TXT, Markdown, SRT, WebVTT or JSON, open any audio file with Zerm from Finder, and reopen transcripts from History.
  • A wider, current model catalog.
    • New on-device models: Parakeet Unified (English), Parakeet 110M for 8 GB Macs, and ivrit.ai Whisper models tuned for Hebrew.
    • Recommended picks the best English, multilingual and Hebrew models for your Mac. Local models can be filtered by English only, Multilingual and Great in Hebrew.
  • Cloud models are up to date and filterable by provider.
    • Current models: OpenAI gpt-transcribe, Gemini 3.5 Transcribe, AssemblyAI Universal-3.5 Pro and Mistral Voxtral, plus the new Gladia provider.
    • Model settings such as Output Format and Dictionary terms are sent the way each provider supports them.
    • Custom OpenAI-compatible endpoints are verified when you save them.
  • History has a copy button on every row.
  • Hebrew throughout the app.

Removed

  • Meetings. Replaced by Transcribe File. Meeting recordings saved by earlier versions are permanently deleted when 2.8.6 first launches. Dictation recordings and History are not affected.
  • Outdated models. Whisper tiny, base and small and Parakeet V2 are retired. If you used one, Zerm moves you to its replacement automatically.

Compatibility

  • macOS 14.4 or later
  • Apple silicon (arm64)

Zerm remains local-first. Audio leaves the Mac only when you choose a cloud provider.

v2.8.5 Download .dmg 24.2 MB

2.8.5

Read Aloud was failing on its most common path. This release makes it always speak, and speak every language you have a voice for.

Fixed

  • Read Aloud works again. Selecting text and pressing the hotkey now always speaks: if the on-device AI rewrite cannot produce a trustworthy version, Zerm reads the selection exactly and tells you why instead of failing silently.
  • Read Aloud speaks every language you have a voice for. Hebrew, Russian, Spanish and any other text is detected and routed to a matching installed Apple voice automatically — even when Kokoro or Deepgram, which are English-only, is your chosen provider.
  • "Read exactly" is the default reading mode. The AI modes (Retell, Summarize, Explain, Simplify) remain available in Read Aloud settings, and selected text is treated strictly as content to be spoken, never as instructions.
  • Longer selections keep their reading instructions when Enhancement and Read Aloud share the same on-device model.

Compatibility

  • macOS 14.4 or later
  • Apple silicon (arm64)

Zerm remains local-first. Audio leaves the Mac only when you choose a cloud provider.

v2.8.4 Download .dmg 24.2 MB

2.8.4

Enhancement was completely broken in 2.8.3 and this release fixes it, along with meeting recording.

Fixed

  • AI enhancement works again. On-device enhancement was returning nothing at all and pasting the raw transcript instead.
  • History rows that appeared blank now show their text. No transcript was ever lost — the words were always there, the row just drew the empty enhancement over them. Affected rows are repaired when you launch this version.
  • Enhancement is about 3.5x faster. Short dictation now finishes in roughly a tenth of a second.
  • Mixed-language dictation keeps every word in the language it was spoken in. English, Hebrew and Russian in one sentence stay that way.
  • Enhancement no longer answers a question you dictated, and no longer pastes stray markup into your document.
  • The on-device model list has been rebuilt on measurement. Models that produced empty or unusable output are gone, and are removed from your Mac to reclaim the space.
  • Meeting recordings capture the call correctly. The call track could previously end up half its real length and drift out of sync with your microphone.
  • Meetings transcribe after you stop instead of during the call, so recording no longer competes with everything else for CPU.
  • The installer window is properly branded.

Compatibility

  • macOS 14.4 or later
  • Apple silicon (arm64)

Zerm remains local-first. Audio leaves the Mac only when you choose a cloud provider.

v2.8.3 Download .dmg 28.2 MB

2.8.3

What’s new

  • Recordings are saved even if transcription or AI fails. Retry from History.
  • English and mixed dictation no longer collapse to Hebrew. Enhancement keeps the original script.
  • Instant + Refine pastes once. Enhanced no longer copies the selection after you finish speaking.
  • Enhancement defaults to Qwen3. Gemma stays on Read Aloud.
  • Word replacements respect Unicode word boundaries. Microphone detection no longer under-allocates the Core Audio buffer list.

Compatibility

  • macOS 14.4 or later
  • Apple silicon (arm64)

Zerm remains local-first. Audio leaves the Mac only when you choose a cloud provider.

v2.8.2 Download .dmg 28.2 MB

2.8.2

What’s new

  • Rebuilt Meetings with explicit Room, selected-app, and all-system-audio capture; durable recovery; speaker attribution; and synchronized playback.
  • Improved Hebrew and automatic language transcription across local and supported cloud models.
  • Rebuilt Read Aloud with Retell, Summarize, Explain, and Simplify modes, local history, and improved Hebrew voice support.
  • Fixed repeated clipboard operations, safer in-place replacement, and clearer paste/refinement feedback.
  • Improved technical terminology for software, cloud, codecs, models, and CLI tools.
  • Reduced local-AI memory pressure with the official memory-efficient Gemma 4 E2B model, idle unloading, and hardware-aware recommendations. Larger models remain manual opt-ins.
  • Reorganized navigation, Settings, permissions, and recording preflight for a clearer native macOS workflow.
  • Added stability, privacy, accessibility, localization, and release-infrastructure fixes throughout the app.

Compatibility

  • macOS 14.4 or later
  • Apple silicon (arm64)

Zerm remains local-first. Audio is sent off-device only when you explicitly select a cloud transcription or AI provider.

v2.8.1

What's new

  • The Recording tab is rebuilt as two dedicated screens — Record for the meeting in progress, Library for everything you have recorded — instead of one crowded pane.
  • Capture, transcript and automation settings moved into their own panel behind a Settings button, so the recording screen shows only what is happening now.
  • The Library is searchable, and shows each recording's length, speakers, size and summary at a glance. Open any of them for playback with the transcript following along.
  • Speaker colours in a transcript are stable and distinct — the same voice keeps its colour between launches, instead of every speaker coming out the same shade.
  • A transcript no longer highlights its first line as if it were playing before you press play.

Install

Download Zerm_2.8.1_aarch64.dmg (Apple silicon). The build is Developer ID signed and notarised by Apple.

Existing installs update themselves through Sparkle.

2.8.0 — Meeting recording

Records whole conversations — the room through your microphone, the call through your Mac — and transcribes them as they happen.

New: the Recording tab

  • Two separate tracks. Your side and theirs are captured independently, not mixed, so the transcript can tell them apart and either side can be re-transcribed on its own.
  • Live transcription while the meeting is still running, using whichever model you already use for dictation.
  • Speaker identification for the people in the room with you, on-device.
  • Summaries with action items and chapters when the meeting ends, through local Ollama by default — a meeting transcript never leaves your Mac unless you choose a cloud model.
  • Offers to record when Zoom, Teams, Meet, Signal, Slack or FaceTime opens.
  • Import existing recordings — wav, mp3, m4a, mp4, flac, ogg, aac, caf, aiff.
  • Playback with a synced transcript. Click any line to jump to the moment it was said.
  • Survives a crash. Transcript lines are written as they land, so an interrupted meeting keeps everything up to the last few seconds.

System audio is captured with a Core Audio process tap: no driver to install and no screen-recording permission. It does need the Audio Recording permission, which macOS will ask for the first time you record a call.

Also in this release

  • The dashboard no longer stutters when scrolled quickly.
  • Power Mode's app picker no longer walks through symlinked directories, which could make it hang.
  • Apple's built-in dictation model now downloads missing language assets instead of failing with an error you could not act on.

2.7.1

Headphones

Dictating while wearing headphones no longer mutes your audio.

Muting exists for one reason: stopping audio from your speakers bleeding into the microphone and being transcribed as words you never said. On headphones that cannot happen, so the mute was cutting your music every time you dictated, for no benefit.

Bluetooth headsets and the built-in headphone jack are detected as headphones. Speakers, USB audio interfaces, HDMI and AirPlay keep muting, because those really can bleed into the mic. You can turn the behaviour off under Mute Audio While Recording → Keep Playing on Headphones.

Fixed

  • Transcribing a recorded audio file failed on some MP4 and M4A files, including Teams meeting recordings.
  • Large uploads to cloud transcription providers could stall until they timed out when connected to a VPN.
  • Dictation could hang at startup if a browser was unresponsive or showing an automation permission prompt.
  • Text from a finished recording could briefly appear in the next one.
  • A memory bug in audio-device handling, a data race in the recorder, and several errors that were being silently swallowed.
  • The site changelog was not being published on release.

Under the hood

Ports the relevant fixes from upstream VoiceInk v2.0 and v2.1, and clears every compiler warning in the project (110 to 0).

Installation: download Zerm_2.7.1_aarch64.dmg. Apple Silicon, macOS 14.4 or later. Signed and notarized.

2.7.0 — honest metrics, working enhancement, real docs

Your dashboard no longer resets

Statistics were recomputed from your surviving transcripts, while the cleanup service deletes those on a retention timer defaulting to one day. So the numbers only ever reflected whatever history had not been swept yet.

They now live in their own store that transcript retention cannot reach, holding counts and durations only — no transcript text. Your existing history is imported once on first launch. There is a new Reset Statistics button in Privacy settings, because clearing transcripts used to clear the numbers as a side effect and that was the only way to erase them.

Also fixed: if you had auto-delete enabled but had never picked a retention period, the cutoff was computed as now and your entire transcript history was deleted at launch.

AI enhancement works again

Enhancement was switched off product-wide by a hidden setting with no interface, which reset your toggle on every launch and blocked the pipeline even when the toggle read on. The same setting had pinned the response timeout to two seconds, so anything that did run timed out and quietly pasted the raw transcript.

Enhancement was only slow because it ran before the paste. It now runs after, so you can have both:

  • Instant — paste immediately, no AI. The fastest path.
  • Instant + Refine — paste immediately, then improve the text in place a moment later.
  • Enhanced — wait for the AI, paste once.

Refine rewrites text in place in native macOS apps such as Mail, Notes and TextEdit. In Electron apps (Slack, VS Code, Cursor, Notion), browser text fields and terminals, macOS does not allow it — there, the refined version is offered to you instead and what you already typed is left untouched. Nothing is ever silently rewritten in an app that cannot support it.

Dictation speed is unchanged: measured at 85 ms from hotkey to recording.

Every setting explains itself

Around a hundred options now carry an info icon, each linking to a real documentation page. Previously every documentation link in the app pointed at a domain that was never registered.

Also

  • The support button was opening a mail draft addressed to an unrelated person, with your system information attached. It now opens the project issue tracker.
  • Power Modes no longer silently override your global enhancement setting; each one can inherit it, force it on, or force it off.
  • The recorder's enhancement indicator is now legible at a glance.

Requires macOS 14.4 or later. Apple Silicon.
Signed with a Developer ID certificate and notarized by Apple.

2.6.1

Fixes three things that made the app feel broken, plus a large latency cut.

Fixed

  • Shortcut recording did not work anywhere in the app. Pressing the modifier you wanted to bind fired Zerm's own hotkey instead, which could raise an invisible modal sheet and deadlock the app. Zerm's triggers now stand down while you are recording a shortcut.
  • Read Aloud could stop working for an entire session. Left and right Option were indistinguishable, so the key-state tracking could stick 'pressed' and never fire again.
  • Quitting the app crashed with an abort in the Metal teardown of the speech engine.
  • Unsigned/local builds crashed at launch in CloudKit.
  • An empty AI enhancement no longer replaces your transcript with nothing.
  • With language set to Auto, the Whisper initial prompt is no longer silently discarded.

Faster

before after
hotkey → recording 431 ms 79–86 ms
stop → transcription starts 270–449 ms 20–42 ms
dictation after a >2 min gap ~800 ms model reload none

Read Aloud starts sooner and no longer stalls after the first sentence. The on-device language model is no longer loaded at launch, cutting several GB of idle memory.

2.5.1

A cleaner menu bar

  • New tray icon. The menu bar now shows the Zerm microphone — the same glyph as the app icon — rendered as a template image, so it stays crisp on both light and dark menu bars.
  • Decluttered tray menu. The model, prompt, AI provider, AI model, language, and audio-input pickers no longer crowd the menu; all of them remain available in the main window. The menu keeps quick actions: Toggle Recorder, AI Enhancement, Retry/Copy Last Transcription, History, and Settings.

Install: download the DMG, drag Zerm to Applications. The app is Developer ID signed and notarized. Existing users on 2.4.0+ will be offered this update automatically.

2.5.0

Smarter on every Mac

  • Hardware-aware model recommendations. Zerm now rates your Mac (chip + memory) and tunes the Recommended model list to it. Models that would strain your machine show a clear warning, and models that need more memory than your Mac can spare can no longer be installed.
  • Faster, cooler local transcription. Whisper now runs on your Mac's performance cores only (efficiency cores were slowing it down), and automatically backs off under Low Power Mode or thermal pressure.
  • Less battery drain. Model prewarming now skips Low Power Mode and brief lid-close cycles, and the recording audio meter uses half the wakeups — visually identical.

Permissions that just work

  • The Permissions screen now detects when macOS has granted Screen Recording but needs a relaunch, and offers a one-click "Quit Zerm to Finish".
  • Permission cards refresh automatically after you return from System Settings.
  • Enabling Accessibility now shows the proper system prompt so the grant sticks to the right binary.

Install: download the DMG, drag Zerm to Applications. The app is Developer ID signed and notarized. Existing users on 2.4.0+ will be offered this update automatically.

2.4.0

What's new

Sidebar

  • Renamed the AI Models tab to Dictation Models (it manages your speech-to-text models — Read Aloud has its own).
  • Removed the unused Transcribe Audio tab.

Transcription models

  • Dropped slow, outdated local models and added a compact Small option.
  • Refreshed cloud models and added OpenAI and AssemblyAI as new providers.

Enhancement

  • The on-device model is now a downloadable list — choose from several Google Gemma tiers (1B through 27B) for local, private enhancement.
  • New Coding and Chat prompts: Coding cleans dictated code and technical talk (without writing or answering), Chat keeps casual messages natural.

Install

  • New install: download Zerm_2.4.0_aarch64.dmg (Apple notarized, Apple Silicon).
  • Existing users: update in place via Zerm's built-in updater.

2.3.3

Faster dictation start

Pressing the dictation hotkey now shows live audio in about 150 ms, down from ~610 ms — roughly 4x faster.

The recorder's waveform also animates as soon as the microphone is actually live. Previously it appeared instantly but sat flat and lifeless for over half a second, then jumped into sync, because the UI switched to "recording" before the audio hardware had started.

Details

Two issues on the record-start path, both fixed:

  • The microphone permission check was scheduled behind the recorder window's first draw, costing ~311 ms on every trigger even when permission was already granted — and delaying audio startup behind it. This was a regression introduced in 2.2.0 alongside the (correct) addition of a real permission check.
  • The waveform was mounted against a silent audio meter ~283 ms before any sound existed, rendering bars identical to the idle state until real audio arrived.

Measured over 115 real recording starts before the change, and re-measured after.

Install

  • Existing users: update in place via Zerm's built-in updater.
  • New install: download Zerm_2.3.3_aarch64.dmg below (Apple notarized, Apple Silicon).

2.3.2

Fixes

  • Sticky “Update available” banner after install/check — only shows when the feed is strictly newer than the running build; clears on session end, abort, or relaunch.
  • Empty transcripts with a working mic (e.g. Yeti) — Whisper retries without VAD when the first pass is empty but the clip has energy; better empty-result messaging.

Install

Open Zerm_2.3.2_aarch64.dmg or update from 2.3.1 via Check for Updates / sidebar banner.

2.3.1

Highlights

Auto-update that actually surfaces in the UI.

  • Sparkle background checks publish updateAvailable state
  • Sidebar shows a bottom Update available card with Install update
  • Settings and menu bar show the same pending-update action
  • Reliable find / not-found / error logging for field diagnosis

Upgrade path

  • From 2.3.0: Check for Updates (or wait for the background check) — you should see 2.3.1 offered and a sidebar banner.
  • From 2.2.x or earlier: Install this DMG once (EdDSA key was rotated in 2.3.0). Future updates work automatically.

Install

  1. Open Zerm_2.3.1_aarch64.dmg and drag Zerm to Applications.
  2. Apple Silicon only. Developer ID signed and notarized.

2.3.0

Highlights

Deep-review release: privacy, accuracy, stability, UX, and a working auto-update feed.

Privacy

  • Transcripts are no longer written in cleartext to the macOS unified log (character counts only)
  • Audio retention defaults on (14 days)

Accuracy

  • Custom dictionary reaches local Whisper and Deepgram batch
  • Continuous phase-carrying resampler + smarter multi-channel mix
  • Language default auto
  • Safer hallucination filter (keeps legitimate parentheticals)
  • Streaming failures fall back to full-file batch transcription
  • Offline dictation commands: “new line”, “scratch that”, spoken punctuation

Stability & UX

  • System mute no longer sticks after a quick cancel during the start sound
  • Session tokens prevent stale paste; busy-state watchdog
  • Hard transcription failures show a real error notification (with deep links where useful)
  • Model must be downloaded / API key present before recording starts
  • Mic test in Settings; history workspace (edit, re-transcribe, local AI actions)

Ops

  • First real unit tests
  • MetricKit diagnostics
  • Sparkle appcast signed for auto-updates (new installs of 2.3.0+)

Install

  1. Open Zerm_2.3.0_aarch64.dmg and drag Zerm to Applications.
  2. First launch: approve the standard macOS prompts (microphone, accessibility).

Apple Silicon (arm64) only. Developer ID signed and notarized.

2.2.0

Highlights

  • First notarized release. The DMG is Developer ID signed and notarized by Apple — no more "Zerm is damaged" warnings and no xattr workaround needed. If a previous copy crashed at launch on download, this release fixes that too.
  • Microphone permission is now checked properly. If macOS has microphone access denied (or it was reset), Zerm now tells you and points to System Settings instead of silently recording nothing and reporting "Nothing transcribed".
  • Debug Logging mode. If recordings ever come back empty on your machine, enable Settings → Diagnostics → Debug Logging, reproduce once, and send us the log (Show in Finder reveals it). It records capture health — device, format, audio levels, driver errors — but never your words: transcripts are logged as character counts only, and no audio is kept.

Fixes

  • Recording failures that previously vanished without a trace (audio-driver render errors, dropped devices, silent input) are now detected and reported in the debug log, with a per-recording summary that pinpoints the cause.
  • Release pipeline: all embedded helper executables are now Developer ID signed with hardened runtime, fixing the launch crash some users hit with the 2.1.3 download.

Install

  1. Open Zerm_2.2.0_aarch64.dmg and drag Zerm to Applications.
  2. First launch: approve the standard macOS prompts (microphone, accessibility).

Apple Silicon (arm64) only.

2.1.3

Two chronic bugs fixed — hands-free dictation losing recordings, and Read Aloud misbehaving in terminals.

Fixed

Hands-free dictation no longer loses recordings on auto-stop

When the silence auto-stop ended a recording, the transcription and paste sometimes never happened — the capture simply vanished (or, on older versions, left a "Transcription Failed: Swift.CancellationError" history entry). Root cause: the auto-stop monitor triggered the stop from inside its own task, and the stop's first step cancelled that task — so the entire stop → transcribe → paste pipeline ran in an already-cancelled task and discarded the capture as if the user had cancelled. The monitor now triggers the stop from a fresh task. Manual hotkey stops were never affected.

Read Aloud now works reliably in terminals and TUI apps

Two independent causes:

  • The simulated ⌘C used to grab the selection was posted while the Read Aloud hotkey's modifier keys were still physically held, so terminals received ⌃⌥⌘C instead of Copy — and the pasteboard was only polled for 100 ms, too short for embedded terminals and Electron panes. Zerm now waits for the hotkey to be released and retries with its own ⌘C (immune to held modifiers) with a longer pasteboard window, restoring your clipboard afterward.
  • The on-device smart-reading model (Gemma) occasionally spoke a self-introduction ("…an advanced model by Google") instead of the selection. That happened when a terminal selection reduced to symbol-only junk (TUI borders), which provoked the model into chatting — and the phrasing slipped past the output validator. Symbol-only text is no longer sent to the model, and any output mentioning the model's identity is rejected in favor of the plainly-cleaned text.

Install

Download Zerm_2.1.3_aarch64.dmg, drag Zerm to Applications, then clear the quarantine flag (build is not notarized yet):

xattr -dr com.apple.quarantine /Applications/Zerm.app

Grant Accessibility and Screen Recording when prompted (Settings → Permissions).

2.1.2

Fixes

  • Long recordings no longer drop mid-dictation. The silence auto-stop window was too short (1.1s), so a natural thinking pause could cut a long recording off. It's now 2.5s — forgiving of pauses, still prompt when you're done.
  • No more stuck "recording already running" after a drop. If the microphone is dropped mid-recording (unplugged, USB glitch, audio glitch), Zerm now detects the dropped capture within a few seconds, transcribes whatever was captured, tells you what happened, and returns to idle — so the next press starts a fresh recording immediately instead of being ignored.
  • Selecting text in terminals/TUIs now works. Reading or handling highlighted text in terminal apps and full-screen terminal UIs (e.g. Claude Code, cmux) reported "No text selected". Added a clipboard-based fallback (⌘C, auto-restored) that those apps handle reliably. This also restores Read-Aloud of selected text in those apps.

Install

  1. Open Zerm_2.1.2_aarch64.dmg and drag Zerm to Applications.
  2. Ad-hoc signed (not notarized). If macOS blocks first launch:
    xattr -dr com.apple.quarantine /Applications/Zerm.app
    

Apple Silicon (arm64) only.

2.1.1

Fixes

  • Transcriptions no longer fail spuriously with "The operation couldn't be completed. (Swift.CancellationError error 1.)" — when a transcription was cancelled mid-flight (starting a new recording, toggling the hotkey again, engine teardown, or the safety timeout), the app mistakenly recorded it as a failure. Cancellation is now handled cleanly: the empty pending entry is discarded and no scary "Transcription Failed" record is written.

Install

  1. Open Zerm_2.1.1_aarch64.dmg and drag Zerm to Applications.
  2. This build is ad-hoc signed (not notarized). On first launch, if macOS blocks it:
    xattr -dr com.apple.quarantine /Applications/Zerm.app
    
    then open it again.

Apple Silicon (arm64) only.

v2.1.0 — Smart Read Aloud + on-device AI

Zerm v2.1.0 — Smart Read Aloud + Zerm's third on-device model.

Install (important)

This build is signed ad-hoc (not yet notarized), so macOS Gatekeeper warns on first launch. After opening the DMG and dragging Zerm to Applications:

Terminal (recommended):

xattr -dr com.apple.quarantine /Applications/Zerm.app
open /Applications/Zerm.app

Or right-click Zerm.app → Open → Open.

Then grant Accessibility, Screen Recording, and Microphone in System Settings → Privacy & Security. Apple Silicon only.

What's new

Three on-device models — Zerm now downloads and manages three local AI models: Whisper (speech-to-text), Kokoro (text-to-speech), and Gemma 4 (the agentic LLM). Each runs fully offline; cloud providers remain optional.

Read Aloud — select text anywhere, press your shortcut, and Zerm reads it in a natural voice (local or cloud), using the same animated widget as dictation.

Smart Reading — text is cleaned instantly (acronyms, URLs, code, errors, emoji, tables) and can be rewritten by the on-device LLM so it sounds human, not robotic. The widget shows the real phase (Thinking… / Preparing… / speaking).

On-device AI Enhancement — dictation cleanup can run on the local Gemma model, no API key, recommended by default.

Notes

  • No auto-update in this build (requires a Sparkle signing key). Download new releases manually.
  • A fully notarized, warning-free build is planned once Developer ID notarization is set up.

Full architecture docs: https://github.com/arcusis/Zerm/wiki

v2.0.1 — Read Aloud fixes & instant feel

Patch release on top of v2.0.0, focused on Read Aloud.

🐛 Fixed

  • App froze when stopping Read Aloud. The audio-level metering tap removed itself from two threads at once (your stop + the playback-finished handler), which deadlocked AVAudioEngine and hung the whole app. The tap is now installed once and never removed in those paths.

⚡ Faster / instant feel

  • Instant widget: the recorder widget now appears the moment you trigger Read Aloud — before fetching the selected text or synthesizing — matching dictation's immediacy.
  • Chunked streaming: the first sentence plays as soon as it's synthesized while the rest synthesize in the background (no more waiting for the whole selection).
  • On-device pre-warm: the Kokoro model loads in the background when selected, so the first read is instant instead of a cold load.

Merged PRs

#231 (stop-hang fix + streaming), #232 (instant widget + 2.0.1). Epic #220.

v2.0.0 — Read Aloud (TTS)

✨ Read Aloud (text-to-speech) — the mirror of dictation

Select text anywhere, press your trigger, and Zerm reads it aloud. Configure in Settings → Read Aloud.

Voices

  • On-device: Kokoro 82M (sherpa-onnx) — download once (~330 MB), then fully offline, no API key.
  • Cloud: Deepgram Aura-2 (default) · Inworld TTS-1.5 Mini · ElevenLabs v3 · Gemini 3.1 Flash · OpenAI gpt-4o-mini-tts · Cartesia Sonic-3.5.

Trigger + widget

  • Modifier-key trigger dropdown (Right Command, Right Option, Fn, …) just like the dictation hotkey, plus a Custom key-combo option.
  • Triggering Read Aloud shows the same animated mini/notch widget as dictation (a "Speaking" state), with double-Escape to cancel.
  • Fully mutually exclusive with dictation — they can't run at once, no collisions.

Dashboard

  • New Words Read Aloud stat alongside the dictation metrics.

Fixes included

  • Voxtral streaming 400 (#36) and Whisper hallucination/temperature (#150).
  • Fixed the recurring stuck/unfocusable window in menu-bar-only mode (#223).

Merged PRs

#206, #221, #222, #223, #224, #225, #226 · Epic #220.


Status: Draft pending a notarized Developer ID build + Sparkle appcast entry (notary issuer ID / Sparkle EdDSA key required to publish the auto-updatable DMG). All source is on Production, tagged v2.0.0.

Coming next: local downloadable models for AI Enhancement (#227).

v1.0.4

Zerm v1.0.4

Maintenance release syncing the relevant fix from upstream VoiceInk v1.79 plus two transcription fixes.

Fixes

  • macOS 26 paste crash (#204). The AppleScript paste path ran on a background queue, but the TIS input-source APIs assert they must run on the main queue on macOS 26, causing a SIGTRAP crash on paste. Now runs on the main actor (matches upstream VoiceInk v1.79's fix for #737).
  • Deepgram "Auto detect" language. Auto-detect sent no language parameter, so Deepgram silently defaulted to English. It now maps auto → multi for the multilingual Nova-3 model in both batch and streaming paths.
  • Idle transcription failure. The first FluidAudio inference after the app sat idle could fail with "Unable to compute asynchronous prediction" (stale CoreML handles after memory paging). Added a one-shot reload-from-disk and retry.

Notes

Brings Zerm current with upstream VoiceInk stable v1.79.

v1.0.3

First launch

  1. Open Zerm. The dashboard shows a setup banner.
  2. Zerm streams the hash-pinned Whisper model (~466 MB)
    into your app-data dir automatically.
  3. Click Install Ollama. Zerm downloads the signed
    installer from github.com/ollama/ollama/releases and
    hands it to your OS's own signature check before it
    runs.
  4. Gemma 3 4B is then pulled through your local Ollama.
  5. macOS: grant Accessibility and Microphone permission when prompted.

Use it

Tap Right Option (macOS) / Ctrl+Shift+Space (Windows

  • Linux), speak, stop talking. Text lands in your clipboard.

Notes

  • macOS builds are signed with Developer ID. Stable releases wait for notarization + stapling; prereleases submit for notarization with --skip-stapling so Apple polling outages do not fail the release.
  • macOS builds include the audio-input entitlement required for microphone capture.
  • Windows installers are Authenticode signed when WINDOWS_CERTIFICATE is configured. Prereleases continue unsigned if that certificate is absent.
  • Stable release jobs fail if required signing credentials are not configured.
  • Linux .AppImage built on Ubuntu 22.04.
  • SHA-256 checksums are attached as SHA256SUMS.txt.

Changes

https://github.com/arcusis/Zerm/commits/v1.0.3

v1.0.2

First launch

  1. Open Zerm. The dashboard shows a setup banner.
  2. Zerm streams the hash-pinned Whisper model (~466 MB)
    into your app-data dir automatically.
  3. Click Install Ollama. Zerm downloads the signed
    installer from github.com/ollama/ollama/releases and
    hands it to your OS's own signature check before it
    runs.
  4. Gemma 3 4B is then pulled through your local Ollama.
  5. macOS: grant Accessibility and Microphone permission when prompted.

Use it

Tap Right Option (macOS) / Ctrl+Shift+Space (Windows

  • Linux), speak, stop talking. Text lands in your clipboard.

Notes

  • macOS builds are signed with Developer ID. Stable releases wait for notarization + stapling; prereleases submit for notarization with --skip-stapling so Apple polling outages do not fail the release.
  • macOS builds include the audio-input entitlement required for microphone capture.
  • Windows installers are Authenticode signed when WINDOWS_CERTIFICATE is configured. Prereleases continue unsigned if that certificate is absent.
  • Stable release jobs fail if required signing credentials are not configured.
  • Linux .AppImage built on Ubuntu 22.04.
  • SHA-256 checksums are attached as SHA256SUMS.txt.

Changes

https://github.com/arcusis/Zerm/commits/v1.0.2

v1.0.1

First launch

  1. Open Zerm. The dashboard shows a setup banner.
  2. Zerm streams the hash-pinned Whisper model (~466 MB)
    into your app-data dir automatically.
  3. Click Install Ollama. Zerm downloads the signed
    installer from github.com/ollama/ollama/releases and
    hands it to your OS's own signature check before it
    runs.
  4. Gemma 3 4B is then pulled through your local Ollama.
  5. macOS: grant Accessibility and Microphone permission when prompted.

Use it

Tap Right Option (macOS) / Ctrl+Shift+Space (Windows

  • Linux), speak, stop talking. Text lands in your clipboard.

Notes

  • macOS builds are signed with Developer ID. Stable releases wait for notarization + stapling; prereleases submit for notarization with --skip-stapling so Apple polling outages do not fail the release.
  • macOS builds include the audio-input entitlement required for microphone capture.
  • Windows installers are Authenticode signed when WINDOWS_CERTIFICATE is configured. Prereleases continue unsigned if that certificate is absent.
  • Stable release jobs fail if required signing credentials are not configured.
  • Linux .AppImage built on Ubuntu 22.04.
  • SHA-256 checksums are attached as SHA256SUMS.txt.

Changes

https://github.com/arcusis/Zerm/commits/v1.0.1

v1.0.0

First launch

  1. Open Zerm. The dashboard shows a setup banner.
  2. Zerm streams the hash-pinned Whisper model (~466 MB)
    into your app-data dir automatically.
  3. Click Install Ollama. Zerm downloads the signed
    installer from github.com/ollama/ollama/releases and
    hands it to your OS's own signature check before it
    runs.
  4. Gemma 3 4B is then pulled through your local Ollama.
  5. macOS: grant Accessibility and Microphone permission when prompted.

Use it

Tap Right Option (macOS) / Ctrl+Shift+Space (Windows

  • Linux), speak, stop talking. Text lands in your clipboard.

Notes

  • macOS builds are signed with Developer ID. Stable releases wait for notarization + stapling; prereleases submit for notarization with --skip-stapling so Apple polling outages do not fail the release.
  • macOS builds include the audio-input entitlement required for microphone capture.
  • Windows installers are Authenticode signed when WINDOWS_CERTIFICATE is configured. Prereleases continue unsigned if that certificate is absent.
  • Stable release jobs fail if required signing credentials are not configured.
  • Linux .AppImage built on Ubuntu 22.04.
  • SHA-256 checksums are attached as SHA256SUMS.txt.

Changes

https://github.com/arcusis/Zerm/commits/v1.0.0

v0.1.0-alpha.18 Pre-release Download .dmg 4.3 MB

v0.1.0-alpha.18

First launch

  1. Open Zerm. The dashboard shows a setup banner.
  2. Zerm streams the hash-pinned Whisper model (~466 MB)
    into your app-data dir automatically.
  3. Click Install Ollama. Zerm downloads the signed
    installer from github.com/ollama/ollama/releases and
    hands it to your OS's own signature check before it
    runs.
  4. Gemma 3 4B is then pulled through your local Ollama.
  5. macOS: grant Accessibility and Microphone permission when prompted.

Use it

Tap Right Option (macOS) / Ctrl+Shift+Space (Windows

  • Linux), speak, stop talking. Text lands in your clipboard.

Notes

  • macOS builds are signed with Developer ID. Stable releases wait for notarization + stapling; prereleases submit for notarization with --skip-stapling so Apple polling outages do not fail the release.
  • macOS builds include the audio-input entitlement required for microphone capture.
  • Windows installers are Authenticode signed when WINDOWS_CERTIFICATE is configured. Prereleases continue unsigned if that certificate is absent.
  • Stable release jobs fail if required signing credentials are not configured.
  • Linux .AppImage built on Ubuntu 22.04.
  • SHA-256 checksums are attached as SHA256SUMS.txt.

Changes

https://github.com/arcusis/Zerm/commits/v0.1.0-alpha.18

v0.1.0-alpha.5 Pre-release Download .dmg 4.1 MB

v0.1.0-alpha.5

First launch

  1. Open Zerm. The dashboard appears with a "Set up Zerm" banner.
  2. Click Download Whisper model in the banner — Zerm streams
    the multilingual ggml-small.bin (~466 MB) into your app data
    directory and loads it automatically.
  3. Install Ollama, then in a
    terminal: ollama pull gemma3:4b. Zerm picks it up live.
  4. macOS: grant Accessibility permission when prompted so the
    Right Option hotkey can fire and auto-paste can simulate ⌘V.

Use it

Tap Right Option, speak, stop talking. The cleaned text
lands in your clipboard and pastes into whatever app is focused.

Notes

  • macOS builds are unsigned for this prerelease — first launch
    requires xattr -cr /Applications/Zerm.app or right-click → Open.
  • Windows .msi is unsigned; SmartScreen will warn — click "More
    info" → "Run anyway".
  • Linux .AppImage is built on Ubuntu 22.04 — should run on
    22.04+.

Changes

https://github.com/arcusis/Zerm/commits/v0.1.0-alpha.5

v0.1.0-alpha.4 Pre-release Download .dmg 4.1 MB

v0.1.0-alpha.4

First launch

  1. Open Zerm. The dashboard appears with a "Set up Zerm" banner.
  2. Click Download Whisper model in the banner — Zerm streams
    the multilingual ggml-small.bin (~466 MB) into your app data
    directory and loads it automatically.
  3. Install Ollama, then in a
    terminal: ollama pull gemma3:4b. Zerm picks it up live.
  4. macOS: grant Accessibility permission when prompted so the
    Right Option hotkey can fire and auto-paste can simulate ⌘V.

Use it

Tap Right Option, speak, stop talking. The cleaned text
lands in your clipboard and pastes into whatever app is focused.

Notes

  • macOS builds are unsigned for this prerelease — first launch
    requires xattr -cr /Applications/Zerm.app or right-click → Open.
  • Windows .msi is unsigned; SmartScreen will warn — click "More
    info" → "Run anyway".
  • Linux .AppImage is built on Ubuntu 22.04 — should run on
    22.04+.

Changes

https://github.com/arcusis/Zerm/commits/v0.1.0-alpha.4

v0.1.0-alpha.3 Pre-release Download .dmg 4.1 MB

v0.1.0-alpha.3

First launch

  1. Open Zerm. The dashboard appears with a "Set up Zerm" banner.
  2. Click Download Whisper model in the banner — Zerm streams
    the multilingual ggml-small.bin (~466 MB) into your app data
    directory and loads it automatically.
  3. Install Ollama, then in a
    terminal: ollama pull gemma3:4b. Zerm picks it up live.
  4. macOS: grant Accessibility permission when prompted so the
    Right Option hotkey can fire and auto-paste can simulate ⌘V.

Use it

Tap Right Option, speak, stop talking. The cleaned text
lands in your clipboard and pastes into whatever app is focused.

Notes

  • macOS builds are unsigned for this prerelease — first launch
    requires xattr -cr /Applications/Zerm.app or right-click → Open.
  • Windows .msi is unsigned; SmartScreen will warn — click "More
    info" → "Run anyway".
  • Linux .AppImage is built on Ubuntu 22.04 — should run on
    22.04+.

Changes

https://github.com/arcusis/Zerm/commits/v0.1.0-alpha.3

v0.1.0-alpha.2 Pre-release Download .dmg 4.0 MB

v0.1.0-alpha.2

First launch

  1. Open Zerm. The dashboard appears with a "Set up Zerm" banner.
  2. Click Download Whisper model in the banner — Zerm streams
    the multilingual ggml-small.bin (~466 MB) into your app data
    directory and loads it automatically.
  3. Install Ollama, then in a
    terminal: ollama pull gemma3:4b. Zerm picks it up live.
  4. macOS: grant Accessibility permission when prompted so the
    Right Option hotkey can fire and auto-paste can simulate ⌘V.

Use it

Tap Right Option, speak, stop talking. The cleaned text
lands in your clipboard and pastes into whatever app is focused.

Notes

  • macOS builds are unsigned for this prerelease — first launch
    requires xattr -cr /Applications/Zerm.app or right-click → Open.
  • Windows .msi is unsigned; SmartScreen will warn — click "More
    info" → "Run anyway".
  • Linux .AppImage is built on Ubuntu 22.04 — should run on
    22.04+.

Changes

https://github.com/arcusis/Zerm/commits/v0.1.0-alpha.2