Who said what, and when.
Native macOS transcription with speaker diarization. Drop in an interview or a recorded call and get a timestamped transcript that knows who was talking — on the Neural Engine, no account, no upload, no per-minute meter.
A reading room, not a waiting room.
Drop a file in the library, pick an engine, and the transcript builds while you keep working. Play it back and the transcript follows along, word by word. Below: the library, the queue with its parallel lanes, the engine manager, and the side-by-side comparison that reads two engines' takes on the same recording.




“Four people on a call and every other transcript reads like one person talking to themselves. This one hands me the turns, stamped, and never asks me to upload anything.”
Tell it what you care about.
Accuracy, download size and language coverage pull against each other, and no single engine wins all three. Cadence ships five and lets you switch per recording — each engine keeps its own transcript, so nothing is overwritten when you change your mind.
Built into macOS. Transcribes on first launch with nothing to fetch, in whichever locales the system already has.
The most accurate engine Cadence ships. English only — reach for v3 or Whisper if the recording isn't.
Nearly v2's accuracy across 25 European languages. The default choice for multilingual European work.
Two thirds the download of full v3 with the same language set. The sensible middle when disk is the constraint.
For recordings in something Parakeet never saw. It detects the language itself — nothing to configure.
Apple Speech has no language count on purpose: it transcribes whichever locales macOS has downloaded, you can add more in System Settings, and the set changes under the app. Cadence asks the system at launch rather than printing a number that goes stale.
Parakeet TDT 0.6B — NVIDIA (CC-BY-4.0) · Whisper large-v3-turbo — OpenAI, Core ML conversion by Argmax (MIT) · pyannote Community-1 — packaged by FluidInference (MIT)
Turn-taking is the subject.
Cadence separates the voices before it writes anything down, so the transcript arrives already attributed. 10.6% diarization error rate, 22 MB, any language — and it is part of the free app, not a paywalled extra.
Every other transcriber wants your audio. This one wants 22 MB.
The only network request Cadence ever makes is fetching the engine you chose, from Hugging Face, once. After that it works on a plane, on a locked-down network, and with a recording you are contractually not allowed to hand to anybody's cloud.
Pay once. Own the tool.
No subscription, no seats, no minutes. The free app is a complete transcriber — every engine, every language, speakers included. Pro is for the day you have thirty recordings instead of one.
A whole transcriber, not a trial. No watermark, no time limit, no sign-up.
For a backlog rather than a file. Queue the library, compare engines, export in whatever format the rest of your pipeline eats.
Questions, answered.
The ones that come up most. If yours isn't here, write to us.
Do I need an internet connection?
Only to install the app and to download an engine the first time you pick one. The default engine is already inside macOS, so a fresh install transcribes with nothing downloaded and nothing online.
Which Macs does it run on?
macOS 26.5 or later. The downloadable engines need Apple Silicon, since they run on the Neural Engine; Apple's built-in engine works everywhere.
Can I use it on confidential recordings?
That is the reason it exists. Nothing is uploaded, so there is no processor to add to a DPA and no terms of service quietly reserving the right to train on your interviews.
Does Cadence record?
No. It reads files you already have — audio or video, dropped into the window. Keep recording with whatever you record with.
What is actually in Pro?
Three things: Transcribe All, which works through every untranscribed recording in the library one at a time; model comparison, which runs one recording through several engines in a single action and shows two of them side by side; and the full export set — WebVTT, Markdown, JSON and CSV, with their options.
What's coming next?
Search across transcripts; a speaker library that recognises the same voices between recordings; Shortcuts actions; and per-recording settings that override the defaults. None of it ships today — listed because it's where Cadence is going.
Batch background removal for Mac, with marketplace presets, real edge control and readiness checks before you upload.
Prompt-driven image segmentation with COCO, YOLO and per-frame video masks. Dataset-ready, entirely on-device.
Who said what, and when. On your Mac.
Free forever for single-file transcription, every engine and every language included. Unlock the backlog with a one-time Pro upgrade.