Cadence
sec · 01 / heroturn-taking · resolved
New · macOS 26.5 · Apple Silicon

Who said what, and when.

Native macOS transcription with speaker diarization. Drop in an interview or a recorded call and get a timestamped transcript that knows who was talking — on the Neural Engine, no account, no upload, no per-minute meter.

Apple Silicon nativeSpeakers = 22 MBv1.0
research-call-04.m4aParakeet TDT v2 · 3 speakers · 47 min
00:00 / 01:56Live demo — click the waveform
Maya · interviewer
Dan · participant
Priya · participant
DER 10.6% · Diarization 22 MB
ExportTXTSRTVTT · ProMD · ProJSON · ProCSV · Pro
Click a line to jump
Core framework
Neural Engine + Core ML
Speech engines
Five engines
Hour of audio
Under a minute
Data leaves Mac
Never
sec · 02 / screenshots
Screenshots

A reading room, not a waiting room.

Drop a file in the library, pick an engine, and the transcript builds while you keep working. Play it back and the transcript follows along, word by word. Below: the library, the queue with its parallel lanes, the engine manager, and the side-by-side comparison that reads two engines' takes on the same recording.

The library — Every recording, its status, its engine.
01The libraryEvery recording, its status, its engine.
The queue — Transcription and speakers, in parallel lanes.
02The queueTranscription and speakers, in parallel lanes.
Engine manager — Download once, run offline.
03Engine managerDownload once, run offline.
Model comparison — Two engines, one recording. Pro.
04Model comparisonTwo engines, one recording. Pro.
sec · 03 / voice
“Four people on a call and every other transcript reads like one person talking to themselves. This one hands me the turns, stamped, and never asks me to upload anything.”
— Priya Shah
Staff researcher · Lisbon
sec · 04 / engines
Five engines

Tell it what you care about.

Accuracy, download size and language coverage pull against each other, and no single engine wins all three. Cadence ships five and lets you switch per recording — each engine keeps its own transcript, so nothing is overwritten when you change your mind.

I care most about
→ Use
Parakeet TDT v2. The lowest word error rate Cadence can reach, 5.4 on English, still entirely on the Neural Engine. If the recording isn't English, take the next option instead.Whisper large-v3-turbo. One hundred languages, detected automatically. It is a little slower and a little less accurate than Parakeet on English, and it is the only engine that will not be surprised by your recording.Apple Speech. Already inside macOS, so the app transcribes the moment it opens. It uses whichever locales the system has installed — add more in System Settings and Cadence picks them up.v3 compact. 299 MB for all 25 European languages and the fastest throughput of the downloadable engines. Accuracy sits a step below full v3, which most interview audio will not notice.
e-01DEFAULT
Apple Speech
no download

Built into macOS. Transcribes on first launch with nothing to fetch, in whichever locales the system already has.

Accuracygood
Speedinstant
Languagessystem
Footprintnone
e-02Recommended
Parakeet TDT v2
464 MB · EN

The most accurate engine Cadence ships. English only — reach for v3 or Whisper if the recording isn't.

Accuracy5.4 WER
Speedvery fast
LanguagesEnglish
Footprint464 MB
e-03Recommended
Parakeet TDT v3
483 MB · 25 EU

Nearly v2's accuracy across 25 European languages. The default choice for multilingual European work.

Accuracy5.7 WER
Speedvery fast
Languages25
Footprint483 MB
e-04Recommended
v3 compact
299 MB · 25 EU

Two thirds the download of full v3 with the same language set. The sensible middle when disk is the constraint.

Accuracygood
Speedfastest
Languages25
Footprint299 MB
e-05Recommended
Whisper large-v3-turbo
627 MB · 100 lang

For recordings in something Parakeet never saw. It detects the language itself — nothing to configure.

Accuracy7.0 WER
Speedfast
Languages100
Footprint627 MB

Apple Speech has no language count on purpose: it transcribes whichever locales macOS has downloaded, you can add more in System Settings, and the set changes under the app. Cadence asks the system at launch rather than printing a number that goes stale.

Parakeet TDT 0.6B — NVIDIA (CC-BY-4.0) · Whisper large-v3-turbo — OpenAI, Core ML conversion by Argmax (MIT) · pyannote Community-1 — packaged by FluidInference (MIT)

sec · 05 / speakers
Diarization

Turn-taking is the subject.

Cadence separates the voices before it writes anything down, so the transcript arrives already attributed. 10.6% diarization error rate, 22 MB, any language — and it is part of the free app, not a paywalled extra.

Fig. 01 — speaker turns over 47 minutes
Maya
Dan
Priya
00:00:0000:23:3100:47:12
What that buys you
Attribution, not guesswork. Each line carries the speaker it came from, in colours that stay consistent between the legend and the transcript.
Timecodes you can trust. Every turn is stamped, so a quote can be found again in the audio in seconds. Click one to play from there.
Play it back and read along. A transport under the transcript plays the file from 0.75× to 2×, marking the word as it is spoken and tinting the turn it belongs to.
Any language. Diarization works on the sound of the voices, so it doesn't care what they're saying.
Video too. Screen recordings of meetings go in the same window as audio files.
sec · 06 / on-device
Stays on your Mac

Every other transcriber wants your audio. This one wants 22 MB.

The only network request Cadence ever makes is fetching the engine you chose, from Hugging Face, once. After that it works on a plane, on a locked-down network, and with a recording you are contractually not allowed to hand to anybody's cloud.

No account. Nothing to sign in to before you can transcribe.
No per-minute billing, no quota, no queue behind other people's files.
No telemetry, no crash reporter, no “help improve the product”.
Sandboxed, notarized, App-Store-shipped. Your files stay in your container.
Fig. 02 — where the audio goes
File inresearch-call-04.m4a
ANEdiarize → transcribe → align
Disktranscript, on your machine
NetworkNot used
sec · 07 / pricing
Free, with a Pro unlock

Pay once. Own the tool.

No subscription, no seats, no minutes. The free app is a complete transcriber — every engine, every language, speakers included. Pro is for the day you have thirty recordings instead of one.

Free
$0always

A whole transcriber, not a trial. No watermark, no time limit, no sign-up.

+Unlimited single-file transcription
+All five engines, every language
+Speaker diarization
+Playback with word-level follow-along
+Re-run one recording on a second engine, free
+TXT, SRT and clipboard export
+Free updates, forever
Download free
Pro · one-timeLaunch price
$9.99$29.99pay once

For a backlog rather than a file. Queue the library, compare engines, export in whatever format the rest of your pipeline eats.

+Everything in Free
+WebVTT export, with options
+Transcribe All
+Markdown export
+Model comparison
+JSON export
+Two engines side by side
+CSV export
Get Pro — $9.99 on the Mac App Store
sec · 08 / faq
Honest answers

Questions, answered.

The ones that come up most. If yours isn't here, write to us.

Q · 01

Do I need an internet connection?

Only to install the app and to download an engine the first time you pick one. The default engine is already inside macOS, so a fresh install transcribes with nothing downloaded and nothing online.

Q · 02

Which Macs does it run on?

macOS 26.5 or later. The downloadable engines need Apple Silicon, since they run on the Neural Engine; Apple's built-in engine works everywhere.

Q · 03

Can I use it on confidential recordings?

That is the reason it exists. Nothing is uploaded, so there is no processor to add to a DPA and no terms of service quietly reserving the right to train on your interviews.

Q · 04

Does Cadence record?

No. It reads files you already have — audio or video, dropped into the window. Keep recording with whatever you record with.

Q · 05

What is actually in Pro?

Three things: Transcribe All, which works through every untranscribed recording in the library one at a time; model comparison, which runs one recording through several engines in a single action and shows two of them side by side; and the full export set — WebVTT, Markdown, JSON and CSV, with their options.

Q · 06

What's coming next?

Search across transcripts; a speaker library that recognises the same voices between recordings; Shortcuts actions; and per-recording settings that override the defaults. None of it ships today — listed because it's where Cadence is going.

sec · 09 / also
Also from Magenta Creations
Silhouette — clean cutouts, in bulk.

Batch background removal for Mac, with marketplace presets, real edge control and readiness checks before you upload.

Contour — segment anything, on your Mac.

Prompt-driven image segmentation with COCO, YOLO and per-frame video masks. Dataset-ready, entirely on-device.

sec · 10 / download

Who said what, and when. On your Mac.

Free forever for single-file transcription, every engine and every language included. Unlock the backlog with a one-time Pro upgrade.

macOS 26.5Apple Siliconv1.0