Choose the diarizing model
Select GPT-4o Transcribe Diarize and add your own OpenAI API key. This is an optional cloud model: the recording is sent directly to OpenAI for transcription.
Meeting recording · macOS
Capture your mic and the call’s system audio together. Transcribe locally by default, or choose optional cloud speaker labels — without a bot joining the meeting.
VoiceToText records Zoom, Google Meet, Microsoft Teams, FaceTime, or any call playing through your Mac, keeps running in the background while you work, and transcribes everything locally on the Apple Neural Engine by default — saved to your on-device history.
How it captures the call
Most meeting tools join your call as a participant. VoiceToText doesn’t. It records locally from macOS itself: your microphone plus the system audio coming out of your Mac, mixed into one recording.
Start a recording from the Conversations pane, then carry on. VoiceToText streams the audio straight to disk in the background — no RAM bloat, no window to babysit — and transcribes it the moment you stop, on-device by default. It works with whatever is making sound:
Speaker 1: Let's kick off the weekly sync.
Speaker 2: Design is signed off, and
the team is on the last endpoint.
Speaker 1: Any blockers before Friday?
Speaker 2: None. I'll send the notes
right after this call.Speaker diarization
Local Whisper and Parakeet models produce a standard transcript. When you need to distinguish speakers, the optional GPT-4o Transcribe Diarize model adds labeled turns that you can rename.
Select GPT-4o Transcribe Diarize and add your own OpenAI API key. This is an optional cloud model: the recording is sent directly to OpenAI for transcription.
The transcript is split into turns labeled Speaker 1, Speaker 2, and so on, instead of one uninterrupted block of meeting text.
Open Name speakers on the saved recording and assign display names. Those names appear when you view or copy the transcript and persist across transcript versions.
Three steps
Open Conversations in VoiceToText and click Start Recording. The first time, macOS asks once for Microphone and Screen Recording — that's what lets it hear you and capture system audio.
It records in the background, streaming straight to disk. Switch apps, take notes, share your screen — the recording keeps running with barely any footprint.
Stop, and the recording is transcribed with your chosen model — on-device by default, long calls in segments — and saved to your history with the audio you can replay and a transcript you can copy.
What you get
No second app subscription and no plugin in the call. Choose a local model to keep transcription on your Mac, or explicitly select a cloud provider when you want one of its features.
Records both directions at once through ScreenCaptureKit, so everyone on the call is captured — not just your side.
Local models transcribe the recording on your Mac. With a local engine, your audio never leaves the device — transcription makes zero network calls.
Recordings are saved on-device with audio and transcript — a rolling history of your 200 most recent. Play back, copy, favorite; deletes come with an undo, and a crash-interrupted recording is recovered on the next launch.
Re-run the transcript with a more accurate model and keep both versions side by side, so you can pick the better one and drop the other.
Already have a recording? Drop in any audio or video file — VoiceToText extracts the audio and transcribes it the same way.
Nothing joins your call, and VoiceToText does not charge by the minute. Local transcription has no provider fee; optional cloud providers may charge under your API key.
Permissions & privacy
Recording a conversation is sensitive, so VoiceToText keeps it local by default. Here is exactly what it needs and where your audio goes.
FAQ
Ready to record
One DMG, drag to Applications, grant Microphone and Screen Recording. Then record any call and get an on-device transcript.
The same app also does hotkey dictation into any text field. See everything VoiceToText does →