POUR LES ÉQUIPESDictée privée sur l'appareil pour toute votre équipe — sièges supplémentaires dès 15 $.
A PRACTICAL, HONEST GUIDE

How to transcribe audio to text

There are three real ways to convert audio to text — built-in phone tools, free online converters, and on-device apps. Each has a place. This guide walks through all of them honestly, then shows you the fastest, most private method for long or sensitive recordings.

The three ways to transcribe audio to text

Before the how-to, here's the honest landscape. The right choice depends on how long your recording is, how sensitive it is, and how often you do this.

1. Built-in tools

Your phone's Voice Memos, live transcript features, and some note apps can turn short speech into text instantly. Great for quick memos — but limited to brief clips, and quality and privacy vary by app.

2. Free online converters

Upload an audio file to a website and get text back in seconds. Convenient for a one-off short clip — but your recording is sent to their server, and you'll usually hit length or monthly-minute limits.

3. On-device apps

A desktop app like Whisper transcribes the file directly on your computer. No upload, no length cap, works offline — the right call for long, confidential, or bulk recordings.

Methods compared at a glance

No method wins on everything. Match the method to the recording.

MethodSpeedPrivacyLength capBest for
Phone / built-in toolsInstant, liveVaries by appShort memosQuick voice notes on the go
Free online convertersFast (small files)Uploaded to serverMinutes / monthly limitA one-off short, non-sensitive clip
Whisper (on-device)~1 hr audio / minuteNever leaves deviceNo capLong, sensitive, or bulk files

The pattern is clear: online tools are the quick option for a small, throwaway clip, but the moment a recording is long or sensitive, an on-device app is faster, cheaper over time, and far more private.

Step-by-step: transcribe an audio file with Whisper

This is the method for real recordings — interviews, lectures, meetings, podcasts. Everything runs on your own machine.

  1. 1

    Get Whisper on your computer

    Whisper is a desktop app for Mac, Windows, and Linux. Install it like any other program — no cloud account to create, no admin console. It activates with a one-time license key.

  2. 2

    Drop in your audio file

    Drag a recording straight into the app. Whisper reads WAV, MP3, FLAC, M4A, AAC, and OGG directly, so you don't need to convert the file first.

  3. 3

    Let it transcribe on-device

    The speech model runs locally on your CPU/GPU. Long files are processed in chunks with a progress bar, and GPU acceleration means about an hour of audio finishes in roughly a minute on Apple Silicon — with automatic punctuation.

  4. 4

    Copy, edit, and save the text

    When it's done you have clean, punctuated text ready to copy into your notes, document, or editor. The audio and the transcript both stay on your machine.

Bonus: dictate live, not just from files

Whisper also does live dictation — press a shortcut, speak, and your words are typed straight into any app: your email, editor, chat, or notes. Same on-device engine, so live speech stays private too. See the voice-to-text app overview.

Why on-device matters for sensitive audio

The convenience of an online converter comes with a real cost: your recording is uploaded to someone else's server to be processed. For a patient interview, a legal deposition, a confidential meeting, or an unpublished manuscript, that's often a dealbreaker. Whisper removes the upload entirely.

Your audio never leaves your device

The speech model runs locally. The recording and the transcript both stay on the machine that made them — no cloud copy to breach or subpoena.

No subprocessors, works offline

There's no third-party AI in the transcription path and no server pipeline. After setup, Whisper runs with no internet at all.

Fast and unlimited

GPU-accelerated — about an hour of audio in a minute on Apple Silicon — with no per-minute fees and no length cap on your files.

Whisper carries no certification badge — instead it gives you an architecture where your audio simply never leaves your control. Read more on the security & privacy page.

Supported formats & limits

Audio formats you can drop in

WAVMP3FLACM4AAACOGG

Drag the file straight in — no need to convert your recording first.

Length & performance

  • No length cap — long files transcribe in chunks with a progress bar.
  • GPU-accelerated: about an hour of audio in a minute on Apple Silicon.
  • Automatic punctuation, so the output reads cleanly.
  • Runs on Mac, Windows, and Linux.

How to transcribe audio to text — FAQ

How do I transcribe audio to text?

Pick a method based on the file. For a quick voice memo, your phone's built-in transcript or Voice Memos may be enough. For a one-off short clip you don't mind uploading, a free online converter works. For long recordings, sensitive material, or bulk files, use a desktop app like Whisper that transcribes on your own device: drop in the audio file (WAV, MP3, FLAC, M4A, AAC, or OGG) and it writes the text out locally — no upload, no length cap.

What is the best way to transcribe a recording?

There is no single best method — it depends on privacy, length, and volume. Online tools are fastest for a tiny, non-sensitive clip. But they upload your audio to a server and usually cap the length. For anything long, confidential, or repeated, an on-device app is better: it's faster on a real recording (about an hour of audio in a minute on Apple Silicon), has no size limit, and the audio never leaves your machine.

Free vs paid transcription — what's the difference?

Free online converters cost nothing up front but trade privacy (your audio is uploaded), impose length or minute limits, and often add watermarks or paywall the export. Paid on-device software like Whisper is a one-time purchase ($29 for the first seat) with no per-minute fees, no uploads, and no caps — you own it and use it as much as you want, offline.

Is transcribing audio to text private?

It depends entirely on the tool. Web-based converters send your recording to their servers to process it, so the file leaves your control. Whisper runs the speech model directly on your device — your audio never leaves the machine, there are no subprocessors, and it works fully offline. For interviews, medical or legal recordings, or anything confidential, on-device is the safe default.

How long can the audio file be?

Online converters usually limit you to a few minutes or a set number of monthly minutes. Whisper has no length cap: it processes long files in chunks with a progress bar, so a full lecture, interview, or podcast episode transcribes in one pass on your own hardware.

What audio formats can I transcribe?

Whisper accepts the common formats directly: WAV, MP3, FLAC, M4A, AAC, and OGG. Drag the file in and it handles the rest — no need to convert the recording first.

Transcribe any recording — privately, on your own device

Drop in a file, get clean text back. No upload, no length cap, works offline. One-time license — $29 for the first seat.

Get Whisper