Guide
Your transcript is inaccurate — here is what actually fixes it
People reach for a bigger model first. It is the least effective change available, and the microphone you already own is the most effective one.
When a transcript comes back full of errors, the instinct is to blame the tool and go looking for a better one. Sometimes that is right. Usually the problem is upstream, and the fix is free.
Here is what actually moves accuracy, in order of effect.
1. Microphone distance (by far the biggest) #
Speech recognition degrades with the ratio of voice to everything else. Distance destroys that ratio faster than anything: sound pressure falls off with distance while the room's noise stays constant, so every step back makes the speech quieter relative to the air conditioning, the traffic, and the coffee machine.
A phone eighteen inches from a speaker's mouth produces a transcript that is dramatically better than the same phone six feet away, using the same app and the same model. This is not a small effect — it is the difference between a transcript you fix in five minutes and one you retype.
If you can change exactly one thing, change this.
2. Where the recorder sits #
Related but distinct, and it is where meeting recordings go wrong.
A device in the middle of a table is equidistant from everyone, which means it is far from everyone and close to the table — which conducts every tap, shuffle and coffee cup straight into the microphone. Better: raise the recorder off the surface, and point it at whoever talks most.
For a lecture or presentation, if the room has a PA system or an official recording, that feed is better than anything you can capture from a seat.
3. Set the language explicitly #
Automatic language detection samples the beginning of the file. If the first thirty seconds are room noise, music, or someone speaking a different language from the rest of the recording, detection can settle on the wrong language and stay wrong for the entire file.
The symptom is unmistakable and confusing: a transcript in a language nobody spoke, or bursts of it. If you know the language, set it rather than letting the tool guess. In Offline Transcription you can hold a file's Transcribe button and choose its language, which matters most when you have several recordings in different languages.
4. Do not record over-compressed #
Below roughly 64 kbps, lossy compression removes the high-frequency detail that distinguishes similar consonants — the difference between "fifteen" and "fifty" lives up there. No transcription tool can put back information the encoder threw away.
Record uncompressed or at a normal bitrate. If you are handed a low-bitrate file, that is your accuracy ceiling and no setting changes it.
5. Cut the silence #
Long stretches with no speech are not neutral. Whisper generates text for whatever it is handed, and given silence it repeats itself, sometimes for pages.
Voice activity detection solves this properly, by never showing the model the silent parts. Failing that, trim the recording before transcribing.
6. Then, maybe, a bigger model #
This is last on the list for a reason.
Model size does improve accuracy, and the improvement is real but smaller than people expect — and it costs a lot of time and disk. More importantly, a bigger model does not fix any of the problems above. It will transcribe a distant, noisy, over-compressed recording slightly better, and still badly.
The honest framing: model size is a multiplier on your audio quality. Multiplying poor audio gets you slightly-less-poor results. Our guide to choosing a model covers the trade-off.
What does not help at all #
Re-running the same file. Sampling is not fully deterministic, so you may get slightly different errors. That is variance, not improvement.
Boosting the volume. Amplifying a quiet recording amplifies the room noise identically. The ratio — the thing that matters — is unchanged.
Noise reduction, usually. Aggressive denoising introduces artefacts that confuse the model as much as the noise did. Gentle high-pass filtering to remove rumble can help slightly. Heavy processing usually hurts.
Paying more per minute. Cloud services run the same family of models. Price is not accuracy.
What good looks like #
Even with everything right, automatic transcription is not perfect and no tool will make it so. Expect these to remain wrong at some rate: unfamiliar proper nouns, technical jargon, numbers spoken quickly, and moments when two people talk over each other.
That is why the workflow matters as much as the accuracy: use a segment view with timestamps so you can play the audio behind any line you doubt, and check the parts you intend to rely on. Our interview guide covers a fast way to do that.
The short version #
Move the microphone closer. Get the recorder off the table and pointed at the speaker. Set the language explicitly instead of trusting detection. Do not record at low bitrates. Trim or detect silence. Only then think about model size, which is the smallest lever of the six.
Offline Transcription
Transcribe audio and video entirely on your own Mac. Nothing is uploaded, there is no account, and there is no per-minute meter.