Guide

How to transcribe an interview you need to quote accurately

Automatic transcription is now good enough to work from and still not good enough to quote blind. The difference is a checking pass that takes minutes, not hours.

Automatic transcription changed interview work in a specific way. It did not remove the need to check your quotes. It removed the need to type out ninety minutes of conversation to find them.

That is a big change, and it only works if you keep the checking step. Here is a method that assumes the transcript is a searchable index of the audio rather than a finished document.

Before you record #

Two decisions here save more time than anything you do afterwards.

Put the recorder near the person, not near you. The most common cause of an unusable transcript is a phone flat on a café table between two people, capturing the espresso machine at the same volume as your subject. A phone propped against something, pointed at them, three feet away, is dramatically better. An external mic is better again, and rarely necessary.

Say who is who, out loud, at the start. "It's the 30th, I'm speaking with Dr Chen about the trial results." Ten seconds, and it survives into the transcript as a header you can search for. If you record many interviews, this is the difference between a folder of files named New Recording 47 and an archive.

Transcribing it #

The mechanics are simple. The decision worth making deliberately is where the audio goes.

Most transcription services upload your recording to a server. For an interview, that is a real question rather than a philosophical one:

  • If you promised a source confidentiality, sending the recording to a third party is something you should be able to describe accurately if asked.
  • If your subject is discussing anything covered by an NDA, medical privacy, or an ethics approval, "we uploaded it to a US service that keeps it for 30 days" may breach terms you agreed to.
  • If you are in the field with no signal, an upload-based tool does not work at all.

Local transcription sidesteps all of it: the file is read from disk, transcribed on your machine, and never sent anywhere. If you would like to verify that claim rather than take anyone's word for it — including ours — we wrote three ways to test it yourself.

With Offline Transcription, drop the recording in, set the language if it is not English, and transcribe. A ninety-minute interview takes a few minutes on Apple Silicon.

The checking pass #

This is the part that matters, and the part people skip.

You are not proofreading the transcript. You are verifying the specific sentences you intend to publish, plus the places errors are known to cluster.

Check every quote you plan to use, against the audio. Not the paragraph around it — the words inside the quotation marks. This is non-negotiable if the quote is going in print with someone's name on it.

Check these four things everywhere:

  1. Names and organisations. Whisper will confidently render an unfamiliar surname as a common word that sounds like it. It does this without any signal that it is unsure.
  2. Numbers. "Fifteen" and "fifty" are one phoneme apart and a factual error apart. So are "million" and "billion" in fast speech.
  3. Negations. "We did not consider that" and "we did consider that" differ by one short word that is easy to swallow in speech. This is the error that gets people sued.
  4. Crosstalk. Where two people speak at once, transcripts blend them, and a sentence can end up attributed to the wrong person entirely.

The mechanical trick that makes this fast: work in a segment view where every line carries its own timestamp, and play the audio for just that line rather than scrubbing the waveform. Checking twelve quotes then takes twelve clicks instead of twelve hunts through ninety minutes. In our app, the Segments view does exactly this — select a line, press play, hear only that line.

Speaker labels #

Nothing built on Whisper does reliable speaker identification. Tools that offer it are running a separate diarisation model, and the results are decent for two people in a quiet room and poor for four people in a lively one.

For a standard two-person interview, the practical approach is to fix it yourself at the point of use: you know which voice is yours, and by the time you are pulling quotes you are listening to those moments anyway. Do not build a workflow that depends on automatic labels being right.

Keeping the transcript useful later #

Two habits, both cheap:

Export with timestamps, not just text. Plain text is what you write from; a timestamped export is what lets you find the moment again in six months when a fact-check comes back. CSV is good for this if you want to sort or filter in a spreadsheet.

Keep the audio. The transcript is derived work. If a quote is ever disputed, the recording is the evidence and the transcript is not.

The short version #

Record close to the subject and slate the interview out loud. Transcribe locally, so the recording stays yours and works without signal. Then verify every quote you intend to publish against the audio, plus names, numbers, negations, and crosstalk everywhere. Use a segment view with per-line playback so that checking takes minutes.

The transcript is a fast index of what was said. Your quotes still come from the recording.

Offline Transcription

Transcribe audio and video entirely on your own Mac. Nothing is uploaded, there is no account, and there is no per-minute meter.

macOS 13.0 or higher · iOS 16.4 or higher