Skip to content
Interview transcriptionVerbatim transcriptionQualitative researchAI transcription

How to Transcribe an Interview in 9 Steps

From consent and recording to verbatim style, AI vs human transcription, speaker labels, timestamps and anonymising the finished interview transcript.

Sarath Chandran

Updated 7 min read
How to Transcribe an Interview in 9 Steps

To transcribe an interview, get the interviewee's consent to record, record clean audio, and decide up front whether you need clean verbatim (what was said) or full verbatim (every "um" and false start). Then run the recording through an AI transcription tool or send it to a human service, check the draft against the audio, label the speakers, add timestamps where you'll quote, and anonymise the transcript before anyone else sees it. AI is the fast, cheap route when you'll review the text yourself. Human services cost more but come checked: Rev charges $1.99 a minute and says its human transcripts are 99%+ accurate and delivered in 12 hours or less (Rev pricing).

Prices and guidance below were checked on September 30, 2026.

This guide covers getting from recording to finished transcript. If you already have transcripts and want to code and analyse them, see AI-transcribed interviews for qualitative analysis.

Tell the person you're recording, what the recording is for, and who will hear it or read the transcript. Then ask again once the recording is running, so their answer is on the tape.

The law depends on where you and the interviewee are. In the US, federal law requires the consent of at least one party to record a conversation, but about 11 states, including California, Florida, Illinois and Pennsylvania, primarily require the consent of everyone in it. When people are in different states, the Reporters Committee for Freedom of the Press advises assuming the stricter law applies (RCFP). Outside the US, check local law.

For research interviews, your consent form sets the limits. CESSDA's data management guide lists "whether consent is in place to allow the fullest use of recordings" among the things to settle before you record (CESSDA). If a transcription service will hear the audio, say so in the consent form.

Step 2: Record audio that can be transcribed

Audio quality limits every transcript, whether a person or software does the typing. CESSDA notes that good sound quality prevents mis-transcription and reduces the chance of sections staying untranscribed (CESSDA).

  • Pick a quiet room. Hard, empty rooms echo, and cafés add noise a transcriber has to talk over.
  • Get the mic close to both voices. A phone on the table between you works better than a laptop across the room. Use one mic per person if you have them.
  • Record a 10-second test and listen back before you start.
  • Don't talk over the answer. Overlapping speech is the hardest part to transcribe. A nod works as well as "yeah, yeah".
  • Keep a spelling list. Note names, places, product names and jargon during the interview. You'll need them in Step 6.
  • Remote interview? Record in the meeting tool. We cover the steps in how to transcribe a Zoom meeting and how to transcribe a Teams meeting.

Step 3: Choose a transcription style

Decide this before anyone starts typing, because it changes what counts as a mistake.

StyleWhat it keepsUse it for
Clean verbatimThe words, without fillers ("um", "like"), stutters, false starts, unintentional repetition or listener interjections ("uh huh")Articles, reports, most business and user research
Full verbatimEvery word as spoken, including fillers, false starts, stutters and interjectionsLegal records, printed Q&As, anything where how it was said matters
Naturalised (with notation)Everything above plus pauses, laughter, overlaps and intonation, marked with symbolsConversation analysis and other linguistic research

The clean and full verbatim definitions follow Rev's own description of the two styles, and Rev recommends clean verbatim unless you need a word-for-word record (Rev). The third row comes from CESSDA, which describes three research approaches: content-focused, a "naturalised" style that captures how things were said, and one that also notes emotional and physical language (CESSDA).

Write the rules down, even if you're the only transcriber: how you label speakers, how you mark inaudible words, whether you keep laughter, how you format timestamps. CESSDA recommends that everyone transcribing agrees on the rules first and that you write transcriber instructions.

Step 4: Decide between AI and human transcription

OptionWhat it costsWhat you getBest for
AI transcription toolFree plans with limits, then a subscription; check each tool's pricing pageA draft in the tool's editor, often with speaker labels and timestampsMost interviews, when you'll review the draft yourself
Rev human transcription$1.99 per minuteRev states 99%+ accuracy, delivery in 12 hours or less, rush and timestamp options, and a verbatim add-onTranscripts you'll publish or rely on without re-checking every line
Happy Scribe human proofreadingFrom $2.00 per minuteA person proofreads the transcriptA checked transcript, when you already use Happy Scribe
Self-hosted WhisperNo licence fee; you need the hardware and some codeA raw transcript that never leaves your machineSensitive audio you can't upload
Typing it yourselfYour timeFull controlShort interviews, or when you need to know the material closely

Two things push the choice. The first is how much checking you can do: an AI draft still needs a full listen-through (Step 6). The second is where the audio is allowed to go. CESSDA advises that if someone else transcribes, you sign a non-disclosure agreement with them and encrypt the files before you transfer them (CESSDA). The same logic applies to a software vendor, so read its data and privacy terms before uploading anything personal.

For a side-by-side of AI tools, see the best transcription software in 2026.

Step 5: Run the transcription

The steps are much the same in any AI tool. Here they are in Exemplary AI, which is our product:

  1. Upload the recording. Exemplary AI transcription accepts MP3, WAV, M4A, MP4, WEBM, MOV and MKV files.
  2. Pick the spoken language. Transcription covers 99 languages, which matters if your interview isn't in English.
  3. Name the speakers. The transcript comes back with speaker labels. Rename them to "Interviewer" and a participant code.
  4. Edit in the transcript editor while you listen (Step 6).
  5. Export. Download a Word (.docx) or text (.txt) file, with speaker names and timestamps switched on or off, or an SRT or VTT file if the interview is a video you want to subtitle.

Exemplary AI is AI-only. If you need a person to check the transcript, use a human service like Rev or Happy Scribe.

Step 6: Check the draft against the audio

This is the step that makes a transcript usable, and it can't be skipped with AI or human work. Play the audio (1.25x speed is often fine) and read along:

  • Names, places and jargon. Use your spelling list from Step 2. Unfamiliar names are easy for software and people alike to mishear.
  • Numbers, dates and amounts. "Fifteen" and "fifty" sound alike.
  • Who said what. Rev's style guide tells its transcribers to "always attribute what is being said to the correct speaker" (Rev Support). Short interjections and crosstalk are where labels slip.
  • What was said, not what was meant. The same guide says: "Do not type what you think the speaker meant to say." Don't tidy up a quote beyond the style you chose in Step 3.
  • Gaps. Mark words you can't make out with a timestamp, like [inaudible 00:14:32], so you or someone else can find them again.

Step 7: Format speakers and timestamps

Pick one layout and use it in every transcript of the project. A simple, readable one:

1Interview: P04 | 12 February 2026 | Clean verbatim
2
3[00:00:05] Interviewer: Thanks for doing this. Can you start with how you got into the job?
4
5[00:00:11] P04: Sure. I started on the night shift at [Hospital A] about six years ago...
  • Put a timestamp at every change of speaker, or at least every few minutes, so any quote can be traced back to the audio.
  • Use participant codes (P01, P02) rather than first names if the transcript will be shared.
  • Keep formatting plain. CESSDA warns that headers, bold and italics can be lost when transcripts are imported into qualitative analysis software, and that two-column speaker/utterance tables can cause problems (CESSDA).

Step 8: Anonymise before you share

If the transcript leaves your hands, for a co-author, a client, an archive or a quote in print, remove what identifies people. CESSDA's best practices for qualitative data (CESSDA):

  1. Replace, don't blank out. Use pseudonyms or generic descriptors ("[a hospital in the north]") rather than deleting the information.
  2. Plan it at transcription. Anonymise as you transcribe, or mark sensitive passages for later.
  3. Stay consistent. Use the same pseudonym for the same person across the whole project, including publications.
  4. Be careful with search and replace. It can change words you didn't intend and miss misspellings.
  5. Mark every change with [brackets], so readers know the text was edited.
  6. Keep an anonymisation log of every replacement, stored securely and separately from the anonymised files.

Look beyond names. CESSDA separates direct identifiers (name, address, phone number) from indirect ones that can identify someone in combination, such as occupation, age and location.

Step 9: Export and store the transcript

  • For editing and sharing: Word (.docx).
  • For analysis software: plain text (.txt) or Word, depending on what your software imports.
  • For long-term archiving: the UK Data Archive lists XML, RTF and plain text as its preferred formats for qualitative text, with HTML and Word among the acceptable ones (UK Data Archive).

Name files consistently (project, participant code, date, version), and keep the raw audio and the un-anonymised transcript in a separate, access-controlled place from the version you share.

FAQ

Should I use clean verbatim or full verbatim?

Clean verbatim for most interviews: it keeps every meaningful word and drops fillers, stutters and false starts. Use full verbatim when how something was said matters, such as legal records, printed Q&As or linguistic research.

With consent, generally yes, but the rules vary. US federal law requires at least one party's consent, while about 11 states primarily require everyone's consent (RCFP). The simplest rule is to ask everyone, every time, on the recording.

How long does it take to transcribe an interview?

It depends on the method and the audio. Rev's human service promises delivery in 12 hours or less. With an AI tool, the draft is the quick part and most of your time goes into the review pass in Step 6, which takes longer with noise, accents or crosstalk.

Can AI transcribe interviews in other languages?

Yes, but coverage varies a lot by tool. Exemplary AI transcribes 99 languages; other tools support far fewer, so check the vendor's language list for your language and dialect before you commit.

How do I anonymise an interview transcript?

Replace identifying details with consistent pseudonyms or generic descriptors in [brackets], think about indirect identifiers like job and location as well as names, and keep a separate, secure log of every change (CESSDA).

One upload. Every format you need.

Drop in a recording and work from the transcript — clips, captions, chapters, show notes and posts all come from the same file.