All posts
Last updated 12 min read

Transcribing Interviews: A Complete Guide for Journalists and Researchers

Anyone who conducts interviews for journalistic research, academic work or market research faces the same task afterwards: the conversation has to be transcribed. Manual transcription takes three to six hours per hour of audio - with AI support it takes minutes.

This guide walks you through the entire process: from proper preparation, through the choice of transcription method, to the post-processing of the finished text.

Preparation: the recording determines the quality

The most important rule: the better the recording, the better the transcript. Invest in preparation, not in post-processing.

  • Microphone: An external microphone (lavalier or tabletop) delivers significantly better results than the built-in laptop microphone. Position it close to the speakers.
  • Environment: Choose a quiet room. Street noise, air conditioning or background music significantly reduce transcription accuracy.
  • Format: Record in high quality (WAV or M4A). Compressed formats such as low-bitrate MP3 degrade speech recognition.
  • Backup: Always record in parallel with a second device. A technical failure cannot be undone.

Transcription methods compared

There are three methods: verbatim for academic work, smoothed for journalism and summarizing for internal notes. The verbatim method captures every word including filler words; the smoothed method removes these and corrects the grammar.

There are three basic methods for transcribing an interview. Which one you choose depends on the intended use:

  • Verbatim: Every word is captured exactly, including filler words (uh, mhm), pauses and slips of the tongue. The standard for academic work and qualitative research.
  • Smoothed: Filler words and sentence breaks are removed, and grammar is corrected. The content is preserved, but the text reads more fluently. The standard for journalism.
  • Summarizing: Only the key statements are captured, the rest is condensed. Suitable for internal notes and quick evaluations.

AI transcription: automatic and fast

AI transcription converts an hour of audio into text in under five minutes - with over 95 % accuracy for English. The most efficient approach is a hybrid: the AI creates the rough draft, and the human corrects proper names and technical terms. This saves 80 to 90 percent of the manual transcription time.

Modern AI models transcribe an hour of audio in under five minutes with over 95 % accuracy for English. The typical workflow:

  • Upload the audio file (MP3, WAV, M4A, MP4 and many other formats)
  • The AI automatically creates a timestamped transcript with speaker diarization
  • Review the transcript in the browser and correct it if needed
  • Export as PDF, Word, TXT or a subtitle format (SRT, WebVTT)

Hybrid approach: The most efficient method is to take the AI transcript as a basis and post-process it manually. This saves 80 to 90 percent of the time compared with purely manual transcription.

Post-processing: from raw text to finished transcript

Even the best AI transcript needs a review. Pay attention to these points:

  • Assign speaker names: The AI automatically recognizes speakers as “Speaker 1,” “Speaker 2.” Rename these with the actual names.
  • Check technical terms: Industry-specific terms, proper names and abbreviations are sometimes misrecognized by the AI. Use the audio playback function to listen to unclear passages directly.
  • Adjust the formatting: Paragraphs, headings and emphasis make the transcript easier to read.

Transcribing interviews for your thesis: step by step

Bachelor’s and master’s theses have specific requirements: the transcript forms part of the research documentation and must comply with your department’s guidelines. Here is how to proceed:

  1. Clarify the requirements: Ask your supervisor which transcription rules apply. Qualitative research usually requires verbatim transcription (see above), with consistent speaker labels and timestamps. Also clarify whether the transcripts should be included in the appendix or submitted separately.
  2. Obtain consent: Before recording, and in writing - you will find sample wording in the next section.
  3. Record: Follow the guidance in the preparation chapter: use an external microphone, choose a quiet room and have a backup device ready. For remote interviews, you can also record directly in the browser - on desktop or mobile, without an external tool.
  4. Create a raw transcript with AI: Upload the audio file and the AI will deliver a timestamped transcript within minutes. This replaces the raw version that would otherwise have taken three to six hours of manual typing per hour of interview.
  5. Review it for academic use: Check the transcript against the audio in the editor - clicking a passage takes you to the corresponding point in the recording. Label the speakers according to your system, for example “I” for the interviewer and “R1” for the respondent. Your corrections are encrypted directly in your browser before they are saved.
  6. Document your methodology: Explain in your thesis how the transcription was produced: AI-assisted raw transcript, manual review and the transcription rules applied. Many universities require you to name the software used - ask your department if you are unsure.
  7. Export and analyze: Export the reviewed transcript in the appropriate format for your analysis - more on this below.

Plan realistically: the AI delivers the raw transcript within minutes, but the academic review still requires manual work. The hybrid approach saves 80 to 90 percent of the time - make sure you allow for the remaining work on each interview in your schedule.

Data protection for interview transcriptions

Interviews contain personal data - voices, names, opinions. In the EU, the processing falls under the GDPR. Inform your interview partners before the recording and obtain explicit consent.

Especially for sensitive topics (whistleblower interviews, patient interviews, confidential sources), you should use a transcription service with client-side encryption, where the provider has no access to the plain text.

Interviewee consent: sample wording

For a thesis, consent should be obtained in writing - it forms part of the research documentation, and many ethics committees explicitly require it. Here is a template you can adapt:

“I consent to the interview with [name] on [date] being recorded as an audio file for the purposes of the [bachelor’s/master’s thesis, working title, university]. The recording will be converted into text using an AI-assisted transcription service [name the provider]. The transcript will be pseudonymized: my name and any details that could identify me will be replaced or removed. The audio file will be deleted once transcription is complete, and no later than [date]. Anonymized excerpts from the transcript may be quoted in the thesis. Participation is voluntary; I may withdraw this consent at any time, without giving reasons, with effect for the future.”

Important: This template is not legally binding and does not replace legal advice. Adapt it to the requirements of your university or ethics committee - many departments provide their own templates. The information to include about the transcription service depends on the provider: with scryp, you can state that the recording is encrypted in your browser before it reaches our servers, that processing takes place in Austria and that your data is not used to train AI models.

How much does it cost to transcribe interviews for a thesis?

Here is a sample calculation for a typical qualitative master’s thesis: 10 interviews of 45 minutes each add up to 7.5 hours of audio. Typing them out manually takes between 23 and 45 working hours at three to six hours of work per hour of audio - weeks of effort alongside your studies or job.

With a monthly subscription, the calculation looks different: the Nano plan costs €13.20 per month (billed monthly, incl. VAT, cancel anytime) and includes unlimited transcription. Transcribing all ten interviews within the same month costs €13.20 in total. If your data collection takes two months, the total is €26.40 - after that, you can simply cancel. Automatic speaker diarization is usually useful for interviews; it is included from the Pro plan (€19.87 per month), while the Nano plan delivers a continuous transcript without speaker separation.

All plans come with a free 14-day trial. The prices quoted apply on scryp.at - the pricing page shows the current prices and currencies for your region.

Working with your transcript: from export to analysis

For your analysis, export the transcript in one of five formats - with speaker labels and timestamps if required:

  • Word (DOCX): The standard format for academic analysis. You can add comments, highlight passages, import the document into QDA software or format it as an appendix to your thesis.
  • Text file (TXT): Plain text with timestamps and speaker labels - the most compatible format. Common qualitative data analysis programs can import TXT files directly.
  • PDF: A formatted document with speaker labels and timestamps - suitable for submission and archiving because it cannot be changed accidentally.
  • SRT and WebVTT: Subtitle formats - useful if you want to subtitle video interviews or show excerpts in a presentation.

You do not need to email files back and forth when coordinating with your supervisor: transcripts can be shared via an encrypted, password-protected link.

Checklist: transcribing an interview

  • Use a good microphone, choose a quiet environment
  • Obtain consent for the recording
  • Backup recording on a second device
  • Decide on a transcription method (verbatim, smoothed, summarizing)
  • Use AI transcription as a basis, post-process manually
  • Assign speaker names, check technical terms
  • Export in the desired format

Conclusion

Interview transcription does not have to be hours of tedious work. With the right preparation and an AI-assisted tool, the effort is reduced to a fraction of the manual time. The key lies in the combination of good recording quality, automatic transcription and careful post-processing.

Frequently asked questions

How long does it take to transcribe a one-hour interview?

Manually, it takes three to six hours per hour of audio - depending on your typing speed, the audio quality and the level of detail required. AI transcription converts one hour of audio into text in under five minutes. You then need to review the result by assigning speaker names and checking technical terms and proper names. The hybrid approach - an AI-generated raw transcript followed by manual correction - saves 80 to 90 percent of the time compared with fully manual transcription.

How accurate is automatic interview transcription?

Modern AI models achieve more than 95% accuracy for English. The result depends on the audio quality, the number of speakers and the dialect - a good microphone and a quiet environment make the biggest difference. You can correct the remaining errors, typically in proper names and technical terms, in the editor: clicking a passage takes you to the corresponding point in the recording, so you can replay unclear sections immediately.

Does AI transcription also work with Austrian dialects?

Yes. scryp's AI model is particularly strong in the German-speaking region and supports German, including Austrian speech patterns; it reliably recognizes dialects, proper names and natural colloquial language. As with any speech recognition system, results can vary depending on the dialect, audio quality and number of speakers. For heavily dialectal passages, ensure good recording quality and check unclear sections in the editor using synchronized audio playback.

Should you have an interview typed up or transcribe it automatically?

Typing it up yourself takes three to six hours per hour of interview. If you give the recording to a third party for transcription, you need the interviewee's consent - sharing a recorded conversation is relevant under data protection law. AI transcription delivers the raw transcript within minutes while you retain control: with a service that uses client-side encryption, the recording is encrypted in the browser, so the provider has no access to the plain text. You complete the review yourself in the editor.