Interviews, webinars, phone calls, voice memos, lectures — five kinds of recording, one transcription pipeline. Upload the audio, get text back with speaker labels and timestamps, clean it up, export it.

What changes is everything around the pipeline: where the microphone sits, how many people are talking, what tends to go wrong, and what you actually do with the transcript. A two-person interview across a table and a 200-attendee webinar with a chat-based Q&A fail in completely different ways. So this guide states the shared workflow once, then works through what is specific to each of the five.

## The Workflow, Stated Once

Every recording type runs through the same four steps.

1. **Record.** A phone, a digital recorder, a conferencing platform, a VoIP service — anything that produces an audio or video file.
2. **Upload.** MP3, M4A, WAV, MP4, and other common formats are accepted. A two-hour recording comes back in minutes, not hours.
3. **Review.** You get a transcript with speaker labels and word-level timestamps. Rename "Speaker 1" and "Speaker 2" to real names, then check proper nouns and technical terms — that is where AI transcription slips most often.
4. **Export.** TXT, SRT, VTT, DOCX, or PDF, depending on what happens next.

Blazescribe transcribes audio in 50 languages and can translate a finished transcript into 106 — useful when you interview a source in one language and publish in another.

### Speaker labels are what make a transcript usable

[Speaker diarization](/blog/what-is-speaker-diarization) labels each segment by voice, so you can tell the interviewer from the interviewee, or a presenter from an audience member in a Q&A. It cannot know anyone's name — you rename the generic labels once, and the whole transcript updates.

### The audio rules that apply to all five

- **Record in the quietest space you can get.** Background noise costs you accuracy on every recording type.
- **Do a 10-second test recording** and listen back before the real thing starts.
- **Avoid crosstalk.** Overlapping speech is the hardest thing for any transcription system to untangle. Let each person finish.
- **Back up anything you cannot repeat.** Two devices. An interview, a call, or a lecture happens once.

Bad audio is best fixed at the source — see [how to improve audio quality for transcription](/blog/how-to-improve-audio-quality-for-transcription).

### Choosing an export format

- **TXT** — plain text for qualitative analysis software (NVivo, ATLAS.ti) or for pasting anywhere else.
- **DOCX** — when the transcript still needs editing and sharing in a word processor.
- **PDF** — archival, read-only copies, or a transcript you hand out.
- **SRT / VTT** — captions, if the recording was video. Captions also serve people who prefer to read and people with hearing impairments.

That is the entire shared workflow. Everything below is what actually differs.

## How to Transcribe an Interview

**Mic setup.** One external mic placed between you and your subject, equal distance from both. Phone and laptop mics produce noticeably lower-quality audio, and a phone lying flat on a table favors whoever sits closest to it.

**Speaker count.** Two, usually — the easiest case for diarization, and the one where labels matter most. A quote is worthless if you cannot show who said it.

**Decide on a transcription style before you start.** Three conventions, and the right one depends on the job:

- **Verbatim** captures every word exactly as spoken, including filler words ("um", "uh"), false starts, and pauses. Used for legal proceedings, qualitative research, and depositions.
- **Clean verbatim** removes filler words and false starts while preserving the exact meaning. The standard for journalism, content creation, and business meetings.
- **Intelligent verbatim** restructures sentences for readability while keeping the meaning intact. Used for blog posts, articles, and public-facing content.

**Common failure modes.** No backup recording — an interview is rarely repeatable, so record on two devices. Forgetting to tell the interviewee they are being recorded. Crosstalk, which gets worse the better the conversation is. And publishing without checking proper nouns and technical terms, which in an interview are usually the words carrying the story.

**What you do with the output.**

- **Journalists** want clean verbatim for accurate quotes. Speaker labels identify the source; timestamps let you jump back to the exact moment while writing.
- **Academic researchers** usually need full verbatim. AI transcription takes weeks of manual work out of a multi-interview project, and TXT drops straight into NVivo or ATLAS.ti.
- **HR teams** get a consistent record for hiring decisions — it helps with compliance, reduces bias, and lets colleagues who were not in the room review the candidate.
- **Podcasters** turn one interview episode into a blog post, show notes, and social quotes.

## How to Transcribe a Webinar

**Get the recording out of the platform first.** Most save an MP4:

- **Zoom Webinars** — Zoom web portal > Recordings
- **GoToWebinar** — the webinar dashboard
- **Webex Events** — the Webex recordings page
- **Microsoft Teams Live Events** — Stream or SharePoint

**Mic setup.** Presenters on wired internet with a real microphone. Your transcript is only as good as the worst presenter's audio.

**Speaker count.** One or two presenters, plus however many people speak during Q&A. Ask presenters to introduce themselves by name at the top — that gives you an anchor for renaming the speaker labels afterwards.

**The failure mode nobody sees coming.** Audience questions submitted through chat never appear in the transcript, because nobody said them out loud — you end up with a string of answers to invisible questions. Read each question aloud before answering it.

**What you do with the output.** A single 60-minute webinar holds enough raw material for five to ten pieces:

- **Blog post** — restructure the presentation into an article with the key takeaways and data points.
- **Training documentation** — format it into step-by-step guides or SOPs for new hires.
- **Email sequence** — pull three to five insights, one per email, each linking back to the recording.
- **Social posts** — quotable moments, stats, and tips.
- **FAQ page** — the Q&A section, cleaned up and published.
- **Lead magnet** — the transcript as a downloadable PDF.
- **Searchable archive** — a library of past webinars your team can actually search.

A webinar with 200 attendees creates value once. A transcribed webinar keeps creating it, through search traffic, onboarding, and lead capture. Two rules: structure the webinar into clear sections, so the transcript arrives half-organized, and repurpose within a week, because timely content performs better. More on that in [how to repurpose audio content](/blog/how-to-repurpose-audio-content).

## How to Transcribe a Phone Call

**Consent comes before anything else.** Recording phone calls without proper consent is illegal in many jurisdictions. Know your local laws before you record.

- **One-party consent** — only one person on the call (you) needs to know it is being recorded. This applies in most US states, the UK (with some exceptions), and several other countries.
- **Two-party (all-party) consent** — everyone on the call must be informed and agree to the recording. This applies in California, Illinois, Florida, Germany, and many other jurisdictions.

Best practice, regardless of local law: always tell the other party. A simple statement works — *"I'd like to record this call for my notes. Is that okay with you?"* For the fuller picture — consent rules outside the US, and what to do when a recording contains health information — see [recording and transcribing legally](/blog/hipaa-compliant-transcription-guide).

**How to capture the audio.**

- **iPhone** — the Voice Memos app on speakerphone (it records through the phone's mic), a third-party call recording app that complies with Apple's guidelines, or a conference calling service with built-in recording.
- **Android** — call recording apps from the Play Store; Samsung, Xiaomi, and other manufacturers also build recording into the Phone app. Quality varies by device and app.
- **VoIP / softphone** — Zoom Phone, Google Voice, RingCentral, and other services offer built-in recording. Download the file from the service's dashboard.
- **A separate device** — put the phone on speaker and record with something else. Simple, but ambient noise costs you quality.

**Mic setup and speaker count.** Two people, one on each end. Use speakerphone or a headset, keep your own side quiet, and prefer a landline or VoIP over cellular where you can — cell quality varies, and the transcript follows it down. Speaker labels then separate the two callers.

**What you do with the output.** Sales calls: review what was promised, catch the objections, sharpen the pitch against real conversations. Client calls: requirements and deadlines become a reference document for the whole team. Support calls: quality assurance, training, documentation. And in some industries — financial services, healthcare, insurance — call recording and transcription are a compliance requirement outright.

## How to Convert Voice Memos to Text

**Where the files live.** iPhone Voice Memos saves M4A — in the app, or under Files > iCloud Drive > Voice Memos. Android's built-in recorder saves M4A, MP3, or OGG, usually under Audio or Recordings in the Files app. Digital recorders, smartwatches, and desktop apps produce MP3, WAV, WMA, or AMR. All of it transcribes.

**Getting the file out.** On iPhone: tap the recording, tap the share icon, save to Files or share it onward. On Android: open the recorder app, tap the recording, share. Or email the file to yourself — the simplest method, and it always works.

**Mic setup.** The phone *is* the mic. Hold it six to twelve inches from your mouth, speak clearly, and step away from crowds, traffic, and wind.

**Speaker count.** One: you. That makes voice memos the easiest thing to transcribe and the fastest — most are one to ten minutes and come back almost immediately.

**Common failure modes.** Rambling across four topics in one memo. And recording no context at all, so a month later you open a transcript that starts mid-thought. Two fixes: state the context in the first sentence ("Meeting with John about the Q3 budget"), and keep one topic per memo.

**What you do with the output.** This is the category people under-use. Record a meeting on your phone and transcribe it later instead of splitting your attention between listening and typing. Talk through an idea while walking or commuting, then turn the transcript into actionable notes. Capture a quick field interview for journalism, user research, or hiring. Dictate a to-do list faster than you can thumb-type one. Voice-journal and get a searchable personal archive out of it.

If you have a backlog of dozens of old memos on your phone, upload them together rather than one at a time.

## How to Convert a Lecture Into Study Notes

**Recording.** Your phone's voice recorder is fine; if the lecture is virtual, record it in Zoom or Teams. Ask permission first if you are recording in person.

**Mic setup.** Sit near the front. Distance is the whole game in a lecture hall — an external mic on the desk in row three beats a phone in row twenty.

**Speaker count.** One, mostly: the professor. The complication is student questions, which come from across the room, far from your mic, and often half-audible. Those are the segments that come back garbled — and often the ones worth having, since the professor's answer is a second run at explaining the concept.

**Common failure modes.** Far-field audio in a big room, and two-hour recordings you never revisit because scrubbing back through audio is slow. The transcript solves the second problem outright: you read at your own pace instead of rewatching.

**What you do with the output.** A raw transcript is not study notes yet. Restructure it into:

- **Topic headings** based on what the professor actually covered
- **Key concepts**, pulled out and defined
- **Examples**, preserved with their context
- **A summary** of each section

Then pick a format. The **Cornell method** — main notes, cue column, summary — maps cleanly onto a transcript. An **outline** of hierarchical bullets by topic and subtopic is the sensible default. Or use the extracted concepts and their relationships as input for your own **mind map**.

Then do the parts the AI cannot. Review within 24 hours, which dramatically improves retention (the spacing effect). Highlight what will be on the exam. Connect concepts back to the textbook chapters. The AI captures what was said; you add what it means.

**Two things a transcript unlocks that handwritten notes do not.** Share it with a study group and split it up — each person takes different sections, adds annotations, and writes practice questions, with everyone working from the same complete source material. And at exam time, search across every lecture transcript at once: find every mention of a concept and compare how the explanation changed between sessions.

One habit worth forming: record every lecture. Storage is cheap, and you never know which one turns out to matter.

## Start With Whatever You Have

Five recording types, one pipeline, and a different set of details for each. The details are what decide whether the transcript is usable or merely correct.

[Sign up for Blazescribe](/signup) and transcribe your first recording — interview, call, webinar, memo, or lecture — in minutes.
