Use case

Interview transcription

Turn interview recordings into searchable, timestamped transcripts on your own computer. Built for user research, recruiting, journalism and academic fieldwork.

  • Keep participant and candidate recordings on your device (Local mode)
  • Timestamped transcripts you can quote and cite
  • Markdown export for Obsidian, Notion or your research repository
Interview transcript in the Noteo desktop app

How it works

1) Import the recording
Open an interview file (MP3, M4A, WAV, MP4, MOV and other common formats) in the Noteo desktop app.
2) Transcribe locally
In Local mode, Noteo extracts the audio with ffmpeg and runs whisper.cpp on your machine.
3) Review and export
Skim the timestamped text, then export Markdown, TXT, SRT or VTT to annotate in your own tools.
4) Optional cloud features
If your consent allows it, switch to Cloud mode for speaker labels, AI summaries or chat.

Why interviews are hard to transcribe responsibly

Interviews are usually the most sensitive audio a team records. A user research session may include a participant’s workplace details, a candidate interview includes salary expectations, and a journalist’s source may be relying on you to protect their identity. Consent forms often promise that recordings will only be handled by the research or hiring team.

Uploading those files to a cloud transcription service can quietly break that promise, or at least add a new data processor you have to disclose. Transcribing by hand keeps the data in-house but costs hours per interview. Local transcription removes that trade-off: the speech model runs on your computer, and the transcript never leaves it unless you export it.

Who uses local interview transcription

The workflow fits anyone who records one-to-one or small-group conversations and has to account for where the audio goes:

  • UX and product researchers running customer interviews and usability tests
  • Recruiters and hiring managers reviewing candidate conversations
  • Journalists and podcasters working with sources and off-the-record material
  • PhD students and academics whose ethics approval restricts data storage
  • Market researchers and consultants synthesising stakeholder interviews

Worked example: a round of user research interviews

Imagine a product designer who runs eight 30-minute customer interviews over Google Meet and records each one. After each call, she imports the recording into Noteo in Local mode. ffmpeg converts it to the 16 kHz mono audio that whisper.cpp needs, and the transcript appears in her local library with timestamps next to every segment.

She then exports each interview as Markdown. The file includes a small front-matter block with the title and date, followed by timestamped lines, so it drops straight into an Obsidian vault or a research repository. There she tags quotes, clusters pain points across the eight sessions and links each insight back to the exact timestamp. The raw recordings and transcripts stayed on her laptop the whole time.

Local vs cloud for interviews

Both modes are available in Noteo. Pick per project, based on what your consent form and data policy allow:

  • Local mode: audio and transcripts stay on your device, which makes it easier to stay within the terms participants agreed to.
  • Local mode has no speaker diarization, so interviewer and interviewee are not separated automatically.
  • Cloud mode adds speaker recognition, AI summaries and chat over the transcript, but the recording is uploaded for processing.
  • Local speed depends on your hardware and the whisper model you choose. Cloud speed does not depend on your laptop.

Compared with the usual alternatives

Manual transcription is private and precise but slow. Human transcription services are accurate but send your files to outside transcribers. Cloud-only AI note takers are fast and label speakers, but uploading is built in. Local transcription with Noteo gives you a fast first draft without a new data processor. You still review it, but you review instead of typing from scratch.

Honest limitations

Local mode requires the Noteo desktop app plus ffmpeg, whisper.cpp and a model file that you install and select. Accuracy depends on the model and the recording: remote calls with poor microphones, overlapping speech and strong accents all produce more errors, so always check quotes against the audio before publishing or citing them. Because there is no local speaker separation, you will need to mark who is speaking yourself if your analysis depends on it.

Data flow (Local vs Cloud)

Stays on your device
  • Interview audio/video
  • Transcript saved in the local library
  • Markdown, TXT, SRT and VTT exports
May use cloud (optional)
  • Optional: speaker labels, AI summary and chat in Cloud mode

FAQ

Can I transcribe interviews without uploading them?

Yes. In Local mode, the Noteo desktop app runs whisper.cpp on your computer, and ffmpeg handles the audio conversion. The recording and the transcript are stored locally and are not uploaded. You only share the text when you export a file and send it somewhere yourself.

Will the transcript separate interviewer and interviewee?

Not in Local mode. Local transcripts currently have no speaker diarization, so all text appears under one speaker with timestamps. Cloud mode includes speaker recognition, but it requires uploading the recording, so check that your consent terms allow it first.

What file formats can I use for interview recordings?

Local mode uses ffmpeg to extract audio, so common formats such as MP3, M4A, WAV, FLAC, OGG, MP4 and MOV work, including video recordings of remote calls. WAV files are passed to whisper.cpp directly, and other formats are converted to 16 kHz mono WAV first.

How do I get interview transcripts into my research notes?

Export the transcript as Markdown. The file starts with a front-matter block (title, creation date, mode) followed by timestamped lines, which works well in Obsidian, Notion or any Markdown-based research repository. TXT, SRT and VTT exports are also available if your analysis tool expects them.

Is local transcription accurate enough to quote from?

It is a strong first draft, not a final record. Accuracy depends on the whisper model you choose, the audio quality and how much people talk over each other. Use the timestamps to check every quote against the original recording before you publish, cite or share it.