Use case

Podcast transcription for long recordings

Turn full podcast episodes, lectures, webinars and workshops into timestamped transcripts. Run it on your own computer with whisper.cpp, or in the cloud with speaker labels and AI summaries.

  • Local mode: whisper.cpp runs on your computer, the audio is never uploaded
  • Cloud mode: speaker labels and AI summaries for show notes
  • Export as TXT, JSON, SRT, VTT or Markdown
Noteo transcript of a long podcast episode with timestamps and export options

How it works

1) Choose a mode
Local mode runs in the Noteo desktop app with whisper.cpp. Cloud mode uploads the file and adds speaker labels and AI summaries.
2) Add the episode
Pick the final mix as MP3, M4A, WAV, FLAC, OGG or Opus, or a video file such as MP4, MOV, MKV or WebM.
3) Follow the progress
In Local mode Noteo shows each stage: extracting audio with ffmpeg, loading the model, transcribing, then parsing the result.
4) Export and publish
Download SRT or VTT for your video or web player, Markdown for your archive, or TXT and JSON for other tools.

Why long recordings are hard to transcribe

A two-hour episode holds tens of thousands of spoken words. Typing that up by hand takes far longer than the recording itself, yet you still need text for show notes, chapter markers, quotes for social posts, and an accessible transcript page that search engines can read.

Noteo is built for podcasters, lecturers, webinar hosts, researchers and anyone who records long conversations and wants a searchable, timestamped transcript without retyping it.

Worked example: a 2-hour interview episode to show notes

Say you host an interview podcast and have just exported a two-hour episode as an MP3. Here is a typical workflow with Noteo:

  • Upload the MP3 in Cloud mode. The transcript is split by speaker, so host and guest turns are easy to follow.
  • Skim the transcript and click any passage to jump to that moment in the audio player, which helps when checking names and quotes.
  • Open the Summary tab and generate an AI summary as a first draft of your show notes, then edit it in your own voice.
  • Use the timestamps next to each paragraph to pick chapter markers and pull quotes.
  • Download a .vtt file for your web player, a .srt file for your video edit, and a .md file for your episode archive.

Local vs Cloud for podcast transcription

Both modes produce a timestamped transcript and use the same export menu. The difference is where the audio is processed and what you get on top.

  • Local mode: the file stays on your computer and is transcribed by whisper.cpp with the model you choose. Speed depends on your hardware and model size. There are no speaker labels, and the transcript is grouped into roughly 30-second paragraphs.
  • Cloud mode: the file is uploaded for transcription, with automatic language detection and speaker labels. AI summaries and chat with the transcript are only available here, and usage counts against your plan minutes.
  • Unreleased episodes or sensitive interviews suit Local mode. When you want speaker turns and draft show notes, use Cloud mode.

How Noteo compares with other options

Manual transcription gives you full control, but it takes hours for every hour of audio. Auto-captions on a publishing platform are convenient, but they stay tied to that platform and are not designed for show notes or an archive.

You can also run whisper.cpp yourself from the command line, which is what Noteo uses in Local mode. What Noteo adds is a file picker, progress stages, a library of past transcripts, a synced audio player and one-click exports, so you do not need a terminal.

Limitations to know before you start

Local mode needs ffmpeg, a whisper.cpp binary and a whisper model file on your computer, set up in the desktop app settings. Long files on slower machines can take a while. Cloud transcription is subject to your plan limits.

No automatic transcript is perfect. Always check guest names, brand names and technical terms before you publish. Subtitle cues follow transcript paragraphs, so they are longer than typical caption lines.

Data flow (Local vs Cloud)

Local mode runs entirely in the Noteo desktop app. Cloud mode uploads the file for transcription.

Stays on your device
  • Audio/video file
  • Transcript + timestamps
  • SRT/VTT/Markdown exports
May use cloud (optional)
  • Cloud mode: upload, speaker labels and transcription
  • Optional: AI summary and chat in Cloud mode

FAQ

Can Noteo transcribe a two-hour podcast episode?

Yes. Noteo is designed for long-form audio such as full episodes, lectures and webinars. In Local mode whisper.cpp runs on your own hardware, so processing time depends on your computer and the model you pick. In Cloud mode the episode counts against the transcription minutes of your plan.

Does Noteo separate the host and guest voices?

In Cloud mode, yes. Uploaded recordings are transcribed with speaker labels, so each turn is attributed to a speaker, and you can rename speakers in the transcript view. Local mode does not separate speakers; it groups the transcript into roughly 30-second paragraphs with timestamps.

Which audio formats can I upload?

Common podcast formats work, including MP3, M4A, WAV, FLAC, AAC, OGG and Opus. You can also add video files such as MP4, MOV, MKV and WebM. Noteo extracts the audio track before transcribing, so you do not need to convert the file first.

Can Noteo write my show notes?

Noteo can draft them. In Cloud mode, the Summary tab generates an AI summary of the transcript, which makes a solid starting point for show notes. Treat it as a draft: review facts, add links and rewrite it in your own voice. AI summaries are not available in Local mode.

Which export should I use for publishing?

Use VTT for HTML5 web players and SRT for most video editors and video platforms. Use Markdown if you keep an episode archive in a notes app, and TXT for a plain transcript you can paste into your website or CMS as an episode transcript page.