Use case
Transcribe Zoom, Google Meet and Teams recordings on your own computer with whisper.cpp. The audio is never uploaded, and the transcript is saved on your disk.

Most meeting transcription tools work the same way: the recording goes to a server, a cloud model transcribes it, and the text lives in someone else’s database. That is convenient, but it is a poor fit for board discussions, HR conversations, legal strategy calls, unreleased product plans, or any meeting where a client asked you not to share the recording.
In those cases the real question is not “which tool is most accurate?” but “can I get a usable transcript without the audio leaving my laptop?” Local transcription answers that: the speech-to-text model runs on your own hardware, so there is no upload step to approve, audit or explain.
Local mode is built for people who record meetings themselves and want the text without handing the audio to a third party:
Say you record a 45-minute Zoom call with a client using Zoom’s local recording, which leaves an MP4 on your computer. In Noteo, you switch to Local mode and import that file. Noteo calls ffmpeg to strip the video and write a 16 kHz mono WAV (the input format whisper.cpp expects), then runs whisper.cpp with the model you selected, for example a base or small multilingual model. Progress is shown while it runs, including the detected language when set to auto-detect.
When it finishes, the meeting appears in your local library with the full text and timestamps. You export a Markdown file into your project notes and an SRT file if you want captions on the recording. At no point was the MP4 or the transcript sent to Noteo’s servers.
Local and Cloud mode solve different problems, and Noteo lets you switch between them. Choose based on the meeting, not on habit.
Cloud-only note takers are the fastest way to get summaries and speaker labels, but every recording is uploaded by design. Running whisper.cpp yourself from the command line is fully private, but you end up juggling WAV conversion, output files and subtitle formats by hand. Noteo’s Local mode sits in between: the same open-source engine running on your machine, with a library, a transcript viewer and one-click exports on top.
Local mode is a desktop feature, so it does not run in the web app or on mobile. You install and point Noteo to ffmpeg, whisper.cpp and at least one ggml model file yourself. Transcription quality depends on the model you pick and on the audio: crosstalk, heavy accents and poor microphones reduce accuracy with any engine. There is no speaker separation locally, so for multi-speaker meetings you may need to add names while reviewing. During processing, ffmpeg writes an intermediate WAV file to your operating system’s temporary folder.
Local mode is designed for offline transcription. Cloud is optional and only needed for AI features.
Yes, the transcription itself runs on your computer. ffmpeg extracts the audio and whisper.cpp converts it to text using a model file stored on your disk, so the recording is not uploaded. Cloud features such as AI summaries and chat are disabled while Local mode is on.
You need the Noteo desktop app, ffmpeg, a whisper.cpp binary (whisper-cli) and at least one ggml model file in a folder you choose. The Offline settings page checks each dependency and tells you what is missing before you start a transcription.
Local mode can auto-detect the spoken language, or you can pick one in Offline settings. The list includes English, French, German, Spanish, Italian, Portuguese, Dutch, Polish, Japanese, Chinese, Arabic and around twenty more. Use a multilingual model, since English-only models only transcribe English.
Yes, as long as you have a recording file. Record the call with the platform’s own recorder or a screen recorder such as OBS, then import the audio or video file in Local mode. Noteo converts it with ffmpeg and transcribes it with whisper.cpp on your machine.
Not yet. Local transcripts are produced without speaker diarization, so the text appears under a single speaker. You can add names while reviewing the transcript. If you need automatic speaker labels, Cloud mode includes speaker recognition, but it requires uploading the recording.
Made with ❤ in Paris, France
Copyright © 2025 Noteo.ai. All Rights Reserved