Choose a saved recording
Drop in an MP3, WAV, M4A, or MP4 file up to 200 MB or 20 minutes. The file is read locally before transcription begins.
Audio to text is useful when your source is already saved as a file rather than spoken directly into an app. Upload a recording, choose a language or keep auto detection, and let the local Whisper model create timestamped text in this browser workspace.
Inspect a saved audio or video file, review its timing, and export the text locally.
The workflow accepts common audio formats including MP3, WAV, and M4A. MP4 files are also supported when the spoken content is stored in the video track. The browser checks the file before processing and keeps the original recording on your device.
Use the result as a first draft for notes, research, classes, interviews, or meeting follow-up. Timestamps let you check an important sentence against the source, while the editor lets you correct names, assign manual speaker labels, and export TXT, JSON, SRT, or VTT.
Drop in an MP3, WAV, M4A, or MP4 file up to 200 MB or 20 minutes. The file is read locally before transcription begins.
Select the main language and add names, brands, or specialist terms the recording contains. The list guides the model but does not overwrite the result.
Use timestamps to replay the relevant section, edit the transcript, and export only after the words that matter have been reviewed.
MP3 is convenient for voice notes and podcasts, WAV is useful for uncompressed recordings, and M4A is common for phone or messaging app audio.
The output is organized into timestamped segments instead of one undifferentiated paragraph, so you can move between text and audio while editing.
The first model download needs an internet connection. After the model is cached, the local workflow can be reused without sending audio to a transcription server.
Example output only. The actual transcript depends on the recording, language, microphone, and background noise.
[00:01:12.300] The product review will be ready after the final audio check.
Yes. WAV and M4A are supported alongside MP3 and MP4, subject to the 200 MB and 20-minute limits.
No. The browser decodes and transcribes the file locally. The original audio is not uploaded for processing.
Yes. SRT and VTT exports use the timestamps returned by the transcript, while TXT and JSON are available for notes and structured workflows.
Start with names, numbers, quotes, and technical terms. Click a timestamp or word to replay the approximate source position before exporting.