Bring your file
Drop in an MP3, WAV, M4A, or MP4. The browser checks the format, size, and duration before processing begins.
Turn voice recordings and spoken audio into searchable, editable text with private voice to text and speech to text processing in your browser.
Run Whisper locally in your browser. Nothing is uploaded and no account is needed.
People search these phrases from different starting points. The workflow is the same: choose a recording, transcribe it locally, and review the words against the source before sharing.
Convert voice notes, personal recordings, and spoken ideas into text you can search, edit, and organize.
Explore voice to textUse multilingual speech recognition for English, Chinese, Spanish, and other supported languages with auto detection or a manual choice.
Explore speech to textUpload a saved MP3, WAV, M4A, or MP4 file and create an editable transcript with timing information.
Explore audio to textUse the format you already have. MP4 video is decoded for its audio track, while the transcript keeps timing information for review and subtitle export. The current local workflow is designed for one file at a time.
MP3, WAV, M4A, and MP4
Auto detect or choose a supported language before processing
Up to 200 MB or 20 minutes per file
Desktop Chrome or Edge with JavaScript enabled
Everything needed to prepare, review, and export a file stays in one browser workspace.
Drop in an MP3, WAV, M4A, or MP4. The browser checks the format, size, and duration before processing begins.
Select a language and add names, brands, or technical terms the model may otherwise mishear. These terms provide context and are not inserted automatically.
Jump from timestamped segments to the audio, assign speaker labels when needed, edit the text, and export the version you trust.
A transcript should stay connected to the recording, especially when names, quotes, or technical terms matter.
Move from the text back to the relevant point in the audio instead of searching through the whole recording.
Mark segments as Speaker 1 through Speaker 4 during review when a conversation has multiple participants. Automatic speaker diarization is not enabled in this local build.
Add people, brands, products, or specialist vocabulary before transcription and check glossary suggestions afterward.
Download TXT for notes, JSON for structured data, or SRT and VTT for caption and subtitle workflows.
This is a sample format. Your exported transcript uses the words and timestamps returned by the local model.
[00:00:04.120] Speaker 1: Let us review the launch timeline. [00:00:08.640] Speaker 2: The first draft is ready for Tuesday.
Each workflow page explains the relevant format, review steps, and limits before you start.
Short guides explain the model setup, file preparation, and review decisions that matter when accuracy is important.
The local model runs in a browser worker. Your audio and transcript are not sent to us, and the FAQ explains what the first version can and cannot do.
They describe closely related starting points. Voice to text usually begins with a voice note or personal recording, speech to text focuses on spoken language, and audio to text covers saved files such as MP3 or WAV. This browser tool supports all three workflows.
No. The local workflow decodes your file and runs the Whisper model in your browser. Your audio and transcript are not sent to this website for processing.
The first version accepts MP3, WAV, M4A, and MP4 files up to 200 MB or 20 minutes. MP4 video is read for its audio track and can be exported with timestamps for subtitle workflows.
Yes. The multilingual Whisper model supports many languages, including English, Chinese, Spanish, French, German, Japanese, Korean, Portuguese, Italian, Hindi, Arabic, and Russian. Auto detection works best when one language is dominant in the recording.
No account is required and there is no desktop application to install. The first run downloads the local model into your browser cache; later runs can reuse it, including offline after the model has been cached.
You can edit transcript text, jump from timestamped segments to the audio, assign manual Speaker 1 through Speaker 4 labels, and add names, brands, or technical terms before transcription. Export is available as TXT, JSON, SRT, or VTT.
The current local version does not promise automatic speaker diarization. It provides manual speaker labels for review, while overlapping speech, background noise, and distant microphones still need human checking.