Check the MP3 first
Use an MP3 up to 200 MB or 20 minutes. A clear recording with the microphone close to the speaker usually gives the editor more useful material to review.
MP3 is one of the most common formats for saved recordings, podcasts, interviews, voice notes, and exported meeting audio. This MP3 to text workflow lets you use the file you already have instead of re-recording it or sending it to an online upload service.
Convert a podcast, voice note, or saved MP3 locally and keep its timestamps attached to the draft.
Choose a language before processing, then add important names, brands, or specialist terms. The local Whisper model uses that information as prompt context. It does not automatically replace the transcript with your glossary, so the final spelling still needs to be checked against the audio.
The result is organized into timestamped segments with editable text. Use the player and timestamp controls to inspect a quote, assign manual Speaker 1 through Speaker 4 labels for a conversation, and export TXT, JSON, SRT, or VTT depending on how you will use the recording.
Use an MP3 up to 200 MB or 20 minutes. A clear recording with the microphone close to the speaker usually gives the editor more useful material to review.
Select the main language and list people, companies, products, or technical words from the recording before starting the model.
Correct the segments that matter, add manual speaker labels if useful, and choose plain text, JSON, SRT, or VTT after listening back.
MP3 files are compact and widely supported, making them practical for interviews, saved calls, lectures, podcasts, and personal recordings.
A transcript is easier to trust when each segment can take you back to the source. This is especially important for quotes and action items.
An MP3 can be clear or difficult depending on the original microphone, compression, background noise, and speaker distance. File format alone cannot guarantee accuracy.
Example output only. The actual transcript depends on the recording, language, microphone, and background noise.
[00:02:31.420] Speaker 1: The customer asked for a shorter onboarding flow.
The initial browser workflow accepts files up to 200 MB and 20 minutes long.
The model may still detect speech, but music, applause, and loud background sound can reduce accuracy and should be checked carefully.
Yes. The SRT export uses the timestamped transcript segments and can be used as a starting point for caption work.
No. The file is selected and processed in your browser. The local workflow does not create server-side storage.