Choose a supported file you can inspect

The local workflow accepts MP3, WAV, M4A, and MP4 files. MP3 and M4A are convenient for phone recordings and exported voice notes. WAV can preserve an uncompressed source, while MP4 is useful when the spoken audio is attached to a video. The current limit is 200 MB or 20 minutes for one file.

Do not rename an unsupported container and assume the browser can read it. The extension and media contents both matter. If a file fails validation, export or convert it using a tool you trust, then check the resulting file before starting a long model run.

Improve the recording before you transcribe

The most helpful improvement is usually microphone position. A speaker close to the microphone gives the model a stronger voice signal than a speaker recorded from the other side of a room. Reduce continuous fan noise where possible, and avoid placing a phone on a surface that creates vibration or handling noise.

Do not over-process an already clear recording. Aggressive noise removal, clipping, or repeated lossy exports can remove consonants and make words less distinct. If you need to edit the source, keep an untouched original so you can compare a questionable transcript with the first recording.

  1. Listen to the first and last minute to check whether speech is present and intelligible.
  2. Confirm that the file is within 200 MB and 20 minutes before opening the model.
  3. Keep a clean source copy so important quotes can be checked later.

Select the language deliberately

Auto detect is a useful default when one language is clearly dominant. A manual language choice can be more predictable for a short recording, a strong accent, or a file that begins with a long pause. The multilingual model supports common languages including English, Chinese, Spanish, French, German, Japanese, Korean, Portuguese, Italian, Hindi, Arabic, and Russian.

Mixed-language recordings need more caution. If speakers switch languages frequently, one setting may not describe every part of the recording equally well. Use the generated text as a draft, listen to the transitions, and correct words that carry meaning rather than assuming detection was perfect.

When to split or simplify a recording

A file near the 20-minute limit may be easier to review if it is divided into meaningful parts, such as an interview introduction and the main conversation. Splitting does not improve the audio signal, but it can make retries, review, and export easier when one section contains a problem.

Keep a small overlap or note the split point when exact continuity matters. Otherwise, a sentence at the boundary may be harder to place in the final transcript. Name the parts clearly so the exported TXT, JSON, SRT, or VTT files can be joined in the correct order later.

If a recording contains long music-only sections, silence, or unrelated room noise, consider whether those sections need to be included. The local model can only respond to the signal it receives; removing irrelevant material may shorten review, but keep the original file for reference.

Why preprocessing cannot guarantee exact text

Volume normalization can make quiet speech easier to hear, but it cannot reconstruct a syllable that was never recorded. Heavy noise reduction can also remove parts of consonants. In most cases, a clear original with a nearby microphone is more useful than an aggressively processed copy.

Treat preprocessing as a way to make listening easier, not as a way to promise a perfect model output. After transcription, the important check is still the source audio. This is especially true for proper names, numbers, acronyms, and phrases that will be quoted or used to make a decision.

Prepare names, brands, and specialist terms

Proper names and domain vocabulary are common sources of transcription mistakes. Before processing, list the people, companies, products, places, acronyms, or technical terms that appear in the recording. Separate them with commas or new lines in the names and terms field.

This list is context for the local model and a checklist for review. It is not a silent find-and-replace operation. If a term does not appear in the recording, it should not be forced into the transcript; if the model uses a different spelling, compare it with the audio before accepting the correction.

Plan the review before you start

Decide what accuracy means for your use case. A private outline may only need the main ideas, while a published interview needs exact names, numbers, and direct quotes. This decision tells you where to spend listening time after the first draft is ready.

The editor keeps timestamps attached to segments and lets you jump to an estimated word position. Use that connection rather than reading the text in isolation. If the recording has multiple people, assign manual labels after listening to each segment; do not treat a label as automatic speaker identification.

Choose the right export

TXT is easiest to read and paste into notes. JSON preserves segments, timing, and the current transcript structure for a later script or workflow. SRT and VTT are useful starting points for captions, but subtitle timing and line breaks still need a final check in a video editor.

For a direct start, use the audio to text converter. If the source is specifically an MP3, the MP3 to text page explains that format and its common use cases. The more specific page can help you confirm that the file and review workflow match your goal.