Select the main language
Keep Auto detect for a recording with one clear dominant language, or choose English, Chinese, Spanish, French, German, Japanese, Korean, Portuguese, Italian, Hindi, Arabic, or Russian.
Speech to text focuses on the spoken language in a recording. Use this page when language recognition, language selection, or multilingual support is the main reason you need transcription. The local Whisper model can process many languages, including English, Chinese, Spanish, French, German, Japanese, Korean, Portuguese, Italian, Hindi, Arabic, and Russian.
Select the spoken language first, then create a private timestamped transcript in your browser.
Auto detect is convenient when one language dominates the file. For a short recording, a strong accent, or a file with specialist vocabulary, selecting the language manually can make the result easier to review. Mixed-language speech may still require manual correction after transcription.
The tool does not send your audio to a speech API. It downloads the model to the browser during first-time setup, decodes the file on your device, and presents timestamped segments for editing. The model and browser still have practical limits, so important names, figures, and quotes should be checked against the recording.
Keep Auto detect for a recording with one clear dominant language, or choose English, Chinese, Spanish, French, German, Japanese, Korean, Portuguese, Italian, Hindi, Arabic, or Russian.
Add names, organizations, products, and technical terms that are likely to matter. This context helps review without automatically forcing a spelling into the transcript.
Play the source from a timestamp, correct the text in place, assign manual speaker labels if needed, and export the reviewed transcript.
Detection works best when one language is dominant. Recordings with frequent code-switching or overlapping speakers may need a manual language choice and closer review.
The speech recognition model runs locally after setup. This is useful for recordings that should not be placed in a third-party cloud queue.
Recognition quality depends on microphone distance, background noise, accents, speaking speed, and the vocabulary in the recording. Review the words that carry meaning.
Example output only. The actual transcript depends on the recording, language, microphone, and background noise.
[00:00:06.900] The next section compares the Chinese and English product names.
The multilingual model includes English, Chinese, Spanish, French, German, Japanese, Korean, Portuguese, Italian, Hindi, Arabic, Russian, and other supported Whisper languages.
Use Auto detect when one language is clearly dominant. Choose a language manually when the recording is short, noisy, or uses several languages.
This workflow creates text in the spoken language. It does not promise a separate translation step.
No. Mixed-language speech, overlapping voices, and technical vocabulary require careful human review even when the model recognizes the main language.