Choose an MP4 video
Select a video with a clear spoken audio track. The browser validates the file and uses the local decoder when native playback cannot read the format.
Video to text is useful when the words are inside an MP4 recording and you need a searchable transcript, a caption draft, or a way to review a presentation without watching the entire file. The browser reads the audio track and leaves the visual track untouched.
Create a timestamped caption draft from an MP4 video without uploading the video to a server.
The current workflow focuses on MP4 files up to 200 MB or 20 minutes. It does not claim to describe on-screen visuals or identify people from the image. Its job is to extract spoken audio, transcribe it, and keep timing information attached to the words.
After processing, review the transcript against the video or audio player. You can edit a segment, assign manual speaker labels, check names and terms, and export SRT or VTT as a starting point for subtitle editing. Caption timing and punctuation should still be checked before publication.
Select a video with a clear spoken audio track. The browser validates the file and uses the local decoder when native playback cannot read the format.
Use Auto detect for a dominant language or choose one manually. Add names, brands, and technical words that appear in the video.
Play from each timestamp, edit the wording, and export SRT or VTT when the text is ready for a separate subtitle or video editing workflow.
The transcript is based on spoken audio. It does not automatically describe slides, on-screen text, gestures, or visual scenes.
The editor keeps timestamps close to the words so you can check whether a caption starts and ends at a useful point in the video.
The video is decoded in this browser. There is no cloud upload queue or server-side storage in the first version.
Example output only. The actual transcript depends on the recording, language, microphone, and background noise.
[00:04:08.210] The next slide shows the three steps in the migration plan.
It extracts the spoken audio track from an MP4 and transcribes that audio. It does not produce a visual description of the video.
Yes. SRT and VTT exports include the transcript timing, but you should review caption breaks and timing in your video editor.
The model may return no useful transcript for music or silent footage. The result should be checked before exporting.
The first version accepts MP4. Other video containers are outside the current upload boundary.