What video to text actually extracts
The video workflow reads the spoken audio track from an MP4 and returns timestamped text. It does not read text shown on screen, identify who is visible in the frame, or describe visuals. If the value of a caption depends on an on-screen chart or a sign, add that information in the video editor after transcription.
The current local tool accepts MP4 files up to 200 MB or 20 minutes. A longer video should be split into chapters or parts before processing, with a note at each split point so the caption files can be joined in order.
SRT versus VTT at a glance
SRT and VTT are both plain-text formats that pair a short caption with a start time and an end time. SRT is the older, most widely supported format and works in most video players and editors. VTT is the WebVTT format used by browsers and many web players, and it supports a few extra features such as cue settings.
For most caption work the difference is small: both give you timecoded text that a video editor or platform can import. If the destination is a web page, start with VTT. If the destination is a desktop editor or a platform that only accepts SRT, start with SRT. The important part is reviewing the wording and timing either way.
Prepare the video for a caption draft
Check that the spoken audio is clearly audible before processing. Music beds, loud sound effects, and heavy background noise lower the quality of the draft. If a section is mostly music with a few spoken words, the model may produce text from the music; mark that section for manual deletion or correction during review.
Add presenters, product names, places, and slide vocabulary to the names and terms field when they appear in the soundtrack. This context helps the local model and gives you a checklist. Remember that terms on the screen are not part of the recognition, so captions for spoken words only need spoken-word review.
From transcript to subtitle file
Upload the MP4, choose the spoken language, and start transcription. The video preview lets you replay any timestamp while reviewing. When the draft is ready, work through the segments that contain names, numbers, and sentence boundaries, then export.
The export produces a caption file from the returned timing. Caption breaks, punctuation, and line lengths are still your responsibility: the tool provides the raw segments, not a finished broadcast caption track.
- Upload an MP4 with clear spoken audio and select the dominant language.
- Review each segment against the video, correcting names, numbers, and sentence breaks.
- Export VTT or SRT, then import the file into your video editor for final timing and styling.
Fix timing and line breaks
Readable captions usually contain one or two short lines, a complete thought, and enough time for a viewer to read them. A raw segment from a speech model may be too long, too short, or split mid-phrase. Adjust breaks in the editor so each cue contains a coherent phrase rather than a word count target.
If the video has an intro delay or a cut, you may need to shift the whole caption track by a constant offset in the video editor. Check the first and last cue against the picture, then verify a few cues in the middle before publishing.
Multi-speaker videos need manual labels
When a video contains two or more speakers, the local workflow lets you assign manual labels such as Speaker 1 through Speaker 4 after transcription. It does not automatically identify voices. Listen to each segment and assign the label based on what you hear.
For caption files, decide whether speaker names belong in the text at all. Some platforms prefer plain captions without names; interview videos often benefit from a name prefix. Keep the decision consistent across the whole file, and remove prefixes in the export that is meant for accessibility-only captions if the platform requires it.
Common caption workflows
For a course video, captions are usually part of the published lesson: export VTT or SRT, check timing, and upload the file with the video. For social clips, captions are often burned into the picture for silent viewing, which means you still need to review line breaks after the video editor renders them.
For YouTube, you can upload the SRT or VTT as a starting point and then adjust timing inside YouTube Studio if the auto-sync needs changes. Keep the original transcript export as a searchable archive alongside the caption file, because the two serve different purposes.
Limits to expect
Recognition quality depends on the recording, not on the export format. Strong accents, quiet speech, overlapping voices, and music can all produce errors that look fluent. On-screen text, gestures, and visuals are outside the transcript, so captions may need extra notes added manually for accessibility.
Treat the caption file as a first pass for timing and wording. The local model does not return a reliable confidence score for every word in this version, so do not rely on an automated highlight to find errors. Use the audio, the glossary checklist, and a human listen for the segments that matter.
Keep the reviewed source
An export is a snapshot, not a backup. Save the final caption file where your video project expects it, and keep the original video and the reviewed transcript in the same place so a later correction does not require re-transcribing from scratch.
For the full workflow, use the video to text tool. If you are captioning an interview and need speaker attribution, the interview review guide shows how to keep quotes and labels verifiable.