What the exported files contain
The editor works with segments, and each segment has a timestamp range and text. Depending on your edits, a segment may also have a manually assigned speaker label. Exports are generated from that reviewed state, so the text you export should match what you see in the editor.
All four formats are plain text files that can be opened and moved between devices. They do not contain the audio itself, and they are not linked to the processing session after download. Keep the original recording separately.
A short example makes the difference visible. The same segment exported as TXT reads “00:01:12 [Speaker 1] The product review will be ready after the final audio check.” As SRT it becomes a numbered cue with separate start and end times, and as JSON it keeps the segment fields a script can read. The words are identical; the structure is what changes.
TXT for reading and notes
TXT is the simplest export: readable sentences, usually with timestamps where the segment timing is useful, and no markup to manage. It is the right choice for pasting into a notes app, drafting show notes, or sharing a plain-language summary with someone who does not need timing.
If you need the timestamps as part of the readable text, TXT keeps them inline. If you only need clean prose, remove or ignore the timestamps during the final copy. TXT has no structure for speaker labels or word-level timing, so use it for reading, not for machine processing.
JSON for structured workflows
JSON preserves the structure of the transcript: segments, their time ranges, and the text in each segment. This makes it the right export when a script, a website, or an archive needs to read the transcript as data rather than as a document.
The exact fields come from the transcript the tool returns, so inspect one export before writing code against it. JSON is also a good format for keeping a searchable personal archive because the timestamps survive future processing.
SRT for video and players
SRT is the most widely supported subtitle format. Video editors, media players, and most publishing platforms accept it. Each cue has a number, a start and end time, and one or more caption lines.
Use SRT when the destination is a desktop editor, a course platform, or any system that does not advertise WebVTT support. Expect to adjust line breaks and timing in the video editor even if the words are correct, because caption readability is a formatting task.
VTT for web captions
VTT is WebVTT, the caption format used by browsers and most modern web video players. It looks similar to SRT but follows WebVTT rules and supports a few extra settings that browsers can interpret.
Choose VTT when the captions will live on a web page, such as an embedded video in a course or a blog post. If the platform accepts both, either works; when in doubt, the platform documentation usually names the preferred format.
Choose by your next step
The deciding question is what happens after the download. A note-taking workflow wants TXT. A script, archive, or analysis workflow wants JSON. A video edit wants SRT. A web page wants VTT. Exporting more than one format is normal: many users keep TXT for reading and JSON for the archive.
Export order matters less than review order. Export after the important names, numbers, and quotes have been checked, because the download is a snapshot of the current editor state. If you edit after exporting, export again before distributing the file.
- Review names, numbers, and quotes against the audio while the editor is open.
- Choose TXT for notes, JSON for structure, SRT for desktop video, or VTT for web video.
- Download, keep the original recording, and re-export if you make later edits.
Common export mistakes to avoid
A common mistake is treating one export format as a universal file. A team that needs readable notes will not open an SRT comfortably, and a developer importing captions will not want a TXT with timestamps glued into sentences. Choose the format for the destination, and when a workflow genuinely needs two outputs, export both from the same reviewed editor state.
Another frequent issue is distributing a caption file without checking the first and last cue. A missing final line, a shifted timecode after a video cut, or a caption that overruns the video length is easy to catch by playing the first minute and the last minute. The same check applies when you edit the transcript after exporting: the old file does not update itself, so re-export before sharing.
Check before you export
A short pre-export check prevents the most common mistakes: confirm the language of the first and last segment, verify every proper name and number that will be quoted, and make sure sentence boundaries read naturally. If the recording had music or overlapping speech, decide how those sections are represented in the text.
Speaker labels, where you assigned them manually, appear in the exports that support them. Check one sample file after downloading: open the TXT or JSON, or import the SRT or VTT into the player you plan to use. A file that opens cleanly now will save you a support question later.
Keep the source recording
Exports are copies of the reviewed text, not backups of the processing session. The transcript lives in the browser until the tab closes, so save the files you need and keep the original audio with them. Local processing keeps the recording private, but only your file management keeps it available.
For a direct workflow, use the audio to text tool. If the file is video and you need captions, the video to text guide walks through the caption-specific review before export.