From speech to timed captions

Create SRT and VTT subtitles on Windows

Use VTA Dictate to transcribe audio or video, refine the words and timing, and export an SRT or VTT subtitle file.

4 minute read
Audio transcription workspace with a microphone, headphones, laptop, and monitor

Begin with a timed transcript

Open a supported media recording and create a local transcript with timing information. Correct the spoken words first, because a beautifully timed subtitle is still wrong if the underlying transcript misheard a name, number, or key phrase.

Break long passages at natural pauses and sentence boundaries. Short, readable cues are easier to follow than dense blocks that force the viewer to choose between reading and watching the picture.

Choose SRT or VTT for the destination

SRT is a simple, widely supported sequence of numbered cues with start times, end times, and text. WebVTT uses a related structure and adds features used by many web players. The publishing platform or editor normally tells you which format it accepts.

Keep a clean text encoding, avoid overlapping cue times, and check line breaks on a small screen as well as a desktop monitor. Names, sound descriptions, and speaker changes should be consistent throughout the file.

Watch the finished video with subtitles on

Import the subtitle file into the target editor or player and watch from beginning to end. Check entrances, exits, fast dialogue, music, silence, scene changes, and any place where the speaker is difficult to hear.

Subtitle creation combines automation with editorial judgment. Timing can be generated from speech, but readability, emphasis, translation choices, and accessibility still benefit from a person who understands the audience and the media.

Meet VTA Dictate

Turn more of what you say into work you can use.

Voice type across Windows, record conversations with permission, transcribe audio, shape meeting notes, and export subtitles from one desktop workspace.