Subtitle field guide

SRT, VTT, captions, and subtitle workflows

A transcript says what was spoken. Subtitles add time, reading constraints, line breaks, sound cues, and a playback review. SRT and WebVTT solve related—but not identical—jobs.

4 minute read
Audio transcription workspace with a microphone, headphones, laptop, and monitor

What SRT and VTT contain

A basic SRT file is a sequence of numbered cues. Each cue has a start time, an end time, one or more text lines, and a blank separator. Hours, minutes, seconds, and milliseconds use a comma before milliseconds. WebVTT begins with a WEBVTT header, uses a period for milliseconds, and supports additional cue settings and web-oriented features.

Neither extension proves the content is accessible or synchronized. Players differ in accepted styling and encoding. Keep an unstyled master and test the exact delivery platform.

Timing is an editorial task

Start a cue close enough to the speech to feel connected, but avoid flashing text too briefly to read. End it at a natural pause without colliding with the next cue. Split on phrases and sentence structure rather than arbitrary character counts. Avoid orphaning a short word on a second line.

Reading speed, line length, number of lines, speaker changes, music, important sound effects, and shot changes all influence a good cue. Platform specifications can be stricter than a general guideline.

From automatic timings to a publishable file

Begin with the best available transcript and timecodes. Correct the words against the media first. Then merge fragments, split overlong cues, repair overlaps and gaps, add speaker or sound information required for the audience, and review the complete program at normal speed.

The free formatter on this site can turn prepared paragraphs into a starter SRT timeline using estimated durations. It does not listen to the media, so its times are drafting scaffolding and must be aligned before publication.

Final checks

Open the file in the target player, check the beginning and end, scrub across cue boundaries, and inspect names, numbers, punctuation, music, and overlapping voices. Confirm UTF-8 encoding and that line breaks survive upload. Watch with the sound off to test whether the text carries the essential information.

For accessibility compliance or a platform delivery contract, follow the applicable specification and involve a qualified captioner where needed.

Sources

  1. W3C official WebVTT specification
  2. Library of Congress — SubRip subtitle format description

Meet VTA Dictate

Turn more of what you say into work you can use.

Voice type across Windows, record conversations with permission, transcribe audio, shape meeting notes, and export subtitles from one desktop workspace.