Skip to content

Media utilities

Free transcript to SRT or VTT converter

Turn a plain transcript into a timed caption file. Paced by speaking rate, anchored to any timestamps you have.

Free, no account needed. Runs in your browser: nothing is uploaded.

4.9G2CapterraProduct Hunt

What it does

Turn text into timed cues

The converter does the mechanical work of caption creation: sizing, wrapping, numbering, and timing.

Honest pacing

Cue durations come from word count at your chosen speaking rate: estimation, clearly labelled.

Anchors when you have them

Timestamps inside the transcript become exact anchors, and pacing fits between them.

Caption-shaped cues

Lines wrap at your limit, two lines per cue, breaking at sentence seams.

How to use it

How to convert a transcript to captions

  1. Paste the transcript. Timestamps on their own lines (0:30) become anchors, but plain text works.
  2. Pick the speaking pace and line length.
  3. Generate, then choose SRT or VTT.
  4. Download the file and fine-tune the timing in your caption editor.

In depth

How estimated timing works

A transcript knows what was said; it does not know when. Real caption timing comes from the audio, which is why professional workflows align captions with speech-recognition tooling. What this converter offers is the honest middle path: statistically sound timing from speaking rate, ready in one paste, labelled as the estimate it is.

Most narrators speak between 140 and 170 words per minute. At a stable pace, generated cues can land within one or two seconds of the audio. Each timestamp in the source becomes a fixed anchor, and the converter adjusts the cues between anchors.

Caption guidelines use about 42 characters per line and two lines per cue. Line breaks should follow grammatical boundaries. The converter applies these rules while it builds the cues, then runs the same checks as the caption validator.

Use the output as a strong first draft: import it into your editor, nudge the cues that drift, and ship. For a file that must be frame-accurate without manual work, transcribe from audio instead: then come back here when all you have is text.

Questions

Common questions

How accurate is the timing?

It is an estimate from your selected speaking rate. For steady narration, it can be within one or two seconds of the audio. Source timestamps become fixed anchors and improve the cues between them.

What does a timestamp anchor look like?

A time at the start of a line: "0:30 Start at the forks…". The cue at that point starts exactly there, and pacing scales to land on the next anchor.

Which output formats are supported?

SRT and WebVTT, switchable after generation. The output passes standard caption validation: sequential timing, sized lines, no overlaps.

Is my transcript uploaded?

No. The conversion runs entirely in your browser.

With Mindstamp

Captions made: now put them to work

Mindstamp plays your captions on the video and adds what captions cannot: questions, chapters, buttons, and per-viewer reporting.

Works on the video you already have. No re-encoding.