If you read a video from a script, that script can provide the text for captions, while the take provides timing. TellyPrompter combines the script with available take timing to prepare SRT or VTT files with Pro or during a complete free first session. It does not transcribe improvised words, and both text and timing need checking against the recording. This guide explains where that method helps and where transcription or manual editing is still needed.
Why captions are worth the bother
Two audiences read your captions, and both are larger than most creators assume.
The first watches with the sound off — on phones, in public, in feeds where video autoplays muted and the first seconds decide everything. For them, captions aren't an accessibility extra; they're the only version of your words that exists.
The second relies on captions to follow speech at all: deaf and hard-of-hearing viewers, people watching in a second language, anyone in a loud room. For them, caption accuracy is the difference between being included and being guessed at.
Platforms will happily auto-generate captions, and auto-captions are far better than nothing. But they're transcription — and for scripted video specifically, transcription is solving your problem backwards.
SRT and VTT in plain terms
A caption file is simpler than its reputation. The two formats you'll meet:
SRT is the old workhorse: numbered blocks of text, each with a start and end time. Nearly everything accepts it — YouTube, editing software, most platforms.
VTT (WebVTT) is the web-native sibling: much the same idea with more styling and positioning options, and the format HTML5 video players speak natively. If you embed video on your own site, VTT is the one you want.
Both are plain text with timestamps. The entire job of captioning is producing two things accurately: the words, and the times at which they're spoken. Keep those two halves in mind, because they're where the two approaches differ.
Where transcription goes wrong on scripted delivery
Automatic transcription listens to your audio and infers the words statistically. On improvised speech it can be a useful starting point, alongside manual transcription. On scripted delivery, it has a structural absurdity to it: the exact words already exist, in a file, written by you — and the transcriber ignores that file and guesses from sound.
The guesses fail in predictable places:
Proper nouns and product names. Your name, your channel's name, the tool you're reviewing, the town you're in — precisely the words that matter most in your video are the words a general speech model has seen least. They come back mangled, differently each time.
Technical and niche vocabulary. Every field's jargon is a transcription minefield, and a channel is nearly always in some field.
Homophones resolved by guesswork. The model picks whichever spelling is statistically likelier, which is not always the one you wrote.
Your stumbles, faithfully preserved. Transcription captures what you said, including the false start you'd rather captions smoothed over.
The result is a familiar workflow: generate auto-captions, then spend twenty minutes per video re-reading them line by line, fixing the mistakes — a proofreading pass against your own script, done by eye. The words you're correcting them to were on your screen the whole time.
The other half: timing
Suppose you skip transcription and use the script directly — paste it into a caption tool. Now you have the written words, but still need to check that they match the performance, and the times are missing: nothing knows when each line was actually said in the take. Manual syncing means dragging caption blocks along a waveform, and it's tedious enough that most people give up and go back to fixing auto-captions.
So the real requirement for script-accurate captions is both halves at once: the script's exact words, and per-line timing from the actual take. Which is where the teleprompter, of all things, turns out to hold the answer.
Captions from the prompter: script plus available timing
A voice-following teleprompter works by matching your speech against the script as you read — that's how the highlight stays on the word you're saying, through pauses, ad-libs and backtracking. (The full comparison of scroll modes is in voice-following vs constant scroll.)
Matching can produce timing positions for words in the script during a take. Those positions depend on the browser recogniser and the app's alignment; they are not a guarantee that every word was spoken or timed correctly.
TellyPrompter can prepare SRT or VTT caption files from your script and available take timing with Pro or during a complete free first session. On the free-first-session offer, complete your first session free and continue with Pro. Existing Free users and visitors on the earlier offer keep their ongoing Free access. Review and original downloads remain accessible after you finish. Check the exported text and timing before publishing: recognition errors, skipped lines and improvised words can require correction. See what Pro includes.
On a batch day this compounds: export captions per take while everything's fresh, and the edit gets caption files to check alongside the footage. The take itself is hands-free too — recording starts with the mic, and "Telly, stop" and its siblings run the session from your mark — so usable timing can be collected alongside the read. Check it before exporting.
The honest scope
Worth stating plainly, because a captioning claim should be exact about what it covers.
The export uses the script and available timing. It can preserve spellings from your script, but skipped or repeated passages and recognition errors can affect the result. Compare the prepared captions with the actual take.
Heavy ad-libbing is the transcription case. Words you improvise aren't in the script, so no script-based export can caption them — a long unscripted aside needs a transcription pass for that stretch, or a quick manual addition. Script-accurate captioning is for scripted delivery; that's the honest boundary of the idea.
The timing depends on the following. Voice-following needs browser speech recognition, whose support and accuracy vary by browser, language and microphone. Firefox does not currently provide the recognition API used here. Review a real take and its export before relying on this workflow.
Review captions against the recording. Check wording, omissions, repeated lines, pauses and timestamps. Edit or transcribe any passages the script-based export does not represent correctly.
A workflow worth stealing
For a scripted talking-head video, the captioning step can shrink to almost nothing:
- Write the script for the mouth, and read it once aloud before recording.
- Record with voice-following on — the timing accrues as you read.
- After the keeper take, prepare SRT or VTT with Pro or within your complete free first session when the required timing is available. Keep the original recording download too.
- Compare the file with the recording and correct text and timing before uploading it alongside the video.
The script is a useful starting point for captions when the performance follows it. The recording remains the reference for what was actually said. Choose script-based export, transcription or manual editing according to the take, and check the result before publishing.