A script you can say out loud is one written for the mouth, not the eye — and the reliable way to get there is a simple structure plus one ruthless edit pass where you read the draft aloud and cut every phrase you would never say to a friend. Most video scripts fail on camera not because the ideas are weak but because they were written in written English, which is a different language from the one you speak.
This piece gives you the four-block structure that underpins almost every talking-head video that works, the read-aloud pass that makes it performable, the formatting habits that help you breathe, and a worked example taking one script through the whole process.
The four blocks: promise, idea, proof, next action
Strip the platform advice down to what it agrees on, and nearly every effective short video has the same skeleton:
- The promise. What the viewer gets, stated immediately. Not a greeting, not who you are, not "in today's video" — the first sentence delivers or previews the value the title promised. If someone reads only your opening two lines, they should know what they'll learn.
- One useful idea. One. The most common script problem isn't bad writing, it's three videos trying to share one runtime. If you find a second idea in the draft, you've found your next video.
- Proof. An example, a demonstration, a before-and-after, a moment from your own experience — something concrete that shows the idea rather than restating it. This is the block that separates advice from assertion.
- One next action. A single closing beat: try this today, watch the follow-up, tell me what happened when you tried it. Not every video needs a conversion ask; a clean takeaway is a perfectly good close. But two asks compete and both lose.
The blocks are a checklist, not a straitjacket — a personal story might spend most of its time in proof; a quick tip is mostly idea. What the structure buys you is the ability to diagnose a limp draft: it's almost always a missing promise, a second idea, or proof that's really just the idea again in different words.
The read-aloud pass
This is the edit that matters most, and it takes ten minutes.
Read the whole draft aloud, at performance volume, and change every phrase your mouth resists. Not skim-mumble — actually say it, as you would on camera. The test for each sentence is brutal and simple: would I ever say this to a friend across a table?
Written English fails that test constantly, in predictable ways:
- Uncontracted verbs. "Do not", "it is", "you will" — fine on the page, stilted aloud. If you'd say "don't", write "don't".
- Sentences with three clauses. You'll run out of breath, or your intonation will go strange trying to voice the commas. Split them. Full stops are free.
- Words you'd never speak. "Utilise", "in order to", "individuals", "however" at the start of a sentence. Your speaking vocabulary is smaller than your writing vocabulary, and the audience clicked for the speaking one.
- Throat-clearing. "So, before we get into it, I just want to quickly mention" — six words of it survive in speech ("quick note first"), the rest is the page talking.
Everything that trips you in this pass will trip you harder on camera, with a light on and a take running. Fix it now, in the cheap medium.
Short lines and breaks as breathing marks
How a script looks on the prompter shapes how it sounds out of your mouth.
Format in short lines — one phrase or one clause per line — rather than dense paragraphs. Each line break becomes a natural micro-pause, so the layout does some of your pacing for you, and a short line is far easier to take in at a glance from two metres away. A paragraph break is a bigger breath: use one wherever you'd pause to let a point land.
Two related habits: put a deliberate break before a hard word or a number, so you arrive at it settled rather than mid-flow; and if you plan to depart from the script — an aside, a story you'll tell freely — write the departure in as a short cue line ("story: the first attempt") rather than scripting it. A script you're allowed to leave stops feeling like a cage, and a voice-following prompter will keep your place and pick you up when you return to the written words.
Two or three hooks, not three scripts
The opening seconds do disproportionate work, so give the opening options — but only the opening.
Once the body is written, draft two or three alternative first lines for the same script: a question ("why do your takes sound nothing like you?"), a direct promise ("this is the edit that makes a script performable"), a contrarian statement ("most script advice makes you sound worse"). Record the body once; record the hooks as separate short takes if you like, and choose in the edit.
This is far cheaper than generating whole alternative scripts, and it respects what's actually uncertain: the body either works or it doesn't, but which door into it suits the platform and the mood is genuinely hard to predict. Keep the hooks honest — a question the video answers, a promise it keeps.
Durations are presets, not promises
You'll see confident claims about the ideal length for a Reel, a Short, a LinkedIn video. Treat all of them as targets to write towards, not rules that predict anything — the platforms themselves say appropriate length depends on the content, and their guidance changes.
The practical use of a target duration is as a writing constraint: aiming at 60 seconds forces the one-idea discipline better than any editing advice. But the number that actually matters is your read time, at your pace, which is usually slower than you'd guess once real pauses are in. Measure it by reading aloud — TellyPrompter shows a measured read-time for a script and has a rehearsal mode for exactly this pass — and then cut or expand against reality rather than against a words-per-minute rule of thumb. When the read runs long, cut the idea's second example before you speed up your delivery; a rushed read costs more than a longer video.
Worked example: one script through all five steps
Say the video is "why your camera batteries die faster in winter", aiming at a 60-second target.
First draft, first lines (written English): "Hi everyone. Today I want to talk about something that affects a lot of content creators during the colder months, which is the fact that battery performance can be significantly reduced in low temperatures."
Blocks check: the promise arrives in sentence three; there's a second idea lurking in the draft about storing batteries, which becomes its own video; proof is currently a restatement ("they really do drain faster") and needs the concrete version — the shoot where two batteries died in twenty minutes; the close asks viewers to like, subscribe and comment, which is three asks.
After the read-aloud pass: "Your camera batteries aren't broken. Cold just drains them faster — and there's a cheap fix. Last January I lost two batteries in twenty minutes on one shoot. Here's what changed. Lithium batteries slow down as they get cold. Keep the spare in an inside pocket, not the bag. Body heat is enough. Swap warm for cold; the cold one recovers as it warms back up. Try it on your next cold shoot, and tell me how long they lasted."
Formatted for the prompter: each of those sentences on its own line, a paragraph break before "Here's what changed", and a cue line — "(show the pocket)" kept visually distinct from the spoken text.
Hooks drafted: the direct version above; a question ("why did both my batteries die in twenty minutes?"); a contrarian ("stop buying more camera batteries").
Measured: read aloud once, it comes to about 50 seconds at a real pace — room to breathe inside the 60-second target, so nothing gets rushed.
That's the whole method. Structure to diagnose, the read-aloud pass to translate, layout to breathe, options where the uncertainty really lives, and your own measured read as the only duration authority. The script that survives it will be shorter, plainer and less impressive on the page than your first draft — and it will sound like you, which is the entire job.