The usual answer to controlling a teleprompter mid-take is a Bluetooth remote or a foot pedal — one more device to buy, charge, pair and hold. The alternative is a prompter you can talk to: TellyPrompter runs the whole take by voice, from starting and stopping to retakes and stumble markers, so the remote you need is the one you brought — your voice.
This piece explains why hands-free matters more on a prompter than almost any other tool, what the voice commands actually do, where a physical remote still wins honestly, and how the pieces add up to a take you never touch.
Why a prompter has to be hands-free
Picture the situation the tool actually lives in. A take is running. Your hands are in frame — gesturing, holding a product, resting naturally. Your eyes are on the lens. The phone is on a rig at your mark, possibly sealed behind beam-splitter glass, and walking over to poke it means breaking the shot, the framing and your own settle.
Mid-take, the prompter is effectively untouchable. Which is why the design assumption has to be that it cannot be looked at or touched while it matters — anything that needs a tap during a performance is a feature that belongs on the setup screen instead.
The traditional patch is hardware: a Bluetooth clicker to nudge the scroll, a foot pedal to pause it. These work, and for some setups they remain the right call — more on that below. But they treat the symptom. Most of what a solo creator needs mid-take isn't scroll adjustment at all; it's take control — start, stop, again, note that stumble — and that's a job speech does better than thumbs.
The scroll that needs no controlling
Start with the reason the usual remote exists, because voice-following removes most of it at the root.
On a constant-speed prompter, the scroll is a guess you made before the take, and the remote is how you correct the guess live — nudge it faster, pause it for a breath. The remote is compensation for a scroll that doesn't know where you are.
A voice-following prompter knows where you are: it listens as you read and keeps the highlight on the word you're saying, through pauses, ad-libs, skipped lines and backtracking. There's no speed to correct, so the single biggest job of a prompter remote simply evaporates. Pause, and the script waits — no pedal required. The full comparison of scroll modes covers this properly; the short version is that the best remote control for scrolling is a scroll that follows you.
What's left is controlling the take — and that's where the commands come in.
The take, run by voice
One press starts everything. In TellyPrompter, the mic button is the take: recording starts and the script starts following you, one decision, one press — made before your hands are in frame, so it doesn't need a command. Arming the camera is separate, decided between takes; rolling is decided at the top of one.
From your mark, the session then runs on a small vocabulary, each command prefixed with the wake word so ordinary script text never triggers anything:
"Telly, stop" ends the take — recording written, done, from wherever you're standing. Ending a take writes the file; so does disarming the camera or leaving for the library, because a perfect performance should never end with no file.
"Telly, again" starts the retake. No walk to the rig, no re-settle, no gap for the doubt to creep into — the fluffed take ends and the fresh one begins from your mark. On a batch day this single command is most of the tempo.
"Telly, marker" drops a marker at that moment and the take keeps rolling. This is the quiet workhorse: a stumble mid-take stops being a decision ("restart? push on? will I remember where that was?") and becomes an entry on a list. Afterwards, the post-take stumble report shows exactly where the rough spots were — so "was that take clean?" is answered at a glance rather than by scrubbing back through footage.
"Go back" and "go next" handle navigation, and tapping any word jumps the script to it between takes.
The wake word is deliberate friction: "stop" appears in scripts constantly, so the bare word does nothing — the command is exact, and a script that happens to contain the phrase is checked against, not obeyed.
Where a physical remote honestly wins
A voice-controlled prompter has real edges, and pretending otherwise would undo the point of this piece.
Constant mode still benefits from hardware. If you deliberately run a fixed-speed scroll — rehearsed pace, timed material — the speed-nudging job returns, and that's a genuine clicker-or-pedal job. Voice commands control takes; they aren't a speed dial.
Firefox has no voice at all. Firefox doesn't implement the Web Speech API, so neither voice-following nor voice commands exist there — the prompter itself still works, on constant and sound-level modes, and there a remote earns its place traditionally.
Recognition is the browser's. Command recognition rides the same engine as the following, so its reliability varies by platform, language and microphone, and a very loud room degrades it. The practical test is the usual one: say the commands once during setup and watch them land.
And some rooms need silence. If you're recording in a context where a spoken "Telly, stop" would sit awkwardly at the end of every take — it's easily trimmed, but it exists in the raw audio — a hardware stop, or simply ending from the device between takes, may suit you better.
None of these are reasons to buy a remote first. They're the honest boundary of not needing one.
What it adds up to
Put the pieces together and the whole session runs from your mark: press the mic, perform, mark the stumbles as they happen, say "again" until it's right, say "stop", and the file is written with a stumble report and per-take timing attached — timing that also gives you script-accurate captions as a by-product, no extra step.
There's a cost comparison buried in here, but it isn't really about the price of a clicker. It's about moving parts. A remote is one more battery, one more pairing, one more thing in your hand that isn't a gesture. A prompter that follows your voice and answers to it needs nothing in your hands at all — which, on a tool whose entire job is to let you talk to a lens like a person, is how it should have worked from the start.
Speech recognition runs on the device and the audio is never uploaded.