TellyPrompter blog

Recording videos alone: the one-person studio workflow

Performance · Published 12 August 2026 · 4 min read

A one-person studio workflow for recording videos alone: batching scripts, running takes hands-free, and ending with usable files.

Recording videos alone means holding three jobs at once: you are the presenter, the camera operator and the person who has to make sense of the files afterwards. The workflow that makes it sustainable is the one that keeps those jobs from interrupting each other — everything operable is decided before the take, the take itself needs nothing but your voice, and the sorting-out happens after, with the machine having taken its own notes.

This is the workflow guide. The craft of solo filming — framing, light, sound, eye line — has its own article; this one is about the system around it.

Before: load the day, not the take

The costliest habit in solo recording is setting up per video: write a script, rig the camera, record, tear down, repeat next week. The setup is the expensive part, so make it pay for several videos at once. Batch recording covers the case for shooting a week in a session; the short version is that your third take of the day is better than your first, and the room is already right.

Practically, that means arriving at the shoot with every script already written and loaded. TellyPrompter's library holds them with one tag each — a "This week" tag is enough — so the gap between videos is picking the next script, not hunting for it. Two preparation tools earn their place here: the measured read-time tells you what each script actually costs in minutes before you commit the afternoon to it, and rehearsal mode lets you do the one read-aloud pass that catches the phrases your mouth refuses, while fixing them is still a text edit rather than a retake.

During: a take is one press, and your voice is the operator

Mid-take is the worst possible moment to operate anything — your eyes are on the lens, your hands are in frame, and the phone is out of reach on the rig. So the session has to run on the smallest possible set of controls.

In TellyPrompter, starting a take is one press: the mic button records and sets the script following you in the same act, because they are the same decision. Arming the camera is separate — that is decided between takes, not at the top of one. From there, the session is voice-operated. "Telly, stop" ends the take. "Telly, again" bins the attempt and rolls the next one without you leaving position — which changes the economics of the retake, because the cost of "let me just do that once more" drops to two words. And when you stumble but want to keep going, "Telly, marker" drops a stumble marker at that point in the take and moves on, so you do not have to choose mid-performance between stopping and trusting your memory.

The honest limits: voice commands and voice-following ride on the browser's speech recognition, which Firefox does not provide — everything else works there, with constant or sound-level scrolling. And recognition quality is the browser's, not ours; it varies by platform and microphone.

One habit completes the during-phase: end every take deliberately, but know that the app is defensive about your files. Ending a take, disarming the camera, and leaving for the library all write the recording — there is no path where a finished performance ends with no file.

After: the take took its own notes

The reason solo review is miserable is that the only record of what happened is your memory plus an unlabelled video file. This is where the voice-following engine quietly does its second job: because it knows which word you were on at every moment, each take carries its own timing data.

So after a session, the post-take stumble report shows you where the markers fell and where the read broke — you review the four flagged moments, not the full nine minutes. Per-take timing shows what each attempt actually ran, which settles "was take two faster?" with a number instead of a feeling. And when a take is the keeper, the same timing exports SRT or VTT captions built from your script — script-accurate, no transcription pass, no mis-heard words to correct — ready for the platforms that want caption files. There is a full guide to script-timed captions if that is the part you came for.

Export is free on every tier, for scripts as well as recordings — the way out is not a paid feature.

The shape of a good solo session

Put together, a session looks like this. The night before: scripts written for the mouth, read aloud once, loaded and tagged, read-times checked against the time you actually have. On the day: room rigged once, camera armed, then script after script — one press to roll, voice to stop, retake or mark, no touching the rig between videos beyond picking the next script. Afterwards: stumble reports to choose keepers, captions exported from the winners, and the rig left standing if you can spare the corner, because a standing rig is the difference between "I could record today" and "I would have to set everything up".

None of this makes you a crew. It makes the crew unnecessary — which was always the point of recording alone. For the wider set of hands-free controls, and for what makes the talking head format work once the workflow is solved, carry on there.

Blog Home Terms Privacy Refunds Accessibility Open the app