Talking on camera well is a learnable skill, not a personality type, and the core of it is one reframe: the lens is a person. Everything that makes someone watchable on camera — steady eye contact, natural pacing, warmth that survives compression — follows from treating recording as a conversation with one viewer rather than a performance to an empty room.
This piece covers the mechanics: where to put your eyes, how to calibrate energy for a medium that flattens it, pacing and pauses, what to do with your hands, and the warm-up that stops take one being the practice run.
The lens is a person
The strangeness of talking to camera is that there's nobody there. Every instinct you've built across a lifetime of conversation — reading a face, adjusting to nods, pausing when someone leans in — gets no input, and the common result is a voice that goes flat and formal, the way people sound leaving a voicemail.
The fix that actually works is concrete, not motivational: picture one specific person on the other side of the lens. Not "the audience" — one viewer. A friend who'd genuinely want this video, a past version of yourself who needed it, a subscriber whose comment you remember. Talk to them. The difference is audible within a sentence, because your speaking apparatus already knows how to talk to one person; it just needs to be told that's what's happening.
A useful side effect: one person changes your word choices. You stop saying "hello everyone, welcome back" — you'd never say that to a friend — and start saying the thing you'd actually say.
Eyes: where they go and why it shows
The camera is merciless about eyes. A glance that feels tiny to you — down at notes, off to a second monitor — reads on screen as distraction or evasiveness, because the viewer experiences your eye line as eye contact with them.
Three working rules:
Look at the lens, not the screen. If you record with your phone or a webcam, the temptation is to watch yourself in the preview. On camera this reads as looking slightly below the viewer's eyes, endlessly. Find the lens, and if it helps, put a small dot or arrow sticker next to it for the first few sessions.
If you read from a script, put the script as near the lens as your setup allows — and the further back the camera, the less any remaining glance shows. This is the entire logic of teleprompter placement, and if reading is part of your workflow, how to read a teleprompter without sounding wooden covers the craft of it in full.
Breaks in eye contact are allowed — if they're human ones. Nobody stares unbroken in real conversation. Looking away to think, glancing down at a prop, turning to gesture at something: these read as natural because they're motivated. It's the unmotivated flick — to notes, to the clock — that reads as reading.
Energy: the camera takes a cut
Here's the unfair physics of the medium: the camera flattens delivery. Energy that feels normal in the room arrives on screen slightly deadened — a combination of compression, small screens, and the missing presence of an actual human body. Speak at genuine conversational level and you'll often watch it back sounding tired.
The working correction is to aim a notch warmer than conversation — not performance, not presenting, just the version of you telling a friend something you're genuinely interested in. More vocal variety than feels necessary, slightly more animation in the face. On camera it lands as normal.
The calibration tool is simple and unglamorous: record thirty seconds, watch it back, adjust. Everyone's correction factor is different, and thirty honest seconds of playback teaches you yours faster than any rule. Most people need to come up a notch; a few natural over-projectors need to come down. You only find out by looking.
One caution at the other edge: pushed too far, energy tips into a presenting voice — the bright, salesy register that audiences flinch from precisely because it isn't talking, it's performing talking. The test is always the same: would you say this sentence, this way, to the one person you're picturing?
Pace, and the pause you're afraid of
Nerves speed people up. The racing read is the most common single fault in self-recorded video: no air between ideas, sentences piling into each other, the audible sound of someone trying to get this over with.
The counterintuitive craft is that pauses are a gift to the viewer. A beat of silence after a point is where the point lands. A breath before a new section tells the audience the gears are changing. On playback, pauses that felt endless in the room read as confidence — the speaker who pauses looks like they know the silence won't hurt them.
Real speech also varies its speed constantly: quick through the setup, slow into the thing that matters. That variation is most of what "engaging delivery" means at the mechanical level. If you read from a prompter, this is exactly where the tool choice bites — a fixed-speed scroll penalises every pause and every slow-down, which is why prompted reads so often sound metronomic. A prompter that follows your voice makes the pause free: stop, and the script waits.
Practical drill: mark two or three deliberate pause points in your script or notes — after the key claim, before the turn — and honour them on the take. Externalising the decision beats trying to feel your way to it under nerves.
Hands, posture, and the rest of you
The question "what do I do with my hands" has a short answer: what you do in conversation, which is gesture. Hands visible and moving naturally read as engaged; hands pinned to your sides read as a hostage video. If you gesture when you talk to friends — most people do — let it happen in frame.
Posture does quiet work. Sitting or standing tall, shoulders open, a slight lean toward the camera: together they read as interest in the viewer. A collapsed posture reads as low energy before you've said a word, and it genuinely constrains your breathing and therefore your voice.
And movement is allowed. Shifting weight, turning slightly, leaning in for the important line — small motion reads as alive. The frozen head-and-shoulders lock that nervous speakers adopt is the thing that looks strange, not the moving.
The warm-up, and why take one is a rehearsal
Voices and faces need a minute to arrive. The first read of any session is reliably the stiffest — most careful, most formal, least you — and the expensive mistake is spending it on the take that counts.
A two-minute warm-up before recording: say your first paragraph out loud, at full performance volume, off camera. Roll your shoulders. Say something ridiculous at volume to loosen the voice. Then treat the first recorded take as a rehearsal by design — press record, do the whole thing, and expect to use the second one. Sometimes the rehearsal take turns out to be the keeper, which is a pleasant surprise rather than a plan.
If nerves are the bigger obstacle — if the problem isn't technique but the feeling in your chest when the red light comes on — that deserves its own treatment, and it has one: nervous on camera? how to record anyway.
The playback habit
The single practice that improves on-camera speaking fastest is also the one people most avoid: watch your own footage. Not to cringe — to calibrate. Once a week, watch one take and ask three questions only: where did my eyes go, where did I rush, what did my energy look like from out here? Fix one thing per week.
Everyone dislikes their own recorded voice at first; it's a known artefact of hearing yourself without the resonance of your own skull, and it fades with exposure. What's on the other side of the discomfort is the thing every good on-camera speaker has: an accurate picture of what the viewer actually receives — one person, talking to them, like it's a conversation. Which is what it was always supposed to be.