Skip to content

DOUBLE CREDITS on your first month, or your first 3 months on annual

Claim now
Kyndrify
Watch

A talking-head video made from three inputs

This is the whole clip, sound on, nothing cut out of it. A presenter was picked, a voice that was already saved was picked, three lines were typed, and the render matched her mouth to the voice. The script is printed under the video so you can check every word against what you hear.

Eleven and a half seconds, one take. A presenter, a saved voice, three lines.

One shot, start to finish. No cut, no music, no camera move.

Part of Showcase.

Captions are on from the first press. The three lines are written out below.

What does an AI talking-head video actually look like?

A talking-head video is one person speaking to camera, rendered instead of filmed. This one runs about eleven and a half seconds. A presenter was picked, a saved voice was picked, three lines were typed, and the render matched her mouth to the voice. Nobody in it is a real person.

  • Who it is for

    Anyone who has read that you can make a video without filming one and wants to see a whole one before they believe it.

  • What you need

    To make one yourself: a presenter, a voice, and something to say. A free Workspace gets you all three. To watch this one: nothing.

  • What it costs

    A video like this is charged by the second. On a paid plan a talking-head render is 152 credits a second, with a floor of 120 credits a render, and the Free plan starts you with 8,400 credits that do not expire.

  • What it does not do

    It does not cut between shots, add music, or move the camera. It is one person, one take, one angle. And nobody in it is a real person.

The door this went through is talking-head video, and pricing shows what a plan's credits buy.

Every word she says

Three lines. They are what was typed before anything was rendered, and they are what you hear, in that order.

There was never a step where anything decided for itself what she would say. That is why the words under the video, the captions on it and the script that made it are the same three sentences.

The three typed lines, with the time each one lands
AtThe line
0:00The lineHey, welcome! I'm really glad you're here.
0:03The lineThis is where I share what I'm learning, in short, simple videos.
0:07The lineIf you're new, start from the top and follow along. See you inside.

One person, one take, no edit

A woman in a grey hoodie stands in an open-plan office and records a welcome message for her channel. She talks straight down the lens. Behind her the room is soft and out of focus, the way a room is when a phone is close to your face and the far wall is not. The frame is tall, the shape a phone holds.

There is no cut in it. That matters more than it sounds, because a cut is where a flaw goes to hide. If a mouth drifts out of time with the words, an editor's instinct is to cut away for half a second and come back. Nothing is cut away here. You get the mouth for the whole clip and you can judge it.

The last thing to notice is the ordinary stuff: a blink, a small shift of weight, the hands staying where hands stay when somebody is talking and not performing. None of that was directed. It came out of a short piece of base footage of this character moving, which is the part of the job a person usually spends a morning on.

The voice came first, and the picture was matched to it

That order is the whole technique, so it goes first.

  1. A presenter

    Not a person. She is an Avatar, a character with no real human behind her, which is why there is no consent record on this page and no likeness to protect. A real person's face is a Digital Twin, and that is a different thing with a consent ceremony in front of it.

  2. A voice

    A voice was saved in the Workspace before any of this started, and then it was picked from a list, the way you would pick a font. Saving your own voice takes about ten seconds of clear audio. Or you skip that and pick one of the Featured voices, which are ready to use on any plan.

  3. A script

    The three lines above, typed. Not a prompt, not a brief, not a description of what she should say. The exact words.

  4. Then the render

    The three lines became speech in the saved voice. A short piece of base footage of this character moving was extended to the length of that speech. And the mouth was matched to the audio, frame by frame.

One thing we did not do: we did not put the audio on afterwards. The sound you hear came back inside the video file, from the render that made the picture. Laying a separate audio track over a finished video is how a clip ends up a few tens of milliseconds out of time, which you do not notice in the first second and cannot unhear in the tenth. So we do not do it, anywhere, to any file. If a render cannot return its own sound, we would rather it fail than hand you something we glued together.

Priced by the second, before it runs

A video like this is charged by the second of finished video, and the price is on the screen before anything starts. On a paid plan a talking-head render is 152 credits a second, with a floor of 120 credits a render, so a very short clip does not cost a very short clip's worth of credits.

The Free plan starts you with 8,400 credits, once, and they do not expire. Two limits are worth knowing before you plan around them: on the Free plan one talking-head video runs to one minute and renders at 720p, and 1080p is on the Plus plan and above.

See what a plan's credits buy

Plans, credits and what each one covers in video, voice and images.

The parts we are not claiming

Nobody in this video is a real person, and the voice is nobody's real voice. She is an Avatar and the voice was made, not recorded from a human being. We are saying that out loud on the page rather than in a footnote, because a video of a person who does not exist should say so.

It is not a customer's video. We made it, to show what the tool does. It is one shot: no edit, no music, no b-roll, no graphics and no camera move. Those exist as separate tools in the Studio, and none of them were used here.

And it is not a promise about how yours will look. Your presenter, your voice and your words are different from these, and so is the result. This is one honest example, not a guarantee.

AI talking head
One person speaking to camera, rendered rather than filmed. The picture comes from a still or from a character's base footage, the words come from a script, and the mouth is matched to the narration.

The transcript

The same words the caption file carries.

Read the transcript

0:00Hey, welcome! I'm really glad you're here.

0:03This is where I share what I'm learning, in short, simple videos.

0:07If you're new, start from the top and follow along. See you inside.

FAQ

Questions about this clip

Is that a real person?
No. She is an Avatar, a character that was generated, with no real human behind her. When a real person's face is used it is a Digital Twin instead, and that route needs a recorded consent step, checked against the photo, before anything renders. Two different things, on purpose.
Is that a real voice?
No. It is a saved voice, and it was not taken from a person. You can save your own voice from about ten seconds of clear audio, which is on every plan, Free included, or use one of the Featured voices on any plan.
Was the sound added afterwards?
No. It came back inside the video file, from the render that made the picture. We never lay a separate audio track onto a finished video, because that is how a clip ends up slightly out of time with its own mouth.
How long did the render take?
We are not going to give you a number, because we have not measured one we would stand behind. What we can tell you is what it looks like from your side: you press render, the video shows up in your Workspace when it is done, and you can close the tab and come back to it.
Can I make one of myself instead?
Yes. That is a Digital Twin rather than an Avatar: you give it a photo and record a short consent clip, and every free Workspace gets one Starter Twin for life. Saving your own voice is on every plan, Free included, and rendering at 1080p is on the Plus plan and above.

Your turn to make one

Pick a presenter, pick a voice, type what you want said. No card needed, and the Free plan starts with 8,400 credits that do not expire.