Some of the faces pulling millions of views were built by four tools handed off in a line. Here is the exact pipeline, what each stage costs you in time, and the one rule that keeps it honest. A first pass takes about twenty minutes.
You have seen the clips. A calm, well-lit person talks straight to camera, the voice is smooth, the footage looks like a real studio, and the clip has millions of views. None of it is a real person.
The face was generated. The voice was generated. The footage was generated. What looks like one creator with a studio and a camera is four separate tools passing work down a line.
Here is the whole line before we go deep. You generate a character, give it a voice, animate it into video, then stitch the clips together. That is the entire method, and by the end of this you can run a version of it for your own business content.
What You Are Actually Building
A synthetic presenter. One consistent face and voice that you own, reading any script you write, without booking a studio or standing in front of a camera yourself.
Businesses already do this with hired spokespeople and stock actors. The difference is that this presenter never reschedules, never renegotiates a day rate, and says exactly what you scripted on the first take. For an owner who wants a steady stream of short video but hates being on camera, that is the appeal.
It is also the part that gets abused, which is why the honesty section near the end is not optional reading.
The Four Tools
Four jobs, four tools. You can swap the specific tool at each stage, but the roles stay the same.
| Stage | Tool | Its job | What to know |
|---|---|---|---|
| Face | Higgsfield or Midjourney | Generates the still portrait of the presenter | Both are paid subscriptions with limited trials. Midjourney runs in Discord or its own web app. |
| Voice | ElevenLabs | Turns your written script into speech | A free tier covers short clips. Longer or higher-quality output needs a paid plan. Cloning a real voice needs that person's consent. |
| Motion | VEO or Runway | Animates the still into lip-synced video | Credit-based. A few seconds of clean footage often takes several attempts. |
| Edit | CapCut | Cuts the clips, adds captions, exports the final video | Free for the basics. This is the one stage with no AI in it and no AI risk. |
Stage One: Generate The Face
Open Higgsfield or Midjourney and describe the person you want in plain words. You get back a photoreal still.
The hard part is not making one good image. It is keeping the same face across every shot so it reads as one person and not a slideshow of strangers. Lock the description down and reuse it word for word each time. If the tool supports a character reference or a seed, use it.
Copy The Character Prompt
Copy this.
Photorealistic headshot of a [AGE]-year-old [PERSONA, e.g. approachable operations manager], [HAIR AND WARDROBE, e.g. short dark hair, plain charcoal sweater], seated at a plain desk, soft window light from the left, neutral grey background, 50mm lens, sharp focus on the eyes, looking directly into the camera, identical face and hair in every frame
Your first real run: generate four versions of the same person and pick the one whose eyes look most natural. The eyes are where a fake face gives itself away first.
The mistake that makes it fail: changing the wording between generations. Swap "charcoal sweater" for "grey jumper" and you get a slightly different person. Paste the same block every time.
Stage Two: Give It A Voice
Take your script into ElevenLabs, pick or build a voice, and it reads the lines back as audio.
Most people rush the script and it shows. A generated voice reading weak writing still sounds like weak writing, just cleaner. Write the spoken lines properly first. This is a good job to hand to Claude or Codex, since a language model will draft tight, plain lines faster than you will from a blank page.
Copy The Script Prompt
Copy this.
Write a [30]-second spoken script for a short video. Topic: [YOUR TOPIC]. Audience: [WHO IT IS FOR]. Rules: one idea per line, plain words, no hype language. The first line is the single most useful sentence in the script. The last line is one clear next step for the viewer. Keep it under [90] words so it reads in [30] seconds at a calm pace. Return only the spoken lines, nothing else.
If the read comes back too fast or flat, tell the tool exactly what to fix. Say: slow the pace and add a short pause after the first line. Plain correction beats re-generating blind.
Stage Three: Animate It
Feed the still and the audio into VEO or Runway. It moves the face and syncs the mouth to the voice, so you get a talking clip instead of a photo.
This is the stage that eats the most credits and patience. Short shots animate cleanly. Long ones drift, and fast movement warps the face. Keep each clip to a few seconds and generate more of them rather than asking for one long take.
If you want the presenter identical across a whole video, animate every clip from the same locked still. If you only need one short clip, one generation is fine and you can skip the consistency worry entirely.
Stage Four: Stitch It Together
Bring the clips into CapCut. Trim them, order them, add captions, and export.
This is ordinary video editing, the same as cutting any short. Burned-in captions matter here because most people watch muted, and captions also carry the video when the audio is the weakest link. Nothing on this stage can be faked or go wrong in the way the earlier stages can.
Done. Four tools, one presenter, one finished clip.
Where This Actually Earns Its Keep
The viral version is the flashy demo. The useful version is quieter and lives inside a business.
- Repetitive explainer videos. The same onboarding walkthrough or policy explainer, refreshed whenever the details change, without re-filming anything.
- Founders who will not go on camera. Plenty of good operators freeze on video. A disclosed AI presenter gets the content made instead of stalling for another quarter.
- Multi-language content. Rewrite the script, regenerate the voice, keep the same face. One presenter across several markets.
- Testing a format before you commit. Try a video series with a synthetic host first, and only put a real person on camera once you know the format works.
Notice that none of these depend on tricking anyone. The moment the value depends on the viewer believing a real human is speaking, you have walked into the next section.
Honest: The Line You Should Not Cross
This workflow is built to make fake footage look real. That is the whole point of it, and it is exactly why you have to hold a hard line on how you use it.
- Disclose it. Put "AI presenter" in the caption or on screen. The instant a viewer thinks a real person is personally vouching for something, honest content has become deception.
- Never rebuild a real person's face or voice without written consent. Cloning someone's likeness to make them say things is impersonation, and in a growing number of places it is now illegal. Use a generated face that belongs to no one.
- Do not fake proof. A synthetic presenter reading a scripted testimonial is a fabricated testimonial. The same goes for reviews, client results, and credentials. Say only what is true.
- It still looks slightly off. Hands, teeth, and fast motion are where the illusion cracks. Keep shots short and mostly still, and it holds up better.
- It is not free and not instant. Between subscriptions and re-rolls, a polished thirty-second clip is an afternoon of work and a paid plan or two, not a single button.
Close
If you do one thing this week, generate one face and one fifteen-second clip about a topic you already know cold. Do not publish it.
Watch it back once. Decide whether the finished quality is worth the afternoon a real video would take, and whether you would put your name next to it with the AI label attached. If the answer is no, you spent twenty minutes and learned where the ceiling sits today. If it is yes, you now own a presenter that shows up every time you have something to say.