Module 05 · AI Avatars and Presenter-Style Video
Syncing Avatar Delivery With Script Timing
Open lesson + course map
On this lesson
Course outline
Module 1 · The AI Video Stack in 2026
Module 2 · Scripting for AI Video
Module 3 · AI Voice and Narration
Module 4 · Generative Video With Veo
Module 5 · AI Avatars and Presenter-Style Video
Module 6 · Editing and Assembly in CapCut
Module 7 · Video Service Packaging
An avatar feels unnatural when its sentences, pauses, expression, and edit points fight each other. Fix the script before rendering: divide it into delivery beats, preview the voice, direct each beat, and time every B-roll insert.
By the end, you will produce a synced 45–60 second avatar clip, a segmented script, and a timing log that records pronunciation, delivery direction, B-roll coverage, and any retake. Use only an avatar and voice you own or have documented permission to use, and disclose synthetic presentation when the client or publishing platform requires it.
// concept
Segment the Script Before You Render
Treat one HeyGen scene as one delivery beat. A beat expresses one idea and ends where the presenter could breathe, change energy, or hand the screen to B-roll. In AI Studio, use Add Scene → Blank and place one beat in its script box. For bulk import, the official guide currently maps each .txt line or .csv cell to a segment; confirm this before client work.
For a 45–60 second clip, start with roughly 105–140 words, then measure the preview. Use four to six beats. Do not split a subject from its verb.
Use this planning syntax outside HeyGen:
[S1 | avatar | calm hook]
Your online order is confirmed — but do you know what happens next?
[S2 | avatar | clear explanation]
First, the seller verifies the item and prepares the parcel.
[S3 | B-roll over continuing voice]
Then the courier scans it, moves it through a sorting hub, and sends a tracking update.Preview every scene’s audio before generating the avatar. Rewrite or split a long beat now: changing A-roll later can require another credit-consuming render. Check current credit rules on HeyGen’s official page.
// concept
Control Pauses, Pronunciation, Pace, and Direction
Use commas for short separation and periods for firmer stops. For an intentional break, place the cursor at the boundary and choose Pause; the official guide currently describes 0.5-second adjustments. Preview before submitting.
Rewrite an overloaded sentence before adding pauses. Write numbers as spoken words if the voice mishandles symbols.
For names or Roman-Urdu terms, preview the word in context. If wrong, double-click it, choose Pronunciation, and enter a phonetic spelling. Saved rules can be reused through Brand System → Brand Glossary; confirm the current path if your interface differs.
| Written term | Sample pronunciation entry | Listen for |
|---|---|---|
| easypaisa | ee-zee pai-sa | four clear syllables, no English “easy-pay-sa” ending |
| Karachi | kuh-raa-chee | stress not forced onto the last syllable |
| PKR 2,500 | two thousand five hundred Pakistani rupees | no reading of “P-K-R” as one word |
These are sample spellings; voices interpret phonetics differently. Save only a tested version.
Set delivery per scene. Start at the voice’s normal Speed baseline. When available, open Voice Director from the megaphone icon and add a short instruction. Supported Avatar IV/V workflows also accept a Motion Prompt for expression, gaze, posture, or one gesture. Confirm current feature and plan availability.
Voice Director: Calm and direct. Land the final question, then pause.
Motion Prompt: Looks at camera, then nods once with a warm expression.Limit each scene to one gesture and one expression. If TTS stays flat, Voice Mirroring can use an authorized recorded performance; check availability and consent requirements.
// concept
Build the Timing Map Around the Voice
The approved voice preview is the timing master. Measure each scene, then assign visuals. B-roll may cover continuing voice; it must not force speech to rush.
| Field | What to record |
|---|---|
| Scene | HeyGen scene number and script beat |
| Audio | Preview duration and intentional pause |
| Direction | Voice speed, Voice Director text, motion prompt |
| Pronunciation | Glossary rule used, or “none” |
| Visual plan | Avatar visible or exact B-roll insert |
| QC/retake | Problem heard or seen, change made, new duration |
In AI Studio, select a canvas element and drag its timing markers against the script. For later CapCut assembly, copy those in/out times into the edit plan. Avoid cuts on a frozen mouth or mid-gesture frame.
// worked_example
Worked Example
This fictional sample is a 50-second vertical parcel-tracking explainer; no performance result is claimed.
| Beat | Sample script | Avatar time | Visual plan | Direction |
|---|---|---|---|---|
| S1 | “Your online order is confirmed — but what happens before it reaches your door?” | 0:00–0:07 | Avatar | Calm curiosity; look at camera |
| S2 | “First, the seller checks the item, packs it securely, and prints the shipping label.” | 0:07–0:17 | Avatar | Clear, instructional; one open-hand gesture |
| S3 | “Next, the courier scans the parcel. [0.5s pause] That scan creates the first tracking update.” | 0:17–0:28 | Packing and scan B-roll from 0:19–0:27 | Slight emphasis on “first tracking update” |
| S4 | “At the sorting hub, the parcel is routed toward your city. Updates can pause between scans, so keep the tracking number.” | 0:28–0:42 | Sorting B-roll from 0:29–0:37; avatar returns | Steady pace; no gesture during warning |
| S5 | “If delivery is delayed, contact the seller or courier with the order and tracking numbers — never share your OTP.” | 0:42–0:53 | Avatar; “Never share your OTP” text at 0:47 | Serious; slow final clause |
Sample settings and log:
Voice speed: 1.0 sample baseline (confirm current control and preview)
Pronunciation rule: "courier" → "kur-ree-er" after preview error
S3 pause: 0.5 seconds using Pause control
S5 Voice Director: "Firm and careful; slow down after the dash."
S5 Motion Prompt: "Looks at camera with a serious expression, hands still."
Final measured duration: 53 secondsDraft one combined S4 and S5 at a faster speed. The warning rushed, and mouth movement looked late on “tracking numbers.” The editor split S5, restored the baseline speed, inserted the dash, previewed, and rendered only that segment. The log recorded S5 v2 — pace/lip-sync fix — 11s.
// failure_cases
Failure Cases to Diagnose
6 cases to diagnose
The lips appear late on one phrase
isolate that scene, remove an awkward phonetic spelling, preview the corrected audio, and re-render only the segment.
Every sentence has the same energy
give each beat a distinct direction—curious hook, neutral explanation, firm warning—rather than one instruction for the whole script.
The avatar freezes under a B-roll cut
move the cut so it covers a complete spoken phrase and leave the avatar hidden until the next clean scene boundary.
Roman-Urdu names change pronunciation between scenes
apply one tested Brand Glossary rule to all relevant scenes and preview each occurrence in context.
The voice races to meet the edit
keep the natural read, extend or replace the B-roll, and update the timing map; never time-stretch speech until it sounds artificial.
A gesture lands after the matching word
shorten the scene and simplify the motion prompt to one action; do not specify a multi-step routine.
// pakistan_angle
Pakistan Angle
Roman-Urdu pronunciation varies by speaker and TTS voice. Build a glossary for names, cities, easypaisa, JazzCash, and English acronyms used in Pakistani commerce. Test on a phone speaker, not headphones alone. Keep CNICs, phone numbers, addresses, order IDs, and OTPs out of scripts and recordings.
On unstable connections or during load-shedding, finish previews, pronunciation, and timing before rendering. Regenerate only a failed scene, download approved files on a stable connection, and keep the log locally. Verify current PK payment and plan access; never base a client delivery on an unconfirmed paid feature.
// hands_on
Hands-On Exercise
6 steps
Choose a rights-cleared avatar and voice, then write a 105–140 word script for a 45–60 second explainer.
Split it into four to six HeyGen scenes at changes in idea, emotion, or visual coverage.
Preview every scene. Add only necessary punctuation and Pause controls; create pronunciation rules for every word that fails the preview.
Record the voice baseline, Voice Director instruction, motion prompt, and measured duration in the timing log.
Assign exact avatar and B-roll in/out times. Render a low-risk test scene first if your current plan allows it, then generate the clip.
Watch once muted for gesture and lip timing, once audio-only for pace, and once on a phone with sound and picture. Re-render only failed scenes and record each change. Done means you have a synced 45–60 second avatar clip plus a timing log that another editor could use to reproduce its pauses, pronunciations, directions, B-roll cuts, and retakes.
// completion_rubric
Completion Rubric
6 checks — tick as you verify
// sources
Sources
5 official sources — check every claim yourself
// check_yourself
Check yourself
4 questions · answers and options are taken word-for-word from this course
1 / 4 · diagnose
Your work shows this failure mode: “The lips appear late on one phrase.” What does the lesson tell you to do about it?