AI Video Production
0/24 complete

Module 05 · AI Avatars and Presenter-Style Video

Syncing Avatar Delivery With Script Timing

20 minfocused lesson6practical steps4grounded questions5source links
Open lesson + course map

On this lesson

Course outline

An avatar feels unnatural when its sentences, pauses, expression, and edit points fight each other. Fix the script before rendering: divide it into delivery beats, preview the voice, direct each beat, and time every B-roll insert.

By the end, you will produce a synced 45–60 second avatar clip, a segmented script, and a timing log that records pronunciation, delivery direction, B-roll coverage, and any retake. Use only an avatar and voice you own or have documented permission to use, and disclose synthetic presentation when the client or publishing platform requires it.

// concept

Segment the Script Before You Render

Treat one HeyGen scene as one delivery beat. A beat expresses one idea and ends where the presenter could breathe, change energy, or hand the screen to B-roll. In AI Studio, use Add Scene → Blank and place one beat in its script box. For bulk import, the official guide currently maps each .txt line or .csv cell to a segment; confirm this before client work.

For a 45–60 second clip, start with roughly 105–140 words, then measure the preview. Use four to six beats. Do not split a subject from its verb.

Use this planning syntax outside HeyGen:

// prompt — copy me8 lines
[S1 | avatar | calm hook]
Your online order is confirmed — but do you know what happens next?

[S2 | avatar | clear explanation]
First, the seller verifies the item and prepares the parcel.

[S3 | B-roll over continuing voice]
Then the courier scans it, moves it through a sorting hub, and sends a tracking update.

Preview every scene’s audio before generating the avatar. Rewrite or split a long beat now: changing A-roll later can require another credit-consuming render. Check current credit rules on HeyGen’s official page.

// concept

Control Pauses, Pronunciation, Pace, and Direction

Use commas for short separation and periods for firmer stops. For an intentional break, place the cursor at the boundary and choose Pause; the official guide currently describes 0.5-second adjustments. Preview before submitting.

Rewrite an overloaded sentence before adding pauses. Write numbers as spoken words if the voice mishandles symbols.

For names or Roman-Urdu terms, preview the word in context. If wrong, double-click it, choose Pronunciation, and enter a phonetic spelling. Saved rules can be reused through Brand System → Brand Glossary; confirm the current path if your interface differs.

Written termSample pronunciation entryListen for
easypaisaee-zee pai-safour clear syllables, no English “easy-pay-sa” ending
Karachikuh-raa-cheestress not forced onto the last syllable
PKR 2,500two thousand five hundred Pakistani rupeesno reading of “P-K-R” as one word

These are sample spellings; voices interpret phonetics differently. Save only a tested version.

Set delivery per scene. Start at the voice’s normal Speed baseline. When available, open Voice Director from the megaphone icon and add a short instruction. Supported Avatar IV/V workflows also accept a Motion Prompt for expression, gaze, posture, or one gesture. Confirm current feature and plan availability.

// prompt — copy me2 lines
Voice Director: Calm and direct. Land the final question, then pause.
Motion Prompt: Looks at camera, then nods once with a warm expression.

Limit each scene to one gesture and one expression. If TTS stays flat, Voice Mirroring can use an authorized recorded performance; check availability and consent requirements.

// concept

Build the Timing Map Around the Voice

The approved voice preview is the timing master. Measure each scene, then assign visuals. B-roll may cover continuing voice; it must not force speech to rush.

FieldWhat to record
SceneHeyGen scene number and script beat
AudioPreview duration and intentional pause
DirectionVoice speed, Voice Director text, motion prompt
PronunciationGlossary rule used, or “none”
Visual planAvatar visible or exact B-roll insert
QC/retakeProblem heard or seen, change made, new duration

In AI Studio, select a canvas element and drag its timing markers against the script. For later CapCut assembly, copy those in/out times into the edit plan. Avoid cuts on a frozen mouth or mid-gesture frame.

// worked_example

Worked Example

This fictional sample is a 50-second vertical parcel-tracking explainer; no performance result is claimed.

BeatSample scriptAvatar timeVisual planDirection
S1“Your online order is confirmed — but what happens before it reaches your door?”0:00–0:07AvatarCalm curiosity; look at camera
S2“First, the seller checks the item, packs it securely, and prints the shipping label.”0:07–0:17AvatarClear, instructional; one open-hand gesture
S3“Next, the courier scans the parcel. [0.5s pause] That scan creates the first tracking update.”0:17–0:28Packing and scan B-roll from 0:19–0:27Slight emphasis on “first tracking update”
S4“At the sorting hub, the parcel is routed toward your city. Updates can pause between scans, so keep the tracking number.”0:28–0:42Sorting B-roll from 0:29–0:37; avatar returnsSteady pace; no gesture during warning
S5“If delivery is delayed, contact the seller or courier with the order and tracking numbers — never share your OTP.”0:42–0:53Avatar; “Never share your OTP” text at 0:47Serious; slow final clause

Sample settings and log:

// prompt — copy me6 lines
Voice speed: 1.0 sample baseline (confirm current control and preview)
Pronunciation rule: "courier" → "kur-ree-er" after preview error
S3 pause: 0.5 seconds using Pause control
S5 Voice Director: "Firm and careful; slow down after the dash."
S5 Motion Prompt: "Looks at camera with a serious expression, hands still."
Final measured duration: 53 seconds

Draft one combined S4 and S5 at a faster speed. The warning rushed, and mouth movement looked late on “tracking numbers.” The editor split S5, restored the baseline speed, inserted the dash, previewed, and rendered only that segment. The log recorded S5 v2 — pace/lip-sync fix — 11s.

// failure_cases

Failure Cases to Diagnose

6 cases to diagnose

  • The lips appear late on one phrase

    isolate that scene, remove an awkward phonetic spelling, preview the corrected audio, and re-render only the segment.

  • Every sentence has the same energy

    give each beat a distinct direction—curious hook, neutral explanation, firm warning—rather than one instruction for the whole script.

  • The avatar freezes under a B-roll cut

    move the cut so it covers a complete spoken phrase and leave the avatar hidden until the next clean scene boundary.

  • Roman-Urdu names change pronunciation between scenes

    apply one tested Brand Glossary rule to all relevant scenes and preview each occurrence in context.

  • The voice races to meet the edit

    keep the natural read, extend or replace the B-roll, and update the timing map; never time-stretch speech until it sounds artificial.

  • A gesture lands after the matching word

    shorten the scene and simplify the motion prompt to one action; do not specify a multi-step routine.

// pakistan_angle

Pakistan Angle

Roman-Urdu pronunciation varies by speaker and TTS voice. Build a glossary for names, cities, easypaisa, JazzCash, and English acronyms used in Pakistani commerce. Test on a phone speaker, not headphones alone. Keep CNICs, phone numbers, addresses, order IDs, and OTPs out of scripts and recordings.

On unstable connections or during load-shedding, finish previews, pronunciation, and timing before rendering. Regenerate only a failed scene, download approved files on a stable connection, and keep the log locally. Verify current PK payment and plan access; never base a client delivery on an unconfirmed paid feature.

// hands_on

Hands-On Exercise

6 steps

  1. Choose a rights-cleared avatar and voice, then write a 105–140 word script for a 45–60 second explainer.

  2. Split it into four to six HeyGen scenes at changes in idea, emotion, or visual coverage.

  3. Preview every scene. Add only necessary punctuation and Pause controls; create pronunciation rules for every word that fails the preview.

  4. Record the voice baseline, Voice Director instruction, motion prompt, and measured duration in the timing log.

  5. Assign exact avatar and B-roll in/out times. Render a low-risk test scene first if your current plan allows it, then generate the clip.

  6. Watch once muted for gesture and lip timing, once audio-only for pace, and once on a phone with sound and picture. Re-render only failed scenes and record each change. Done means you have a synced 45–60 second avatar clip plus a timing log that another editor could use to reproduce its pauses, pronunciations, directions, B-roll cuts, and retakes.

// completion_rubric

Completion Rubric

6 checks — tick as you verify

0/6

// sources

Sources

// check_yourself

Check yourself

4 questions · answers and options are taken word-for-word from this course

0/4
  1. 1 / 4 · diagnose

    Your work shows this failure mode: “The lips appear late on one phrase.” What does the lesson tell you to do about it?