Module 03 · AI Voice and Narration
Pacing, Emphasis, and Emotion in AI Narration
Open lesson + course map
On this lesson
Course outline
Module 1 · The AI Video Stack in 2026
Module 2 · Scripting for AI Video
Module 3 · AI Voice and Narration
Module 4 · Generative Video With Veo
Module 5 · AI Avatars and Presenter-Style Video
Module 6 · Editing and Assembly in CapCut
Module 7 · Video Service Packaging
AI narration direction begins in the text box. Punctuation controls phrasing, sentence length shapes pace, and paragraph breaks separate ideas. On compatible ElevenLabs models, audio tags and voice settings add another layer of direction—but they cannot rescue a script with no spoken rhythm.
By the end, you will have one delivery-marked script, three controlled narration variants, and a listening table that shows which take handles pace, emphasis, and emotion most effectively.
// concept
Mark the Delivery in the Script
Read the paragraph aloud first. Underline one to three meaning-carrying words, mark needed pauses, and label each emotional beat. Then convert those marks into TTS-friendly text.
| Text control | Likely delivery effect | Good use | Risk |
|---|---|---|---|
| Full stop | Clear stop and reset | Separate one claim from its consequence | Too many create a choppy read |
| Comma | Light phrase break | Short lists and introductory phrases | A comma is not a reliable long pause |
Em dash (—) | Interruption or deliberate turn | Set up a contrast | Repeated dashes sound theatrical |
Ellipsis (…) | Hesitation or suspended thought | One reflective beat | Can become melodramatic or inconsistent |
| Question mark | Rising or questioning contour | A real question in the script | Rhetorical questions can sound exaggerated |
| Exclamation mark | Stronger energy | A rare reveal or call | Several in a row sound like shouting |
| CAPITALS | Possible emphasis in Eleven v3 | One contrast word: “not faster—CLEARER” | Whole sentences reduce control |
| Paragraph break | New meaning or emotional beat | Split problem, reveal, and action | A break may produce too much silence or a breath artifact |
Sentence length is a dependable pacing tool. Use a short sentence for impact, then a longer one for context. Equal-length sentences tend to produce a repetitive cadence.
Eleven v3 supports bracketed tags such as [thoughtful], [curious], and [excited]; other models may speak a direction aloud. Confirm the selected model's current prompting guide.
// concept
Direct One Emotional Beat at a Time
In ElevenLabs Text to Speech, paste the text, select the voice and model, open voice settings, and generate. Confirm the current interface and plan access in the official guide. Keep voice and model fixed during a text test.
For Eleven v3, Stability currently uses Creative, Natural, and Robust. Natural is a comparison baseline; Creative is more expressive but less predictable, while Robust is steadier and less responsive. V3 does not currently expose Speed or Similarity and does not support SSML <break> tags.
For models with numeric controls, test Speed 1.0, Stability 50, Similarity 75, and Style 0. These are official starting points, not guarantees. Lower stability can widen emotion; high stability can flatten it. High Similarity may reproduce flaws in a poor reference. Style can add instability, so test it only with a reason.
| Beat | Text direction | Settings test | What to listen for |
|---|---|---|---|
| Calm explanation | Medium sentences; full stops; no exclamation marks | Baseline or v3 Natural | Even pace without sounding detached |
| Important contrast | Short setup, em dash, one key word in capitals | Keep settings unchanged first | Emphasis lands on the intended word |
| Concern or warning | Separate paragraph; slower phrasing; [concerned] only if supported | Compare Natural with Robust | Serious, not melodramatic |
| Energetic payoff | Shorter sentences; one exclamation mark; [excited] only if supported | Compare Natural with Creative | Energy without clipped words |
Split narration at emotional turns. When pronunciation or accent drifts, ElevenLabs suggests sections under roughly 800–900 characters as a troubleshooting target. Name each export, for example 03_warning_take-b.mp3, before assembly.
// concept
Score the Audio, Not the Markup
Make three takes with one voice and model. A is unmarked. B changes punctuation, sentence length, and paragraphs. C adds supported tags or one documented settings change.
Measure duration in the player and count only audible target hits.
| Measure | Take A | Take B | Take C | Pass rule |
|---|---|---|---|---|
| Duration in seconds | Fits the edit slot | |||
| Intended pauses heard | At least 2 of 3 | |||
| Emphasis targets heard | At least 2 of 3 | |||
| Emotion fit (1–5) | 4 or 5 | |||
| Mispronounced words | 0 | |||
| Clear on phone speaker | Yes |
Use headphones for breaths, clicks, and emotion, then an ordinary phone speaker for rushed consonants, weak emphasis, and lost pauses.
// worked_example
Worked Example
This hypothetical sample for a Lahore online stationery seller needs reassurance, then a firm warning, then an energetic action. No result is claimed.
Take A — flat control
Your parcel is packed but one detail can still delay delivery. Check the phone number before you confirm the order. A correct number helps the rider reach you. Confirm it now and your order is ready for dispatch.Generate with one Eleven v3 voice and Stability: Natural. Save take-a-control.mp3. Diagnose whether the warning and action receive the same weight as the setup.
Take B — punctuation and paragraph direction
Your parcel is packed. But one detail can still delay delivery — your phone number.
Before you confirm the order, check it once more. A correct number helps the rider reach you.
Check. CONFIRM. Then your order is ready for dispatch.Keep voice, model, and Stability fixed. Save take-b-text-directed.mp3. Full stops shorten the close; paragraph breaks separate the three beats.
Take C — emotional direction for Eleven v3
[reassuring] Your parcel is packed. [concerned] But one detail can still delay delivery — your phone number.
[thoughtful] Before you confirm the order, check it once more. A correct number helps the rider reach you.
[confident] Check. CONFIRM. Then your order is ready for dispatch.Keep Stability: Natural and save take-c-tag-directed.mp3. Tags depend on the voice. If [concerned] overacts, remove it before testing Robust on only that beat; record the change.
Do not add exclamation marks, capitals, tags, and a settings change together. If “phone number” lacks emphasis, regenerate only: But one detail can still delay delivery. Your PHONE number.
// failure_cases
Failure Cases to Diagnose
7 cases to diagnose
The voice speaks every sentence at one speed
sentence lengths are too uniform. Break the key line into a short sentence and compare it without touching settings.
Pauses become awkward gaps
paragraph breaks, ellipses, and dashes are stacked together. Keep one pause device and regenerate the sentence.
The emotion tag has little effect
test a compatible voice or compare Natural with Creative; tags do not work equally across voices.
The voice reads a direction aloud
the model does not support the tag syntax used. Remove it and use phrasing, punctuation, or a supported model.
Delivery changes unpredictably between takes
too many variables changed, or stability is low. Return to the control settings and change one item only.
A Roman-Urdu word changes accent
isolate and respell it, then test a shorter segment.
Headphones sound fine but the phone sounds rushed
slow or rewrite the dense sentence; do not simply raise volume, which will not restore articulation.
// pakistan_angle
Pakistan Angle
Test words your Pakistani audience will hear: easypaisa, Karachi, bijli, rupay, neighbourhood names, and English–Roman-Urdu switches. An English-trained voice may carry its accent into Roman Urdu, while long mixed-language text may drift. Use short beats for respelling. Clone only with documented permission, and remove phone, address, CNIC, and order data from scripts.
Under mobile-data and load-shedding constraints, prepare scripts and filenames offline, then generate one short beat before the full track. Download accepted takes and test them through the compression used in your publishing workflow. Before committing USD costs in PKR, check current official pricing and payment options; no plan or Pakistani payment rail is assumed here.
// hands_on
Hands-On Exercise
7 steps
Choose a 60–90-word paragraph from a script you are authorized to use.
Mark three emphasis words, three intended pauses, and the emotional goal of each paragraph.
Generate Take A with unmarked text. Record the voice, model, and every visible setting.
Generate Take B after changing only punctuation, sentence length, and paragraph breaks.
Generate Take C by adding compatible direction tags or changing one setting for one emotional beat.
Fill the review table using player time, audible counts, a 1–5 emotion score, pronunciation errors, and the phone-speaker check.
Regenerate only the lowest-scoring sentence, then write one line explaining why the final take won. Done means you have the marked script,
take-a,take-b, andtake-caudio files, plus a completed review sheet with a defensible selection.
// completion_rubric
Completion Rubric
6 checks — tick as you verify
// sources
Sources
3 official sources — check every claim yourself
// check_yourself
Check yourself
4 questions · answers and options are taken word-for-word from this course
1 / 4 · diagnose
Your work shows this failure mode: “Pauses become awkward gaps.” The lesson describes it like this: “Paragraph breaks, ellipses, and dashes are stacked together.” What does the lesson tell you to do about it?