AI Video Production
0/24 complete

Module 03 · AI Voice and Narration

Pacing, Emphasis, and Emotion in AI Narration

20 minfocused lesson7practical steps4grounded questions3source links
Open lesson + course map

On this lesson

Course outline

AI narration direction begins in the text box. Punctuation controls phrasing, sentence length shapes pace, and paragraph breaks separate ideas. On compatible ElevenLabs models, audio tags and voice settings add another layer of direction—but they cannot rescue a script with no spoken rhythm.

By the end, you will have one delivery-marked script, three controlled narration variants, and a listening table that shows which take handles pace, emphasis, and emotion most effectively.

// concept

Mark the Delivery in the Script

Read the paragraph aloud first. Underline one to three meaning-carrying words, mark needed pauses, and label each emotional beat. Then convert those marks into TTS-friendly text.

Text controlLikely delivery effectGood useRisk
Full stopClear stop and resetSeparate one claim from its consequenceToo many create a choppy read
CommaLight phrase breakShort lists and introductory phrasesA comma is not a reliable long pause
Em dash ()Interruption or deliberate turnSet up a contrastRepeated dashes sound theatrical
Ellipsis ()Hesitation or suspended thoughtOne reflective beatCan become melodramatic or inconsistent
Question markRising or questioning contourA real question in the scriptRhetorical questions can sound exaggerated
Exclamation markStronger energyA rare reveal or callSeveral in a row sound like shouting
CAPITALSPossible emphasis in Eleven v3One contrast word: “not faster—CLEARER”Whole sentences reduce control
Paragraph breakNew meaning or emotional beatSplit problem, reveal, and actionA break may produce too much silence or a breath artifact

Sentence length is a dependable pacing tool. Use a short sentence for impact, then a longer one for context. Equal-length sentences tend to produce a repetitive cadence.

Eleven v3 supports bracketed tags such as [thoughtful], [curious], and [excited]; other models may speak a direction aloud. Confirm the selected model's current prompting guide.

// concept

Direct One Emotional Beat at a Time

In ElevenLabs Text to Speech, paste the text, select the voice and model, open voice settings, and generate. Confirm the current interface and plan access in the official guide. Keep voice and model fixed during a text test.

For Eleven v3, Stability currently uses Creative, Natural, and Robust. Natural is a comparison baseline; Creative is more expressive but less predictable, while Robust is steadier and less responsive. V3 does not currently expose Speed or Similarity and does not support SSML <break> tags.

For models with numeric controls, test Speed 1.0, Stability 50, Similarity 75, and Style 0. These are official starting points, not guarantees. Lower stability can widen emotion; high stability can flatten it. High Similarity may reproduce flaws in a poor reference. Style can add instability, so test it only with a reason.

BeatText directionSettings testWhat to listen for
Calm explanationMedium sentences; full stops; no exclamation marksBaseline or v3 NaturalEven pace without sounding detached
Important contrastShort setup, em dash, one key word in capitalsKeep settings unchanged firstEmphasis lands on the intended word
Concern or warningSeparate paragraph; slower phrasing; [concerned] only if supportedCompare Natural with RobustSerious, not melodramatic
Energetic payoffShorter sentences; one exclamation mark; [excited] only if supportedCompare Natural with CreativeEnergy without clipped words

Split narration at emotional turns. When pronunciation or accent drifts, ElevenLabs suggests sections under roughly 800–900 characters as a troubleshooting target. Name each export, for example 03_warning_take-b.mp3, before assembly.

// concept

Score the Audio, Not the Markup

Make three takes with one voice and model. A is unmarked. B changes punctuation, sentence length, and paragraphs. C adds supported tags or one documented settings change.

Measure duration in the player and count only audible target hits.

MeasureTake ATake BTake CPass rule
Duration in secondsFits the edit slot
Intended pauses heardAt least 2 of 3
Emphasis targets heardAt least 2 of 3
Emotion fit (1–5)4 or 5
Mispronounced words0
Clear on phone speakerYes

Use headphones for breaths, clicks, and emotion, then an ordinary phone speaker for rushed consonants, weak emphasis, and lost pauses.

// worked_example

Worked Example

This hypothetical sample for a Lahore online stationery seller needs reassurance, then a firm warning, then an energetic action. No result is claimed.

Take A — flat control

// prompt — copy me1 line
Your parcel is packed but one detail can still delay delivery. Check the phone number before you confirm the order. A correct number helps the rider reach you. Confirm it now and your order is ready for dispatch.

Generate with one Eleven v3 voice and Stability: Natural. Save take-a-control.mp3. Diagnose whether the warning and action receive the same weight as the setup.

Take B — punctuation and paragraph direction

// prompt — copy me5 lines
Your parcel is packed. But one detail can still delay delivery — your phone number.

Before you confirm the order, check it once more. A correct number helps the rider reach you.

Check. CONFIRM. Then your order is ready for dispatch.

Keep voice, model, and Stability fixed. Save take-b-text-directed.mp3. Full stops shorten the close; paragraph breaks separate the three beats.

Take C — emotional direction for Eleven v3

// prompt — copy me5 lines
[reassuring] Your parcel is packed. [concerned] But one detail can still delay delivery — your phone number.

[thoughtful] Before you confirm the order, check it once more. A correct number helps the rider reach you.

[confident] Check. CONFIRM. Then your order is ready for dispatch.

Keep Stability: Natural and save take-c-tag-directed.mp3. Tags depend on the voice. If [concerned] overacts, remove it before testing Robust on only that beat; record the change.

Do not add exclamation marks, capitals, tags, and a settings change together. If “phone number” lacks emphasis, regenerate only: But one detail can still delay delivery. Your PHONE number.

// failure_cases

Failure Cases to Diagnose

7 cases to diagnose

  • The voice speaks every sentence at one speed

    sentence lengths are too uniform. Break the key line into a short sentence and compare it without touching settings.

  • Pauses become awkward gaps

    paragraph breaks, ellipses, and dashes are stacked together. Keep one pause device and regenerate the sentence.

  • The emotion tag has little effect

    test a compatible voice or compare Natural with Creative; tags do not work equally across voices.

  • The voice reads a direction aloud

    the model does not support the tag syntax used. Remove it and use phrasing, punctuation, or a supported model.

  • Delivery changes unpredictably between takes

    too many variables changed, or stability is low. Return to the control settings and change one item only.

  • A Roman-Urdu word changes accent

    isolate and respell it, then test a shorter segment.

  • Headphones sound fine but the phone sounds rushed

    slow or rewrite the dense sentence; do not simply raise volume, which will not restore articulation.

// pakistan_angle

Pakistan Angle

Test words your Pakistani audience will hear: easypaisa, Karachi, bijli, rupay, neighbourhood names, and English–Roman-Urdu switches. An English-trained voice may carry its accent into Roman Urdu, while long mixed-language text may drift. Use short beats for respelling. Clone only with documented permission, and remove phone, address, CNIC, and order data from scripts.

Under mobile-data and load-shedding constraints, prepare scripts and filenames offline, then generate one short beat before the full track. Download accepted takes and test them through the compression used in your publishing workflow. Before committing USD costs in PKR, check current official pricing and payment options; no plan or Pakistani payment rail is assumed here.

// hands_on

Hands-On Exercise

7 steps

  1. Choose a 60–90-word paragraph from a script you are authorized to use.

  2. Mark three emphasis words, three intended pauses, and the emotional goal of each paragraph.

  3. Generate Take A with unmarked text. Record the voice, model, and every visible setting.

  4. Generate Take B after changing only punctuation, sentence length, and paragraph breaks.

  5. Generate Take C by adding compatible direction tags or changing one setting for one emotional beat.

  6. Fill the review table using player time, audible counts, a 1–5 emotion score, pronunciation errors, and the phone-speaker check.

  7. Regenerate only the lowest-scoring sentence, then write one line explaining why the final take won. Done means you have the marked script, take-a, take-b, and take-c audio files, plus a completed review sheet with a defensible selection.

// completion_rubric

Completion Rubric

6 checks — tick as you verify

0/6

// sources

Sources

// check_yourself

Check yourself

4 questions · answers and options are taken word-for-word from this course

0/4
  1. 1 / 4 · diagnose

    Your work shows this failure mode: “Pauses become awkward gaps.” The lesson describes it like this: “Paragraph breaks, ellipses, and dashes are stacked together.” What does the lesson tell you to do about it?