Module 03 · AI Voice and Narration
Cloning and Directing a Voice in ElevenLabs
Open lesson + course map
On this lesson
Course outline
Module 1 · The AI Video Stack in 2026
Module 2 · Scripting for AI Video
Module 3 · AI Voice and Narration
Module 4 · Generative Video With Veo
Module 5 · AI Avatars and Presenter-Style Video
Module 6 · Editing and Assembly in CapCut
Module 7 · Video Service Packaging
A voice clone lets a text-to-speech system produce new speech in a person's vocal identity. That power creates a hard boundary: clone only your own voice or a voice whose owner has given documented, purpose-specific permission. A public video, podcast, or WhatsApp voice note is not permission.
You will prepare source audio, create and direct an ElevenLabs Instant Voice Clone, repair weak sentences without wasting credits, and package the narration with its consent record, test script, settings sheet, and disclosure. Confirm current interface labels and plan gates in the official guides linked below.
// concept
Get Permission Before Audio Enters ElevenLabs
Write permission before recording or uploading. Name the voice owner, project, approved subject, channels, commercial status, access, expiry, and revocation contact. Both parties keep a dated copy. Do not collect a CNIC image merely to prove consent; that creates unnecessary identity-data risk.
Use clear language, not a blanket release. This sample is a template, not legal advice:
VOICE-CLONE CONSENT — SAMPLE
Voice owner: [name]
Project: [specific video or channel]
I permit [producer] to upload my approved recordings to ElevenLabs and use
the resulting voice only for [purpose], on [channels], until [date].
Scripts require my approval: [yes/no]. Commercial use allowed: [yes/no].
I may withdraw permission by contacting [email/phone]. After withdrawal,
new generation stops and the clone is deleted, subject to any written contract.
Agreed by voice owner: [signature/email confirmation and date]
Agreed by producer: [signature/email confirmation and date]If permission is unclear, use your own voice or a licensed/default synthetic voice. Never make someone appear to say unapproved words. Where a listener could mistake the output for a new human recording, disclose: AI-generated narration using a voice clone with the speaker's written permission.
// concept
Record and Create the Clone
Record in a quiet, soft-furnished room at one microphone position. Turn off fans and notifications. Keep distance and delivery consistent; the clone can reproduce echo, mouth clicks, speed, accent, and emotion. ElevenLabs' current guide recommends roughly one to two minutes of clean audio and prioritizes capture quality over extra duration. Confirm this before recording.
Match the intended performance: record a calm explainer calmly, and include every target language. A Roman-Urdu channel needs natural Roman-Urdu source delivery, not only English.
The current official path is Voices → plus icon → Instant Voice Clone. Upload the authorized sample, name it for the project, confirm rights and consent, and save. Open Voices → My Voices → Use voice for Text to Speech. Confirm the current UI and plan gate. Instant Voice Cloning is currently listed on a paid plan; on a free account, practise directing with an included voice rather than bypassing the gate.
// concept
Direct the Narration With Text and Settings
Voice choice and source audio matter more than slider hunting. Save one baseline and change one variable per test. If a model omits a control, record not shown.
| Control (confirm current UI) | What moving it changes | Practical diagnostic |
|---|---|---|
| Stability | Lower allows more variation; higher is steadier but may sound flat | Raise for pacing swings; lower for rigid delivery |
| Similarity | Higher follows the source more closely, including possible recording flaws | If artifacts grow, fix the source instead of maxing this control |
| Style exaggeration | Amplifies source performance but can reduce stability | Start at zero; test only for a deliberate style |
| Speaker Boost | May add subtle similarity at processing cost | Compare on/off with identical text |
| Speed | Changes rate where the model exposes it | Start at 1.0; avoid quality-damaging extremes |
The official guide currently gives a common baseline near Stability 50, Similarity 75, Style 0, and Speed 1.0 where available. These are test values, not universal best settings. Save the model name because newer models may direct delivery differently.
Use commas for grouping, full stops for endings, and paragraph breaks for scene changes. For a local name, test only its sentence with a phonetic spelling such as Ra-wal-pin-dee; keep correct spelling in captions. Confirm current Studio pronunciation-dictionary support before relying on it.
Treat the credit or character allowance as finite. Lock the script, generate a short diagnostic, and split the approved script into sections. Regenerate only a failed sentence with surrounding context. Check the displayed cost and current pricing before Generate; never assume a regeneration is free.
// worked_example
Worked Example
This hypothetical sample follows an Islamabad editor making an authorized property explainer. The narrator consents to this one client video. The editor names the clone Property explainer — consent expires 30 Sep and tests:
Assalam-o-alaikum. Before you visit the property, confirm three things:
the exact location, the written payment schedule, and who will verify the documents.
This sample listing is in Rawalpindi. Save the agent's number, but do not send
your CNIC photo on WhatsApp until you have confirmed why it is required.
Ready? Let us check the details, one step at a time.Sample settings sheet: model [exact current model name]; Stability 50; Similarity 75; Style 0; Speed 1.0; Speaker Boost off; disclosure required yes. Controls not displayed by the chosen model are marked not shown.
In sample output one, Rawalpindi is rushed and the last line sounds flat. The editor tests Ra-wal-pin-dee in isolation and runs another take at Stability 45, changing nothing else. The log says: pronunciation improved; delivery more varied; one breath artifact remains. They assemble the clean takes and keep correct spelling in captions.
The folder contains consent.pdf, source-audio.wav, test-script.txt, settings-v02.md, narration-v02.wav, and disclosure.txt. This is a reviewable production record, not a claim of campaign results.
// failure_cases
Failure Cases to Diagnose
6 cases to diagnose
The clone carries hiss or echo
compare it with the source sample. If the noise is already there, re-record; a higher Similarity value may reproduce it more strongly.
Roman-Urdu words drift toward an English accent
confirm the source contains natural speech in the target language, then test one phonetic respelling at a time.
Every take sounds flat
the source performance may be flat or Stability may be too high. Record a representative performance before pushing Style exaggeration.
Pacing changes across the full script
split narration at paragraph boundaries, keep the same model/settings, and regenerate only the inconsistent section.
A fix consumes credits without changing the fault
stop changing several sliders together. Use identical diagnostic text and alter one setting, punctuation mark, or spelling per take.
Consent is impossible to prove
pause production. Obtain a dated, purpose-specific record or delete the clone; a verbal “the client is fine with it” is not an audit trail.
// pakistan_angle
Pakistan Angle
ElevenLabs charges in foreign currency, and its current page lists card/platform-wallet methods rather than JazzCash or easypaisa. A Pakistani card may add exchange spread, taxes, or international charges. Confirm the PKR total with the bank. If cloning is too costly, rehearse with an included voice and subscribe only when project and consent are ready.
Roman Urdu has no standard spelling, so keep a pronunciation sheet: Rawalpindi → Ra-wal-pin-dee, easypaisa → easy-paisa, plus approved Urdu-English switches. These are generation aids, not caption spellings. Test through an ordinary phone speaker used for WhatsApp, YouTube, or TikTok.
For load-shedding and mobile-data limits, clean audio offline, upload one approved sample, test by sentence, and download on stable Wi-Fi. Keep CNIC scans and unapproved WhatsApp voice notes out of the folder. Restrict consent-record access and follow its deletion date.
// hands_on
Hands-On Exercise
7 steps
Choose your own voice or obtain the signed/email consent record above. Define the project, channels, expiry date, disclosure, and revocation route.
Record one to two minutes of clean, representative speech. Listen with headphones for echo, clipping, background voices, and inconsistent distance.
Create the clone through the current Voices workflow, or use an included voice if your plan does not offer cloning. Record the actual path and plan gate you saw.
Generate the worked-example test script or an equally demanding approved test containing a local name, a number, a pause, and an Urdu-English switch.
Complete a settings sheet with model, every visible control, text version, observed fault, one-variable change, selected take, and credit cost shown before generation.
Generate the approved script in sections, repair only failed sentences, assemble one narration track, and save the disclosure alongside it.
Test revocation: identify the current My Voices → More actions (three dots) → Delete voice path. Do not delete a needed clone, but document who will perform deletion and when. If consent is withdrawn, stop new use, remove sharing where applicable, delete the clone, and record completion; contact ElevenLabs Support when the UI prevents deletion. Done means the folder contains documented consent, source/test audio, the test script, a complete settings-and-regeneration log, one assembled narration track, and the disclosure text.
// completion_rubric
Completion Rubric
5 checks — tick as you verify
// sources
Sources
4 official sources — check every claim yourself
// check_yourself
Check yourself
4 questions · answers and options are taken word-for-word from this course
1 / 4 · diagnose
Your work shows this failure mode: “The clone carries hiss or echo.” What does the lesson tell you to do about it?