GOamplify. 15 Day AI Mastery, Day 7
x2 streak Observer
0 XP
Phase 2, Visual, Audio and Cinematic Media

Day 7. AI Sound and Voice Generation

Choosing a voice is a casting decision. Treating it as a settings dropdown is the most common mistake today.

Day narration
2 Why this matters

Voice carries meaning the words do not.

Identical text in two voices is two different messages. Age, accent, pace and warmth tell the listener who is speaking before a single fact lands. Cast for the listener, not for your own taste: a technical explainer for engineers and a launch film for investors do not want the same instrument.

Tools today: ElevenLabs and Suno. The documented settings spec, not the audio file, is the professional deliverable.

3 Watch first

These teachers did this work publicly. Watch them, then come back.

How to Use ElevenLabs AI, Complete Beginner Tutorial
by Kevin Stratvert
ElevenLabs Complete Platform Tutorial, Beginner to Pro
by Emma’s Productivity Lab
How to Use Suno AI Better than 99% of People
by Isa does AI
4 The core

The written lesson. Read it slowly, it saves you later.

NOVA, key ideaWrite the settings down. A narration pack recorded at different settings will not cut together, and in six months nobody remembers what you used.

The three settings that matter

Stability, low gives expressive, variable, occasionally erratic reads, high gives consistent, controlled, sometimes flat ones. For a narration set that must match across twenty clips, go higher, for a single emotional line, go lower. Similarity, how tightly it clings to the source voice, push it too far and you inherit the artefacts of the sample along with its character. Style and speed, small moves, large changes here usually mean the wrong voice was cast.

Motion graphic, stability, the trade
LOW, EXPRESSIVEMID, BALANCEDHIGH, CONSISTENTtwenty matching clips want high, one emotional line wants low

Directing the read

Punctuation is your direction. A full stop is a beat, a comma is a breath, a paragraph break is a pause. Rewrite text for the ear: shorter sentences, one idea each, no subordinate clauses stacked three deep. Read it aloud yourself first, if you stumble, the model will too.

Numbers, acronyms and place names are the reliable failure. Write them the way they should sound, twenty twenty six, not 2026, when you want it spoken that way. Spell an unfamiliar name phonetically and fix it in the script, not with fifteen regenerations.

Music, and the structure that makes it usable

Text to music gives you something listenable from a bad prompt. Getting something usable takes structure: genre, instrumentation, tempo, mood, and above all a section plan, intro, verse, build, drop, outro. Without a section plan you get ninety seconds of texture that never resolves, which is useless under a film that has an arc. Generate more than you need and choose, this is closer to a music search than to composition.

Motion graphic, a track that resolves
INTROVERSEBUILDDROPOUTROno section plan means texture that never resolves

Sound design is the missing eighty percent

Most AI video is music plus dialogue and nothing else, and it reads as fake for a reason people cannot name. What is missing is the world: room tone, footsteps, cloth, distant traffic, the specific quiet of an empty office. A single well placed atmosphere track does more for believability than a better generated shot. Carry this into Day 9, it is the difference between a demo and a film.

Voice cloning without documented consent is a legal and reputational problem, and in a relationship driven market it is a relationship problem, which is worse. Get it in writing, keep the file, note the scope and the expiry. Check the licence tier on generated music before it goes near a client campaign, particularly paid media, free tiers routinely forbid commercial use and nobody reads that until the invoice.

5 Checkpoint

Three quick questions. Not the exam, just a pulse.

NOVA reacts, nothing is scored, nothing is stored against you.

6 Do the work

Two tracks. Pick yours, produce something.

Student mode

Write a sixty second script explaining something you actually know. Generate it in two different voices and listen back. Then generate a short background track and mix them. Notice how much the voice choice changes the meaning of identical words.

Pro mode

Produce a full narration pack for the brand you built on Day 6. One master voice, consistent settings recorded in a document so the next person can match it, a set of clips with locked durations, and a brand sting. If you clone a voice, get written consent and file it.

7 Exercises and brainstorm

Tick them when they are actually done.

Brainstorm, no ticks, just think
Today's badges
Voice Castercomplete Day 7
Three in a rowcheckpoint streak
Day close

Day 7 in progress

0
XP today
0/3
Exercises
no
Deliverable

Tomorrow, Day 8: avatars. A delivery mechanism for information, never a substitute for a human moment.