Choosing a voice is a casting decision. Treating it as a settings dropdown is the most common mistake today.
Identical text in two voices is two different messages. Age, accent, pace and warmth tell the listener who is speaking before a single fact lands. Cast for the listener, not for your own taste: a technical explainer for engineers and a launch film for investors do not want the same instrument.
Tools today: ElevenLabs and Suno. The documented settings spec, not the audio file, is the professional deliverable.
Stability, low gives expressive, variable, occasionally erratic reads, high gives consistent, controlled, sometimes flat ones. For a narration set that must match across twenty clips, go higher, for a single emotional line, go lower. Similarity, how tightly it clings to the source voice, push it too far and you inherit the artefacts of the sample along with its character. Style and speed, small moves, large changes here usually mean the wrong voice was cast.
Punctuation is your direction. A full stop is a beat, a comma is a breath, a paragraph break is a pause. Rewrite text for the ear: shorter sentences, one idea each, no subordinate clauses stacked three deep. Read it aloud yourself first, if you stumble, the model will too.
Numbers, acronyms and place names are the reliable failure. Write them the way they should sound, twenty twenty six, not 2026, when you want it spoken that way. Spell an unfamiliar name phonetically and fix it in the script, not with fifteen regenerations.
Text to music gives you something listenable from a bad prompt. Getting something usable takes structure: genre, instrumentation, tempo, mood, and above all a section plan, intro, verse, build, drop, outro. Without a section plan you get ninety seconds of texture that never resolves, which is useless under a film that has an arc. Generate more than you need and choose, this is closer to a music search than to composition.
Most AI video is music plus dialogue and nothing else, and it reads as fake for a reason people cannot name. What is missing is the world: room tone, footsteps, cloth, distant traffic, the specific quiet of an empty office. A single well placed atmosphere track does more for believability than a better generated shot. Carry this into Day 9, it is the difference between a demo and a film.
Voice cloning without documented consent is a legal and reputational problem, and in a relationship driven market it is a relationship problem, which is worse. Get it in writing, keep the file, note the scope and the expiry. Check the licence tier on generated music before it goes near a client campaign, particularly paid media, free tiers routinely forbid commercial use and nobody reads that until the invoice.
NOVA reacts, nothing is scored, nothing is stored against you.
Write a sixty second script explaining something you actually know. Generate it in two different voices and listen back. Then generate a short background track and mix them. Notice how much the voice choice changes the meaning of identical words.
Produce a full narration pack for the brand you built on Day 6. One master voice, consistent settings recorded in a document so the next person can match it, a set of clips with locked durations, and a brand sting. If you clone a voice, get written consent and file it.
Day 7 in progress
Tomorrow, Day 8: avatars. A delivery mechanism for information, never a substitute for a human moment.