Anyone can generate a beautiful five second clip. Almost nobody can generate thirty seconds that hold a viewer.
This is the day the course stops being about tools, because holding a viewer for thirty seconds requires the things film schools teach and prompt guides do not: sequence, continuity, escalation and rhythm.
Tools today: Higgsfield, Google Flow, Runway. Generation is expensive in time and credits, and every regeneration tempts you to accept whatever came back, so the decisions get locked before any tool is opened.
Shot list, six shots for fifteen seconds is a workable density, subject, camera, light and duration per shot, written before any tool is opened. Character reference, one locked image per character, identity drifts between generations unless you carry a reference through every shot. Location bible, the same wall, the same window, the same time of day, audiences forgive a lot and do not forgive a room changing shape between cuts. Colour grade, one grade across the whole piece, decided up front, mixed grades are the fastest way to make good shots look like a mood board.
They learned from captioned footage, so they respond to real terminology: dolly in, track left, crane up, handheld, static locked off, rack focus, 24mm wide, 85mm portrait, shallow depth of field. They are weak on complex compound moves, a shot that dollies while craning while whip panning will come back as something else. Ask for one clear move per shot. If you need complexity, get it in the cut, not in the generation.
Establish which side of the action the camera lives on and stay there. Cross the line and your two characters appear to swap places, and the viewer feels a wrongness they will not be able to name. This is the single most common structural error in AI film work, because each shot is generated in isolation with no memory of where the camera was. The model will not enforce it. You enforce it, in the shot list, before generation.
Cut length is meaning. Long holds create weight and unease, short cuts create energy and, held too long, exhaustion. Vary deliberately, a thirty second film cut entirely at two seconds per shot is flat regardless of how good the shots are.
The mute test settles it: play the cut with no sound. If the story still reads, the sequencing works. If it only works with music, the music is carrying a film that does not exist.
Day 7 said sound design is the missing eighty percent, and it is more true here: room tone, footsteps, cloth, wind, the specific quiet of an interior. Most AI films fail on the audio bed, not the imagery. And frame for the crop from the start, a sixteen by nine master reframed to nine by sixteen at the end will cut the subject’s head off. Decide the aspect ratios in the shot list and protect the frame while generating.
NOVA reacts, nothing is scored, nothing is stored against you.
Write a six shot sequence with a beginning, a turn and an end. For each shot specify subject, camera move and light. Generate them, cut them together, add the music you made on Day 7. Watch where the cuts feel wrong and write down why.
Produce a thirty second brand film as two fifteen second segments of six shots each. Lock a character reference, a location bible and a colour grade before you generate a single frame. Respect the 180 degree rule across cuts. Include sound design, not just music. Deliver a sixteen by nine master and a nine by sixteen social cut.
Day 9 in progress
Tomorrow, Day 10: vibe coding. Software used to require a builder. Now it requires a specifier.