Module 6. The half of your film that decides whether anyone believes it. It arrives after picture for a reason, it is the cheapest thing in your pipeline, and almost nobody in this market does the layer that actually works.
Cut your best shot against silence. Then cut it against a room tone, a coat rustle and a distant street.
Same frames. One of them is footage and one of them is a film.
Nothing else in this pipeline changes perceived production value as cheaply as sound does, and nothing else is skipped as often.
You score to picture. You do not generate picture to fit a track you already love.
The moment the music exists first, every timing decision in the film is hostage to it, and you will cut a good shot to protect a bar.
| Layer | What it costs you | What it buys |
|---|---|---|
| Voice | the most attention | the story, if there is dialogue or narration |
| Music | the most taste | the emotion, and the pace |
| Effects | the least of both | the belief |
Effects are the highest return and the most skipped. A room tone under every interior and a footstep on every step will do more for your film than a better music cue.
Your dialogue is already in shots.json if you wrote it there in Module 2. If it is not, it belongs there now, not in a separate script document that will drift.
"Add a dialogue field to every shot that has a line, taken from the script. For each one, note who says it and the delivery in three words or fewer. Flag any shot where the line is longer than the shot's duration can hold."
That last clause catches the thing you would find in the edit. A nine second line in a four second shot is a problem you want in Module 6, not Module 7.
Voice models take direction badly when the direction is vague and well when it is physical, which is the same rule as the camera notes in Module 4.
| Vague | Survives |
|---|---|
| "emotional" | quiet, almost whispered, breath before the last word |
| "angry" | clipped, faster than normal, no pauses |
| "tired" | slower, dropping at the end of every phrase |
Generate three takes of every line, not one. Voice is the cheapest thing you will generate and the one where take three is most often the one.
A voice is a person. Cloning one you do not have permission to use is not a technical question, it is a rights question, and it is the kind that follows a film.
Get permission in writing for any real voice, including your own actors, and keep it with the project. If you are using a synthetic voice, note which model and which settings in the log, the same as a seed. You will be asked, at delivery, what is in your film.
The mistake is generating sixty seconds of "cinematic emotional" and cutting your film to whatever came back.
"From shots.json, write me a music map: where a cue starts, where it stops, what changes in the film at that point, and the total seconds of each cue. One row per cue."
Then generate to that map. You are commissioning a cue with a length and a job, which is what a composer would have asked you for.
This is where the file pays off completely. Every shot already describes its own world.
"For every shot, read the action line and list the sounds that must exist for it to be believable. Separate them into room tone, foreground actions and background world. Write it to audio/spot-list.md, grouped by shot."
That is a spot list, the thing a real sound editor builds first, and you now have one for a forty shot film in about a minute.
Same limit as Module 5, in the other sense.
It cannot tell you the read is flat. It cannot tell you the cue is wrong for the scene, or that the footstep sounds like a different floor.
It can tell you:
"Check the audio: every shot with dialogue has a take, nothing is zero-length or clipping, everything is 48 kHz, and every cue matches the length in the music map."
Every platform has a delivery loudness, and "it sounds fine on my laptop" is not a spec.
"Read my delivery spec and check the mix against it: integrated loudness, true peak, and whether we are inside tolerance. Tell me the numbers and whether it passes. Do not change anything."
Numbers, then a pass or fail, then you decide. Your studio ships a loudness check for exactly this, and it reports rather than fixes, on purpose.
Ask for the measurement before you ask for the fix. A tool that normalises without showing you the number has taught you nothing and may have flattened your dynamics.
Everything in this module prepares, lists, generates, checks and reports.
The balance between her voice and the street is a creative decision, and it is yours.
Your film from Module 5. Sound is the fastest module in this course.
dialogue to every shot that has a line, with a three word delivery note. Flag lines toolong for their shot."*
shots.json, one row per cue, with lengths."audio/spot-list.md: room tone, foreground, background, grouped by shot."shots.json with the picture, not in a document beside itaudio/spot-list.md, a genuine spot list for the whole filmaudio/<shot_id>/ with takes, versioned, and a log of what made each oneYour film has a sound track that was designed rather than found.
Edit. Assembling the thing, the paper edit that gets you there in one sitting, and the point where the tool hands the film back to you for good.
Verified against Claude Code 2.1.220 · 2026-07-30.