Module 4. Your approved frames become shots. Same file, same shot ids, one queue, and the expensive part runs while you sleep instead of while you watch.
images/approved/ and status: approved in shots.json. That is the whole input.
Nothing gets animated on a maybe. Video is the most expensive generation you will run, in credits, in minutes, and in the attention it costs to review.
If you skipped Module 3 and you are animating straight from text, this is where that bill arrives.
You stop describing a scene and start describing a movement.
The frame is already decided. The only new information is what moves, and how the camera behaves while it does.
Your prompts get a second field. Keep it short and physical:
| Vague | Survives the round trip |
|---|---|
| "cinematic" | slow push in, 20mm, handheld micro-drift |
| "dynamic" | whip pan left to right, motion blur |
| "emotional" | static, subject turns toward camera, everything else still |
Physical beats poetic. A model can execute "slow push in". It cannot execute "cinematic". It can only guess what you meant, differently each time.
"Add amotionfield to every approved shot inshots.json. Write it from that shot'scamerafield in physical terms: what moves, which direction, how fast. Flag any shot where the camera note is too vague to execute."
A video generation takes minutes. Twelve of them takes an evening, an evening most people in this market spend watching a progress bar and pasting the next prompt.
"Build a queue from every approved shot: submit, poll until done, download tovideo/<shot_id>/v<n>.mp4, log prompt, model, seed and duration tovideo/log.jsonl, then take the next one. Keep going until the queue is empty."
Start it and leave. This is the "works at 3am" line from Module 1, cashed in. It is also the first thing in this course that is genuinely unattended.
A queue that dies at shot 4 and reports success is worse than no queue. Ask for the boring parts explicitly:
Then kill it halfway and run it again. If it picks up where it stopped, it is real. If it starts from shot 1, you found out now instead of at 3am.
Unattended and unbounded are different things. The queue spends real money while you are asleep, so it gets told the number:
"Before you submit anything, price the queue: how many shots, how many seconds, what that costs at my rate. Show me the total and wait. Then stop at that ceiling and report, even if the queue is not empty."
A ceiling and a dry run, every time. The expensive mistake is not a bad generation. It is forty of them, made correctly, from a prompt you would have fixed in one line at shot two.
The limit from Module 1 has not moved. It cannot watch the video.
It can tell you the file exists, its duration, its resolution, its codec, whether every approved shot produced one. It cannot tell you the hands are wrong or the face swims.
"Check every approved shot has a video. Report anything missing, zero-length, under-duration, or the wrong resolution. Then build me a contact sheet: first frame, middle frame, last frame, per shot."
It checks what is checkable, then hands you the looking.
Shot 7 came back wrong. You do not start over. You change one row.
"Shot 7's motion is too fast. Changemotionon that row to a slow push, regenerate that shot only, keep the old take asv1and file the new one asv2."
One row, one regeneration, both takes kept.
This is the compounding difference between a pipeline and a session. A client note on Thursday costs you one row, not a lost afternoon reconstructing what you did on Monday.
Your approved frames from Module 3. Start it before dinner.
motion field to every approved shot, physical terms, from the camera field. Flagthe vague ones."*
video/<shot_id>/, log everything, retry twice,skip anything already downloaded."*
Step 4 is not optional. Watch a queue do two shots before you trust it with twenty.
shots.json carrying motion notes alongside the prompts: one file, the whole shotvideo/<shot_id>/: every take, versioned, never overwrittenvideo/log.jsonl: prompt, model, seed and duration for every second of footage you ownYou have footage, and you were not in the room when most of it arrived.
Consistency. Holding the face, the coat and the light across every shot, and what to do with the ones that drifted, without regenerating the whole film.
Verified against Claude Code 2.1.220 · 2026-07-30.