AI Video Generation
Sora, Veo, Flow — the loudest tools in the kit and the most misunderstood. Here's what an AI video generator genuinely delivers today, where it falls apart mid-scene, and the image-first workflow that separates people who get shots from people who get slot-machine luck.
Every conversation I have about AI tools eventually arrives at the same question: "But have you seen the video stuff?" Yes. I've seen it, I've used it, and I've burned an embarrassing number of credits learning what it actually is. So let me save you the tuition: an AI video generator is neither the end of cinema nor a gimmick. It's a very specific instrument that does a few things brilliantly and a few things so badly you'll laugh out loud.
Text-to-video tools like Sora and Veo — and workflow layers built around them like Flow — take a written prompt or a still image and return a few seconds of moving footage. Sometimes those seconds are jaw-dropping. Sometimes the coffee cup melts into the actor's hand. The skill isn't prompting; the skill is knowing which jobs to give the machine.
What generators are genuinely good for today
Forget generating your whole film. Here's where these tools earn their keep on real projects right now:
- Mood pieces — a thirty-second tonal sketch that shows an investor or collaborator the feel of your film before a frame is shot.
- Previz in motion — animating your boards so you can test pacing, camera moves, and blocking ideas before you're on the clock with a crew.
- Impossible establishing shots — the aerial of a city you can't fly to, the period street you can't build, the storm you'd never schedule.
- Pitch sizzle — proof-of-concept footage that makes a deck feel like a film instead of a document.
- Texture and inserts — abstract backgrounds, dream fragments, screens-within-screens, the connective tissue that never survives a budget meeting.
Notice what every item on that list has in common: none of it requires the same character to appear twice. That's not an accident. That's the current boundary of the technology.
Where it falls apart — and what it costs
Three failure modes will find you fast. First, character consistency: get a face you love in shot one, and shot two will hand you their slightly-wrong sibling. Cutting between generations of "the same" person is where most AI short films visibly break. Second, physics: hands, liquids, doors, anything with weight — the generators approximate motion rather than understand it, and audiences feel the wrongness before they can name it. Third, dialogue scenes: two people talking in a room, the bread and butter of drama, is exactly what these tools handle worst. Lip sync drifts, eyelines wander, and the emotional continuity that makes a scene a scene simply isn't there.
Then there's the cost reality nobody's marketing department mentions: generation is a slot machine. You will not get your shot on the first pull. You'll get it on the eighth, or the twentieth, and every pull costs credits. Whatever a platform's pricing looks like when you read this, budget for many attempts per usable shot — plan your credits the way you'd plan film stock, and give yourself room to fail toward the good take.
That pullquote is the single most useful workflow habit I can give you. Text-to-video hands the machine every decision at once — composition, light, lens, subject, movement — and you get back an average of all of them. The pro move is image-to-video: build your frame first as a still (in Midjourney or wherever your image work from Chapter 6 lives), iterate cheaply until the composition and lighting are exactly yours, and only then feed that locked frame to the video model with a simple, specific motion instruction. You've reduced the slot machine to one variable. Your hit rate goes up, your credit burn goes down, and — this is the part that matters — the frame is your taste, not the model's.
Last year I needed a night exterior of a lighthouse in a storm for a pitch reel — a shot that would've cost more than the entire short. I spent one evening generating stills until I had a frame that looked like my film, then animated that frame. Fourteen pulls for three usable seconds. The producer watching the reel asked where we shot it. That's the tool working: not replacing the film, but getting the film in my head onto a screen fast enough to get it funded.
Treat the generator as a second-unit crew that works for cheap, never sleeps, and occasionally hallucinates. Give it the shots no one will scrutinize frame-by-frame; keep the human moments for humans. Next chapter we move to the other half of the sensory equation — voice and sound, where AI's wins are quieter and, honestly, more valuable.
You're already using AI when you use the Filmmaker Toolbox — script breakdowns, shot lists, budgets, and pitch decks, generated in minutes. It's the producer side of everything this course covers, built for filmmakers.
