Module 2 — Core CraftChapter 7 · 9 min read
AI Tools for Filmmakers · Core Craft

AI Video Generation

Sora, Veo, Flow — the loudest tools in the kit and the most misunderstood. Here's what an AI video generator genuinely delivers today, where it falls apart mid-scene, and the image-first workflow that separates people who get shots from people who get slot-machine luck.

WR
Will Roberts
Working filmmaker · Written from the set
Video Lesson — Coming Soon

Every conversation I have about AI tools eventually arrives at the same question: "But have you seen the video stuff?" Yes. I've seen it, I've used it, and I've burned an embarrassing number of credits learning what it actually is. So let me save you the tuition: an AI video generator is neither the end of cinema nor a gimmick. It's a very specific instrument that does a few things brilliantly and a few things so badly you'll laugh out loud.

Text-to-video tools like Sora and Veo — and workflow layers built around them like Flow — take a written prompt or a still image and return a few seconds of moving footage. Sometimes those seconds are jaw-dropping. Sometimes the coffee cup melts into the actor's hand. The skill isn't prompting; the skill is knowing which jobs to give the machine.

What generators are genuinely good for today

Forget generating your whole film. Here's where these tools earn their keep on real projects right now:

  • Mood pieces — a thirty-second tonal sketch that shows an investor or collaborator the feel of your film before a frame is shot.
  • Previz in motion — animating your boards so you can test pacing, camera moves, and blocking ideas before you're on the clock with a crew.
  • Impossible establishing shots — the aerial of a city you can't fly to, the period street you can't build, the storm you'd never schedule.
  • Pitch sizzle — proof-of-concept footage that makes a deck feel like a film instead of a document.
  • Texture and inserts — abstract backgrounds, dream fragments, screens-within-screens, the connective tissue that never survives a budget meeting.

Notice what every item on that list has in common: none of it requires the same character to appear twice. That's not an accident. That's the current boundary of the technology.

Where it falls apart — and what it costs

Three failure modes will find you fast. First, character consistency: get a face you love in shot one, and shot two will hand you their slightly-wrong sibling. Cutting between generations of "the same" person is where most AI short films visibly break. Second, physics: hands, liquids, doors, anything with weight — the generators approximate motion rather than understand it, and audiences feel the wrongness before they can name it. Third, dialogue scenes: two people talking in a room, the bread and butter of drama, is exactly what these tools handle worst. Lip sync drifts, eyelines wander, and the emotional continuity that makes a scene a scene simply isn't there.

Then there's the cost reality nobody's marketing department mentions: generation is a slot machine. You will not get your shot on the first pull. You'll get it on the eighth, or the twentieth, and every pull costs credits. Whatever a platform's pricing looks like when you read this, budget for many attempts per usable shot — plan your credits the way you'd plan film stock, and give yourself room to fail toward the good take.

Don't prompt for a shot. Generate a still you love first, then animate it. Control the frame, then buy the motion.

That pullquote is the single most useful workflow habit I can give you. Text-to-video hands the machine every decision at once — composition, light, lens, subject, movement — and you get back an average of all of them. The pro move is image-to-video: build your frame first as a still (in Midjourney or wherever your image work from Chapter 6 lives), iterate cheaply until the composition and lighting are exactly yours, and only then feed that locked frame to the video model with a simple, specific motion instruction. You've reduced the slot machine to one variable. Your hit rate goes up, your credit burn goes down, and — this is the part that matters — the frame is your taste, not the model's.

◆ From the set

Last year I needed a night exterior of a lighthouse in a storm for a pitch reel — a shot that would've cost more than the entire short. I spent one evening generating stills until I had a frame that looked like my film, then animated that frame. Fourteen pulls for three usable seconds. The producer watching the reel asked where we shot it. That's the tool working: not replacing the film, but getting the film in my head onto a screen fast enough to get it funded.

Treat the generator as a second-unit crew that works for cheap, never sleeps, and occasionally hallucinates. Give it the shots no one will scrutinize frame-by-frame; keep the human moments for humans. Next chapter we move to the other half of the sensory equation — voice and sound, where AI's wins are quieter and, honestly, more valuable.

Pairs with this chapter
Filmmaker Toolbox

You're already using AI when you use the Filmmaker Toolbox — script breakdowns, shot lists, budgets, and pitch decks, generated in minutes. It's the producer side of everything this course covers, built for filmmakers.

Open Filmmaker Toolbox

Key takeaways

AI video generators excel at mood pieces, previz in motion, impossible establishing shots, and pitch sizzle — jobs that don't need a character twice.
The failure modes are predictable: character consistency across shots, physics, and dialogue scenes.
Generation is a slot machine — budget credits for many pulls per usable shot, like film stock.
The pro workflow is image-first: lock a still you love, then animate it — image-to-video gives you control text-to-video can't.
← Previous
Image Generation for Previz
Next Chapter →
AI Voice & Sound