
Oct – Dec 2025
Ingredients
Three vendors converge on the same input contract in ten weeks.
Within ten weeks, the big three video tools converged on the same idea: give the model ingredients (a few images, a clip, a sound) and get back a multi-shot scene with its own audio. One model, many inputs, one coherent result.
Tutorials4
How to make a professional AI animated film
Curious Refuge · YouTube · Oct 2025
How it worked
- Veo 3.1 Ingredients to Video: up to three reference images (character, object, scene, style), now with audio; Frames to Video for first and last frame; Extend to chain past a minute.
- Storyboards in Nano Banana Pro (fourteen references, identity for five subjects, real text) become first frames.
- Kling O1, the first unified multimodal video model: seven simultaneous inputs,
@Element1 / @Image1prompt syntax, prompt-driven editing of uploaded video. - Runway Gen-4.5 with native audio and minute-long multi-shot generation; Wan 2.6 reference-to-video where a clip carries appearance and voice.
- Node canvases go mainstream (Runway Workflows, Freepik Spaces, Krea Nodes) and open weights keep pace: LTX-2 with native audio, HunyuanVideo 1.5 on a single 4090.
- Use the image model as a coverage engine: generate every angle of a scene in Nano Banana Pro before touching a video model, then animate only the frames you would actually cut to (blizaine, Framer).
- Let the language model write the prompts, five self-contained shots at a time, each written as if the video model has no memory of the last (PJ Ace).
- Cut hard between only the beats that matter: short clips and jump cuts read as intent and sidestep long-clip drift (Framer).
- Bolt open-source voice cloning (E2/F5-TTS) onto the stack for dialogue the video model cannot deliver.
Example Films1
More making-ofs3
How films were actually made, from the people who made them.
Related on the timeline7