Ingredients

Oct – Dec 2025

Ingredients

Three vendors converge on the same input contract in ten weeks.

Within ten weeks, the big three video tools converged on the same idea: give the model ingredients (a few images, a clip, a sound) and get back a multi-shot scene with its own audio. One model, many inputs, one coherent result.

Tutorials4

How to make a professional AI animated film

Curious Refuge · YouTube · Oct 2025

How it worked

  1. Veo 3.1 Ingredients to Video: up to three reference images (character, object, scene, style), now with audio; Frames to Video for first and last frame; Extend to chain past a minute.
  2. Storyboards in Nano Banana Pro (fourteen references, identity for five subjects, real text) become first frames.
  3. Kling O1, the first unified multimodal video model: seven simultaneous inputs, @Element1 / @Image1 prompt syntax, prompt-driven editing of uploaded video.
  4. Runway Gen-4.5 with native audio and minute-long multi-shot generation; Wan 2.6 reference-to-video where a clip carries appearance and voice.
  5. Node canvases go mainstream (Runway Workflows, Freepik Spaces, Krea Nodes) and open weights keep pace: LTX-2 with native audio, HunyuanVideo 1.5 on a single 4090.
  6. Use the image model as a coverage engine: generate every angle of a scene in Nano Banana Pro before touching a video model, then animate only the frames you would actually cut to (blizaine, Framer).
  7. Let the language model write the prompts, five self-contained shots at a time, each written as if the video model has no memory of the last (PJ Ace).
  8. Cut hard between only the beats that matter: short clips and jump cuts read as intent and sidestep long-clip drift (Framer).
  9. Bolt open-source voice cloning (E2/F5-TTS) onto the stack for dialogue the video model cannot deliver.

Example Films1

More making-ofs3

How films were actually made, from the people who made them.

Related on the timeline7

Sources & further reading7