The prompt becomes a bundle

Apr – May 2025

The prompt becomes a bundle

Kling 2.0's multimodal visual language and the open-weight answer.

The prompt grew up. Instead of a paragraph of text, you handed the model a bundle: reference images, short video clips, a style, a camera move. Open-source models caught up with the same idea, so the technique was no longer locked inside one app.

Tutorials1

Kling 2.0 is here: multi-elements and prompt structure

Theoretically Media · YouTube · Apr 2025

How it worked

  1. Convey identity, style, scene, action and camera with image references and video clips rather than text alone (Kling 2.0 MVL).
  2. Add, swap or delete subjects in existing video with the Multi-Elements editor.
  3. Kling 2.1 added tiers, ten-second clips and start/end-frame conditioning.
  4. Open weights caught up: Wan 2.1-VACE for reference-to-video and video-to-video in ComfyUI; HunyuanVideo-Avatar for audio-driven multi-character dialogue.
  5. Performance capture from real actors with SAG-AFTRA-approved likeness models, as in Echo Hunter.

Example Films1

Related on the timeline3

Sources & further reading3