
Apr – May 2025
The prompt becomes a bundle
Kling 2.0's multimodal visual language and the open-weight answer.
The prompt grew up. Instead of a paragraph of text, you handed the model a bundle: reference images, short video clips, a style, a camera move. Open-source models caught up with the same idea, so the technique was no longer locked inside one app.
Tutorials1
Kling 2.0 is here: multi-elements and prompt structure
Theoretically Media · YouTube · Apr 2025
How it worked
- Convey identity, style, scene, action and camera with image references and video clips rather than text alone (Kling 2.0 MVL).
- Add, swap or delete subjects in existing video with the Multi-Elements editor.
- Kling 2.1 added tiers, ten-second clips and start/end-frame conditioning.
- Open weights caught up: Wan 2.1-VACE for reference-to-video and video-to-video in ComfyUI; HunyuanVideo-Avatar for audio-driven multi-character dialogue.
- Performance capture from real actors with SAG-AFTRA-approved likeness models, as in Echo Hunter.
Example Films1
Related on the timeline3