"Introducing H3 Max Live. Video generation is now faster than real time. An infinite broadcast where every frame is generated on the fly and every scene is directed by chat. Type !prompt and it's on screen in seconds." Two days after fal shipped H3 Max — its post-training of MiniMax H3 running at roughly 35× the official endpoint's throughput and real-time factor 1 at 768p — the company pointed the model at a live stream and let the audience drive. The engine is H3 Max Director, an autoregressive, natively continuous version of the model designed for unbroken, audience-steerable streams with a context window of up to two minutes. That context window is the release. Text-to-video and image-to-video stitch independent clips and hope a first frame carries the continuity; Director remembers what happened up to two minutes prior, so characters, styles, and narratives persist natively while an LLM turns chat prompts and upvotes into the next beat, connecting each scene to the footage before it. Variations of the Director framework fold in voice and audio control — automatic voice cloning and soundscape continuity across extended sequences — so the stream holds together in sound as well as picture. Latent Space called it breaking the infinite-videogen barrier: once generation outpaces viewing, a broadcast can simply never end. Ethan Mollick, testing H3 Max through nothing but the web interface two days earlier, had already called it: "A line in AI video was crossed… H3 Max can now create reasonably high quality AI video in less time than it takes you to watch it." The first experimental streams got kicked off Twitch and YouTube almost immediately — so fal built its own home for the format, fal.live, two days later.