Vibe Video: The Coding Model That Makes Films by Writing Frontend Code

Over the past month, the most talked-about AI video in the world was not made by a video model. It was made by a coding model that wrote frontend code until an animation appeared.
The clips came from Claude Opus 5.5. In one widely shared case, a creator used a reference prompt, access to Midjourney, and a mood board to produce a cinematic dystopian short in about twelve hours. The post passed 20 million views. Another creator generated a two-minute-sixteen history-of-civilization piece with a single instruction; it cleared ten million views in two days. A GitHub collection of Opus 5.5 videos has gathered more than 400 entries, most with public prompts and code.
People started calling it Vibe Video. The name is clumsy, but the technique is worth understanding, because it is a genuinely different path to moving images than the one everyone else is on.
Code as the rendering engine
The usual AI video pipeline is a diffusion model predicting pixels. You feed it a prompt or a still, it returns frames. The model is trained on video, and its sense of motion is learned from that data.
Code rendering works differently. A language model writes HTML, CSS, SVG or JavaScript that draws an animation, and a browser renders it. Nothing is predicted frame by frame. The motion comes from code that explicitly describes position, timing and easing.
That distinction matters because it makes some things easy that diffusion finds hard. Precise graphics, typography, charts, geometric motion and transitions between screens are trivial when they are code. They are exactly the things that come out muddy from a video model. The trade is that code rendering cannot produce a photorealistic human face in motion. It draws what code can describe.

Why creators glommed on to it
Two properties explain the enthusiasm. The first is reproducibility. A diffusion clip is a sample; run it again and you get a cousin, not a twin. A coded animation is a program. Change one value and you know what happens. That makes revision cheap, and revision is where most of the work actually lives.
The second is that the output is already in a format the rest of a workflow can use. A browser animation can be captured, embedded in a site, or wired into a data source. It sits naturally inside a product page or a dashboard, which is where a lot of commercial demand for motion actually comes from.
There is also a supply-side reason. Claude Opus 5.5 launched on September 22 with a 230-page system card and API pricing at four dollars per million input tokens and twenty per million output. It benchmarked near the previous flagship on most tasks at meaningfully lower cost. Cheaper frontier models make long, iterative code generation practical for individuals, which is what the twelve-hour builds actually require.
The other lane
The same week, another general model was moving into video from a different direction. GPT-6 Astra was reported to act as a kind of AI director, handling storyboard planning and white-model pre-visualization and reaching into Blender, Unreal Engine and DaVinci Resolve to build scenes, block shots and handle edits.
Put the two together and a pattern shows up. General-purpose models are not trying to beat Seedance, Kling or MiniMax at photoreal motion. They are moving upstream, to the planning and the production pipeline, and sideways, into motion that is programmatic rather than photographic. That leaves the specialized video models to keep pushing on realism while the general models take the parts of the process that look like software.
How the workflow actually runs
The process described by the creators getting good results is closer to software development than to filmmaking. You write a prompt that sets the mood and the constraints, the model produces code, the code renders in a browser, and you look at the result. Then you change one thing and run it again. The loop is short and the output is inspectable, which is why twelve-hour builds are possible at all. Twelve hours of handing a paragraph to a video model would produce noise. Twelve hours of editing code produces a finished piece.
The mood board and the reference model matter for a reason that has nothing to do with composition. They give the stills a consistent palette and subject, so that when the code places them and animates around them, the pieces look like they belong to one film. The code supplies the motion and the typography. The stills supply the world. Neither half works without the other.
Where the cost really lands
It is tempting to think code-rendered video is free, because the rendering happens in a browser on your own machine. The model calls are not free, and the loop is long. Every rewrite is a paid request, and a piece with hundreds of edits is hundreds of requests with a large context attached. Anthropic's price cut to four dollars per million input tokens is what made that loop affordable for individuals rather than just studios. The cheaper the model, the more freely a creator can iterate, and iteration is the entire method.
That has a strange consequence. The bottleneck for this style is not compute for rendering. It is review. Someone has to watch each render and decide what to change, and that attention does not scale with a lower API price. The creators succeeding with Vibe Video are, functionally, directors staring at dailies and giving notes, except the notes go into a diff.
Code rendering and generated video, side by side
It helps to be precise about which style does what. Code rendering is reproducible, cheap to revise, sharp on graphics and text, and comfortable with data. It cannot generate a photorealistic actor, a crowd, or a candid moment. Diffusion video does those well and treats everything else as an approximation. A logo rendered by a video model comes out almost right. A logo rendered by code is exact.
That split suggests the two will live together rather than compete. A piece might open with a generated establishing shot, cut to a coded animated chart, and close on generated footage. The interesting creative question stops being which tool to use and becomes where to switch. The teams that figure out those seams will make work that reads as deliberate.
Where it does not work
Code rendering is not a replacement for video generation, and treating it as one will disappoint people. It cannot shoot an actor. It cannot produce a documentary frame or a believable crowd. It is strong on design, abstract motion, information and interface, and weak on anything that needs to look like a camera was there.
It also has a cost profile nobody talks about. You can burn a lot of model tokens before a single frame looks right, because the loop is write, render, inspect, rewrite. Cheaper models lower that cost, but they do not remove it. The creators getting the striking results are not just prompting; they are reviewing code.
What it tells us about the field
The interesting thing about Vibe Video is not that a coding model can make a nice clip. It is that the boundary between "video generation" and "software that makes things move" turned out to be softer than anyone assumed. When a language model can write the animation directly, some of the demand for generative video is answered by code instead.
For people who make visuals, the practical question is no longer whether to pick one tool. It is which parts of a piece are best served by a predicted frame and which are best served by a written one. The teams figuring that out now are the ones whose work will look deliberate instead of merely generated.
Related articles
AI Short Drama Hit the Hot Search, and Live-Action Shoots Fell 70%
Only one of the top 20 titles on a major Chinese drama chart was made with real actors. AI costs about a tenth as much to produce, and it is rewriting the whole pipeline.
Decagon's PACT Protocol Wants Consent to Be a Standard
Decagon open-sourced a protocol for verifying a personal agent's identity and the permissions a customer granted it. It is plumbing, and it decides whether the agent economy works.
The AI Reunion Wave China Can't Decide How to Feel About
AI tribute films brought departed public figures back to Chinese screens and pulled in hundreds of thousands of likes. Then the backlash arrived, and it was about consent.
Apple Published an Open Multimodal Model and Barely Told Anyone
Apple's research-first release slipped past the mainstream, but the model's fine-grained visual grounding says a lot about where its AI stack is heading.