← Back to blog
AiAbout 7 min read

Kling 4.0 Ships: Ten Keyframes, Thirty Seconds, and the End of the Slot-Machine Edit

Published Oct 7, 2026
Kling 4.0 Ships: Ten Keyframes, Thirty Seconds, and the End of the Slot-Machine Edit

Kuaishou's Kling 4.0 opened internal testing on September 28, with the lightweight Kling 4.0 Flash tier going out to members the same evening and the full model promised for October. The company has been running a specification war with ByteDance's Seedance for the better part of a year, and this release is its clearest statement yet about what the fight is actually over.

The headline numbers are straightforward. Native single-pass generation doubles from Kling 3.0's 15 seconds to 30, with video extension promising up to two minutes across multiple passes. Output reaches 4K and 1080p in 10-bit HDR, plus a 21:9 ultrawide frame. Audio moves to two-channel stereo, and the lip-sync drift that dogged the previous generation is the specific defect the release notes claim to have fixed. Prompt input rises to 8,000 tokens.

None of those numbers is the interesting one.

Ten keyframes is a different product category

Kling 3.0 accepted a start frame and an end frame. Two images, and the model interpolated everything between them. Anyone who has tried to shoot a product video that way knows what it feels like: you can set the bookends and then you watch what you get. Chinese creators have a word for it, 抽卡, drawing cards, the same term gacha games use for randomized pulls.

Kling 4.0 accepts up to ten keyframes. That sounds like a quantity change. It is not.

With ten control points you can describe a sequence of beats inside one generation: character enters, walks to the counter, camera pushes in, product rotates, hand reaches for it, cut to a wider framing. Each of those beats gets its own anchor image, so the model is no longer guessing at the shape of the arc between two endpoints. It is filling in the gaps between ten of them.

For a team producing short commerce video at volume, that changes the unit of delivery. Before, a 30-second product spot meant a shoot or a stock-footage edit, with AI generation confined to a few seconds of filler. Now the entire timeline can be produced from a script with keyframes mapped to shot beats. The work shifts from "generate and hope" to "storyboard, then generate."

Five blank brass storyboard cards laid in a neat row on a dark walnut table beside a slim steel ruler

The comment section on Kuaishou's own announcement is worth reading, because the top-voted question had nothing to do with capability, and everything to do with price. That is a fairly honest signal about where the constraint sits in 2026: the capability ceiling keeps rising, and the cost floor has not moved nearly as fast.

What the multimodal references actually control

Kling 4.0 supports up to 15 multimodal references in a single task, up to 10 images, 5 video clips, and 7 subject combinations. Kuaishou has branded this Omni Reference, and it also lets you add, modify, or delete subjects and backgrounds inside existing footage.

Read that list again and it starts to look less like a generation feature and more like a consistency contract. The recurring failure mode in AI video has never been the single frame, it is the fourth shot, where the actor's face has drifted, the jacket has changed shade, and the product label has morphed into something that would not survive a legal review. Locking identity to 15 reference slots across a 30-second generation is an attempt to answer that directly.

There is a real production constraint hiding here. Reference material has to exist before you can use it, and building a character sheet, a product turnaround, and a voice sample is itself a production job. The teams that get value from Kling 4.0 will be the ones that already have that library assembled, or that have made assembling it part of their pipeline.

Speed, cost, and where Kling actually stands

Kling 4.0 Flash is the high-frequency tier: faster generation, lower cost, capped at 720p during early access. The full model is the professional tier for film and advertising work.

It is worth being precise about the competitive picture rather than repeating Kuaishou's framing. On the Artificial Analysis text-to-video leaderboard, Kling 3.0 sat around twelfth, behind Google's Gemini Omni Flash, Alibaba's Wan 3.0, and MiniMax's H3 Max. Kling 4.0 is a serious capability step, and whether it closes that ranking gap is an open question until independent evaluations land.

The market context matters more than the ranking. Kuaishou reported about $3 billion in independent funding in July 2026 at a roughly $18 billion post-money valuation, reportedly against a five-year listing commitment. Citigroup's read, per Securities Times, centers on two questions: whether Kling can win users back from Seedance, which reportedly holds over 80 percent of the domestic market, and whether the pricing is competitive enough to matter.

The price question is not incidental. The same week Kling 4.0 was announced, HeyGen launched a universal video model at one cent per second, Creatify shipped a post-trained MiniMax H3 derivative at four cents per second, and SpaceXAI pushed a Grok Imagine Video lite tier out through third-party platforms with a claim of being four times cheaper than the previous option. The generation layer is being commoditized from several directions at once, and a model that leads on 4K HDR output is competing on a dimension its rivals have decided not to lead on.

That reframes what Kling 4.0 is for. The ambition is to be the generator that a production team can hand a brief to and get a deliverable back from, which is a different claim and a harder one to price.

The workflow is the actual upgrade

The unglamorous part of the release may be the most consequential: text-to-video, image-to-video, reference generation, video extension, and video editing now sit in one native multimodal workflow. Images, video, and audio references can be combined in a single input field. The generation page supports timeline preview and editing, grid and list layouts, and what Kuaishou calls a canvas Agent experience.

AI video tooling has spent two years being a collection of disconnected steps, generate here, upscale there, edit in a third application, sync audio in a fourth. Consolidating those into one surface is what makes a 30-second timeline practical. It is also where the tool starts to overlap with the editor, which is a different competitive fight than the one against Seedance.

Consolidation carries its own cost. A tool that owns generation, editing, and timeline review is a tool that wants to own the project file, and teams that already have an established post-production stack will have to decide how much of their pipeline to hand over. Kuaishou has been building in this direction for a while with MCP and CLI support for agent-driven batch creation, and the canvas framing suggests the company sees orchestration as the product, not the clip.

Where the bottleneck actually moves

The framing Kuaishou is pushing is that AI video goes from slot machine to storyboard. That is a fair description of what ten keyframes enables. It also moves the bottleneck: the teams that struggled to produce a coherent 30-second piece will now struggle to write a coherent 30-second script. The tool has handed over storyboard control. Whether there is a storyboard to hand it is a separate question, and not one a model release can answer.

There is a version of this that is genuinely good news and a version that is a trap. The good version: teams with a script and a shot list get a production multiplier, because the expensive part of their process (coordinating a shoot) collapses into a generation pass. The trap version: teams without either generate a great deal more content of the same quality, at higher volume, and discover that the constraint was never the render.

Kling 4.0's answer to that is essentially a bet that the market has already made the transition. Ten keyframes, 15 references, and 8,000-token prompts are features for people who have something specific to say. Asking for that much control is only worth it if you know what you want the shot to look like before you press generate.

Related articles