← Back to blog
AiAbout 6 min read

Kandinsky 6.0 Open-Sources Video Weights That Generate Sound in the Same Pass

Published Oct 6, 2026
Kandinsky 6.0 Open-Sources Video Weights That Generate Sound in the Same Pass

--- title: Kandinsky 6.0 Open-Sources Video Weights That Generate Sound in the Same Pass meta_title: Kandinsky 6.0 Open Sources Video With Native Audio meta_description: Kandinsky Lab released 3B and 29B video models that generate synchronized audio in one pass, under an MIT license, with day-zero vLLM-Omni support. ---

Most open text-to-video systems treat sound as somebody else's problem. You generate silent footage, then bolt on a soundtrack afterward, and the lip movement rarely lands where the dialogue does. Kandinsky Lab's answer, released this week on fal, is to model both at once.

Kandinsky 6.0 Video ships in two sizes: a 29B Pro model aimed at cinematic output and a 3B Lite version for quick iteration. Code and weights are out under an MIT license, which matters more than the parameter counts for anyone planning commercial use.

What the two tiers actually deliver

Both models target short-form generation. Clips run up to roughly five seconds, and audio comes out at a 44kHz sample rate. Pro is the quality tier; Lite is positioned for teams that need to iterate quickly and are willing to trade fidelity for speed and cost.

The headline capability is joint generation. Language, vision, and audio are modeled together rather than in a pipeline, so footsteps, ambient effects, and mouth shapes line up with what is happening on screen. That kind of native synchronization has mostly been the domain of closed commercial systems, which is what makes an open release notable.

Overlapping translucent ripples in dark water lit by warm amber and cool teal light

vLLM-Omni provides day-zero inference support. That matters beyond convenience. When an inference stack ships support on release day rather than weeks later, teams can evaluate a model against their own workload before the initial enthusiasm fades.

On fal, the model is bundled with integrated upscaling and a standalone video super-resolution tool. So the practical pipeline is text or image in, five-second clip out, then upscale.

The licensing needs a careful read

Here is the part that deserves attention. Coverage of the release describes it as MIT, and Kandinsky Lab's own announcement says code and weights are provided under an MIT license. An independent model tracker, though, lists the licensing as custom rather than standard permissive.

That discrepancy is worth resolving before anyone builds a product on top of these weights. A genuinely permissive license means commercial use, modification, and redistribution without a separate agreement. A custom license can look equally open at a glance and still carry restrictions that surface later, often around territory, scale, or downstream redistribution.

The pattern this year has been for Chinese labs to publish open weights with permissive-sounding names and carve-outs buried in the terms. MiniMax's H3 license, for instance, excludes the United States, the European Union, the United Kingdom, and South Korea. Anyone moving from a model card to a deployment plan should read the actual terms.

MIT-licensed video weights would be a meaningful step if they hold up as described. Permissive video licenses are rarer than permissive language model licenses, partly because video models have more complex provenance questions around training data and partly because the labs that produce them are more often commercial. A genuinely open video model that also handles audio would give smaller studios a path that does not require an API contract.

Why synchronized audio is the harder problem

Generating plausible motion from a text prompt is solved enough that the interesting work has moved elsewhere. The unsatisfying part of open video generation has been audio.

Consider what silent generation pushes onto the user. A character speaks, but the jaw moves on its own schedule. Someone walks, but the footsteps have to be added in post, frame-matched by hand. A door closes, but there is no sound, and dropping in a stock effect rarely lands on the right frame. These are small problems individually and a steady tax collectively, because each one requires an editor's attention.

Joint modeling changes the shape of the task. The model handles timing internally, so the output arrives as something closer to a finished clip. For short-form content, social ads, and character animation, that is the difference between a demo and a deliverable.

Whether Pro actually achieves tight synchronization is a separate question from whether the architecture supports it, and the independent evaluations that would settle it are not in yet. The release makes the claim and provides demonstrations; treating those as proof would be premature.

Where it sits in the open video cohort

Short five-second clips place Kandinsky 6.0 squarely in the current open-source video generation field rather than ahead of it. Alibaba's Wan 3.0 holds the top spot on Artificial Analysis's rebuilt text-to-video board with an Elo of 1157, and the open-weight options worth running yourself, notably MiniMax H3, are already established.

So the fair framing is competitive rather than revolutionary. What Kandinsky 6.0 adds is native audio at an open license and a size range that spans a serious quality tier and a lightweight one. The 3B Lite model, in particular, opens the door to teams that could never justify hosting a 29B diffusion model.

The two-tier structure also gives teams a tradeoff to reason about explicitly. Run Lite while iterating on prompts and storyboards, where a five-second clip needs to exist but does not need to be beautiful. Move to Pro once the sequence is locked. That is a workflow shape the closed providers have supported for a while and open weights have largely lacked.

What to watch

Three things would tell you whether this release matters in six months.

First, the license terms. If they turn out genuinely permissive, the model becomes a default option for commercial open video work. If the restrictions are meaningful, adoption will concentrate in territories where they do not bite.

Second, independent benchmarks. Artificial Analysis is rebuilding its image-to-video board on a new scale, and every Elo moves when it lands. Until Kandinsky 6.0 appears on a board like that, the quality claims rest on vendor demonstrations. The current text-to-video leaderboard has Wan 3.0 at the top with an Elo of 1157 and a cost of $12 per minute, with 10 of 20 category wins, which gives a sense of how crowded the top of the field is. Kandinsky 6.0 is not competing at that level, and the honest question is whether it needs to.

Third, whether the audio holds up under scrutiny. Synchronization is easy to show in a five-second clip of a person talking to camera, and hard to show in a crowded scene with overlapping sound. The demos published so far are the easy version.

One more consideration applies to any team planning to build on this. Video generation models are expensive to run, and a 29B diffusion model is a serious deployment. Renting capacity through fal removes that burden but also removes the main advantage of open weights, which is running them yourself. The 3B Lite tier is the more interesting option for most teams, because it is the one that can plausibly run on hardware a small studio already owns.

For now, the interesting fact is structural. An open model that produces picture and sound together, under a license that may or may not be fully permissive, with inference support arriving the same day. The gap between open and closed video generation narrowed this week, and the licensing question will decide how much of that gap stays closed.

Related articles