← Back to blog
AiAbout 7 min read

A Seventeen-Minute Chinese AI Film Won at Astana, and the Team Was a Dozen People

Published Oct 7, 2026
A Seventeen-Minute Chinese AI Film Won at Astana, and the Team Was a Dozen People

--- title: A Seventeen-Minute Chinese AI Film Won at Astana, and the Team Was a Dozen People slug: a-seventeen-minute-chinese-ai-film-won-at-astana meta_title: A Dozen People, a Month and a Half, One Award meta_description: 2xLabs won best story at the inaugural Astana AI Film Festival with a 17-minute short made by a dozen people and under $14,000 in compute. category: ai tags: AI film,Astana AI Film Festival,2xLabs,The Last Span,generative video,Chinese animation,production cost,storytelling ---

The inaugural Astana AI Film Festival handed its open competition best-story award to 2xLabs, a Chinese team, for a 17-minute short called "The Last Span" (合龙). The film follows a girl, Ayu, crossing a river to find her mother while people on both banks build a bridge that finally closes at the center.

The premise comes from a real object. The painting "Dwelling in the Fuchun Mountains" was burned and separated into two scrolls, held in different collections, then exhibited together again. The film turns that history into a story about reunion.

The production numbers

Executive director Wang Yuchen told China News Service the project took a team of a dozen people about six weeks, covering writing, direction, art, AI visual generation, editing and post. Compute cost came in under 100,000 yuan, roughly $14,000.

The split of that time is the part worth noting. Actual video generation took just over a week. Everything else went into polishing the script and locking the visual style, so that image and story reinforced each other.

The workflow itself is unremarkable in 2026. The team handed styled images and audio to the model, then used prompts to describe shot type, character action and vocal timbre, and generated the footage. When results missed, they went back to re-examine their reading of the scene and adjusted the prompt.

That last detail describes a loop most AI video teams recognize. The failure is rarely the model ignoring the prompt. It is the creator having an imprecise idea of the scene, which the model faithfully reproduces. Wang's account suggests the team treated a bad generation as feedback about their own direction rather than as a tool malfunction.

What the director says is the differentiator

Wang's answer to the quality question is blunt. AI technology keeps updating, so striking visuals are easy to copy. A creator's own thinking and distinctive cultural expression are what give a work a clear character.

The choice of source material supports the argument. The team did not pick a premise because it was easy to generate. It picked one that already carried a structure: two separated halves of a painting, reunited. That gives the film a shape a model cannot supply.

Wang also observed that creators in different countries share a common cinematic language. Using that language well to make a quality film is the precondition for any cross-cultural understanding. He cited a recent Chinese film, Welcome to the Dragon Restaurant, as an example of a work that earned its audience on craft before its message.

Kazakh media covering the festival praised the film's distinctive art direction and tight narrative structure. That is the reception a crew would want regardless of how the footage was made.

The craft decisions that made it hold together

Seventeen minutes is a long time to hold an audience with generated imagery, and the team's choices read as craft decisions rather than technical ones. Limiting the cast to a small number of characters reduces the drift problem that plagues longer AI work, where a face rendered in shot three stops matching the same face in shot thirty. Keeping the setting to a river and its two banks does the same for environments: fewer locations means fewer chances for the world to change between shots.

The bridge itself is a useful narrative device for a generated film. A construction that visibly progresses from shot to shot gives the audience a way to read time passing without dialogue, and it gives the production a natural reason to show the same location repeatedly at different stages. That is the kind of structural choice that plays to what the pipeline handles well.

Sound is the other place the film earns its reception. The team generated audio alongside visuals and described vocal timbre in prompts, which is the workflow that keeps a character's voice stable across scenes. Festival reports noted the music and voice work as strengths, which matters because uneven audio is one of the most common tells in AI video.

The economics of a small AI film

Under 100,000 yuan for a seventeen-minute film is a number that changes conversations with financiers. Traditional 2D animation at that length typically runs into the millions, and live action with a river build and period costuming would not be cheaper. The gap is wide enough that it changes what kinds of stories get told, because a premise that would never justify a seven-figure budget can justify a hundred-thousand-yuan one.

The catch is that compute was not the largest cost. Six weeks of a dozen people's time is the real line item, and it stays the same regardless of the price of generation. That is why the team's report that generation took only a week matters: the expensive part of AI filmmaking is still the part that was always expensive, which is the thinking.

Why this matters more than the festival

Festival awards for AI work are easy to dismiss as novelty. The Astana case is more interesting because the constraints were ordinary. A dozen people, six weeks, a modest compute budget. That is a shape many independent teams can actually attempt.

The film also shows where the effort concentrated. Generation was fast. The slow parts were deciding what the story was and holding the visual style steady across seventeen minutes, which is the persistent difficulty in AI video. Long-form work exposes consistency failures that short clips hide.

There is a translation problem specific to this category of film. A story rooted in a specific historical object has to work for an audience that has never encountered that object. Wang's answer is that quality comes first: the film has to be watchable before its values can travel. Whether that reasoning holds is something the festival jury has already answered once.

The broader context

The announcement landed in the same week that Chinese AI video tooling hit several milestones. Companies in the category are pushing native generation to 30 seconds, adding keyframe control measured in tens of frames, and expanding multimodal references. The 2026 domestic market for AI drama and comic series is projected to pass 40 billion yuan.

Production is getting cheaper in ways that show up in the numbers. One documented setup generated 40 seconds of continuous 1080p footage from scratch on a 16GB GPU in 42 minutes, and a distilled model variant runs 10-second clips on a 12GB consumer card.

Individual creators have proven the same point at larger scale. A history series built by one person accumulated over 100 million plays, with the creator handling writing, direction, editing and art direction alone. Reported projections put Chinese AI drama and comic series markets past 40 billion yuan this year, and platform incentives have followed.

Against that backdrop, a twelve-person team winning an international jury prize is a useful counterweight. Tooling improvements lower the cost of one more iteration. They do not supply a reason to make the film. Wang's summary, that creators' aesthetic judgment, narrative skill and understanding of people increasingly separate one work from another, is the case the Astana award actually supports.

The next test is whether the pattern repeats. A single award can be an outlier. A dozen films made the same way, at the same scale, would be evidence that the barrier to international recognition has genuinely moved.

Related articles