Vidu Q4 Preview Makes Reference-Heavy Video Cheap

Shengshu Technology put Vidu Q4 Preview into public preview on October 7, 2026, and the headline number is the price. Launch pricing starts at about a cent and a half per second, which the company describes as up to five times as much output for the same budget under comparable conditions. That is a company claim, and the fine print says the price is promotional and may vary by plan and region. Still, the trajectory is clear: the cost of a second of generated video keeps falling, and it is falling fastest where it is hardest to defend.
Price alone would not be interesting. What makes this release worth a closer look is that the cheapness arrives with controls that used to sit in the expensive tier.
Fifteen references, three audio tracks
Q4 Preview supports up to 15 image references and up to 3 audio references in a single job. Output goes up to 2K or 4K with 10-bit color depth, and the available tiers start at 540p for people who just want to check a composition before committing. Shengshu positions the model around three things: how characters perform, how the camera speaks, and how complex effects hold together.
The reference count is the part creators will notice first. Steering a video model with one reference image is a blunt tool. It can lock a face or a look, but it struggles to hold several things at once. Feeding in many references lets a creator pin a character's face, a costume, a prop, and a setting separately, then ask the model to keep all of them consistent as the shot moves. That is the difference between generating a clip and planning a scene.
Audio references matter for a similar reason. If you can hand the model a voice and a music bed, you get closer to a finished shot in one pass instead of stitching picture and sound together afterward. The company frames this as lowering the barrier for people, teams, and small studios that cannot afford a full production crew.

Why Chinese video labs keep cutting prices
Vidu is not alone in this. Shengshu's move lands in the same window as xAI's cheaper Lite tier and ByteDance's ongoing work on longer single shots. The pattern across Chinese video labs is a race down the price curve, paired with a push on controllability. Two things drive it.
First, inference efficiency has improved enough that a lower price per second can still make sense on the vendor's side. Second, the market that pays top prices is small. A handful of studios and agencies will pay for the best output, but the volume comes from everyone else: e-commerce sellers testing creative, short-drama teams, ad agencies iterating on concepts, educators producing explainer content. That audience is sensitive to price per attempt, because their workflow is built on trying many versions and keeping the one that works.
When Shengshu talks about five times the output for the same budget, it is really talking about attempts. A creator who could afford ten generations yesterday can afford fifty today, and more attempts means a better chance of landing a usable shot. The cost of a rejected take is the hidden line item in every AI video budget, and cheap generations shrink it.
What to check before you switch
A preview label deserves respect. Q4 is described as the first public look at a next-generation flagship, released ahead of the full model, which means the stable version could differ in quality, speed, and price. The consistency and lip-sync claims come from the company's own materials, and independent testing at this scale is not yet public. Anyone planning a production pipeline on top of it should test on their own footage first.
The other thing to watch is what "cheap" does not cover. Reference-heavy jobs cost more to process than single-image ones, and 4K output is not the same product as a 540p draft. Reading a rate card by its lowest number tells you very little. The useful comparison is the cost of a finished, accepted shot after however many attempts it took.
How it stacks up against the rest of the field
The honest comparison is not between price cards. It is between finished shots. Google's Veo 3.1 produces short clips, around four to eight seconds, with native audio, and it has the benefit of a broad ecosystem of integrations. ByteDance's Seedance family has been pushing single-shot length toward half a minute, which matters for anyone telling a story rather than making a loop. Vidu's own angle is different again: a high reference count plus high resolution, aimed at creators who need one shot to look exactly right rather than many shots that look good enough.
For an agency, those differences map onto different jobs. Concept boards and internal reviews want speed and volume, so the cheapest tier wins. A client-facing hero shot wants control and resolution, so reference count and output size matter more than pennies per second. A long narrative scene wants a model that can hold a take together. The mistake is picking one model and forcing every job through it, which is common because switching costs feel high. In practice the cheap preview tiers make it easy to test several models on the same brief before committing.
There is one more thing to weigh, and it is not on the rate card. Switching video models means re-learning how to prompt. A model that leans on reference images wants a different kind of instruction than one that leans on text. Creators who have spent months building a feel for one system pay a real tax when they move, which is why price alone rarely wins a market. The models that hold onto users are the ones that make the reference workflow feel obvious, so the jump from one tool to another doesn't cost a week of relearning.
The real shift is who gets to iterate
For most of the last two years, high-quality video generation was a tool for people with a budget and a reason. The interesting thing about the current wave of price cuts is that it moves the tool toward anyone with an idea and a laptop. That changes the kind of work that gets made. Teams stop rationing generations and start using them the way they use a sketch: cheap, disposable, and fast.
Shengshu's own framing is telling. Its chief executive talks about turning efficiency gains into a lower barrier rather than a higher margin, and about solving real problems in real settings. Whether the preview lives up to that is an open question. But if the price holds and the controls work, the more important outcome is not a cheaper model. It is a bigger number of people who can afford to make the thing they imagined.
Related articles
AI Short Drama Hit the Hot Search, and Live-Action Shoots Fell 70%
Only one of the top 20 titles on a major Chinese drama chart was made with real actors. AI costs about a tenth as much to produce, and it is rewriting the whole pipeline.
Decagon's PACT Protocol Wants Consent to Be a Standard
Decagon open-sourced a protocol for verifying a personal agent's identity and the permissions a customer granted it. It is plumbing, and it decides whether the agent economy works.
The AI Reunion Wave China Can't Decide How to Feel About
AI tribute films brought departed public figures back to Chinese screens and pulled in hundreds of thousands of likes. Then the backlash arrived, and it was about consent.
Apple Published an Open Multimodal Model and Barely Told Anyone
Apple's research-first release slipped past the mainstream, but the model's fine-grained visual grounding says a lot about where its AI stack is heading.