The Video Ad Stack Broke Into Specialists, and Boreal-H3 Shows Why

Creatify Labs launched Boreal-H3 on October 2, a video model built by post-training MiniMax H3 in partnership with MiniMax and aimed at a single job: producing advertising clips. The interesting part sits in the benchmark Creatify invented to sell it, and what that benchmark implies about where video generation is heading.
The pitch, in numbers
Boreal-H3 scores 133.3 on Creatify's Ads Quality Index, where MiniMax H3 is normalized to 100. The comparison set put it ahead of Gemini Omni 1.1 Flash at 122.2, Seedance 2.5 at 118.1, and Wan 3.0 Prime at 111.1.
On reference fidelity (measured at a stricter threshold of 7 out of 10) Boreal-H3 reached 85.3 percent, ahead of MiniMax H3 at 82.4 percent, Wan 3.0 Prime at 79.4 percent, Seedance 2.5 at 71.9 percent, and Omni 1.1 Flash at 70.6 percent.
Against its own base, the post-training loop moved subject consistency from 83.3 to 94.4 percent, raised brief completion from 27.8 to 50.0 percent on a set of difficult production briefs, and cut visible defects from 1.11 to 0.33 per clip, a 70 percent reduction.
The cost and speed figures are where the pitch sharpens. At 768p, Boreal-H3 generates at 1.1 seconds per video-second, so a 10-second clip takes about 11 seconds and costs $0.40. Seedance 2.5 runs at 46.5 seconds per video-second and roughly $0.47 per second. On a 15-second product demo, Boreal-H3 rendered in 84 seconds against 411 seconds for Seedance 2.5.
Why the benchmark matters more than the model
The Ads Quality Index measures something specific: how often a model produces a clip that works as an ad. A clip passes only when prompt following, reference consistency, and visual realism each score 6 out of 10 or higher. Each model's pass rate is then indexed against MiniMax H3.
That is a different question from the one most video benchmarks ask. Artificial Analysis and similar leaderboards rank outputs on general quality and aesthetic appeal. Whether a clip keeps a product label readable, holds the packaging silhouette steady, and preserves the creator's face from first frame to last is not what those rankings measure, and it is what an advertiser checks first.
The honest caveat is that Creatify built the index, runs it, and is selling the model that wins it. The company says its automated judges were validated against human preferences and that human studies used anonymized, randomized outputs. That is a reasonable methodology claim and it is still a first-party claim, from the vendor, about the vendor's product. No independent replication exists yet.
There is also a quieter detail: the comparison includes Seedance 2.5 and Wan 3.0 Prime, both of which are strong general-purpose video models being evaluated on advertising-specific criteria. Losing a specialized benchmark to a specialized model is expected. The number that would be alarming is if Boreal-H3 lost to a general model on its own home turf.
The specialist split is real and it is structural
Where Boreal-H3 fits is more interesting than whether it wins.
Video generation in 2026 has visibly split. On one side are the cinematic generalists (Google's Veo, Kling, Runway's Gen-4.5) selling quality, camera control, and native audio to anyone with a story to tell. On the other are the advertising specialists: Creatify, Arcads, HeyGen, all built around UGC-style creator video and product demos, sold to performance marketers who need volume and iteration speed rather than a hero shot.
The economics explain the split. A performance ad campaign needs dozens of variants to find the hook that works. If each variant costs $5, the math fails. If each costs 40 cents and renders in 11 seconds, you can test a dozen hooks before lunch. Cinematic quality and iteration throughput are different products, and trying to sell both from one model means compromising on whichever the buyer cares about more.
HeyGen took a similar path the same week, launching a universal video model at one cent per second, also post-trained on MiniMax H3. Two companies independently decided that MiniMax's base was the right substrate to specialize, which says something about where the open-weight frontier sits.
What the model still cannot do
The section most vendors skip: product deformation.
A bottle's shape can shift subtly between frames, or across separately generated variants of the same product. Brand marks render as plausible-looking but technically wrong, a logo with the wrong proportions, an extruded wordmark with a letter that does not quite exist. Packaging consistency across generations is common enough that serious workflows lock a reference image to constrain it rather than trusting a text description.
Creatify claims to have addressed this and published a defect-rate reduction. What the claim does not cover is the failure mode that matters most in a paid campaign: a defect that survives review because it looks right at thumbnail size and wrong at full size. Automated judges scoring visual realism at 6 out of 10 will not reliably catch a misdrawn logo. A human comparing output against the physical product will.
The workflow advice that follows is unglamorous. Start every product shot from a photograph of the real product rather than a text description, because a text-only prompt makes the model invent the label along with the object. Write prompts for the motion (what moves, how the camera moves, what the light does) because the photo already carries the product. Put each clip next to the physical object before it ships.
What specialization buys, and what it costs
The case for a model like Boreal-H3 is straightforward: if your output is advertising, a model tuned on advertising will save you iteration cycles that a general model spends on capability you do not need.
The case against is lock-in of a specific kind. A model post-trained on one company's evaluation loop is optimized for that company's definition of a good ad, and that definition is not published in a form anyone else can audit. The Ads Quality Index is a reasonable proxy for "does this work as a commercial." Whether it correlates with "does this campaign perform" is a question that only media buyers can answer, and they will answer it with their budgets over the next two quarters.
There is a second-order effect worth anticipating. When a model is post-trained on a specific commercial objective, it tends to get better at the patterns that objective rewards and worse at the ones it does not measure. An ad-focused model that has been optimized for product fidelity and creator consistency may drift away from the visual inventiveness that makes a campaign memorable. The Ads Quality Index has no component measuring whether a clip is interesting, and that is a deliberate scoping choice with consequences.
The reference-material requirement is the other practical constraint. Keyframe control and multi-reference control both require the advertiser to supply good inputs: a clean product shot, a consistent creator, an established visual style. That is fine for brands that already have a photography library and a design system. For a small seller generating creative from scratch, the specialization is less useful, because the inputs the model needs are the inputs the seller does not have.
What Boreal-H3 really demonstrates is that the video generation market has matured past the point where a single leaderboard settles anything. The relevant question for a buyer in late 2026 is no longer which model is best. It is which model is best at the specific thing you are making, and whether you can measure that yourself rather than trusting the vendor's index.
That last clause is the one that matters most. The vendors building advertising models are also the vendors selling the benchmarks that rank them. The only reliable evaluation is a controlled test on your own product set, with your own review criteria, and a human comparing each output against the physical object. It is slower than reading a leaderboard and considerably more useful.
Related articles
Meta Is Licensing Midjourney's Image and Video Tech, and the Reason Is Telling
Benchmarks move every quarter. A community's judgment about what looks good does not.
AssemblyAI Cut Real-Time Speech Latency to 91 Milliseconds. Here Is Why That Number Matters
The constraint on voice has never been word error rate. It has been turn-taking.
Google's Nano Banana 2.1 Is a Feature Drop, Not a Flagship
A mid-tier release tells you what a lab thinks most of its users actually need, which is a more useful signal than a flagship.
OpenAI Published 722 AI-Derived Mathematics Results. Mathematicians Are Asking Who Gets the Credit.
A proof assistant checks that an argument follows, not that the argument is the one someone thinks it is.