← Back to blog
AiAbout 6 min read

Gemini Omni 1.1 Flash Chained Four Requests Into a 40-Second Shot. The Workflow Is the Product.

Published Oct 5, 2026
Gemini Omni 1.1 Flash Chained Four Requests Into a 40-Second Shot. The Workflow Is the Product.

Google's Gemini Omni 1.1 Flash is a strange product to evaluate, because the model itself is not new. It reached general availability on August 27, 2026 under the id gemini-omni-1.1-flash, and the older preview endpoint was retired on September 30. Everything around the generation changed; the generation quality is largely where it was.

The capability list reads like a production checklist. Scene extension now analyses up to ten seconds of prior context, which Google describes as a jump from previous models that only referenced the final second. Videos extend in ten-second increments to a cumulative ceiling of forty seconds. First-and-last-frame control lets developers specify both ends of a shot and generate the movement between them, aimed at camera orbits, zoom transitions, and loops with no visible seam. A 360p draft mode generates up to 60 percent faster at a third of the cost of standard 720p. Final output upscales to 1080p or 4K. Up to three seconds of video can be supplied as reference material for character and scene consistency.

Read the extension feature carefully, because the marketing rounds it off. A forty-second Omni clip is four chained requests, each using the last ten seconds as context, not a single forty-second pass. Individual generations still land between three and ten seconds. That distinction matters for anyone budgeting latency and inference cost against a deliverable, and it is the kind of detail that decides whether a forty-second shot is practical or merely possible.

A single metal film reel standing upright on a dark table with warm light through it

The pricing is designed to make you iterate

Google's price structure rewards drafting. API pricing runs $1.50 per million input tokens and $17.50 per million video output tokens, billed at 5,792 tokens per second of 720p video, which works out to roughly $0.10 per second or $6 per minute. Per ten-second clip, the tier ladder goes from about $0.30 at 360p, to $1.00 at 720p, to $1.50 at 1080p, to $3.00 at 4K.

That gradient is not accidental. The gap between a 360p draft and a 4K master is an order of magnitude, which pushes teams toward a working method where you iterate cheaply in a low-resolution tier and only pay for the final take. Compare that with the habit most people brought to video generation, which was to write a prompt, wait, look at the result, rewrite the prompt, and pay full price for every roll of the dice. Google is pricing the second behavior out of the workflow.

Compare the surfaces too, because the distribution is uneven. Google frames the full capability set as developer-facing and API-available: five capabilities ship to developers through AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform, while the consumer rollout in the Gemini app is scoped to scene extension only. The same model sits inside Google Flow for AI Plus, Pro, and Ultra subscribers, and it is free inside YouTube Shorts Remix and the YouTube Create app. One model, five access tiers, four of which do not get the interesting features.

The arithmetic of a forty-second shot

Chained extension changes what a shot costs, and the change is not linear.

A single three-to-ten-second generation is cheap and fast. A forty-second sequence is four requests in series, and each one fails independently. If request three produces a take that breaks continuity, you do not regenerate forty seconds. You regenerate from the first frame of the broken segment, then re-chain everything after it, and the cost of a mistake multiplies with position in the sequence. The first ten seconds are cheap to redo. The last ten seconds are expensive, because redoing them means redoing everything downstream of them.

That is the practical argument for the keyframe control. Being able to specify a start and end frame turns each segment into a bounded problem with a verification step at both ends. Instead of judging a whole sequence by watching it and hoping continuity held, a team can check whether the segment landed on the frame it was supposed to land on. For a camera orbit or a zoom transition, that is a far more testable unit of work than a free-form prompt.

The 360p draft tier fits the same logic. Prototyping a forty-second sequence at roughly a dollar instead of six dollars makes iteration affordable, and the upscale to 1080p or 4K becomes a separate, deliberate step. What Google has effectively shipped is a production pipeline with a review gate built into the pricing.

What the benchmarks say now

Gemini Omni 1.1 Flash leads Arena's text-to-video board at 1516, just ahead of its own predecessor Gemini Omni Flash at 1513, and Arena's own rank band for the newer model spans one to four, which is a polite way of saying the two are tied inside the confidence interval.

The leaderboard picture elsewhere is less flattering than the launch coverage implied. The model swept all four Artificial Analysis boards by mid-July 2026. That sweep is over. On the rebuilt text-to-video board, as of early October, Gemini Omni Flash 1.1 sits at eighth with audio and eighth silent, and on Arena's image-to-video board it runs second behind MiniMax H3. Artificial Analysis has not rated 1.1 Flash directly on its text-to-video board, where the older Flash model leads and Alibaba's Wan 3.0 and MiniMax H3 sit close behind.

That is a normal outcome in a fast-moving category, and it does not undercut the workflow additions. It does mean the honest framing of this release is that Google is competing on controllability and price tiers rather than on raw output quality, and the leaderboard is no longer the place where it wins.

Two things the release does not fix

Native synchronized audio is not confirmed. Google's launch materials do not detail soundtrack generation alongside video, and audio output is listed as arriving in later releases. Input audio is limited to voice references at launch. For an any-to-any model family with that name, the missing audio output is the gap that stands out, particularly for teams producing short-form content where sound is half the deliverable.

The second is duration. Forty seconds assembled from four chained requests is a real capability, and it is also a ceiling that a narrative video format hits quickly. Google has teased a higher-end Gemini Omni Pro with no release date, and the extension mechanics suggest the path to longer shots runs through better context handling rather than through a single longer generation.

Every output carries SynthID watermarking and C2PA Content Credentials by default, and the retired preview endpoint no longer appears in Google's model list. Google now recommends Omni 1.1 over the Veo 3.1 preview endpoints, which is a notable internal hand-off. Veo 3.1 remains the current flagship Veo, and no Veo 4 has been announced.

What to actually build with

The useful way to read this release is as a shift in what the product is. A video model that produces a ten-second clip is a novelty generator. A model that analyses ten seconds of prior context, accepts a start and end frame, drafts in 360p for a third of the price, and upscales to 4K is a component you can put inside a timeline.

That is the version of generative video that production teams can plan around, and it is the version where the interesting engineering moves from prompt writing to pipeline design. On this release, the assembly matters more than the clip.

Related articles