Google Halved Its Image Prices, and Tripled the Cost of Sending References

Google released Nano Banana 2.1 on October 6. The headline is a price cut: a 1K image falls from 6.70 cents to 3.36, a 4K image from 15.10 cents to 7.56. Batch pricing halves those again, bringing a 1K generation to 1.68 cents. For anyone running image generation at volume, that is the kind of change that alters what is worth building.
The part that gets less attention is on the input side. Google halved output tokens per million, from $60 to $30, and raised input tokens from $0.50 to $1.50 per million. Text and thinking output went from $3 to $7.50 per million. The per-image figure, though, is a fraction of what a customer actually pays.
Why the input change matters
An image request carries far more than an output. It is a prompt plus, increasingly, reference images. Nano Banana 2.1 accepts up to 14 reference images in one request, and each input image consumes roughly 1,120 tokens. A workflow that fuses a dozen references to keep a product consistent across variations is sending a lot of input for every output.
Triple the input price and a heavy-reference workflow can see a saving much smaller than the per-image number suggests. The teams that benefit most from this release are the ones doing text-to-image at scale with short prompts. The teams that will need to model their own numbers first are the ones doing multi-reference editing, because that is where the input side dominates.
Google's own guidance does not hide this. The pricing page lists 3,780 tokens per 4K image, which works out closer to 11 cents than the 7.56 cents circulating in coverage. Run the arithmetic against your own workload before assuming the bill halves.
The practical way to model it is to count images per month at each resolution, multiply by the new per-image output price, then add input cost, which is the average number of reference images per request times tokens per image times the new input price. For a text-to-image workload with short prompts, the saving lands close to the headline. For an edit-heavy workload that sends a dozen references with every call, the input side can absorb most of it. The release rewards one style of work and taxes another, and the pricing page is the only place that distinction is visible.
What actually improved
The capability story is more concrete than the price story. The model holds four characters and ten objects consistent while fusing up to 14 reference images, which is the feature that matters for product catalogues, storyboards and brand campaigns. Subject consistency across editing turns has been one of the more frustrating failure points in image tooling, and this directly targets it.
Text in images improved too. Google's own infographic factuality score jumped from 0.179 to 0.521, and the company says it fixed tiling artifacts that appeared in wide aspect ratios. Output goes to 1K, 2K and 4K, with aspect ratios as wide as 8:1.
Grounding is available through Google Web and Image Search, and thinking levels are configurable at minimal, medium and high. The model also carries C2PA Content Credentials for origin tracking, which becomes more relevant as platforms and regulators push for provenance on generated images.
There are limits worth knowing before porting an agent over. Audio generation, caching, code execution, file search, function calling, Maps grounding, the Live API, structured outputs and URL context are not supported. If your pipeline relies on function calling or structured outputs around the image step, that logic has to live in a separate text model call. The Batch API is supported.
The distribution list is broad from day one: the Gemini app, AI Mode in Search, Google AI Studio, Google Flow, Stitch, Google Ads and the enterprise platform all get the model at launch. That matters because it means the improvements reach consumer products and the developer API at the same time. A developer can ship a feature against a model that millions of people are already using through a different surface, which shortens the feedback loop between a product decision and a user reaction.
One restriction carries over from the rest of the industry. Generating real people is not available on the enterprise platform tier. That will matter to some use cases, and it is a reminder that the capability question and the permission question are separate, even for a major platform vendor.
The migration deadline is softer than it looks
Several write-ups give October 29 as the day the previous model, gemini-3.1-flash-image, stops answering. Google's own changelog says it is deprecated with no shutdown date announced. Anyone scheduling a migration around October 29 is working from a number the vendor is not publishing.
That does not make migration optional. A deprecated model keeps working until the vendor says otherwise, and the cost of staying on it has doubled relative to the replacement. The old model is now a standing 50% surcharge on the same output.
The bigger picture is a price war
Google, OpenAI and the Chinese open-weight labs are all chasing developers who build image features into apps. OpenAI removes gpt-image-1 from its API on October 23 and points users to GPT Image 2.5 Sunburst or Flare. On the independent Arena text-to-image board, Nano Banana 2.1 ranks fifth behind several versions of OpenAI's models.
That ranking is the context the price cut has to be read against. On the independent board, Google trails the quality leaders. It leads on distribution, with the model shipping across consumer surfaces and the developer API on the same day, and it has chosen to compete on price rather than on the top spot in a preference vote. That is a coherent strategy for a company that already has the users. It is also a signal that for the middle of the market, being good enough and cheap beats being best and expensive.
Meanwhile the open-weight options keep closing the gap from below. Models like Qwen-Image and the layer-decomposition releases do not carry a per-image bill at all, which sets a floor no hosted vendor can undercut. The per-image price war is being fought in the space between free open weights and the premium tier, and that space is getting narrower every quarter.
Cheaper output and better text rendering push AI images further into work that used to require a designer for every revision: localised banners, menu cards, simple infographics. That is the direction the price cut points. The question each team has to answer for itself is whether the input side of their own pipeline lets them collect it.
The wider effect is on what counts as a reasonable product feature. When a generated image costs three cents and arrives in a few seconds, a feature that produces one image per user per session becomes affordable in a way it was not last year. Personalised illustrations, per-location marketing assets and generated placeholders all move from expensive ideas to implementation details. Price changes at this level do more than reduce costs on existing work. They make new categories of work possible, because the arithmetic that ruled them out no longer applies.
Google's decision to cut output prices while tripling input prices is a bet about which style of work will dominate. It assumes teams will do more generation and less reference-heavy editing, and it prices accordingly. Teams that do the opposite should test before migrating rather than assume the headline applies to them.
Related articles
Satellite Photos Are Now Robot Training Data
The bottleneck in physical AI training stopped being compute or model capability. It became the quality of the synthetic world.
Google's Gemini 3.5 Live Translate Removes the Pause
Translation that runs continuously, in the speaker's own voice, on a phone already in your pocket, moves the feature from something you open to something simply on.
China Wrote the First Mandatory Safety Standard for AI Agents
Safety moves from a feature you advertise to a gate you pass. The risk inventory sits at 13 categories and 97 items.
AI-Generated Content Now Has to Declare Itself
This step doesn't solve every problem, but it turns “AI-generated” from an option you could hide into a question you have to answer.