OpenAI Split Its Flagship Image Model in Two. That Says More Than the Benchmark Does.

On September 8, OpenAI shipped ChatGPT Images 2.5, and buried inside the announcement was a decision that matters more than any speed number. The company took the single flagship image model and broke it into two.
For developers, the API now offers GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst. Flare is the fast one. OpenAI says it runs two to four times faster than GPT-Image 2, and early partner Manus backed that up with its own testing, along with better transparent background output. Sunburst is the slow, precise one. It is built for editing that has to keep everything else exactly as it was. Both cost the same: $8 per million input tokens and $30 per million output tokens.
That pricing detail is the interesting part. When two models do different jobs but cost identical money, the vendor is not selling you more compute. It is asking you to decide what you actually need, which is a different kind of product conversation than "here is our best model, use it for everything."

What Actually Changed
On the consumer side, the head line number is latency. OpenAI claims generation delay dropped by up to 50 percent compared with Images 2.0, while sharpness, natural lighting, and texture detail improved. That combination is unusual. Usually you trade one for the other.
The editing story is where Sunburst earns its name. The pitch is that the model understands what not to change. Higgsfield AI, one of the partners OpenAI named, put it that way, and it is a fair summary of what the update is really about. Earlier image editors were good at following an instruction and bad at leaving the rest of the frame alone. You would ask for a new color on a jacket and get a subtly different face, a moved horizon, a shifted light source. Sunburst is aimed squarely at that failure mode, holding faces, products, and brand elements steady across multiple rounds of editing.
There are three new tools in the ChatGPT app. Sketch lets you type @Sketch and draw a rough layout or silhouette inside the chat, treating your doodle as a reference for the final render. Templates give a starting point for common formats like posters and merch. Comments let you pin a note to a spot on an image, so you can say "delete this" or "make this blue" without rewriting the prompt.
OpenAI also leaned on a safety statistic. Sunburst generates unsafe content at a rate of 1.09 percent, Flare at 1.41 percent, and the previous model at 1.64 percent. The direction is good, though these are self-reported and a talking point rather than a guarantee.
The Adobe Move
The same day, Adobe announced it had integrated GPT-Image-2.5 into Firefly. That is not a minor partnership note. Adobe has spent two years repositioning Firefly as an orchestration layer, a place where models from Google, Runway, and others get plugged into a professional workflow. Adding OpenAI's newest image model on day one is that strategy running as designed.
Adobe's Matt Chotin framed it as giving creators flexibility from idea to delivery. What it means in practice is that the fight has moved off raw model quality. Both sides now distribute through each other. More than 70 Adobe tools are embedded inside ChatGPT, so a designer can call Photoshop, Premiere, or Firefly features from inside OpenAI's interface. Each company becomes a channel for the other, and neither can own the start of the creative task alone.
The Numbers Behind the Release
There is a reason this launch reads as a product-strategy move rather than a pure capability jump. When GPT Image 2 shipped in April 2026, it held the top spot on the Artificial Analysis text-to-image leaderboard with an Elo of 1178 and ranked second on the editing board. Images 2.5 is an increment on that base, and as of the days after launch, neither of the two new API models had an entry on the Artificial Analysis leaderboard, and no third party had published a reproducible benchmark. The speed and quality improvements are all self-reported, or reported by the named partners Adobe, Manus, and Higgsfield. That does not make them wrong, but it means any judgment about whether the split delivers should wait for independent testing.
What is not self-reported is the price. Both models sit at eight dollars per million image input tokens and thirty dollars per million output tokens, identical to each other. The output resolution range runs from 1024 by 1024 up to 3840 by 2160, with total pixel counts capped between 655,360 and 8,294,400. For a developer, the decision on a given feature now comes down to a single question: does this workload need to go fast, or does it need the fourth edit to leave the subject untouched? The model that answers that question gets the call, and the fact that both cost the same removes price as an excuse for choosing wrong.
Why the Fork Is the Real Story
Image generation has stopped being an experiment. OpenAI says the platform produces more than three billion images a week, and that number is the tell. At that volume, the interesting question is no longer "can it make a beautiful picture" but "which kind of work is this being asked to do."
Marketing teams pushing out hundreds of social variants care about throughput. A brand team finalizing one campaign key visual cares about whether the logo drifts after the fourth edit. Those two jobs pull in opposite directions, and a single model forces one of them to compromise. Splitting the line is an admission that the customer base has split too.
The response from competitors has already been shaped by this. Midjourney still wins on a particular cinematic polish, but it has no native chat interaction and a steeper learning curve. Adobe Firefly leans on commercial licensing and its software ecosystem, but its standalone conversational creation is thin. Google's in-app image model is strong on some tasks and inconsistent on brand-level control. Nobody has the whole board.
For anyone choosing a tool, the practical takeaway is to stop shopping for a single best image model and start matching the model to the job. High-volume iteration and precision finishing are now separate purchases even when they come from the same company at the same price. The benchmark rankings will keep moving, and they will matter less and less. What matters is whether the workflow around the model can survive the next release, because the model underneath it will be replaced again in a few months.
That last point is worth sitting with. Images 2.5 arrived roughly five months after Images 2, and both replaced a version that was state of the art last year. The teams that build a durable process, with their own prompts, style references, and review steps, are the ones who can absorb that churn. The teams that hardwire a single model into their pipeline inherit every one of its shutdowns and price changes. On that score, the dual-model release is a useful reminder from the loudest vendor in the room: even the leader is telling you not to depend on one model.
Related articles
Step 5 Skipped Four Version Numbers, Landed Second Among Open Models, and Nobody Is Calling It
Benchmark position is a signal producers use to judge a model. Call volume is a signal consumers produce by using it.
Gemini Omni 1.1 Flash Chained Four Requests Into a 40-Second Shot. The Workflow Is the Product.
Google shipped a production pipeline with a review gate built into the pricing.
Microsoft Cut Partial-Transcript Latency to 100 Milliseconds. That Changes What Transcription Is For.
Below 100 milliseconds, a transcription product can respond while someone is still talking.
Synthesia Spent a Decade Recording Avatars. Now It Wants Them to Talk Back.
A training tool people voluntarily repeat is a different product from one they complete because someone assigned it.