← Back to blog
AiAbout 6 min read

SenseTime Wants Images Delivered, Not Rolled

Published Oct 2, 2026
SenseTime Wants Images Delivered, Not Rolled

Ask anyone who uses image models for work what the job is really like, and you will hear about the same friction. You get a picture you like. Then someone asks you to change one word in it, and you start over. Ten attempts later you have a worse result than the one you had, and the file from twenty minutes ago is gone.

SenseTime put its production SenseNova U1 Pro model on its Raccoon app and SenseNova API on 21 September, and the pitch is aimed squarely at that friction. The company calls it a delivery-grade multimodal agent rather than an image model. The distinction it is drawing is between generating a picture and finishing a task.

From rolling to delivering

The workflow U1 Pro describes is longer than a prompt. It starts with understanding the request, then searches for and fills in missing information, processes the material it has, generates and edits, checks its own output, fixes what it finds, and verifies the result before handing it over. The claim is that the checking and correcting happen before delivery rather than after you notice something is wrong.

That is a different product category from a text-to-image endpoint. A text-to-image model answers a prompt and stops. U1 Pro is trying to answer a brief, which is a messier thing. A brief usually implies a set of related outputs, like a poster plus a matching banner plus a version for a vertical screen, and it usually comes with requirements about text that must appear exactly as written.

The model's technical approach follows from that. SenseTime describes an interleaved text-image chain of thought, where reasoning steps and image formation alternate instead of running as separate stages. Layout hierarchy, typography placement, and graphic elements are meant to stay aligned through generation rather than being fixed up afterward. SenseTime also lists output resolution up to 8K with custom aspect ratios, aimed at large-format work instead of square social crops.

The target workloads are infographics, posters, and architectural visualisations. Those share one property: they carry structured information, and getting the structure wrong is worse than getting the pixels slightly off.

The benchmark it is claiming

SenseTime has a leaderboard position to point at. On the SuperCLUE Image text-to-image board for August 2026, SenseNova U1 Pro sits first at 93.81, ahead of GPT Image 2 at 93.23, with ByteDance's Doubao Seedream 5.0 Pro and Alibaba's Qwen Image 3.0 Pro close behind. A Chinese model taking the top spot on a Chinese benchmark is the expected result, and it still tells you something: the gap between the Chinese labs and the American frontier is now small enough that the ordering depends on the benchmark.

What SenseTime has not published is more telling. There is no model card with parameter counts, no licence terms, and no independent benchmark tables. The 8K ceiling and the layout claims are company-stated until third parties reproduce them. That is not unusual for a Chinese enterprise model shipped through a commercial API, and it does mean a buyer is trusting the vendor's numbers rather than testing them.

The API is also gated. SenseNova U1 Pro is available to enterprise customers as a whitelist model, requested by email rather than self-serve. That is how a lot of Chinese enterprise AI reaches the market, and it shapes who can evaluate it. You cannot spin it up on a Friday afternoon and see for yourself.

Two doors into the same model

SenseTime is shipping U1 Pro through two channels with very different characters. The consumer side is Raccoon, its office suite, where the model appears as a one-image-explains-it mode: paste a document or a set of product photos, ask for a poster or an explainer graphic, get a finished frame back. Anyone can try that today.

The enterprise side is the SenseNova API, and it is whitelist-only. Companies request access by email, and SenseTime onboards them one at a time. That gatekeeping is common among Chinese enterprise AI vendors, and it shapes who can evaluate the model. A startup cannot spin it up on a Friday to see whether it fits a pipeline. A large customer with a procurement relationship can, and that customer is the one SenseTime is built to serve.

The split tells you what the company thinks the product is. The consumer door is a demo and a funnel. The enterprise door is where the revenue lives, and the friction at that door is a feature rather than a bug: it lets SenseTime price by customer instead of by token, and it keeps the model out of the hands of anyone who would stress-test the claims in public.

Why the reframing matters

The interesting thing about U1 Pro is not the resolution or the benchmark. It is the claim that the unit of output should be a finished deliverable rather than a generated image. If that claim holds, it changes what a designer's job looks like.

Today, a large part of using an image model well is knowing how to roll. You learn which prompts produce usable output, how to phrase an edit so the model does not wreck the rest of the picture, and when to give up and fix it by hand. That skill is real, and it is also a workaround. It exists because the models are bad at finishing things.

A model that checks and corrects its own work would move that skill from the human to the machine. The designer's job would shift toward specifying what the deliverable needs to be true, rather than steering a generation toward a usable state. That is a bigger change than a benchmark jump, and it is the direction several labs are pushing at once. Alibaba's Qwen-Image-2.1 merged generation, editing, and transparent output into one model for the same reason. Ant's Ming-Image-Layer splits a finished design back into editable pieces for the same reason.

The competition in image generation is no longer about which model draws the prettiest picture. It is about which one leaves the least cleanup for the person who has to ship it.

What to check before trusting it

If you are evaluating U1 Pro, three things matter. Whether the 8K output stays coherent at full size, since a large canvas is where text and layout usually break. Whether the model handles text in the languages you actually publish in, because an infographic is only as good as its smallest legible line. And whether the self-checking loop catches real errors or just smooths over them, which you can only learn by feeding it briefs with known-wrong details and seeing whether it flags them.

Those are the questions a vendor benchmark cannot answer. They are also the questions that decide whether a model goes into a production pipeline or stays in a demo folder.

Related articles