← Back to blog
AiAbout 6 min read

Ant Group's Ming-Image Designs in Layers, Not Flat Pictures

Published Oct 2, 2026
Ant Group's Ming-Image Designs in Layers, Not Flat Pictures

Most image models get judged on how good one picture looks. Ant Group shipped two models on 22 September that invite a different test: whether the thing you get back can still be edited when you are done looking at it.

The inclusionAI lab at Ant published Ming-Image-0.1-Design and Ming-Image-0.1-Design-Layer on Hugging Face, both six billion parameters under the MIT licence. That licence matters more than usual here, because it lets a company use the output commercially without negotiating anything first. The pair is aimed at design work rather than photography. Interface screens, infographics, posters, and the sort of layout where the words actually have to be readable.

Small type is where open image models usually fall apart. Ming-Image-0.1-Design generates the whole composition, text included, and it can write RGBA files, so the background comes out transparent instead of a flat white rectangle. The model card suggests 2048 by 2048, or 1024 by 1024 if you want it faster, at 12 sampling steps and a CFG scale of 1.0. Twelve steps is low for a diffusion model. Once the weights have loaded, output comes quickly. It runs on the standard diffusers library and supports vLLM-Omni for serving.

The second model is the interesting one. Ming-Image-0.1-Design-Layer takes a finished, flattened design and splits it back into transparent layers, following a layer plan you hand it that says how many pieces you want. What you end up with is closer to a working design file than a single flat PNG.

A flat graphic design splitting into five transparent panes stacked in three-dimensional space, each holding a different visual element

That gap is the reason most image models stop short of real design work. If a client wants the headline moved and the logo swapped, a flat PNG forces a full regenerate, and everything else on the canvas goes with it. Layers let you change the one thing. Ant evaluates this on the Crello test set, measuring how close each recovered layer is in colour and in shape.

The numbers Ant did not publish

There is no figure you can quote. The model card points at the Artificial Analysis UI/UX Design leaderboard and at the Crello results, but both appear only as chart images. No numbers in the text, no named competitors. You cannot tell from the release how Ming-Image-0.1-Design compares to Qwen-Image, Seedream, or any commercial model, and neither claim can be checked without rerunning the evaluation yourself.

The hardware guidance has a gap too. Ant validated the 6B model on a single CUDA GPU with 80 GiB of VRAM, which is far more headroom than the parameter count suggests. The card gives no guidance for 24 GB consumer cards, even though community quantisations built for exactly those cards were up within hours. INT4, INT8, FP8, GGUF, and ComfyUI builds all appeared the same day.

That last detail is worth sitting with. A model lab publishes an 80 GiB requirement, and the community responds by making the model run on a card a third that size. The official number describes what Ant tested, not what the model needs. Any team weighing adoption should treat the validated configuration as a starting point and check what the quantised builds actually do on their own hardware.

Why layers change the workflow

Design work is not a sequence of finished images. It is a sequence of revisions. A client sees a draft, asks for the headline bigger and the background cooler, then decides the logo should sit on the other side. Each of those requests is small. On a flat image, each one is also expensive, because changing one element means regenerating a composition that was already approved everywhere else.

Layered output turns that around. Once the design exists as separate pieces, a revision touches the piece it is about and leaves the rest alone. For an agency billing by the hour, the difference shows up in margins. For a solo designer, it shows up in whether the tool is faster than doing it by hand.

The other place layers pay off is handoff. A flattened PNG is a dead end. Anyone downstream who needs to adjust it starts over. A layered file can move into the tools a design team already uses, which is the difference between an AI tool that sits beside a workflow and one that joins it.

The editing race this belongs to

Ant is not shipping into empty space. Across the image market, editing has become the place models separate. GPT Image 2.5 Sunburst sits at the top of the editing leaderboards, with its Flare sibling close behind. Nano Banana Pro remains the easy route to identity lock without fine-tuning, and it blends up to fourteen images and five people. Seedream 5.0 pairs editing with live web search, which helps when an edit needs a real reference. On the open side, FLUX.1 Kontext handles layout-preserving edits you run yourself, and Qwen-Image-Edit leads the open-weight editing ranks.

Against that field, Ming-Image is making a narrow bet. It is not trying to win the general editing arena. It is betting that text-rich design is a big enough slice of the market to support a specialised model, and that layer decomposition is a feature the general models will not bother to build because it does not show up in a demo reel.

That bet has a history behind it. Open image models have spent two years chasing photorealism, and the quality gap between open and closed weights has closed a lot. What has not closed is the gap between a model that produces a picture and a model that produces an asset. A picture is done when it looks right. An asset is done when it can be handed to someone else and changed. Ming-Image is aimed at the second category, which is a harder thing to demonstrate in a launch post and a more useful thing to have on a Tuesday afternoon.

Who this is actually for

If you produce design assets at volume and the text inside them has to be legible, this is worth testing now. A small marketing team turning out social cards, ad variants, and internal decks is the clearest fit, and the MIT licence puts nothing in the terms between you and commercial use.

The layered model is where I would start, because layered output is the thing your existing image tool almost certainly cannot do. Everything else here, a text-to-image model that renders readable type, has competition. The ability to take a finished flat design and hand back an editable file does not.

The first test worth running is text rendering at small sizes in a language other than English. That is where models in this class usually fail, and the release says nothing about it. Multilingual typography is the hidden requirement behind most real design work, and a model that only handles Latin script well is a model with a ceiling.

If your output is photographic rather than typographic, none of this applies. Ming-Image is not trying to compete with the models that render beautiful people and landscapes. It is aiming at the unglamorous half of design, the part where a poster needs a headline in the right place and a product card needs a price that a human can read. That half has been waiting a while for a model that takes it seriously.

Related articles