← Back to blog
AiAbout 7 min read

Alibaba Open-Sourced a Model That Cuts a Flat Image Into Editable Layers

Published Oct 6, 2026
Alibaba Open-Sourced a Model That Cuts a Flat Image Into Editable Layers

Alibaba's Qwen team put Qwen-Image-Layered on ModelScope and Hugging Face in early October, with a licence that allows commercial use. The Chinese announcement on October 2 described it as the first image model with native, Photoshop-style layer understanding and editing. That claim is worth unpacking, because the mechanism behind it is not the one most image editors have been building toward.

The standard way to edit a generated image is to select a region and regenerate inside it. Inpainting works this way. FLUX 3 Image, released the day before, works this way too, with bounding boxes that define exactly which pixels the model is allowed to touch. The approach has a known weakness: the model has no idea what a layer is. It sees a rectangular region and a prompt, and it fills the hole.

Qwen-Image-Layered starts from a different representation. It decomposes a finished image into a set of RGBA layers, each full-resolution and each with its own transparency. An object becomes one layer, a piece of text another, the background a third. You can move a layer, recolour it, or swap it out, then recompose the stack. The training objective is what makes this measurable: the model learns to produce layers whose recomposition reproduces the input image, so decomposition quality has a direct score.

Three thin frosted glass panes stacked with a slight offset on a pale stone slab

Why this is a different problem from segmentation

Segmenting an image into objects is an old task, and models have been good at it for a while. Layer decomposition asks for something more specific. The output has to be usable as an editing primitive, which means each layer needs clean alpha edges, correct stacking order, and no double-counting of pixels that appear in two layers at once.

Both of those failure modes are common. A cut-out of a person standing in front of a wall will, on a naive segmentation, leave a ghost of the person behind when the layer is removed. The wall layer has to have learned that the person occludes it, and it has to reconstruct what the wall looked like behind the person. That is a generation problem, not a labelling problem, and it is why the RGBA representation matters. Each layer carries its own alpha channel, so the model has to decide, pixel by pixel, which layer owns what.

The Chinese report credits two pieces of architecture for this: a purpose-built RGBA-VAE encoder, and layer-level 3D positional encoding. The second one is the interesting part. If layers are stacked in depth, the model needs some notion of where each one sits in that stack, which is what the positional encoding supplies.

The "zero drift" claim

The pitch is that edits cause almost no drift: change one layer and everything else stays put.

Drift is the failure everyone who has edited a generated image knows. You ask for a background swap, and the model returns a slightly different face, or a shirt in a different shade, or a shadow that moved. Every diffusion-based edit carries some of this, because the model regenerates more than you asked for. Adobe's new Prompt to Edit feature in Lightroom has exactly this problem, and it is the reason photographers are wary of it.

A layer model sidesteps the issue structurally. If the object you are not editing exists as a separate RGBA layer, and the edit only touches a different layer, the untouched layer gets copied rather than regenerated. Drift becomes a bug in the compositor rather than a property of the model.

That is a cleaner guarantee than any inpainting approach can offer, and it is why the architectural choice is the story here rather than the benchmark numbers.

What it changes for actual work

Three workflows get shorter.

Reuse of existing assets. A product shot embedded in a banner can be pulled out and reused without someone tracing a clipping path. The layer is already there, with alpha.

Transparent output. A model that emits RGBA layers by default can produce transparent PNGs as a side effect of its normal operation. That removes a step that currently requires either a segmentation model or manual masking.

Batch variation. If the master image is a stack of layers, a team can regenerate variations by swapping one layer and recomposing. The rest of the image is untouched, which matters when the brand guidelines fix everything except the product.

A fourth one is less obvious. Layers are structured data. A stack of named RGBA components feeds into an editor or an automated pipeline far more cleanly than a flat bitmap does, which makes the model useful as a component rather than as a destination.

Where the limits are

The model decomposes an image. It does not know what the layers should be called, and it does not infer intent. Ask it to split a photograph of a cluttered desk and the result may separate the objects well and the lighting badly. Semi-transparent materials are the hard case, since a glass of water arguably belongs to several layers at once.

Depth ordering also stays ambiguous for reflections and shadows. A shadow cast on a wall is a property of an object, but it lives on the background layer, and deciding which side of that line to put it on is a judgement the model has to make without being told.

None of this makes the release less significant. It makes it a foundation rather than a finished tool.

Why the layer format travels better than a better inpainter

For the past two years, the image model race has been run on the same course. Each generation produced sharper inpainting, better prompt adherence and fewer obvious artefacts.

That course has limits, and they are structural. No matter how good the inpainter gets, the interaction model stays the same: the user defines a region, the model fills it, and the user judges the result. Quality improves, but the workflow does not change. The evidence is that the same complaint keeps recurring across model generations, which is that a local edit produces changes somewhere else.

Layers attack the workflow instead of the quality metric. Once an image is a stack, the unit of work becomes the layer rather than the selection, and the user manipulates components directly. That is how designers have worked in dedicated tools for thirty years, and the reason image models did not do it was that the representation was not available.

Building that representation is expensive. A model has to learn occlusion, depth ordering and alpha from data that mostly does not contain layer decompositions, which is why the training objective, reconstruction of the input from the predicted stack, is the load-bearing part of the design.

If the approach works, the consequences reach past Alibaba's own products. A compositing tool that accepts a layer stack can be wired into a pipeline that previously needed a human to cut masks. A stock image library that ships decompositions becomes more valuable than one that ships flat JPEGs. The model is the first piece of that chain, and the ecosystem does not have the rest of it yet.

The competitive picture

Layer decomposition is not unique to Alibaba. Ant Group's Ming-Image takes a related position, generating designs as layered compositions rather than flat pictures. Stability and others have shipped segmentation models that approximate parts of the workflow. What differs is the commitment to RGBA layers as the native output format, and the decision to release weights that anyone can run.

That last choice matters for how quickly the idea spreads. The image tooling market has spent two years converging on the "select a region, describe a change" interaction, and the results have plateaued in quality for the simple reason that the underlying representation did not change. Handing the ecosystem an editable stack instead of a bitmap is the kind of shift that shows up in other people's products before it shows up in a benchmark table.

The thing to watch is whether compositing tools start accepting layer stacks as an input format, and whether editors begin treating decomposition as the first step of a project rather than a repair operation. The benchmark score is a side issue.

Related articles