← Back to blog
AiAbout 6 min read

FLUX 3 Image Turned Image Generation Into a Layout File

Published Oct 3, 2026
FLUX 3 Image Turned Image Generation Into a Layout File

Black Forest Labs shipped FLUX 3 Image on October 1, and most of the coverage led with resolution: native output up to 4K, roughly 16.8 megapixels, no upscaler in the loop. That is a real upgrade. It is also the least interesting thing about the release.

The part worth reading the docs for is the input format. FLUX 3 Image does not want a paragraph and a wish. It wants a table.

Every element gets an address

A generation request is a JSON array of elements. Each element carries a unique ID, a bounding box written as `[y_min, x_min, y_max, x_max]` on a normalized 0 to 1000 grid, and a plain-language description of what belongs there. A background block spans the whole canvas. A product sits at `[400, 200, 950, 800]`. A logo lives in the top left corner. The scene prompt still exists, but it is now one input among several rather than the whole brief.

The grid is resolution independent. The same layout JSON renders the same composition at 768 pixels wide or at native 4K, which means a layout you tested on a cheap draft is the layout you get on the expensive final.

That sounds like a convenience feature. In practice it changes who can call the model. BFL designed the element table so a language model can write it from a one-line brief and a target aspect ratio. An agent does not have to translate intent into prose and hope for the best. It can lay out the image the way it lays out a spreadsheet, and the element IDs survive as handles it can reference in later turns.

A blank sheet of paper taped into an even grid on a wooden desk, with a few small abstract objects resting in some of the squares

There is a cheaper consequence too. Because the grid is resolution independent, a layout can be validated at 768 pixels for a fraction of the price and then rendered at 4K with the composition intact. Iteration and final delivery stop fighting each other over budget, which matters to anyone producing dozens of images against the same template.

Pixel locking and the drift problem

Editing follows the same logic. You specify a source bounding box and a target bounding box, pick the element to move, replace, or remove, and the pixels outside that region are supposed to stay exactly where they were. BFL's own measurement puts untouched areas at 67.8 to 89.7 percent bit-identical fidelity, a range that depends on how complex the edit is.

The honest part is the wording. The launch demo promises edits "without changing any other pixel." The API documentation says untouched pixels "generally" stay unchanged, and BFL acknowledges that edits can extend past the specified box. Boundary artifacts are real. Any team putting this into a pipeline should build an acceptance test rather than trust the demo.

Still, the comparison to ordinary inpainting is not close. Traditional inpainting leaves subtle drift even in regions it was told to leave alone. Over one edit, nobody notices. Over a ten-step approval cycle, the background has shifted, the typography has softened, and the composition no longer matches what a client signed off on three rounds ago. That accumulation is what makes long editing chains unusable, and it is the specific thing bounding boxes are meant to stop.

What creators actually found

A creator who ran the same editing instructions across Midjourney, Ideogram 4.5, and FLUX 3 Image posted a useful comparison. On a test that asked for a garment recolored to pink silk with white flowers and a custard yellow bow, then for macro close-ups of the face and hands, Midjourney handled the close-ups but missed the white flowers and the exact shade. Ideogram 4.5 matched the garment changes, though the lighting in the macro shots looked off. FLUX 3 Image matched the garment changes and drew praise for lighting and context in the close-ups. In a separate test, it kept a tiny-column edit confined to the intended objects while the other two spread the change further.

One creator's thread is not a benchmark. Treat it as a signal about where the model is strong, which appears to be fine-detail instruction following rather than broad stylization.

What bounding boxes do not fix

Placement and drift are two problems, not all of them. A box tells the model where something goes. It does not tell the model what the thing should look like, and it does not stop a reference element from changing style once it is placed. Teams that generate a single product against thirty different backgrounds still have to keep the product itself consistent, and that work lives in the reference images and the prompt, not in the coordinates.

Text rendering is the other open seam. BFL's own demo images come from a prompt that a human curated. Whether a layout model reliably places legible type inside a small box, in a specific font, is a separate question, and the launch material does not answer it. For magazine-layout and e-commerce work, which is the audience this release is aimed at, that is the detail that decides adoption.

Nor does a structured input make the model any less of a black box about aesthetics. You can specify that a logo goes in the top left corner. You cannot specify that the composition should feel calm.

The trade-offs nobody put in the headline

Bounding boxes cost more setup than typing "make the background blue." You have to know where things go. For a one-off image, that overhead probably is not worth it. For a pipeline that regenerates the same layout across twenty product photos, or an agent that has to place a logo somewhere reliable, it pays back fast.

Launch pricing runs 50 percent off through October 8: about $0.0205 for a 768-pixel image up to $0.3035 for 4K, with references and prompts included and failed generations not charged. Commercial self-hosted weights are available by contacting BFL, and an open-weights edition is promised in the coming weeks.

The reference handling is worth a note on its own. A single call accepts up to 10 reference images, and each element can be told which reference it comes from and where it belongs. Ten references alone is not unusual. Ten references that each map to a specific box on the canvas is a workflow for assembling a composition out of existing assets, which is closer to how a designer works than to how a prompt works. Combine that with the layout grid and the model starts to look less like a painter and more like a compositor that has opinions.

The bigger picture is that image APIs have been optimized for years around one question: how good does the picture look? FLUX 3 Image adds a second question that agents care about more. Can you address one part of the picture, change it, and leave everything else alone? Once the answer is yes, image generation stops being a slot machine and starts behaving like a file you can edit.

That distinction is why the comparison to Midjourney and Ideogram matters less than it first appears. Those models are competing on how beautiful a fresh image is, and they remain very good at that. BFL is competing on whether a generated image can be treated as a document with stable regions, and that is a different contest with a different winner. The first contest is close to settled and the second one just started, which is roughly where an infrastructure bet wants to be.

Related articles