The $1.50 House: Precision Editing Turns Into a Production Line

A demo making the rounds combines two models into something that looks less like an AI art trick and more like a factory. It wires a segmentation model called SAM 3.1 to Ideogram 4.5, an editing model that shipped on September 30, and uses the pair to restyle entire real-estate listings. The reported cost was about $1.50 per house, with individual edits landing near six cents.

Whether that exact number holds across a real brokerage's image set, the shape of the claim is the interesting part. Image editing has crossed from "generate something new" to "change one thing and leave everything else alone," and that shift changes what a creative workflow costs.
Why "change one thing" is the hard part
Ask a model to invent a scene and it has enormous freedom. Ask it to change a jacket's color and keep the body, the pose, the background, the lighting and the fabric texture exactly as they were, and the problem gets harder, not easier. Most editing models drift. Every pass adds a small error, and after several edits the image has degraded: colors shift, textures smear, edges soften. Artists call it artifact buildup, and it is the reason multi-step editing has been more of a promise than a habit.
Ideogram 4.5 is built specifically against that drift. The company's pitch is that it touches only the region you designate and preserves the pixels elsewhere. It ships in four quality tiers between 0.8 and 22 cents per image at native 2K resolution, available on Ideogram's own platform, its API, and launch partners including Runway, Pika and Leonardo AI. An open-weight release is promised but not yet out.
The headline capability for commercial work is what Ideogram calls zoom editing. You can crop into a large, high-resolution image, edit just that region, and stitch the result back into the original without downscaling the whole file first. The company's demo uses a source around 4,016 by 6,016 pixels, a little over 24 megapixels. For a product photographer fixing a reflection in one corner of a large-format shot, or an interior image where a single piece of furniture needs replacing, that is the difference between a five-minute change and a full regeneration.
How the pipeline is assembled
The real-estate demo matters because it shows how these pieces combine. Segmentation identifies what is in a room, the sofa, the walls, the windows, the floor, and produces masks. An editing model then acts inside those masks. What took a human retoucher an afternoon of masking and cloning can be scripted into a sequence: segment the furniture, replace it with a staged version in a chosen style, keep the room's architecture and light exactly as they were.
Run that across a vacant property's photo set and you have virtual staging, the practice of furnishing an empty home in images so that buyers can picture living there. The economics do the work here. Physical staging requires moving furniture into a home, insuring it, and paying for labor and storage. Virtual staging was already cheaper and faster, and the new generation of editing models pushes the per-image cost down far enough to make it routine rather than a premium add-on.
The savings are real but easy to overstate. Current market pricing for virtual staging runs from roughly nothing for occasional free tools up to several hundred dollars a month for professional plans, and some tools charge per room or per export. When a low per-edit price is quoted, the number that matters is the cost per acceptable image, not per generation. Anyone doing this at volume budgets extra attempts, in the neighborhood of twenty percent, because not every output passes review.
The part vendors leave out
There is a compliance layer that the sleek demos tend to skip. Images that advertise a real property are subject to listing rules and, in some places, disclosure requirements. Industry guidance has moved toward allowing AI-created listing images only when the representation clearly discloses that AI was used. A virtual staging image is meant to help a buyer understand a space, not to misrepresent it.
That puts a limit on how far the automation can go. A model that replaces furniture is fine. A model that widens a room, raises a ceiling, or invents a window that does not exist is producing a misleading listing, and no amount of cost savings justifies that. Human verification of every architectural feature remains the step that decides whether an image gets published, and the final package on a real listing should carry the source files, a record of AI use, and any required disclosure.
Product photography is the quieter market
Real estate is the flashy example because the cost comparison is dramatic, but the same precision-editing pipeline applies wherever a product image has to meet a brand's rules on a schedule. Online sellers shoot thousands of items, and each one needs a clean background, consistent lighting, correct colors and often a small correction: removing a scuff, straightening a label, swapping a colorway. That work has historically gone to retouchers or to a studio reshoot, and both are slow relative to the pace of a catalog.
A model that changes only the designated region fits this workflow better than a generative model that redraws the whole frame. If the goal is to fix a reflection in the corner of a large image, redrawing the whole thing risks altering the product itself, which is unacceptable when the image is what a buyer uses to judge what they are ordering. Precision editing keeps the parts that must stay true and touches only the part that needs to change.
The same reasoning extends to two adjacent jobs that used to require specialized tools. One is editing text that is baked into an image, a price tag or a sign, without disturbing the pixels around it. The other is photo restoration, where the point is to repair damage while preserving the authenticity of the original. Both are cases where a broad generative edit would destroy the thing that gives the image its value, and a targeted edit is the only acceptable answer.
The role of the human in this workflow does not disappear, it moves. Instead of doing the masking and the cloning by hand, the retoucher defines what should change and then judges whether the result is faithful. As the cost of each edit falls toward a few cents, the bottleneck shifts from production to review, and the people who can tell a good edit from a subtly wrong one become more valuable, not less.
The line between tool and pipeline
What makes this moment different from the last few years of image editing is that precision editing crossed the threshold where it can be trusted inside a longer workflow. When a model reliably changes one thing and leaves the rest, it stops being a novelty you fire and forget and becomes a step you can chain. Segment, edit, verify, publish. Each stage is automatable, and the output of one feeds the next.
That is why the $1.50 number, wherever it lands exactly, is worth noticing. It is the price of a completed task rather than the price of a picture, in a workflow where the image is a means to an end. The teams that win here will not be the ones chasing the lowest generation cost. They will be the ones who build the verification and disclosure steps around the model so that a cheap edit can be shipped without becoming a liability.
Related articles
A Wheeled Semi-Humanoid Finished an Hour of Laundry Without Help
Individual tasks can succeed while a workflow still fails. Dyna changed the metric.
LTX 2.5 Wants to Render Your Blocky Blender Draft Into a Finished Shot
You do not control what happens in text-to-video. This tries to fix that.
ServiceNow Turns Agent Failures Into Training Data
Generation without verification is noise. The gates are the product.
One Framework for Language and Vision: Horizon's 1.6B Open Model
A bet that the bridges between language and vision were never needed.