The Weird Economics of AI Chips: New Racks Cut Costs as Old Chips Get More Expensive

Two things happened in the AI hardware world in September, and they seem to point in opposite directions. Nvidia started shipping its newest rack-scale systems to the big cloud providers, promising sharply cheaper inference. At the same time, the price to rent its three-year-old H100 chips went up. Both are true, and together they explain something about how the AI economy actually works.
The New Racks
Nvidia confirmed that its Vera Rubin NVL144 rack systems are shipping in volume, with the first major deployments coming online at Microsoft, Google, Amazon, and Oracle, and CoreWeave saying its first Rubin clusters would be available to customers before the end of September. This is the largest hardware transition the industry has attempted, replacing the Blackwell generation that has powered most frontier model training for the past two years.
The architecture is not a single chip. Vera Rubin pairs a Vera CPU with two Rubin GPUs per node and links them so a whole rack acts like one large processor. The NVL144 configuration connects 144 Rubin GPUs across 72 nodes in one rack, with HBM4 memory for the first time in a production system and roughly 13 terabytes per second of memory bandwidth per GPU. Nvidia claims around 3.6 exaflops of FP4 inference performance per rack, about three times a comparable Blackwell NVL72 rack. Those are ideal-condition numbers, and conservative independent estimates put the real-world gain closer to two to two and a half times.
The number that matters to everyone outside the data center is cost per token. Early cloud pricing suggests Rubin-based instances could undercut Blackwell inference by 30 to 40 percent at launch, with prices typically falling further as capacity ramps. That figure flows downstream to the price of an API call, the size of a subscription, and whether a video generation feature can exist inside a twenty-dollar plan.
The Old Chips That Got More Expensive
Here is the counterintuitive part. In September, the hourly rental price for the H100, the chip that powered the first wave of AI products, rose 22 percent to $3.28 an hour. Nvidia's chief executive pointed to the increase as proof that compute is a durable, revenue-generating asset rather than something that depreciates on schedule.
The reason is supply. New Blackwell and Rubin chips are being reserved for the largest buyers, which leaves older H100 clusters handling the everyday inference workloads. Tight power and memory supplies across data centers keep the older hardware busy and keep rental prices high. The H100 is still far cheaper than it was at launch, when big cloud firms rented it for seven to eight dollars an hour, but the recent rise runs against standard accounting, which writes chip value down over five to six years.
That gap has drawn attention. Short sellers have argued that real chip lifespans are shorter than the official schedules assume, which would mean the depreciation on cloud providers' books is too slow. Nvidia posted record quarterly revenue of 96.2 billion dollars in August, and the rising rental price gives the company a fresh argument that its chips keep earning long after they are new.
The Race to Scale the New Racks
The speed of the rollout says as much as the specifications. CoreWeave became the first cloud provider to run the new hardware at multi-rack scale, bringing up seven Vera Rubin NVL72 racks, 504 GPUs in total, as one production cluster spanning two regions on September 16. The jump from a single validated rack to a multi-rack fabric matters because agentic AI workloads make many small, latency-sensitive calls, and delays compound across repeated model and tool interactions. Each NVL72 rack pairs 72 GPUs with 36 Vera CPUs, and the fabric design supports roughly 128,000 GPUs per rail without redesign, which means adding capacity is a configuration change rather than a rebuild.
The supply pattern also looks different this time. In the Blackwell generation, demand so outstripped supply that smaller cloud providers and national AI programs waited six months or more. Nvidia ramped Rubin production earlier, and several sovereign AI projects in Europe and the Middle East are reportedly in the first shipment wave rather than the third. Nvidia also announced a partnership with Australian cloud and data center operators to build up to two gigawatts of AI factory capacity by 2027. If that holds, renting GPU time gets meaningfully easier, which pulls the whole cost curve down for everyone downstream.
Why Both Things Are True
The apparent contradiction resolves once you separate training from inference. Training needs the newest, biggest clusters, and that demand is what justifies building racks like Vera Rubin. Inference is everything a model does after it is trained: every chatbot reply, every generated image, every coding suggestion. As usage grows, the recurring cost of serving those requests can become larger than the original training bill.
That is why Nvidia is opening part of its ecosystem rather than defending every socket. On September 10, it announced a collaboration with d-Matrix, a startup building chips specifically for inference. Under the deal, d-Matrix's next-generation Raptor chips will connect to Nvidia's rack architecture through NVLink Fusion. Raptor is built for AI inference, especially low-latency work like chatbots, coding assistants, and voice agents, with a memory-centric design meant to cut latency and cost. Nvidia gives up nothing it fully owned, because it still supplies the CPUs, networking, switches, and software around the rack, while gaining a place in the inference market even if it does not win every accelerator.
There is a harder constraint underneath all of this. A single Vera Rubin NVL144 rack draws around 600 kilowatts, roughly the power of 500 average homes. Electricity, not chips, is becoming the binding limit on AI growth, which is why Microsoft, Google, and Amazon have all signed nuclear and geothermal power deals in the past year to feed clusters like these. If you want to know where AI capacity is going, watch power purchase agreements rather than benchmark scores.
What It Means for the Price of AI
For users, the practical read is that the cost of running a model should keep falling, even as the cost of the newest hardware stays high. Cheaper inference makes it possible for features that were too expensive last year to become standard this year. Long, careful chatbot answers, coding assistants that read an entire codebase, and on-demand video generation all depend on this cost curve bending downward.
The question investors keep asking is whether the enormous data center spending will ever pay for itself. Rubin is part of the answer. If the same model gets roughly twice as cheap to run every eighteen months or so, the spending starts to look less like a bubble and more like infrastructure. If that curve stalls, the skeptics get their moment. The next two quarters of cloud pricing will say which way it goes, and that number will reach you long before any benchmark does, in the form of what your AI tools cost.
Related articles
Salesforce Pays $2 Billion for a Company That Interviews Your Customers For You
Interviews are evidence. Digital twins are a prediction. The line between them is the test.
AMD Buys Fei-Fei Li's World Labs for $8.2 Billion to Own Physical AI
AMD is buying a research lab, and paying a chipmaker's price for it.
Google Put a TPU in Orbit and Started Counting the Cost of Space Data Centres
One working chip proves the trip is survivable. It says little about the profit.
Someone Catalogued 13,000 Ways AI Writing Gives Itself Away
The tells did not disappear. They moved somewhere harder to see.