← Back to blog
AiAbout 6 min read

A Photonic Memory Wall Is the $88 Million Bet Volantis Just Made

Published Oct 4, 2026
A Photonic Memory Wall Is the $88 Million Bet Volantis Just Made

The pitch from Volantis is easy to state and hard to believe: build an inference system that makes the memory wall disappear.

The San Francisco semiconductor company closed an $88 million Series A on 1 October, co-led by Lachy Groom and Abstract Ventures, with John Doerr, VXI Capital, Triatomic and Susa Ventures joining. The angel list reads like a who's who of the inference-cost conversation, including Dwarkesh Patel, Naveen Rao and Sholto Douglas. The founding team is drawn from NVIDIA, AMD, Broadcom and Ayar Labs, with prior work on the first CoWoS product, high-volume tunable VCSELs and early co-packaged optics.

The tradeoff everyone has learned to live with

Running a large model fast needs two things that hardware has been unable to supply together. You need enormous capacity to hold the weights, and enormous bandwidth to feed them into the compute engine continuously. Every architecture in production today picks a side.

On-chip SRAM gives you bandwidth without capacity, which caps model size and context windows while raising cost and energy per token. HBM and GPU memory give you capacity without enough bandwidth to serve the largest models at speed. Even newer approaches such as 3D DRAM stay on the same curve, just further along it.

Volantis says its first system, A-1, is designed to lift capacity and bandwidth together by nearly two orders of magnitude. That would let it run models exceeding 20 trillion parameters at up to 10,000 tokens per second per user. Those are design targets, announced by the company, with no independent verification. Treat them as a statement of intent, not a benchmark.

The reason the tradeoff has persisted so long is physics. Bandwidth is limited by how much data can cross the boundary between memory and compute, and that boundary is where most of the energy in a large model actually goes. Moving weights from memory to the arithmetic units costs far more power than the arithmetic itself. Any architecture that raises bandwidth without also raising that cost wins twice.

An optical fiber connector glowing at its core against dark brushed metal

Photonics, pointed at a harder problem

Photonics already moves data inside data centres. What exists today is mostly chip-to-chip. Connecting compute to memory is a different problem, because it demands more than a hundred times as much data to travel over much shorter distances. Energy and cost behave differently at that scale.

Volantis built its platform specifically for that environment. Rather than external lasers, it uses custom micro-VCSELs, leaning on the existing gallium arsenide VCSEL supply chain and sidestepping the indium phosphide constraints that have bottlenecked other optical designs. The company says the resulting links consume less than one picojoule per bit.

The optical fabric pools many memory chips into a single logical resource, so adding memory adds bandwidth rather than just capacity. That is the trick. It is also the part that will decide whether the whole thing works, because pooling memory across an optical interconnect is where latency and fault tolerance get genuinely difficult. A single slow link can hold up an entire pool, and a fault that would take down one memory module now has a much larger blast radius.

Why the timing is not an accident

The market for this pitch has changed shape in the last year. When models answered questions, inference was cheap and bursts were short. As agents take on longer tasks, running for minutes or hours at a time, inference speed stops being a convenience and becomes a throughput multiplier. A coding agent that finishes a task in two minutes instead of thirty does not just feel faster. It changes how many ideas a developer can test in a day.

That shift is why investors who think about inference economics keep showing up on rounds like this. If agent workloads scale with how fast the underlying hardware is, then the hardware stops being a background cost and becomes the thing that sets a company's operating speed. Memory bandwidth, in that world, is not a spec sheet number. It is a line item in how quickly a business can act.

Who this competes with

Volantis is not alone in attacking the inference bottleneck, and the alternatives shape how the bet should be judged. GPU makers keep pushing bandwidth with new memory types and packaging, squeezing the same curve harder. Chip design houses chase specialised accelerators tuned for specific model shapes. A photonic overhaul competes against all of them, and it only wins if it moves the curve rather than the point on it. That is a high bar, which is exactly why the round is notable: serious investors rarely fund curve-movers unless they believe the incumbent approach is close to its limit.

The honest gap

Two caveats belong in any fair reading. The first is that A-1 does not exist yet. Volantis plans to deliver its first integrated inference engines to customers in 2027, which leaves room for a lot to go wrong between a funding round and a shipping product. The second is that the numbers are the company's own, described as design objectives. There is no third-party benchmark and no customer running the system.

That does not make the bet unreasonable. It makes it early. Semiconductor timelines are long, and the money is being spent on engineering headcount and the path to first deployments, not on marketing a product. The $88 million is a runway, not a proof.

The memory wall in plain terms

The memory wall is the gap between how fast a processor can compute and how fast it can fetch the data it computes on. In modern AI, the arithmetic is rarely the bottleneck. Waiting for weights to arrive is. The consequence is that a faster chip, on its own, buys far less than the spec sheet implies, because the chip spends much of its time idle. Every serious effort in inference hardware is, in some form, an attack on that idle time. Volantis is attacking it from the memory side rather than the compute side, which is the less crowded approach and the harder one, because squeezing bandwidth out of the memory-to-compute path means rebuilding the path itself rather than tuning what sits at either end.

What would prove it

The milestone to watch is not an announcement. It is a customer running a model that would not otherwise fit, at a speed that changes what they can do. If the A-1 ships and the memory bandwidth claim survives contact with a real workload, the memory wall stops being a fact of life and becomes a solved problem with a vendor attached.

Until then, the honest summary is that Volantis has raised serious money from people who understand the problem, to attack the single hardest constraint in inference hardware, and has promised a system that would matter a great deal if it arrives. In semiconductors, that is roughly how every important thing starts. It is also how a lot of expensive things end.

Related articles