← Back to blog
AiAbout 6 min read

Two Voice Models Just Reset the Bar: 50ms to First Audio, and a 99M Model on a Laptop CPU

Published Oct 6, 2026
Two Voice Models Just Reset the Bar: 50ms to First Audio, and a 99M Model on a Laptop CPU

Two speech synthesis releases in the same week point at the same conclusion from opposite directions. One chased latency to the edge of what is perceptible. The other chased size down to what a laptop can hold in memory. Together they mark how quickly the text-to-speech race has moved from sounding human to fitting somewhere specific.

The first is from Gradium, a voice startup that announced a TTS model with roughly 50 milliseconds to first audio. The company says that is the lowest latency among frontier TTS models. The second is Supertonic, an open-source model with 99 million parameters that runs locally on a laptop CPU, produces studio-quality audio in under a second, and, in the company's testing, correctly pronounced phone numbers and other hard text where ElevenLabs, OpenAI, and Gemini TTS all failed on the same input.

Why first-audio latency is the metric

For a voice assistant, the number that matters is how long before the first sound reaches the user, rather than how long the full clip takes to render. A conversational agent that pauses a third of a second before speaking already feels slow; one that manages 50 milliseconds keeps the rhythm of an ordinary back-and-forth. That threshold is why the frontier pricing tiers, the interruption handling in real-time voice models, and the latency claims all cluster around the same few tenths of a second.

A translucent glass sphere on a dark mirror surface with concentric rings of light

Smaller models reach that range more easily, which is the honest trade. A 99M-parameter model that runs on a CPU gives up the expressiveness of a frontier cloud model and gains the things that model cannot offer: no network round trip, no per-call cost, no audio leaving the device, and predictable behavior on a machine that is already sitting on the desk.

The hard text problem is a dictionary problem

The Supertonic result about phone numbers points at a quieter failure. Synthesis models handle ordinary sentences well and stumble on strings where the correct reading depends on context or convention. Phone numbers, product names, addresses, and abbreviations routinely come out mangled, which is why they are the classic failure case in a demo.

A third release approaches this from a different angle. Onepin positions itself not as a model but as a quality-assurance step after synthesis. It checks every generated line against a pronunciation dictionary of about four million words, covering names and products, scores naturalness and accuracy, and fixes a single mispronounced word without regenerating the whole line. For teams producing long audio, being able to repair one word instead of rerendering a five-minute segment is a real change in the cost curve.

The dictionary approach also separates two jobs that had been fused. The model handles prosody, emotion, and flow. The checker handles the finite set of proper nouns that a general model was never going to pronounce reliably. That split is more useful than a marginally better encoder.

A fourth model shows the streaming bar

Sopro V2 Turbo, a 120M model, is being prepared with a build that targets roughness and break-up on cloned voices. It reports about 300 milliseconds to first audio on a laptop CPU, true streaming, Apache 2.0 licensing, and coverage of English, European Portuguese, French, and German. Its author is candid that thin or cartoon voices, noisy reference clips, and some out-of-distribution speakers still fail, and has been collecting those examples.

Read across all four and the frontier has split into distinct lanes. Cloud models keep pushing expressiveness and multilingual coverage. Small models own the on-device lane, where latency, privacy, and zero marginal cost matter more than the last few percent of naturalness. And a middleware layer is emerging to catch the errors that neither lane handles well on its own.

What to measure before you commit

Latency and size claims in this category are almost always vendor-reported, so the practical test is your own input. Run the model on the text you actually need, including the names and numbers that break demos. Measure time to first audio on the hardware you will ship on, not on the vendor's benchmark machine. And check whether the license allows the deployment you have in mind, since for local models the license is often the deciding constraint.

The pattern across these releases is familiar from image models a year ago: quality converged, then the competition moved to where the model runs, how fast it starts, and what it costs per call. Voice is now in that phase. The model that sounds best matters less than the one that starts speaking before the user notices the pause.

Where the frontier voice models are looking instead

The latency race is happening at the same time as a different push, and the two can be confused. Frontier real-time models, such as the reasoning-capable speech systems that can be interrupted and can call tools mid-conversation, are chasing the ability to hold a conversation that feels like talking to a person on a phone call. That requires understanding intent, deciding when to speak, and handling being cut off, which are hard problems on a different axis from raw synthesis speed.

For most applications, though, the advanced conversational layer is overkill. A voiceover for a video, a narrated product walkthrough, or a bedtime story needs good pronunciation, consistent style across many takes, and a predictable cost. Those needs are served better by a small, fast model than by a frontier system with a chat loop attached, which is why on-device synthesis is becoming the default for pre-recorded audio while the cloud stays the home of live conversation.

There is also a regulatory thread pulling at the same area. Marking synthetic audio so listeners and platforms can tell it apart is becoming a requirement in more jurisdictions, and voice cloning has drawn its own wave of court rulings. A local model that never uploads a sample sidesteps some of that exposure, while a cloud model that clones a voice inherits it. For teams shipping voice features, the choice of where synthesis runs is becoming a compliance decision as much as a latency one.

A practical checklist

Test candidates on your own scripts, not on a vendor's demo text, and include the names, acronyms, and numerals that break general models. Measure time to first audio on the hardware you will actually use, since a laptop CPU result and a data center result are not comparable. Confirm the license covers commercial use and the deployment shape you want, because for open models the license is often the constraint that decides the choice. And decide early whether you need a QA pass after synthesis, since a model that pronounces 95% of words correctly will still fail on the one brand name that matters to your client.

None of these four releases is a breakthrough on its own. Read together, they show a field that has stopped arguing about whether synthetic speech sounds real and started competing on the terms that production teams actually care about: how fast, how small, and how cheap. That is what a maturing tool looks like.

Related articles