← Back to blog
AiAbout 6 min read

ElevenLabs Solved the Voice and Then Raised the Price of Being Heard

Published Oct 3, 2026
ElevenLabs Solved the Voice and Then Raised the Price of Being Heard

On September 28, ElevenLabs released two speech models. Eleven v4 is built for narration and prepared dialogue. Eleven v4 Turbo is built for live conversation, with a median inference latency near 100 milliseconds and roughly 150 milliseconds to first speech. Two days later, the company closed an employee tender offer that valued it at $22 billion, double its $11 billion mark from February.

Those two announcements are the same announcement. Once a voice model is fast enough that people stop noticing the delay, the remaining question is who sounds real, and that is a product question with a price tag attached.

Latency was the last technical excuse

Voice agents have had a stubborn failure mode for years. The words are correct, the timing is wrong. A pause that runs a beat too long turns a helpful assistant into something users hang up on, and the fix was always more hardware or a shorter sentence. Turbo's numbers sit below the threshold where humans notice. Around 100 milliseconds is roughly the length of a normal conversational gap, which means the delay stops reading as a machine thinking and starts reading as a person considering.

The new architecture also carries over 90 languages, up from more than 70 in v3, with a detail that matters more than the count. A voice cloned from an English recording will speak Japanese with a natural Japanese accent rather than dragging the source accent along. Brands that localize a single spokesperson across markets get a cleaner result, and fiction that depends on a character's accent needs a deliberate decision rather than a default.

Requests now accept up to 10,000 characters, double the previous limit. Inline delivery tags like [laughs], [whispers], and [said angrily] still work and can be stacked. Instant Voice Clones still need about 10 seconds of audio.

The jump from 70-odd languages to over 90 is smaller than the jump in how a cloned voice carries across them. A single recording can now front a product in a dozen markets without each version sounding like the same person straining to imitate a local accent. For a company that localizes one spokesperson, that removes a step that used to require a second recording session per language.

Expressiveness is where the money is

Artificial Analysis ranked v4 first for expressiveness, and ElevenLabs reports that about 75 percent of listeners preferred it in blind tests. Both figures come from the company or its partners, so treat them as marketing until someone reruns the test. The direction is still telling. Every major lab can now produce intelligible speech. The differentiator has moved to whether the speech can carry tone, pacing, and a recognizable personality across a long recording without the voice drifting between chapters.

The commercial picture makes the same point from a different angle. ElevenAgents, the platform for conversational agents, now accounts for 55 percent of ElevenLabs revenue. Voice is not being sold mainly as narration for audiobooks. It is being sold as the interface layer for software that talks.

That shift explains the price structure. Eleven v4 costs $0.08 per 1,000 characters through the API, and Turbo costs $0.04. Both are discounted until October 12, to $0.022 and $0.011 respectively. The reference rates are the ones to budget against, because the launch discount is temporary and corrections add characters. Turbo is also restricted in one specific way that is easy to miss: the Text to Dialogue WebSocket accepts up to ten registered voices per connection on v4, but only one on Turbo.

The rest of the market is not standing still

ElevenLabs is not alone in treating voice as the agent interface. On September 30, Inception Labs launched Mercury Voice, a diffusion language model built specifically for low-latency voice agents, with a median time to first answer token of 320 milliseconds and pricing around nine-tenths of a cent per minute of conversation at launch. Suno shipped a speech model in beta. Vercel added Microsoft's audio models to its AI Gateway with zero data retention.

The approaches differ in where they put the hard part. ElevenLabs treats synthesis as the product and models the words separately. Inception folds the language model into the latency budget, arguing that a fast synthesis layer is useless if the model behind it takes two seconds to think. Both are coherent, and they will keep colliding, because a user experiencing a voice agent cannot tell which component was slow.

The uncomfortable parts

Cloning is guarded by mandatory owner identity verification and filters that block unauthorized clones of prominent public figures. That is a reasonable floor, not a solution. A ten-second sample is now enough for a high-fidelity clone, and the number of people who have ten seconds of someone's audio is effectively everyone.

There is a second-order problem no filter addresses. As expressive models get better at sounding like a specific person, the value of voice as a verification signal collapses. A voiceprint used to be weak evidence and is now close to no evidence at all. Banks and call centers that built identity checks on "we recognize the customer's voice" have been quietly retiring that assumption for two years, and this release makes the retirement overdue.

The latency figure is also narrower than it sounds. It is measured on ElevenLabs' own benchmark protocol with network latency removed. An actual assistant still has to recognize the request, decide on an answer, and deliver it across a network. Turbo removes one source of delay from a chain that has several.

A gentle waveform of light travelling through a curved glass tube on a dark polished surface, teal with warm amber highlights

What the valuation is really saying

Doubling a valuation in eight months, on a $300 million secondary sale rather than new primary capital, is a statement about retention and about a category that has not commoditized yet. Voice models are the rare AI product where users notice quality immediately and switch for it. Hyperscalers and specialized competitors are all shipping, and the quality gap has been narrowing.

The bet behind the $22 billion number is that the gap stays wide enough to charge for, and that the agent platform keeps growing faster than the raw synthesis price falls. Turbo moving to $0.011 per 1,000 characters during launch suggests the second half of that bet is the harder one. When the marginal cost of a voice approaches zero, the thing you can still sell is the voice nobody else has.

There is also a structural reason a voice company can hold a valuation like this while image and video companies struggle to. A voice agent runs continuously, in every call and every session, so usage scales with the product's success rather than with one-off generation. Narration and agents have different economics: one is a project, the other is a meter that never stops. ElevenLabs built toward the second, and 55 percent of revenue says the strategy worked. The question the next year answers is whether the quality edge survives long enough for the meter to keep running.

The company's own framing points the same way. The launch material leads with expressiveness rather than accuracy or speed, because those two stopped being differentiators several releases ago. What a voice agent has to do to feel real is carry hesitation, emphasis, and a consistent identity through a forty-minute session, and that is a harder target than producing a clean paragraph. The model that wins it will not necessarily be the cheapest per character. It will be the one whose output people cannot tell apart from a person they have spoken to before, which is a bar that raises every time it is cleared.

Related articles