← Back to blog
AiAbout 6 min read

Alibaba Split the Voice Agent in Two, Then Cut the Price Up to 95 Percent

Published Oct 2, 2026
Alibaba Split the Voice Agent in Two, Then Cut the Price Up to 95 Percent

Alibaba's Qwen team released Qwen-Audio-3.1 in the last week of September 2026, a five-model audio stack covering speech recognition, text to speech and realtime interaction. The numbers that travelled fastest were the price cuts: roughly 85 percent off the realtime model, about 70 percent off text to speech, and up to 95 percent off automatic speech recognition.

The numbers matter, because they land in a week when the whole industry was cutting prices. Microsoft, OpenAI and Anthropic all shipped cheaper or rebalanced models in the same stretch, and Alibaba's move reads as a deliberate escalation in audio specifically. But the more interesting thing about Qwen-Audio-3.1 is not the discount. It is how the flagship model is built.

Two models behind one voice

The headline model is Qwen-Audio-3.1-Realtime, a full-duplex speech system built for voice agents that need to call tools in the middle of a conversation.

Its central design decision is a split. Instead of one large end-to-end system that listens and speaks, Qwen runs two models with a shared audio encoder and language-model backbone. A decision model predicts whether to keep listening, start speaking, stop, or resume. A separate generative model handles the actual speech. In plain terms, one model decides when to talk and the other decides what to say.

That split is a practical answer to a specific failure. Voice agents feel broken most often not because they misunderstand you, but because they misjudge the moment. They talk over you, or they go silent, or they answer before you finish. Treatment has usually been to bolt turn-taking onto a bigger model and hope the behaviour emerges. Qwen is treating "when to speak" as its own engineering problem, with its own model, separate from content generation. The company reports replies aimed at the wrong speaker dropping from 0.13 to 0.03 on one benchmark, and a filler-word rate falling from 0.759 to 0.296 on another. It also reports a trade-off, with the unwanted resume rate after an interruption rising from 0.035 to 0.130, which is the kind of disclosure released notes rarely include.

The model is available as a managed API. Qwen-Audio-3.1-Realtime-Plus runs on QwenCloud over WebSocket. There are no open weights, which matters for anyone tracking Qwen's open-image work, because the audio stack is being run as a closed, managed product.

The specs and the pricing

The realtime endpoint lists text and audio as both input and output. Context runs to 262K tokens, with 245K maximum input and 16K maximum output. Default throughput limits are 60 requests and 100,000 tokens per minute. Pricing is $6.4 per million audio input tokens and $0.8 per million text input tokens, with combined text and audio output at $24 per million and output text not charged at all. Function calling, web search, structured outputs, context caching and fine-tuning are all built in.

A companion model, Qwen-Audio-3.1-ASR-Flash-Filetrans, targets offline transcription of long audio. It supports hot words, speaker separation, punctuation, and recognition across multiple languages plus Chinese dialects, at $0.15 input and $0.47 output per million tokens. For anyone who has priced long-form transcription at scale, that is a steep drop in what used to be a fixed cost.

The training story is organised into three layers, which Qwen labels Think, Act and Speak. The Think layer re-anchors the audio model to a source text model using million-hour-scale paired data, then applies on-policy distillation rather than imitation of pre-written answers. The Act layer builds executable environments, each bundling a tool pool, a stateful database and a written business policy, with tasks scored on terminal state first. The Speak layer decides whether, when and how to talk.

The detail in the Act layer is the one worth pausing on. Scoring checks the resulting state before it checks the wording, so a fluent reply cannot rescue a failed action. That is the right way to grade an agent that is supposed to do something rather than chat about it.

Why the price war is the real story

Strip away the architecture and Qwen-Audio-3.1 is a clear strategic move: Alibaba is racing OpenAI and Google on the cost and latency of voice infrastructure rather than on openness.

The comparison that makes this legible is latency. Qwen reports an interruption stop latency of about 1.116 seconds, against 0.383 seconds for GPT-Realtime-2. Being slower to stop when a user interrupts is a visible weakness, and the fact that Qwen published it alongside the wins says something about the state of the benchmark culture in audio, where honest numbers are still a differentiator.

For builders, the practical effect is a cheaper backend. If realtime speech is 85 percent cheaper and recognition is up to 95 percent cheaper, then voice features that were previously reserved for premium tiers become viable in ordinary products. Always-on listening, live translation, and spoken interfaces for agents all sit downstream of that cost curve. When the cost of a spoken minute collapses, product decisions that were unaffordable last quarter become obvious this one.

The caveat is the one that applies to every managed API. You get the behaviour, not the internals. You cannot inspect the turn-taking model, fine-tune the decision logic on your own data in a way you fully control, or run the stack on your own hardware. For most startups that is an acceptable trade. For anyone building something where the latency profile or the data path is the product, it is a dependency worth pricing in.

The competitive context sharpens the point. In the same stretch of September, OpenAI and Anthropic both shipped models framed around doing more work at lower cost, and Microsoft added a streaming transcription model and two text to speech models to its in-house line. Audio is no longer a quiet corner where a smaller lab can be best without anyone noticing. It is a category the largest players are competing on price and latency in public, which is good news for buyers and a squeeze for anyone charging a premium for voice.

What to watch

The obvious question is whether Qwen eventually open-sources any part of this stack, and how the two-model design holds up under real multi-turn tool-calling load rather than benchmark conditions. Splitting turn-taking from content is elegant on paper. Whether it stays elegant when a user interrupts, changes their mind, and switches tasks mid-sentence is the kind of question demos do not answer.

For now the takeaway is narrower and clearer. Alibaba has taken the part of voice agents that felt most broken, named it, and given it a dedicated model. Then it cut the price until the whole category got cheaper. If the next year of voice products is built on cheaper, more interruptible speech, this release will look like one of the reasons.

Related articles