← Back to blog
AiAbout 7 min read

Microsoft Cut Partial-Transcript Latency to 100 Milliseconds. That Changes What Transcription Is For.

Published Oct 5, 2026
Microsoft Cut Partial-Transcript Latency to 100 Milliseconds. That Changes What Transcription Is For.

On October 2, 2026, Microsoft AI announced MAI-Transcribe-2-Streaming, a speech recognition model that returns partial transcripts in just over 100 milliseconds across 60 languages. It ranks first on Artificial Analysis for both final and partial transcript accuracy. Introductory pricing is $0.54 per hour of audio through the end of the year.

The company also announced MAI-Voice-2.1 and MAI-Voice-2.1-Flash, covering 23 languages and 26 locales. The model is reachable through Azure Speech, Microsoft Foundry, the MAI Playground, OpenRouter under the id microsoft/mai-transcribe-2, Vercel, and Azure Voice Live, with a LiveKit integration listed as coming soon.

One hundred milliseconds is a threshold, not a benchmark curiosity. It sits at the edge of what people perceive as instantaneous. Below it, an interface can respond while someone is still speaking, and that is the difference between a product that produces a document and a product that participates in a conversation.

An extreme close up of a microphone diaphragm in fine gold mesh under cyan light

Why final accuracy was never the whole story

Speech recognition has been a mature category for years, and the leaderboards have mostly measured one thing: how accurately a model renders a completed utterance once the speaker has stopped. That metric answers the question a transcription service cares about. It does not answer the question a live product cares about.

A meeting assistant that only shows words after the speaker finishes is a recorder with better formatting. A meeting assistant that shows words as they arrive can highlight the current speaker, surface a name the moment it is mentioned, trigger a search while the subject is still on the table, and let someone scroll back mid-sentence without losing their place. All of that depends on partial latency rather than final accuracy.

Partial accuracy is the harder problem, and it is the one Artificial Analysis now tracks separately. Generating a stable guess before the sentence is complete means committing to an interpretation you may have to revise. Models that revise constantly produce text that flickers, which is worse than waiting. Models that commit too aggressively leave errors on screen. Ranking on partials rewards the model that gets the interim text right often enough that a reader can trust what they are looking at.

Microsoft claiming first place on both boards means it is not trading interim stability for final correctness, which is the usual way models buy partial speed. That combination is the actual announcement.

The 100 millisecond number

End-to-end latency in streaming recognition is a chain, not a single figure. Audio arrives in frames. An endpointing model decides when an utterance may have ended. The recognition model produces text. A downstream system renders it. Every link adds delay, and the total determines whether the experience feels live.

Microsoft's 100 millisecond figure covers the model's own output, based on the company's own measurement and Artificial Analysis placement rather than an independent latency audit. The relevant comparison concerns whether the number holds under the conditions a live product creates, which is a different question from a stopwatch reading in isolation: several speakers, background noise, a poor conference microphone, and 60 languages arriving on the same pipeline. Vendors rarely publish latency distributions across those conditions, and they are the ones that determine whether a meeting product is usable in an actual meeting.

The pricing structure is the other half of the story. $0.54 per hour of audio is an introductory rate that runs through the end of the year, which tells you Microsoft expects to move it. The strategic intent is legible. Streaming transcription is the input layer for every meeting assistant, call center analytics product, and live captioning service. Being the cheapest fast option for 60 languages is a distribution play, and distribution in an input layer is sticky, because the second team to switch providers has to rebuild their handling for a different set of interim behaviors.

What this competes with

Microsoft is not entering an empty field. Deepgram, AssemblyAI, and Speechmatics have all built businesses on streaming accuracy and low latency, and Google and Amazon bundle competitive recognition into their own cloud stacks. The realistic buyer is a developer who already runs on Azure, values the fact that MAI-Transcribe-2-Streaming appears in Foundry and OpenRouter alongside the rest of the MAI family, and would rather not maintain a second vendor relationship for one component.

The MAI-Voice-2.1 announcement matters here too, and the pairing is deliberate. Speech recognition and speech synthesis are the two ends of the same conversation, and a platform that sells both can pitch a single pipeline: audio in, transcript out, response synthesized, speech returned. Adding 23 languages and 26 locales on the synthesis side widens the number of markets where that pipeline can be sold as one product.

What fast partials change in product design

Once interim text arrives inside a tenth of a second, a set of interface decisions that used to be unreasonable become ordinary.

Consider interruption. If a transcript settles only after a speaker stops, a system has to wait for silence before it can act on what was said, which means it cannot respond until the conversation has already moved on. With fast partials, a system can recognize a command the moment it is spoken and cut in, which is the behavior people expect from a person in the room. Voice agents that support barge-in depend on this. So do live captioning systems that need to keep up with a fast speaker rather than trailing three sentences behind.

Search is the second beneficiary. A meeting assistant that can search across the transcript as it is being written can surface a document the moment a project name comes up, while the discussion is still happening. That is a fundamentally different product from a recorder that produces searchable notes an hour later.

Structured extraction is the third. If partial text is reliable enough, a downstream model can begin reading it before the meeting ends, tagging decisions and action items as they are spoken. The gain here comes from being able to review the output while the participants still remember what they meant, rather than from raw speed.

Each of those behaviors raises the stakes on partial stability. A transcript that rewrites itself two or three times a second makes all three features annoying rather than useful, which is why the deeper story in Microsoft's announcement is the accuracy placement on partials rather than the latency figure by itself.

Where the interesting risk sits

Transcription accuracy figures are usually reported as aggregate percentages, and aggregate percentages hide the distribution underneath. Performance on accented speech, code-switching between languages mid-sentence, children's voices, and non-standard microphones varies enormously between models, and those are precisely the conditions where a live product is most likely to be used. A model that posts a strong average can still be the wrong pick for a specific call center whose customers switch languages halfway through a sentence.

The second risk is more strategic. Once transcription is fast and cheap enough to run continuously, the natural next step is to run it always. A model that costs half a dollar an hour makes continuous listening affordable in a way that a per-minute pricing model did not, and continuous listening is an ambient recording of everything said near a device. That is a data governance decision that no accuracy leaderboard will make for the teams buying this.

MAI-Transcribe-2-Streaming is a genuinely useful piece of infrastructure, and the partial-accuracy ranking is the part worth watching, because it measures the metric that live products actually depend on. What remains unresolved is whether a 100 millisecond number measured by the vendor holds up in a room with bad acoustics and four people talking over each other. That is the test every buyer will run, in private, before they commit a production pipeline to it.

Related articles