Microsoft Cut Partial-Transcript Latency to 100 Milliseconds. That Changes What Transcription Is For.

On October 2, 2026, Microsoft AI announced MAI-Transcribe-2-Streaming, a speech recognition model that returns partial transcripts in just over 100 milliseconds across 60 languages. It ranks first on Artificial Analysis for both final and partial transcript accuracy. Introductory pricing is $0.54 per hour of audio through the end of the year.
The company also announced MAI-Voice-2.1 and MAI-Voice-2.1-Flash, covering 23 languages and 26 locales. The model is reachable through Azure Speech, Microsoft Foundry, the MAI Playground, OpenRouter under the id microsoft/mai-transcribe-2, Vercel, and Azure Voice Live, with a LiveKit integration listed as coming soon.
One hundred milliseconds is a threshold, not a benchmark curiosity. It sits at the edge of what people perceive as instantaneous. Below it, an interface can respond while someone is still speaking, and that is the difference between a product that produces a document and a product that participates in a conversation.

Why final accuracy was never the whole story
Speech recognition has been a mature category for years, and the leaderboards have mostly measured one thing: how accurately a model renders a completed utterance once the speaker has stopped. That metric answers the question a transcription service cares about. It does not answer the question a live product cares about.
A meeting assistant that only shows words after the speaker finishes is a recorder with better formatting. A meeting assistant that shows words as they arrive can highlight the current speaker, surface a name the moment it is mentioned, trigger a search while the subject is still on the table, and let someone scroll back mid-sentence without losing their place. All of that depends on partial latency rather than final accuracy.
Partial accuracy is the harder problem, and it is the one Artificial Analysis now tracks separately. Generating a stable guess before the sentence is complete means committing to an interpretation you may have to revise. Models that revise constantly produce text that flickers, which is worse than waiting. Models that commit too aggressively leave errors on screen. Ranking on partials rewards the model that gets the interim text right often enough that a reader can trust what they are looking at.
Microsoft claiming first place on both boards means it is not trading interim stability for final correctness, which is the usual way models buy partial speed. That combination is the actual announcement.
The 100 millisecond number
End-to-end latency in streaming recognition is a chain, not a single figure. Audio arrives in frames. An endpointing model decides when an utterance may have ended. The recognition model produces text. A downstream system renders it. Every link adds delay, and the total determines whether the experience feels live.
Microsoft's 100 millisecond figure covers the model's own output, based on the company's own measurement and Artificial Analysis placement rather than an independent latency audit. The relevant comparison concerns whether the number holds under the conditions a live product creates, which is a different question from a stopwatch reading in isolation: several speakers, background noise, a poor conference microphone, and 60 languages arriving on the same pipeline. Vendors rarely publish latency distributions across those conditions, and they are the ones that determine whether a meeting product is usable in an actual meeting.
The pricing structure is the other half of the story. $0.54 per hour of audio is an introductory rate that runs through the end of the year, which tells you Microsoft expects to move it. The strategic intent is legible. Streaming transcription is the input layer for every meeting assistant, call center analytics product, and live captioning service. Being the cheapest fast option for 60 languages is a distribution play, and distribution in an input layer is sticky, because the second team to switch providers has to rebuild their handling for a different set of interim behaviors.
What this competes with
Microsoft is not entering an empty field. Deepgram, AssemblyAI, and Speechmatics have all built businesses on streaming accuracy and low latency, and Google and Amazon bundle competitive recognition into their own cloud stacks. The realistic buyer is a developer who already runs on Azure, values the fact that MAI-Transcribe-2-Streaming appears in Foundry and OpenRouter alongside the rest of the MAI family, and would rather not maintain a second vendor relationship for one component.
The MAI-Voice-2.1 announcement matters here too, and the pairing is deliberate. Speech recognition and speech synthesis are the two ends of the same conversation, and a platform that sells both can pitch a single pipeline: audio in, transcript out, response synthesized, speech returned. Adding 23 languages and 26 locales on the synthesis side widens the number of markets where that pipeline can be sold as one product.
What fast partials change in product design
Once interim text arrives inside a tenth of a second, a set of interface decisions that used to be unreasonable become ordinary.
Consider interruption. If a transcript settles only after a speaker stops, a system has to wait for silence before it can act on what was said, which means it cannot respond until the conversation has already moved on. With fast partials, a system can recognize a command the moment it is spoken and cut in, which is the behavior people expect from a person in the room. Voice agents that support barge-in depend on this. So do live captioning systems that need to keep up with a fast speaker rather than trailing three sentences behind.
Search is the second beneficiary. A meeting assistant that can search across the transcript as it is being written can surface a document the moment a project name comes up, while the discussion is still happening. That is a fundamentally different product from a recorder that produces searchable notes an hour later.
Structured extraction is the third. If partial text is reliable enough, a downstream model can begin reading it before the meeting ends, tagging decisions and action items as they are spoken. The gain here comes from being able to review the output while the participants still remember what they meant, rather than from raw speed.
Each of those behaviors raises the stakes on partial stability. A transcript that rewrites itself two or three times a second makes all three features annoying rather than useful, which is why the deeper story in Microsoft's announcement is the accuracy placement on partials rather than the latency figure by itself.
Where the interesting risk sits
Transcription accuracy figures are usually reported as aggregate percentages, and aggregate percentages hide the distribution underneath. Performance on accented speech, code-switching between languages mid-sentence, children's voices, and non-standard microphones varies enormously between models, and those are precisely the conditions where a live product is most likely to be used. A model that posts a strong average can still be the wrong pick for a specific call center whose customers switch languages halfway through a sentence.
The second risk is more strategic. Once transcription is fast and cheap enough to run continuously, the natural next step is to run it always. A model that costs half a dollar an hour makes continuous listening affordable in a way that a per-minute pricing model did not, and continuous listening is an ambient recording of everything said near a device. That is a data governance decision that no accuracy leaderboard will make for the teams buying this.
MAI-Transcribe-2-Streaming is a genuinely useful piece of infrastructure, and the partial-accuracy ranking is the part worth watching, because it measures the metric that live products actually depend on. What remains unresolved is whether a 100 millisecond number measured by the vendor holds up in a room with bad acoustics and four people talking over each other. That is the test every buyer will run, in private, before they commit a production pipeline to it.
Related articles
Step 5 Skipped Four Version Numbers, Landed Second Among Open Models, and Nobody Is Calling It
Benchmark position is a signal producers use to judge a model. Call volume is a signal consumers produce by using it.
Gemini Omni 1.1 Flash Chained Four Requests Into a 40-Second Shot. The Workflow Is the Product.
Google shipped a production pipeline with a review gate built into the pricing.
Synthesia Spent a Decade Recording Avatars. Now It Wants Them to Talk Back.
A training tool people voluntarily repeat is a different product from one they complete because someone assigned it.
A Model That Refuses to Write Sentences Is Now Handling a Trillion Tokens a Day
A classifier with a much better interface, aimed at the work software already wanted structured.