StandardsAboutContact
The Weights
MAI-Transcribe-2-Streaming Teardown: Wait

MAI-Transcribe-2-Streaming Teardown: Wait

Microsoft AI says its streaming speech-to-text model returns first partial transcripts in just over 100ms and ranks first on Artificial Analysis. The price is the only fully checkable number, and it sits 1.5x above the earlier batch model's listed rate.

Wait. Microsoft's MAI-Transcribe-2-Streaming claims first partial transcripts in just over 100ms at $0.54 per audio hour, 1.5x the $0.36 per hour Artificial Analysis lists for the earlier batch MAI-Transcribe-1. The latency and its number-one ranking are vendor-reported, with no disclosed audio conditions, so test on your own calls first.

The Weights Desk · 4 min read

Microsoft AI released MAI-Transcribe-2-Streaming, a real-time transcription model aimed at voice agents, alongside MAI-Voice-2.1 text-to-speech models. The announcement claims first partial transcripts in just over 100ms and a number-one accuracy ranking on Artificial Analysis. Our call is Wait: the only figure we could verify end to end is the price, and the headline latency comes with no stated test conditions.

The deciding number: just over 100ms to first partial

Microsoft says the model produces its first hypotheses, called partials, in just over 100ms of receiving audio. For a voice agent this sets how early downstream reasoning can start. But the announcement gives no percentile, audio type, chunk size or network path, so a reader cannot rerun it. A median on clean studio audio and a p95 on telephony audio are very different numbers, and the page does not say which this is.

The accuracy ranking cannot be reproduced from public text

Microsoft says the model ranks no. 1 for accuracy on both final and partial transcripts on Artificial Analysis, and that words appear 2x faster than its closest competitor, without naming that competitor. The ranking page is not linked from the announcement. Artificial Analysis's own MAI-Transcribe-1 write-up reports AA-WER scores, 3.0% and fourth place for that model, but does not describe the datasets or normalization behind the metric.

Price: 1.5x the earlier batch model, and only through year-end

Microsoft lists MAI-Transcribe-2-Streaming at $0.54 per audio hour as introductory pricing through the end of the year. Artificial Analysis listed MAI-Transcribe-1 at $6 per 1,000 minutes on Microsoft Foundry, which is $0.36 per hour. The streaming model is 1.5x that rate, for a different job, and the post-promotion rate is not stated. A cost model built on $0.54 is a floor, not a forecast.

The text-to-speech side: tiered pricing and a Flash variant

Microsoft lists MAI-Voice-2.1 at $22 per million characters and MAI-Voice-2.1-Flash at $15 per million characters, calling Flash about 60 percent cheaper than comparable models, a comparison it does not itemize. The voice models cover 23 languages and 26 locales, and the transcription model is listed at 60 languages with automatic language detection. These are two tier prices on one page, and the vendor's own comparisons are unverified.

What a real test has to measure

A useful pilot replays a few hundred of your own calls and records time to first partial at p50 and p95, how often partials are rewritten, and final word error rate split by accent and noise level, on identical audio against your current provider. The model is listed on Foundry, MAI Playground, Vercel, OpenRouter and Azure Voice Live, with LiveKit coming soon, so the harness needs little new infrastructure.

The bottom line: Wait

Bottom line: Wait. The deciding number is the roughly 100ms to first partial, and it carries no published conditions, while the ranking has no linked method and the price is promotional. Run the replay test now if latency is your bottleneck, and switch only if your own p95 holds up. Watch for independent latency measurements, the post-promotion hourly rate, and a documented AA-WER method.

Should a voice-agent team adopt MAI-Transcribe-2-Streaming now?
Pilot it, but do not commit. The 100ms first-partial figure and the accuracy ranking are Microsoft's own claims, and the price is introductory. Replay your own calls and compare time to first partial and partial revisions against your current provider before switching.
What does it cost compared with Microsoft's earlier model?
Microsoft lists MAI-Transcribe-2-Streaming at $0.54 per audio hour through the end of the year. Artificial Analysis listed MAI-Transcribe-1 at $6 per 1,000 minutes on Microsoft Foundry, which works out to $0.36 per hour. The streaming model is therefore 1.5x the price, for a different capability.
Is the accuracy ranking independently verified?
Not from the material we could read. Microsoft cites Artificial Analysis for the ranking but links no methodology. Artificial Analysis's earlier MAI-Transcribe-1 page reports AA-WER results without describing the datasets or normalization, so the ranking cannot be reproduced from public text alone.
Where is it available?
Microsoft lists Microsoft Foundry, MAI Playground, Vercel, OpenRouter and Azure Voice Live, with LiveKit coming soon.
  1. Microsoft AI releases new transcription and text-to-speech models for voice agents — The Decoder
  2. Our first streaming transcription model — Microsoft AI
  3. MAI-Transcribe-1: everything you need to know — Artificial Analysis