Microsoft MAI Transcribe 2, Voice 2.1 Explained
Microsoft AI announced three new audio models on October 1, 2026: a speech-to-text model called MAI-Transcribe-2-Streaming, and two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Together they cover the two halves of a voice assistant: hearing what someone says and saying something back.
The headline number is about 100 milliseconds, and it needs careful reading. That is Microsoft's claim for how quickly the transcription model produces its first partial result after audio arrives. It is not how long a whole conversational reply takes. This article separates what Microsoft has actually claimed, what its documentation confirms, and what is still unproven, including what it could mean for Hindi and other Indian languages.
⚡ Quick facts
- Announced: October 1, 2026, by Microsoft AI
- Models: MAI-Transcribe-2-Streaming (speech to text), MAI-Voice-2.1 and MAI-Voice-2.1-Flash (text to speech)
- Speed claim: first partial transcripts in just over 100ms of receiving audio, per Microsoft
- Languages: 60 for transcription; 23 languages for voice, including Hindi and English (India)
- Price: $0.54 per audio hour (introductory, through 2026); $22 and $15 per million characters for the two voice models
- Status: public preview in Microsoft Foundry, per Microsoft's documentation
What is MAI-Transcribe-2-Streaming?
MAI-Transcribe-2-Streaming is a real-time speech recognition model. Instead of waiting for a recording to finish, it works on audio as it arrives and sends back partial transcripts that are refined into final text. That is the behaviour needed for live captions, dictation, meeting notes and voice agents. It follows Microsoft's earlier batch model, which we covered in our explainer on MAI-Transcribe-2.
Microsoft says the model produces its first partial hypotheses in just over 100ms of receiving audio. Its Foundry blog also says words appear in the transcript as early as 320 milliseconds after they are spoken, against more than 500 milliseconds for a competitor baseline, and that words appear about twice as fast as with its closest competitor. These are different measurements, and all of them are Microsoft's own claims.
On accuracy, Microsoft says the model ranks first on both final and partial transcripts on Artificial Analysis's streaming word-error-rate leaderboard. Treat that as a Microsoft-reported ranking: benchmark leaderboards change quickly and we did not independently re-run it.
What is MAI-Voice-2.1?
MAI-Voice-2.1 is Microsoft's highest-fidelity and most expressive text-to-speech model. Microsoft's documentation positions it for long-form audio such as audiobooks, podcasts and voice-overs, and says plainly that it prioritises naturalness and expressiveness over latency-critical scenarios.
Its documented features include a single voice that can speak across languages with a native accent, emotion and style control through SSML, and instant voice cloning from 5 to 60 seconds of reference audio. Cloning is gated: it requires Microsoft's Limited Access review, consent and licensed voices.
What is MAI-Voice-2.1-Flash?
MAI-Voice-2.1-Flash is the low-latency sibling, built for real-time agents, phone menus and call-centre use. It covers the same languages and supports cloning. Microsoft says end-to-end latency is about 150ms, and that it is 55% faster and 60% less expensive than competing models in its class. The comparison set was not published in the material we reviewed, so those percentages remain Microsoft's claim.
Why 100ms matters, and what it does not mean
A voice agent has three jobs in sequence: transcribe the speech, decide what to say, then speak it. People notice delay quickly in conversation, so every stage counts. A faster transcriber lets the language model start working sooner and makes live captions feel closer to real time.
But a 100ms first partial does not make a whole conversation respond in 100ms. The language model's thinking time and the speech model's startup both add to the total. The 100ms figure also describes a first, unfinished hypothesis, not the final corrected text. Real latency in a product depends on network distance, region and how the app is built, which is why a model hosted in a nearby region matters. Microsoft lists Central India among the Azure Speech regions for the voice models.
Languages and locales
A language is not the same thing as a locale. English is one language, while en-IN, en-US and en-GB are three locales with different accents and conventions. Microsoft's counts mix both terms, so here is what is stated:
| Model | What Microsoft states | What the documentation shows |
|---|---|---|
| MAI-Transcribe-2-Streaming | 60 languages, automatic language detection | Hindi plus Assamese, Bengali, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu, Urdu and Nepali listed |
| MAI-Voice-2.1 and Flash | 23 languages and 26 locales | 28 locale codes across 23 languages listed on October 2, 2026, including hi-IN and en-IN; more are added over time |
We note the 26 versus 28 discrepancy rather than pick one number. It is most likely a documentation update after the announcement, but Microsoft has not said so.
What it could mean for India
Everything in this section is a potential application, not an announced deployment. Microsoft has not announced Indian customers or India-specific products for these models.
The documentation lists Hindi voices named Arjun, Dhruv, Grant, Harper, Kavya and Priya, and English (India) voices Dhruv and Priya, in a neutral style only. On the transcription side, the listed language set covers most major Indian languages. Microsoft's batch-transcription documentation also describes code-switching, such as mixing Hindi and English in one sentence, plus speaker labelling and word-level timestamps, but it is not confirmed that streaming supports all of those yet.
Potential uses include live captions for lectures and regional-language video, Hindi voice agents for banking or support lines, and audiobook narration. Whether the models handle Indian accents, noisy phone audio and Hinglish well is something only independent testing will show. Anyone building for India should run their own audio through the preview before committing. For cost-free tools that work in India today, see our list of free AI tools that work in India.
How it fits among other voice models
Microsoft is entering a crowded field. Reporting from heise names ElevenLabs, Deepgram, AssemblyAI, Google, OpenAI and xAI as competitors in audio. We are not declaring a winner: the figures that matter, accuracy on your own accents and latency from your own region, are not published in a comparable way across vendors, and the speed and price comparisons in this launch come from Microsoft.
For context on the rivals, see our coverage of OpenAI's GPT Live voice, Google's Gemini 3.8 Live Avatar and the ElevenLabs and Universal Music platform. Microsoft's push also fits its broader move toward in-house models, which we touched on in our Copilot apps story.
Pricing and availability
| Model | Price (per Microsoft) | Notes |
|---|---|---|
| MAI-Transcribe-2-Streaming | $0.54 per audio hour | Introductory, through the end of 2026; later price not found |
| MAI-Voice-2.1 | $22 per 1M characters | Highest fidelity, long-form |
| MAI-Voice-2.1-Flash | $15 per 1M characters | Low latency, real-time agents |
The models are available through Microsoft Foundry (Azure Speech), the MAI Playground, Azure Voice Live and direct APIs, and Microsoft lists third-party routes such as OpenRouter and Vercel. Microsoft's Learn pages mark both Transcribe-2 and Voice-2.1 as public preview with no service-level agreement and recommend against production use. Full pricing is on Azure's Speech pricing page.
What Microsoft is building
With text, image and now three audio models under the MAI name, Microsoft is assembling its own model family instead of relying only on partners. heise reports this as part of an effort to reduce dependence on OpenAI. That is reported context, and Microsoft's announcement itself focuses on the models, not on strategy. What it does show is a company treating voice as a core interface, from transcription at the input to speech at the output.
Frequently asked questions
What is MAI-Transcribe-2-Streaming?
Microsoft's real-time speech-to-text model, announced October 1, 2026. It returns partial transcripts while a person is still speaking, supports 60 languages with automatic language detection, and is in public preview in Microsoft Foundry.
Does 100ms mean a voice agent replies in 100ms?
No. Microsoft's figure is for the first partial transcript after audio arrives. A full voice-agent turn also includes the language model thinking and the speech synthesis, so total response time is longer. Microsoft separately quotes about 150ms end-to-end latency for MAI-Voice-2.1-Flash on the speech output side.
Does MAI-Transcribe-2-Streaming support Hindi?
Microsoft's documentation lists Hindi among the supported languages, along with Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia, Assamese and Urdu. The documentation page that describes code-switching and diarization covers Microsoft's batch transcription API, so streaming-specific support for those features is not confirmed here.
How much does MAI-Transcribe-2-Streaming cost?
Microsoft lists an introductory price of $0.54 per audio hour through the end of 2026. A standard price after that was not available when this article was written, so check Azure's Speech pricing page.
What is the difference between MAI-Voice-2.1 and MAI-Voice-2.1-Flash?
MAI-Voice-2.1 is the highest-fidelity, most expressive model, aimed at long-form audio such as audiobooks, podcasts and voice-overs, priced at $22 per million characters. MAI-Voice-2.1-Flash is the low-latency variant for real-time agents and call-centre use, priced at $15 per million characters.
Are these models ready for production?
Microsoft's own Learn documentation marks both MAI-Transcribe-2 and MAI-Voice-2.1 as public preview, without an SLA and not recommended for production workloads. Teams should test them, not depend on them yet.
Can MAI-Voice-2.1 clone my voice?
It supports instant voice cloning from 5 to 60 seconds of reference audio, but access is gated behind Microsoft's Limited Access review, with consent and licensed-voice requirements.
How many languages does MAI-Voice-2.1 speak?
Microsoft says 23 languages and 26 locales. Its documentation voice table currently lists 28 locale codes across 23 languages, including hi-IN and en-IN, and notes that more locales are added over time.
Sources
- Microsoft AI: our first streaming transcription model
- Microsoft Foundry blog: new MAI voice models
- Microsoft Learn: MAI voices
- Microsoft Learn: MAI-Transcribe
- heise: Microsoft expands MAI with three audio models