Microsoft Rolls Out Streaming Transcription and Voice Models to Challenge OpenAI Dependency
Microsoft has released three new models in its MAI family, including its first real-time transcription engine and two text-to-speech variants, as the company works to reduce reliance on costly third-party AI providers.

Microsoft Corp. has grown its MAI artificial intelligence model portfolio with the introduction of a streaming transcription model alongside two text-to-speech alternatives. The trio targets developers building voice agents capable of understanding spoken input and generating immediate responses that mimic natural human conversation.
The centerpiece is MAI-Transcribe-2-Streaming, which processes human speech transmitted via WebSocket and produces continuously refreshed transcripts as the speaker talks. Once speech ends, the model confirms the transcript is complete. This capability enables applications featuring real-time captions or systems that begin handling requests before users finish speaking, according to Microsoft.
Available through Microsoft's Vercel AI Gateway, MAI-Transcribe-2-Streaming carries a price tag of 54 cents per audio hour. The model recognizes more than 60 languages with automatic language detection and typically delivers initial transcript predictions within 320 milliseconds on average. Microsoft cautioned that actual performance varies based on network conditions and the AI system generating responses.
https://x.com/MicrosoftAI/status/2105693024013467905?ref_src=twsrc%5Etfw
Microsoft previously released MAI-Transcribe-2 last month at 10 cents per audio hour, representing more than a five-fold cost reduction compared to the Streaming version. The price difference reflects the additional computational work required by the Streaming model, which continuously returns partial results as audio arrives, whereas the standard MAI-Transcribe-2 waits for the speaker to complete their statement before processing begins.
Complementing the transcription capability are two text-to-speech models designed to generate natural-sounding speech from written text. MAI-Voice-2.1 prioritizes expressive and high-fidelity audio output, while MAI-Voice-2.1-Flash sacrifices some quality for faster processing and reduced expenses. The Vercel AI Gateway lists MAI-Voice-2.1 at $22 per million characters and the Flash variant at $15 per million characters. Both support 23 languages.
The announcements signal Microsoft's strategic shift toward reducing dependence on external model providers including OpenAI Group PBC and Anthropic PBC, despite substantial investments in both companies. In July, reports indicated that Microsoft AI Chief Executive Mustafa Suleyman was growing concerned about expenditures tied to OpenAI's and Anthropic's advanced frontier models.
In response, Suleyman directed Microsoft's AI research teams to intensify efforts around the MAI model family, with plans to eventually deploy these models across Copilot agents integrated into applications like Excel and Outlook. "We pay a lot of money to Anthropic, so our goal is to reduce and ultimately eliminate that cost," Suleyman stated in remarks to Bloomberg.
The new model releases provide Microsoft with the foundational components needed to construct sophisticated voice AI agents capable of engaging users through natural, conversational interactions. Functional voice agents require three core capabilities: recognizing and comprehending spoken input, determining appropriate responses based on that input, and producing audible replies.
MAI-Transcribe-2-Streaming addresses the first requirement while the MAI-Voice models handle the third. The intermediate step relies on Microsoft's Mai-Thinking-1, a standard large language model that analyzes speech transcripts to determine agent behavior. This modular approach grants developers granular control over voice agent quality, response speed and operational expenses.

