Skip to content

UPDATED 21:31 EDT / OCTOBER 01 2026

AI

Microsoft targets ultra-realistic voice agents with its first streaming transcription model

Microsoft Corp. today expanded its MAI artificial intelligence model family with its first streaming transcription model, debuting alongside two others focused on text-to-speech.

They’re designed for developers who want to build voice agents that can listen to people’s voices and reply instantly, similar to how humans talk to one another.

The most consequential of the three is MAI-Transcribe-2-Streaming, which the company said accepts human speech via a WebSocket, transforming it into a transcript that’s continuously updated as the person keeps on talking. Once the person has finished speaking, it will confirm that the transcript is final. In this way, the model can power applications that display live captions or start processing a user’s request before they’ve finished saying it, Microsoft said.

MAI-Transcribe-2-Streaming is listed on Microsoft’s Vercel AI Gateway and priced at 54 cents per audio hour, the company said. It supports more than 60 languages, and can detect which one someone is speaking automatically. It delivers its first transcript hypotheses within 320 milliseconds on average, although Microsoft said it cannot guarantee this kind of performance in every scenario, because speed of response is also determined by the network connection and the AI system that generates the reply.

Microsoft isn’t new to transcription models. Last month, it debuted the MAI-Transcribe-2 model with a price of 10 cents per audio hour, which is more than five times cheaper than the Streaming variant. That’s mainly because the Streaming model does a lot more processing, continuously returning provisional results while the audio is still arriving. The regular MAI-Transcribe-2 model simply waits for the speaker to finish what they’re saying before processing begins.

Moving away from transcription, there are two models focused on generating speech from text, which are equally important for conversational AI agents. They include MAI-Voice-2.1, which is said to be the best choice for more expressive and higher-fidelity outputs, and MAI-Voice-2.1-Flash, which tones down those attributes in favor of a speedier response and lower cost.

According to the Vercel AI Gateway listings, MAI-Voice-2.1 is priced at $22 per million characters, while the Flash version costs $15 per million characters. Both of the text-to-speech models support 23 languages, Microsoft said.

The new releases demonstrate that Microsoft is accelerating its transition away from model providers such as OpenAI Group PBC and Anthropic PBC, despite being a major investor in both of those companies. In July, it was reported that Microsoft AI Chief Executive Mustafa Suleyman was becoming increasingly concerned about the costs associated with OpenAI’s and Anthropic’s powerful frontier models.

In response, he instructed Microsoft’s AI researchers to double down on the MAI model family, with the goal of eventually using those models to power its Copilot agents in platforms such as Excel and Outlook. “We pay a lot of money to Anthropic, so our goal is to reduce and ultimately eliminate that cost,” Suleyman told Bloomberg in an interview.

With the launch of the new models, Microsoft has everything it needs to build powerful voice AI agents that can converse with their human users in a natural, humanlike way. Voice agents need to be able to do three things to function: They must be able to recognize and understand what someone is saying to them, then decide what to do based on what was said, and finally generate an audible response.

MAI-Transcribe-2-Streaming takes care of the first problem, and the MAI-Voice models solve the third. What sits in the middle is Microsoft’s powerful reasoning model, Mai-Thinking-1, which is a standard large language model that reads through the transcripts of what was said to decide what the agent should do. By using three separate models, developers have much more control over the quality, latency and costs of their voice agents.

Image: Microsoft AI

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

 

About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

Send us a news tip

Send us a News Tip

  • This field is for validation purposes and should be left unchanged.
  • Max. file size: 244 MB.

Sign in

SIGN IN

Bio

Ethics statement

Extract the signal from the noise

Get SiliconANGLE updates and analysis.

Contact us

Partner with us

Contact us

Guest inquiry