Microsoft MAI-Transcribe-2-Streaming Targets Voice Agents

Microsoft MAI-Transcribe-2-Streaming Targets Voice Agents

Microsoft has added three new models to its MAI family, all aimed at the same goal: voice agents that can hold a conversation without awkward pauses. The headline release is MAI-Transcribe-2-Streaming, the company's first streaming transcription model. It ships alongside two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash.

The target audience is developers. Microsoft wants them to build agents that listen to a speaker and respond almost immediately, closer to the rhythm of a human conversation than the stop-and-wait pattern of many current assistants.

How the streaming model works

MAI-Transcribe-2-Streaming receives speech over a WebSocket connection and returns a transcript that keeps updating while the person is still talking. When the speaker stops, the model marks the transcript as final.

This matters for two kinds of applications, according to Microsoft:

  • Live captions, where text needs to appear while someone speaks.
  • Early processing, where an agent can start working on a request before the user has finished the sentence.

The model supports more than 60 languages and detects the spoken language on its own. Microsoft says it produces its first transcript hypotheses in 320 milliseconds on average. The company adds a caveat: it cannot promise that speed in every case, because overall response time also depends on the network connection and on the AI system that writes the reply.

The model is listed on Microsoft's Vercel AI Gateway at 54 cents per audio hour.

Why streaming costs more

Microsoft already has a non-streaming transcription model. Last month it launched MAI-Transcribe-2 at 10 cents per audio hour, which makes the new variant more than five times as expensive.

The gap comes down to workload. The standard model waits until the speaker is done and then processes the audio in one go. The streaming version keeps working while audio is still arriving and sends provisional results back continuously. More processing means a higher price, and developers will need to decide whether lower latency is worth it for their use case.

Two voices, two price points

On the output side, Microsoft is offering a choice between quality and speed:

  • MAI-Voice-2.1 is positioned as the option for more expressive, higher-fidelity speech. It costs $22 per million characters on the Vercel AI Gateway.
  • MAI-Voice-2.1-Flash gives up some of that expressiveness for faster responses and a lower price of $15 per million characters.

Both support 23 languages, according to Microsoft. That is well below the 60-plus languages covered by the transcription model, so an agent may be able to understand more languages than it can answer in.

The full voice stack

A working voice agent has to do three things: understand what was said, decide what to do about it, and speak a reply. With this release, Microsoft now covers all three steps with its own models.

MAI-Transcribe-2-Streaming handles the listening. The MAI-Voice models handle the speaking. In between sits Mai-Thinking-1, Microsoft's reasoning model, a standard large language model that reads the transcript and decides how the agent should respond.

Splitting the pipeline into three separate models gives developers more control. They can tune quality, latency and cost at each step instead of accepting a single package. That kind of modular design is showing up elsewhere too, for example in AWS's open decision model for agents.

The strategy behind the launch

The release also fits a broader shift at Microsoft. The company is a major investor in both OpenAI and Anthropic, yet it is building its own models to rely less on them.

In July, reports said Microsoft AI chief executive Mustafa Suleyman had grown concerned about the cost of frontier models from OpenAI and Anthropic. He reportedly told his researchers to focus on the MAI family, with the aim of eventually running Copilot agents in products such as Excel and Outlook on Microsoft's own models. "We pay a lot of money to Anthropic, so our goal is to reduce and ultimately eliminate that cost," Suleyman told Bloomberg.

Our Take

For developers, the most useful part of this launch may be the pricing transparency. Having batch transcription at 10 cents, streaming at 54 cents and two voice tiers listed side by side makes it easier to model what a voice agent will actually cost before building it. The trade-off is also clear: real-time responsiveness costs much more than waiting for a speaker to finish.

The release also suggests that voice is turning into a competitive layer in its own right. Specialists such as ElevenLabs have drawn heavy investor interest, with the company recently doubling its valuation to $22 billion. Microsoft offering a complete in-house stack puts pressure on that market, especially for customers already working inside its ecosystem.

There are open questions. Microsoft's 320-millisecond figure is an average, and the company itself says real-world latency depends on factors it does not control. The difference in language coverage between transcription and speech output could limit some international deployments. It is worth watching whether independent developers report similar speeds in production, and whether Microsoft follows through on moving Copilot onto MAI models. If it does, that would be a strong signal that the company sees its own models as good enough to replace its partners' in core products, not just in developer tools.