- Microsoft
- MAI-Transcribe-2-Streaming and MAI-Voice-2.1
- Audio
Microsoft pairs streaming transcription with multilingual MAI voices for conversational agents
Microsoft AI launched MAI-Transcribe-2-Streaming alongside MAI-Voice-2.1 and its Flash variant. Streaming recognition supports 60 languages with continuously revised partial transcripts; the voice models support 23 languages and 26 locales, with a consistent speaker identity across languages. Microsoft lists Foundry, Playground, Vercel and Azure Voice Live access, with LiveKit still coming soon. Personal voice cloning requires gated approval and recorded talent consent. These are audio components for an agent loop, not an autonomous agent or independently validated speed record.
Published Updated 4 sources
Primary source
Microsoft AI · October 1, 2026
Read the original: Streaming transcription and new MAI voice modelsOpens Microsoft AI in a new tab. Read it there before you rely on the summary above.
Unlock the full brief free.
- What to check before you trust this story, written down
- Each verified source, with why it matters and who published it
- A note whenever a source has been withdrawn
- Thirty days of stories to browse, not seven
Related on Rise Productive
AI model picker
A free tool for choosing the AI setup that fits your work.
The newsletter
What I built and what changed in AI, about once a week.
