Skip to main content

News

  • Microsoft
  • MAI-Transcribe-2-Streaming and MAI-Voice-2.1
  • Audio

Microsoft pairs streaming transcription with multilingual MAI voices for conversational agents

Microsoft AI launched MAI-Transcribe-2-Streaming alongside MAI-Voice-2.1 and its Flash variant. Streaming recognition supports 60 languages with continuously revised partial transcripts; the voice models support 23 languages and 26 locales, with a consistent speaker identity across languages. Microsoft lists Foundry, Playground, Vercel and Azure Voice Live access, with LiveKit still coming soon. Personal voice cloning requires gated approval and recorded talent consent. These are audio components for an agent loop, not an autonomous agent or independently validated speed record.

Published Updated 4 sources

Primary source

Microsoft AI · October 1, 2026

Read the original: Streaming transcription and new MAI voice models

Opens Microsoft AI in a new tab. Read it there before you rely on the summary above.

Unlock the full brief free.

  • What to check before you trust this story, written down
  • Each verified source, with why it matters and who published it
  • A note whenever a source has been withdrawn
  • Thirty days of stories to browse, not seven

This is not an account: there is no password, and the unlock is a cookie in this browser. You also join the Rise Productive newsletter from Demetri Panici, about once a week: what I built and what changed in AI. We'll email you a link to confirm, and you can unsubscribe in one click. The same signup unlocks every free tool on the site. How your email is handled.