Skip to content
Microsoft AI Launches New Transcription and Text-to-Speech Models

Microsoft AI Launches New Transcription and Text-to-Speech Models

First seen 2 Oct 2026, 10:07 UTC • •

Article Content

Browse articles
ThreatCluster AI
ThreatCluster •October 2, 2026 at 15:08 UTC
  • •MAI-Transcribe-2-Streaming offers real-time transcription in 60 languages with low latency.
  • •MAI-Voice-2.1 supports multilingual text-to-speech with a consistent voice across languages.
  • •Pricing for MAI-Transcribe-2-Streaming is $0.54 per hour through the end of the year.

On October 1, 2026, Microsoft AI introduced MAI-Transcribe-2-Streaming, a real-time transcription model that supports 60 languages and boasts a top accuracy ranking on Artificial Analysis. The model generates its first partial results in just over 100 milliseconds, allowing voice agents to respond while users are still speaking. Additionally, Microsoft released two text-to-speech models: MAI-Voice-2.1, which can speak 23 languages with a consistent voice, and MAI-Voice-2.1-Flash, optimized for low-latency applications. The introductory pricing for MAI-Transcribe-2-Streaming is $0.54 per hour of audio, while MAI-Voice-2.1 is priced at $22 per million characters and MAI-Voice-2.1-Flash at $15 per million characters. Both voice models include safeguards against misuse and are available on platforms like Microsoft Foundry and OpenRouter.

Start a free Starter trial for enhanced analysis

Ask AI about this cluster

Timeline

2026-10-01
Launch of MAI-Transcribe-2-Streaming
Microsoft AI released a new transcription model that supports 60 languages and offers real-time results.
News.Microsoft
2026-10-01
Release of new text-to-speech models
Microsoft introduced MAI-Voice-2.1 and MAI-Voice-2.1-Flash, enhancing multilingual voice capabilities.
News.Microsoft

More articles in this cluster (4)

Common questions

What languages does MAI-Transcribe-2-Streaming support?
It supports real-time transcription in 60 languages.
How fast does MAI-Transcribe-2-Streaming provide partial results?
It delivers its first partial results in just over 100 milliseconds.
Where can I access these new models?
The models are available through Microsoft Foundry, the MAI Playground, and OpenRouter.