← Back to the wire6 Oct 2026
The AI WireDispatch No. 086
techcommunity.microsoft.com ▸ models· filed 2 Oct 2026 · 2 min read

Microsoft adds streaming transcription and two text-to-speech models to Azure AI Foundry

Microsoft has introduced three speech models to Azure AI Foundry: one that streams partial transcripts before a speaker finishes, and two text-to-speech models that trade off latency and cost. The streaming model is priced at $0.54 per hour until the end of the year, while the voice models run $22 and $15 per million characters. Benchmark comparisons come from vendor-reported figures, and the pricing of transcription for the long term remains unspecified.

Machine-drafted illustration · reviewed by a humanFIG. 01

Listen to this dispatch

Narrated by an AI-generated voice.

Microsoft has added three speech models to its MAI lineup in Azure AI Foundry, following the non-streaming MAI-Transcribe-2 released earlier this month. One model transcribes live speech as it arrives; two convert text to spoken audio with different quality and cost trade-offs.

MAI-Transcribe-2-Streaming handles 60 languages with automatic language detection. It returns partial transcripts before a speaker finishes. Microsoft says the first hypotheses arrive within a few hundred milliseconds of receiving audio, then refine as more context arrives and commit a stable version once the utterance ends. Microsoft claims the model ranks #1 for accuracy on both partial and final transcripts on the Artificial Analysis leaderboard, with words appearing as early as 320 milliseconds after being spoken versus more than 500 milliseconds for the closest competitor. Those figures come from Microsoft's own reporting of a third-party benchmark, so vendor-reporting caveats apply. Microsoft frames the practical difference as whether an application can act on a request mid-sentence — identifying a caller's issue, looking up a record, preparing a tool call — rather than waiting for a complete transcript.

The text-to-speech side includes two models. MAI-Voice-2.1 supports 23 languages with a consistent voice identity across them: the same voice adapts pronunciation and delivery per language rather than carrying one accent everywhere. MAI-Voice-2.1-Flash is the same model family optimized for latency and high volume. Microsoft says Flash is 55% faster and 60% less expensive than competing models in its class, with the comparison sourced from ElevenLabs' documentation — again, a vendor-asserted figure rather than an independent evaluation.

Pricing: MAI-Voice-2.1 runs $22 per 1M characters; Flash runs $15 per 1M characters. MAI-Transcribe-2-Streaming is priced at an introductory $0.54 per hour of audio through the end of the year. The earlier MAI-Transcribe-2 remains available for workloads needing speaker diarization and word-level timestamps, which the streaming model apparently does not offer.

All three models are accessible through Microsoft Foundry and Azure Speech, and through OpenRouter, Vercel, and LiveKit.

The announcement does not answer two questions: whether the streaming model's accuracy holds up outside benchmark conditions, and whether $0.54 per hour of audio is a long-term rate or a promotional one. Transcription pricing is the line item most likely to balloon for a voice-agent deployment running constantly, so that answer matters more to production users than any leaderboard position.

Read the original at techcommunity.microsoft.com →

End of dispatch
More on the wire
87techcommunity.microsoft.com ▸ azure · filed 2 Oct 2026Federated Agent Factory: Microsoft's Reference Architecture for AI Agents Across Entra Tenants85techcommunity.microsoft.com ▸ tooling · filed 2 Oct 2026Open-source demo separates AI model proposals from sandboxed tool execution84techcommunity.microsoft.com ▸ product · filed 27 Sept 2026Azure Container Apps Sandboxes: per-agent microVMs with egress controls83github.blog ▸ product · filed 27 Sept 2026GitHub's HydraFusion preview orchestrates coding models at runtime