MAI has released three new models that change how voice agents can be built. MAI-Transcribe-2-Streaming, a real-time transcription model, launched October 1 and immediately ranked first on Artificial Analysis for both final and partial transcript accuracy. Alongside it came MAI-Voice-2.1 and MAI-Voice-2.1-Flash, two text-to-speech models that support 23 languages with native accents while keeping the same speaker identity across them. For customer support teams, healthcare providers, and sales professionals who rely on voice interactions, these models shrink the time between hearing and responding - the core loop of any conversational AI.
Transcription that starts before the speaker finishes
MAI-Transcribe-2-Streaming produces its first partial transcript roughly 100 milliseconds after receiving audio. It then revises those partials as more context arrives and commits a stable transcript immediately when the speaker stops. This means voice agents can begin reasoning or calling tools mid-sentence. For real-time dictation or subtitling, internal evaluations show words appearing on screen twice as fast as the closest competitor. The model works across 60 languages with automatic, continuous language detection. It sits on the Pareto frontier for accuracy versus latency, proving faster responses do not require sacrificing precision. The introductory price runs $0.54 per hour of audio through the end of the year.
These partial transcripts open practical workflows for live captioning during events, real-time note-taking in healthcare consultations, and instant response triggers in customer service calls. The model was designed so voice applications act on speech as it happens, not after a pause.
One voice, 23 languages, no accent drift
MAI-Voice-2.1 lets a single voice speak English, Mandarin, German, and 20 other languages - each with a native accent rather than dragging one accent across all of them. A tutoring app can switch languages mid-lesson without changing the teacher's voice. A multilingual assistant replies in whatever language the user speaks while sounding like the same person. The model covers 26 locales and costs $22 per 1 million characters.
MAI-Voice-2.1-Flash targets high-volume, latency-sensitive workloads. It generates 45 seconds of audio with 150 milliseconds of end-to-end latency. Inference runs 55% faster than comparable models, and pricing sits at $15 per 1 million characters - roughly 60% cheaper than alternatives. Both voice models support voice cloning from a few seconds of reference audio and include built-in consent guardrails.
Closing the conversational loop
A voice agent must hear, understand, decide, and speak within the window where a human still feels they are in a conversation. Pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash buys time on both ends of that loop. The transcription model delivers words faster, and the Flash voice model returns spoken responses with minimal delay. The time saved gives the agent more room to reason, check facts, and use tools while keeping the interaction at conversational speed.
Developers can already build customer service agents that transcribe requests as they are spoken and respond in natural speech, multilingual assistants that auto-detect language and reply in the same voice, and interactive learning tools that use distinct speakers for tutoring, role-play, and narration across supported languages. MAI built a demo called Chatter in the MAI Playground to show the models working together in a live agent. Access is available through OpenRouter, Microsoft Foundry, Vercel, and the MAI Playground, with LiveKit support coming soon.
Why this matters for creatives, customer support, healthcare, and sales professionals
Voice agents built on these models can handle real-world conversational rhythms - interruptions, mid-sentence corrections, and fast turn-taking - without the awkward pauses that make automated calls feel robotic. Customer support teams can deploy agents that begin resolving issues while the caller is still describing the problem. Healthcare providers can use real-time transcription during patient consultations and generate spoken follow-up instructions in the patient's preferred language, all in the same voice. Sales teams making international calls can maintain a consistent brand voice across languages. Writers and content creators gain faster dictation and subtitling tools. The core shift is practical: the technology no longer forces a tradeoff between speed, accuracy, and voice quality.
Your membership also unlocks: