Microsoft is testing MAI Realtime, its first native real-time voice model, through limited access in the company's MAI Playground platform.
The system operates as a full-duplex model that listens and speaks simultaneously, rather than taking turns. This positions it as a direct competitor to OpenAI's GPT Live voice capabilities.
MAI Realtime currently offers two voices — Victoria and Grant — which deliver more natural speech than Microsoft's existing Copilot voice mode. The model supports automatic language detection and can switch between languages mid-conversation without interruption.
Language support and technical features
The system supports 17 languages including English, German, Spanish, French, Italian, Portuguese, Japanese, Korean, Chinese, Dutch, Hindi, Indonesian, Arabic, Russian, Turkish, Vietnamese, and Thai.
Two listener configurations are available: Switchboard mode uses an MAI-Ears endpointer with inline control tokens, while a deterministic setup combines silence-based endpointing with a Whisper semantic endpointer.
Both modes handle interruptions cleanly and maintain low response latency. The model does not produce singing or non-speech audio, focusing purely on conversational interaction.
The limited preview suggests Microsoft is preparing to challenge established players in the real-time voice AI market. The company has not disclosed when MAI Realtime will receive broader availability or commercial pricing.
View tweet from @testingcatalog
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.