Microsoft MAI-Realtime Leak Explained 2026: Full-Duplex Voice, 17 Languages (Not Official)
On August 2, TestingCatalog spotted MAI-Realtime in Microsoft's MAI Playground — a full-duplex voice model that listens and speaks at once, with two voices, 17 languages (including Korean), mid-conversation language switching and tool use. But there's no official Microsoft announcement, model card or pricing yet.
Let’s be clear up front: this is a leak / hidden preview, not an official announcement. On August 2, TestingCatalog spotted a hidden early-access entry called MAI-Realtime inside Microsoft’s MAI Playground. Microsoft has not announced it, published a model card, disclosed pricing, or named launch regions. So read this as “spotted,” not “confirmed.” Still, the direction is interesting enough to lay out.
What was spotted?
The core is full-duplex — it listens and speaks at the same time. Most voice AI today is turn-based (“I speak, then it replies”); this overlaps like a human conversation.
- Appears to be Microsoft’s first native full-duplex speech model
- Two voices, Victoria and Grant — reportedly noticeably more natural than Copilot’s current voice
- 17 languages (English, German, Spanish, French, Japanese, Korean, Chinese and more)

Microsoft’s earlier voice/image moves are covered in Microsoft MAI image and voice.
What else showed up?
From the leaked working listing:
| Item | Detail |
|---|---|
| Mid-conversation language switch | Pinned or auto-detect, switches without losing its footing |
| Turn-taking modes | Two selectable modes |
| Latency | Live debug panel |
| Tools | Web search and more |

This tracks with competition shifting from raw capability toward naturalness and real-time feel, a pattern also visible in AI rankings via OpenRouter.
So is it confirmed?
No. To repeat — there is no official Microsoft announcement. It surfaced as a hidden early-access entry in the MAI Playground; launch, pricing and regions are undecided and it’s not in the public catalogue. As a preview, final specs can change. The low-cost, low-latency race also shows in cases like Grok 4.5.
Frequently Asked Questions
Q. Can I try it now? Not as a general user. It was spotted as a hidden early-access entry in the MAI Playground; it’s not in the public catalogue and hasn’t officially launched.
Q. What makes full-duplex different? Existing voice AI replies after you finish (turn-taking). Full-duplex listens and speaks simultaneously, so it can interrupt and overlap like a human conversation.
Q. Does it support Korean? The spotted listing includes Korean among 17 languages. But this is a preview observation, not an official spec, so the final version may differ.
Related posts
HotThe AI Memory Crunch Explained: RAM Up 89%, HBM Sold Out, Shortage Until 2030?
The real bottleneck of the AI era isn't GPUs — it's memory
HotDeepSeek Is Building Its Own AI Inference Chip — Breaking Free From Nvidia and Huawei
The budget-AI champion's next move is silicon — vertical integration, Chinese style
HotThe AI Developers Actually Use Most? Chinese Models Now Dominate
The benchmark #1 and the 'most-used' model are not the same