Microsoft is testing an AI that speaks while you speak, without announcing anything
On August 2, someone spotted a hidden entry in Microsoft's internal playground, an interface reserved for testing. There is a model called MAI-Realtime. It would be offered in early access to a small group of partners. No press release, no technical sheet, no pricing page, not even a blog post. Just a model hidden in a corner.
And it wouldn't be just another voice model. It would be Microsoft's first model capable of listening and speaking at the same time.
What we know, and why to take it with a grain of salt
Let's start with the essential point: this is a leak, not an announcement. The information comes from TestingCatalog, which had access to the trial version. Microsoft has not confirmed anything or provided a date. The company could very well bury the project next week.
This version offers two voices, Victoria and Grant. They would be significantly more natural than those of the current voice mode of Copilot. The model would support seventeen languages, including French. A debugging panel also allows tracking processing steps and latencies, that is, response times. It quickly becomes clear that the tool is aimed at developers, not the general public. It also has a rather funny limitation: it does not sing and produces no sound other than speech. No sound effects, no humming.
The walkie-talkie and the phone
To understand what this changes, we first need to look at how current voice assistants work. The complicated word of the day is full duplex. Two objects are enough to explain it.
With a walkie-talkie, each person speaks in turn. You press, you speak, you release, then you say "your turn." As long as you are speaking, you hear nothing. With a phone, both people can talk and listen at the same time. They can interrupt each other or say "mmh" while the other continues. Full duplex is the phone. Half duplex is the walkie-talkie.
The voice assistants you use work like walkie-talkies. Their processing resembles a relay race with three runners. The first converts your voice into text. The second prepares the response. The third converts this text back into voice. Each one waits for the previous one to finish.
Three runners pass the baton. Only one runs directly. The difference is measured in waiting time.
A full duplex model eliminates these relays. The audio comes in and goes out in a single pass. Most importantly, the model continues to listen to you while it speaks. If you interrupt it, it can stop instead of finishing its sentence into the void, like your current assistant.
The red line is us. Linguists measure about 200 milliseconds of silence between two turns of speech in humans.
The numbers show the difference. A well-optimized classic chain responds between 800 milliseconds and 2 seconds. Without particular optimization, it takes between 2 and 4 seconds. A full duplex model like Moshi, published by Kyutai, operates around 200 milliseconds. This is also, very precisely, the silence we leave between two turns of speech in a human conversation.
That's why a conversation with a voice assistant always seems a bit strange. The problem doesn't necessarily come from the voice. It comes from the gap before the response. In a real conversation, a second of silence is already enough to create discomfort.
The hardest part isn't speaking, it's knowing when to be silent
Eliminating the relays creates another problem. The machine must guess something we do without thinking: have you finished your sentence or are you just searching for your word?
All the difficulty of the live interaction lies in this question. A human solves it without thinking. A machine struggles with it.
The Microsoft model would offer two methods for cutting in. With the first, it decides alone by analyzing the meaning of your words. A component designed by Microsoft then takes care of identifying the ends of sentences. The second method combines a duration of silence with a meaning analysis. It is more predictable and aims at those who want the model to react consistently, rather than letting it improvise.
This setting shows who the product is aimed at. You don't offer two turn-taking management modes to someone who just wants to ask for the weather. You offer them to a company that plans to install this system in its customer service and precisely adjust when the machine should be silent.
Except that OpenAI got ahead a month ago
It's impossible to talk about MAI-Realtime without looking at what everyone can already use. On July 8, OpenAI launched GPT-Live. The principle is the same: a voice model that listens and speaks at the same time. No early access or closed lab. It's already in ChatGPT, everywhere and right now.
A month apart, but above all two very different places. On the left, it’s already useful. On the right, it’s visited in secret.
The detail that hurts Microsoft is availability. GPT-Live-1 is offered to paying subscribers, GPT-Live-1 mini to free accounts. In other words, even without paying a cent, you can test full duplex tonight on your phone. Full duplex is simply the ability to listen and speak at the same time. You just need to press the voice button on ChatGPT.
The video that made the rounds, and what it really tells
OpenAI released a demonstration that is worth watching, “This is the new ChatGPT Voice”. We see three elderly women chatting with the assistant. They interrupt each other, hesitate, then continue. And it works. The machine says “mmh” and “yeah” while they talk. It goes silent when they think, then resumes when they finish. The result is strikingly natural.
The choice of these three ladies is obviously not innocent. The demonstration shows that the voice also works for people who do not type on a keyboard. It’s hard to find a better argument for this technology. The Register however pointed out, with the delicacy we know, another detail. This age group is also the target of phone scams facilitated by this kind of tool. Both observations are true at the same time. That’s the problem.
So, does Microsoft do better?
Short answer: no. On the only criterion that matters to you, Microsoft is actually very far behind. GPT-Live runs in the pockets of hundreds of millions of people. MAI-Realtime remains hidden in an internal tool, with no date, no price, and no confirmation.
In terms of capabilities, the two models are very similar. Microsoft still has two or three arguments to make.
- What OpenAI has in addition: availability, of course, but also a well-thought-out trick. When your question requires real research, GPT-Live hands it off in the background to a more powerful model, GPT-5.5. In the meantime, it continues to chat. This avoids a ten-second silence as soon as you ask a complicated question.
- What Microsoft has in addition: seventeen announced languages, two methods for managing turn-taking, thus deciding who speaks and when, as well as a panel that displays response times live. None of this excites someone who simply wants to chat with their phone. However, a company that wants to connect the model to a phone standard will see the immediate interest.
- The real gap: GPT-Live does not yet offer an interface for developers, meaning no official way to integrate it into their own applications. This interface is promised, but it does not exist yet. Microsoft can therefore still access the professional door, not the public one. With its settings and debugging panel, which is used to identify technical problems, its model strongly resembles a product designed for this purpose.
My impression, and I give it for what it’s worth: OpenAI has won the showcase, Microsoft plays the back office. It’s not the most glorious position. It’s often the most profitable.
The other way to talk to your machine is to make it write
While Microsoft hides its model and OpenAI makes grandmothers chat, another use of voice is already working very well on your machine. And curiously, almost no one talks about it. It’s dictation.
These are not the same difficulties. A full duplex model, capable of listening and speaking at the same time, must manage turns of speech. It must understand when you have finished and know when to be silent. Dictation only goes one way: you speak, then a clean text appears where your cursor blinks. No conversation to hold, so no turn of speech to negotiate. The problem is simpler. That’s why it is already solved.
What Wispr Flow does, and why it surprises
The most well-known of the bunch is called Wispr Flow. The principle can be summed up in one sentence: you hold down a key, you speak, you release it, then the text appears in the application you were using. Any application. A search field, an email, VS Code, or the chat window of your favorite AI.
The selling point is speed. 220 words per minute by voice, compared to 45 by keyboard. The number is nice, but it’s not what really changes the usage.
The real difference is the cleaning up. You dictate as you speak, with your “uhs,” your “well,” your “no wait, let me start over.” None of that appears on the screen. The AI model recognizes that you corrected yourself, discards the first version, and keeps the correct one. It also adds punctuation and adjusts the tone to the application you are writing in. The old dictation faithfully copied every hesitation. You then had to correct everything, line by line. You might as well type.
The whole product is in this funnel. What goes in is a spoken draft, what comes out is written.
The difference with the old dictation does not come from the recognition of words. This part has been working for years, and I discussed it in detail in the new transcription models from OpenAI. The difference comes from what the software does with the words once it has recognized them.
Wispr Flow works on Mac, Windows, iPhone, and Android, in over a hundred languages. The company raised 81 million dollars for this. Personally, I have been using it for a few weeks to write my prompts, that is, my instructions to the AIs, my emails, and a good part of my notes. Going back to the keyboard for a long text now feels like a chore. It’s typically the simple tool whose usefulness you understand the day you uninstall it.
Two things that still pose problems
The first is the cloud. Wispr Flow does not have an offline mode: every sentence you pronounce goes to servers, then comes back as text. There is a privacy mode that guarantees that the audio and transcription are not stored. Very well. But that doesn’t change the route: your voice still goes out from your home. Before dictating a password or a medical report, think for three seconds.
The second is the price. And here, I complain. The subscription costs 15 dollars per month, or 12 dollars per month if you pay annually. That’s 144 dollars a year to transform voice into text. Taken alone, it’s not expensive for the service provided. The problem is that nothing is ever taken alone.
Each card is reasonable. It’s the stack that poses the problem.
It’s the subscription too many, just after the previous too many subscription. An AI for chatting, another for coding, online storage, an office suite, music, and now a subscription for speaking. Each one justifies itself very well alone. Once added up, they end up costing a rent. The software has shifted from “I buy a tool” to “I rent the right to use it.” I don’t remember being asked for my opinion.
Exit strategies, because there are some
Good news, you are not obliged to pay to dictate.
- The simplest solution is already installed at your place. On Windows 11, the keys
Win + Hopen voice input. On Mac, dictation can be found in the keyboard settings. It's free, it works, and it even adds punctuation. However, the system does not clean up your hesitations. Your "uhs" will therefore appear on the screen. - To enjoy cleaning without a subscription, VoiceInk is open source, which means its code is accessible. It runs entirely on your machine with a locally installed Whisper model and costs 29 dollars, one-time only. Nothing leaves your computer. Its drawback is its compatibility: it only works on Macs equipped with an Apple chip.
- Between the two, SuperWhisper allows you to choose between local processing and cloud processing. It offers more settings but also requires a bit more configuration before it becomes truly useful.
My advice, since I'm here: pay for a month of subscription to understand why people become addicted. Then, see if a local solution doesn't do 90 percent of the work with a one-time payment. In both cases, you will never type a long prompt on the keyboard again.
And what does this have to do with MAI-Realtime? Dictation is a simplified version of the problem that Microsoft and OpenAI are trying to solve. Getting a machine to write under dictation is settled. Making it hold a real conversation, with pauses, interruptions, and changes of mind, is much more complicated. The simple version already works, and you can use it starting tonight.
What does it change for you
Three things. The third one is not good news.
1. You will be able to interrupt your machine
Today, when your assistant launches a thirty-second response that does not match your question, you have two options. You either wait politely, or you stop everything to start over. With a model that listens to you while it speaks, you simply say "no, wait" and it stops. Like a human. On paper, it seems tiny. In practice, it changes everything.
2. Voice becomes usable when your hands are busy
This is the real promise. In the car, while cooking, doing DIY, or holding a child in your arms, the keyboard is not an option. Yet, using voice remains a chore today. We articulate like a radio presenter, we wait, then we repeat. Dropping to two-tenths of a second is the moment when talking to a machine stops being an exercise. We just talk.
3. Call centers are the real target, and it must be said
Let's look at who pays for this kind of model. It's not you. It's the companies that answer the phone all day. An assistant that can be interrupted naturally and does not leave awkward silences resembles a human on the other end of the line. That is precisely the goal.
I have an opinion on this, and I stand by it. Technical progress is real. The question of whether you are informed that you are talking to a machine is just as important. It is no longer a laboratory problem. It is a problem that can call you directly at home.
What I think about it
What strikes me the most is the discretion. According to this leak, Microsoft is testing a model capable of holding a natural conversation in seventeen languages. Yet, it appears without a technical sheet, in a corner of an internal tool. Two years ago, we would have had a conference, a polished video, and three hours of demonstration.
There are two possible readings. Either Microsoft is not yet convinced that the model holds up and prefers to test it discreetly. Or this kind of release has become so mundane that it doesn't even deserve a press release. This second hypothesis seems the most probable to me. It is also the most dizzying.
I will not rejoice too quickly. A model spotted in an internal playground is not an announced product. Between the two, there remain pricing, safeguards, availability in Europe, and that good old habit of burying projects without warning. Let's meet in six months to see if Victoria and Grant still exist.






Join the conversation
You need an account to comment on this article. Creating one is free and takes under a minute.
No comments yet.