Mustafa, so good to see you. Thanks so much for doing this. So you have an interesting announcement, a new model that is an AI transcription model called My Transcribe One, right? Even though there's a second version of some of the other language models, right? Tell me about it. I want to know what it is and what sets it apart from other AI transcription.
So, we now have a family of models, actually. So, we've got My Transcribe One, and this is now the best transcription model in the world on the top 25 languages. On the top 11, we're number one. The remaining 14, we're better than Gemini and OpenAI. Not only are we state of the art, best quality, but we're also two x more GPU efficient, which is a super important achievement. I mean, our primary objective is that we want to constantly save on cost, so that we can pass that on to all of our consumers and our enterprises. And so, in Azure Foundry, this is actually going to be the cheapest of any hyperscaler for transcription. And this is a huge workload that is blowing up across the company. Many agentic developers now who are using AI coding assistants are actually managing teams of agents, and often they're doing so in voice. We also know that on the consumer side, voices is going through the roof. We also use transcription in Teams, as well as facilitator and many other features that are using transcription. So, it's pretty key.
And so, there's tons of, as you may, I mean, tons of AI transcription software. I mean, now if I do a voice memo on my iPhone, I get a transcript. So, tell me what is different about this. What makes it better? What sets it apart in these benchmarks?
So, I think the key thing is that it's really good at dealing with noisy environments. So, if you have background noise because you're in an office, it's really good at detecting who the speaker is and pushing out that background noise. It's also really good at detecting different accents. So, it's much more diverse in that sense, regional accents and different languages. And, you know, I think that is very much testament to what the team did in terms of collecting a very large, very diverse, very high-quality data set. So, the model has seen a tremendous amount of past experience. As well as that, they've also trained an extremely efficient architecture. So, a lot of these speedups that we've been able to deliver come from the modeling team, in addition to also our inference serving stack as well, which has been optimized specifically for this. And we've done the same thing with voice generation as well. So, we now have state of the art voice generation models in many different languages and accents. And we've got state of the art image generation, too. So, we actually achieve the same quality as ChatGPT image 1.5, which is what we have live in our products. But we do it with two x fewer GPUs. So, it's lightning fast, best quality, priced at the very best in the market for any hyperscaler. And so, I think we've really managed to deliver on the three most important variables to enterprises, developers, startup founders, etc.
So, the idea is like you could be using something like Cursor, right? And you're a developer and you're building software, and you could say, I actually want to use these models for just the voice, like maybe I'm using Claude Code or something, or Claude on the back end, but I'm going to use these models for the voice back and forth because I think it's picking up the nuances of my accent or my language. Is that sort of the idea?
Exactly right. So, if you're a developer and you're building a new product or service, you'll obviously select which voice model to use, and the main things that matter to you is: is it the best in the world? Is it the fastest that anyone can deliver? This is two x faster, and is it the cheapest that's available? So, you know, we fully expect this to drive maximum adoption because of that. Everybody is now using voice, not just transcription, but generation because, you know, if you're coding up your own agents, you want to be able to talk to them on your headphones, when you're out busy, when you're in between meetings. And this is the new production function, and that's why we actually focused on it as the first step in our super intelligence journey.
Yeah, it's really interesting. And when you, I mean, we've all had these experiences, right? When you're talking to any kind of voice assistant, it's almost like you start to speak in a different way, right? It's not your natural voice because you want it to understand you. And that's almost like a mental load in itself, and sometimes can be kind of, I don't know, annoying, and it's just easier, you know, it's easier just to type it in. Are you able to, is the goal to sort of like be able to have a much more natural kind of like a conversation you'd have with a human being? Is that ultimately what you're trying to get to?
Exactly. Like it's already got really good expressiveness and prosody. So, prosody is like the fluency between words, so that one word rolls into the next, so that the pauses are correct. You know, all of that we've paid very careful attention to the smoothness of the delivery of the generation. And that's because we can see, even in our own team, many of our top developers are managing teams of agents, and you want to coordinate those just as you would with anyone else. So, the interrupt is key, but also the latency is critical. Like you want it to immediately stop generating or start generating at the right time. And of course, you want it to be affordable. So, you know, getting those three right has been the big challenge.
You know, when you think about this, it's like there's a lot of, you know, when you take voice and translate it to text, you realize like how compressed and like lossy that process is. Like there's so many nuances in voice, and I'm just wondering like where we are now. I mean, can it pick up, you know, sarcasm or if I say a word more loudly or emphasize it in a way, does it get that? Does it understand? Or is that something for the next version of these models?
Yeah, I mean, I think the way to think about it is that these are components in people's stacks. If you have a Claude agent, whatever your setup is, you know, you basically want to be able to embed transcription and voice generation in various different parts. Like in Teams, for example, like having real-live translation and transcription is critical across many languages. We also are going to have, you know, we have a facilitator agent that documents meetings, summarizes them at the end, and then is able to sort of send you a little voice-based summary of what happened if you missed a meeting, for example. So, these are the sorts of product features that we're applying this to later in the year, and I think it's going to really feel like a step change because your entire development environment is going to kind of start to be much more interactive and much more animated. You'll leave your desktop, you'll go to your laptop, you'll go to your phone, and you'll make a call and follow up on a task that you've been pursuing for a little while. So, they're very fundamental capabilities, and my motivation in starting here was that we wanted to deliver absolute state of the art. Like we are a frontier lab building our own AI self-sufficiency mission, pursuing full super intelligence in the future. And this is the first kind of waypoint along that journey, which we kicked off in earnest last autumn. You know, we founded the team basically in October or September. And that was really after we were able to renegotiate the contract with OpenAI, so that we could pursue our own independent super intelligence and self-sufficiency missions. And that's really what we've started in earnest over the last 6 months.