There's a lot of talk about agents right now as well. In fact, Microsoft Windows just now got OpenAI's Claude sort of direct integration. What can they do with agents that we couldn't do with just a regular chatbot before?
You can use those to organize your work on a day-to-day basis, track your health, run background tasks for you, find you the best price of things that you care about, things that you would previously write code for. Obviously an agent is now going to help you do lots of things. Lots of my friends have been using it to open their gates at home, control their HVAC systems. Someone was telling me the other day that he asked his agent to help him lose weight. And so it's going to be easier than ever before to be a small developer with a few people and a bunch of agents building high-quality stuff. The only test that we can put technology to is a very simple one: does it make our collective lives healthier and happier? Does it enhance human progress? This is something that would take two or three people to do pretty much full time. That is kind of a little superpower. People are going to be very excited that they're going to have a perfect health assistant in their pocket. I really think that we're close to medical super intelligence.
Mustafa Suleyman is one of the most influential people shaping the future of AI. He co-founded DeepMind and now he leads Microsoft AI overseeing Copilot and Microsoft's rapidly growing family of AI models. In this conversation, we talk about agents that can work for days, AI moving onto your devices, Microsoft's push for human-centered super intelligence, and why he believes world-class AI healthcare can be available to everyone within just a few years. And we talk about so much more. So, please enjoy this wide-ranging and fascinating conversation with Mustafa Suleyman. So, I'm excited to do this conversation again with you. This is the second time we've spoken. We spoke at the 50th anniversary and so it'll be exciting to see how things have evolved since we had that conversation.
Yeah, it's been a crazy year. It's wild, wild times in the industry.
Yeah. So, there was just a whole bunch of announcements. The Microsoft Build event, the main keynote just happened and it was just announcement after announcement after announcement. I'm curious, what do you think are the announcements that have the most impact on the most amount of people?
Yeah. I mean, I think our main motivation was to try to land the idea that we as a platform company are building a hill climbing machine. And that means at every layer of the stack, from our own silicon to our orchestrators, our harnesses, and of course our own models, we're trying to give developers all the flexibility they need to pick whatever they want so that they feel like they can build to the absolute frontier. And obviously we have our own internal Microsoft offering now between the Maia 200 chip and the MAI models that we're very proud of, but we're also partnering across the board with everyone you can imagine to make sure that developers feel like they've got the best ecosystem with us.
Mhm. Can you explain the hill climbing machine a little bit more? What does that mean? So in AI, once you have a clear evaluation and you know exactly what your prediction target is, you want to make sure that the data that you curate is tightly coupled to the evaluation that drives it. And the harness that sits around an agent that tries to improve itself with respect to that data towards that evaluation objective
is tightly coupled again to the full cloud environment that you use. And that's what we mean by the hill climbing machine. You can hill climb once you have those four components of the stack tightly integrated into an RLE, a reinforcement learning environment that ultimately produces an agent that is customized to your objective that you fundamentally control.
Got it. And so it's something that gets smarter and smarter as it learns more about your business and what you're doing and the types of things you use it for.
Yeah. And that's kind of the evolution that we've been on in the industry. The last three or four years have been about producing pre-trained models that have a kind of base level information about the world. They generally understand all the web data and a good amount of the books and PDFs and general purpose knowledge.
The next phase has then been about adapting those models to specific tasks using reinforcement learning, and that's where they kind of hill climb towards a specific objective over time.
I see. I see. So there's a lot of talk about agents right now as well. In fact, Microsoft Windows just now got OpenAI's Claude sort of direct integration. What does that open up for normal people now that we use agents? What can they do with agents that we couldn't do with just a regular chatbot before?
Well, increasingly you're going to be able to step out of the CLI and just use your agent as an app. I mean, you saw the companion app for OpenAI's Claude, but for many other agents that'll be in there too. And you can use those to organize your work on a day-to-day basis, track your health, run background tasks for you, find you the best price of things that you care about. Anything that you would think to give to an assistant or things that you would previously write code for, obviously an agent is now going to help you do.
Have you seen any killer use cases in your own experimentation?
Lots of my friends have been using it to open their gates at home, control their HVAC systems, set up automatic timers for things. Someone was telling me the other day that he asked his agent to help him lose weight, and so he has cameras inside of his kitchen and it was letting him know that he was heading back to the kitchen for a snack a third time in two hours or something. So look, there's something out there for everybody. I'm not sure that would be my choice, but it's pretty cool to see the flexibility.
Gotcha. Yeah. So he's going and opening his fridge and maybe he's getting a notification on his watch: are you sure you want to be doing that right now?
Very cool. You guys also talked about Web IQ. Can you quickly explain what Web IQ is and what that opens up for people?
I mean, obviously right now we have Work IQ, which means that for big enterprises, your model should be grounded on the knowledge that you've already got in your own workflows. Web IQ provides the flip side of that, which is that we want to have super fast, very efficient access to web content so that you could point your agent at any grounded external knowledge base and it should just immediately be able to provide good citations and accurate retrieved information from the external world.
Gotcha. And when you say the external world, does that also include the Microsoft suite of products? So it's pulling from Excel and PowerPoint and Word and all those things as well?
Yeah, from your personal OneDrive of course, so your corpus of documents. But also accessing the external web, so that'll just be regular real-time information from the web.
Gotcha. Okay, cool. And then also there was talk about Project Solera. That was something that I hadn't previously heard about. I don't even think that was in the brief that I was on earlier on. Can you explain that a little bit, what it opens up, because it kind of feels like it's putting agents all over the place. So, like your kitchen example, this is putting agents maybe even in your kitchen and you had the little badge thing as well. Can you explain what that is and what you see that opening up for people?
I mean, it's pretty incredible. The badge has been around for like 30 years or more and basically hasn't evolved at all, and yet most of us who go to work every day in a big company have to carry around a badge to open the doors and so on. So I think it's really exciting that there could be an ever-present, aware kind of sensor system on your badge that basically makes it easier for you to ask any question of your ever-present AI, with obviously camera and vision. So it's pretty cool and I like the bold experimental nature of it. I'm very excited to try and use it, and it's a very flexible platform so anybody can ship their own models, ship their own harnesses, integrate it in their own products. So it's really supposed to be a general purpose platform.
Right. Now, some of the examples you guys showed during the keynote were like in a doctor's office, which makes a lot of sense. That example makes a lot of sense. But do you see examples of just normal everyday consumers using it around their house, or is it mostly designed for business use cases?
I think it's an open platform, so who knows what people are going to do with it. But the first hypothesis was, there is this device that people carry around on their belt every day that is actually pretty important. I mean, it's secure, it grants you access in a pretty sensitive way, and so a lot of the basic infrastructure has already been built. And so the goal was basically to build on top of that and see what developers and different companies choose to do with it.
Gotcha. Now, my follow-up question around that is, we all have phones in our pocket. But what can we do with that that we can't do with our phone?
Well, I think it's ease of access, fast, stable, reliable. Obviously it's secure. I mean, technically speaking, you could also use your phone to unlock a door anyway, right? So I think that many of these deployments are going to be somewhat overlapping because clearly all your devices do some portion of all the capabilities of other devices. I think the main motivation is it would just be great to add a few more features and create an open platform with this project that people can experiment with.
Yeah, that makes sense. So you guys also mentioned seven new models, right? You've got a new thinking model, you've got a new image model, you've got a new voice model. Seven different new models. What can people do today now that these models are available that they couldn't do yesterday?
Yeah. So for us, this is about taking first steps towards true self-sufficiency in AI, and that means that we have to have the capacity to build our own models from scratch and show that we can achieve the absolute frontier. So for example, in transcription, it's the very best model in the world. It's also the fastest and the cheapest. Our image-to-image and image editing models are now number two and number three in the world, beating out everybody other than OpenAI at the top for GPT-4 image 2, and a couple of Google models for editing, but we're still better than Nano Banana Pro Nano Banana 2. So it's a pretty big deal, and I think that what we're trying to show is that taking steps towards building this hill climbing machine shows that we can actually create models that are of the absolute state-of-the-art, so that as models diverge in different directions, we can actually build exactly what's needed for Microsoft and our developers. One example of that is our code models are actually very very small and very very strong in highly inference efficient general purpose agentic use cases, and we've chosen that use case in our pre-train model, in our RL climb, and in our post-training, and that means that they're optimized for agentic coding and for the enterprise, which is different to Google and say OpenAI that also have to worry about general purpose consumers. So I think over the next few years, you're going to start to see more and more divergence in the lineages of these models. Because clearly general purpose will deliver, which means that one model that can deliver across all the domains, but highly specialized models are always going to be able to push that a step further. And you saw that in the work that we did on Excel and with McKinsey and with a bunch of other companies, showing that we can drive really outsized performance for 10x lower cost with very efficient in-house models.
Right. Right. Now, I think you used the wording, you mentioned it's all commercially licensed lineage.
What does that mean? And with it being commercially licensed lineage, does that have any negative impact on the output of the model compared to companies that might just use whatever they want in the training?
I mean, look, I think part of the challenge is that we don't always know exactly what's in other people's pre-trained sets. Today we've also released a 109-page technical paper sharing in great detail exactly what has been used across the stack: algorithms, training infrastructure, compute, GPUs, data content, data composition. We've shared actually a lot on the model design side of things as well. So we tried to be as transparent as possible, knowing that our main job is to build trust so that people come build on our platform, and that also extends to how we buy the data. We have paid a great deal for this data. We've licensed it very carefully. We've been very deliberate about not using open source data sets, right? We don't know what security bugs could be introduced through some of the more open source data sets that maybe haven't gone through the same rigor. Whereas for our customers, we have some of the largest governments in the world on our platform. We have many of the largest companies in the world. So we want to give them trust and confidence that the model that we give them is absolutely clean from top to bottom and has been acquired in exactly the right way.
Yeah. Amazing. So I'm curious about the bigger roadmap for your various platforms. You've got Copilot, you've got the new GitHub app that you guys showed off today, which is really cool. I'm actually really excited to go play with that one. But you obviously have deals with OpenAI, you guys have deals with Anthropic, you guys are kind of working with all these companies. So I'm curious what the bigger picture is here, developing your own models. Is it to be the new state-of-the-art, to try to pass them? What is the roadmap for all the models and you guys working with the other companies as well?
I mean, we are a global company that provides for small developers, enterprise developers, governments, and every kind of institution you could think of around the world. So as a platform company, our job is to provide optionality. If you come to Foundry, I think there's something like 11,000 models that are available. You also get the absolute frontier, so the best models in the world from Anthropic and OpenAI, and you also get the option to use our new in-house models which are growing and getting better and better. So the thesis is we shouldn't restrict anybody or try to lock anybody in or reduce their optionality. People want to make different choices for different types of experimentation in different settings. And that's really the core mission: optionality first.
Right. So people might find that the OpenAI model does really well for this specific use case for them, but maybe the MAI state-of-the-art model is better for this use case. So they could switch between them based on the use case that these models are best at.