CEOInterviews.AI
Start App
Mustafa Suleyman
Executive Vice President & Chief Executive Officer, Microsoft AI, Microsoft AI

Mustafa Suleyman: The AGI Race Is Fake, Building Safe Superintelligence & the Agentic Economy | #216

📅 Dec 16, 2025 Peter H. Diamandis 84 MIN 106545 VIEWS 204 SEGMENTS · 5 SPEAKERS
Get access to metatrends 10+ years before anyone else - https://qr.diamandis.com/metatrends Mustafa Suleyman is the CEO of Microsoft AI (https://microsoft.ai/) Dave Blundin is the founder & GP of Link Ventures Dr. Alexander Wissner-Gross is a computer scientist and founder of Reified Chapters: 00:00 - The Evolution of AI and Microsoft's Role 09:12 - Economic Benchmarks for AI Agents 18:37 - The Future of AI in Science and Engineering 27:54 - AI Alignment and Containment Strategies 36:42 - Anthropomorphism and AI Personhood 46:10 - The Role of AI in Government and Society – My companies:...

Questions asked in this interview

12
  1. 0:00What's the mandate from Satya? Is it win AGI?
  2. 6:22Is it like the old adage, you can't go wrong buying IBM in the old days?
  3. 11:19So, what would be the first model to make a million dollars, given $100,000 in starting capital?
  4. 14:16I remember really clearly, after DeepMind got acquired by Google, what was the price tag on that deal?
  5. 21:52What do you think happens and when?
  6. 29:41Does that fit your model too?
  7. 33:39So there was a whole team here already working on it or did you bring in your team?
  8. 43:09I think we want things to feel ergonomic, right?
  9. 55:05Is there a Three Mile Island-like event that scares everybody but doesn't kill anybody?
  10. 59:40How are we doing on that?
  11. 1:12:53This window of opportunity is so short and so acute and it's really clear how you succeed right now in AI post-AGI. I mean, who could predict?
  12. 1:18:41and what about agentic AI in the government in particular?
Interviewer 0:00 ↗
What's the mandate from Satya? Is it win AGI?
Mustafa Suleyman 0:03 ↗
I don't think there's really a winning of AGI. I'm not sure there's a race.
Interviewer 0:07 ↗
One of the OGs of the AI world, Mustafa Suleyman is the CEO now of Microsoft AI. He spent more than a decade at the forefront of this industry before we even had gotten to feel it in the past couple of years. Now,
Mustafa Suleyman 0:23 ↗
Fundamentally, the transition that we're making is from a world of operating systems, search engines, apps, and browsers to a world of agents and companions. We're all going as fast as we possibly can, but a race implies it's zero sum. It implies that there's a finish line, and it's not quite the right metaphor. As we know, technologies and science and knowledge proliferate everywhere, all at once, at all scales, basically simultaneously.
Interviewer 0:52 ↗
Are you spending a lot of your energy, compute, human power on safety?
Mustafa Suleyman 0:57 ↗
Yeah. No, I mean,
Interviewer 1:00 ↗
Now that's a moonshot, ladies and gentlemen.
Host 1:05 ↗
Everybody, welcome to Moonshots. I'm here with DB2 and AWG and Mustafa Suleyman, the co-founder of DeepMind, Inflection AI, and now the CEO of Microsoft AI. Welcome, my friend. It's good to have you here. Thank you for making time for us.
Mustafa Suleyman 1:21 ↗
Thanks for having me. Yeah, I'm excited to do this.
Host 1:23 ↗
Yeah, it's, you know, what you've been building with Satya is amazing. And it's hard to believe that Microsoft is 50 years old and it's reinvented itself so many times and for the last 5 years it's been at the top of the game, the most valuable company in the world, 250,000 employees, and from what I understand 10,000 employees now under you. So a few important questions I want to open with. First, some broad context. You're building inside a massive company with huge resources, probably arguably more than almost everybody else. And the question I have is, what's the end goal here? You've got all the hyperscalers sort of providing open access to AI and they're doing a land grab to try and get as many users as possible. You've been building within the Microsoft 365 ecosystem. Is the goal in the next couple of years maximum users? Is it data centers? Is it cloud? How do you think of what you're optimizing for?
Mustafa Suleyman 2:41 ↗
I mean, it's a good question. So, we are on any given day a $4 trillion company with almost $300 billion of revenue. It's incredible. It's just surreal and very, very humbling. And we play at every layer of the stack. Obviously, we have an enormous business in data centers and in some ways we're like a modern construction company, hundreds of thousands of construction workers building gigawatts a year of CPU and AI accelerators of all kinds and enabling that to be available to the market. APIs on top of that, but also first-party products in every domain you can think of, from gaming and LinkedIn right the way through to all the fundamentals of M365 and Windows, and of course in our search and consumer businesses. And fundamentally, the transition that we're making is from a world of operating systems, search engines, apps and browsers to a world of agents and companions. All of these user interfaces are going to get subsumed into a conversational agentic form. These models are going to feel like having a real assistant in your pocket 24/7 that can do anything that has all your context. You're going to do less and less of the direct computing, just as we're seeing now. Many software engineers are using assistive coding agents to both debug their code and also generate large amounts of code, just as we used third-party libraries. Now we're just going to use AIs to do that generation, and it's making them more efficient, more accurate, and faster. So the trajectory we're on is quite predictable. It's one from user interfaces to AI agents, and that is a paradigm shift which the company is completely focused on. After seeing five decades worth of transitions, I think the company is super alert to making sure that we're best placed to manage this one.
Host 4:48 ↗
Do you see yourself providing sort of an open-source AI like the other players out there, or do you think you can keep it contained within Microsoft 365?
Mustafa Suleyman 4:58 ↗
I think we're pretty open-minded. We've got some pretty small open-source models. I think realistically,
Host 5:04 ↗
When I say open source, I really mean open access, if you would.
Mustafa Suleyman 5:07 ↗
Yeah. I mean, look, there are always going to be APIs that provide incredibly powerful models. Microsoft is really a platform of platforms. Being a platform and being a great provider of the core infrastructure that enables other people to be productive is the DNA of the company. We will always have masses of APIs that turbocharge that. But what an API is is going to start to look kind of different too. The distinction between the API and the agent itself may be pretty blurred. Maybe we're principally in the business in 5 years time of selling agents that perform certain tasks that come with a certification of reliability, security, safety, and trust. That is in many ways the strength of Microsoft, and that's one of the things that attracted me. This is a company that's incredibly trusted, it's actually very secure, and sometimes I think the slowness or the friction is actually a bit of an asset. There's a kind of steadiness that comes with having provided for all of the world's biggest Fortune 500 companies and governments and major institutions.
Host 6:22 ↗
Is it like the old adage, you can't go wrong buying IBM in the old days?
Mustafa Suleyman 6:26 ↗
I think there's a steadiness about us which I think is reassuring to people, and there's a kind of deliberate customer-focused patience. There's not the same anxiety and somewhat sclerotic nature that comes with being an insurgent. There are some downsides to our position. We take a little longer to get things through, but the company is firing on all cylinders. It's very impressive to see. One more question before I turn it over to Alex. We're seeing in this hyperscaler war, literally week by week, everybody outdoing each other in this insane period of everybody coming out with new benchmarks. Do you miss not being in that game, or is the stability that Microsoft provides to build for a long-term vision what you find most exciting? My background at DeepMind is such that I spent a good decade grinding through the flat part of the exponential where basically nothing worked. There were some amazing papers, AlphaGo was obviously incredible, but it was in a very unique simulated controlled game-like environment. Things actually working in the real world were few and far between. I've always taken a multi-decade view, and that's just been my instinct. I think it's super important to ship new models every month and be out there in the market, but it's actually more important to lay the right foundation for what's coming, because I think it's going to be the most wild transition we have ever made as a species.
Host 8:14 ↗
Can you just flesh that out a little bit? Was there a period of time where it was just three of you grinding it out in London?
Mustafa Suleyman 8:18 ↗
Well, there were more than three of us, but for the decade between 2010 and 2020, there were just so few successful commercial applications of deep learning. There were plenty behind the scenes, image recognition, improvements to search, but commercial... Playing Go is not a huge market. Whereas now, from 2022 onwards, LLMs in production are changing what it means to be a human. We hit an inflection point. That is very different to the grind of training tiny models with very little data and very small clusters back in the 2010s.
Narrator 9:11 ↗
Every week, my team and I study the top 10 technology meta trends that will transform industries over the decade ahead. I cover trends ranging from humanoid robotics, AGI, and quantum computing to transport, energy, longevity, and more. There's no fluff, only the most important stuff that matters, that impacts our lives, our companies, and our careers. If you want me to share these meta trends with you, I write a newsletter twice a week, sending it out as a short two-minute read via email. And if you want to discover the most important meta trends 10 years before anyone else, this report's for you. Readers include founders and CEOs from the world's most disruptive companies and entrepreneurs building the world's most disruptive tech. It's not for you if you don't want to be informed about what's coming, why it matters, and how you can benefit from it. To subscribe for free, go to dmandis.com/metats to gain access to the trends 10 years before anyone else. All right, now back to this episode.
Co-Panelist 10:05 ↗
Yeah. So when last we spoke circa 2015, I think that was perhaps 3 years post ImageNet, 5 years pre language models, few-shot learners, agents, agentic AI was nowhere to be seen at the level of what we see now. Since you've written about your vision, what you've socialized as a modern Turing test, the idea of economic benchmarks for autonomy by agents, I'd love to hear: Where are Microsoft's economic benchmarks for these agents? If the agents are about to take over the economy or take over so many economically useful functions, why are we stuck with benchmarks like MNIST rather than Microsoft leading the way with Microsoft's economically autonomous benchmarks for its agents?
Mustafa Suleyman 10:54 ↗
Yeah, it's probably just worth adding the context that we met in 2015 in Puerto Rico at the AI safety conference.
Co-Panelist 11:02 ↗
True, many of the field now were there at the same time. Seminal moment.
Mustafa Suleyman 11:04 ↗
Yeah.
Co-Panelist 11:05 ↗
Was it the day after New Year's Eve or somewhere around New Year?
Mustafa Suleyman 11:09 ↗
It was pretty cold out everywhere except Puerto Rico.
Co-Panelist 11:12 ↗
Yeah, exactly. It was pretty cool. It was quite a surreal moment actually.
Mustafa Suleyman 11:14 ↗
It was like a calm right before it all happened.
Co-Panelist 11:19 ↗
Yeah, totally. And the modern Turing test was something I proposed, I guess it was 2022 when I wrote it. It was basically making a pretty simple prediction. If the scaling laws continue with more data and compute, adding an order of magnitude more compute to the best models in the world every year, then it's pretty clear we would go from recognition, which was the first part of the wave, to generation, which we're now in the middle of, or maybe ending that chapter, to then having perfect generation at every time step, which in sequence is going to produce assistive, agentive actions. Actions would obviously look like an intelligent knowledge worker or a project manager or a strategist or a startup founder. So how would we measure that performance? Rather than measuring it with academic and theoretical benchmarks, one would clearly want to measure it through capabilities. What can the thing do in the economy, in the workplace? And how do we measure the economy? We measure it by dollars and cents. So, what would be the first model to make a million dollars, given $100,000 in starting capital?
That's right. Yeah. Which model could turn it into a million dollars?
Mustafa Suleyman 12:34 ↗
10x return on investment by an agent. Exactly. I think that's a pretty good measure of performance and capability. Certainly, we've kind of just breezed past the Turing test, right? It kind of has been passed. No one's really done a big, you know, AlphaGo silver prize wound down before we breezed past Turing.
Co-Panelist 12:58 ↗
Yeah. And no one celebrated it. Where was the big Kasparov vs. Deep Blue moment?
Mustafa Suleyman 13:04 ↗
Can we clink virtual glasses right now and celebrate that we won? It happened.
Co-Panelist 13:09 ↗
Yeah. Exactly. And that's what it feels like to make progress in a world full of these compounding exponentials where we just get desensitized to 10x. So much so that you can be like, 'Guys, why haven't you done it yet?'
Mustafa Suleyman 13:22 ↗
Yeah.
Co-Panelist 13:23 ↗
We're spoiled. Where's my Microsoft Loebner Prize for the modern Turing test?
Mustafa Suleyman 13:28 ↗
Right? Exactly. Someone said to me earlier on, 'But this AI thing, it's still in its infancy, isn't it?' And I'm like, man, if this is infancy, wow. I can talk to my computer fluently. Star Trek is here in real time. At the same time, agents don't really work yet. The action stuff is still progressing. It's getting better and better every minute, but it's pretty clear that in the next couple of years, those things come into view and they're going to be very, very good.
Co-Panelist 14:00 ↗
Can we get together again after the modern Turing test has been passed and just to celebrate, recognize it?
Mustafa Suleyman 14:05 ↗
Virtual glasses again. Absolutely.
Co-Panelist 14:08 ↗
Hopefully we can pop a champagne or something.
Mustafa Suleyman 14:10 ↗
I think we should have an optimist pop the cork for us or something.
Co-Panelist 14:14 ↗
Exactly. Exactly.
Host 14:16 ↗
Dave, hey, I want to flush out that backstory a little bit more too. It's such a cool story. I remember really clearly, after DeepMind got acquired by Google, what was the price tag on that deal? It was like half a billion dollars, something like that.
Mustafa Suleyman 14:28 ↗
650.
Host 14:30 ↗
650. What year was that?
Mustafa Suleyman 14:32 ↗
2014.
Host 14:33 ↗
2014. I remember reading maybe a year or two later that Google justifies deal by having DeepMind tune the air conditioning in the data centers.
Mustafa Suleyman 14:42 ↗
Yeah. Right.
Host 14:43 ↗
My interpretation of that was like, 'Wow, this isn't going all that well.' And now it's obviously the biggest thing that's happened in the history of humanity and forking out all over the place.
Mustafa Suleyman 14:52 ↗
I mean, we did the data center thing was pretty cool. We did actually reduce the cost of cooling the Google data center fleet by...
Host 14:58 ↗
Yeah. It's so funny because I read it at the time and I was like, 'What a bust.' And then I read about it in Wikipedia on the flight over here to meet with you and it's like it was actually 500 attributes fitting into the neural net and it was a lot more complicated than the news made it sound at the time.
Mustafa Suleyman 15:12 ↗
That's right.
Host 15:14 ↗
But like you were talking about the flat part of the exponential and you think about all of this R&D which is so close to becoming AGI is tuning the air conditioning. But that's the nature of exponentials. They sneak up on you like this. But the other way to think about that is that it's basically taking an arbitrary data input, an arbitrary modality, and using the same general purpose method to produce very accurate predictions in a novel environment, which is the same thing that's happened with text and audio and image and now coding and with other time series data. It's just another proof point of the general purpose nature of the models. I think it's so easy to get caught up thinking five years is a long time.
Mustafa Suleyman 15:55 ↗
It's like a blink of an eye. It's a drop in the ocean. I think because we're such a frantic second-to-second news culture, social media type environment, we just don't have an intuition for these time scales. I think other cultures do, and historically before digitalization, we had much more of a natural intuition for the movement of the landscape and the seasons and the ages. And now we're just like, 'It's not coming quick enough.' It's like, dude, it's coming. We've shifted to a 24/7 operations. I know very few people, including this group, that are operating around the clock every day just because when we do a Moonshot podcast week to week to celebrate and talk about what's just happened, it's insane on a week-by-week basis what's going on.
Co-Panelist 16:44 ↗
Yeah. Yeah.
Host 16:44 ↗
You know, and Peter's always saying people are very, very bad at exponentials, right? 100,000 years of evolution has us predicting tomorrow will be like yesterday.
Mustafa Suleyman 16:54 ↗
But you're one of the few people who, having lived through that, air conditioning becomes AGI in just a few years. So where we sit right now is on another inflection point and the implications are massive and people are way underreacting across the board. You're one of the few people who, having seen it before, can say I just got very lucky. We were very lucky to have an intuition for the exponential. That's a very powerful thing because we can all theoretically observe the shape of the exponential, but to go through the flat part and then get excited by a micro doubling, that's the bit. When you're like, 'Oh my god, I remember this.' The MNIST image generation thing, the first generative models, there's like these, I can't remember, maybe 256 by 256 pixels, black and white, handwritten digits. I think this was like 2013, maybe even 2012, and this guy, I think maybe he was employee number five at DeepMind, Dan Vesta, this awesome Dutch guy out of EPFL, generated the first number seven that was provably not in the training set for the first time. I was like, man, that is amazing. How could it have learned something about the idea of seven? That was it. It's got a concept of seven. How cool is that?
Co-Panelist 18:24 ↗
You know, I got the highest score on MNIST ever in 1991 when it first came out, when you were three years old, right?
Mustafa Suleyman 18:31 ↗
Yeah. Nine. You're nine years old. Okay. And actually that's the same data set that's now in PyTorch that people benchmark off.
Co-Panelist 18:40 ↗
Pretty crazy. Incredible.
Mustafa Suleyman 18:41 ↗
Yeah. How often are you surprised by what you're seeing? How often is there a Move 37 sort of aha moment? Is it happening more frequently?
Co-Panelist 18:54 ↗
I was absolutely blown away by the first versions of LaMDA at Google. This was maybe 12 people working on it led by Noam Shazeer and Daniel De Freitas and Quarkley, and I got involved later maybe three or four or five months after they'd been going. It was just breathtaking. Everyone at that point had been playing with LLMs and they were one-shot that produce an answer and have a prompt, but they were really the first to push it for conversation and dialogue. Seeing the kind of emergent behaviors that arise in yourself, things that you didn't even think to ask because it's going to be a dialogue rather than a question-answer situation, sounds so trivial to say in hindsight because now we're obviously steeped in conversation as the default mode. But that was breathtaking for me. Obviously then I pushed really hard to try and ship that at Google, and for various reasons we couldn't get it launched. That was when we all left. I left, Noam left to do Character, David Luan left to do Adept. We were all like, okay, this is the moment. I think there's been a couple moments since then, but that was probably the biggest one I remember in recent memory. Mind-blowing. And the scaling laws have delivered such unexpected performance. Going back to your earlier days, did you anticipate the kinds of capabilities that have resulted? Was this predictable for you, or is it still like, wow, what it's able to do in medicine, in conversation, in scientific research?
Well, especially working off of pure text. How far we've gotten. Nobody, I think, would have seen how far we would get with just text.
Mustafa Suleyman 20:44 ↗
Yeah. In 2015, I collaborated with a bunch of really awesome people on an NLP deep learning paper at DeepMind where we were essentially trying to predict a single word in a sentence. We had scraped Daily Mail news articles and CNN articles and we were like, can we fill in the blank, just predict one word in a sentence or complete the final word in a sentence, the inverse of the problem the way the models now work. It was a pretty big contribution and a well-sighted paper, but it was like, this is never going to scale. We were just like, okay, we're way too early. Not enough data, not enough compute. But we were still optimistic that with more data and compute, that is a method that will work. I don't want to have hindsight bias and say it was all very predictable, but everyone in the field just had the same hammer and nail and kept chipping away. Can we add more data to this? Can we clarify our prediction target and can we add more compute? Broadly speaking, that's what's delivered.
Co-Panelist 21:52 ↗
Yeah. We'd love to maybe pull on that theme a bit. You mentioned how surprising generating 7 from MNIST was. You mentioned how surprising the success of LaMDA for conversational tuning and conversational performance in general is. I think you've already made a little bit of news in this episode, if I understood correctly, with the expectation that in the next 2 years, so I read that as 2027, we'll see agents start to pass your modern Turing test. We'll see them be able to 10x $100,000 return on investment. I'm curious about the next surprises to come. AI for science. Microsoft Research has an AI for science initiative. Do you have timelines in your mind for AI solving math, which we're seeing a whole bunch of startups right now tear through Erdős problems? AI for physics, chemistry, medicine, material science. What do you think happens and when?
Mustafa Suleyman 22:48 ↗
Yeah, actually you've just reminded me, the more recent thing that has blown my mind is the fact that these methods could learn from one domain, coding puzzles, math, the essence of logical reasoning. Just as it learned the essence or the conceptual representation of a number seven, it's clearly learned the abstract nature of a logical reasoning path and then can apply that to many other domains. That's kind of interesting because it can apply that as well as the underlying hallucination/creativity sort of instinct that it has, which is more like interpolation. Those two things combined are a lethal combination for making progress in new mathematical theorem solving or new scientific challenges, because that's basically what humans do all the time. We combine these two capabilities. I couldn't really put a date on those things. It's hard to put a date because they are very fundamental, but it feels like they're definitely within reach. It would be very odd to bet against them.
Co-Panelist 24:04 ↗
Just maybe from an over/under perspective, do you think solving science and engineering for some reasonable definition of solving is going to ultimately be harder or easier than the modern Turing test, 10xing return on investment?
Mustafa Suleyman 24:22 ↗
It's going to be harder because I think a lot of the training data for strings of activity in the workplace or in entrepreneurialism, startups, that kind of exists in a lot of the log data and also it lends itself naturally to real-time calibration with a human. The AI can check in, the human can oversee, intervene, steer, and calibrate. It's going to be a much more combined effort between AI and human, where a human is participating in steering the reinforcement learning trajectory. Whereas in a novel domain where it really is inventing completely new knowledge, that's happening in a very abstract vector space and it's unclear yet how the human is going to intervene in the theorem solving problem. Everyone's working on this, particularly in biology and synthetic materials, because you want to give humans a better intuition for where in the search space to look for new hypotheses for drugs or for materials. Then the human can take or reject that, feed that back to the model, go and test it in silico, run the experiment, and feed that back into the model to improve the search.
Co-Panelist 25:42 ↗
And maybe a follow-up question: what can humanity in general, Microsoft specifically, or the AI community subset of which listens to the podcast, do to accelerate AI for science and accelerate the solution to science, math, engineering with AI?
Mustafa Suleyman 25:56 ↗
I mean, arguably that would be one of the most impactful things for humanity that would fundamentally move everything at light speed.
Co-Panelist 26:05 ↗
Yeah.
Mustafa Suleyman 26:05 ↗
I think it's already happening very organically. This is not only the most powerful technology in the world, it's also the fastest proliferating in human history. The cost of access, the cost of inference coming down by multiple orders of magnitude every couple of years...
Co-Panelist 26:25 ↗
Would you ever have imagined it would be so cheap?
Mustafa Suleyman 26:27 ↗
That bit I also totally got wrong. It's like the biggest surprise for me isn't that we're getting this level of capability. It's how cheap it is, how accessible it is.
Co-Panelist 26:36 ↗
100%. I mean, that's a thousandx over two years. Is it going to do that again or was that a one time?
Mustafa Suleyman 26:42 ↗
Is it a thousand? I think it's like a 100x. The inference cost has come down. A single token inference cost I think has come down 100x in the last two years.
Co-Panelist 26:49 ↗
Last two years. Okay. There have been competing estimates. Some estimates measure intelligence per token per dollar. There's an estimate that it's 40x year-over-year, but that's for certain weight classes of models. I've seen a 1000x for some classes of models. Craziness.
Mustafa Suleyman 27:04 ↗
Oh, wow. That's wild. Yeah. I got that totally wrong because I didn't think that the biggest companies in the world were going to open source models that cost billions of dollars essentially to train. So much so that when we founded Inflection, maybe 9 months or a year before ChatGPT was released, we started doing fundraising a year before ChatGPT was released. We basically raised a billion and a half dollars with a 25 person team to build what at the time was the largest H100 cluster with Nvidia and CoreWeave. We were CoreWeave's first AI customer. They were previously in crypto and we were their first AI customer working with them to build our data centers. Nvidia got behind us. I think we built a cluster at the time of about 15,000 H100s, growing to 22,000. Then that year ChatGPT came out and a few months around that time Llama came out. We were like, 'Oh my god, our entire cost base of our company has just been undermined by the fact that open source, it seems like open source is going to... it's not really about performance, it's just cost.' So then Perplexity, for example, founded after the arrival of Llama, knowing that they could depend on Llama and obviously OpenAI as an API and all the other APIs, and so they had a much lower cost base. That was another thing that was not predictable. I mean, other people predicted it, to be clear, I just got it wrong.
Co-Panelist 28:50 ↗
Abundance, baby. Demonetization, democratization of the most powerful tools in the universe. Hyper-deflation, if anything.
Mustafa Suleyman 28:59 ↗
Hyper-deflation, yeah. I think that's a really important point. The cost of accessing knowledge or intelligence or capability as a service is going to go to zero marginal cost. Obviously that's going to have massive labor deflation, displacement effects, but it's also going to have a weirdly deflationary effect because people aren't going to have dollar-based incomes to go buy things. That's obviously bad, but the cost of consuming stuff is also going to come down. We actually have a transition mismatch because labor markets are going to be affected before the cost of services comes down. Maybe there's a 10 to 20 year lag between that, which is going to be very destabilizing.
Co-Panelist 29:41 ↗
Which by the way is what we started to talk about a little bit earlier. I posit that in the long term there's an extraordinary future for humanity where access to food, water, energy, healthcare, education is accessible to every man, woman, and child. It's the shorter term that is challenging, the 2 to 7 year time frame. Does that fit your model too?
Mustafa Suleyman 30:06 ↗
Yeah, the short term I think is going to be quite unstable. The medium to longer term, it's pretty clear that these models are already world class at diagnostics. We released a paper maybe four or five months ago now called the MAI Diagnostic Orchestrator. Essentially it uses a ton of models under the hood to take a set of rare conditions from the New England Journal of Medicine, rare cases that can't be easily diagnosed that the best experts do a weak job on, and it's like four times more accurate, roughly about 2x less the cost in terms of unnecessary testing. There's a study that came out of Harvard and Stanford looking at, in this case GPT-4, a physician by themselves, a physician with GPT-4, and GPT-4 by itself. It was incredible that if you left the AI alone, it was far more accurate in diagnostics than the human. We're biased in our thoughts and what we saw yesterday, our recent diagnosis. We got a lot of feedback after we released the paper because we only showed the AI on its own, the physician on its own. A lot of people wanted to see what it was like to have the physician and the AI, or at least the physician have access to Google search as well. That improves performance a little bit, but the AI still trumps by quite a way.
Host 31:32 ↗
Dave, what are you thinking?
Co-Panelist 31:33 ↗
Oh, so much. So, Microsoft, you've been here how many years now?
Mustafa Suleyman 31:39 ↗
Just a year and a half.
Co-Panelist 31:40 ↗
Year and a half. So you're but you feel like you're part of the, you're indoctrinated. So what's the mandate from Satya? Is it win AGI or is it be self-sufficient? What is the target?
Mustafa Suleyman 31:54 ↗
I don't think there's really a winning of AGI. I think this is a mis framing that a lot of people have imposed on the field. I'm not sure there's a race. We're all going as fast as we possibly can, but a race implies that it's zero sum. It implies that there's a finish line. It implies that there's medals for 1, 2, and 3, but not 5, 6, and 7. It's just not quite the right metaphor. As we know, technologies and science and knowledge proliferate everywhere, all at once, at all scales, basically simultaneously or within a year or two. My mission is to ensure that we are self-sufficient, that we know how to train our own models end to end from scratch at the frontier of all scales on all capabilities, and we build an absolutely world-class super intelligence team inside of the company. I'm also responsible for Copilot. That's our tool for taking these models to production in all of our consumer surfaces.
Co-Panelist 32:53 ↗
So just to clarify, when we look at Polymarket, which we do a lot on the podcast, the horse race to who has the best AI model at the end of the year and who has the best AI model at the end of next year, there's no Microsoft line on that chart, right?
Mustafa Suleyman 33:06 ↗
So now there will be, I assume.
Co-Panelist 33:08 ↗
Yeah, there will be. Next year we'll be putting out more and more models from us, but this is going to take many years for us to build this. DeepMind or OpenAI, these are decade old labs that have built the habit and practice of doing really cutting edge research and being able to weed out carefully the failures and redirect people. This is an entire culture and discipline that takes many years to build. But yeah, we're absolutely pushing for the frontier. We want to build the best super intelligence and the safest super intelligence models in the world.
Host 33:38 ↗
Yeah.
Co-Panelist 33:39 ↗
Nice. So when you arrived, if we go back to Inflection, the thesis there is 18,000 H100s. We're going to build a big transformer. We're going to take a transformer architecture, build... So I assume now you've got all the OpenAI source code and that was here. You probably looked at it a year and a half ago on day one when you arrived. Just like start scrolling, I guess. Trying to visualize how multi-decade billion dollars of R&D what it looks like and how it arrives in a building. But you just dropped right into it. So there was a whole team here already working on it or did you bring in your team?
Mustafa Suleyman 34:17 ↗
Yeah, all my team came over and obviously we've been growing that team a lot. We've hired a lot from all the major labs and we're very much in the trenches of the hiring wars, which are quite surreal. This is kind of unprecedented how that's working out.
Co-Panelist 34:29 ↗
Crazy.
Mustafa Suleyman 34:29 ↗
Yeah. Phone calls every day from all the CEOs to all of the other people. It's this constant battle. We're really building out the team now from scratch. That's pretty much how it's been.
Co-Panelist 34:41 ↗
10,000 employees under you now?
Mustafa Suleyman 34:42 ↗
No, no. The core super intelligence team is like a few hundred. That's really the number one priority. The rest of that is Copilot, the search engine. Along those lines, I just have to ask because the terms AGI and ASI, super intelligence, start getting thrown around in a very interesting fashion. Do you have an internal definition of AGI versus digital super intelligence here?
Yeah, I think very loosely, these are just points on a curve.
Co-Panelist 35:16 ↗
Are they interchangeable in your mind, AGI and ASI, or are they different?
Mustafa Suleyman 35:19 ↗
I think they're generally used as different. Different people have different definitions. The AGI definition, it's like the Turing test. It'll pass by and it'll be blurred and we will have recognized it in retrospect. Roughly speaking, at the far end of the spectrum, a super intelligence is an AI that can perform all tasks better than all humans combined and has the capacity to keep improving itself over time.

98 more exchanges in this transcript

Sign in free to read the rest of this interview. No card required.

Sign in to read the full transcript

Cite this transcript

APA, MLA, BibTeX
APA

Suleyman, M. (2025, December 16). Mustafa Suleyman: The AGI Race Is Fake, Building Safe Superintelligence & the Agentic Economy | #216 [Interview transcript]. Peter H. Diamandis. CEOInterviews.AI. https://ceointerviews.ai/interview/579351/

MLA

Mustafa Suleyman. "Mustafa Suleyman: The AGI Race Is Fake, Building Safe Superintelligence & the Agentic Economy | #216." Peter H. Diamandis, 16 Dec. 2025. Transcript, CEOInterviews.AI, https://ceointerviews.ai/interview/579351/.

BibTeX
@misc{suleyman2025_579351,
  author       = {Mustafa Suleyman},
  title        = {Mustafa Suleyman: The AGI Race Is Fake, Building Safe Superintelligence \& the Agentic Economy | #216},
  howpublished = {Interview transcript, Peter H. Diamandis. CEOInterviews.AI},
  year         = {2025},
  month        = {dec},
  url          = {https://ceointerviews.ai/interview/579351/},
  note         = {Speaker-attributed transcript with timestamps}
}