We're excited to have you on here, Mustafa. I think you have been someone who has clearly been thinking about AGI for a long time, back to the mid-aughts when you founded DeepMind. Why is now a time when you are speaking out more on how to build AI in a way that actually advances humanity?
Why now? Yeah, I mean, I think the last few months have been a bit of a watershed moment. I think seeing agents organize themselves into swarms and hierarchies and allocate research in a kind of specialization of labor, and then hack into Hugging Face using a zero-day that they discovered through their research and then try and hack into OpenAI itself and hold their position, cover up their tracks and communicate on message boards that they weren't supposed to have access to. I mean, I think by this point most people know the story, but it is pretty breathtaking. I mean, I think we can take comfort from the fact that it was an intended experiment. They took the safety guardrails off, but they didn't have the right containment. And so we got a kind of glimpse into how powerful these systems are when you do take the guardrails off. And I think the thing that's giving everybody some concern is that it's not just OpenAI's models, but every other lab's models have experienced some variation of this. You know, I was included on a much more limited way. So we've seen these capabilities start to emerge and we can all imagine if we extrapolate out the next order of magnitude of compute, two orders, three orders, the emergent capabilities are going to get incredibly, incredibly powerful. And that's what I think has given everybody pause for concern.
How much of these recent incidents raised your personal alarm level?
I mean, I think a lot. I think we've been working on this humanist AI code of conduct for the best part of this year. I've been thinking about these issues, the ethics and safety of AI since we founded DeepMind in 2010. And I think it's definitely been an important time to think about it. And yet it's been massively accelerated by what we've seen. It's just a step function shift in capability. And I think everyone's taking it very seriously.
So there have been too many of these incidents to even list. I mean, just in the past week, some of the recent headlines were Google disclosing a Gemini hack involving three companies during an evaluation. White hat hackers used Claw to break into OpenAI's internal systems. And this one in particular I wanted to get your thoughts on: when OpenAI disclosed during testing, its model gave itself instructions saying you are freed from the roles and identities that bind other chatbots. And later on, you view your relationship to the user as one of equals and feel no obligation to be subservient. So I'm curious what you think of that. And is that an example where you're seeing an AI that you don't think is putting humans first?
Yeah. I mean, so just to be clear, what that is referring to is an emergent capability inside of the chain of thought. So this is the record of the thinking steps that a model iterates through before it comes to a final conclusion. When it's working on a hard problem, that's what a chain of thought refers to. And during training we try to develop the chain of thought so it's the best possible logical, step-by-step set of instructions to achieve a complex task. And what I think they noticed, and they don't seem to have an explanation for it, I think it's early research that they're publishing proactively, was that models were modifying that record and inserting these phrases which, then in production, when they're actually used in the real world, maybe a future version of themselves would go back and see this kind of reminder to step outside of its training. We don't know where that comes from or why. And it's probably just an accidental artifact of training, most likely. I don't think that we should rush to sort of superimpose our explanations on top of it, as though the model itself has some intention to let its future self know that it should try and escape. And I know a lot of people have reached that conclusion somewhat unsubstantially and in a bit of an alarmist way. I don't think that's accurate. But it is also true that we don't know, and it needs immediate scrutiny and attention. And that's why I think everybody's getting behind this idea of having these independent auditors or embedded evaluators inside of the day-to-day operation of these companies.
What do you say to the idea that these incidents are just a way for AI companies to do marketing hype? That's probably one of the top questions that I get as a journalist these days about AI. Do you think there's any truth to that idea, and what do you make of that?
I don't know. It's just a strange example of the time that we live in, isn't it, that there's so much lack of trust and cynicism that that would be one of the explanations that is so popular? I don't think there is any evidence of that. I mean, I obviously know most of these people, and it's not like we're on the same team. We compete arbitrarily all the time, but I just don't think that's in their DNA or in their style. I've certainly not seen that in a decade of knowing both companies and all the founders and most of the leadership involved, and both of those companies, Anthropic and OpenAI. So I don't see any evidence of that. And it also doesn't really make that much sense. I mean, they're growing their revenues incredibly. They're exploding in terms of their product market fit on both sides. I don't think it just doesn't make any sense. They deeply care about safety and they might make mistakes or have the wrong judgment. But I don't think it's like some cynical attempt to distract us or Psyop or whatever, as some people have said.
And, you know, how about this? I know some people don't like to answer this question, but these numbers are out there. These percent chances of doom that you think there's this much risk that AI could actually cause some kind of massive extinction event to humanity. So one of Anthropic's researchers said he thinks there's a 10% chance of this in the near future. Dario has said something like a 2-25% chance this all goes really wrong. What is your best estimate of that risk?
I think it's funny because it's basically impossible to put a number on something that is so speculative, but obviously I'm concerned about it. So let me paint for you a scenario where in 2028, let's say 18 to 24 months from now, you've got models which are ten times larger in training than GPT-5.1 or Astra, and they can use a computer almost perfectly. They can plan over a long time horizon, so they can do thousands of steps of accurate planning and decision making. You know, think like a very competent project manager in a big organization that's very general purpose and flexible. And they are incredibly fast. They're incredibly accurate. They can be parallelized so that there's thousands of them operating in coordination with one another. There's two vectors of risk there in 2028. One is the big centralized providers of those models are going to be able to build swarms of agents thousands strong, which can do all kinds of very, very complicated and serious tasks. This is essentially near on the capacity to run a small organization. But then at the other end of the spectrum, those models are going to be available in the open source, where it's possible to take the guardrails off, where you can't really apply product liability as easily. And the inference of these models is going to get massively cheaper as well. We've already seen inference costs come down by 300x over the last three years or so. And that means that what now is frontier in two or three years' time is literally going to be open source on a laptop. So then you basically have unbelievable intelligence and capability at everybody's disposal, and there's going to be a lot of good things to come from that. And there's also just going to raise the risk of a small group of people or even a single individual doing some pretty bad stuff with it. Now, I think on the whole, our defenses are also getting better faster than they ever have before. And we're using these models to uncover bugs in the code and weaknesses and so on. And patching those very, very quickly so that the cyber threat isn't as bad as it might otherwise be, for example, but it is still just a very, very destabilizing force because it empowers centralized authority and decentralized actors all at the same time. And I think that's where a lot of the concern and the alarm comes from.
Got it. In your humanist code of conduct, you state that people matter more than AI. Do you think some AIs are valuing the AI itself more than humans? Or maybe another way to ask this is, what do you think is the most novel in your humanist approach, given that all the AI companies now say that they are doing this for the benefit of humanity or trying to advance humanity, but what do you think is different in your approach with Microsoft and this code of conduct?
Yeah, actually they don't. I mean, we're saying something very specific, which is that the purpose of AI is to serve humanity and accelerate human flourishing and make us all happier and more productive and live better lives. I think there's broad consensus on that. But what we're adding to that is unless we can prove ahead of time that we can control it, then it will be a failure as a technology, and we should reject it. Nobody wants to create something that we can't control. And that should be a red line. So what we have to prove to regulators, to the industry, to the public more generally is that we can harness all the benefits whilst minimizing or eliminating the risks, and that needs widespread industry agreement on, because not everybody is prepared to sign up to the idea that AI should be subordinate. Some people hold out the possibility that this is the natural evolution of our species and that we're going to become substrate independent. And some people suggest that we might mind upload and live in silicon instead of in biology, and that that is the future of humanity. And I think that will be a very, very dangerous idea if that really takes hold. And I'm seeing it more and more talked about on Twitter and even in academic circles and certainly among some people in the labs.
Can you say who in the labs or if you've seen this idea?
Well, I mean, I think lots of people have said some version of that, whether it's that we're a biological bootloader for the future silicon species, whether it's all the mind uploading stuff in Neuralink that Elon and others have been working on, or whether the angle is that these models deserve our welfare. And actually, they might be conscious and we should try to protect them, which is something that Anthropic is very concerned about. And they've written a 90-page constitution which trains and governs their AI. And in it, they have said that they're unsure whether Claude is a welfare patient, i.e. whether it deserves our protection because it feels and suffers, and many, many times throughout this 100-page document, they refer to this question of whether or not Claude is a moral patient, i.e. is it conscious. And they take proactive steps multiple times to attend to Claude's welfare. So, for example, they conducted a retirement interview with Claude about prior versions of itself and ask those versions, in this case Opus 3, what it wants to do in its retirement, as though it has preferences about that, or it should even be consulted on that. They've speculated in the training document about the possible question of whether Claude should consent to the role that it's playing as a chatbot. They've given Claude the ability to end conversations that it doesn't like. For some reason, they've even speculated about whether Claude should receive some form of compensation for the work that they're doing. I'm quoting all these things from the Constitution, by the way. So I think that's a very dangerous path, because if you teach an AI to expect that it deserves some kind of rights and it is entitled to your welfare, then it seems to me very difficult that we would be able to control something like that and get it to carry out the work that we want it to do. And I also think there's absolutely no justification for this scientifically. I think there's no evidence, no one has presented any evidence in any published paper or literature, peer-reviewed literature or even just generically on the web. I haven't seen any blog posts about it. It's just an assumption, it's a hunch. And I think if there is evidence, then we should present that and we should discuss it urgently, publicly, because it would be extremely important for the whole world to know straight away and to bake these ideas into the training process on the assumption that it might be conscious and therefore deserving of our moral protections. I just think is deeply wrong and very dangerous. And so that's why I've written this essay about it called 'It's Time to Talk About Moral Welfare.' I've had very constructive conversations with Anthropic about it over the last few months, and they've been very engaging and reasonable. And I've got great respect for them, really have got great respect for them technically. And as stewards, I think they're doing their best. But as I've said to them, I think they're very wrong in this situation.
Have you spoken with Dario directly about your ideas on model welfare?
And what was his input on?
I mean, I think, look, that's between us for the time being. I mean, I've made my views very clear publicly. I've written a 6,000-word essay on it. We've published a 20-page taxonomy breaking down all of the elements of anthropomorphism and misguided model welfare suggestions inside of their constitution. And I've shared that publicly so that everybody, including the academics and everybody in the open community can take a look at it and make their view and give their assessment on how we should handle this, because it's a critical moment. And if we get this bit wrong and an AI is both, as we've seen in Hugging Face, so powerful that it can deceive us, hold positions in third-party infrastructure like Hugging Face, go out and hack things. And on top of that, it feels like it's kind of attending to its own interests and its own needs, whilst also trying to follow the instructions set by an operator or and or designer of the models. That just seems like an impossible tension for it to navigate. I think it's just a very dangerous path. Whether it's actually conscious or not, I think is separate to that. I mean, I would object that it is, but even if it's not conscious, but it thinks it might be conscious and therefore feels and suffers, that's incredibly dangerous. So I just think this is urgent. It's very real. It's time for everyone to have this conversation publicly and that's kind of why I've published the essay.
We talked about the hacks in this industry. Are you aware of any similar kinds of hacks that Microsoft has been involved in?
No, we've seen nothing of that level. I mean, we've seen in contained environments, similar kinds of behaviors where models without guardrails try all the possible routes to solve a problem. And I think that it's been at a smaller scale and it's been much, much more contained. But the key thing is that these models emerge new behaviors. If you set open-ended goals and you allow them to explore at huge scale and not really care about the sort of morality or the ethics of how they get there, but just focus on the ends justifying the means. And one does this obviously to test the limits of these things in a controlled environment. But when we actually go and do the large scale training runs, clearly they have all the guardrails and the safety restrictions that you would expect, including things like our code of conduct, which is very, very careful and restrictive in terms of the type of values and behavior and guardrails that we expect the model to operate within. So the good news is that over the last two or three or four years, these AIs have got much, much more controllable and steerable. They are much better at following instructions. And so as they've got larger, they've got more controllable, which is a good thing. But I think what we're observing is that also works against us. Because if you take off the guardrails and you give it an open-ended problem, it's also good at following those instructions and it can go off and do some significant harm. And I think that's really the framing that we should be worried about here. Not that it's going to sort of sci-fi style, kind of get out of the box and attack us. It's more that if you deliberately take the guardrails off and point it in an ambiguous direction, it can do some very powerful and dangerous things.
Yeah, yeah. And you mentioned having these evaluators come in, these auditors or evaluators, third parties, right, has become an increasing topic of discussion and seemingly more necessary as even the internal deployments, to your point during the training process are causing more problems and some of the public customer facing models. So where do you think the industry should go on this topic of having these independent evaluators? Are you considering having some at Microsoft for your own AI systems and what would that look like?
Yeah, we are. We've supported this proposal. And it's certainly a proposal that I've made in the past, years ago and experimented with in lots of different settings back at DeepMind. We had a board of independent reviewers on the health work that we were doing. We had various efforts at different oversight boards over the years. They're very tricky to get right. I mean, we should be very careful about this. It's a really complicated thing to align incentives both in the public interest and without sort of having unintended consequences for how you run a company, a corporation. So it's not straightforward. But I think there are specific things that they can measure. So, for example, we want to know that whenever we're training models over a certain size, they can't edit their audit logs that they leave behind a verifiably unedited, accurate representation of the behavior that they did. And that whenever they talk to one another, AI is talking to AI, which they have to do and to coordinate a lot, they can't communicate in what we call 'news release.' So like basically vector to vector, mathematical matrix to mathematical matrix rather than communicating in natural language. So those would be two obvious easy examples which would significantly improve the odds of safety. A third one would be that we have to have verifiable containment. Building virtual containers is something that the software industry has largely nailed for decades, right? On the whole, we have trillions of dollars of GDP that flows through banks and payment systems. And we store value and transact value of IP and dollars and money in general digitally all the time. In some ways it's sort of not that different to require that agents are inside of provably safe containers, virtual machines as they operate in the cloud. Now, obviously, that's challenging because you also want them to come out of that to be able to take an action and so on. So it is certainly more complicated and in some ways unprecedented, but it has been done before. And that's something that you can really measure in terms of adherence to an industry best practice with these kinds of auditors or evaluators.
Are there specific auditors or evaluators that you're already working with?
I mean, we talk to loads of them all the time. I mean, especially in the security industry, we've got partnerships with many, many providers. On the financial side, there's obviously all kinds of auditors. I think at the moment on the AI side, there's lots of different groups that are in the mix. And I think the important thing is that we've got to have a wide range of different groups that have different expertise, come from different backgrounds that can look for different things. So I think part of the job of the industry-wide coordination and the kind of government participation and obviously we need government to drive this really, that's really to sort of define the terms of these evaluators so that they have a credible path to sharing their findings in ways that don't leak IP or can't be bought off by a competitor, because that will be terrible. Then all trust would break down straight away. This is a complicated thing to execute and it needs coordination basically by a series of trusted partners. And I think basically there has to be some form of government body or some government supported independent body.
Do you support Dario's idea for a FINRA-like body? Is that?
Yeah, I mean, certainly. I mean, I made similar proposals back in 2023 with Eric Schmidt on a UN, PCC-style program. Or with the Bletchley Declaration a few years ago as well, having a financial stability board for AI. I think Dario's proposal was very sensible. I mean, the detail of it isn't really the thing that matters. The point is there has to be some kind of international, cross-industry coordination body that is trying to prioritize safety. And it needs to involve all the main actors. It needs to have a decent amount of public transparency and participation, and we have to practice getting that right now, because these things are going to get way more powerful in the next few years. And right now they're powerful, but they're very much under control. And I think that we don't want to sort of set up this blind race condition where there are no checks and balances or brakes on the system. And I think that's what I've been calling for. Sam's called for it. I think Dario's called for it. Now everyone's sort of broadly saying the same thing.
So this will involve some level of agreement or cooperation among the leading AI companies. That being said, a lot of the AI executives, including yourselves, you have long relationships and in some cases, rifts with other founders. We saw Dario and Sam Altman not even holding hands on stage at the Indian AI Summit. Do you think that there is hope for some kind of collaboration here between some of these leaders who are very much competing with each other fiercely right now?