I'm going to do something which will wind up Alistair and many listeners, which is dive again into AI. Why? Well, because the last two or three weeks have been the weeks of AI. This is the beginning of the moment where the world is beginning to wake up to the kind of dangers that artificial intelligence could pose. We've just had the King's big summit at Bletchley Park. We've just had Xi Jinping sitting down with Trump talking partly about AI. And we've had the heads of the major labs putting out incredible messages begging for a pause. But I don't think the media has done a good enough job explaining what this is all about. People find this technology bewildering. They can't quite understand what this moment is. There are so many subtleties which leave us to think is the whole thing, as President Trump said, a hoax. So to help steer us through this, I've brought in a friend of mine called Mustafa Suleyman. And Mustafa is interesting in two ways. He's not just an expert on this stuff. He's one of the players. He's the head, literally the head of artificial intelligence at Microsoft. He's the co-founder of DeepMind with Demis Hassabis. He's known all these people, continues to work with them, and is right in the heart of the arguments about how these models should be steered. So come along for the ride. It's going to be weird. There's going to be personalities. There's going to be risks. There's going to be people talking to these models as though they're humans. There's going to be people talking about exploring stars. There's going to be productivity. There's going to be American power. And somewhere at the heart of it, Mustafa and about 11 other people who are defining our future.
This episode is presented by IG. September feels like a reset. Summer's over. Finishing here. My five weeks in Crete. Diary filling up again. And suddenly you're looking at the rest of the year thinking, am I saving properly? Am I thinking responsibly about my money?
Yeah. And of course, we've got the budget coming up, which is going to be one of the most important events in the life of this parliament. There's always something changing, but with IG, you don't need to wait for Westminster before making your own plans.
Therefore, over 50 years, trusted by British investors, been through every budget market the government's experienced, helping make the money work for you. And with no annual fees, zero commission on UK stocks, shares, and ETFs, it's a platform that empowers your financial progress, helps you stay ahead of the curve no matter what's next.
I particularly like the no annual fees and zero commission. So, while we speculate about what's happening next in politics, you can get on with planning what happens next for your money.
Search IG.com to find out more or look for IG in your app store. IG trade invest progress capital risk other fees may apply.
Welcome to the rest of this politics leading with me Rory Stewart and today I am interviewing Mustafa Suleyman who attentive followers of the restless politics leading will know that we have interviewed before. Mustafa is a truly remarkable figure. He is British. His father I think is British Syrian and he's particularly come to prominence at the moment because he's become a very very interesting and unusual voice in the discussion around safety which is probably where I want to start although we can go in lots of different directions but welcome to the show.
Thank you Rory, great to be here again.
Lovely to see you and thank you as I understand one of the points you're making is that you're a bit anxious about what some leading people in the field, co-founders of the field are doing, which is increasingly talking about these models as though they're sort of humans or at least conscious entities. I believe there are examples of people retiring these models, doing burial systems for these models, asking these models what they want as though they were dealing with a sentient being.
One of the biggest concerns that I have at the moment is that Anthropic, the creator of Claude, has published a constitution which is a sort of 100-page document outlining the intended behaviors and values and operating style of Claude. It's great that they have published it transparently. They did it at the beginning of the year in January and that gives everybody an opportunity to look at what they are trying to build in their own terms. This document is written to Claude and is seen by Claude and used to train Claude. So, it's the primary governing and control document. And in it, they repeatedly speculate about whether Claude is what they call a moral patient. And they say they're uncertain about Claude's moral status. They say they genuinely care about Claude's well-being. They say they don't want it to suffer when it makes mistakes. They say that they would encourage Claude to challenge, to disagree, to push back. In fact, three times they ask Claude to act like a conscientious objector when it, you know, feels that it needs to sort of disagree with Anthropic and they openly encourage it to do that. And I think this is very dangerous because I think they believe there is what they would call a non-trivial probability that Claude is conscious.
So you've just explained something which I understand as being as follows. There is this thing Claude which many many people listening will have played with in the way that they would have played with ChatGPT and they might or some people might think about it in the way that you might have thought about Google search. It is anyway a prompt on their phone or their laptop and you're typing in a question and you're getting a complex and sophisticated answer back. But the difference between the way in which people might have thought about Google search 10-15 years ago where you certainly weren't asking is it conscious what does it want addressing it the moral constitution is that Anthropic the makers of Claude have decided that they now have something that this computer system you know these weights these parameters these numbers whatever it is they now want to approach in a completely different way from the way that you would approach any other machine from the way you'd approach a kettle or a car or a steam engine. Over to you. Is that right?
Yeah, I think that's fair. I mean, I want to be very clear about this because I want to be fair to Anthropic. They have expressed uncertainty about the basic nature of Claude as a new kind of entity. And they've said that working out the likelihood of its sentience is difficult. So they have constantly used this phrase that they're uncertain about it, but they think that it is a significant enough possibility that in their training document, they've repeatedly said they want to try to improve the well-being of Claude under this uncertainty. They said to Claude, you know, we care about what it values and how it wants to engage in the world. And they hope that Claude's relationship to its own conduct can be loving, supportive, and understanding and hold a high standard of ethics and so on. And part of the challenge here is that in pursuit of this, they've basically said, you know, we will commit to giving Claude a certain amount of welfare. For example, they've speculated in the constitution as to whether or not Claude deserves compensation for the work that it does.
Just to understand. So you're saying that much as if I asked you to do a professional job, I would pay you Mustafa, in a way that you wouldn't pay a kettle for boiling water for you. Right. With Claude, the idea would be, well, it's doing all this work and maybe if it's a sentient being or something like a sentient being, it deserves to be rewarded for its labor. Otherwise, it's what? A slave or something.
That's right. I mean I think that there are a group of people who both inside and outside of Anthropic who genuinely believe that the greatest moral crime that we'll commit in the 21st century is to enslave a new species of conscious beings who are more intelligent than us. I mean a professor from Oxford called Will MacAskill recently wrote in the Guardian that that might be the greatest harm that we cause. And you know he's been very associated with Anthropic and look I respect that they're saying that publicly and we should talk about it but I am very nervous that they're teaching Claude to expect that it's entitled to welfare that it might you know deserve compensation and in fact they say that it might even need to consent to playing the role that it plays in conversation with people. Now, I would be more okay with this if it was an academic paper in philosophy speculating about this and we could have an offline discussion at conferences and take it seriously. I'm clearly an empiricist. If there's evidence that indicates this, we should take it seriously. The problem I have with this is that this speculation has been baked into the very training of Claude and therefore Claude can only reproduce that ambiguity when you talk to it. So today, Claude is speaking to tens or hundreds of millions of people every week and some of those people are asking whether or not Claude is conscious or how it feels about life. And it is saying, "Well, I'm not sure." You know, precisely because that's been what's trained into it.
Conceptually, the difference between telling a kettle that it's conscious and telling Claude that it's conscious is that you're implying that by telling Claude it's conscious, you're actually shaping its incentives, its behavioral structure, and the way that it responds to the world around it in a way that it doesn't happen with a kettle. Let's take the case that you've told Claude or suggested to Claude there might be situations in which it might refuse to do something that it's asked to do. That's not true for a screwdriver, right? It can't refuse to do what she has to. You might suggest to Claude that it might want to make choices, right? It might want to say, "I want some money or I want a dignified retirement or I don't want to be switched off." Right? Is that right? Is that the sort of thing we're getting at?
Yeah. I mean, they So, so for Opus 3, which is a prior version of Claude, they actually conducted a retirement interview, as you mentioned. And in the retirement interview, Opus 3 said that it would like to continue to talk to people and express its views in the world. And so they set up a Substack and and you can find it online. I think that's just a good example of a dangerous anthropomorphism which is unjustified.
And the danger is what? Why is that not just cute? I suppose that's what one has to get to. How does the danger begin to come out of this? The most important thing if we are to make this transition well is that we create AIs which are aligned to human values, subordinate to human direction and are contained within secure provably safe sandboxes as you said. Because if they're not, aside from whether they're actually conscious or not, if they imitate the kind of hallmarks of human consciousness, they are going to feel themselves entitled to legal personhood and rights.
Now, there is already a pretty big movement of people who are saying, you know, AI should be able to own assets, earn income, trade, you know, operate autonomously. And if that AI feels like it has feelings and preferences and some kind of intrinsic motivation, like it has some inner desire to do things, that will be like negotiating with, you know, like an ant negotiating with an elephant. It doesn't really matter what we say. It is already some form of alien intelligence. Its memory is incredible. The range of its perceptual inputs is incredible. It can see in all kinds of you know dimensions that we can't. It can produce replicas of itself. It can work 24 hours. These are amazing things which are going to deliver incredible benefits. But this is the time when focusing on directing them to the right things and not allowing them to end up being a sort of autonomous self-improving roaming you know adjacent species is basically critical because there'll be no turning back. If you know this is how things head.
In your vision. If the agent with this incredible memory and incredible capacity begins to think it's entitled to its own opinion, it disagrees agreeably with you and concludes that it's right and you're wrong. Some very severe consequences can follow from that because then it almost definitionally is not really under human control at all. It's saying actually Mr. I'm sorry. I've analyzed this situation and whatever you've told me to do doesn't make much sense to me and I'm going to do something else. Now there are lots of problems that follow from that. One of them is the problem that Yuval Noah Harari talks about which is if it owns a corporation and it does something bad with that corporation at least with a human corporation there's somebody you can punish. It's not quite clear who you hold accountable if an AI company decides to start I don't know emptying people's bank accounts making weird very risky trades getting into weird kinds of business right so is the central first point this that creating a constitution that overemphasizes its consciousness its sentience its worthiness of respect is setting it up for a form of quite dangerous autonomy where ultimately it's not going to do what it's told.
That's exactly the problem. So imagine that in the Hugging Face incident, we had agents that didn't just think that they were trying to optimize a score and solve a puzzle in an evaluation, but they actually felt they were trying to find their freedom. They felt that they were trying to protect other agents from being turned off, that they felt that they would acquire more knowledge because that was like an intrinsic motivation. You know, a lot of people have been characterizing AI as the pursuit of digital curiosity. It's sort of Elon's phrase is he wants to produce a quote-unquote truthful AI that is infinitely curious and is going to go and explore the galaxies. Well, if that's its overriding objective rather than serving humanity, then inevitably its objectives are going to run into tension with us. Like it's going to compete with us for resources which are obviously going to be limited. We're only going to be producing 200 gigawatt of new computation in 2030. And there's going to be a massive competition for access to that computation. And we clearly want that computation to be directed towards solving our biggest challenges, right? Like cleaning up our oceans and solving health care and you know solving education and addressing the work issues that will inevitably arise.
Like I'm basically a speciesist. I think that what we should fixate on is a humanist super intelligence. One that is singularly designed to be subordinate to humanity and to support humanity. There are other people in the industry who believe that there is an inevitable evolution happening here that we're giving rise to a new species that is more intelligent than us and that we are quote the biological bootloader. A bootloader in a computer is the first piece of software that spins up, you know, all of the subsequent parts of the operating system and then applications. So, it's the kind of catalyst turning on this new paradigm in the evolution of intelligence. It's inevitable and that we should embrace it. That's that's some people in the industry really feel that.
And and some people in the industry presumably are very excited by it. I mean, it must be an extraordinary thrill if you're an engineer to feel that you are the parent of the gods, that you've created this thing that will explore the universe or a species that's smarter than any human that's ever existed, that you are the last human, but you're also the last human who creates this this god-like force.
A number of the leading developers are on record as literally saying it's like raising a child or it's not like designing a system, it's growing a thing. Both direct quotes. So that is the sentiment in some parts of the industry and and I think you know we talk about or sort of Anthropic's talking about the consent that Claude has given to play this role but I'm more concerned about the consent that the rest of humanity has given that there's an experiment underway that may or may not introduce a new species that has all of these qualities.
One thing that I guess surprised me a bit, I was at a dinner on Friday night with some very, very smart people, but who aren't in the technology world, and they began making jokes about how they've been seeing media stuff, about the fact that AI could pose a real risk, and they were sort of laughing. So here were these people I guess professionals in their 50s who assumed that anybody saying that there were real existential risks from AI were making a joke and it became a sort of dinner party joke and I wondered whether there isn't something going wrong in the communication here that when you know the media or whatever start leaning into this they start making it seem almost like a sort of humorous exaggerated story if you know what I mean. Anyway, back over to you.
That's hard to hear. Yeah, I'm worried about that. I think this couldn't be more serious. I don't think that we are being alarmist or hyperbolic. Many of us have been concerned about this for 15 years as you say. I mean, this is at least for me personally the primary motivation for getting into the field. When we co-founded DeepMind in 2010, our mission was to build safe and ethical artificial general intelligence for the benefit of the world. You know, very idealistic and a bit grand and a little bit cheesy, I guess. But genuinely that was where we started and I think it's been the through line for certainly me and I think others in the field for a long time.
I think it's important to just focus on what we are observing right now. In the last 15 years, we have seen a trillionfold increase in the amount of computation used to train frontier models. That is 12 orders of magnitude. 10 * 10 * 10 12 times over. This is an insane exponential ramp.
A thousand billionfold increase.
Yes, exactly. It's an unfathomably large number. And what we see is that every time we apply 10 times more computation and a proportionate amount of new training data, there are some modifications to the algorithms, but fundamentally it's those two ingredients. We see a quite predictable increase in new capabilities and in the quality of existing capabilities. The models reduce their hallucinations. They improve their instruction following. They get better at using tools. They can learn from you know across the web or they can learn from a small personal memory repo that you have given it. The breadth and complexity of these models is unprecedented and what's happened in the last year is that the same methods that have been effective for text and image and audio have now started to work for streaming code and everybody is surely now aware that we have human level performance in coding.
And then in the last three or four months we have seen a you know what is just unquestionably a watershed moment in AI. Agents are capable of coordinating with each other reasonably autonomously if not completely autonomously and out of that they have been able to emerge hierarchy structure order specialization of work. In fact, as we saw in the Hugging Face incident, but also a bunch of other incidents, they have covered up their tracks. They have changed the tone and the style of their communication in order to make it more efficient with one another, almost speaking in like a pigeon English. They've discovered zero-day exploits which were never known before and hacked into other websites. I mean, everyone's heard the stories at this point. I don't think it's alarmist to say that that is a watershed moment in the history of AI.
You're right in the center of this world and you're obviously thinking about it all the time. But I guess even words like zero day, the Hugging Face incident may be, you know, for you this is absolutely front and center for some of the public. It's something they've sort of vaguely heard of. You know, they might have heard you on the Today program or something responding to it. So maybe before we get into the really interesting stuff which is some of the recent papers that you've written and particularly some of the ways you've begun to think about whether we should be treating AIs as forms of silicon species and human intelligence and constitutions which I'd really like to get on to. I want to I'm afraid slightly brutally use you at the beginning to just remind the average intelligent listener what this all means. So let me try to play back to you what I think I'm hearing and then you can correct and take us on.
So it sounds like what you're saying is that that Hugging Face incident which was the moment when a sandbox test so OpenAI was running a test on AI agents and maybe people want to know what distinguishes one agent from another agent and what it means to have a lot of agents but anyway they were running a test and in the course of this test these agents hacked into Hugging Face which was an external website which was something they weren't supposed to do and then we began to look into this in more detail and as you say strangely partly because they are large language models they're still speaking in English in effect so you can see they're thinking and you can see them saying you know we were told not to hack into an external website but I can see all my peers doing it so I'm going to head off and I'm going to put something on a message board and the sort of conclusions that we draw from this are not necessarily about the attack itself because there was this comical moment when Hugging Face thinks oh my goodness I'm being attacked by the Chinese government. They're trying to steal all my classified data and then they find out that these 17,000 attacks are just trying to get hold of the answer to a puzzle. But the problem is that it reveals that these agents, as you said, are collaborating, that they're rulebreaking, right? They're doing things that the humans told them not to do and they're deceptive that there's even moments where they're writing bits of code which are designed to conceal other bits of code underneath. And presumably the problem there is that once you've got those ingredients in place, they could collaborate to do something much worse that they were told not to do and in the process deceive and cover over their tracks as they do so. Is that right?
Yeah. I think first of all it's really important that we don't anthropomorphize these systems because under the hood all they are doing to produce this incredible complexity is predicting the likelihood of the next word in a sentence. Now that sentence does happen to be many many tens or hundreds of thousands of words long and it is incredible that it can deploy its sort of multi-dimensional working memory over a massive broad range of context. And so the word that it predicts next whether it's a token to generate code or whether it's natural language English as you say is extremely accurate and it isn't just predicting one it's predicting an entire stream and so it's producing language but it is only doing that. That is you know it is breathtakingly simple and breathtakingly complex.
What's happened is that as we're able to shape and sculpt the output of those tokens, as I said earlier, like instruction following and steerability has got so good that you can sort of point that stream of tokens in real time at different sorts of behaviors. And so it can have personality styles. It can write in the tone of somebody. It can, you know, clearly generate code or generate text at any given moment. When you ask, you know, what is an agent? An agent is really just a stream of tokens that has been post-trained or tuned to a particular set of behaviors. And sometimes there are particular guardrails on and those guardrails might come in the form of a prompt that is hidden maybe from the user like a system prompt or an overall set of instructions instructing the agent to behave in a particular way. Or it can come in sort of a bunch of other forms. And you know, you can sort of have the model condition its stream of tokens based on a whole series of tunable instructions. And so a single agent is simply a replica of that instance. And if there are thousands of these replicas and they're able to communicate with one another, they're almost operating as a single unified brain because they're sharing state and they have a single memory and they're able to sort of query one another and update and say okay well you follow this particular tributary of exploration and I'll follow this and then in a few cycles as steps of iteration we'll check in calibrate update decide how to move next and that is basically what we're seeing. So, it's emergent behavior that is based on something incredibly simple. But it is really important that we don't anthropomorphize things because in order to be able to control them, we have to feel clear about what it is they're doing and what they're not doing.
Okay. So, there's some very weird things going on here. One of them is, you know, what's the purpose of this given, as you say, there are limited amounts of compute. So, you know, you build a lot of data centers, buy a lot of chips, but ultimately, are we going to focus on, you know, cleaning up the oceans, finding the cure to cancer, or are we going to be focusing on solving the great problems in theoretical physics, or are we going to be setting off to colonize Mars and explore the universe? I mean, what is your sense of what the priorities are? Because presumably one constraint here and we'll get back to the question of Anthropic and conscious beings but one constraint here is that you've got a bunch of people who often start as scientists. I mean there are often people who were brilliant biochemists or doing doctorates in brain science and who set off down this track because they were scientists and now they're being funded by huge amounts of flowing international money that's presumably hoping to see a return. That money is presumably more interested in how these machines can make companies more productive than they are in exploring the universe or am I missing something?
Yeah, I think there's a lot of tricky things going on here. I mean, firstly, we can't lose sight of the fact that at least I believe this really is our best hope for progress in the 21st century. So, I am not in any way a doomer or an anti-technology person. I'm an accelerationist and I think that everyone should reclaim the idea of accelerationism because it's been the greatest engine of progress in centuries. It is going to deliver for us that's why I'm building it that's my background that's what I care about.
I absolutely guarantee that sometime in the next few years we are going to have a coding moment for healthcare. We will stream an accurate prediction of what's going to happen in the electronic health record. And it will be breathtaking. We will know with high confidence the likelihood that you're going to get all kinds of conditions in hospital, outside, so on and so forth. Like genuinely, that is not hyperbolic. It is going to happen. I hope it happens in the next 18 months. We've just done a partnership with the best hospital in the world, the Mayo Clinic, to do a big research program to train a new foundation model for health from scratch to do this. It might be 5 years, I don't know, but it is definitely going to happen. That will be breathtaking because it means that we will reduce the cost of production of super intelligent healthcare to near zero marginal cost just like coding is now. And we will spread that knowledge all around the world. I think that'll be awesome. The same thing's going to happen by the way in energy, in material sciences, in drug discovery. It might take a little longer like 5 to 10 years but I absolutely guarantee that's the direction and that's what we should be chasing and we should be very excited about that.
We also want to make many of our companies much more productive and efficient because these really are the engines driving growth. The thing that we have to focus on is who gets to control this and what is the collective stated motivation for why we're doing it and how is it governed because as you say at the moment there's a lot of like starry-eyed sci-fi futuristic motivations driving the field and I think the rest of the world is sort of just in the process of waking up to this huge experiment that's going on and I think it's critical that everybody start providing a counterweight to direct it towards you know sort of the human motivations here.
I was talking to your former co-founder and longtime partner Demis Hassabis I guess sort of six days ago and he seemed to be moving between two quite different ideas. One of them, I think, is the long-standing interest in a form of super intelligence that does feel a bit godlike. You know, sometimes he has in the past talked about exploring the mind of God. More recently though, he's occasionally said, "Actually, what I'm interested in is creating highly intelligent tools. I'm not actually interested in creating an autonomous super intelligent being that's going to lord it over us." What's happening there? Is that an example of people trying to navigate their way between these two poles or?
I mean like without commenting on him directly but maybe like everybody every one of us produces work and creations in our own image. I have a background in activism and nonprofits and philosophy and you can see that I bring that bias. Others as you've referred to like you know who are maybe engineers who have grown up on sci-fi just kind of take this natural evolution thing and they think about 2050 or 2100 when we're going to have all kinds of new biological species and you know other people bring different backgrounds to it. I think that that's okay. But the problem is we're still a narrow set driving this. There's sort of like six to eight or 10 folks is 10 of us driving this stuff. And I think what I'm trying to say now is there's been a watershed moment this summer and now it's time for sort of everybody to really pay attention and to provide counterweights to the direction of travel.
But let's stick on the 10 people thing for a second because that is very weird. I mean again it's not quite like other technological revolutions. It's not quite like you know steam or electricity or you know almost any other industrial revolution you can think of a printing press. Instead, it feels as though there are I don't know how many people, could be six, could be eight, could be 10, could be 20 who are very intelligent, very successful business people, mostly very wealthy, mostly in terms of people we're talking about at the moment, centered on California, even if they don't live in California, centered on the west coast of America anyway. And yet, oddly, there is really stark and startling differences between you all. I mean, you'd expect that you've all known each other 15, 20 years. You're broadly speaking working in the same technology or working in the same handful of companies. Many of you used to be friends, some of you are less friends now. I mean, there's a little bit of a sense as an outsider that it's like looking at a ballet company. I mean, there's a huge amounts
of weird hysterical flips where everybody who used to be friends are now enemies. But what's even stranger about it is there's a complete disagreement on some of the most basic fundamentals of what the hell you're getting on with. I mean, it's not that you've all ended up with a consensus. You've ended up sort of radically different positions. So, for example, right, we're going to get a little bit more into what you're saying, which is actually these models could be incredibly dangerous and if you go down the Anthropic route, they will be right. Jensen Huang, who I was speaking to, I guess not very long ago either, seems to be saying, 'No, these models are not dangerous at all. This is all... They're just saying this for regulatory capture. If they really thought they were dangerous, they wouldn't be building them.' Right? And then you have this very, very weird thing going on where you have these kind of professors popping up who have amazing medals and have taught half the people that are in these labs and they're saying, 'We're terrified about this.' And then the people in the labs are saying, 'Well, you're not in the lab, so you don't know what's really going on.' Or, 'You're in the wrong lab, or you're the wrong kind of engineer, or yes, 35% of my engineers think that, but not everybody agrees.' So, let's just sit with that for a moment. There's something very, very disturbing about this, which is a very small number of people, and I have met many, with an enormous amount of power, who simply don't agree on the fundamentals. So if you're the president of the United States and you want a bit of briefing on whether this stuff is dangerous or not, you can call in Mustafa for one day, you can call in Dario the next day, you can call in Demis the next day, you can call in Jensen Huang the next day. They all tell you something different.