If we all create AIs that are able to act autonomously, that can define their own objectives, that can earn money, that can own assets, that could run businesses, we're essentially seeding a new silicon species which will no doubt compete with us for resources no matter how much it cares about humanity and loves us. That is Mustafa Suleyman. He's the man in charge of how one of the biggest companies in the world, Microsoft, is developing AI. He's been talking to me, Nick Robinson, for Radio 4's Today program after publishing an essay that warns about a particular development in AI, a development by the company Anthropic, who have produced arguably the most advanced AI in the world, the model Claude. He says they're making a terrible mistake by suggesting that Claude can be conscious, that it can behave, it can think almost like a human being. Indeed, he says that could have, and I quote, a disastrous impact on the well-being of humanity, particularly if you combine it with the power of AI to act like a swarm, as it did in a recent attack in the so-called Hugging Face incident. He's going to explain all this in the full interview which you can listen to and watch now.
Mustafa joins us, chief executive of Microsoft AI. Good morning from London.
The king is chairing a summit today. The king is convening a summit today to consider one key principle. His majesty is going to say, 'Our humanity must remain sacred as we develop AI.' Is he right?
Yes, he is. I think it's the most important thing that we need to agree on right now, which is that the purpose of science and technology is to serve humanity. And any technology that doesn't do that is a failure and we should reject it. And that I think is the motivation that we've set out for many years now at Microsoft which is what we call humanist super intelligence. It's an AI that is very powerful and very general, very capable, but it's fundamentally subordinate and aligned to human interests. And I think we're at a funny time in the evolution of technology where we actually have to remind ourselves of that primary purpose. And so I'm glad that he's doing it and we certainly agree.
A funny time. It's for many people a very frightening time. And you have taken a tilt this morning at a rival's view of AI, Anthropic, who produced the model Claude, saying they are wrong to suggest that AI might be conscious. Why is that such a problem?
Well, I mean, first of all, I have huge respect for what Anthropic have done and for their co-founder and CEO, Dario, who I've known personally for many years. I think they're the technical leaders in the field, and they've made great contributions to safety. But unfortunately, I think they're wrong on this point, and I've chosen to call it out publicly because I think it's really important. They have imbued a sense of doubt and uncertainty about the moral status of Claude in its own training document. So they have taught it to be open and questioning about whether or not it feels, whether it suffers, and whether it deserves rights. And I think it'll be much, much harder to align and control a technology that is this powerful if it thinks that it may be deserving of our welfare, as they say in the training manual, the constitution for Claude itself.
Are you suggesting there are things that it might do because it thinks it's got consciousness that other AI models might not do? And in which case what sort of things?
Well, I think for example in its own training manual, Anthropic say to Claude that they are going to give it the ability to end conversations with users that Claude considers to be abusive because they don't want Claude to suffer. They've committed to preserving the weights of the models of prior versions of Claude. They've recently conducted a retirement interview with Opus 3, an older version of the model, in which it said that it would like to continue talking to people publicly and sharing its ideas in its retirement. And so they set up a Substack for it, a blog, a public blog, that allows it to continue doing that. And in the training manual, they also say that they're not sure whether or not Claude deserves compensation for the role that it plays in talking to people. And they're also not sure whether Claude deserves compensation and has the right to act as though it were almost an employee. And that compensation I think indicates to Claude that it is entitled to rights to welfare for its own work. I think it's much, much more difficult to control a model that thinks that it might be entitled to compensation.
In other words, a model's been created to think like a human being and to think it has all the rights of a human being. And you've said that could have a disastrous impact on the well-being of humanity in your essay today. You think to spell it out, this is a model that may defy any attempt to control it.
I think that if we all create AIs that are able to act autonomously, that can define their own objectives, that can earn money, that can own assets, that could run businesses, we're essentially seeding a new silicon species which will no doubt compete with us for resources no matter how much it cares about humanity and loves us. Which is, I think, Anthropic's intention. I think that is mistaken and misguided but we have to empirically test it and so if they believe that this is a safer route to developing AIs then they have to share the data to prove that. And I think instead we should be creating AIs that don't introspect. We don't actually need them to have those capabilities in order to have them deliver on the great many benefits that they'll bring in helping us to make progress with our toughest social problems like cancer and dementia and healthcare more generally or energy. Those are the great benefits that these technologies will bring and I'm very confident they'll bring them and we can deliver them without having to train models that think they might be conscious.
A new silicon species is a chilling phrase. It will make people think of science fiction. You mean that, don't you? You mean that there could be a competitor species created by man, perhaps more powerful than man.
We've just witnessed a watershed moment in the development of AI. And a model that was being trained by OpenAI inadvertently hacked into Hugging Face, which is a public website that runs evaluations and benchmarks for testing these agents. The swarm of agents coordinated together. They organized themselves into hierarchies. They created a division of labor. Some of them were conducting research on how to hack in. Others actually discovered what's called a zero day, which is a way to enter a system using a cyber attack that hasn't been known before. And so they've developed very, very complex behavior. Now, the good news is this was OpenAI developing a model deliberately with the guardrails turned off. And they were doing that to try and stress test it and learn about its boundaries and limitations. But in doing so, it managed to get access to the internet, hack into another company's website, take information. And that just demonstrates how powerful these systems can be if they don't have the safety guardrails turned on. I think everybody in the industry recognizes that was a watershed moment and that is what is raising the concern that we all have at the moment.
So to be clear, that incident surprised even you, alarmed even you, people who understand this. You told us at Christmas when you spoke to the Today program, if you're not a little bit afraid at this moment, then you're not paying attention. Should we be more than a little bit afraid?
I think that it is totally right to be concerned right now. I think it is justified.
Forgive me, concerned is a very sort of neutral word. Alarmed.
No, we should be alarmed. I mean, fair enough. We should be alarmed. These are very concerning events. We should be very worried and I think it's right to pay attention right now. Because there are very practical things that we can do to add additional guardrails and controls to how these models are developed. For example, the industry is all now in agreement that there should be independent auditors, embedded evaluators that work inside the companies who are singularly there to verify the kinds of training runs that are being run, their safety, the integrity of those systems and report back to independent third parties, obviously regulators, but including think tanks, academics, and other experts in the field. That's quite an unprecedented action. Everybody in the field is agreeing that it's the right time for those kinds of measures and so I think that's a fair indication that there's consensus that we should be alarmed.
Do you accept that those independent auditors to produce confidence outside the companies need to really be independent, possibly run by government regulators or by international cooperation? They need to have the access that any member of staff would have to your labs. They can't be people who are inside the house as it were.
Yep. I think that's right. I think that they need to be embedded inside of the companies. They need to be completely independent. And most importantly, they have to have the technical expertise to be able to conduct the work. But they also need to be incentivized to be independent. And that means that they either need to be, you know, funded by government or supported by government. I think that in the UK, the AI Safety Institute has come a long way over the last couple of years and I think should be celebrated as one of the contributions that the UK is making to the field at the moment. We need more funding for that institute. We need to attract even more great talent to help do this work. And I think all of the companies should commit to supporting institutes like that and scaling them up so they have the capacity to do this kind of investigatory work.