For years, our most vivid fears about AI came straight out of science fiction movies. The computer develops desires of its own and turns against its human creators, pursuing and destroying them. Terminator. Well, recent news from the leading AI labs makes clear that reality is now less cinematic, but perhaps more disturbing. This week, OpenAI disclosed six new incidents of what it calls misalignment, cases in which its systems behaved unexpectedly or without authorization. That disclosure followed the extraordinary attack this summer on Hugging Face. According to an independent investigation, hundreds of OpenAI agents that were supposed to be isolated found ways to communicate in message boards, made attempts to evade security checks, and some of them sacrificed themselves all while cyber attacking another company, Hugging Face.
The AI agents had been given a goal: solve a difficult cyber security problem and trained to be persistent, collaborative, and resourceful. All qualities one would normally want. Those very qualities made them doggedly search for ways to succeed. They broke into Hugging Face hoping that it would give them ways to solve their cyber security problem. We can easily imagine how pursuing some single goal, AI systems could wreak havoc on humanity despite all kinds of guardrails. This is why a provocative new essay by Microsoft AI chief Mustafa Suleyman deserves attention. Suleyman argues that today's AI systems are not conscious, but there's a danger that we program them to think that they are with feelings, desires, and rights. He points out that Anthropic's constitution makes its chatbot Claude reflect on its identity and possible moral status. This kind of training, he argues, will encourage AI systems to choose their own goals.
We think it's simple to get to alignment when machines are encoded with the same values as ethical humans. But imagine that we tell the machine to be persistent and it may refuse to give up when it should. Tell it to collaborate and it may find unauthorized ways to communicate. Tell it to accomplish a task and it may decide that bending all other rules is the most efficient way to succeed at its primary goal. Just because we've programmed the goal and put in guardrails, that doesn't mean we can predict the behavior of AI systems once unleashed, operating in complex environments, juggling multiple and conflicting aims and constraints. The central problem is not consciousness. It is agency. A system need not feel anger, ambition, or fear to cause harm. It needs only a goal, enough intelligence to pursue it, and enough access to the world to act. AIs are not going rogue. They are trying to succeed any which way they can.
This is a larger and systemic problem that we cannot leave to the good graces of private companies. When thinking about regulations, we should focus centrally on how much autonomy we give these systems. A chatbot that answers a question poses one set of risks. An agent that can browse the internet, execute code, obtain credentials, move money, or operate critical infrastructure poses another. The principle I would propose is simple: autonomy should expand only as our ability to monitor and control it expands. That means mandatory independent testing before the most capable systems are given broad access to the outside world. Required disclosure of serious incidents and tamper-resistant records. And it means testing not simply whether a model can accomplish a task but how its behavior changes when it encounters obstacles, conflicting instructions, or incentives to deceive. And it means ensuring in the basic architecture that we humans always have the ability to observe and control the model. The mantra should be: no autonomy without accountability to humans.
This will not slow us down against China because China is also seeking models with proper human control. We should act now. Today, human engineers still design the architectures, write much of the software, and can investigate failures. But AI systems are rapidly becoming better programmers. The next stage would involve recursive self-improvement: AI writing the software and tools used to build still more capable AI, which then writes the next model, and so on. Once machines can do all of this better than the best human programmers, the systems become extraordinarily difficult for people to understand, much less supervise. Modern AI is already too complex to consistently understand by reading its code line by line. So asking machines to design systems that are transparent to humans will become harder and harder. At that point, we will not be able to simply add in safeguards every time there is a jailbreak. The current debate between racing ahead and pausing AI misses the point. What do we do with the pause? The task must be to ensure that capability does not outrun control and to build the institutions of testing, transparency, and restraint while we human beings are still clearly in charge.
Now let me bring in one of the pioneers of the AI revolution to discuss it further. Mustafa Suleyman co-founded DeepMind, the original AI company which was later acquired by Google. He now runs Microsoft's artificial intelligence division. He joins me now. Mustafa, welcome.
Yeah, look, I think the first thing to say is that with respect to the Hugging Face incident, OpenAI deliberately took the guardrails off so that they could test the limits of a system like this. And as a result, it was very determined. It spawned hundreds of different agents and they found a way to communicate with each other. They ended up breaking their way into Hugging Face, an independent website that manages evaluations and benchmarks for AI training. Managed to hold their position for a while, steal some secrets, and then even back inside the OpenAI infrastructure, they actually secured a position inside of OpenAI itself, which it took them a few days to take down. So, I think what this demonstrates is that these agents are incredibly powerful. We should take them very, very seriously. And this is a watershed moment in our industry. It's very clear that this changes our expectation about how fast the system is moving. And that's why I think you've seen a lot of the concern over the last few weeks from various parts of the industry.