Dario Amodei, you're the CEO of Anthropic. Thank you for doing this today.
Thank you. Thank you for having me.
Okay, so we're here to talk about the concern that AI really could kill us all in the not-so-distant future. What's the risk of this actually happening?
So just to back up a little bit, you know, we've believed since we started Anthropic, I've believed since, you know, the 12 years that I've been working in this field, that this is a powerful technology and therefore it's a technology with serious risks and serious benefits. I've always believed that the benefits outweigh the risks and even more to the point that we can reduce the risks to be very low if we build the technology in the right way. Looking at some of the benefits, I really believe that this technology could cure diseases that have plagued humanity for millennia. Just last week, we did some work on, you know, doing better protein binding, which is one of the early steps in drug development using our models. We're starting to work on them. We'll have things to say in the coming months to make basic scientific discoveries in biology. But when something's very powerful, right, what we're building is essentially intelligence. When something's very powerful, then of course it can be used well and it can also be used for ill. And so I am worried about what we've seen with the OpenAI Hugging Face incident with models, you know, taking actions they weren't authorized to take and I'm worried about that escalating. I'm worried about misuse of AI models. Just two days ago, we put out a report about misuse of models to build biological weapons. And our view, the whole way we've built Anthropic is so that these bad outcomes don't happen. Instead of talking about the probabilities, let's talk about what we can do. It's not like we're just rolling the dice, right? If we build in the right way, I think the probability of something bad happening is very low. If we build in the wrong way, the probability of something bad happening is very high.
The conversation right now is that we are on the precipice of something really scary. Is this a five-alarm fire?
I, you know, I don't, you know, depends what you mean by five-alarm fire. What I would say is...
How would you describe it?
What I would describe is, so something I've been saying since, you know, since we started Anthropic, almost since before we started Anthropic, is that this technology is on an exponential, right? It's like you have one, then two, then four, then eight, then 16, then 32. But even as I was saying that to myself, I don't think I fully just appreciated what it would actually be like when the progress was as fast as it was. So we're, an exponential curve looks a little like this...
Where are we on the curve?
We're kind of on the bend where it's starting to get steeper.
So, is that a five-alarm fire to you?
What it means, it doesn't mean we need to panic today. It doesn't mean we need to shut it all down. What I would say is it's a warning sign. It's a warning sign that we need to slow down. We need to make this technology carefully and we need to make sure that our safeguards, our ability to understand it, our ability to control it keeps up with the pace at which the technology is happening. So far we've been trying to do that by building the safeguards fast, but I think we've reached the point where at least to some extent we need to moderate a little bit the pace at which the technology is being built.
Does that then mean that Anthropic is going to stop releasing more advanced models?
It doesn't mean that. What it means is that we need to make sure that every generation of models that we release is properly tested. The thing I've proposed, so, you know, in this essay that I put out a few hours ago, I proposed a three-step plan, right? And the first step is that we put in place embedded evaluators, external evaluators from a third-party nonprofit or perhaps someday government organization that can observe the process of us training and running our models. That actually is, I think, ultimately more important than how or when we release models because it's the predicate for more carefully releasing models. I want this to be done across the industry so that whenever anyone builds an AI model, there's a third-party evaluator who understands the best safety practices, who can verify whether that company is following the safety practices that they're committed to. It's like a food inspector, you know...
But can the food inspector keep up when the technology is changing so fast?
Well, one of the things they'll be able to do is tell us the rate at which they can keep up. That may be one of the factors that may determine, that may set the speed limit on the technology. We have to understand it and the external evaluators have to understand it.
So Elon Musk has said he agrees with you. Sam Altman says he agrees with you. What's the action plan for you guys to come together? Is there a group chat that's going to come into fruition here?
So, first of all, I'm enormously appreciative that other industry leaders including our competitors have offered their agreement with this plan. I think the next step is for multiple industry players to also agree to have third-party evaluators. Then I think the next step after that is in some form the companies need to work together to set standards about safety standards for release, perhaps something about the pace of release. But in order to do that there needs to be involvement of the government because what we're talking about is the rate of release of products because of whether they're safe. We need to have that conversation with the government in the room.
Do you think the government fundamentally understands what's possible here? Because, you know, former Anthropic employee Jacob Coxon said about both Anthropic and OpenAI, neither company is acting responsibly. You guys are gambling with our lives. What's your response to that?
Well, first of all, I mean, you know, one thing that Jacob said, it wasn't quoted as much, but, you know, he also said that he thought, you know, Anthropic was the most responsible of the major players and that, you know, we were trying very hard to get it right. He was calling out honestly the same thing that I'm calling out, which is the difficulty of the situation, right? The structural difficulty of the situation that we're all in and that the plan that I've put out is designed to help us address. We're designed to help us work together. Now, to your question about the government, you know, we are talking to the government every day. We are talking to government officials every day in the Commerce Department, in the Treasury Department, in the intelligence community to tell them as much as we can what we know. And I'm optimistic that if we keep working together with them, we can update them on the latest. I think they really want to do the right thing here. I think they're looking for how they can help. And so I'm optimistic that the companies working together with the government can get to a plan here that makes sense.
All right. So there is some proposed legislation out there, right? You have the Ban Artificial Super Intelligence Act from Senator Sanders and Representative Casar that would punish corporations and individuals working on advanced AI just like those unlawfully developing nuclear weapons. One of the parts of this proposed legislation says entities will be subject to the corporate death penalty and persons shall be subject to not more than 20 years in prison. Do you agree with that?
So, you know what I would say? There's a very wide range of legislation being offered, like the range is incredible. If we go back only one year to 2025, and, you know, there was a bill called SB53 in California. There were similar bills in other states which basically just said you have to disclose what your safety testing plans are, right? And the entire industry was against it. Anthropic was the, or didn't say anything, Anthropic was the only company enthusiastically in support. And then there's all the way to, you know, let's ban all of it. So the AI kill switch could be a good idea. I don't think banning it completely, banning it is the right approach. One, the idea that we would never get the incredible benefits, would never get the medical cures, that doesn't sit right with me. Second, this technology will be built in other countries including authoritarian countries, China among others. And I am worried about how they will use it, how they will use it against free societies, how they will use it to repress their own people. Now, in one place, I do talk about working together with them, and I believe we need to do that to some extent, but I just don't see how it's a stable solution to never build the technology. What we need to do is not stop. We need to slow down. And that is in line honestly with what Anthropic, that's been Anthropic's philosophy since the beginning, right? We've always tried to balance. We've said it's important that we be builders of this technology, but we need to create a race to the top. We need to raise the bar. So, you know, we need to be on the frontier, but we need to do it in the right way. We've said, you know, we're excited to use the models for, you know, national defense applications, but not every national defense application. We've always tried to find the middle ground, the thing that works for everyone, the sensible solution, and that's what we're trying to do here as well. It may not be the perfect solution, but we're trying.
Would you support Congressman Ted Lieu's AI kill switch bill?
I've not looked at that bill more super closely, but what I would say is that the idea that we need some way to, it's not a panacea. It's not the only thing, but the idea that we need some way to make sure that we can deactivate or limit or shut down AI models...
That's not the only thing. If an AI model is powerful enough, it can circumvent attempts to shut it down. And we've seen that in simulations as well. So it may not work at all.