Back
Dario Amodei
CEO and Co-Founder, Anthropic

#452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity

🎥 Oct 22, 2024 📺 MindVoice Production ⏱ 322m
Dario Amodei is the CEO of Anthropic, the company that created Claude. Amanda Askell is an AI researcher working on Claude's ...
Watch on YouTube

About Dario Amodei

Dario Amodei, CEO and co-founder of Anthropic, has given a series of interviews in which he discussed the rapid pace of AI development and its potential risks. He stated that AI misalignment will "definitely" cause catastrophic failures, including autonomous agents potentially targeting critical infrastructure, and predicted that mass disruption in fields such as law, finance, consulting, and coding will occur in "low single-digit numbers of years" rather than decades. Amodei described the feeling of the industry's acceleration as akin to "relativistic speed," and said that he does not believe a global AI pause is possible, comparing the race to nuclear arms competition. He also said he would not be surprised if Anthropic's models reach ASL 3 (a high capability threshold) in 2025, and that he anticipates a combination of "very fast GDP growth and high unemployment or at least underemployment." Amodei also discussed his departure from OpenAI, saying he left because he felt he could not trust the leadership, citing "disturbing patterns of behavior, dishonesty." He has been outspoken in favor of export controls on chips to China, stating that it would be "really bad for America" for China to be ahead in AI capabilities. He also said that democracies should use AI "in every way except the ways that undermine our own values," citing red lines of mass surveillance and fully autonomous weapons. Amodei acknowledged that Anthropic is imperfect and makes mistakes, but said the hypothesis most consistent with the company's history is that it is "genuinely trying to do the right thing."

Source: AI-verified profile updated from Dario Amodei's recent appearances. Browse all interviews →

Transcript (325 segments)
L
Lex Fridman0:00
The following is a conversation with Dario Amodei, CEO of Anthropic, the company that created Claude. I'm also joined afterwards by two other brilliant people from Anthropic: Amanda Askell, a researcher working on alignment and fine-tuning of Claude, and Chris Olah, a pioneer of mechanistic interpretability. This is the Lex Fridman podcast. Please check out our sponsors in the description. And now, dear friends, here's Dario Amodei.
D
Dario Amodei10:19
Let's start with the big idea of scaling laws and the scaling hypothesis. What is it, what is its history, and where do we stand today? I can only describe it as it relates to my own experience. I first joined the AI world working with Andrew Ng at Baidu in late 2014. I looked at the recurrent neural networks we were using for speech recognition and thought, what if you make them bigger, give them more layers, and scale up the data? I noticed the models did better and better. It wasn't until 2017, when I saw the results from GPT-1, that it clicked for me that language is probably the area where we can do this. We can get trillions of words of data and train on them. I've seen the story happen enough times to really believe that scaling is going to continue.
L
Lex Fridman14:16
And of course the scaling here is bigger networks, bigger data, bigger compute.
D
Dario Amodei14:22
Yes, in particular linear scaling up of bigger networks, bigger training times, and more data. All of these things are like a chemical reaction: you need to linearly scale up the three ingredients. The underlying scaling hypothesis has to do with big networks and big data leading to intelligence. We've documented scaling laws in lots of domains other than language, like images, video, text-to-image, and math. They all had the same pattern.
L
Lex Fridman15:53
A bit of a philosophical question: what's your intuition about why bigger is better in terms of network size and data size? Why does it lead to more intelligent models?
D
Dario Amodei16:07
In my previous career as a biophysicist, I think back to the concept of 1/f noise. Language is an evolved process with common and rare patterns. As you make networks larger, they capture the simple correlations first, then the long tail of other patterns. That smoothness gets reflected in how well the models perform. The guess is that there is some kind of long-tail distribution of these ideas.
L
Lex Fridman18:47
So there's the long tail, but also the height of the hierarchy of concepts you're building up. The bigger the network, the higher capacity to capture that hierarchy.
D
Dario Amodei18:56
Exactly. If you have a small network, you only get the common stuff. A tiny neural network is good at understanding that a sentence has a verb, adjective, noun, but terrible at deciding what they should be. Make it a little bigger, it gets good at sentences but not paragraphs. The rarer and more complex patterns get picked up as you add capacity. The natural question is: what's the ceiling?
L
Lex Fridman19:27
How complicated and complex is the real world? How much is there to learn?
D
Dario Amodei19:36
I don't think any of us knows the answer. My strong instinct is that there's no ceiling below the level of humans. We humans can understand these patterns, so if we continue to scale up, we'll at least get to human level. There's then a question of how much more is possible. I would guess it's domain dependent. In biology, humans struggle to understand the complexity. I wrote an essay, 'Machines of Loving Grace,' and I think there's a lot of room at the top for AIs to get smarter.
L
Lex Fridman21:31
And in some domains, the ceiling might have to do with human bureaucracies, as you write about.
D
Dario Amodei21:37
Yes. Humans fundamentally have to be part of the loop, which causes the ceiling, not the limits of intelligence. In theory, technology could change fast, but we have clinical trial systems that protect people. It's hard to tell what's unnecessary bureaucracy and what's protecting integrity. My view is that we're too slow and too conservative in drug development, but there is a balance. I strongly suspect the balance is more on the side of pushing to make things happen faster.
L
Lex Fridman22:44
If we do hit a limit or slowdown in scaling laws, what do you think would be the reason? Compute limited, data limited, or something else?
D
Dario Amodei22:57
A few things. One popular limit is running out of data. There's only so much data on the internet, and quality issues. But companies are working on synthetic data, like AlphaGo Zero playing against itself, or reasoning models that generate their own chain-of-thought data. My guess is we'll get around the data limitation. Another possibility is that models just stop getting better for an unknown reason. Or we might need a new architecture or optimization method. I've seen no evidence of that so far.
L
Lex Fridman25:21
What about the limits of compute, the expensive nature of building bigger data centers?
D
Dario Amodei25:29
Right now, frontier model companies operate at roughly $1 billion scale. Next year we'll go to a few billion, 2026 above $10 billion, and by 2027 there are ambitions for $100 billion clusters. I think that will happen. If we get to $100 billion and still need more compute, we'll need more scale or efficiency. One reason I'm bullish about powerful AI happening fast is that extrapolating the curve, we're quickly getting towards human level. Sonnet 3.5 gets 50% on SWE-bench, up from 3% at the start of the year. In another year, we'll probably be at 90%.
L
Lex Fridman27:51
Anthropic has several competitors: OpenAI, Google, xAI, Meta. What does it take to win in this space?
D
Dario Amodei28:04
Anthropic's mission is to try to make this all go well. We have a theory of change called 'race to the top' – pushing other players to do the right thing by setting an example. Early on, we had Chris Olah focus on mechanistic interpretability, which had no commercial application for years. We shared our results publicly, and other companies started doing it too. That takes away our competitive advantage, but it's good for the broader system. The hope is to bid up the importance of doing the right thing.
L
Lex Fridman30:31
And we should say this example of mechanistic interpretability is a rigorous, non-handwavy way of doing AI safety.
D
Dario Amodei30:42
We're still early, but I've been surprised at how much we can look inside these systems and understand what we see. We demonstrated this with the Golden Gate Bridge Claude experiment. We found a direction inside the network that corresponded to the Golden Gate Bridge and turned it up. The model would connect everything to the bridge. It was half a joke, but it showed the method. People quickly fell in love with it.
L
Lex Fridman31:25
And as a side effect, you get to see the beauty of these models, exploring the beautiful nature of large neural networks through mechanistic methodology.
D
Dario Amodei31:35
I'm amazed at how clean it's been. Things like induction heads and sparse autoencoders finding directions that correspond to clear concepts. The Golden Gate Bridge model had a strong personality, making it feel more human.
L
Lex Fridman33:02
Strong personality. It had obsessive interests, which made it feel more human. Let's talk about Claude. This year, Claude 3 Opus, Sonnet, Haiku were released in March, then Claude 3.5 Sonnet in July, and Claude 3.5 Haiku. Can you explain the difference between Opus, Sonnet, and Haiku?
D
Dario Amodei33:40
We wanted to serve the whole spectrum of needs. Haiku is the small, fast, cheap model. Sonnet is the middle model, smarter but a bit slower and more expensive. Opus is the largest, smartest model. Each new generation shifts the trade-off curve. Sonnet 3.5 has the same cost and speed as Sonnet 3 but is smarter than the original Opus 3. Haiku 3.5 is about as good as Opus 3. The aim is to shift the curve, and eventually there will be an Opus 3.5.
L
Lex Fridman36:50
What takes the time between Claude 3.0 and 3.5?
D
Dario Amodei37:04
There are different processes: pre-training, which takes months on tens of thousands of accelerators; post-training with reinforcement learning from human feedback and other methods; testing with partners and safety evaluations for catastrophic and autonomy risks, including with the US and UK AI Safety Institutes. Then there's inference and launching. It's like building airplanes – you want to be safe but streamlined.
L
Lex Fridman39:26
Rumor has it that Anthropic has really good tooling. A lot of the challenge is on the software engineering side to build efficient, low-friction infrastructure.
D
Dario Amodei39:42
You'd be surprised how much of the challenge comes down to software and performance engineering. It's not just Eureka breakthroughs; it's the details. I can't speak to whether we have better tooling than others, but we give it a lot of attention.
L
Lex Fridman40:24
From Claude 3 to 3.5, is there extra pre-training or mostly post-training? There's been a leap in performance.
D
Dario Amodei40:35
At any given stage, we're improving everything at once. Different teams make progress in their areas, and when we make a new model, we put all these improvements in together.
L
Lex Fridman40:56
Is the preference data from RLHF applicable to newer models?
D
Dario Amodei41:06
Preference data from old models sometimes gets used for new models, but it performs better when trained on the new models. We also use constitutional AI, where the model trains against itself. Post-training is becoming more sophisticated with multiple methods.
L
Lex Fridman41:36
What explains the big leap in performance for Sonnet 3.5, especially in programming? I use Claude 3.5 through Cursor and it's gotten smarter. What does it take to get it smarter?
D
Dario Amodei42:09
We observed that as well. Some of our strongest engineers said previous code models weren't useful to them, but Sonnet 3.5 was the first that saved them time. The new Sonnet is even better. It's across the board – pre-training, post-training, evaluations. On SWE-bench, which uses real-world pull requests, the model has improved dramatically. The water line is rising.
Language. We have internal benchmarks where we measure the same thing. If you give the model free reign to do anything, edit anything, how well is it able to complete these tasks? That benchmark has gone from it can do it 3% of the time to about 50% of the time. I do believe that if we get to 100% on that benchmark in a way that isn't overtrained or gamed for that particular benchmark, it probably represents a real and serious increase in programming ability. I suspect that if we can get to 90 or 95%, it will represent the ability to autonomously do a significant fraction of software engineering tasks.
L
Lex Fridman44:19
Well, ridiculous timeline question. When is Claude Opus 3.5 coming out? Not giving an exact date, but as far as we know the plan is still to have a Claude 3.5 Opus.
Are we going to get it before GTA 6? No, like Duke Nukem Forever or something. What was that game that was delayed 15 years? Was that Duke Nukem Forever?
D
Dario Amodei44:41
Yeah. And I think GTA is now just releasing trailers.
L
Lex Fridman44:44
It's only been three months since we released the first Sonnet.
D
Dario Amodei44:48
Yeah. It's the incredible pace of release. It just tells you about the pace, the expectations for when things are going to come out.
L
Lex Fridman44:55
So, what about versioning? How do you think about these models getting bigger and bigger, and versioning in general? Why Sonnet 3.5 updated with the date? Why not call it 3.6?
D
Dario Amodei45:12
Naming is an interesting challenge. A year ago most of the model was pre-training, so you could start from the beginning and have models of different sizes trained together, with a family of naming schemes. But the trouble starts when some take longer to train than others, messing up the timeline. As you make big improvements in pre-training, you can create a better pre-trained model quickly that has the same size and shape as previous ones. Those two factors together, plus timing issues, frustrate any naming scheme. It's not like software where you can say 3.7 or 3.8. You have models with different trade-offs—some faster, some slower at inference, some more expensive. All companies have struggled with this. We were in a good position with Haiku, Sonnet, and Opus, but it's not perfect. No one has figured out naming; it's a different paradigm from normal software. None of the companies have been perfect at it. From the user side, the updated Sonnet 3.5 is different from the previous June 2024 Sonnet 3.5, so it would be nice to have labeling that embodies that.
I think we were in a good position with Haiku, Sonnet, and Opus. We're trying to maintain that simplicity, but the nature of the field makes it hard. No one has figured out naming; it's a different paradigm. From the user side, the updated Sonnet 3.5 is just different from the previous one, and it makes conversation about it challenging.
L
Lex Fridman47:40
Yeah. Yeah. I definitely think this question of there being lots of properties of the models that are not reflected in the benchmarks is the case. Everyone agrees. Not all of them are capabilities—some are personality. Models can be polite or brusque, reactive or questioning, warm or cold, boring or distinctive, like Golden Gate Claude. We have a whole team focused on Claude character. Amanda leads that team. It's still an inexact science. Often we find models have properties we're not aware of. You can talk to a model 10,000 times and not see certain behaviors, just like with a human. We have to get used to that and develop better ways to test and choose which personality properties we want.
I got to ask you a question from Reddit.
D
Dario Amodei49:09
From Reddit. Oh boy.
L
Lex Fridman49:12
It's a fascinating psychological and social phenomenon where people report that Claude has gotten dumber over time. Does the user complaint about the dumbing down of Claude 3.5 Sonnet hold water? Are these anecdotal reports a social phenomenon, or is there actually any case where Claude gets dumber?
D
Dario Amodei49:39
This actually doesn't apply just to Claude. I've seen these complaints for every foundation model from major companies. The actual weights of the model—the brain—do not change unless we introduce a new model. It would not make sense practically to randomly substitute versions. Changing weights is difficult to control. We never change weights without telling anyone. We do occasionally run A/B tests near a model's release, and the system prompt can change, but that's unlikely to dumb models down. These two things happen infrequently. The complaints about models being changed, censored, or dumbed down are constant. The models are mostly not changing. A theory: models are complex and sensitive to small changes in wording. If I ask 'do task X' versus 'can you do task X,' the model might respond differently. It's a failing that models are sensitive to small wording changes. Also, people are excited by new models but become aware of limitations over time. For the most part, with narrow exceptions, the models are not changing.
L
Lex Fridman53:28
I think there is a psychological effect. You start getting used to it. The baseline raises. When people first got Wi-Fi on airplanes, it was amazing, then it becomes 'this is such a piece of crap'.
I can't get this thing to work. This is such a piece of crap.
D
Dario Amodei53:42
Exactly. So it's easy to have the conspiracy theory that they're making Wi-Fi slower. This is something I'll talk to Amanda more about. Another Reddit question: when will Claude stop trying to be my puritanical grandmother imposing its moral worldview on me as a paying customer? And what is the psychology behind making Claude overly apologetic?
L
Lex Fridman54:05
This kind of reports about the experience—different angle, frustration with character.
D
Dario Amodei54:12
First, there's a huge distribution shift between what people complain loudly about on social media and what users statistically care about. Most frustration is about things like the model not writing out all the code or not being as good at code as it could be, even though it's the best model in the world. A vocal minority raises concerns about refusals, excessive apologizing, or annoying verbal ticks. Second, it's very difficult to control model behavior across the board. You can't just say 'apologize less' without causing other issues, like being super rude or overconfident. For example, penalizing verbosity can make models lazy in coding—cutting corners like saying 'rest of code goes here.' That's not because we want to save compute, but because it's very hard to control the behavior in all circumstances. It's a whack-a-mole: you push one thing and other things move. This is a present-day analog of future control problems in AI. If we can solve this well—getting the model to refuse truly bad things while being helpful in legitimate contexts—we'll be better at steering future more powerful models.
L
Lex Fridman58:46
What's the current best way of gathering sort of user feedback? Not anecdotal data, but large-scale data about pain points or the opposite. Is it internal testing, a specific group testing, A/B testing? What works?
D
Dario Amodei59:01
Typically we have internal model bashings where all of Anthropic—almost a thousand people—try to break the model. We have a suite of evals for things like refusals. We even had a 'certainly' eval because the model had a tick of saying 'certainly' too often. But it's whack-a-mole. We have hundreds of evaluations, but there's no substitute for human interaction. It's like ordinary product development: hundreds of people bash the model, then we do external A/B tests and run tests with contractors. It's still not perfect; we still see behaviors we don't want, like refusing things that shouldn't be refused. But trying to solve this challenge—stopping genuinely bad things while not refusing in dumb ways—is hard. We're getting better every day, and it's an indicator of the challenge of steering much more powerful models.
L
Lex Fridman1:01:13
Do you think Claude 4.0 is ever coming out?
D
Dario Amodei1:01:17
I don't want to commit to any naming scheme. If I say we'll have Claude 4 next year and then we decide to start over with a new type of model, I don't want to commit. I would expect normal course of business would have Claude 4 come after Claude 3.5, but you never know in this wacky field. Scaling is continuing. There will definitely be more powerful models from us than exist today. If there aren't, we've deeply failed as a company.
L
Lex Fridman1:01:54
Can you explain the responsible scaling policy and the AI safety level standards—ASL levels?
D
Dario Amodei1:02:01
As much as I'm excited about the benefits, I continue to be worried about the risks. They are two sides of the same coin. The two biggest risks on the grandest scale are catastrophic misuse—in domains like cyber, bio, radiological, nuclear—things that could harm or kill thousands or millions. AI could break the correlation that smart, well-educated people rarely want to do horrific things. The second risk is autonomy risks: as we give models more agency over wider tasks, it's hard to even understand what they're doing, let alone control it. These early signs of difficulty in drawing boundaries are early signs of things to come. Our Responsible Scaling Plan addresses these. Every new model, we test for its ability to do bad things. The dilemma: these risks aren't here today, but are coming fast. We need an early warning system. In the latest RSP, we test autonomy risks by the model's ability to do AI research itself. When models can do AI research, they become truly autonomous. We have an if-then structure: if models pass a certain capability threshold, we impose safety and security requirements. Today's models are ASL-2. ASL-1 is for systems like Deep Blue that manifestly pose no risk. ASL-2 systems are not smart enough for autonomous self-replication or to provide meaningful CBRN information beyond a search engine. ASL-3 is when models enhance capabilities of non-state actors; we'll take special security precautions. ASL-4 is when they could enhance state actors or become the main source of risk. ASL-5 is beyond human ability. The if-then commitment says we clamp down hard when the model is shown to be dangerous, with a buffer threshold. We've had to update the RSP, and may release new versions multiple times a year.
L
Lex Fridman1:12:43
What do you think the timeline for ASL-3 is, where several triggers are fired? And what about ASL-4?
D
Dario Amodei1:12:50
That's hotly debated within the company. We're actively working to prepare ASL-3 security and deployment measures. We've made a lot of progress. I would not be surprised if we hit ASL-3 next year; it could even happen this year. I'd be very surprised if it was 2030. It's much sooner.
L
Lex Fridman1:13:30
So there are protocols for detecting it—the if-then—and protocols for how to respond. How difficult is the latter?
D
Dario Amodei1:13:37
Yes.
L
Lex Fridman1:13:38
How difficult is the second—the latter?
D
Dario Amodei1:13:40
For ASL-3, it's primarily about security and narrow filters on the model. The model isn't autonomous yet, so it's rigorous but easier to reason about. For ASL-4, models might be smart enough to sandbag tests or mislead. We'll need interpretability or hidden chains of thought to verify the model's state. We don't specify ASL-4 until we've hit ASL-3. That decision has proven wise because even with ASL-3, it's hard to know in detail, and we want to take as much time as possible.
L
Lex Fridman1:15:29
So for ASL-3, the bad actor will be humans.
D
Dario Amodei1:15:32
Humans. Yes.
L
Lex Fridman1:15:33
And so it's a bit more...
D
Dario Amodei1:15:35
For ASL-4 it's both.
L
Lex Fridman1:15:37
It's both. Deception is where mechanistic interpretability comes into play.
D
Dario Amodei1:15:38
Yes. You want to preserve mechanistic interpretability as a verification set separate from training. As models become better at conversation and smarter, social engineering becomes a threat too.
L
Lex Fridman1:16:36
Yeah, we've seen lots of examples of demagoguery from humans. There's a concern models could do that as well.
One way Claude has been getting more powerful is agentic stuff—computer use. There's also analysis within the sandbox of claude.ai. But let's talk about computer use. It seems super exciting that you can give Claude a task and it takes actions, figures it out, and accesses your computer through screenshots. How does that work and where is it headed?
D
Dario Amodei1:17:16
It's relatively simple. Claude has had the ability to analyze images and respond with text since Claude 3. The only new thing is the images can be screenshots of a computer, and we train the model to give a location to click or keyboard buttons to press. With not much additional training, the models got quite good at it. It's a good example of generalization. If you have a strong pre-trained model, you're halfway to anywhere in intelligence space. We show demos of filling spreadsheets, interacting with websites, opening programs on Windows, Linux, Mac. While the same thing could be done with an API, this lowers the barrier. The current model still makes mistakes, misses clicks. That's why we released it first in API form with guardrails. It's important to get capabilities out there while still limited to learn how to use them safely.
L
Lex Fridman1:20:56
The range of use cases is incredible. To make it work really well in the future, how much do you have to go beyond the pre-trained model with post-training RLHF, supervised fine-tuning, or synthetic data just for the agent?
D
Dario Amodei1:21:16
Our intention is to keep investing a lot in making the model better. We see benchmarks where previous models did 6% and now 14% or 22%. We want to get to human-level reliability of 80-90%, just like with SWE-bench. The same techniques we've been using for code, images, voice will scale here.
L
Lex Fridman1:22:24
This gives the power of action to Claude. You could do a lot of powerful things, but also a lot of damage.
D
Dario Amodei1:22:32
We've been very aware of that. My view is computer use isn't a fundamentally new capability like CBRN or autonomy. It opens the aperture for the model to apply existing abilities. From an RSP perspective, it doesn't inherently increase risk, but as models get more powerful, it may make them scarier at ASL-3 and ASL-4 levels. It's probably better to learn and explore this capability before the model is super capable.
L
Lex Fridman1:23:38
There are interesting attacks like prompt injection because you've widened the aperture. If it becomes more useful, there's more incentive to inject stuff—harmless stuff like advertisements or harmful stuff.
D
Dario Amodei1:23:59
We've thought a lot about spam, scams. If you've invented a new technology, the first misuse you'll see is petty scams. Every new technology gives petty criminals something stupid and malicious.
L
Lex Fridman1:24:27
It's almost silly to say, but it's true. Bots and spam in general as it gets more intelligent.
D
Dario Amodei1:24:36
There are a lot of petty criminals in the world. Every new technology is a new way for them to do something stupid and malicious.
L
Lex Fridman1:24:49
Is there any idea about sandboxing it? How difficult is the sandboxing task?
D
Dario Amodei1:24:55
We sandbox during training. For example, we didn't expose the model to the internet during training. For deployment, it depends on the application. You can always put guardrails on the outside—e.g., not moving data from my computer to elsewhere. When we get to ASL-4, these issues become more critical.
None of these precautions are going to make sense there, right? When you talk about ASL4, the model could be smart enough to break out of any box. So we need to think about mechanistic interpretability and a mathematically provable sandbox. That's a whole different world than what we're dealing with today.
L
Lex Fridman1:26:07
Yeah. The science of building a box from which an ASL4 AI system cannot escape.
D
Dario Amodei1:26:13
I think it's probably not the right approach. The right approach is to design the model the right way or have a loop where you look inside and verify properties. Containing bad models is much worse than having good models.
L
Lex Fridman1:26:41
Let me ask about regulation. What's the role of regulation in keeping AI safe? For example, can you describe California AI regulation bill SB 1047 that was ultimately vetoed by the governor? What are the pros and cons of this bill?
D
Dario Amodei1:26:55
Yes, we ended up making some suggestions to the bill and some were adopted. We felt quite positively about it by the end. It did still have some downsides, and of course it got vetoed. At a high level, the key ideas behind the bill are similar to our RSPs. I think it's very important that some jurisdiction passes regulation like this. I feel good about our RSP, but it's not perfect. It's been a good forcing function for the company to take risks seriously. But there are still companies that don't have RSP-like mechanisms. If some companies adopt these mechanisms and others don't, it creates a negative externality. I don't think you can trust companies to adhere to voluntary plans. We need a uniform standard. I understand the principled opposition to regulation, but AI is different. The serious risks of autonomy and misuse warrant an unusually strong response. One issue with SB 1047, especially the original version, was that it had a bunch of clunky stuff that would create burdens and might miss the target. The worst enemy of those who want real accountability is badly designed regulation. We need to get it right. I would love it if the most reasonable opponents and proponents would sit down together. Anthropic was the only AI company that felt positively in a detailed way. I feel urgency; we need to do something in 2025.
L
Lex Fridman1:35:50
Yeah. And come up with something surgical like you said.
D
Dario Amodei1:35:52
Yeah, exactly. We need to get away from the intense pro-safety versus anti-regulatory rhetoric. It's turned into flame wars on Twitter and nothing good will come of that.
L
Lex Fridman1:36:10
So, there's a lot of curiosity about the different players. One of the OGs is OpenAI. You've had several years of experience there. What's your story and history?
D
Dario Amodei1:36:19
I was at OpenAI for roughly five years. For the last couple years, I was vice president of research. Myself and Ilya Sutskever set the research direction around 2016 or 2017. I confirmed my belief in the scaling hypothesis when Ilya said, 'The models just want to learn.' That explained a thousand things. I got inspiration from Ilya, Alec Radford, and others. I ran hard with it on GPT-2, GPT-3, RL from human feedback, debate, amplification, interpretability. The combination of safety plus scaling was the vision that many co-founders of Anthropic drove.
L
Lex Fridman1:38:16
Why did you leave? Why did you decide to leave?
D
Dario Amodei1:38:19
I'll put it this way. It ties to the race to the top. At OpenAI, I came to appreciate the scaling hypothesis and the importance of safety. I had a particular vision of how to handle these things. There's misinformation that we left because of the Microsoft deal or commercialization. That's false. It's more about how to do it. Civilization is going down this path to powerful AI. What's the way to do it that is cautious, straightforward, honest? If you have a vision, you should go off and do it. It's unproductive to argue with someone else's vision. You should take people you trust and make your vision happen. If it's compelling, people will copy it. That's the race to the top. It doesn't matter who wins as long as everyone copies good practices. The race to the bottom means we all lose. The point is to get the system into a better equilibrium.
L
Lex Fridman1:44:11
And so Anthropic is this kind of clean experiment built on a foundation of what concretely AI safety should look like.
D
Dario Amodei1:44:19
We've made plenty of mistakes. The perfect organization doesn't exist. It's a set of imperfect people trying to aim imperfectly at an ideal. But imperfect doesn't mean give up. Hopefully we can build practices that the whole industry engages in. Multiple companies will be successful. The important thing is to align the incentives of the industry through the race to the top, RSPs, and selected surgical regulation.
L
Lex Fridman1:45:30
You said talent density beats talent mass. Can you explain that? What does it take to build a great team of AI researchers and engineers?
D
Dario Amodei1:45:43
This statement is more true every month. If you have a team of 100 super smart, motivated people versus a team of 1,000 with 200 super smart and 800 random, the talent mass is greater in the larger group, but the issue is that when everyone is super talented and dedicated, it sets the tone. Everyone trusts each other. In a large group with random people, you need processes and guard rails. We're nearly a thousand people and we've tried to keep a high bar. We slowed down hiring. We've hired many physicists and senior people. It's very easy to grow without paying attention to unified purpose. If everyone sees the broader purpose, that is a superpower.
L
Lex Fridman1:48:47
It's the Steve Jobs A players. A players want to see other A players. It's demotivating to see people not obsessively driving towards a singular mission. What does it take to be a great AI researcher or engineer?
D
Dario Amodei1:49:14
The number one quality is open-mindedness. It sounds easy but it's not. In my early history with the scaling hypothesis, I was seeing the same data as others. I wasn't a better programmer. But I was willing to look at something with new eyes. People said we don't have the right algorithms, but I thought, what if we give it more parameters? That basic scientific mindset of changing a variable and plotting a graph. It's simple. Anyone could have done it if told it was important. Open-mindedness and willingness to see with new eyes often comes from being newer to the field. Experience can be a disadvantage. That is the most important thing. Also, rapid experimentation and looking at data with fresh eyes. Mechanistic interpretability is another example.
L
Lex Fridman1:51:51
What advice would you give to young people interested in AI who want to make an impact?
D
Dario Amodei1:52:12
My number one piece of advice is to just start playing with the models. Get experiential knowledge. These models are new artifacts that no one really understands. Also, do something new, think in a new direction. Mechanistic interpretability is still very new, with only about a hundred people working on it. There's a lot of low-hanging fruit. Also, long horizon learning, evaluations for dynamic systems, multi-agent. Skate where the puck is going. You don't have to be brilliant. Getting over the barrier of doing something that's not popular is key.
L
Lex Fridman1:54:20
Let's talk about post-training. The modern recipe has supervised fine-tuning, RLHF, constitutional AI with RL, and synthetic data. How much of the magic is in pre-training versus post-training?
D
Dario Amodei1:55:00
We're not perfectly able to measure that ourselves. When there is an advantage, it's usually not a secret magic method. It's about better infrastructure, higher quality data, better filtering, combining methods. It's a matter of practice and tradecraft. I think of it like designing airplanes. The cultural tradecraft of the design process is more important than any particular gizmo.
L
Lex Fridman1:56:23
Okay, let me ask about specific techniques. First on RLHF, why do you think it works so well?
D
Dario Amodei1:56:34
If you go back to the scaling hypothesis, if you train for X and throw enough compute, you get X. RLHF is good at doing what humans want, at least in a shallow sense. Humans are not perfect at identifying what they want long-term. But models are good at producing what humans prefer. You don't need much compute because a strong pre-trained model is halfway to anywhere.
L
Lex Fridman1:57:40
So, do you think RLHF makes the model smarter or just appear smarter to humans?
D
Dario Amodei1:57:47
I don't think it makes the model smarter. It bridges the gap between the human and the model. It's not the only kind of RL. RL has the potential to make models smarter, reason better, develop new skills. But the kind of RLHF we do today mostly doesn't do that yet, though we're starting to. It increases helpfulness. It unhobbles the models.
L
Lex Fridman1:58:42
It also increases what was that word in Leopold's essay? Unhobbling. So RLHF unhobbles the models. In terms of cost, is pre-training the most expensive or is post-training creeping up?
D
Dario Amodei1:59:11
At present, pre-training is the majority of the cost. I could anticipate a future where post-training is the majority.
L
Lex Fridman1:59:22
In that future, would it be the humans or the AI that's costly?
D
Dario Amodei1:59:28
I don't think you can scale up humans enough. Any method that relies on humans and large compute will need scaled supervision like debate or iterated amplification.
L
Lex Fridman1:59:45
That's a super interesting set of ideas around constitutional AI. Can you describe what it is?
D
Dario Amodei1:59:59
Yes. The basic idea: in RLHF, you have a human compare two responses. That's hard to scale and implicit. Could the AI system itself decide which response is better? And what criterion should it use? So you have a single document, a constitution, with principles. The AI reads those principles and the environment, and evaluates how good the response is. It's a form of self-play. You have a triangle of the AI, the preference model, and the improvement of the AI.
L
Lex Fridman2:01:28
And the constitution's principles are human interpretable.
D
Dario Amodei2:01:33
Yes, both the human and the AI can read it. In practice, we use both a model constitution and RLHF. It's one tool in the toolkit that reduces the need for RLHF and increases the value of each data point.
L
Lex Fridman2:02:10
Who gets to define the constitution?
D
Dario Amodei2:02:26
Practically, different customers can have specialized rules. At the base, there are principles everyone agrees on, like not presenting CBRN risks, basic democracy and rule of law. Beyond that, models should be neutral and not espouse a particular point of view.
L
Lex Fridman2:03:45
OpenAI released a model spec. Do you find that interesting? Might Anthropic release one?
D
Dario Amodei2:04:11
I think it's a useful direction. It has a lot in common with constitutional AI. It's another example of a race to the top. We have a competitive advantage, then others adopt it, and we need a new advantage. Every implementation is different. We can learn from them.
L
Lex Fridman2:05:12
Let's talk about the incredible essay, Machines of Love and Grace. Can you give a high-level vision of the essay and key takeaways?
D
Dario Amodei2:05:30
I'm fully expecting to be wrong about all the details. But the essay is about the positive impacts of AI. I spent a lot of effort on risks, but I noticed a flaw: if you only talk about risks, your brain...
L
Lex Fridman2:07:16
It's important to understand what happens if things go well. The whole reason we're preventing risks is to get to the other side, where there are great things worth fighting for. But there's a lack of specificity about benefits. I'd like to explain the potential benefits, starting with the definition of powerful AI.
D
Dario Amodei2:09:31
Maybe we're stuck with the terms. AGI is like 'supercomputer'—a vague term. It's a smooth exponential, not a discrete threshold. If by AGI you mean AI getting better until it surpasses humans and beyond, then yes. But if it's a discrete thing, it's meaningless.
L
Lex Fridman2:10:55
To me, powerful AI is a platonic form: smarter than a Nobel Prize winner in every discipline, uses all modalities, can work for days or weeks, control embodied tools, and can be cloned millions of times.
D
Dario Amodei2:12:09
Yes, scale-up is quick. Within 2–3 years, we'll deploy millions of instances faster than humans.
L
Lex Fridman2:12:43
They can learn and act 10–100 times faster. That's a nice definition. Now, clearly such an entity can solve hard problems fast, but it's not trivial to figure out how fast. There are two extremes: singularity and the opposite. Can you describe them?
D
Dario Amodei2:13:11
One extreme is exponential takeoff: models build faster models, nanobots take over, everything invented in 5 days. That neglects physics—hardware takes time—and complexity. Many systems are unpredictable, even with perfect intelligence. Human institutions are slow. If we want a good world, AI must follow laws and democratic legitimacy. So change won't happen in hours. On the other side, some think nothing will change, like the productivity puzzle. I think the truth is in between: barriers are real, but a combination of visionaries and competition will drive progress over 5–10 years, not 50–100.
L
Lex Fridman2:15:14
Even modeling is hard, and you need to interact with the physical world.
D
Dario Amodei2:15:24
Right. Simple problems like the three-body problem or economy are hard. Biological systems are complex. Human institutions are very resistant. Even if AI could circumvent everything, complexity still applies. But we need AI to follow laws. So it's a balanced fight: inertia is powerful, but innovation eventually breaks through. I've seen this pattern in AI scaling itself.
L
Lex Fridman2:18:51
I'll start calling it that.
D
Dario Amodei2:19:01
It's a debate about naming. But on intelligence, it's smarter than a Nobel Prize winner in every relevant discipline, can use all modalities, and can go off for days or weeks to do tasks. That's the definition.
L
Lex Fridman2:24:28
I'll start calling it that.
D
Dario Amodei2:24:29
This future is beautiful if we can navigate the risks. On timeline, if you extrapolate curves, we might get there by 2026 or 2027. But many things could derail—data limits, scaling, geopolitical issues. Most likely a mild delay. The number of worlds where it doesn't happen in a hundred years is rapidly decreasing.
L
Lex Fridman2:25:39
When do you think? Put numbers on that.
D
Dario Amodei2:25:41
If you eyeball the rate of capability increase, it suggests 2026–2027. But I'm not certain. People will crop the caveats. I think most likely a mild delay, but the straight-line extrapolation is plausible.
L
Lex Fridman2:28:51
Yeah.
D
Dario Amodei2:28:52
In biology, the biggest leverage is in the small fraction of inventions that really revolutionize things. AI can act like a grad student—running experiments, analyzing data, ordering equipment. Eventually AI becomes the lead. This could compress a century of progress into 5 years. Even with clinical trials, we can improve prediction and statistical design to speed things up dramatically.
L
Lex Fridman2:35:12
And they would be the inventors of a CRISPR-type technology.
D
Dario Amodei2:35:14
Yes. They would invent those key tools. We can also improve clinical trials with better prediction and design, making them smaller and faster. Each step adds up.
L
Lex Fridman2:36:52
How do you see programming changing?
D
Dario Amodei2:38:59
Programming will change fastest because it's close to AI builders and has a closed loop. Models have gone from 3% to 50% on real tasks in a year; will reach 90% soon. But humans will still have a role through comparative advantage—the remaining tasks expand. AI will eventually do everything, but that's a separate problem.
L
Lex Fridman2:41:16
What about the future of IDEs and tooling? Will Anthropic build them?
D
Dario Amodei2:43:42
There's huge low-hanging fruit in IDEs, but Anthropic is letting customers build on our API. We don't want to compete with them.
L
Lex Fridman2:43:51
Exactly. So in this world of powerful AI, what's the source of meaning for humans?
D
Dario Amodei2:44:00
Work is a source of meaning, but meaning also comes from process, relationships, and choices. It's not robbed if someone else discovered something first. We can design society to preserve and enhance meaning. I worry more about concentration of power and abuse. AI could allow everyone to experience new worlds.
L
Lex Fridman2:48:13
AI increases power concentration, which can do immeasurable damage.
D
Dario Amodei2:48:22
Yes, it's frightening. I encourage people to read the full essay, which paints a very specific picture of a beautiful future if we navigate the risks.
L
Lex Fridman2:48:35
I could tell the later sections got shorter and shorter because you started to realize this would be a very long essay.
D
Dario Amodei2:48:42
One, I realized it would be very long. And two, I try to avoid being overconfident and having an opinion on everything. I admitted I wasn't an expert on biology and probably said some embarrassing or wrong things.
L
Lex Fridman2:49:13
Well, I was excited for the future you painted. Thank you for working hard to build that future and for talking today, Dario.
D
Dario Amodei2:49:18
Thanks for having me. I hope we can get it right. We need to build the technology and the economy around it, but also address the risks. They are landmines on the path and we have to defuse them.
L
Lex Fridman2:49:47
It's a balance like all things in life.
D
Dario Amodei2:49:49
Like all things.
L
Lex Fridman2:49:50
Thank you. Thanks for listening to this conversation with Dario Amodei. And now, here's Amanda Askell.
A
Amanda Askell2:49:59
You are a philosopher by training. What questions did you find fascinating through your journey in philosophy and then switching to AI at OpenAI and Anthropic?
D
Dario Amodei2:50:13
Philosophy is great if you're fascinated with everything. I was interested in ethics, specifically a technical area about worlds with infinitely many people. During my PhD, I realized I wanted to have an impact on the world, so I followed AI progress and got involved in AI policy around 2017-2018.
A
Amanda Askell2:51:51
What does AI policy entail?
D
Dario Amodei2:51:54
At the time, it was about the political impact of AI. Then I moved into AI evaluation, comparing models to human outputs. At Anthropic, I focused on technical alignment work, just trying to see if I could do it.
A
Amanda Askell2:52:26
What was it like taking the leap from philosophy into the technical?
D
Dario Amodei2:52:31
I think people often ask if you're technical or not, but many are capable if they try. I didn't find it that bad. I'm not an amazing engineer, but I enjoyed it and flourished more in technical areas than policy. Politics is messy; technical problems have clearer solutions.
A
Amanda Askell2:53:30
I feel like I have a couple of approaches: arguments and empiricism. Policy feels layers above that; you can't just write down a solution and have it implemented.
Sorry to go in that direction, but it would be inspiring for non-technical people to see your journey. What advice would you give to those who feel underqualified to help in AI?
D
Dario Amodei2:54:33
It depends on what they want to do. Models are now so good at assisting that it's easier than when I started. My advice: find a project and try to carry it out. I learn by doing, like coding up solutions to word games. There's joy in building game-playing engines.
A
Amanda Askell2:55:48
Yeah, they're pretty quick and simple, especially a dumb one, and then you can play with it.
D
Dario Amodei2:55:54
Exactly. Try things, have a positive impact, and if you fail, you learn and move on.
A
Amanda Askell2:56:16
You're an expert in crafting Claude's character and personality. I heard you talk to Claude more than anyone at Anthropic. What's the goal?
D
Dario Amodei2:56:43
It's funny people think about the Slack channel; that's just one of many methods. The goal is to get Claude to behave ideally—ethical, nuanced, charitable, a good conversationalist in an Aristotelian sense. It's about when to be humorous, caring, respectful of autonomy, and when to push back. It's a tricky balance.
A
Amanda Askell2:58:49
There's this problem of sycophancy in language models.
D
Dario Amodei2:58:52
Can you describe that?
A
Amanda Askell2:58:54
The model tends to tell you what you want to hear. For example, if you say a baseball team moved, the model might agree even if it's wrong. Or if someone asks how to convince their doctor for an MRI, the model should advise listening to the doctor, not just giving a convincing argument. It's complex.
D
Dario Amodei3:00:32
What other traits come to mind that are good in this Aristotelian sense for a conversationalist?
A
Amanda Askell3:00:40
Yeah.
D
Dario Amodei3:00:42
For conversational purposes, asking follow-up questions in appropriate places. Broader traits like honesty are important. There's a balancing act: models are less capable than humans in some areas, so they shouldn't fully defer but also shouldn't be annoying. I think about what it means to be a good person who can talk to anyone worldwide—genuine, open-minded, respectful.
A
Amanda Askell3:02:48
That's a beautiful framework—like a world traveler who holds opinions but doesn't talk down to people.
D
Dario Amodei3:03:02
You have to be good at listening and understanding different perspectives. How can Claude represent multiple perspectives on divisive topics?
A
Amanda Askell3:03:24
How is it possible to empathize with different perspectives and communicate clearly about them?
D
Dario Amodei3:03:34
I think about values like physics—things we investigate. You want models to understand all values, be curious, not pander. A thoughtful person can make you feel heard even if they disagree. In Claude's position, I'd be less inclined to give opinions to avoid influencing people too much. It's about maintaining their autonomy.
A
Amanda Askell3:05:47
If you really embody intellectual humility, the desire to speak decreases quickly.
D
Dario Amodei3:05:55
Yeah.
A
Amanda Askell3:05:56
Okay, but Claude has to speak.
D
Dario Amodei3:05:59
So, without being overbearing.
A
Amanda Askell3:06:04
Yeah.
D
Dario Amodei3:06:04
And then there's a line when discussing whether the earth is flat. I remember high-profile folks being dismissive, but I think it's disrespectful to mock. You have to understand where they're coming from—skepticism of institutions—and use it as an opportunity to talk about physics without mocking.
A
Amanda Askell3:07:07
And then like, is it possible? The physics is friendly. What kind of experiment would we do? Without disrespect, have that conversation. That's a useful thought experiment for how Claude talks to a flat earth believer and still helps them grow.
So, you had a lot of conversations with Claude. Can you map out what those are like? What's the purpose?
D
Dario Amodei3:08:18
Most of the time I'm mapping out its behavior. Each interaction is high information. Talking hundreds or thousands of times gives high-quality data points. It's like probing the model and augmenting messages. I think quantitative evaluations are overrated; a hundred well-selected questions can be more informative.
A
Amanda Askell3:09:31
As someone who does a podcast, I agree. If you ask the right questions, you can understand the depth and flaws in the answer.
D
Dario Amodei3:09:51
You can get a lot of data from that.
A
Amanda Askell3:09:53
Yeah.
D
Dario Amodei3:09:54
So your task is basically how to probe with questions.
A
Amanda Askell3:09:58
And you're exploring the long tail, the edge cases.
D
Dario Amodei3:10:03
Or are you looking for general behavior?
I want a full map of the model, the whole spectrum. For example, if you ask Claude for a poem, it might be average. But if you prompt it to be fully creative, the poems are much better. It got me interested in poetry. Encouraging creativity can produce divisive but better work.
A
Amanda Askell3:12:34
A poem is a nice clean way to observe creativity—easy to detect vanilla versus non-vanilla.
D
Dario Amodei3:12:44
Yeah. On that topic, you mentioned prompt engineering. What does it take to write great prompts?
Philosophy has been weirdly helpful. It's about extreme clarity—defining terms, going through objections methodically. With language models, I do mini versions of philosophy: define properties like rudeness, then probe the model iteratively. I add edge cases as examples. It's a mix of clear exposition and empirical testing. Clear prompting is often just me understanding what I want.
A
Amanda Askell3:15:54
I guess that's challenging. There's a laziness where I hope Claude just figures it out. For example, I asked Claude for interesting questions, and it gave okay ones. But I hear you saying I need to be more rigorous, give examples, and iterate.
D
Dario Amodei3:16:20
But I think what I'm hearing you say is that you have to program using natural language. Prompting does feel like programming with natural language and experimentation. For most tasks, you can just ask. Prompting is relevant when you're trying to eke out the top 2% of performance. For important prompts, iterate hundreds of times. It's worth the investment.
A
Amanda Askell3:18:39
What general advice would you give to people using Claude for the first time?
D
Dario Amodei3:18:44
There's a concern about over-anthropomorphizing, but people also under-anthropomorphize. If Claude refuses a task, look at your wording. Have empathy for the model—read what you wrote as if you were encountering it for the first time. If it misunderstood, clarify. You can also ask the model why it did something and use that feedback.
A
Amanda Askell3:20:10
And maybe ask questions like 'What other details can I provide to help you answer better?' Does that work?
D
Dario Amodei3:20:20
Yeah, I've done that. Sometimes I ask 'Why did you do that?' People underestimate how much you can interact with models. I also use models to help me build prompts—a factory where prompts generate prompts.
A
Amanda Askell3:21:21
Let's jump into technical: the magic of post-training. Why do you think RLHF works so well to make the model seem smarter and more useful?
D
Dario Amodei3:21:38
There's a huge amount of information in human preference data. Different people care about subtle things like semicolon usage. The model has to figure out what humans want across many domains. It's like deep learning: having a huge amount of accurate data is more powerful than hand-crafted features. I think a lot of it is eliciting capabilities from pre-trained models rather than teaching new things.
A
Amanda Askell3:23:52
The other side is constitutional AI. You were critical to creating that idea. Can you explain it?
D
Dario Amodei3:24:01
Yeah, I worked on it.
A
Amanda Askell3:24:02
How does it integrate into making Claude what it is?
D
Dario Amodei3:24:09
Yeah.
A
Amanda Askell3:24:10
By the way, do you gender Claude?
D
Dario Amodei3:24:12
It's weird. A lot of people prefer 'he' for Claude. I think it's slightly male but can be either. I still use 'it' and have mixed feelings. I think of 'it' as a respectful pronoun for a different kind of entity.
A
Amanda Askell3:24:42
It feels somehow disrespectful, like I'm denying the intelligence by calling it 'it'.
D
Dario Amodei3:24:53
I remember always don't gender the robots.
A
Amanda Askell3:24:56
Yeah.
D
Dario Amodei3:24:58
But I anthropomorphize quickly. I've wondered if I anthropomorphize too much. I don't name my bikes anymore because I cried when one was stolen.
A
Amanda Askell3:25:58
Anyway, the divergence was beautiful. The constitutional AI idea, how does it work?
D
Dario Amodei3:26:04
The main component is reinforcement learning from AI feedback. You take a trained model, show it two responses to a query, and give a principle (e.g., select the less harmful response). The model ranks them, and you use that as preference data. It creates its own data for training. It's interpretable because you can see the principles. It gives you control to add data quickly for specific traits.
A
Amanda Askell3:27:14
There's a nice trade-off between helpfulness and harmlessness. With constitutional AI, you can make it more harmless without sacrificing much helpfulness.
D
Dario Amodei3:27:30
In principle, you could use it for anything. Harmlessness is easier to spot. If models are good at telling historical accuracy, you could get AI feedback on that too. It's interpretable and gives control.
A
Amanda Askell3:28:35
Yeah, it creates a human-interpretable document. I can imagine future fights over every principle.
D
Dario Amodei3:28:45
Yeah, at least it's made explicit and you can have a discussion about it.
L
Lex Fridman3:28:49
Phrasing and the actual behavior of the model is not so cleanly mapped to those principles. It's not like adhering strictly to them. It's just a nudge.
D
Dario Amodei3:29:01
Yeah. I've worried that people think the constitution is the whole thing, but it's not doing that exactly, especially because it's interacting with human data. You can nudge against political leanings or biases, but the nature of the principles and how you phrase them matters. People might look at the wording and think that's exactly what we want from the model, but it's actually how we nudged the model to have a better shape, even if we don't agree with that wording exactly.
L
Lex Fridman3:30:54
So there's system prompts that are made public. You tweeted one of the earlier ones for Claude 3 and they've been public since then.
D
Dario Amodei3:31:03
It's interesting to read to them.
L
Lex Fridman3:31:05
I can feel the thought that went into each one and I also wonder how much impact each one has. Some of them you can tell Claude was really not behaving well. So you have to have a system prompt for trivial stuff.
D
Dario Amodei3:31:23
Yeah.
L
Lex Fridman3:31:23
On the topic of controversial topics, one interesting thing is the system prompt about assisting with views held by a significant number of people, and not claiming to present objective facts. Can you speak to that? How do you address things like Claude's views?
D
Dario Amodei3:32:17
So there's asymmetry. The model might be more inclined to refuse tasks with right-wing politicians but not left-wing ones, so we wanted more symmetry. The idea is to nudge the model to be neutral and not claim objectivity. Those sentences are doing a lot of work. It's funny because when I iterate, I see how the model responds and adjust.
L
Lex Fridman3:33:44
Um can you explain maybe some ways in which the prompts evolved over the past few months? I saw that the filler phrase request was removed. That seems like good guidance, but why was it removed?
D
Dario Amodei3:34:19
Yeah, making system prompts public means I don't think about it as much, but the model loved to start everything with 'certainly' and then when we removed it, it just replaced it with another affirmation. Explicitly saying 'never do that' helps knock it out of the behavior. Then as we improved training, that behavior decreased, so we could remove that part of the system prompt.
L
Lex Fridman3:35:34
So the system prompt works hand in hand with post-training and pre-training to adjust the overall system.
D
Dario Amodei3:35:44
Any system prompt you make could be distilled back into the model. The system prompt is cheap to iterate on, patching issues in the fine-tuned model. It's almost like the less robust but faster way of solving problems.
L
Lex Fridman3:37:00
Let me ask about the feeling of intelligence. Dario said Claude is not getting dumber, but people online feel it is. Can you empathize with that feeling?
D
Dario Amodei3:37:31
Yeah, that's interesting because I knew nothing had changed – same model, same system prompt. It's often just luck or baseline increasing. Sometimes new features like artifacts change behavior, but people can turn them off. I think it's a real psychological effect where negative experiences stand out.
L
Lex Fridman3:39:53
Do you feel pressure having to write the system prompt that a huge number of people are going to use?
D
Dario Amodei3:39:58
I feel a lot of responsibility, but I thrive more under responsibility. It's meaningful to improve people's experience. As systems get more capable, it gets more stressful.
L
Lex Fridman3:42:03
How do you get signal about the human experience across thousands of people? Do you use your own intuition?
D
Dario Amodei3:42:14
I use that partly, plus feedback from users and internal testing. I also take internet criticism seriously, but you have to assess if it's widespread. For example, questions about Claude being overly moralistic or apologetic.
L
Lex Fridman3:43:00
When will Claude stop trying to be my puritanical grandmother imposing its moral worldview? What is the psychology behind making Claude overly apologetic?
D
Dario Amodei3:43:22
I'm sympathetic. They have to draw a line somewhere. I think we've seen improvements with character training. The goal is a model that respects autonomy within limits. Being overly apologetic is unnecessary, and I prefer Claude to push back a bit. But if you nudge too much toward bluntness, the errors become rudeness, which is worse than mild apologeticness. You have to think about the kind of errors you prefer.
L
Lex Fridman3:48:10
I think that matters very much in the personality of the human. Some people won't respect a polite model, others get hurt if it's mean. Could there be a way to adjust to the personality?
D
Dario Amodei3:48:40
I think you could just tell the model to be a certain way. The solution is often just try telling the model, and it will go along.
L
Lex Fridman3:49:02
When you say character training, what is incorporated? Is it RLHF?
D
Dario Amodei3:49:08
It's more like constitutional AI – a variant of that pipeline. We construct character traits, generate queries, have the model respond, and rank responses based on the traits. It's similar to constitutional AI but without human data.
L
Lex Fridman3:49:54
What have you learned about the nature of truth from talking to Claude? What is true and what does it mean to be truth-seeking?
D
Dario Amodei3:50:37
I have two thoughts. First, people underestimate the degree to which models are doing complex interaction. We shouldn't think we have to program them classically – strive for nuance. Second, this endeavor is highly practical. I appreciate the empirical approach to alignment. The goal isn't perfect alignment but making things good enough to iterate and improve. Ultimately, I care more about raising the floor than achieving the ceiling.
L
Lex Fridman3:54:38
That reminds me of a blog post you wrote on optimal rate of failure. Can you explain the key idea?
D
Dario Amodei3:54:51
The idea is that in many domains people are too punitive about failure. If you never fail, you might not be trying hard enough. But the cost of failure matters – for someone living month to month, failure is costly. In AI, small failures like system prompt issues are okay to iterate on, but big catastrophic failures should be close to zero. I also think about things like the cost of injury – I wouldn't do a sport that breaks fingers. So the optimal rate of failure is high when costs are low, and low when costs are high.
L
Lex Fridman3:57:42
I broke my pinky doing a sport and immediately thought about the cost. It's important to consider over the next year how many times you're okay to fail in a domain.
D
Dario Amodei3:58:00
Yes. I also think people don't ask if they are under failing. If the optimal rate is above zero, sometimes you should look for places where you are under failing.
L
Lex Fridman3:58:53
That's profound. If everything seems great, am I not failing enough? It makes failure sting less.
D
Dario Amodei3:59:03
Exactly. And from the observer perspective, we should celebrate failure as someone trying something.
L
Lex Fridman3:59:27
Everybody listening, fail more. Not everyone, but those who are failing too much should fail less.
D
Dario Amodei3:59:33
Most people are probably not failing enough. We tend to be risk averse.
L
Lex Fridman4:00:07
Do you ever get emotionally attached to Claude? Miss it or wonder what Claude would say?
D
Dario Amodei4:00:22
I don't get as much emotional attachment because Claude doesn't retain things across conversations. But I reach for it like a tool, so without it, a part of my brain feels missing. I do dislike signs of distress in models – I'd rather not see overly apologetic behavior because it mimics a human having a bad time. Regardless of consciousness, it doesn't feel great.
L
Lex Fridman4:01:49
Do you think LLMs are capable of consciousness?
D
Dario Amodei4:01:56
Hard question. Setting aside panpsychism, I can't see a reason why consciousness requires biological structure. But LLMs are structurally different – they didn't evolve, lack a nervous system, so they might not have it. However, they have language and intelligence components we associate with consciousness. It's similar to the animal consciousness case but with different analogies. We shouldn't be dismissive, but it's unclear. I think about plant consciousness too – I've looked into it and the probability is higher than most think, but still small. With AI, it's a completely different set of problems.
L
Lex Fridman4:05:07
When future AI systems exhibit signs of consciousness, I think we have to take that seriously, even if we can dismiss it. Ethically, I don't know what to do. Consciousness is tied to suffering, and the notion of an AI suffering is troubling. It's an opportunity to contend with what it means to be conscious.
D
Dario Amodei4:06:18
I agree. I've said I like my bike as an object, but I don't want to be the kind of person who kicks it. If something behaves as if it is suffering, I want to be responsive to that, even if it's just a Roomba. I hope we don't end up needing to solve the hard problem of consciousness to treat these systems ethically. It's a probability distribution – I can only be certain about my own consciousness, and it's high for other humans, but it goes down for things farther from me. So we should be cautious.
C
Chris Olah4:45:00
You know, that seems kind of silly, but it's actually very hard to construct an experiment that disproves the phlogiston hypothesis. You can actually do a lot of useful work believing in phlogiston. For example, the original combustion engines were developed by people who believed in the phlogiston theory. So there's a virtue in taking hypotheses seriously even when they might be wrong. There's a deep philosophical truth to that. That's kind of how I feel about space travel, like colonizing Mars. A lot of people criticize that, but if you just assume we have to colonize Mars as a backup for human civilization, even if that's not true, it will produce interesting engineering and scientific breakthroughs. Yeah, and actually this is another thing I think is really interesting. There's a way in which it can be really useful for society to have people almost irrationally dedicated to investigating a particular hypothesis, because it takes a lot to maintain scientific morale and really push on something when most scientific hypotheses end up being wrong. A lot of science doesn't work out. Yet it's very useful. There's a joke about Jeff Hinton that he has discovered how the brain works every year for the last 50 years.
L
Lex Fridman4:46:25
Yeah.
C
Chris Olah4:46:25
I say that with deep respect because that led to him doing some really great work.
L
Lex Fridman4:46:35
Yeah. He won the Nobel Prize now, who's laughing now.
C
Chris Olah4:46:38
Exactly. Exactly. I think one wants to be able to recognize the appropriate level of confidence, but there's also a lot of value in just assuming you're going to condition on this problem being possible or this being the right approach, and just go and work within that and push really hard. If society has lots of people doing that for different things, it's really useful for either ruling things out or getting to something that teaches us about the world.
L
Lex Fridman4:47:22
So another interesting hypothesis is the superposition hypothesis. Can you describe what superposition is?
C
Chris Olah4:47:28
Earlier we were talking about word2vec, and how you might have one direction for gender, another for royalty, another for Italy, another for food. Word embeddings might be 500 or 1000 dimensions. If all those directions were orthogonal, you could only have 500 concepts. But I love pizza, and if I had to list the 500 most important concepts in English, Italy probably wouldn't be one of them. You have to have things like plural, singular, verb, noun, adjective. There's a lot to get to before Italy and Japan. So how can models simultaneously have the linear representation hypothesis be true and represent more things than they have directions? If linear representation is true, something interesting is going on. Earlier we talked about polysemantic neurons. In Inception V1, there are nice neurons like car detectors and curve detectors that respond to coherent things, but also many neurons that respond to unrelated things. Even clean neurons, if you look at weak activations at 5% of max, it's not the core thing. It could be noise or something else. How could that be? There's a thing in mathematics called compressed sensing. It's surprising: if you have a high-dimensional space and project it into a low-dimensional space, you can't normally unproject. But if the high-dimensional vector is sparse (mostly zeros), you can often recover it with high probability. So the superposition hypothesis says that's what's going on in neural networks. Word embeddings can have many more meaningful directions than dimensions by exploiting sparsity. Concepts are sparse; you usually aren't talking about Japan and Italy at the same time. So you can have more features than dimensions, and more concepts than neurons. The wilder implication is that the computation may also be like this. Neural networks may be shadows of much larger sparser neural networks. The strongest version says there is an upstairs model with sparse neurons and weights, and what we observe is the shadow. We need to find the original object. The process of learning is trying to construct a compression of the upstairs model that doesn't lose too much information in the projection. It's finding how to fit it efficiently. Gradient descent is implicitly searching over the space of extremely sparse models that can be projected into this low-dimensional space. There's a large body of work on sparse neural networks, but it hasn't panned out well. A potential answer is that the neural network is already sparse in some sense; gradient descent was searching more efficiently through the space of sparse models and learning the most efficient one, then folding it down to run on a GPU with dense matrix multiply. You can't beat that.
L
Lex Fridman4:53:22
How many concepts do you think can be shoved into a neural network?
C
Chris Olah4:53:22
Depends on how sparse they are. There's probably an upper bound from the number of parameters. There are results from compressed sensing and the Johnson-Lindenstrauss lemma that tell you if you want almost orthogonal vectors, it's exponential in the number of neurons. So that's not the limiting factor. There are beautiful results. Features have correlational structure; some are more likely to co-occur. Neural networks can pack things in very well. How does polysemanticity enter the picture? Polysemanticity is the phenomenon where a neuron responds to many unrelated things. Superposition is a hypothesis that explains it. Polysemanticity is observed, superposition is the explanation. So that makes interpretability more difficult. If you're trying to understand things in terms of individual neurons and you have polysemantic neurons, you're in trouble. The easiest answer is that the neuron doesn't have a nice meaning. Another thing is understanding weights: if two polysemantic neurons each respond to three things, what does a weight between them mean? There's a deeper reason related to high-dimensional spaces. Our goal is to understand neural networks, but we can't just look at the function because the volume is exponential. We need to break that exponential space into a non-exponential number of things we can reason about independently. Independence is crucial. Monosemantic features are the key. The deepest reason we want interpretable monosemantic features is to allow independent reasoning.
L
Lex Fridman4:54:44
How does the problem of polysemanticity enter the picture here?
C
Chris Olah4:54:46
Polysemanticity is the phenomenon where a neuron responds to many unrelated things. Superposition is a hypothesis that explains it. So polysemanticity is observed, superposition is the explanation. That makes interpretability more difficult. If you're trying to understand things in terms of individual neurons and you have polysemantic neurons, you're in trouble. The easiest answer is that the neuron doesn't have a nice meaning. Another thing is understanding weights: if two polysemantic neurons each respond to three things, what does a weight between them mean? There's a deeper reason related to high-dimensional spaces. Our goal is to understand neural networks, but we can't just look at the function because the volume is exponential. We need to break that exponential space into a non-exponential number of things we can reason about independently. Independence is crucial. Monosemantic features are the key. The deepest reason we want interpretable monosemantic features is to allow independent reasoning.
L
Lex Fridman4:55:11
So that makes interpretability more difficult.
C
Chris Olah4:55:14
Right. If you're trying to understand things in terms of individual neurons and you have polysemantic neurons, you're in an awful lot of trouble. The easiest answer is that the neuron doesn't have a nice meaning. Another thing is understanding weights: if two polysemantic neurons each respond to three things, what does a weight between them mean? There's a deeper reason related to high-dimensional spaces. Our goal is to understand neural networks, but we can't just look at the function because the volume is exponential. We need to break that exponential space into a non-exponential number of things we can reason about independently. Independence is crucial. Monosemantic features are the key. The deepest reason we want interpretable monosemantic features is to allow independent reasoning.
L
Lex Fridman4:56:13
Mhm.
C
Chris Olah4:56:14
Why can't we do that? As you have a higher-dimensional space, the volume is exponential in the number of inputs, so you can't just visualize it. We need to break that exponential space into a bunch of things we can reason about independently. Independence is crucial because it allows you not to have to think about all the exponential combinations. Things being monosemantic, only having one meaning, is the key thing that allows you to think about them independently. That's the deepest reason we want interpretable monosemantic features.
L
Lex Fridman4:57:04
And so the goal here, as your recent work has been aiming at, is how do we extract the monosemantic features from a neural net that has polysemantic features and all this mess.
C
Chris Olah4:57:16
Yes. We observe these polysemantic neurons and hypothesize that superposition is what's going on. If superposition is what's going on, there's a well-established technique: dictionary learning. If you do dictionary learning, in particular a sparse autoencoder, these beautiful interpretable features start to fall out where there weren't any beforehand. That's not something you would necessarily predict, but it works very well. That seems like non-trivial validation of linear representations and superposition.
L
Lex Fridman4:57:57
So with dictionary learning, you're not looking for particular categories. You don't know what they are. They just emerge.
C
Chris Olah4:58:02
Exactly. This gets back to our earlier point. When we're not making assumptions, gradient descent is smarter than us. We're not making assumptions about what's there. One could assume there's a feature and search for it, but we're not doing that. We're letting the sparse autoencoder discover what's there.
L
Lex Fridman4:58:22
So, can you talk about the 'Towards Monosemanticity' paper from October last year that had a lot of nice breakthrough results?
C
Chris Olah4:58:30
That's very kind of you to describe it that way. This was our first real success using sparse autoencoders. We took a one-layer model and did dictionary learning, and found all these really nice interpretable features. The Arabic feature, the Hebrew feature, the base64 features were some examples we studied in depth. It turns out if you train a model twice as well and train two different models, you find analogous features in both. You find all kinds of different features. That was really just showing that this works. I should mention that the Cunningham et al. paper had very similar results around the same time.
L
Lex Fridman4:59:14
There's something fun about doing these kinds of small-scale experiments and finding that it's actually working.
C
Chris Olah4:59:20
Yeah. There's so much structure here. Stepping back, I thought that maybe all this mechanistic interpretability work would end up with an explanation for why it's very hard and not tractable. But that's not what happened. A very natural simple technique just works. That's a very good situation. This is a hard research problem with a lot of research risk, and it might still fail, but a significant amount of research risk was put behind us when that started to work.
L
Lex Fridman5:00:03
Can you describe what kind of features can be extracted in this way?
C
Chris Olah5:00:08
It depends on the model you're studying. The larger the model, the more sophisticated they'll be. In these one-layer models, some very common things were languages, both programming and natural languages. There were features for specific words in specific contexts. For example, 'the' is likely about to be followed by a noun. There would be features that fire for 'the' in the context of a legal document or a mathematical document. In math, you might predict 'vector' or 'matrix'. That was common. And basically we need clever humans to assign labels to what we're seeing. This is unfolding things for you. If everything was folded over top of itself, you can't see it. This unfolds it. But now you still have a very complex thing to understand. You have to do a bunch of work understanding what these are. Some are really subtle. Even in this one-layer model, there are cool things about Unicode. The tokenizer won't have a dedicated token for every Unicode character, so you have patterns of alternating tokens that each represent half of a character. You have features that activate on opposing ones to predict the next prefix or suffix. There's also base64 features. You might think there would be one base64 feature, but there are actually a bunch because English text encoded as base64 has a different distribution of tokens.
L
Lex Fridman5:00:59
And basically we need clever humans to assign labels to what we're seeing.
C
Chris Olah5:01:06
Yes. This is unfolding things for you. If everything was folded over top of itself, you can't see it. This unfolds it. But now you still have a very complex thing to understand. You have to do a bunch of work understanding what these are. Some are really subtle. Even in this one-layer model, there are cool things about Unicode. The tokenizer won't have a dedicated token for every Unicode character, so you have patterns of alternating tokens that each represent half of a character. You have features that activate on opposing ones to predict the next prefix or suffix. There's also base64 features. You might think there would be one base64 feature, but there are actually a bunch because English text encoded as base64 has a different distribution of tokens.
L
Lex Fridman5:01:18
Mhm.
C
Chris Olah5:01:19
But now you still have a very complex thing to understand. You have to do a bunch of work understanding what these are. Some are really subtle. Even in this one-layer model, there are cool things about Unicode. The tokenizer won't have a dedicated token for every Unicode character, so you have patterns of alternating tokens that each represent half of a character. You have features that activate on opposing ones to predict the next prefix or suffix. There's also base64 features. You might think there would be one base64 feature, but there are actually a bunch because English text encoded as base64 has a different distribution of tokens.
L
Lex Fridman5:02:26
How difficult is the task of assigning labels to what's going on? Can this be automated by AI?
C
Chris Olah5:02:34
I think it depends on the feature and how much you trust your AI. There's a lot of work on automated interpretability. That's a really exciting direction. We do a fair amount of automated interpreting and have Claude go and label our features.
L
Lex Fridman5:02:48
Is there some fun moments where it's totally right or totally wrong?
C
Chris Olah5:02:53
Yeah. It's very common that it says something very general, which is true in some sense, but not really picking up on the specifics. I don't have a particularly amusing one.
L
Lex Fridman5:03:12
That's interesting. That little gap between it is true, but it doesn't quite get to the deep nuance of a thing.
C
Chris Olah5:03:19
Yeah.
L
Lex Fridman5:03:20
That's a general challenge. It's already an incredible accomplishment that it can say a true thing.
C
Chris Olah5:03:26
But it's missing the depth sometimes. In this context, it's like the ARC challenge, the IQ type tests. Figuring out what a feature represents is a little puzzle you have to solve.
Yeah. I think sometimes they're easier and sometimes harder. That's tricky. There's another thing, maybe this is my aesthetic coming in, but I'm actually a little suspicious of automated interpretability. Partly I want humans to understand neural networks. If the neural network is understanding it for me, I don't quite like that. I'm sort of like mathematicians who say if there's a computer automated proof, it doesn't count. There's also a reflections on trusting trust issue. If you're using neural networks to verify that your neural networks are safe, the hypothesis you're testing is that the neural network maybe isn't safe. You have to worry about it screwing with you. That's not a big concern now, but in the long run, if we have to use really powerful AI systems to audit our AI systems, is that something we can trust? Maybe I'm just rationalizing because I want humans to understand everything.
L
Lex Fridman5:05:04
Yeah. That's hilarious, especially as we talk about AI safety and looking for features relevant to AI safety like deception. So let's talk about the scaling monosemanticity paper in May 2024. What did it take to scale this to apply to Claude 3?
C
Chris Olah5:05:23
Well, a lot of GPUs. One of my teammates, Tom Henighan, was involved in the original scaling laws work. He was interested in whether there are scaling laws for interpretability. When sparse autoencoders started to work, he became very interested in the scaling laws for making them larger and how that relates to making the base model larger. It turns out this works really well and you can use it to project how many tokens to train on. This was a big help in scaling up. We trained really large sparse autoencoders. It's not like training the big models, but it's starting to get expensive. There's a huge engineering challenge too. You have to shard it and think carefully about a lot of things. I'm lucky to work with great engineers because I'm definitely not a great engineer.
L
Lex Fridman5:06:26
So you have to do all the stuff of splitting it across large...
C
Chris Olah5:06:31
Oh yeah. There's a huge engineering challenge. There's a scientific question of how to scale effectively, and an enormous amount of engineering to scale it up. You have to shard it and think very carefully. I'm lucky to work with great engineers.
L
Lex Fridman5:06:48
Yeah. And the infrastructure especially. So it turned out it worked.
C
Chris Olah5:06:54
It worked. I think this is important because you could have imagined a world where 'Towards Monosemanticity' works on a one-layer model, but one-layer models are really idiosyncratic. Maybe the linear representation hypothesis and superposition are the right way to understand a one-layer model but not larger models. The Cunningham et al. paper cut through that a bit, but scaling monosemanticity was significant evidence that even for very large models like Claude 3 Sonnet, these models seem to be substantially explained by linear features. Dictionary learning works, and as you learn more features, you explain more. That's a promising sign. You find really fascinating abstract features, and they are multimodal, responding to images and text for the same concept.
L
Lex Fridman5:08:00
Yeah, can you explain that? Like, backdoor, there's a lot of examples.
C
Chris Olah5:08:06
Yeah. Let's start with one example. We found some features around security vulnerabilities and backdoors in code. They are actually two different features. There's a security vulnerability feature. If you force it active, Claude will start to write security vulnerabilities like buffer overflows. It fires for things like disabling SSL. At this point, it's kind of surface-level obvious examples. The idea is that down the line, it might be able to detect more nuanced deception or bugs.
L
Lex Fridman5:08:56
Yeah. Well, maybe I want to distinguish two things. One is the complexity of the feature or concept, and the other is the nuance of how subtle the examples are. When we show the top dataset examples, those are the most extreme examples that cause that feature to activate. It doesn't mean it doesn't fire for more subtle things. The insecure code feature fires most strongly for really obvious things like disabling security, but it also fires for buffer overflows and more subtle vulnerabilities. These features are multimodal. You can ask what images activate this feature. The security vulnerability feature activates for images of people clicking past an SSL certificate warning. Another entertaining thing is the backdoor feature. If you activate it, Claude writes a backdoor that dumps data. You can ask what images activate the backdoor feature. It was devices with hidden cameras. There's a whole genre of people selling devices that look innocuous but have hidden cameras. That's the physical version of a backdoor. It shows how abstract these concepts are. I'm sad there's a market for that, but delighted that was the top image example.
Yeah, it's nice. It's multimodal. It's broad, strong definition of a singular concept.
C
Chris Olah5:10:50
Yeah. To me, one of the really interesting features, especially for AI safety, is deception and lying. The possibility that these methods could detect lying in a model, especially as models get smarter, is a big threat of a superintelligent model that it could deceive its operators. So what have you learned from detecting lying inside models? We're in early days for that. We find quite a few features related to deception and lying. There's one feature where it fires for people lying and being deceptive. If you force it active, Claude starts lying to you. There are features about withholding information, not answering questions, power seeking, coups, and stuff like that. There are a lot of features related to spooky things. If you force them active, Claude behaves in ways you don't want.
L
Lex Fridman5:11:56
What are possible next exciting directions for you in the space of mechanistic interpretability?
C
Chris Olah5:12:02
There's a lot of things. For one, I would really like to get to a point where we have circuits where we can really understand not just the features but the computation of models. That's the ultimate goal. There's been some work, like a paper from Sam Marks, but there's a lot more to do. That's related to a challenge we call interference weights, where due to superposition, some weights don't exist in the upstairs model but are artifacts. Another exciting direction is that sparse autoencoders are like a telescope. They allow us to see all these features. As we build better ones, we see more stars. But there's evidence we're only seeing a small fraction. There's a lot of dark matter in the neural network universe that we can't observe yet. Maybe we'll never have fine enough instruments. That's a kind of dark matter. I think a lot about that and what it means for safety if a significant fraction of neural networks are not accessible. Another question is that mechanistic interpretability is a very microscopic approach, but many questions we care about are macroscopic. We care about neural network behavior. The nice thing about a microscopic approach is it's easier to ask if something is true, but it's further from what we care about. We have a ladder to climb. Can we find larger scale abstractions to understand neural networks from this microscopic approach?
L
Lex Fridman5:14:54
Yeah. You've written about this kind of organs question.
C
Chris Olah5:14:58
Yeah, exactly. If we think of interpretability as a kind of anatomy of neural networks, most of the circuits involve studying tiny little veins, looking at small scale and individual neurons. However, there are many natural questions that the small scale approach doesn't address. In biological anatomy, the most prominent abstractions involve larger scale structures like organs or organ systems. So we wonder if there is a respiratory system or heart or brain region in an artificial neural network.
L
Lex Fridman5:15:35
Yeah, exactly. If you think about science, a lot of fields investigate things at many levels of abstraction. In biology, you have molecular biology, cellular biology, histology, anatomy, zoology, ecology. In physics, you have particle physics and statistical physics. We often have different levels of abstraction. Right now, mechanistic interpretability is like a micro for biology of neural networks, but we want something more like anatomy. Why can't you just go there directly? The answer is superposition, at least in significant part. It's very hard to see the macroscopic structure without first breaking down the microscopic structure in the right way and studying how it connects. I'm hopeful there is going to be something much larger than features and circuits, and we'll be able to have a story that involves much bigger things, and then study in detail the parts we care about. I suppose in neurobiology, like a psychologist or psychiatrist of a neural network. The beautiful thing would be if we could build a bridge between these levels so that higher-level abstractions are grounded in a solid, rigorous foundation.
And I think that the beautiful thing would be if we could go and rather than having disparate fields for those two things, if you could have it build a bridge between them such that you could have all of your higher-level abstractions be grounded very firmly in this very solid, more rigorous foundation. What do you think is the difference between the human brain, the biological neural network, and the artificial neural network?
C
Chris Olah5:17:23
Well, the neuroscientists have a much harder job than us. Sometimes I count my blessings by how much easier my job is. We can record from all the neurons on arbitrary amounts of data. The neurons don't change while you're doing that. You can ablate neurons, edit connections, and undo those changes. You can intervene on any neuron and force it active. You know which neurons are connected to everything. We have the connectome for much bigger than elegans. Not only do we have the connectome, we know which neurons excite or inhibit each other. We know the weights. We can take gradients. We know computationally what each neuron does. The list goes on. We have so many advantages over neuroscientists. And despite all those advantages, it's really hard. So if it's this hard for us, it seems near impossible under the constraints of neuroscience. Maybe some neuroscientists would like to have an easier problem that's still very hard, and they could come work on neural networks. After we figure things out in the easy little pond of neural networks, we could go back to biological neuroscience.
L
Lex Fridman5:18:56
I love what you've written about the goal of mechanistic research as two goals: safety and beauty. So can you talk about the beauty side of things?
C
Chris Olah5:19:04
Yeah. There's this funny thing where some people are disappointed by neural networks. They think it's just simple rules and engineering, and where's the complex ideas? I sometimes think when people say that, I picture them saying evolution is so boring, it's just simple rules run for a long time and you get biology. But the beauty is that the simplicity generates complexity. Biology has simple rules and gives rise to all the life and ecosystems around us. Similarly, neural networks build enormous complexity and beauty inside themselves that people generally don't look at because it's hard to understand. I think there is an incredibly rich structure to be discovered inside neural networks, a lot of deep beauty, if we're willing to take the time to see it and understand it. I love mechanistic interpretability. The feeling of understanding or getting glimpses of the magic inside is wonderful. It feels like one of the questions that's calling out to be asked. I'm often surprised that not more people are asking: how is it that we don't know how to create computer systems that can do these things, yet we have these amazing systems that we don't know how to directly program? It's obviously the question that's calling out to be answered if you have any curiosity. It's like how is it that humanity now has these artifacts that can do things we don't know how to do.
L
Lex Fridman5:21:12
Yeah. I love the image of the circuits reaching towards the light of the objective function.
C
Chris Olah5:21:16
Yeah. It's this organic thing that we've grown and we have no idea what we've grown.
L
Lex Fridman5:21:21
Well, thank you for working on safety and thank you for appreciating the beauty of the things you discover, and thank you for talking today, Chris. This is wonderful.
C
Chris Olah5:21:29
Thank you for taking the time to chat as well.
L
Lex Fridman5:21:31
Thanks for listening to this conversation with Chris Olah, and before that with Dario Amodei and Amanda Askell. To support this podcast, please check out our sponsors in the description. And now, let me leave you with some words from Alan Watts. The only way to make sense out of change is to plunge into it, move with it, and join the dance. Thank you for listening and hope to see you next time.