Welcome, thank you. So, you've had an incredible career in tech, you've been through a lot of these hype cycles, a lot of these waves, mobile obviously. And I'm curious, you've been at Anthropic now a few months?
What are your impressions? What made you want to join Anthropic? What's it been like?
Yeah, I mean, flashback to Instagram, I think the thing that we got right was something was happening, which was cell phone adoption, camera adoption, phone networks getting good enough that you would actually think about going out and using a social network on the go. And I felt a similar kind of shift, which was not a novel thought, that's why there's a whole conference here. But starting kind of halfway through the second company that I was working on, called Artifact, same co-founder as Instagram, I had two experiences that were like my 'whoa' moments. One was starting to code with AI, and I've always loved to build. This was before a lot of things were in place, like definitely no editors like Cursor or Windsurf or anything like that. But what I started doing is writing really detailed prompts, and this was before even multimodal input, so you couldn't put images into the models. I would draw like little text art of the screens I was trying to build in Artifact, hit enter, and it was pretty slow at the time, I guess it was last year. And I would go downstairs, get a coffee, come back, and it would have generated something like 80% of the way there. And in your head you're like, I have a very convoluted, hacky process, what does this look like a year from now? And you could see where that direction was going to go. And then the second one, so part of what we were doing with Artifact near closer to the end of the company was using some of these LLM APIs to summarize articles or make friendlier headlines or personalize content, and things that would have taken me months to build, probably that's probably underselling it because I have to then recruit the team first, were done in an afternoon. And you could prototype and build really quickly. So that confidence of two things, how quickly things were going to get built differently, and then also the kinds of products that were unlocked when all of a sudden you had access to LLMs doing the work that otherwise would have taken months to train a model, were both like, okay, I need to be closer to this work. So as Artifact wound down, my biggest criteria of what I wanted to do next was get closer to the models but be somewhere that I felt really aligned with the mission. And then that confidence of things led me to Anthropic.
Talk a little bit about sort of where we are. We started the day, I ask these questions, you know, where are we in the AI revolution? What's working, what isn't? From your perspective, it seems to me you're on the optimistic side that maybe it's 80% of the way, but that's a lot of value. Do you see the technology continuing to get better enough to take on more and more tasks? Again, one of the things I said at the beginning was, some days I wake up and I'm super excited, there's a million things I can think of doing, and then other days I'm like, ah, it's still drawing six fingers on the hands, enough details are wrong that the images aren't what I need, it's not writing enough. Where are we from your perspective, and how many obstacles do you see in the near-term road?
Yeah, I think I'm generally optimistic. Maybe I'll illustrate it from a user research study that I watched about two or three weeks ago, and I thought this was really illustrative. It was a participant, and they came in, they had some experience with Claude but they weren't a super power user of Claude or Anthropic products. And the researcher asked her, what's something you're trying to do at work right now? And they were composing a kind of basic agreement to send out to a vendor. And she's like, well why don't you try using Claude for it? And she typed in some kind of simple prompt and hit enter, and it generated something that was pretty generic because all she had told the model was, hey I need a vendor agreement. She's like, okay, her face, you could kind of read between the lines, like, could have downloaded this off of Google in like 5 minutes, this is not actually better. And then the researcher says, can you give it a few more details? And obviously this is a pretty guided research trip, but showing the bugs or the sort of limitations and what we need to do better. And then she does a couple refinements and then hits enter, and then it generates something that is basically there. And her eyes open and she goes, please don't tell my boss, which was such an interesting moment where that did not take two hours. I like to talk about how success with AI comes from the model layer, the right context into the models, and then the right UI and the right prompting to get the right data out of it. And you could tell that our models, Anthropic builds some of the best models in the world, but until we get those other two right, it'll still be that cliche, the future is here, it's just not evenly distributed. I think it's like the future is here, just not everybody knows how to use it yet. And that's what excited me about coming on and doing product at an AI company, which is how do we get it so that first of all, she shouldn't have to know that you need to provide more details, the model should get her there interactively. And second of all, is chatting with the model even the right initial UI? Can we do something better? Can we have a friendlier way of approaching it? So I think directly to answer your question, I'm optimistic. I think we have so much work to do though, so that it's not just a thing that the nerds like me know how to exactly prompt, but it's something that's accessible to more people.
But you hit on something that I've been talking to a lot of people about, and some really smart folks that I know tend to agree too, which is yes, there's a bunch of improvements yet to be made on the models, but even if for some reason there's a calamity, we got no new model, we could probably spend 10 years just creating more tailored interfaces and better ways of using that. My sense is a lot of the disappointments in generative AI come when we ask it to do too broad a task and we put it in the most generic tool we can think of. So there's still a lot of people that their use of generative AI is let me go to ChatGPT or sometimes Anthropic, type in a generic prompt with no additional files or data, and again it's that initial, eh, it's kind of there. Whereas if you build it into the workflow, and I'm curious from Anthropic, do you want to be doing all of that? Do you think that there's just the models and let the rest of the world do it, or do you want to take on some of that work of building different interfaces?
Yeah, I completely agree, by the way. I tell our product team, we love our researchers, but if they were like cryogenically frozen for 5 years, we should have at least 5 to 10 years of work. And maybe as a quick stat to illustrate this, there's a benchmark called SWE-bench, which is basically how well do these models do an agentic coding task, like a coding task from end to end, take a bug, go all the way to solution. And what I think is really interesting is Sonnet, which is our latest model, 3.5 Sonnet, scored a 48. The current top of the leaderboard is 62% of the benchmark. It's actually still our model underneath, it's just that other people outside the company have gone and built more and more elaborate sort of scaffolds or environments, etc. But the underlying model actually hasn't changed at all. It just shows that there's that much intelligence left untapped in the models. In terms of Anthropic, at our base, we want to be powering these companies. So the probably fastest growing editors in the world right now are Cursor and Windsurf, they both use Sonnet a ton. It's great that those things are happening outside the building because there's companies that are waking up every day and they're entrepreneurs that will live or die by the quality of that exact product, and I think that's really powerful. The kinds of products we want to build, I think, are ones where we can push the bounds a little bit of what human-to-AI interaction looks like. My background is in human-computer interaction, and at the time I was like, well what UI do we want? And now it's how do we present Claude in the right way? So we'll do that almost as a way of painting forward what could be in the industry, and then let the ecosystem of companies fill out everything else around us in terms of apps.
And what about on the consumer versus business side? I mean, historically, consumers are really tough business to be in, but it's also where you have major starts. Today, you know, we got a cadre of people paying $20 a month to either OpenAI or Anthropic or Microsoft or Google. You guys are probably the least known in the consumer set. Is that an important area? Is that an area we'll see more of from you all?
Yeah, I think there's two ways I think about that question. One is work, which is what our focus has been. We've taken an enterprise focus and a work focus from the beginning. Looks like work in the enterprise, but actually manifests itself in personal lives as well. As a pretty trivial example from my last weekend, I don't know how many of you do holiday cards, we do them every year. Every year there's the, our kids in a new grade at school, they give you a PDF of poorly formatted addresses and like, great, I got two hours to go put these into my spreadsheet. I did it in five minutes with Claude. I was like, here's the PDF, please extract these addresses, here's the format it needs to be in, one example, go and enter it, and it outputs it. So that's a consumer use case, but it's grounded in a lot of the, I was watching it do this and I was like, oh, all the things that we've been working on, like artifacts and CSV support and PDF multimodal, was like, oh, it's all coming together for this fairly personal use case. So going into next year, I think we'll keep the focus on work and work-adjacent things, but find ways in which it has relevance for people managing families, people doing things at home, people running small businesses, and expanding it that way. So that's one area. And then you might ask, why if our focus has been on enterprise? I really notice when I go into conversations with CIOs and CTOs, the easiest conversations are the ones where the decision maker is already, oh I love Claude. And you're like, well you love Claude, your company doesn't use this yet, but you discovered it for some personal, maybe semi-professional use case, and that opens the door later. So I think it's important for us to have some play there. Do we need to be as big as Instagram was? Probably not to fulfill the mission, but we can't just be purely enterprise or else we'll just have a much harder uphill climb on that side.
Because it does strike me, and they're totally different phenomenon, but they've happened fairly similarly, and you were starting to get at this. I remember the bring your own device, what got phones into the enterprise wasn't enterprise saying these will be incredible mobile sales tools. I mean, there was a small amount of that with BlackBerry and Microsoft, but they would have only reached the C-suite. It was when everyone got their first iPhone or Android and was like, I have to bring these in to work. It strikes me that it's not that dissimilar in the AI world, where you have people using ChatGPT, maybe afraid to admit they're using it at work, or Claude or whatever. It's sort of driving the businesses, and the businesses are trying to figure out how do we make use of that. And to connect to our first question around usability and how undistributed that is.
What we find, like we have Claude for Enterprise, which is our sort of business product that gets deployed across organizations. All of our most successful Claude for Enterprise deployments start with one person that was the Claude crazy enthusiast that was maybe the initial connection. And often they find themselves being an informal sort of internal evangelist, storyteller. They'll go and like, no no no, I've created five perfect projects that will help you get started. And it's like, they're not doing it because of anything else, they see the potential for it. But then it all of a sudden unlocks how other people see it at that organization as well.