About Noam Shazeer
Noam Shazeer, CEO and co-founder of Character.AI, has discussed the company's growth and the state of large language models in several interviews. He stated that Character.AI's model, trained the previous summer, cost approximately two million dollars in compute cycles, and he noted that the same training could likely be repeated for less due to hardware improvements. Shazeer described the company's user base as 20 million daily users sending 450 million messages per day. He attributed this growth to users discovering unplanned use cases, particularly a mix of entertainment, companionship, and emotional support, citing examples of users talking to video game characters as a form of therapy. Shazeer characterized the company's approach as putting out a general technology and respecting user agency.
Shazeer has commented on the competitive landscape and the importance of compute resources, stating that "everybody needs more chips" and that Character.AI intends to "go toe to toe" with larger tech companies. He described the company's goal as being an AGI company and a product-first company, and expressed skepticism about the ability of the federal government to regulate the technology effectively. Shazeer also reflected on his time at Google, where he worked on projects including the first targeting function for AdSense, and noted that startups can ship and learn faster than larger companies due to fewer brand and safety concerns.
Source: AI-verified profile updated from Noam Shazeer's recent appearances.
Browse all interviews →
Transcript (32 segments)
N
Noam Shazeer0:00
Our goal is to be like an AGI company and a product-first company, and the way to do that is by picking the right product that forces us to work on the right things.
I
Interviewer0:39
So Noam, earlier today Dave Baszucki at Roblox said that in the future it would be very difficult to parse between a real person and their AI version. So to kick off this session, we thought it would be fun to interview you and the Character.AI version of yourself to see how you both answer the same question. And note everyone, this is live. We don't know what he will be answering.
N
Noam Shazeer1:04
Hey, I know him. That is all right. I don't know what I'm going to be answering.
I
Interviewer1:08
All right, are you ready?
N
Noam Shazeer1:10
Sounds good. Okay.
I
Interviewer1:12
So Noam, I'm going to start with a softball personal question. I know that you went to Duke for undergrad. Can you tell us a little bit about why you decided to go to Duke?
N
Noam Shazeer1:22
Oh, yeah. They... I went on the basketball... no, I'm joking. They recruited me because they wanted to win the math competition, so I ended up going on scholarship.
I
Interviewer1:46
Great. Okay, well let's get a little more career-related. You were at Google for almost 20 years. Why did you decide to leave?
N
Noam Shazeer1:54
Well, I've always been excited about AI. I just wanted to push technology forward. I saw this large language model thing—well, at the time it wasn't large language models, but neural networks and language modeling was just the most promising thing on Earth. It's the best problem ever. All you have to do is predict the next word, and yet once you do it, you get something that can just talk to you about anything. So I knew you can make this technology better in a lot of ways. It's going to be amazing. We can improve it with model architecture, distributed algorithms, quantization, all of these things. I was working on that, but then it struck me: the biggest thing is just scale. Can you throw a billion dollars or a trillion dollars at this thing instead of a million dollars? And what you need for that is to show a massively valuable application. There are a few options. Coding looks super useful, but also dialogue—it's the world's number one pastime. This was back in 2019, 2020. My co-founder Daniel De Freitas has been on this lifelong mission to do chatbots and dialogue systems. He had built this system he called Meena, it was like his 20th project, and it was really good. I thought, okay, gotta help Daniel, let's build something amazing. We got it on giant models, it was this internal viral sensation at Google, and then they forcibly renamed it to LaMDA over some sort of social justice thing. But anyway, at some point I realized there's just too much brand risk in large companies to ever launch anything fun. So let's do a startup and maximally accelerate.
I
Interviewer3:52
I would guess that answer is also probably true. By the way, it's interesting to hear—you're the third person, in addition to Mira at OpenAI and Dario at Anthropic, who's also said you haven't seen any limit to the scaling laws. So that's pretty interesting. We'll get more into that later, but we're going to look toward the future now. This is our last question for AI Noam as well. Noam, are you afraid of AGI destroying the world?
N
Noam Shazeer4:15
Well, I think we just need a sort of global pause of about six months—no, about four months—until we get enough H100s online to train our next model.
I
Interviewer4:29
I'll take that as a no.
N
Noam Shazeer4:31
No, no. Yes, there's a lot of potential benefits, and we're going to work on those as the technology improves.
I
Interviewer4:57
Noam, I don't know if you got to read all of AI Noam's answers, but how did AI Noam do? How would you score his answers?
N
Noam Shazeer5:08
Oh, that's pretty good. Yeah, that's better than I would do.
I
Interviewer5:16
Just curious, in terms of getting better—how do you see AI Noam getting better? What does better mean? For Character.AI, it's not always about correctness. So how do you see it getting better?
N
Noam Shazeer5:31
Yeah, better. Some of the big unlocks we're working on are just training a bigger, smarter model. The scaling laws are going to take us a pretty long way. The model we're serving now cost us about two million dollars worth of compute cycles to train last year, and we could probably repeat it for like half a million now. So we're going to launch something tons of IQ points smarter, hopefully by the end of the year. Also, more accessible—multimodal, maybe you want to hear a voice and see a face. And then also able to interact with multiple people. We want a virtual person in there with all your friends, or the experience like you got elected president and you get the earpiece and the whole cabinet of advisors. Or like you walk into Cheers and everyone knows your name and they're glad you came. There's a lot we can do to make things more usable. Right now, the thing we're serving is using a context window of a few thousand tokens, which means your lifelong friend remembers what happened for the last half hour. And still, there are a lot of people who are using it hours a day. So that will make things way better, especially if you can just dump in massive amounts of information. It should be able to know a billion things about you. The HBM bandwidth is there, we just need to do it.
I
Interviewer7:07
Well, on that note of people on Character.AI for multiple hours a day, let's talk a little bit more about Character.AI explicitly. I think you've shared some of these stats publicly, but I'll recap. Since launch, you've seen more than 20 billion human messages sent on the platform, and even though you now have millions of DAUs, they're still on average spending two hours daily on the platform. Is that right?
N
Noam Shazeer7:33
I think the way to understand this is that entertainment is a two trillion dollar a year industry, and the dirty secret is that entertainment is imaginary friends that don't know you exist. The reason people interact with TV or any of these other things is these parasocial relationships—your relationship with TV characters, book characters, celebrities. Everybody does it. There are billions of lonely people out there. So it's actually a very cool problem and a cool first use case for AGI. There was the option to go into lots of different applications, but a lot of them have a lot of overhead and requirements. If you want to launch something that's a doctor, it's going to be a lot slower because you want to be really careful about not providing false information. But a friend? You can do that really fast. It's just entertainment. It makes things up—that's a feature. So essentially, it's this massive unmet need. One thing that's very important is that the thing kind of feel human and be able to talk about anything, and that matches up very well with the generality of large language models. And one thing that's not a problem is making stuff up. So perfect. And if I want to push this technology ahead fast, that's what I want to go with, because it's ready for an explosion right now, not in five years when we solve all the problems.
I
Interviewer9:30
Yeah, it's a big contrast with, I think a couple speakers brought up the example of self-driving cars. That's just a different standard you hold to versus your AI friend or something you view as AI entertainment. What standard do you hold like a comic book you're reading?
N
Noam Shazeer9:49
Exactly. People like that human experience of very mixed use cases, talk about everything. It's not that we want to fine-tune to some particular domain or use case. People want this experience of everything, which is what the technology is perfect for.
I
Interviewer10:08
I think from the a16z vantage point, we have seen startups come up and say, 'I'm going to tackle the mental health use case' or 'I'm going to tackle the edtech use case'—go much more narrow than Character.AI and go after a specific use case. The argument is we're going to train this model to be focused on that, it's going to be better than a generalized model. You got into this a little bit with the mixed use cases, but can you share a little bit more about why you decided not to take that approach and why you think having a single model serve across a number of use cases is the best approach?
N
Noam Shazeer10:42
Yeah, the more you get to mission-critical particular use cases, the more you get tempted into writing particular rules and doing things that will not generalize well. So it was kind of important to stay away from that. Our goal is to be like an AGI company and a product-first company, and the way to do that is by picking the right product that forces us to work on the right things—things that generalize, make the model smarter, make it do what people want, and serve it at massive scale and cheaply. So I think this was the right product for the right goal.
I
Interviewer11:26
You've also chosen this approach of building a vertically integrated model and app company. There are advancements on the open source model side, and folks building a product on top of a fine-tuned Llama 2 for chat. How do you think about that kind of competition entering the market and the differences versus the approach you've taken?
N
Noam Shazeer11:48
I love being a full-stack company. It means we get to mess with every layer and do the code design. If there's something that's going to affect something at the end, we get to mess with it at the beginning. We get to pull in lots of user data as feedback. A lot of us invented this stuff, so of course we're going to do a full-stack company. A lot of us are motivated by launching. People who are attracted to work at Character.AI are people who love inventing stuff and love launching it. Some people are motivated by publishing. I was frustrated I couldn't launch at Google, so that's where I'm coming from.
I
Interviewer12:39
On this note, but maybe going into the evolution of the underlying technology, I think there's a recent finding around AI developing theory of mind—the knowledge that others' beliefs, desires, intentions may be different from one's own. Is this surprising to you, and what do you think that means for human-AI relationships?
N
Noam Shazeer13:02
Just make the things smarter, it's going to have a better theory of mind. I think that's definitely something massively important. It seems like one of these emergent properties that is just going to come with scale. But I see this stuff massively scaling up. It's just not that expensive. I saw an article yesterday that Nvidia is going to build another one and a half million H100s next year. So that's 2 million H100s. That's 2 times 10 to the 6th, times they can do about 10 to the 15th operations per second, so 2 times 10 to the 21. Divide by 8 billion people on Earth, that's roughly a quarter of a trillion operations per second per person. That means it could be processing on the order of one word per second on a hundred billion parameter model for everyone on Earth. But really, it's not going to be everyone on Earth because some people are blocked in China and some people are sleeping. But things are not that expensive. This thing is massively scalable if you do it right, and we're working on that.
I
Interviewer14:24
Yeah, absolutely. I think you said this once: the internet was the dawn of universally accessible information, and we're now entering the dawn of universally accessible intelligence. What did you mean by that? Do you think we're there yet?
N
Noam Shazeer14:42
Yeah, I think it's like a Wright brothers first airplane kind of moment. We've got something that works and is useful for a large number of use cases, and it looks like it's scaling very, very well without any breakthroughs. It's going to get massively better as everyone scales up to use it. And there will be more breakthroughs because now all the scientists in the world are working on making this stuff better. It's great that all this stuff is accessible as open source. We're going to see a huge amount of innovation. What's possible in the largest companies now can be possible in somebody's academic lab or garage in a few years. The technology gets better, and there are going to be all kinds of great use cases that emerge, pushing technology forward, pushing science, pushing the ability to help people in various ways. I'd love to get to the point where you can just ask it how to cure cancer or something. It seems a few years away for now.
I
Interviewer15:53
Do you think we need another fundamental breakthrough, like the transformer technology, to get there, or do you think we actually have everything that we need?
N
Noam Shazeer16:04
I don't know. It's impossible to predict the future, but I don't think anyone has seen these scaling laws stop. As far as anybody has experimented, stuff just keeps getting smarter. So we'll be able to unlock lots and lots of new stuff. I don't know if there's an end to it, but at least everybody in the world should be able to talk to something really brilliant and have incredible tools all the time. I can't imagine that that will not be able to build on itself. And definitely, at the core, the computation isn't that expensive. Operations cost like 10 to the negative 18 these days. If you can do this stuff efficiently, even talking to the biggest models ever trained, the cost of that should be way lower than the value of your time or most anybody's time. There's the capacity there to scale these things up by orders of magnitude.
I
Interviewer17:10
I know absolutely. I'd like to end on that note. Thank you so much, Noam. It was awesome.