Wojciech Zaremba5:29
So I would say that we are at the moment in the place where actually a variety of skills turn out to be extremely relevant. So let me walk you through some examples. Okay. So for instance, one of the endeavors is figuring out what personality the model should have. Okay. And what personality the model should have. Correct. And you can think, you can use philosophical theories, you can think from the perspective of what the users want. Maybe also one framing that I heard which is quite interesting framing. I heard it from my colleagues from Anthropic is a framing saying the personality that I would like the model to have is a personality of a person whom I would be willing to give responsibility of deploying AGI to the world. Okay. Like imagine, so it's like what are the qualities of this person because you know at some point they might be tasked with this task. And you know, you might want to have trust towards how they are acting. And it's kind of interesting that in case of AI, which might be way easier than in case of humans, it's possible to very carefully examine behaviors at the micro scale, every single axis of the behavior. And for instance, you would like models to be able to admit their mistakes. Yes. Okay. That even if they make mistakes, you would want them to be able to admit their mistakes. That'd be good for humans too. But yeah, models for sure. Yeah. There is also something, you know, you might think that maybe the thing that you want from models is to have total honesty. Okay. But then when you look at humans, so okay, like you might think about even yourself when you are in the restaurant and you didn't like a meal. Mhm. And the server is asking you how did you like it and you're like, okay. Yeah. And you know, you can say that this is not fully honest. Yeah. But maybe that's okay. Yeah. Or if somebody's wearing shoes and I don't like their shoes, I don't think they're good shoes, I'm not going to go tell them I don't like your shoes. Even though that's honest, right, but it's not very helpful. So the AI probably has to determine what is actually helpful versus just honest. Yeah. So here you can start to describe that maybe the perspective to think about AI is from the perspective of being helpful. But then the question is, let's say if someone is causing harm to themselves, to what extent AI should help with that versus not. So there were also these framings some time ago, people were speaking about that we should create humanity-loving AI. Ilya was advocating for that and it's, you know, this is like a really compelling vision. It's almost like imagine that you have this super intelligence. It's like a data center and there are thousands or even millions of GPUs buzzing. Yeah. And you want this data center to think nicely about humans. Sounds good. And this framing kind of makes sense, but it also turns out to be somewhat problematic. And I can tell you why. So let's say you have a cigarette selling company. Okay. Should it be the case that when they try to revise their emails, the model is refusing to do so? Because it's like, ah, yeah, yeah, I don't know, yeah. And also who are we to judge what is good for humans versus not. People tried to do it sometime ago during prohibition. And that didn't work. And maybe the closest where we are today is that the AI should empower you. There are also some limits with respect to law that it should respect. Like even if you want to manufacture a pandemic, it shouldn't empower you. Right, right, yeah, I think that would be true. So those are the questions that a team has to grapple with and the team has to build the data set or the language system so that it acts in the way that you're most intended. So you can think, there's a few things. You can think about the personality and this, we have the model spec that even describes publicly how our AI is. You can also think about what are the topics that are not okay to speak about and within even these topics you have a taxonomy. Mhm. That specifies, oh, that's okay but that's not. You know, in case of bio, there's a bunch of things that you can ask, but you know, if you start taking pictures of petri dishes with viruses, maybe it shouldn't be advising you on the next steps. Yeah, that's great.