All right, hello everyone. Okay, Sarah, you're the queen of AI investing. This is a phrase never ever to be used again, but it's great to be here with both of you. So I had two different ideas for our final discussion. The first was a product off because these two men have the merge to prod button, both of them, and I was like, oh please just release everything we know is coming over the next six or 12 months, ignore all internal guidelines. The second was we just redesigned Instagram together since they've both actually run Instagram before. Both of these got fully shot down, and so instead I think we'll just like trade notes among friends. Like lame, I know, but really excited to hear from you both anyway. So this is actually a relatively new role for both of you. Kevin, let's start with you. Like you've done a bunch of really different interesting things. Like what was the reaction you got when you took the job from friends and the team?
Generally excitement. I mean, it's I think it's one of the most interesting and impactful roles. There's so much to figure out. I've never had such a challenging, interesting, sleepless product role in my life. It's got all the challenges of a normal product role where you're trying to figure out who you're building for and what problems you can solve and things like that. But normally when you're building product, you're building off of kind of a fixed technology base, right? You know what you have to work with and you're trying to build the best product you can. Here it's like every two months computers can do something computers have never been able to do before in the history of the world, and you're trying to figure out how that changes your product. And the answer should probably be a fair amount. And so it's just so interesting and fascinating to see on the inside as AI gets developed, but I've been having a blast. Mike, what about you? I remember hearing the news, I was like, oh I didn't know you could convince the founder of Instagram to go work on something that existed already.
Yeah, my favorite three reactions: like people who know me were like, that makes sense, like you're going to like have fun there. The middle people were like, why? Like you don't have to work, like why are you doing this? And like then if you knew me, you know me, like I can't not. And I think that like I couldn't stop myself. And the third was like, oh you could hire the founder of Instagram, which was also fun. And it's like, I mean not many people could, but like there's like probably a list of three companies that would have been interesting. And so yeah, there's like a range of reactions depending on how well you knew me and how like you've seen me in my like semi-retired state, which lasted like six weeks, and I was like, all right, what are we doing next?
So we had dinner together with a bunch of friends recently and I was like impressed by the childish delight that you had around like, yeah I'm learning about all this Enterprise stuff. Like tell me if it's about serving customers that are not all of us with Instagram or just you know working in an organization that's research driven, like what's the biggest surprise so far?
Those are two I think both very worthwhile pieces of this role that are like very new to me as well. Like when I was 18 I made this like very 18-year-old vow which was like every year of my life I wanted to be different, like I don't have the same year twice. And I was like it's why like you know I didn't, there's been times where I was like oh like another social product, I'm like doing that again. First of all like your bar is like really distorted and second of all like just would feel like too much of the same thing. So yeah, Enterprise has been wild. I'm really curious about your experience with that as well. Like your feedback loop, I actually imagine it's a lot more like investing is far longer, right? You're like you have that initial convo and you're like I think they like me and then you're like oh no it's now in like some requisition state and it's going to take like six months before they like even get to deployment before you know whether it's right. And so like getting used to that pace where like what why hasn't this shipped yet? They're like Mike you've been here two months, like this is like it's making its way through the VPS, like it's going to get there eventually. So like getting used to different timelines for sure, but like the part that is fun is actually getting the feedback in the kind of engagement where you're like once it gets deployed you have somebody that can call you and you can call them and be like how's it working for you? Like is this good? Whereas users like you're doing like data science and aggregate and like sure you can bring in like one or two people but it's like they won't, they don't have enough financial incentive riding on telling you where you suck and where you're doing well. And like that's been a different but also like rewarding side of that for sure.
Kevin, you've worked on such a wide range of products before, how much do your instincts like apply?
Yeah, I was going to add on to the Enterprise point too, and then I'll get to like the other interesting thing about Enterprise is it's not necessarily about the product, right? There's a buyer and they have goals and you could build the best product in the world that all the people at the company might be happy to use and it still doesn't necessarily matter exactly. I was in a meeting with one of our big Enterprise customers and they were like, this is great, we're really happy, da da da, you know the one thing we need is we really need you to tell us 60 days before you launch anything. And I was like, I also would like to know 60 days. So very very different actually. And it's interesting right because at OpenAI we have a consumer product and we have an Enterprise product and we have a developer product, so we're kind of doing all at once. Instinct wise, in I'd say in like half the job it works. You know when you have a sense of the product you're trying to build, you know we're getting towards the end of shipping Advanced Voice Mode or something, or you're getting towards shipping Canvas and you're making final touches trying to understand who you're building for and like what exactly what problems you're trying to solve, it works then because that's a little bit more like the tail end of it is shipping a normal product. But the beginning of these things is nothing like that. So like there will just be these capabilities that we don't know. You have some sense as you're training some new model that it might have capability X, you don't really know, nor does the research team, nor does anybody, right? You're like I think this might be possible and it's kind of like coming through the mist, but it's this emergent property of a model. And you know so you don't know whether it's going to really work and you don't know whether it's going to be like 60% good or 90% good or 99% good. And the product that you would build that would make sense with something that works 60% of the time is super different than 90 or 99% of the time, right? So you're kind of just waiting and you're at least like I don't know if you feel this, checking in with the research team from time to time like hey guys how's it going, how's that model training, any insight on this? And they're like it's research, we're working on it, you know it's we don't know either, we're working through this at the same time. And it makes it super fun because you're kind of like discovering things together but very sort of stochastic too. It's the thing it most reminds me of like from the Instagram days where like Apple like WWDC announcements, you're like this could either be awesome for us or could like absolutely cause chaos for it. It's like that but your own company is the one kind of disrupting you from within, which is like very cool but also like oh this might totally end my product roadmap now.
Yeah, what does that cycle look like for both of you? You described it as you know like peering through the mist trying to look at the next set of capabilities. I mean can you plan if you don't know exactly what is coming and what is the iteration cycle to discover new things that should belong in your product?
I think like on the intelligence side you can sort of squint and see like all right it's advancing this way and so the kinds of things that you'll want to do with the model and start building the product around that. So there's three ways, right? Intelligence feels not predictable but at least on like a slope that you can kind of watch. There's the capabilities you decide to invest in from the product side and then do fine-tuning with the actual research teams. So something like Artifacts, we spend a lot of time between research. I think the same was true with Canvas, right? Like you're doing a like co-design, co-research, co-fine-tune, and that's like I think a real privilege of getting to work at this company and getting to do design there. And then there's the capability front, so maybe Voice Mode for OpenAI, for us it's the computer use work that we released this week. You're like all right, 60%, all right, got good, yes, all right. And like so what we try to do is embed designers early in the process but knowing that like you're not placing a bet. Like the experimentation talk was saying like your output for experiment should be learning not necessarily like perfect products you're going to ship every time. I think the same is true when you partner with research, like your outcome is hopefully demos or informative things that could spark product ideas, not like a predictable product process where you're like well it's this derisked by now which means it's going to look that way when research comes along.
I've also one thing that I've really enjoyed because research is at least parts of research are very product oriented, especially on the post-training side like Mike was saying, and then parts of it are really like academic research at some level. And so we'll just also occasionally hear about some capability and we'll be in a meeting and you'll be like oh I really wish we could do this thing and a researcher on the team will be like oh no we can do that, we've had that for three months. And we're like really? What does that, like okay where do I learn more? And they're like oh well we didn't think, we didn't know it was important so you know I'm working on this other thing now. But you do just get like magic happening sometimes too.
One thing like we think a lot about when we're investing is actually like can you do anything with a model if it is 60% successful at a task instead of 99%? And unlike lots of tasks that's closer to 60, right? But the task is really important, valuable. Like how do you think about that internally in terms of evaluating progression on a task and then what types of things you like put in sort of the burden of product to make it graceful failure or to like sort of cross the miles the user versus you know we just need to wait for the models to get better?
I'd argue there are a lot of things that you can actually do when something is 60% right. You just need to really design for it. You have to expect that there's a human in the loop a lot more than there would be otherwise. Like if you look at take like GitHub Copilot, right? That was kind of the first AI product that really opened people's eyes to like this thing can be useful not just as you know Q&A but for really economically valuable work. And that launched I don't know exactly which model that was built off of but I mean it was multiple generations ago so I guarantee you that model wasn't perfect at anything related to coding. I think it's GPT-2 which is like pretty small. So yeah, I mean and so but the fact that it was still valuable for you because if it got the code you know some significant fraction of the way there that was still stuff you didn't have to type yourself and you could edit it. And so there are experiences like that that I think totally work. I think we'll see the same kinds of things happening with sort of the shift towards agents and longer form tasks where you know it may not be perfect but if it can save you five or 10 minutes that's still valuable. And even more if the model can understand where it doesn't have confidence and can come back to you and say I'm not sure about this, can you actually help me with this? Then you know the combination of human and model together can be much higher than 60%.
I also find that 60%, this magic 60% number, like it's kind of lump. I made it up five minutes ago. That was a takeaway. 60%, that is our new, that's the Mendoza Line of AI. Like I think it's often very lumpy where like it'll do very well on some tasks and not well on others. And I think that also helps like when we ever run like pilot programs with customers, it's really interesting when we'll get like the same day feedback from two different companies. One will be like it's solved our whole problem, like we've been trying to do this for three months, thank you. The other will be like it was way off, it's like worse than the other model. And so like it's also humbling to know that you have your own internal evals but like the rubber hitting the road and actually seeing the model out in the world is where it's kind of the equivalent of like you do all this design and then like you put it in front of one user and you're like oh wow I was wrong. The model has that feeling as well. We try as hard as we can to like have a good sense but then people have their own custom data sets, they have their own internal use, they've prompted it a certain way and like so that delivers that sort of almost like vodou nature of when you actually put it out in the world. I'm curious if you feel this.
I think there's a very real sense in which models today are not intelligence limited, they're eval limited. Yeah, they can actually do much more and be much more correct on a wider range of things than they are today. And it's really about sort of teaching them, they have the intelligence, you need to teach them certain specific topics that you know maybe weren't in their original training set but they can do it if you do it right.
Yeah, we've seen that all the time where like there was a lot of like exciting AI deployments that happened in like you know maybe three years ago and now they're like we think the new models are better but we never did evals because all we were doing was just shipping cool AI features three years ago. And like the hardest hump to get people over is like let's step back and like what does success actually look like for you? Like what problem are you solving? Like often the PM has rotated so it's like somebody's inherited it and then be like all right what does that look like? All right let's do some evaluations. What we've learned is like Claude is actually good at writing evaluations and also grading them. So like we can automate a lot of this for you but you have to tell us what success looks like and then let's go and actually iteratively improve our way over there. And like that is often like the difference between like 60% of a task and like 85% of a task. If you come interview at Anthropic, which maybe you should at some point, maybe you're happy in your role, maybe not, you'll see one of the things we do in our interview process is actually like make you get a prompt from like crappy eval to good and like just we want to see you think. But like not enough of that talent exists elsewhere so we're trying to get that like if there's one thing we can teach people that's probably the most important thing, writing evals.
I mean it's I actually think it's going to become a core skill for PMs. We actually had this, and maybe this is like a little inside baseball but I thought this was interesting, like internally we had our research PMs who like work a lot on model capabilities and model development and then we had our like more like product surface PMs or API PMs. And we ended up realizing that like the job of a PM in 2024, 2025 building AI powered features is looking more and more like the former than the latter in a lot of cases. Like we launched like our code analysis and like basically Claude can analyze CSVs and write code for you now. And the PM there was like getting it 80% of the way there and then having to hand it over to the PM that could write the evals then go to like fine tune and like prompt. I was like that's actually the same role. Like the quality of your feature is now gated on how well you have done the evals and the prompts. And so like that PM definition is definitely just merged now.
Yeah, absolutely. We set up a boot camp and like took every PM through writing evals and like what it was like difference between good and bad evals. And you know we're definitely not done there, we've got to keep iterating and getting better on it but it is such a critical part of making a good product with AI.
Yeah, as part of this recruiting call for any of the people who want to be good at building AI product or research product in the future, we can't come to your boot camp Kevin, so how do we develop some intuition for getting good at this eval and iteration?
I actually think it's something you can use the models themselves for. Like you were talking about, you can ask the models at this point what makes a good eval, give me you know I want to do this, can you write me a sample eval? And it will be pretty good. Yeah, I think that goes a long way. I think there's also this question of like, and if you listen to like everybody from like Andrej Karpathy to others who have like spent a lot of time in the field, like nothing beats looking at data. And so like people often get caught up being like well we already have these evaluations and the new model is like 80% there rather than 78%, we can't ship, or like you know it's worse. And I was like have we looked at the cases where it fails? And you're like oh actually this was better, it's just our grader is not as good, you know. Or it's funny like again a little inside baseball, you know like every model release has the model card and some of these evals we've seen like even the golden answer I'm like I'm not sure a human would say it or like I think that math is actually a little wrong. Like getting 100% is going to be really hard because even just grading them is very challenging. So like I'd encourage you to like the way you build the intuition is go look at the actual answers, even to sample them, be like all right yeah maybe we should evolve the evals or maybe like the vibes are good even if the eval is like tied. So like getting real and getting like deep on the data I think matters.
I also think it'll be really interesting to see how this evolves as we go towards longer form more agentic tasks. Because it's one thing when your evals are like I gave you this math thing and you were able to like add four-digit numbers and get to the right answer, you know it's easy to know what good looks like there. As the models start to do more long form, more ambiguous things, go get me a hotel in New York City, you know what's right there? A lot of it will be about personalization. You know if you ask any two humans who are perfectly competent they're going to do two different things. So your grading becomes much softer and it'll just be interesting. I think we'll have to evolve yet again. Like speaking of having to reinvent stuff over and over again, I think a lot like when you think about, and I think both labs have some concept of like this is what capabilities look like as things evolve, like it looks a little bit like a career ladder, like what bigger and longer horizon tasks are you taking? And maybe like evals start looking more like performance review. I'm in performance review season so this is the metaphor that's in my head, sorry. But it's like you know like did the model meet your expectation of like what a competent human would have done? Did it exceed it because it did it twice as fast or like discovered some restaurant you wouldn't have known? It greatly exceeds, meets, most, like it starts being like more nuanced than just like right or wrong. Let alone you have humans writing these evals and the models are getting to the point where they can often beat humans at certain tasks, like people prefer the model's answers to a human's answers. And so if your humans writing your evals like yeah, you know so what does that mean?
Okay, evals are clearly the key. We're going to go spend a bunch of time with these models teaching ourselves to write evals. What other skills should product people be learning now? You're both on that learning path.
I think prototyping with these models is a thing that is underused. Like our best PMs do this where we'll get into some long conversation about like should the UI be this or that and before our designers have even like picked up their Figma, like often our PMs or sometimes our engineers will be like great I prompted Claude, I did like an A/B comparison of what these two UIs could look like, let's try them. And I'm like oh this is really cool. I'm like play that out and like we'll be able to prototype much like a far greater variety and evaluate like on a much faster scale than before. So like that skill of like using these tools to actually be in prototyping mode I think is a really really useful one.
That's a good one. I would also, you sort of said this but I think it's also going to push PMs to go deeper into the tech stack. Yeah, because it's and maybe that changes over the years, like if you were doing like database tech in I don't know 2005 maybe it required you to be able to go really deep in a different way than it would if you were doing database tech now. Like layers of abstraction get built and you maybe don't need to know all the fundamentals. But it's not like every PM needs to be a researcher by any means, but I think having an appreciation for it, spending time and learning the language and gaining intuition for how this stuff works a little bit I think will go a long way.
I think the other piece like you're dealing with this like stochastic non-deterministic system which like evals are our best attempt to do it but like product design in a world where like you're not in control of what the model is going to say, you can try. And so like what are the feedback mechanisms that you need to close that loop? Like how do you decide when like the model's gone astray? How do you collect that feedback in a rapid way? You know like what are the guardrails you want to put in? Like how do you even know what it's doing in aggregate? Like it's a much more like you're understanding like the output of this intelligence across a lot of outputs over a lot of people every single day. It just requires a very different set that like oh the bug report is you clicked on the button and didn't follow the user, it's like that's a pretty knowable kind of problem, right? And maybe this will change, you know 5 years from now when people are used to it, but I think we're all still in the mode of adapting to this sort of non-deterministic user interface ourselves. And certainly people who are not you know tech people here in this room working on tech products who are using AI are definitely not used to it. Like it goes against all of the intuition that we've built up for the last like 25 years of using computers. And so like the idea that you're going to put in the exact same things, normally if you put in the exact same inputs computers give you the exact same outputs and that is no longer true. And it's not just that we have to adapt to it building products, we have to also put ourselves in the shoes of the people who are using our products and think about what this means for them. And there's like I mean there are downsides to it, there are also really cool upsides. And so it's fun to kind of think about how you can use that to your advantage in different ways.