Back
Mark Chen
Chief Research Officer, OpenAI

OpenAI’s Research Playbook: Scaling Laws, Reasoning, Evals, and AGI — Mark Chen | Let Them Cook

🎥 Jun 25, 2026 📺 Latent Space ⏱ 41m 👁 1163 views
In this episode, we have Mark Chen, Chief Research Officer at OpenAI, joining us to make Korean tofu stew and flambé shrimp. From scaling laws and why pre-training is not dead, to the o1 reasoning bet, evals crisis, research taste, long-context learning, and the future of end-to-end AI research, we cover what it takes to push models toward the frontier. We talk about: • Why Mark still believes in scaling laws and the exponential • How OpenAI chooses research bets and allocates compute • Why reasoning became one of OpenAI’s biggest bets • How to develop research taste without a traditional ML...
Watch on YouTube

About Mark Chen

Mark Chen, Chief Research Officer at OpenAI, appeared on the podcast "Let Them Cook" on June 25, 2026, where he discussed the company's research direction. Chen stated that he "firmly believes in being on the exponential and in scaling laws" and said he "fairly strongly disagrees" with views that pre-training is dead, noting that such narratives have recurred throughout the history of developing large language models. He described reasoning as "one of the biggest examples" of a research bet at OpenAI, referencing the o1 model as a breakthrough that was difficult to get off the ground because the pre-training plus post-training paradigm "felt like such a promising paradigm" at the time. Chen also said the field is in an "evals crisis," with a low number of canonical gold standard benchmarks, and noted that tools like Codex have enabled faster iteration of evaluations. On June 16, 2026, Chen appeared alongside SoftBank CEO Masayoshi Son at an event in Tokyo where SoftBank announced a cybersecurity service using OpenAI's technology. Chen described cyber capabilities as a "dual-use capability," stating that "even though the models can get better and better at finding vulnerabilities, we can use that for defense." He added that "the important thing is we can go and try the models and try to find the vulnerabilities before external actors can go and try to find the same vulnerabilities." Son compared the dynamic to a criminal with a knife facing a police officer with a gun, saying defenders must have the "best, most powerful weapon" to defend against bad actors.

Source: AI-verified profile updated from Mark Chen's recent appearances. Browse all interviews →

Transcript (134 segments)
M
Mark Chen0:03
Yeah. Like that. I know. That feels like already the situation I'm in real life. Cheers.
A
Alan0:10
Cheers. Hey guys, welcome to the Leighton Space Cooking Series where we invite founders and researchers and just let them cook. Today we have a very special guest, the chief research officer of OpenAI, Mark Chen. Welcome.
M
Mark Chen0:24
Thanks for inviting me, Alan.
A
Alan0:25
Yeah, thank you for coming. To begin, this all started from the inspiration after hearing a story that Mark Zuckerberg would make soup to try to poach researchers, and in response you brought soup to researchers. Is this true? Did this happen? Did it work?
M
Mark Chen0:39
Oh, you know it's absolutely a true story. And I have brought soup to our own researchers. I think that made us calm down a little bit. I think we came out on top. But yeah, still a very funny story in the craziness of how AI's evolved.
A
Alan0:52
How often do you cook? Is it something you're familiar with?
M
Mark Chen0:55
Well, you know, I do enjoy cooking, but I don't have the luxury of doing that so often. I usually have a work dinner every night of the week. And, you know, maybe post AGI, this is going to be my hobby. I've always joked I'm going to start a noodle stand once it's all over.
A
Alan1:10
Yeah, post AGI, hopefully that'll still be there. Great. Looking at what we have in front of us, do you have an idea generally of what we'll probably be making?
M
Mark Chen1:19
Korean tofu soup, maybe.
A
Alan1:22
Yeah, that's generally what it is. So, inspired by the story of you bringing soup to researchers, we're making a tofu Korean stew, and then we have prawns that we'll be cooking. Are you ready to go?
M
Mark Chen1:32
Yeah, let's do it.
A
Alan1:33
Great. The first thing we should probably do is separate the veggies and then cut them. Basically what we want to do is cut the dirty part off with the dirt, then separate that across.
M
Mark Chen1:44
That I know.
A
Alan1:46
Okay. Just have that there. So we could do that. And while that's going, I guess I could ask more about your background. In a previous life, you were once a trader, and even Sam, I think last year in April, tweeted about how if you're a high-frequency trader, you should consider joining OpenAI because, you know, build AGI. So do you think there's a relation between being a trader and being a researcher, or do you think it's just a very technical and competitive area where a lot of great employees can come from?
M
Mark Chen2:13
Really the most important thing is there are a lot of researchers who just started out without formal training in machine learning or AI research. We've very much believed in training people up to do this. I think the real hard thing is the ability to creatively solve problems and think outside of the box. It's not so much you have to do a PhD, even though that does bring a valuable skill set. With trading in particular, I don't know that it's that special of a profession. I kind of think of it as, we've had great mathematicians join, we had great physicists join. But trading is something where it's very unhackable — you can't cheat the real world. It's a hard metric to optimize. There's also a lot of characteristics like, it's a field where attention to detail really matters, and it's kind of the brutal hard optimization and squeezing out the juice of a system. And some of those skills transfer over.
A
Alan3:17
Gotcha. For people who want to get into research but don't have a PhD, what do you think are the main attributes or things they can learn to develop research taste? Because I guess that's the main part of getting into this field that may be very foreign to them.
M
Mark Chen3:32
I think research taste is a little bit overrated. It is something you have to develop, but the best mechanism I've found for developing that is really just replication. I think you should take papers that you really look up to and try to fully replicate them. A lot of replication stood out in my mind. Back in 2018, there were ResNets, there were Pixel CNNs, and I learned so much just trying to replicate the training curves exactly — get to the exact amount of training loss or perplexity the papers hinted towards. It teaches you a lot of techniques that people don't really talk about, but once you dive in a couple layers deeper, you learn those techniques. Really the first thing that got me into the field was when AlphaGo played Lee Sedol. That was the turning point for so many people. It was inspirational, and the first big project I really went after was: can I get a DQN working?
A
Alan4:40
Yeah.
M
Mark Chen4:40
Yeah, that's true. I think it was move 37 or one of the games. It was pretty insane watching it happen and seeing all that develop and where we have gotten to today, especially with research. I mean, isn't it crazy that you're seeing move 37s in almost every field now? There's move 37s in math, in computer science, and coding. I think even, yeah, it feels like a lot of people woke up at the start of this year and were like, man, agents are working in my profession, and they're essentially realizing that these models can just do long-horizon meaningful work for them.
A
Alan5:17
Yeah. No, that's true. It is very impressive to see, even just using in my own work. But okay, the next thing we could do is just cutting the onion. So what we have to do here is just dicing it. Do you think there's jobs that RL maybe will have a much harder time to break into? For example, coding may be easier since a lot of the context is accessible, whether it be the code bases or even the work you're trying to do. But let's say if you're trying to do the job that a junior consultant may do, where all the context is a little scattered, maybe a little more difficult. How do you view those different scenarios? Is there a way you assess what can be the right approach?
M
Mark Chen5:53
RL has traditionally had headwinds when it comes to fields that are more subjective than objective. If you think of one example, creative writing, where you can take two pieces of creative writing and two experts can have wildly different opinions. It's these fields where things are hard to grade where RL has the least ability to directly apply. I know a lot of people are developing techniques to apply RL in these settings, but for now it's where there's cold hard truth — things like math and computer science where you implement it correctly or wrong. That's where you kind of see it really taking off.
A
Alan6:46
That actually brings up a thought in terms of evaluating those fields. As models get much more powerful and they even saturate, for example, solving the IMO questions, how do you view evaluating superhuman intelligence? Getting to a point where it's so good at things that even the top 0.1% of humans can't do — how can we push past that frontier of intelligence?
M
Mark Chen7:12
It's kind of crazy. I feel like a lot of it centers on interfacing with the real world. When we've thought about how to evolve past things like programming context in the past, a lot of the initial direction we took was you should move it to real-world research. And we've seen that the models have gotten a lot better at discovering novel theorems and pushing the frontiers of hard sciences. But even today, that's no longer a surprise. We almost take it for granted now that these models can solve very difficult problems. They can make contributions and even draw relationships between fields that are novel and insightful. I think we think of coding as really a domain that tests if our models can learn in high-context settings and in real-world long-horizon settings.
A
Alan8:11
Gotcha. Okay, yeah, that makes sense. And since you're done with all the vegetables, we can now do the next step which is sautéing it. Yeah, we can use the induction stove. Let me turn it on. We'll just sauté it with some oil. So I put the pan on the front burner.
Great. Perfect. And then while that heats up, we can just wait and then add the vegetables. But yeah, I guess more so on views for research — are there commonly accepted ideas that you disagree with? Whether it be like 'pre-training is dead' or 'language models will never get us to AGI.' I think there's a lot of takes out there that are very ambiguous and obviously haven't been proven out yet. And from your perspective as the research lead at OpenAI, what would you say?
M
Mark Chen9:13
I firmly believe in being on the exponential and in scaling laws. So I think any of these bare takes I fairly strongly disagree with. When it comes to 'pre-training is dead,' I think the funny thing is this narrative only started spreading more widely after the last one or two years or so. But in many times in the history of developing LLMs, people have been saying this. There have always been some bottlenecks — 'you can't scale past this because of this bottleneck.' And we've always found some kind of technique, whether it be better engineering or some new research insight, that helps you break past the boundary. And so I think it's just more and more of the same — more careful research engineering, more careful data engineering, more careful scaling — and it always unlocks that next ability to scale further. It's held for almost 10 orders of magnitude, but there's no reason it should not keep holding.
A
Alan10:15
Yeah, that's a very fair point. I guess on research bets that have helped you scale beyond — were there specific ideas that you can even remember in the early days that everyone was saying this is not going to work?
M
Mark Chen10:26
Well yeah, I think of reasoning as one of the biggest examples of this. The first breakthrough that we launched to the world was o1, but it wasn't easy to get that off the ground. Because the world we were living in back then — pre-training plus post-training — that felt like such a promising paradigm. And so even at a company like OpenAI, you would have people ask naturally, 'Why do something when you have a machine that works?' And fundamentally, it's to the credit of Jakub, Ilya, many of the people who really had conviction and vision in this space that we started pushing on this in earnest. And even then it took a lot of steering to get the whole company behind this as a fundamental bet.
A
Alan11:12
Gotcha. And how do you kind of develop that ability to motivate researchers? Because I assume that's a big part of taking a lot of bets and some will pan out, but still building the trust in the team to know that eventually some of these will actually have power-law effects.
M
Mark Chen11:24
You know what's really cool about OpenAI is research feels like a meritocracy. Often times the research managers are the people who have done the best research in the past. And so I think a lot of steering can come top-down — like if your manager says, 'Hey, I'm really convinced this is the path forward,' generally people will take that into heavy consideration. It's like, this person who you've respected for their research taste and execution for so long is now very excited by this idea — it's definitely something people take into account. So I think there's good top-down steering. At the same time, one really cool thing about OpenAI is there are bottom-up elements. We like to be convinced that we're wrong, and someone can just come with cold hard evidence. And many things like that have turned into core parts of our research roadmap — just things that no one was really trying to steer but some researcher on the ground had heavy conviction in. And that's also a really big delight to see.
A
Alan12:32
Yeah. No, absolutely. I heard in a recent interview that you gave that your internal research roadmap hasn't really changed, even through all that we've seen with model development and even other companies. How often do you guys assess that, reassess that, even act proactively? I assume it's not a lot of reactive decision-making as other models come out, but how do you think through that process, especially as everything around you just continues to get better?
M
Mark Chen12:59
I think the high-level research roadmap should be stable. People need something to ground in. People need to see a path to what we're building. And I've been very happy that we've stayed the course for a while. But the implementation details can change over time. The sequencing will matter, the relative resourcing will matter, and the exact threads on the ground will matter. What we do is we have kind of points in time that force us to reconsider these things. One example is when we do compute allocation. One of the parts of the job is just figuring out how to allocate compute to projects. And it's a time to question: are we really putting compute and people to use at the highest priority areas?
A
Alan13:47
And I guess, could you clarify what you mean by the higher level versus the more implementation details? Like, high level as general as AGI that's our north star, or is it more granular than that?
M
Mark Chen14:02
At the very highest level, we have an org that focuses on pre-training — giving models a lot of world knowledge. We focus on RL, teaching the models how to reason with that knowledge, how to chain little insights together. And then finally, alignment and post-training. And we're always looking at both how to scale the main line in each of these domains and also new bets that fundamentally unlock either different scaling properties or more aggressive scaling properties.
A
Alan14:33
Gotcha. And so even in that, I heard that every one to two months you go through like 300 research projects that could be followed through on. Is there a way that you kind of hone that decision-making as you decide what to actually double down on and what not to? Because I assume there's a lot of talented researchers who provide possible ideas to pursue.
M
Mark Chen14:52
So, in the spirit of focus — one narrative you might have heard is we're really focusing our bets at OpenAI. And we're also trying to do a little bit more directive compute allocation as well. I don't like micromanaging my managers. I think one important thing is to empower them, but to just give big swaths of compute to the big bets you want to make, and then also give them flexible pools of compute which they can freely allocate to things that they believe in or kind of fudge with the allocations that we prescribe. So I think it's just tying a small number of bets — say three to five bets from each org — into the main research roadmap, and then really letting the managers and org leads take things from there.
A
Alan15:46
Gotcha. Okay, that makes sense. I guess for rising researchers, let's say in an interview setting, are there specific tells or ways that you can identify, okay, this person has some potential of becoming a researcher to impact an org in a specific way? Or is it just looking at the previous research that they've done and that's what heavily dictates whether they can actually continue on?
M
Mark Chen16:09
It's a hard problem before someone comes to OpenAI. I think that's genuinely true. For a lot of the best research managers, they work with so many researchers over time where you kind of develop an intuition — the things that they say, the ideas that they bring up. Do those hit the same mark? Are they the things that you would be thinking about personally too? And so there's this gut check of does their intuition match the same intuition that you have. But it is really hard to tell out of the gate. Usually in, say, six months to a year, it's pretty clear who has the strongest trajectory and who's going to make a lot of impact. Honestly, I think it's a hard problem, but just having seen a lot of people go through research development at OpenAI, you develop an intuition for who's more P0 in different areas. And one thing to mention there is not every researcher is the same. There's a lot of different types of impact. There are the people who just take an idea that's very clear and they'll just implement it before anyone else. There are also the people who just come up with the crazy, almost too crazy, moonshot type — but somehow not that crazy, and they really convince you in a different way of seeing the world or a completely different type of project. So there's a lot of ways to make impact.
A
Alan17:43
Yeah. No, that's helpful. And so I guess elaborating on that, would you say there's similarities that you would see between top engineers and top researchers? I often hear top engineers, even at small companies and startups, are ones who can take an idea and see it all the way through to the product. Or do you think it's more so they're focusing solely on the research, not considering the end design, how it's used by the customer?
M
Mark Chen18:12
Well, the thing about research is many times the path forward is unclear. And so what differentiates researchers is how often they're pointed in the right direction. Like you say taste, right? I think in engineering there are certain patterns that work — like if you want to build a product that looks this way, the engineering principles can be pretty similar. But for research, I think the thing that's slightly different is just this ability to have good research taste, to convince other people that what you're doing is promising, and then to just integrate it into the core research roadmap.
A
Alan18:53
Gotcha. Great. Okay, it seems like we're done with the vegetables. Now we have to multitask. We're going to pour some water into our pots to get the base of the soup going. So, in the top right. Here. Let me pour some here. You can use some as well. And while we have this simmer, we'll add the veg and cook our prawns here. Let me clean this up real quick. Looking great so far. I feel like we got some color on the onions and mushrooms. Let's turn this on.
I guess one aspect or area that seems very interesting is evals. And more specifically, have there been instances where you've seen through just vibe checks that it is really good, but on the actual benchmarks it performs very poorly? Or do you think it's heavily correlated — that if your SWE-bench Pro score is a high number, then your vibe check on it doing coding tasks is also really high?
M
Mark Chen19:53
No, I mean I think there is this phenomenon — internally, I'm not sure if this is an externally used word — but just like benchmaxing. I think you can kind of overfit onto certain distributions and it won't be reflective of how well you generalize. Easy ways to do this are: you take a benchmark and you find very similar types of instances to the benchmark and you overtrain on those instances. Beyond that, the other scary thing in the field is the number of canonical gold-standard benchmarks is low, and we really are kind of in an evals crisis. Where all the really great evals that we all know, like growing up taking the SAT, those are all fully saturated. And we really need to find good new ways to benchmark the models. I think one great thing about tools like Codex is they've really enabled the fast iteration of evals — we're able to have one person very quickly put together a very high-quality eval. Another interesting thing about just being able to deploy your models is you can just see how they perform as people are doing things with them. In math, in coding and software, you get a sense for where they fall over what task horizon they can do from this very broad-based deployment.
A
Alan21:21
Yeah. No, that's helpful. Now we'll just add the prawns to the oil and get some color on it. So I guess double-clicking onto that — how do you balance both doing well on these benchmarks but also not benchmark-maxing as you said? Because I assume you want to be most honest and not kind of game the system. But if you have a lower score than a competitor or other models, the consumer may be like, 'Wait, your scores aren't that great, so the model probably isn't that good.' How do you balance both of those?
M
Mark Chen21:55
You just really have to operate over representative mixtures of evals and always invest in creating new evals. There's this philosophy of once an eval is out in the world, then it's already not a good eval. And one thing is also partnering with external organizations to create evals. In many of the hard math and science evals, we've partnered with external organizations and they've been able to craft gold standards there for us. There's a kind of interesting philosophy of separate the teams that are creating the evals from the teams that are optimizing the models themselves, because that way you don't co-incentivize them. The evals team can work to build evals that are hard for the model, so there's this inherently adversarial process where you're not kind of treating yourself nicely. The incentives are somewhat aligned in the right way between the two teams.
A
Alan23:01
Yeah. And do you kind of also contribute and help in the ideation process, or even deciding what eval you should work with a third party on to develop?
M
Mark Chen23:10
Yeah. So I think a lot of the work that Jakub and I do also involves just steering the direction the evals go. We'll notice certain gaps or certain capabilities that we want — and every capability on the flip side is an eval, right? You need some kind of eval that measures if you've elicited that capability well. So yeah, it takes a lot of steering, and just to get everyone on the same page with evals is also a lot of prep work.
A
Alan23:37
Yeah. No, that's great. I guess on Jakub — you said in a previous interview that he's a very funny guy. Do you have any fun stories that maybe you haven't shared about working with him? Because you also say that you guys align very well, so your discussions even on research are very efficient and help a lot when driving towards a frontier. On the opposite side of being very funny, are there things that you share?
M
Mark Chen24:00
You asked about a funny story. Well, he told me this joke yesterday which I thought was very funny. In many ways we kind of jointly manage the research efforts, and apparently some researcher came up to him and was like, 'It feels like I now just have an army of really dumb, like gold medalists.' And Jakub's like, 'That feels like, Corey, the situation I'm in in real life.' So yeah, he's just brutally sarcastic and funny.
A
Alan24:33
No, that's great. It's great to have humor in the workplace, you know, to balance out, especially as you're pushing the frontier on very important work. But that also brings to mind one kind of weird scenario of how models can perform very well on, let's say, the IMO or even the ILY but may struggle with some more mundane task that a human can easily do. So how do you deal with that?
M
Mark Chen24:55
Ultimately I think what's intuitive for the models is often not that intuitive for humans. There's a lot made of this jagged frontier analogy where there's some things that the model is just inherently good at, maybe based on the data it sees or the things we can teach it more easily. I actually think a lot of it boils down to context — the models don't have a lot of context that a human has. Vision, of course, is something that's more naturally biologically wired for humans. And so yeah, there are just certain jagged capabilities that models are better at than humans and vice versa. But I also think context — just being able to take a single task, learn lessons from it, and apply them to future tasks — that capability is something that a lot of people are working towards right now. But it's very natural for humans.
A
Alan25:55
Yeah. And on the context point, a very low-hanging fruit example that many people say is just to increase the context window to provide more examples so the model can perform. But I assume there's more complexity on how to actually enable this. Even with a large context window and a lot of context, there could be bloat or even just a lot of context rot as people have said. So how do you go through that process of navigating that?
M
Mark Chen26:18
I think there's kind of the canonical way you would solve for very long-horizon learning, which is you just naively increase your context window. And that makes sense. I think there's a difference between implementing long context and implementing long context well, like you said. There's a lot of needle-in-the-haystack style to measure that. But beyond that, there are also a lot of engineering and research shortcuts that you could take. Many coding products today have features like compaction, where you can compress either insights or working state. Stuff like that just shortcuts a lot of the very brutally difficult and expensive primitives that you have to build with just native long context.
A
Alan27:11
Gotcha. Great. Okay, now we're going to do the fun part. Let's lower the heat a little bit and add a little more oil to the pan. And then we'll torch the prawns to get more flavor in there.
M
Mark Chen27:24
Yep. But one-shot learning.
A
Alan27:30
Yeah. So, I'll first do it on my pan to show you what it looks like. Wait, I didn't pour any bourbon. Okay, so let's pour like a fourth. And then pour it in. Heat is off. And then torch it.
M
Mark Chen27:49
Awesome.
A
Alan27:50
Great. So pour it to like half the fourth cup. And then once you have that, I can give this to you. Great. So it's off and then once that's off, we can turn it back on. Do you want to hold on this and just press this button to fire it up?
M
Mark Chen28:14
Yeah. Cool.
A
Alan28:14
Great. Okay. Bon appétit. It's a little light, but yeah. Great. And then we could turn on the heat again. And then we'll just cook off the alcohol.
M
Mark Chen28:25
Great. How are you feeling?
A
Alan28:28
Basically just cooking everything out. Awesome. I guess in terms of research ideas and what to work towards, do you think there's still a lot of low-hanging fruit or ideas that can still be improved a lot through just optimizing small parts of already implemented work? Or do you think right now there has to be a lot of research that are completely new bets that people take?
M
Mark Chen28:51
Yeah, that's a really great question. I feel like there are new bets but probably not that many. Hopefully you feel like AGI is coming soon, right? And everyone sees that these models are getting really capable. If you really imagine the implications of that, we're getting closer and closer to a world where the models can come up with more of the innovations on their own. They can kind of do self-sustained research. This is one of the big goals that we've set for our research work. So I think what really matters is are there big bets before that point in time. And I think the window is small, but there are still some fairly significant ideas we're trying out.
A
Alan29:35
I mean there have been some researchers who have stated that to get to AGI, we still need let's say two or three more breakthroughs — like continual learning or some other ideas. Do you follow that same perspective, or do you think it's not as drastic as coming up with three completely different paradigms?
M
Mark Chen29:55
I don't know. I don't know if that same framing — like continuous learning is a basic primitive that you have to unlock. There's so many different techniques. I don't know what I would consider as a breakthrough versus not, but I think there are clearly many shots on goal and I'm pretty sure it'll work.
A
Alan30:18
Great. Okay. So the shrimp is basically done.
M
Mark Chen30:21
Awesome.
A
Alan30:22
Do you want to do the flambe thing again to get more color?
M
Mark Chen30:26
Let's do it. I'll turn on the heat a little bit and then let me get some more oil.
A
Alan30:31
Great. Yeah, because we want to get some dark color. Great. And then we can hopefully get another shot.
M
Mark Chen30:43
Same amount as before?
A
Alan30:45
Yeah, same amount. We could hopefully get it because I think the heat wasn't as high. Okay, can you put it in? Nice. There you go.
M
Mark Chen30:56
Just touch the button.
A
Alan30:59
Great. Wow. There it is.
M
Mark Chen31:01
It's good stretch.
A
Alan31:04
Indeed. And it adds some good flavor here.
M
Mark Chen31:08
Yeah. Let's see.
A
Alan31:10
Great. All right. Let's see. Let's go. There it is. We're in the final stretch. We have our shrimp all cooked with some fire. So now we can kind of cook it off a little bit and then we should add our veg to the water.
M
Mark Chen31:32
I'm impressed by your multitasking abilities, you know. I think that's actually one thing we need our models to get better at. It should just be able to do some threading like this and also just have a conversation with people.
A
Alan31:44
Yeah. That also reminds me — do you think images and audio and video and even text should all be under one model? Or do you think it'll break through specific specialized models, like a specialized audio model?
M
Mark Chen31:55
Well, for a research lab, I think there are a lot of advantages for it to being under one. You just have to maintain one infrastructure stack, for instance. The cost of maintaining and scaling many infrastructure stacks at once — I think that's something you shouldn't underestimate. So there are a lot of benefits to just, you do some core research in your fundamental stack and that just carries over to whatever modality or whatever thing you want. So I think there's a strong bias for us to keep it in as few different architectures as possible.
A
Alan32:34
Gotcha. Great. That makes a lot of sense. I think the architecture as well is something that isn't often considered but is very important. One term that I've been seeing a lot that you've also kind of mentioned is vibe researcher. You know, we have vibe coders obviously. But on vibe researching, what do you think is the end state? Do you think the main value out of a vibe researcher is just the research taste of coming up with the right idea, or do you think it's more so the execution of going through and following through on the actual research?
M
Mark Chen33:02
I think we're actually moving towards this world very quickly. Both at OpenAI and at other labs, you're starting to see a lot of the work become mostly orchestration focused. The research is coming up with the ideas, and the model's great enough to do the implementation and execution by itself. So when it comes down to the value of coming up with ideas versus execution — both are still important, but it does feel like there's a market shift towards just being able to come up with a lot of ideas and then the model can actually do the execution and orchestration for you. I think it's very much going to be the future of doing research. We also said earlier, the models don't quite have the taste yet, and that's why you still need the researchers coming up with the ideas. It's going to be hard to teach the models good taste. We noticed that. But in terms of actually accelerating the research, there's clear tangible benefits already.
A
Alan34:09
Do you think there'll ever be parity in terms of research taste with models?
M
Mark Chen34:14
I think so. When we look at our three-year roadmap, the end goal that we want to reach is one where the models are just doing end-to-end research. And I think a part of that problem is just being able to have the model come up with good taste. You point at some generic benchmark or something and it finds the right solutions.
A
Alan34:34
Yeah. No, that's helpful. And in terms of research done by humans at OpenAI, how do you guys go about the postmortem process of, let's say, a research bet that didn't turn out well? I assume that a lot of it is taking these best bets and some don't turn out well.
M
Mark Chen34:50
Well, I would say that is a big part of OpenAI's alpha. Because I think one thing that differentiates us from other labs is we take a lot of high-risk bets. I think it's what's allowed us to stay at the frontier so consistently over time. But it also means that some of the bets are not going to pan out. And a hard corollary of that is when a bet doesn't pan out, you have to not delude yourself into thinking that this is something that will work and kind of disconnect from it. So I think there are certain calls you have to make — kind of look back and be like, well, this was a promising idea at the time but actually it's less important than we thought, there's some other approach that works better, or there's something else that we discovered. But much of that work is also very fruitful. What we realized is even sometimes when people fail at proving out a technique, their writeups are very important because it will often be a natural idea and you can save a lot of people from going through the same thing.
A
Alan35:58
Yeah. No, that's helpful. So I guess when it comes to this positive view on failure, how do you balance that with, you know, a researcher who takes a lot of bets, consecutive bets, and none of them pan out? Because I assume at a certain point you'd want a researcher to eventually have contributions that are actually beneficial compared to only taking bets that maybe pan out to being not a good space.
M
Mark Chen36:20
Just through experience, I've definitely seen some people fall into this. But I've also had several cases where it's just bet after bet, it doesn't pan out, and just when you're at the brink of frustration, you have something that's like a mega hit. And this happened enough. So it really just depends on are the ideas themselves sound. They can be ambitious, but they still have to be sound. And there's a certain kind of person who will just take a lot of those ideas and it's okay because they're somewhat on the riskier frontier, but they only have to justify it once in a while for it to make sense. Maybe like a very trading-like lens on the world, but yeah, it's just on expectation — they need to add value.
A
Alan37:05
Yeah. No, that's great. Okay, so we're basically assembled. Now it's the finishing touches. So you can taste your soup and then we just add soy sauce if it's not salty enough. And if it's too salty, we can add some water. We can lower this down. Let's see our final creation. How is it?
M
Mark Chen37:24
It's pretty good.
A
Alan37:25
Good. Healthy. Okay. Mine's good. Mine needs a little bit of water. Could you pass water?
M
Mark Chen37:30
Absolutely. Look great. So how was that? How did that feel? This is a student distillation. You're clearly better than I am at this point.
A
Alan37:39
No, no, no. I feel like you did a very great job, especially even with the shrimp and flaming.
M
Mark Chen37:45
Yeah. Smells great.
A
Alan37:50
Wow. Okay, that sounds good. I guess just more generally, I'm kind of curious. Are there areas in research or topics that you think are right now overrated and underrated? Like what would you categorize under?
M
Mark Chen38:02
Well, I think if you still have a 'pre-training is dead' view of the world, I think pre-training is definitely not dead. It's underrated. And honestly, I think products and kind of thinking about end uses and how you tie all the primitives you build in research to real agentic use cases in the world — that's also underrated. I think you really can't just build everything in a vacuum and not connect things to utility.
A
Alan38:37
Yeah. No, that's great. Great. Awesome. I think we are ready to taste. So, do we want to give it a go?
M
Mark Chen38:42
Let's do it.
A
Alan38:43
We can move this. And I think there should be some plating. Yeah, I do want to take those plates over here. Okay, shrimp looks great. Everything else here. Great. And then we could just use these pots.
M
Mark Chen38:59
So, we have our shrimp.
A
Alan39:00
Yep. You want to try it? Cheers.
M
Mark Chen39:02
Cheers. Let's see. It may be a little sweet.
A
Alan39:06
That's a little too sweet because we found way too much.
M
Mark Chen39:09
It's good for me, I would say. Cheers.
A
Alan39:10
That's good. Okay, great. Sean, do you want to come and taste our tofu soup?
S
Sean39:15
Smells so good. By the way, you guys can't smell it, but um...
A
Alan39:18
We'll pretend you're a researcher that's trying to approach, and I'm Zuck and he's trying to get you. So, do you want to grab the spoon over here?
M
Mark Chen39:26
Soup. Soup is really going to sway a decision here.
S
Sean39:30
Quality of soup.
A
Alan39:33
All right. Try both.
S
Sean39:34
All right. What was the artistic direction here?
M
Mark Chen39:38
Artistic direction? Well, it was mimicry, I think. There's great art in mimicry.
A
Alan39:43
Yeah. Just letting them cook.
S
Sean39:46
Good. Wow. Yeah. Strong. I mean, I think one thing is like savory and spice that go together, but then also like the sort of seafood-ness kind of really goes into it. Great.
A
Alan40:02
Okay. It's fine. Mine is mine.
S
Sean40:04
Am I supposed to pick a winner or what?
A
Alan40:05
No, no, no. You're supposed to try it and...
M
Mark Chen40:07
No, you do pick a winner. Yes.
A
Alan40:09
Okay. This is an eval.
M
Mark Chen40:11
Yes, SwitchBench.
A
Alan40:13
External evals.
S
Sean40:14
Okay. I got to say, I feel like there's too much water in this.
A
Alan40:19
I think it's also a pot because this is a big pot.
S
Sean40:23
Okay. I would say I have to go for this. Just like you're our respected guest, but I want to be objective.
M
Mark Chen40:29
Of course. Yes. The density, I think, really adds flavor.
A
Alan40:35
I'll do half the water. Half the water probably.
M
Mark Chen40:39
Okay. Makes solid sense. I mean, I think it's very personal, right?
A
Alan40:41
Yeah. I think it's also very personal taste. You know, even when you do a lot of cooking, taste...
M
Mark Chen40:45
No, no, no. Okay. I know a couple recipes. I follow them to the T. I can't — if you tell me, 'Oh, cook something slightly different,' I'm completely lost.
A
Alan40:56
Right. Oh, I'm not going to lie. I kind of looked up in ChatGPT a couple things beforehand like just as prep.
M
Mark Chen41:06
No worries.
A
Alan41:06
But yeah, it was great having you. I feel like you're always leading the field with a lot of research tastes as well and it's great seeing the work. So hopefully it was fun.
M
Mark Chen41:15
A lot of fun. Yeah.