Back
Mark Chen
Chief Research Officer, OpenAI

Mark Chen: GPT-5, Open-Source, Agents, Future of OpenAI, and more!

🎥 Aug 18, 2025 📺 Matthew Berman ⏱ 40m
Download (GPT-5 UPDATED) Humanities Last Prompt Engineering Guide (free) http://bit.ly/4m76knm Join My Newsletter for ...
Watch on YouTube

About Mark Chen

Mark Chen, Chief Research Officer at OpenAI, appeared on the podcast "Let Them Cook" on June 25, 2026, where he discussed the company's research direction. Chen stated that he "firmly believes in being on the exponential and in scaling laws" and said he "fairly strongly disagrees" with views that pre-training is dead, noting that such narratives have recurred throughout the history of developing large language models. He described reasoning as "one of the biggest examples" of a research bet at OpenAI, referencing the o1 model as a breakthrough that was difficult to get off the ground because the pre-training plus post-training paradigm "felt like such a promising paradigm" at the time. Chen also said the field is in an "evals crisis," with a low number of canonical gold standard benchmarks, and noted that tools like Codex have enabled faster iteration of evaluations. On June 16, 2026, Chen appeared alongside SoftBank CEO Masayoshi Son at an event in Tokyo where SoftBank announced a cybersecurity service using OpenAI's technology. Chen described cyber capabilities as a "dual-use capability," stating that "even though the models can get better and better at finding vulnerabilities, we can use that for defense." He added that "the important thing is we can go and try the models and try to find the vulnerabilities before external actors can go and try to find the same vulnerabilities." Son compared the dynamic to a criminal with a knife facing a police officer with a gun, saying defenders must have the "best, most powerful weapon" to defend against bad actors.

Source: AI-verified profile updated from Mark Chen's recent appearances. Browse all interviews →

Transcript (95 segments)
M
Matt0:00
There's a lot of excitement leading into GPT-5. I'm wondering what the energy is like internally at OpenAI leading up to such a big launch.
M
Mark Chen0:07
We have a research roadmap. In the last several years, this roadmap hasn't changed much at all.
M
Matt0:13
Interesting. It's able to cut the thinking time that it needs to extract the information by several factors.
You are head of research. So, how do you personally balance that tension between the product side of the organization and the research side?
M
Mark Chen0:27
One thing that we're really excited about is to kind of raise the bar in terms of what we consider acceptable for open source.
M
Matt0:34
Let's talk about some of the lessons that you learned from GPT-4.
M
Mark Chen0:38
And GPT-5 is one of the first models we have that marries both the pre-training paradigm and the reasoning paradigm together. We do believe that this technology is going to raise the quality of life for most people.
M
Matt0:51
Mark, thanks for joining me today. Good to see you again.
M
Mark Chen0:54
Great to see you, Matt.
M
Matt0:56
So I want to start with a broad question which is there's a lot of excitement leading into GPT-5. I'm wondering what the energy is like internally at OpenAI in the months and then more recently weeks leading up to such a big launch, such an important model. What is that energy like? What is it like inside of the company?
M
Mark Chen1:14
Every single launch, you know, it comes with high emotion, right? And you have that feeling when you're starting out to do a project where people are excited and then you have a period in the middle where there's always this kind of internal uncertainty, right? Is this model going to be good? Is it going to kind of hit the expectations? And then kind of really seeing kind of near the home stretch everything come together. The energy picks back up and I would say right now the feelings are very strong. People are excited to get this model out and yeah, we're excited to show it to the world.
M
Matt1:43
Yeah, I've tested it and it is absolutely incredible. And I'm really just starting to figure out where it really excels, where its personality and the tone and everything are. So, but I visited you guys a few weeks ago and Greg Brockman had mentioned that OpenAI still sees itself as a research lab and you are head of research. So how do you personally balance that kind of tension between the product side of the organization and the research side of the organization?
M
Mark Chen2:19
Yeah. I mean I think it's important to take stock of, you know, we're here to do research and the research really is the product, right? Every time we make a big breakthrough, that is something that will lead to a lot of value and utility for people and of course the product itself enables us to do more research, right? So I think these two things go hand in hand. It's a delicate balance, you really don't have one without the other and that's kind of the view that I hold in my head. We want our research to have contact with the world. We want people to be able to experience kind of all this intelligence that we're building and we're very lucky that it's resulted in a very successful product too.
M
Matt2:54
Let's talk about some of the lessons that you learned from GPT-4 and how you applied it to GPT-5. So from an outsider's perspective, there doesn't seem to be a large amount of new publicly available data that can be applied to a brand new foundation model. Is that a correct assumption? And if so, how did you solve that data scarcity problem?
M
Mark Chen3:19
Right. So I would say it's somewhat accurate but not fully accurate. You can always improve and expand the envelope of the data that you put into the model. So we're always looking for more sources of publicly available data, more sources where we can license data from. But I do think there's a couple axes by which GPT-4 and GPT-5 differ. So when you look at GPT-4, it's really this culmination of scaling the pre-training paradigm. And GPT-5 is one of the first models we have that marries both the pre-training paradigm and the reasoning paradigm together. So it really heavily leans on our O series of models as well. And one of our philosophies here is we want people to get reasoning when they need it and get a very fast response when that's the more appropriate mode to be in. And they don't have to kind of pick, hey, do I need reasoning for this? Do I not need reasoning for this? And back to the question of data. One thing that we've also been exploring in GPT-5 is the use of synthetic data. And this is data that isn't written by humans, but is generated by models. And we've had a very healthy synthetic data program led by one of our researchers, Sebastian Bubeck. And, you know, it's bearing fruits and it helps us improve coverage in areas that we really want to shine.
M
Matt4:37
Yeah. Okay. So, I was going to ask you all about synthetic data. So, let's skip to that. One of my questions was, are you using synthetic data? Obviously, the answer is yes. There are some folks in the industry who say synthetic data and models trained on synthetic data from prior generations of models can only really be marginally better than their predecessor. What do you think about that? What's your take there? So how do you think about the synthetic data in future generations?
M
Mark Chen5:08
Right. I mean, we really do believe in the potential for synthetic data to be, you know, higher quality and to improve the model in meaningful ways beyond just kind of really broadening or kind of deepening surface level knowledge in a particular category. Of course, this is a research program we're still pursuing. I think it still has a lot of room to go. But I think we've seen enough signs of life that we've decided to use some of that to power GPT-5.
M
Matt5:36
And are you able to speak to the mix of synthetic versus human data for GPT-5 and maybe how that compares to the mix for GPT-4?
M
Mark Chen5:47
I do think the exact mix is something that, you know, we want to keep to ourselves but over time it's becoming more and more.
M
Matt5:55
Are there any categories of knowledge where synthetic data really excels? Are there any categories of knowledge like math, science, coding? Those seem like maybe the obvious ones, but you tell me like where does synthetic data really work and maybe where does it fall short, right?
M
Mark Chen6:12
Right. And I think fundamentally I believe in the promise of synthetic data in a very broad-based setting. And it's really kind of up to us to choose which domains that we really want to kind of unleash our set of tools on. You know, we care a lot about code in the GPT-5 release and certainly that's one of the areas that we've emphasized but it's by no means the only way that we view synthetic data.
M
Matt6:38
Okay. And just to continue on that, are there areas that synthetic data have fallen short for you or that you maybe wouldn't apply synthetic data to?
M
Mark Chen6:48
I don't think so. And I mean, there's certainly areas that are more amenable, but I don't think that we, you know, consider the techniques in synthetic data to be not general in a deep sense.
M
Matt7:00
So as you know, let's like rewind months, maybe even years when you start thinking about the architecture for GPT-5. What are some of the early bets that you made that maybe at the time you were a little bit nervous that hadn't been proven yet and ended up really working out with a big model like this?
M
Mark Chen7:17
Right. It's going to be the coming together of a lot of advances in architecture, in optimization, in reasoning and even just kind of in the core infrastructure that we're building. So, you know, we have teams that are exploratory teams in all of these domains, right? Exploratory architecture teams, exploratory optimization teams and in the early phases of a project like this, you know, they're coming up with their scattershot set of ideas and over time we refine those to the bets that are really working. And it's nice to kind of see that kind of winnowing of the ideas and the refinement of some of these ideas and you integrate them all together and it creates this thing that's really kind of a combination of a lot of these innovations on all of these axes.
M
Matt8:04
Were there any early bets that really surprised you though? Like you didn't know it was going to work or you had some doubts and then you just saw incredible results from it?
M
Mark Chen8:15
Yeah. I mean, one thing that I want to stress is kind of the coming together of reasoning and pre-training here, and that might sound pretty obvious on the surface, right? Like, why can't you just kind of get the best of both worlds very easily? There's actually a lot of work done by our post-training team led by Max Schwarzer. And, you know, it took a lot of work to make these reasoning models a lot faster, a lot more robust, a lot more reliable so that we could be able to combine both of these two paradigms that we've been working on somewhat independently into one surface that, you know, people can access the best of both worlds with.
M
Matt8:51
When you're thinking about where to invest your compute in pre-training versus RL, how do you think about that mixture as a decision of really compute investment and really monetary investment because of compute?
M
Mark Chen9:06
Right. So I mean today we invest heavily in both, right? And it's really a function of on the RL side it's a new paradigm, there's a lot of very promising work, a lot of kind of like GPT-2 era style work of just exploring all of the different things that influence RL and it feels very early, right? And so there's a lot of things to explore but at the same time on the pre-training side there's also a lot of energy right now, right? There's the work in synthetic data like what I just talked about and I think still a lot of healthy work in optimization and architecture.
M
Matt9:40
As you're training the model, as you're doing post-training on the model, how do you decide when a snapshot is like, hey, this is the one we're going with? Because obviously you're going to continue to improve it. That's what you guys do. But like how do you decide this is the one we're putting out? It's ready. It's baked. And we'll wait for subsequent iterations for the next versions.
M
Mark Chen10:02
Yeah. Well, I mean, I think it's a little bit of an art, right? And you want to strike this balance of, you know, you want to pursue perfection. You want to pursue something that, you know, it doesn't have kind of any kind of nitpicks or any, you know, you play around with the models a lot. It has to pass the vibe check. And it shouldn't have any kind of small pathologies or, you know, like any behavior that you're worried about. And you have to toe the line of kind of, you know, waiting too long and kind of training what you consider the perfect model on all axes and something that just like feels good and ready to deploy. So I think we've struck a good balance here and the post-training team does a great job of kind of vibe checking all the models.
M
Matt10:43
All right. All right. You mentioned vibe check. What is Mark Chen's vibe check? How do you make sure that the model is good for you, for your personal life, for your work life? What does your vibe test look like?
M
Mark Chen10:54
Yeah, I mean, I check it on a couple different axes. There are a couple of math problems that are my go-to math problems.
M
Matt11:01
Can you share?
M
Mark Chen11:02
Oh, sure. Sure. So there's a fun one that, I think I shared on Twitter a while back. But, it has this very simple description. Basically, you want to create a random number generator, uniform random number generator modulo 42 and you have access to random number generators modulo every single prime less than 42. And this is actually still a relatively unknown problem. A lot of models I think still can't get this consistently right. But you can see a progression of solutions until you get to the most optimal solution here. And you can kind of see improvement in models in terms of how creative they are in being able to generate steps towards the best solution here.
M
Matt11:46
Yeah. Yes. Like I understand those cutting edge math problems, the frontier math problems, the ones that are really hard. Do you have anything that is more like day-to-day tests?
M
Mark Chen11:57
So I think I have two others, you know, one are just kind of generating user interfaces, right? A lot of kind of physical simulations that are very visual, right? You can quickly get a test for like does the physics look right, right? I think you see classic examples of this of, you know, balls bouncing around hexagons, but you can also do kind of like fluid simulations or other kind of user interfaces. And that gives you a sense for, you know, how powerful is the model? How good is the underlying physics? And how good is it at kind of generating robust code and aesthetic code. Another thing that I like to test it with is creative writing. You know, I write a lot day-to-day as well. And yeah, sometimes, probably one of the biggest use cases I use personally in my day-to-day is having it comment on, revise, you know, as a thought partner in documents that I write. And so we really just kind of like does it have an intuitive grasp for what good style is, you know, is it persuasive? Is it compelling? Those are kind of the things that I kind of test for.
M
Matt13:03
Did you see a noticeable improvement from GPT-4 to GPT-5 for creative writing?
M
Mark Chen13:08
Yeah, I do and I think other users will notice that, too.
M
Matt13:12
Do you ever test humor, comedy?
M
Mark Chen13:14
I always find the models struggle with that. So, that's like kind of the golden benchmark for me.
M
Matt13:20
Yeah. I mean, I think there was a time in GPT-2, GPT-3 era where the models weren't generating that realistic human text, and it would be funny in this kind of like, you know, it's just a little bit off. But, I think now that, you know, we're in much more realistic text land. It might take a little bit of time before the models become funny again.
M
Mark Chen13:44
Yeah. A lot of dad jokes I've noticed. Yeah. I think it's getting better. The reasoning models are surprisingly good at humor. I do think that ultimately could be a good test of reasoning, right? Can you kind of like figure out deeply what makes something funny?
M
Matt13:59
One thing that I use models for frequently is getting life advice and I don't know if you do this as well, but I, you know, some complex situation in my life where there's different kind of angles I need to think about it on and it's kind of a difficult decision. It just helps me understand all of those angles. Helps me understand maybe something I hadn't thought of, a new way to approach the problem. Do you ever use it for that?
M
Mark Chen14:24
Yeah. No, I love it as a brainstorming partner. And yeah, it's, you know, when you need advice on a situation, when you have some kind of complex problem where, you know, there's pros and cons of every single kind of decision you could take. Sometimes I use it as kind of something you can confide in, right? And it often times provides some interesting perspectives.
M
Matt14:45
You know, over the last few months, even back to the beginning of the year, China has open sourced a number of incredible models. They've made a lot of innovations on efficiency. Did you take any lessons or techniques from those open source models, open source papers as they were coming out and apply them to GPT-5 or like how has it adjusted your thinking and your approach to your own internal models?
M
Mark Chen15:11
Right. So I want to, one thing that I'm very proud of at OpenAI is that we have a research roadmap. It's something that we put a lot of effort into and it scopes out kind of what path that we want to take towards AGI and it prescribes a couple of things you want to do in the short and medium term but also in the long term. And in the last several years this roadmap hasn't changed much at all and it still hasn't changed. Yeah. Yeah. Even in light of model releases from, you know, from DeepSeek and others. And I think that's really one of the strengths at OpenAI, right? We have conviction in the path that we want to take, we know what kind of ideas it implies and we just pursue those ideas and we're not a very reactionary company in our research roadmap. And at the same time, you know, Chinese labs have been doing really great work too, right? Like I think DeepSeek in particular, they do phenomenal architecture research. They write very, very efficient kernels and I think there's kind of lessons you can learn from that. But by and large the research roadmap hasn't changed and we are still executing on kind of the plan that we had 6 months ago, a year ago.
M
Matt16:21
Yeah, that's actually kind of surprising to hear. So I assume the research roadmap is more broad categories of achievements you're looking to make like a certain amount of efficiency or are they specific techniques that you're looking to test and implement?
M
Mark Chen16:36
It really did inform kind of the development of reasoning models to begin with. And, you know, this is really one of I think OpenAI's proudest achievements over the last year. Just the ability to create models that can think deeply, to reason for longer and longer periods of time before coming up with very intelligent answers. We think that's a very core part of developing general intelligence. And, you know, that's something that we've been pushing on and kind of developing independently of what's going on outside.
M
Matt17:07
Okay. All right. So let's talk about GPT-5 a little bit more. And as it compares to GPT-4, were there any emergent capabilities with the new architecture, the new training flow that surprised you as you were starting to get towards the end of GPT-5 post-training?
M
Mark Chen17:25
Right. Again, I think a lot of these things will take some time to tell, right? When we launch a new model into the world, it takes a lot of experimentation and, you know, the world discovers a lot of these things for us. One of the things that, you know, you can notice right off the bat is the coding is much better, right? Like people prefer the coding capabilities of GPT-5 with, you know, I think 70 plus percent win rate over our previous set of models like 4o and o3. So you really see that step up in terms of coding capabilities and so I do think developers, right, they're going to notice the difference here and it's just much more robust code. It's a model that, you know, hallucinates a lot less. It's more reliable. You can trust the reasoning for longer and longer periods of time. It's much better at agentic tool calling. So a lot of these capabilities that, you know, people focus on for knowledge work I think are going to improve and people are going to notice that.
M
Matt18:18
During my testing there were two specific things that I noticed with respect to coding. One, the front end was, it was just able to create much more beautiful front ends, visually appealing front ends. And then it was also outputting much longer code when I was like single turn, much longer code. I think previously when I was using GPT-4o and even the thinking models, it would start to max out around like 500, 600 lines of code, but GPT-5 is certainly pushing well past a thousand lines of code in a single turn. So it's really interesting. Was that a decision? Is there a certain training technique that you needed to elicit longer chunks of code from the model?
M
Mark Chen18:59
Yeah, I mean I think one important thing about GPT-5 is it's tailored for developers as well, right? We care about practical situations which often involve long code bases, large pieces of code that are generated and so we've put a lot of focus into things like long context too, the ability to handle real world long repositories.
M
Matt19:16
Your colleague Noam Brown said he believes the future of AI is likely to be an omni model, one giant model to rule them all versus many smaller more specialized models. So just wondering what's your take on that? Do you believe that? And maybe it's a mixture of both, but I want to hear what you think, Mark.
M
Mark Chen19:40
Funnily enough, I haven't talked to Noam directly about this before. I think it makes a lot of sense, right? Like, I think, you know, having one big brain that's capable of everything. You know, it's going to be able to kind of have sub-modules and learn sub-modules in the right way. But I do think at the same time, you know, one thing that we aspire to build at OpenAI is organizational AI. When you look at the levels of AGI framework, we do kind of think in terms of organizations of AI agents working together to achieve and accomplish high level difficult objectives. And I think it's always been a fascinating question, right? Like do organizations work better or do kind of single entities work better and I think it's still a very active area of research and one that we hope to get more signal on in the near future.
M
Matt20:36
You touched on agents and I want to pull on that thread a bit. I have been talking about scaffolding being such an incredible opportunity for so many developers. But, you know, as you think about a giant or kind of an omni model that's going to start to eat some of that scaffolding, I guess my question is how much improvement, how much headroom for improvement do you think is available for folks who want to build scaffolding on top of GPT-5?
M
Mark Chen21:04
I do think scaffolding will always need to be there in some sense to tailor models to a particular application. But the hope with what we're trying to do with building very general models is that we can kind of chip away at the need to provide very detailed, very complex scaffolding, right? I think a very intelligent model, it should be able to soak in all of that context and just be able to kind of tell what you want to do. And it should be able to thrive even with less information, right? Some of the scaffolding today is built to really engineer around deficiencies in the models. And we really hope that part of the work that we're doing on robustness, on reliability can kind of remove the need for a lot of scaffolding and, you know, allow people to kind of more intuitively tailor models to their applications.
M
Matt21:51
And one element of scaffolding is memory or context management. Do you think we can reach whatever your definition of artificial super intelligence or AGI, whatever these kind of big milestones in AI are, without having memory as internally in the model? Do you think it's possible?
M
Mark Chen22:12
Yeah, I mean I think there's just so much rich work that needs to be done with memory. Like clearly today, right? Like, you know, there's a context window for models and that's one of the big limitations and you would really love the model to be able to maintain memory and context over, you know, a really, really long period of time, right? Like it should be able to fit code bases. It should be able to fit all of your, you know, personal documents and thoughts and maybe even like everything that you see day-to-day, right? Visual signals. So all of these things are going to help the model make better decisions for you to allow you to kind of remove that scaffolding you're talking about, right? And really have the model act on your behalf autonomously without having to ping you all the time. So we believe kind of memory is a huge limitation and a huge thing to overcome to make the models even more useful in the future.
M
Matt23:04
And do you think let's say you had an infinite or like a very, very large context window, is that the kind of the ultimate solution for memory or do you think there's some other architecture where memory is maybe baked more into the core model itself and maybe the weights get updated dynamically?
M
Mark Chen23:23
This is a little bit above my pay grade, but there's a lot of implementation details, right, in terms of how you could do memory and even when you say, hey, if we support a huge long context, is that sufficient? I mean a lot of it rides on how you actually implement that long context, right? You need an architectural primitive that allows you to implement long context that's rich enough to actually integrate all of that past information. And so I think a lot of it comes down to the specific architecture that you develop. And yeah, it's stuff that we're always improving. You want to make the model that much more efficient at both synthesizing and being able to pull from early memory.
M
Matt24:02
Speak for a second about multimodality. What is available today in GPT-5 and then what do you have planned in the coming months if you're able to share about different modalities for inputs?
M
Mark Chen24:14
So we still have the same multimodal feature set as our previous models. You know, they can do images in, they can do audio in. So, they're perceptual models and they're also able to do kind of image generation. I do really kind of think of perception as a core part of intelligence. And one thing that we've highlighted in previous demos in o3 and o4 mini for instance is the ability for a model to take a very complex image, to analyze it, to really find the important parts and extract and kind of foveate around the image and understand what's the most important part for answering the query. GPT-5 is much more efficient at doing the same kind of task. When you give it visual perception tasks, it's able to cut the kind of thinking time that it needs to extract the information by several factors. And that's kind of what we're seeing across the board in terms of GPT-5's reasoning ability. It's just much more efficient. It's faster at getting you the reasoning that you need for your task.
M
Matt25:12
I believe I read you are one of the creators of the original Codex model, the coding model. Yeah. Okay, cool. So by the way when I first saw GitHub Copilot I was absolutely blown away. It was really the first time I'd ever seen something like that. When did you first realize how powerful AI models could be at coding? Was this always something you knew or was there a certain moment in time where you knew it was just incredible at coding because of the amount of data out there?
M
Mark Chen25:39
I think one of the biggest reasons we focused on code when we were producing Codex is we really saw this as a way to accelerating our own work fundamentally because, you know, a lot of the work that we do in research is implement and try out ideas and the medium of that is code, right? And so we all really inspired to work on coding models because that seemed like a way and a fast way to kind of accelerate our own progress. The problem back then is there wasn't any good way to measure progress on code generation back then and part of developing Codex was also developing this evaluation mechanism for figuring out how do you compare code models, like what are the right benchmarks. I think today, you know, it's much more mature and we've also gone from, you know, the problems that we could solve with original Codex which were very basic, you know, five lines of code to the IOI problems that we're solving today, right? Which are very complicated, thousand line programs, involve a lot of creativity. So I think the models have just come so far and, you know, we still believe in them as really this vector for pushing scientific progress forward as well as helping accelerate ourselves.
M
Matt26:54
Yeah. I mean like for all of the frontier labs, coding seems to be such an emphasis, such a focus. Is that because coding and even as an extension math are the keys to more capable all-purpose models? Is that the piece of it that allows them to do reasoning at a much higher level?
M
Mark Chen27:16
Yeah, I mean I think there are a lot of different ways to teach reasoning, right? Mathematics is also a very efficient way to learn reasoning. You know, physics, there are a lot of domains that require a lot of deep reasoning, even stuff like creative writing, right? Or telling jokes like you said. But, you know, we do think of coding as specifically aligned because so much of our day-to-day involves coding and also so much of the value that I think people give to the world today through technology is through code as well. So it's definitely an area of strategic importance and one that I think a lot of us internally benefit from.
M
Matt27:52
It's so scalable because it does have verifiable outputs. And so I want to talk about verifiers for a little bit. As part of GPT-5's training, outside of more verifiable domains like STEM, were you able to figure out ways to verify things like creative writing or humor as we talked about, less or more subjective domains?
M
Mark Chen28:13
Yeah, I mean I think that's always a big part of our research program. We're trying to make RL more generalizable and one way to do that is try to find ways to make verifiers more generalizable as well. You know, kind of any RL system is going to require verifiers and we're trying to make our RL systems as general as possible.
M
Matt28:32
Can you share a little bit about how you thought about verifiers for GPT-5? Are you using LLM as a judge or like anything you can share?
M
Mark Chen28:41
It's a mixture of a lot of approaches. Unfortunately, I think we'll share that at a later date.
M
Matt28:46
Is there anything you want to cover, Mark? I have a couple more questions, but I want to make sure like if there's something especially cool that you're excited to talk about, we get there.
M
Mark Chen28:54
It's really just been an exciting month for us, right? I think we're coming off the heels of not just GPT-5 release, but there's also the open source release. And, you know, there we're just really excited that we've been able to squeeze all this capability into such a small model that's accessible for people who, you know, want to run models on prem. And then kind of even before that, we've had some really good results at competitions like AtCoder and the IMO. And it's really kind of just confirmation that these broad-based reasoning models are able to achieve stronger and stronger results, right? We're continuing to raise the ceiling. And for me, the AtCoder result stands out just because it's the first top three result at a world-class tier kind of a math or programming competition. And I think there's just a world of difference between your like top 100 and top three. So really kind of happy with the progress that the team is making pushing into that top tier level of human level performance.
M
Matt29:52
And so like how does achieving such good performance at IMO or at AtCoder, how does that translate into more everyday use intelligence where people are actually going to see a difference?
M
Mark Chen30:04
Yeah. That's a really, really great question. And when we approached these two competitions, we didn't set out to create like here's an IMO specific model that we're going to pump a ton of compute into and all it can do is the IMO, right? We're creating fairly general models. We take general techniques, layer them on top of our reasoning models and with a very small amount of tuning if any, we create models that can really kind of perform at this level at the IMO and AtCoder. And I think, you know, some of the models that we use for math competitions, it's the same models that we're using in the coding competitions. So they really aren't these like very narrow specific models. We care about a program of producing general intelligence. And so I think what you're seeing is what you're getting, right? These models are very good at these contests, but also you're just very good generally, right? We haven't done too much tuning for them.
M
Matt30:56
Yeah. I mean that kind of speaks to Noam and yourself, what you thought about having kind of a generalized omni model in the future if it's good at all of these different types of very difficult tasks and then those things are all inside one model and then they can apply to more everyday use. So definitely lines up with what you were speaking about.
M
Mark Chen31:16
Absolutely.
M
Matt31:18
I mean let's talk a little bit about the open source models that were just released. Tell me why you are so excited about these two flavors of models, 20B, 120B. Why are you so excited and how do they help progress OpenAI's mission?
M
Mark Chen31:36
Yeah, so I think in terms of just the sizes of the models, right, we're happy to deliver models which can run on a laptop and run on a phone, right? These are very, you know, very commonplace kind of consumer form factors and it's really just going to allow a lot of hobbyists to become more involved. People with kind of very specialized constraints, academics, right? They're going to be able to access models which they have full control over. And I think that's going to be very powerful. And in terms of how it advances our goals, one thing that we're really excited about is to kind of raise the bar in terms of what we consider acceptable for release for open source. And we've done a lot of extensive safety work on this model. What I mean here is, you know, we have a preparedness framework which allows us to think about risk for models like, you know, is it risky from a bio perspective or from a chemical perspective or from a cyber security perspective and what we've done with these open source models is to just test, right? Like if some actor decided to fine-tune one of these models to be maximally dangerous in terms of like cyber security or being able to attack, you know, code offensively, you know, what is kind of the ceiling there, right? And we're making sure that it passes under the bar of what we consider high levels of danger before we release this model. So I think it's really a good chance to set norms for what is safe and responsible to release in open source. At the same time, right, we're releasing a frontier class open source model.
M
Matt33:12
Yeah, no, I'm super excited to test it out. I obviously read all about it already. What went into the decision of the total number of parameters? I think obviously it's what can fit on consumer grade hardware but then also the active parameters. How did you decide what that ratio looks like?
M
Mark Chen33:31
Yeah, I mean I think really it is just motivated by the surfaces and by latency, like, yeah, we just want it to be very easy for people to run. And I think that really just drove all of the decisions in the open source model. We've engaged with developers very early on in the process. We've taken a lot of input even starting from I think 2 or 3 months ago and I think a lot of those conversations were very helpful in terms of parameterizing the model.
M
Matt33:59
The benchmarks are very impressive, right? Comparable to o3 mini, sometimes o4 mini. Let's talk about benchmarks a little bit. You know, you spoke about GPT-5 or a model getting close to winning AtCoder second place, getting gold medal at IMO. How do you think about benchmarks going forward when many of these benchmarks are saturated, but you have things like the ARC-AGI benchmark, which is still pretty far off. What are your favorite benchmarks? How do you think about benchmarks and how these benchmarks kind of tell the story of the model capabilities?
M
Mark Chen34:34
You're absolutely right that there is a bit of a crisis when it comes to benchmarks today, right? And what I mean is that we used to be able to take hard benchmarks written by humans and just use them to test the models as well. And, you know, as the models get really to that frontier of what humans can do, the human benchmarks are not going to cut it anymore, right? Like I think like today, no one will be surprised if it could get a very good score on the SAT anymore, right? And so now you've had to develop a lot of new types of benchmarks just to test the frontier capabilities of the models. And the problem is that there's no kind of standard or agreed upon way to do this, right? People are putting it together in the last just couple of years. We do a lot of that work ourselves too, right? And I think over time some of these things will become more and more standard. But the danger too is that, you know, they tend to have fairly short shelf lives these days. When you see kind of any benchmark that starts with, you know, just a couple percentage points of accuracy, often times within a year, you know, you're already getting into like 30 to 50%. And so I do think kind of the shelf life of benchmarks has really reduced quite a bit and we spend a lot of active effort in just creating orthogonal new types of benchmarks to challenge our models.
M
Matt35:53
Yeah. And what do you think about interactive benchmarks like ARC-AGI-3? Right. It's kind of an interactive game benchmark and what do you think of these benchmarks that are more set up as gaming environments for the AI to play within?
M
Mark Chen36:07
Yeah, I mean it's certainly a good test of our models and I think there's a danger in just indexing too much in one specific benchmark. What we try to do is just develop these broad-based reasoning capabilities and then use all of these specific kind of thermometers to gauge whether we're making progress in a broad-based way. So I wouldn't say our approach really takes one benchmark as front and center and optimizes that but rather, you know, we try to get a pulse on, you know, what are all signal-bearing benchmarks out there, design more ourselves and just make sure that those are all kind of small signals that we're pushing in broad-based reasoning in the right way.
M
Matt36:44
I have a lot of developers and engineers who watch my videos and a lot of them are quite nervous as AI continues to get better, right? Second place at the AtCoder competition. What do you tell let's say new graduates or students that are thinking about getting into coding as a career? What do you tell them? Are you, I assume you're optimistic, but don't let me put words in your mouth. What do you tell them to make them feel more confident about that as a career choice?
M
Mark Chen37:15
Yeah, I mean I would say just lean into using the tools to accelerate yourself. That's what we think about at OpenAI as well, right? We ultimately want to build AI that can accelerate research. And I think part of that is just accelerating your own productivity, right? If you learn how to interface with the tools, if you learn how to make yourself 2x, 3x more effective, there's still a lot of value you bring with your ideas and with learning deeply how the technology works. So really lean into that, understand the mechanics, be able to contribute to the mechanics and just accelerate yourself. I think that's probably one of the most important things going forward.
M
Matt37:49
And then more generally, let's say not coding but knowledge work, there again like a lot of people are very fearful of being automated away. Do you apply that same recommendation to general knowledge workers, just learn the tools and if so why?
M
Mark Chen38:04
I think so. And it's always because I feel like the model's going to automate some surface, but it creates new surface for us to adapt to, right? We're not going to lie, you know, this, we think this is going to transform the economy in significant ways, but we also believe that humans are very adaptable, right? We've always done it in the past. There have been very powerful productivity tools in the past. And I think the more that we adapt to them, we create new surface for us to work in. And I think ultimately, you know, we do believe that this technology is going to raise both through scientific progress and just through economic means, right, the quality of life for most people. And I think that's one of the things that motivates all of us here.
M
Matt38:47
Yeah, I love that. I love the very optimistic view. I'm very optimistic as well. Okay. So let's end on this question. What are you personally most excited about over the next 6 months and then let's say over the next 24 months?
M
Mark Chen39:01
I would say in 6 months, I'm really excited to just continue the reasoning scaling paradigm, right? I think there's so many different ways that we see pumping more test time compute into the models and to have the models leverage test time compute effectively. There's a lot of, you know, different RL objectives that we're trying out. There's a lot of innovations in RL optimization that we're trying out and there's a lot of different ways that we're looking into scaling RL. So, you know, it's still a field that is rich with ideas, that's kind of not mature in the sense that there's just one recipe and we're really excited to kind of figure out all those small details. And I would say in 2 years, I just think, you know, we really can get to the point where these models are just as effective as, you know, I am at doing AI research and I'd love to create this system that, you know, the AI is really kind of driving a lot of the innovations that make future systems more successful.
M
Matt40:06
Yeah, self-improving AI is just such a cool concept to think about and yeah, if it's able to discover new innovations in artificial intelligence, apply it to itself, and then continue to scale up from there, it's really just a function of how much compute you can throw at it.
M
Mark Chen40:21
Absolutely.
M
Matt40:22
Mark, thank you so much for chatting with me today. It's been a pleasure. I really appreciate it.
M
Mark Chen40:28
Yeah. Really, really enjoyed the chat. Thank you so much, Matt.