About Mark Chen
Mark Chen, Chief Research Officer at OpenAI, appeared on the podcast "Let Them Cook" on June 25, 2026, where he discussed the company's research direction. Chen stated that he "firmly believes in being on the exponential and in scaling laws" and said he "fairly strongly disagrees" with views that pre-training is dead, noting that such narratives have recurred throughout the history of developing large language models. He described reasoning as "one of the biggest examples" of a research bet at OpenAI, referencing the o1 model as a breakthrough that was difficult to get off the ground because the pre-training plus post-training paradigm "felt like such a promising paradigm" at the time. Chen also said the field is in an "evals crisis," with a low number of canonical gold standard benchmarks, and noted that tools like Codex have enabled faster iteration of evaluations.
On June 16, 2026, Chen appeared alongside SoftBank CEO Masayoshi Son at an event in Tokyo where SoftBank announced a cybersecurity service using OpenAI's technology. Chen described cyber capabilities as a "dual-use capability," stating that "even though the models can get better and better at finding vulnerabilities, we can use that for defense." He added that "the important thing is we can go and try the models and try to find the vulnerabilities before external actors can go and try to find the same vulnerabilities." Son compared the dynamic to a criminal with a knife facing a police officer with a gun, saying defenders must have the "best, most powerful weapon" to defend against bad actors.
Source: AI-verified profile updated from Mark Chen's recent appearances.
Browse all interviews →
Transcript (78 segments)
U
Unknown0:00
Well, thank you very much. Before we get going, just a massive thank you obviously to the institute for hosting us today. Beautiful space. And also for all of you for turning up. I know you're not here to hear me talk, so I won't talk for too much longer. So really just to say also a massive thank you to you both for coming. It's rare that you get two such great minds in the same place. So we really appreciate the time going into this.
T
Terry Tao0:23
It's our third time.
U
Unknown0:25
Yeah. And a little pattern is starting to build up. And maybe actually that's a good place to start that conversation. You guys had a conversation almost a year ago or so to the day and at the time Terry I think your prognosis for where GBT was for mathematics was something like a very ineffective grad student.
T
Terry Tao0:46
Which remained with me because I'd heard that feedback myself as a human being.
U
Unknown0:50
So it was a clear benchmark.
T
Terry Tao0:51
Guys for Yeah. Okay. Yes. That was my opinion.
U
Unknown0:56
Why don't we start with how you think things have changed since then and then Mark hear your side of the story.
T
Terry Tao1:01
Okay. Yeah. So a lot has happened in the last year. That's not just in AI but yeah so these tools have definitely become a lot more powerful. I think there are now capabilities that basically are now normalized and like we just use them all the time. So deep research tools. Literature search has become really really good. It has surpassed traditional searches. Code generation of course is the big thing but as a pure mathematician I'm not as heavy user of code but it has changed the way I approach a math problem. I will plot something, if there's an inequality I think is true I will ask an AI to try to prove it or disprove it. I already use it. If there's a lemma that I don't think I know how to prove but I just can't be bothered doing the pen and paper calculations. I will just outsource it. I've not yet found it to be useful at the deepest level of when I'm trying to solve a problem and with pen and paper or with a colleague I can't sort of interact with it on a conversational level quite at the level I need yet. But maybe in the future. But I think also socially I think we're beginning to, the mathematician community as a whole is beginning to understand that these tools are here to stay and we have to actually start adapting how we do our research. So certain things that were very tedious and maybe we would force our graduate students to do, you know, we can offload to AI and this opens up lots of new ways to do mathematics, lots of research projects especially at scale that we just could not dream of doing. So while I think we can use AI to assist our current workflows, it's a little bit awkward still to do that, but I think there'll be much more mileage out of creating new workflows which are optimized for AI. It's like when we invented the automobile, we started changing the way we built cities to, you know, and of course you could say that maybe not all the changes were good, but yeah, we're sort of in this intermediate stage where somehow our roads are still for people on horses and we now have automobiles.
U
Unknown3:25
So, would it be fair to say that we've got to the point where occasionally helpful collaborating? But maybe more interesting that all the bigger open spaces, how you change the way you do maths with these tools coming. Mark, would that be true to what you're seeing and what you're building for?
M
Mark Chen3:39
Yeah, honestly, I don't blame Terry for saying it's an ineffective grad student a year ago. I think that's largely the state that we were in back then. And you know I really do think of the backdrop of AI progress as hill climbing this what we call meter plot internally of the models doing autonomous work for longer and longer periods of time. And I think last year we were in the category of minutes and you saw that right it would just the model would hallucinate. It would kind of fall over when you gave it significant chunks of work. But I do think the last year has been a transition for a lot of us in that we've seen the mistakes go down and therefore you can trust the model to do longer periods of work in general and that's you know really kind of allowed us to do away with a lot of the scaffolding that we might have needed to use before and really start to attack you know bigger problems and truly orchestrate with the model and yeah I just think of a year ago we were in the world where we were kind of roughly achieving a bronze medal at the IMO. I think this summer, you know, across all kind of high school mathematics and programming competitions, we are achieving gold medal performance. And I think we've just kind of run out of these human written benchmarks. And that's why you do see people evolving to this sphere of doing mathematical research. And fundamentally, that's always been the goal. We don't find any pride at OpenAI just kind of solving you know IMO problems or anything like that the real ambition is to push the frontier of science and finally the task horizon has caught up to a point where we are actually able to go do that work and again it's not there yet I think the trend and the trajectory is strong but yeah I do think you know it's true that a lot of people are finding utility in it today.
U
Unknown5:30
I mean I'd like to come to maybe first proof and transition as we go into more front-end mathematics, but maybe to stay with the capabilities right now. I think often the Erdos problems are seen as a way of getting a litmus test for where the models are. And that maybe is a representative set and that some of those problems are maybe not as complex as other parts of those problems and were designed that way and that you might say that the success of the models has been in doing lots of the easier ones quickly versus necessarily moving towards the kind of ACON level problems. Is that a fair depiction in your mind of where they are today?
T
Terry Tao6:05
Yeah. So I've been heavily involved in tracking the progress on the Erdos problems in particular. So I mean and it is basically largely true what you say these problems range widely in difficulty. There are some that we desperately want to have solved and they've been worked at for decades. You know I have papers you know making tiny progress on some of these problems. And to date AI has not really helped with the ones that we've already poured a lot of attention to. But there was this very long tail of problems. Erdos posed a thousand problems. They weren't all winners. But he understood that the important thing was to stimulate discussion and interest and he kind of knew that the problems that were going to be important would take on a life of their own. But those are as long unexplored problems where there's maybe almost no follow-up literature. And that's where the AI tools have really made a lot of spectacular progress. You maybe 20 30 of these problems have been solved with fairly minimal human supervision by these AI tools. And we were able to verify them too often with some other AI tools in the form of verification. And we've kind of worked out a kind of a workflow for doing this without being overwhelmed by AI slop incorrect solutions or whatever. So yeah it is a new capability that we hadn't before. Because we can now attack attention bottleneck problems. And so I think what this suggests to me is that we need to start creating more and more broad challenge sets of problems for AI tools also but the general public. So actually in the same period many of these early problems were also solved by amateur mathematicians sometimes with AI tools, sometimes without. The same mechanisms that the same kind of workflows that enable AI to be successful also actually enable amateur mathematicians to be successful as well. So I foresee a change in our culture where instead of only working on a small number of really hard problems and not sharing a longer list of other things we'd care about, we all start as mathematicians, we all start releasing problems of things that we want to get answers to. And 100 problems and maybe this AI can solve 10% of them and maybe this other high school student can solve another 5%. But this we can get a much much more community-driven way of doing mathematics. So I think this is what the only problems I sort of get an early harbinger of.
U
Unknown8:51
Yeah. And it's maybe interesting to contextualize that other domains of science at least in my own world of biology in which the number of people collaborating in any given paper is just exponentially risen over time. And that seems to be the trajectory that science is much more of a team sport. Maths is maybe and to some degree of physics the outlier in that domain. When you're thinking about this Mark is it always just a question of how smart can we make the models and the ever more difficult questions they can answer or is it also a question of how do you empower humans to work collaboratively on these problems?
M
Mark Chen9:25
Yeah, I mean right now we really do see heavy engagement with the community that is a necessary part of driving progress in all of these scientific fields. And Kevin here he runs our OpenAI for science program and part of this is that kind of like you said you know these experiments like first proof or the problems it really is an engagement with the community on figuring out what problems are actually important to tackle. We've done this kind of exercise in physics as well right brought in kind of expert physicists to kind of lay out a program of here are the really important things that feel like they're amenable to AI and that also helps us shape the AI in turn. And it allows us to kind of find the deficiencies, right? We can look at where our models fall over and really shore those things up. What we hope to build is this platform where scientists around the world can just accelerate themselves and we want to empower that community mathematician. We see people like that today empowered you know you have these 20 year old 21-year-old kind of kids using the models to solve some of these problems. It may not be, you know, these sophisticated and very significant leaps, but they're able to do a lot of self-directed work. I kind of had this thought when you were asking the question of I know Terry, you've organized a lot of big community initiatives in math before. I don't know how you think AI changing that world or you know does it enter that world in a significant way.
T
Terry Tao10:50
I think it combines very well actually. So I think what AI will enable is finally a way to use division of labor which is something that like all industry you know since the industrial revolution like every industry has managed to become more efficient through division of labor except mathematics. So you know so to do mathematics traditionally there's several different tasks to you know there's problem generation there's strategy generation and then strategy selection among all the strategies you generated and then execution of the strategy verification of the strategy communication of the results. And so you basically need we've trained our mathematicians to sort of be somewhat good at each of these tasks. So we specialize in field but we have to have some idea where problems come from what are good problems what are good strategies you have we have some technical skill we have to verify and we have to explain now some mathematicians are better at some of these than others and so we have been able to benefit from collaboration because of this. But we can't really specialize the same way that in the sciences you know you can have you kind of have technical staff and you can have people who are project managers and things like that. So but now with AI and other modern collaboration tools and formal verification it has become possible to run math projects where individual participants specialize in just one of these areas and maybe there's some gaps you know among your collaboration no one knows how to do the technical thing but AI can plug in some of the gaps in a collaboration. But you still need the humans because the AI performance is very very jagged. So maybe some of these inputs can now be automated but if you automate that too much you know for example if you can automate strategy generation okay but you can't automate verification then you just get this whole hundreds and hundreds of possible strategies but anyway AI generated that you can't deal with that but if the verification also keeps pace then suddenly you have a new style of doing mathematics that could be extremely effective.
M
Mark Chen13:03
Yeah. Oh just one quick comment on that too. I actually absolutely agree that AI capabilities are super jagged today and so you see this really fruitful collaboration with humans. It's also interesting to explore the flip side of that which is that some of these AI systems are more humanlike than you imagine and you have to like pump a lot of RL in the right way to like not have the problem have the models give up in the same way a human would. You know, if you give a too hard problem, oftentimes the model can just, you know, it'll like run a couple tester, you know, prompts in its own train of thought and be like, ah, you know, this problem's too hard. I don't think I can actually do it. Let me pretend to the user like I tried really hard.
T
Terry Tao13:40
So we've seen with the problems that you get an AI to try to solve a problem. It will first thing you do is go to the problem website, look it up, say, it's an open problem. It's too hard. I'm not going to try. So just say do not use the internet. I try to solve the problem yourself. It's actually pretty easy.
U
Unknown14:00
It's good to know that it's, you know, Frontier Research is actually just about coaxing the models into behaving the way you want to. That vision right now is probably quite a compelling vision for this room and beyond where we're sort of saying the technology fundamentally empowers more people to collaborate on these problems. But is this just a stepping stone to a world, Terry, in which you're only collaborating with many AI agents and slowly but surely they come to dominate the sphere?
T
Terry Tao14:26
I think yes and no. I mean I think the type of math that we do today might slowly kind of move in that direction. But there could be very new types of doing math that we can't even envision right now. Which I think so math is infinite and the difficulty levels are unbounded there even problems in math are unsolvable we know they're unsolvable okay there's an asterisk but I don't want to talk about it okay you know so like there's certain things that even the most powerful AI we can't there's even right now there's certain cryptographic challenges that yeah I mean AI cannot you know mine all the bitcoin right whatever. So I think there will always be a frontier. And I'm pretty sure just because how complimentary human and AI at least current generation LLMs are with human skills that the best combination is always going to be a complex combination of humans. But the nature of the combination may change over time.
U
Unknown15:35
So let's assume even just philosophically there is this frontier beyond which at least the current paradigm of AI wouldn't be able to cross and some centaur like human AI collaboration is needed getting to that frontier in your mind Mark is that a question of much smarter RL training or is it actually just a question of raw computation you know if I could give you an infinite amount of compute today would you be able to accelerate your way to that frontier?
M
Mark Chen16:02
Yeah, I mean I think when I think about the OpenAI research program overall, it is really fundamentally about how do we improve the algorithms such that they scale to the level of compute we have you know next year and the year after. So it's grounded in the reality of what compute we actually have. And I think all the algorithms we know they are simple and they scale but they take a lot of engineering and you know fine-tuning to make sure that they truly scale to the next order of magnitude and the order of magnitude beyond. One really great thing is this is a very multi-dimensional problem today. There are many axes by which we can scale model intelligence. We can, you know, scale up the model, you know, build these bigger brains with just more core knowledge. And this kind of captures the intuition that just like kind of maybe the more math you know broadly, like just internalize deeply, like it's easier to make these connections and these jumps, there's also a reasoning axis which we scale. And this is ability to take all of that base knowledge and chain it together to create new insights. And we have a couple people in this room kind of working on this thing we talked about in the GPT5 livestream which is kind of connecting this a little bit and having the models just generate new knowledge for themselves and really kind of amplify its knowledge in certain domains. So I think there are a lot of different axes that are going to play into bringing the models to the next frontier. Yeah, but overall like all of these things are grounded in this meter plot like we are aggressively hill climbing towards more and more autonomous longer horizon tasks. And we see that trend continuing.
U
Unknown17:36
The term hill climbing, if I've learned anything working with the research team, is that we always must find and define a hill to climb. And perhaps that's where these two worlds come together as defining the right hill. So maybe if we focus on that for the next segment. First proof would seem like an obvious example of where we've tried to codify a hill to climb. In your mind, Terry, is that representative of what you're thinking could be the new emergent math to come or is that the final form of this more classical math that AI has been working on?
T
Terry Tao18:07
There'll be a spectrum. Yeah. So, first proof is a very interesting experiment. And the proofs that the various people with AI tools generated were quite good. What we saw actually that there was a definite verification bottleneck. So we had a lot of proofs generated of some were terrible, some were quite good, some were similar to things in the literature, some were similar to the proofs that the authors themselves had. There was a couple which were actually different from the official proofs. And so that was interesting. But to evaluate carefully exactly how novel and how interesting each proof was like you don't have a way of doing that effectively. So I think that the first proof team are going to create a more structured competition later where they will have some mechanism for verification. So we in order to take full advantage of the new capabilities that AIs have we do need to create challenges that are easily verifiable. So somehow the level of automation and AI power that you can profitably use before it becomes slop is roughly proportional to how stringent your verification is. So yeah so I think initially you're going to see a lot of progress in areas either which are sort of elementary enough that they can be relatively easily formalized so combinatorics I think you're going to see. So the early problems definitely fall in this category there's some numerical type challenges where you want to find a configuration mathematical object that obeys certain properties once you have the objects verifying them are very easy. We saw some examples actually here in physics. There's some similar type problems. I think there we will see a lot of progress. But there are other parts of mathematics where the objective is not to find an object that obeys a certain property but to find a good overarching theory to explain something or a good definition and those we have a lot harder time to verify like you know if you want to propose a new conjecture or a new strategy to answer a problem to you know maybe AI can generate a hundred of these possible strategies but only a human expert can verify or can give an informed opinion. And so that will be a bottleneck. So even if AI drives the cost of sort of creating solutions down to zero there are still other huge bottlenecks which we didn't which were not front and center in our minds. But yeah I think we also need to become very much better at stating goals precisely. So AI is almost too good at fulfilling a goal to the letter. So you ask I want to solve this problem. I want a proof of this theorem. And maybe a future just runs for now. Here's your proof. But actually what you wanted was you wanted people to work hard to fail to find examples to connect to other literature and to communicate all the partial results. And that was actually the value of solving a particular problem. And there is a danger that if you specify your goal to an AI too narrowly, you miss out on most of the benefit. So we'll have to be more careful about goal specification.
M
Mark Chen21:34
Yeah, just two really quick things to add on. I do kind of think of this offline version of first proof. And we were actually discussing this a little bit too where you can imagine that you just train a model with knowledge up to some specific you know very very detailed like this day this time and you can imagine like what a first proof would be at that point in time and now you have the benefit of hindsight. You know kind of what the techniques you're after might be what creativity in the model might look like. And I think those are very interesting thought experiments to run. I think you know there's a thought experiment of what date you would choose as a cutoff to get maximal signal on that experiment. But yeah I also do kind of think about yeah just you know the process of mathematics isn't just answering or proving a theorem it's all this partial progress that you assimilate somewhere and I do kind of think about you know we have AI systems at OpenAI where they're kind of just like central repositories for information. You could imagine that kind of serving a function in mathematics as well. Like you have just this kind of global library in some sense as I know I think Daniel Litt published something online a while ago kind of it just is this agent that mathematicians can interface with and it kind of like fills out this convex hull of mathematical results and you know you can always kind of use it as a source of truth for like what people are exploring and it'll kind of connect a lot of the dots for you and yeah just be this case which stores what we know.
T
Terry Tao23:05
Yeah. Right. Although it may sometimes be useful to turn that off, you know. So I mean I've worked sometimes on a problem where I know too many techniques. Okay. And there's a powerful technique I know will solve the problem, but it requires a lot of technical skill to use and so I do it and I solve my problem and then I publish or something and then someone points out actually you had to use this much simpler tool. You have a much simpler proof. Yes. I do worry a little bit that sometimes having access to every single technique known in the literature is not necessarily the best way forward. But having a diverse array like multiple AI tools and I think there will still be people who will take pleasure and pride in sort of doing things old school and finding sort of more human ways to solve problems too.
U
Unknown23:53
Well I wonder if that pattern even plays the case study you gave Mark. Where it's true if we could go back in time just before a particular paradigm shift in whatever domain of science and then see whether or not the model would predict it that could be one verification tool. But I guess the kind of Kuhnian vision of paradigm shifts could also mean that in fact there is a future paradigm shift to come that would invalidate the prior paradigm shift. So you don't actually want the model to guess the previous one because it might take it off a pathway that doesn't get to the next one as you go. And it throws into relief I think this question of verification validation it's both a philosophical and a practical question of all the domains you could argue that maths and disclaimer here I did work a little bit on lean so big shout out to that team has the actual capacity to do automated verification in a way that very few other domains do not perfectly and not without its own drawbacks. Is your instinct that that structure where there'll need to be a sort of separate validation tool will need to come into existence for all the other domains of knowledge that we want to work on that it will mirror what's happened in maths or it will need some other type of paradigm.
T
Terry Tao25:08
I as I said I definitely believe that there's an upper bound on how much AI you can inject into a workflow before it becomes a net loss that it is causing more errors and problems than it is solving. And one of the biggest upper bounds is the ability to verify. So yeah so in math I think we have the best shot at getting really high levels of automation of being able to effectively use excuse me high levels of automation in a way that you couldn't do in less verifiable domains because we have a high verification bar at least for the specific task of proving things which is not the only thing that we care about but and proving things that we've already specified we want to prove. Yeah but although even for verification does have weaknesses the language itself can be exploited by malicious agents. So and so an AI may sort of you know attempt to be helpful and try to prove as many things as possible just secretly add some axioms to the formal system and things you can try to shut them down but if the AI is too powerful actually yeah I mean at some point you actually have to sort of limit how capable your AI is or you know or have periodically humans involved in the loop. So you know there are other in the other sciences you can do some of this. So for example numerical simulation can be used as a verifier in some cases but again you can't rely on like if you say what you want to model the weather and you have a supercomputer that predicts the weather and you have trained an AI to mimic the numerical simulation it is possible that at some point they will just exploit some feature of the numerical simulation that is not part of the ground truth so it will work up to a point and then it will stop. So we do need to get a lot better at knowing the limits of our verifiers. So you know a lot of verification systems that we have, they work just fine if they're used non-adversarily. But you know if you're training an AI specifically to maximize output based on using this verifier, it will find the exploits. AI is so good at that. Yeah, it's a ruthless cheater. So we do have to be aware of that and just because a human verifier surpasses all human tests it may not be suitable for AI use.
U
Unknown27:40
That makes a lot of sense and intuitively to AI cheating the easiest way to make something measurable is to design it to be measurable from day one from step one. Mark do you think like that when you're trying to make the models ever smarter? Are you thinking in terms of first principle what would have to be true to be measured as being smarter or do you rely purely on generalization to try and get ever smarter models?
M
Mark Chen28:04
Yeah. So I really think when it comes down to it, why do we care about attacking math and physics at a place like and it really comes down to we are out of good evals, good human written evals and science doing science is the eval now and math is particularly exciting because you can you know attack some kind of theorem you can verify it in many cases and you know you feel confident that you are legitimately pushing the frontier forward. I know there are initiatives in physics too. you I know in physics there's a little bit more handwaving around oh you know this constant's too small and so but you can still you know build pretty formal systems right and I think and so you know it allows us to kind of really push the frontiers in both math and physics but fundamentally one of the reasons we cared so much about reasoning and in informal language is we care about generalization right we want to be able to do deep reasoning in fields like biology too and create breakthroughs. Even if it's kind of fuzzy what a breakthrough means, right? I think in math it's much more clear. It's like you solve a Navier Stokes. Yeah, that's a big breakthrough. If you kind of the model says, hey, here's your next breakthrough in machine learning. I mean, I don't know how to verify if that's true. And I think it's just so empirical and, you know, kind of time tells with a lot of these things. So I think what we care about is this fundamental generalizable reasoning layer. Natural language feels like a good way to express this in a way that kind of falls less into this trap of like you have a tool bag of techniques and you just center on the known techniques. I feel like in natural language we are able to kind of express these new techniques at least I think yeah we've been able to do that so far. So, yeah, we really deeply care about generalization and I do think kind of these formal fields they give us, you know, a really rigorous way to test that we're pushing the frontier.
U
Unknown30:04
Beyond the structural nature of maths and therefore the ways that you can formally verify. Is there some other practical benefit to pursuing ever greater capabilities in that space or is it in your mind really more just equivalent of an eval?
T
Terry Tao30:22
Do you want? Well, I think one positive feature for using math as a test bed for other use cases is that so you know we had this quote earlier today of Vladimir Arnold that mathematics is a place where experiments are cheap. It's also the place where failure is cheap. So you know it's related you know so you know if you're an engineer and you ask to build a bridge and the bridge collapses that's an expensive mistake. You know if you're a surgeon and you asked to and you cut the wrong thing that's an
In math, if you try to prove a theorem and your proof strategy doesn't work, that's not an expensive mistake. So I think we have this freedom to fail, which is more so than in other disciplines. And because of this, we have a culture of learning from our mistakes a lot more than in other disciplines. So it's a relatively safer place to experiment with AI than, let's say, bridge building or heart surgery.
M
Mark Chen31:20
Yeah, I love that you say that. It's exactly the way that we think about things at OpenAI as well. Because I think fundamentally what we care about developing AI for, the really inner core goal, is to use it to develop stronger AI. We want to design better experiments to build us stronger models, more intelligent models, which by extension will do even better math. But if it builds an even stronger model and you get that flywheel, that is an expensive thing. If you screw the system up in any way, compute's at stake. You run the wrong experiment, you burn a lot of money, a lot of compute. So I do think about math, physics as safe domains to push the frontier.
U
Unknown32:03
Yeah, it makes a lot of sense. I mean, I even wonder if you can push that, and Kevin, your work touches on this. Assuming a world in which the models may be discovering things that really are beyond the frontier of human knowledge, or even the ability for a human to really conceptually follow, presumably there needs to be some way to re-represent those findings into a logical chain that at least we can follow the steps, if not the actual constituent parts. And that maths, more than any other domain, seems to have invented a workflow that accommodates that. Certainly compared to, say, biology or chemistry, which doesn't have as much by way of kind of formal axioms that you can work against. So potentially it's a necessary precursor to any truly frontier science advancement to have this capability. Regardless though, it does, to your point, Terry, suggest that the way that we think about doing math in the future will change, that we're emphasizing maybe creativity, collaboration, different skills perhaps than what have happened the last 100 years. Does that filter then through into how you teach maths?
T
Terry Tao33:04
Yeah, it's... and so open problems still, how to... yeah, so in the very short term, some things have had to change. So yeah, like homework, weekly homework assignments have been the first casualty. But you know, I think we can push our students to do more ambitious things now. So I switched much more to a project-based type of assessment. In smaller classes, you can do some oral assessment. The skills we need to teach will be different. So validation, independent, the ability to independently verify AI-generated output will become essential. Softer skills, how to work with people, we have not been uniformly good at that in the past, but we'll have to get better. The pace of change is such that education systems are not catching up as rapidly as... but I think by necessity we'll be forced to. I mean, with COVID, for example, we did do some emergency changes to our curriculum and it kind of worked. It was not a great experience. So hopefully this time we can do a bit more planning. But it'll be on that level of change, I think.
M
Mark Chen34:30
Yeah. I think the analog of that is, you know, our interviews became busted very quickly too. You know, I think people, if they have time to do some kind of take-home or some kind of written type of interview, it's very hard. I do think kind of moving to a world where you can have a model also just kind of interact with you and teach you things, and the model itself can judge, you know, how much are you learning, are you uptaking, that actually kind of feels like a directionally good update. I've thought about revamping interviews in the form of, you know, you convince the model that you have the skills necessary to work at OpenAI. So you know, I mean, and of course you have to prevent hacking and jailbreaking and stuff like that. But yeah, I really do think fundamentally teaching has to change in some way. I am kind of curious to get Terry's take on a couple things. First, you know, I've heard from some other professors that, you know, it's... yeah, you really do see this divergence of like, it's like the worst... the best homework and like worst live exam scores in history. I don't know if that's a trend that you see as well. And I think the second thing is like, do you actually see this divergence of students who are like very motivated to learn that get really good using the tools? And is there this cohort that really feels accelerative?
T
Terry Tao35:56
Yeah. So definitely I've noticed homework scores going up and in-person scores going down. Not so much... it's not at a collapse level. I mean, I don't have hard data. I do get a sense that the weakest students are using AI to sort of get to a median level. And the brightest, brightest students generally tend to avoid using AI because they're worried about... they can notice that using it too much atrophies their skills. The weakest students, I think, they feel like they have less to lose. But which... so it in some sense is an equalizing feature in that respect. And yeah, I mean, once you have a certain level of expertise, these tools are great. And so maybe the equilibrium is to actually discourage their use or use them in specific ways. I can certainly see homework assignments of the future where the solution is not the point, because anybody can enter in, but for example, what prompt did you use to get to the solution? And that might be the more interesting assessment tool. So yeah, we have to figure it out. And actually, it's important, just like AIs will optimize whatever reward function, the reward function we give to our students will actually make a lot of difference. Yeah, we have to think this carefully.
U
Unknown37:25
Yeah, it's... in some ways the extreme cases are easy to understand. Total cognitive offloading and therefore no learning is occurring, etc. Maybe the more nuanced though would be something where it is a productive use of AI, and I think you used this metaphor in a recent interview, which is to say you've been helicoptered to the destination as opposed to taking the scenic route there. There's something about the change in workflow corresponds to something at the cognitive level for humans, and that we don't know yet what it may be that you lose if you do that, but we need to be alive and alert to it. Do you have a thesis about what we might lose if we start doing that?
T
Terry Tao38:05
Um, I think we will see empirically pretty soon. So yeah, I think we just need much more awareness of all the different facets of research or any other task. So somehow... so I said before AI allows for a decoupling of many things, which can be good for divisional labor, it's more efficient. But it does mean that goals that... where previously it was okay to set very fuzzy goals because any human attempt to reach these goals would sort of also hit all the nearby goals as well. So yeah, as I said, with an AI, you know, if you want to go see a nice waterfall or something on a mountain, you take a hike and sometimes you see some interesting wildlife or you get a glimpse of an even nicer location that you might want to go to someday. And maybe you meet some other hikers and you have conversation. There's serendipity, which just naturally happens. And so in the past we would just say it's a good idea to go visit this waterfall, but we didn't sort of unpack that carefully enough to say why we do that and what are the actual values, what are the actual benefits. And so but now we have this alternate way to get to these waterfalls. As I said, you can get an AI helicopter to drop you off there. And so yes, you get your little Instagram photo, but maybe that's not the only thing that you wanted. So yeah, I think unfortunately we're going to have to learn this by experience. It's hard to... I mean, you can talk romantically about the journey and things, but I think it's only when we see what happens when we don't have that, that we really understand what we're missing.
U
Unknown39:54
Yeah. I mean, I should say just today we released our learning outcomes measurement suite, which is how we use the models to assess whether humans are learning when they use them. So agreed as sort of a live research question. But I wonder, Mark, for you, serendipity and the idea of an inexact answer from the model in order to create space to explore, is that a quality of the models that you're interested in exploring? Is that a model behavior question, personality question? How do you grapple with that?
M
Mark Chen40:24
Yeah. So actually one of the biggest initiatives we have this year in terms of building a new primitive and a new interaction paradigm with the AI is we started an interactive agents team. And I think it's not sufficient that you just ask an AI a question and then it just comes back, even let's say like a day later, with its best attempt at a solution. I think humans are collaborative, they work in these constructs. And if you take like a completely non-math example, right, you want to create some kind of, let's say, PowerPoint or some kind of artifact like that, that's not the way you operate. You don't just tell some AI like, just make me a PowerPoint, it should be perfect and should kind of address these things. You want it to kind of come back and you shape the direction it's going, and there's multiple rounds of interaction with this agent. And I think that truly is what it's like to co-work with a very intelligent agent. And we want to build that deeply into the model, just something that's very steerable, it feels like a thought partner. And I do hope, you know, within a couple months, at least within a year, that's the way the AI looks.
U
Unknown41:30
Yeah, it's much harder to do reinforcement learning on collaboration. Like, how do you score how good you're vibing with your co...
M
Mark Chen41:39
My thesis is it is possible.
U
Unknown41:42
I agree that it might not be as difficult as you imagine. There are probably quite clear biological signals of what vibing looks like in the real world that you can bring back into the machine world.
M
Mark Chen41:53
Okay. Yeah. So, embody your AI body language.
U
Unknown41:57
There we go. I'm glad that I've got a quote that will be kept after this. I will open up the floor for questions just after this. I'll ask one more just to give you all time to think about it. Perhaps we will meet again in a year, and I'd be interested to know what your predictions look like for where we'll be in one more year's time.
T
Terry Tao42:19
I really hope we're going to see a lot of new types of mathematical projects that are kind of challenge-based, where, you know, so like first-proof type things, where some group of mathematicians, for example, will create a really good set of problems that they would like some proportion of these problems solved, and they have a very good gradation of difficulty, they have a very good verification protocol, and they would just open it up to the community. And so just sort of taking full advantage of, well, not just AI but also just like the internet and Metcalfe's law, you know, if there's N people who can produce problems and people that can solve problems, then there's N squared possible connections. And mathematicians have been very bad at using sort of making this large-scale network. So I think we'll see a different, almost like a marketplace style of doing mathematics. And there I think we'll see AI shine. So this is one thing I want to see, and maybe in a year we'll start seeing that.
M
Mark Chen43:30
Yeah. I hear ML's kind of foreshadowing math here. When you look at how frontier labs operate today with research scientists, we are moving into this world where the strongest research scientists, they're able to kind of pursue a lot of ideas in parallel and just really act as orchestrators. So they can think about this idea, think about a bunch of variations in the experiments, and just kind of have the model go and execute and implement that. I hope there is that kind of similar paradigm in math, where people like Terry and yourselves feel empowered to just go explore a fairly broad set of ideas and strategies with very little handholding. I do think the very little handholding part will also become more true. The task horizon will continue to elongate, just like we were at minutes of, you know, in terms of the horizon a year ago. I think in a year from now we're going to be in multiple days where you can actually trust the model to do tasks that would take you that long. And then I think beyond that, yeah, it's just making sure the interaction is seamless, right? These things should just feel like they interact very naturally with groups of humans and with the communities that you guys operate in. And finally, I really do hope we have some really big breakthrough, whether it be in math, in physics, in biology. I think, you know, today, the things we're proving are good and all, but I do think there's the potential for this to actually produce something that's very beneficial for humanity.
U
Unknown45:05
Fantastic. There was this metaphor in software development of bazaars and cathedrals. A bazaar being a self-organizing thing that springs up and is very diverse, a cathedral being one great mind architects and is therefore very elegant. And that the idea perhaps will be that in math we get both those phenomena occurring, and hopefully that is a flourishing. So I do want to open up to questions. Does anyone have a burning question?
A
Audience Member45:29
Yeah, I wonder whether any of you can talk about the word models which instead of predicting the next token, it predicts the next state. And what I'm reading about the V-JEPA is that it can actually self-correct things because the hallucination can be avoided. Is that all true? Because whatever I could run on my Mac, context model is just a hello world thing. So I don't know much about it.
U
Unknown45:53
For the benefit of those watching, the question was about how do we feel about world models? Are they a paradigm shift, and would they be specifically useful for maths and solving hallucinations?
T
Terry Tao46:04
I think it's potentially a very promising alternate direction. LMs are great, and in some ways they're too great, actually, in that we've kind of routed our entire AI infrastructure around making the LMs as powerful as possible, and it could crowd out some other very complementary ways to create AI assistants that are jagged in a completely different way. So I definitely support research into world models. I think for a long time they will underperform the LMs because we had just, because of all the momentum and infrastructure, have... it's like we have built our cities around the automobile and gasoline, and we have this entire infrastructure, and it's actually making it hard for alternate transportation to break through. But there are definitely people who are pushing that, and I wish them a lot of luck. Yeah.
M
Mark Chen46:59
Yeah. I do think when you think about a pure video, you know, generative video world model, we still seem pretty far from that. You know, I think the existing video models, they're pretty good, you know, physics simulators, but I do think with a little bit of adversarial pressure, they also fall apart. I do imagine that'll get more and more robust over time, but you know, it's not quite there yet. We are pushing fairly hard on that. I think, you know, there's many spectrums of world models. You can see an LLM as a world model too. But you know, I think digital world models where we're interfacing with computers, you know, there's all the rules and the feedback of a computer, that's a very important, interesting system, and I do think we'll tackle and really get a lot of value from that very soon.
U
Unknown47:42
I wonder on that if there is a middle ground, in as much as you can construct RL environments that are based on the laws of physics or follow the laws of physics, which to some degree confers the benefits you otherwise get from a world model to an LLM or anything else. So perhaps it's an intersection rather than two alternate pathways.
T
Terry Tao48:02
Yeah. So AI in science, in many fields, well, is effective when it's predicting very well. For example, protein folding, predicting weather accurately. But in mathematics and in theoretical physics, we're asking for something different, right? We want to kind of understand, get a formula, get a proof.
A
Audience Member48:19
But is it... do you think it's potentially too limiting? Will it be easier to get an AI that will tell you, I have a proof of theorem, hypothesis, but your brain is too narrow to understand it. So I can teach other AIs about it and do more progress with it.
T
Terry Tao48:36
Yeah, that's definitely my relationship with AI already. So, some of us have hit that frontier.
U
Unknown48:42
Again, for the benefit of those watching, the question was in some domains of science, I think we're satisfied with pure simulation. If it can do the thing, we consider the thing to be proven even if you can't formally verify it. So, if it can predict weather accurately, even if we don't know how, we're kind of happy with that outcome. We hold in maths and physics to a different standard, which is that it must be verifiable in the way we already discussed. Is that somehow limiting or a mistake to put that restraint?
T
Terry Tao49:08
Yeah, I think there'll be different types of mathematical tasks that we don't do nowadays which we will entrust AI to, and they could be quite complementary to the task of getting a formal proof of a problem. To give you an analogy, so in chess nowadays, all chess players train using these chess engines. And one thing the chess engine does is it gives you the score at any time in his position. You know, white is three points ahead or whatever. And it's a really great signal to train human chess players. You get instant feedback. Oh, that was a really bad move. I'll try and do this instead. I could imagine an AI which, you know, a human is trying to do a proof and like every time you say I'm going to try to prove a contradiction, your score goes down. Okay, that's a really bad idea. Okay, you should back up and do something else. So, maybe a Swan electric shock. So, that's a pro... the pro model. Yeah. Okay. So, yeah, a good math teacher can do that, but yeah, but maybe so we have to be creative. The type of tasks that maybe AI could help with that we just don't think about today.
M
Mark Chen50:14
Yeah, I mean, I do think verification is important. It doesn't have to be formal verification. I think we deeply want to know why something is true. And I think that's actually part of a deeper alignment problem, right? When the AI is attacking, you know, actual real-world impact tasks, you want to know why it made a certain decision, right? Let's say, you know, decided here's the best strategy to, I don't know, like, grow a business or something, right? You don't want it to do that without having a good justification. And so we have a lot of alignment techniques like debate, right? Where even if you don't get necessarily an airtight formal thing, you can kind of understand the outline and kind of like interact with the proof and kind of question it. So yeah, I do think kind of investment into techniques like debate and alignment, they'll really help us in the future.
A
Audience Member51:04
So on that, can you get anything out of the space stuff, like see what the intermediate might think?
M
Mark Chen51:13
Yeah. Yeah. So that's something we look into a lot, right? I think the top-order thing is, you know, just being able to monitor the reasoning or the train of thought, and you actually get a lot of insight from there. You can actually like do a lot of, even with many attempts on a problem, kind of get a sense for like what strategies the model gravitates towards and just kind of gain a lot of insight into how the model brain works. Yeah, I think you can go a level deeper, like look at the activations and like try to find mechanistic circuits and things like that. But yeah, definitely a very deep field of study.
U
Unknown51:47
These two questions may link up in as much as for interpretability, the way that we collapse the latent space does limit the kind of associations that could come out of the model. At what point do you decide that it's better to maintain the latent space because of the theoretical new connections it could make at the cost of interpretability, or do you just think that we need a different interpretability paradigm that doesn't require that kind of collapsing?
M
Mark Chen52:12
I think the reason we operate in text space today is interpretability buys you so much, right? I think you can debug so many things that go wrong with the model by just being like, oh well clearly it's like reasoning wrong here, so there's, you know, we can go and debug. When you're doing something in just like pure uninterpretable latent space, you lose that. And I don't know that we would switch to something like that in the interim.
T
Terry Tao52:36
Well, ideally we should have a diversity of models. So maybe there are some applications where you just want the answer, you don't care about interpretability, and then you just turn the dial one way. Okay. But there are other applications where you really want to see the process, you really want to see a human-readable chain of thought or whatever, and you turn the dial the other way. Yeah.
U
Unknown52:54
Yeah. I mean, it does feel intuitively like having to compress things into language does come at a cost. So even alongside a plurality of models, you might also want a plurality of expressions or verification methods so that you don't somehow force it into a shape that doesn't make sense. Sorry, please go ahead.
A
Audience Member53:13
I have a question about attribution and incentives a little bit as we look to this kind of future of science. One thing with AlphaFold, you know, sort of the world, I think largely thinks like AI came and solved that problem. And of course, AI in some sense did, but it was sitting on the protein data bank and decades of effort. And then you see the protein data bank loses its funding immediately after and various things like that. And I'm just curious that, you know, as we think about like these large-scale math problems and all of the human effort that's going to need to go into producing the sets of problems and verifying them and etc., it does seem like there are for many of these things, theoretical physics problems, all of them, that there's a lot of danger that in some sense while AI was maybe like the critical enabler that made something happen, it didn't happen on its own, it's really this ecosystem. And that so on one side it's how do we control that narrative, and the other part is in some sense like for OpenAI and for the big companies have a lot of responsibility in some sense about how they navigate that. And I'm just curious about, you know, how do we avoid it going in a bad direction.
T
Terry Tao54:20
Um, well, yeah, this is an important point. A partial solution, so as I said, I do envisage the rise of like challenge problems where people will create these data sets of tasks that they want solved, and there it's kind of win-win because the people who create these data sets, they will get some fractional problems solved, which is what they want, and then but these data sets could be very useful to calibrate AIs. And so there are some cases where it can be win-win. But yeah, there are definitely cases where people have built a data set at great expense for not for this reason, and then it gets absorbed into various AI. And yeah, I don't know how well we can track that. Yeah, so it leads into like intellectual property law, and it is a very tricky problem. Okay, which I will toss to you to deal with.
M
Mark Chen55:13
No, I think as of now, I think the AI doesn't want your credit. So I think... but yeah, I do think the vision we have for Open for Science, it's really not about us claiming the credit here. I do think, I mean, we certainly have the ambitions to move science forward, but Kevin here, he wants to build a platform where mathematicians around the world can just accelerate the field in its totality. I think like we don't know the right questions to ask. We aren't the orchestrators within OpenAI. I do think kind of the credit should just go to you guys. I know that's not exactly the question. I mean, I think it's not like AI is getting credit, but I think the public perception is AI solved this problem on its own.
T
Terry Tao55:56
And they come somehow like humans, experiments, and all this is not so. With the math problems, what we've seen is that there'll be times when there's an open math problem that no one has looked at, and then some AI solution gets a solution, and then this hits social media, AI solved an unsolved problem. Okay, and in many, many cases, you know, like 24 hours later, someone often armed with one of these deep research tools uncovers that this result was already proven by a very similar method in the literature. And we can't say for sure whether the AI used that solution or indirectly was aware of it. But it happened so frequently. I mean, we have this whole table of contribution. It's like the whole section we just... this... and so to some extent, we at least have the capability to detect at least some of this, because we also have these research tools. So we can recover attribution sometimes. It's not perfect. But yeah, it could be that the same technology that allows us to use the literature to solve problems can also use the literature to attribute solutions.
M
Mark Chen57:12
Yeah. Just one thought on top of that, in general it is a very hard problem, just data attribution, and when you generate something, you know, how inspired is it from which data points. One interesting thought here is perhaps like the novelty and contribution is somewhat correlated to the amount of time the models spend thinking about something. Maybe it's not always true, but yeah, I think, you know, modulo stuff it rediscovers in literature. There's a PR element though that just like DeepMind could have gone out of their way to say more about the protein data bank and the narrative. Like there's things that are not about the AI models and data. So it's just about how we talk about it in PR and announcements.
T
Terry Tao57:55
Yeah, that's a very fair point. I understand the incentives. I do hope, you know, anyone here at OpenAI can verify this, but like we care a lot about integrity. I think, you know, at least I would very much strongly fight for the correct narrative there.
U
Unknown58:14
Maybe just to add a flavor from my previous life, maybe one paradigm shift is that the general public tends to underestimate the degree to which the progress of science relies on improving tools fundamentally. You need better microscopes, so on and so forth. These are not glamorous jobs. It's not typically what theoretical scientists like to work on. And that's a narrative we should be saying, is that by building ever better tools for science, that's what enables humans to accelerate the entire field, often, you know, really big steps forward. AlphaFold is really a tool. It's not per se scientific research, though it does do some research component things there too. And if we can stay on the emphasis of it being a tool, that might also flow back to your point about where does the investment flow in society, because it's to the things that need to surround the tool to make it effective, but ultimately in service of the human researchers or scientists that will use the tool to solve the problem. We are just on the edge of time, so I'll take one more if that's all right.
A
Audience Member59:16
Yeah. So I think one of the really interesting things about this conference is, you know, it's accelerating math and physics using AI, and I think there's a certain amount of kind of synergy that comes from like being able to solve math and then how that can help solve like other physics problems. So I'm curious to hear your thoughts on like additional synergies you see coming out of, you know, OpenAI through you guys kind of tackling math, physics, and then kind of what else you guys see coming out from here.
M
Mark Chen59:44
Yeah. Well, yeah, so I think Kevin would probably be the best to speak to this, but we do care about tackling domains outside of math and physics as well. I think some that we have explored are biology, where we have had the AI work on just making the biological procedures in the wet lab much more efficient. I think with one of our partners, Ginkgo Bioworks, we iterated on a lot of their core processes and made the cost for synthesizing proteins, I think, 40% more efficient. And I think that's just the underlying primitive that will drive more progress. So a lot more you could imagine doing in things like material science and other domains.
T
Terry Tao1:00:24
Yeah, and say that IPAM, this institute, I mean our entire core mission is basically to find these synergies. We, you know, Institute for Pure and Applied Mathematics is basically in the name. Yeah, we bring together events like this, you know, where different communities talk together. Yeah, so, and a lot of it is this serendipity. I mean, we have some idea, you know, we don't just randomly smash together fields. Okay, but yeah, we do pick ones where we do believe there will be a lot of unexpected fruitful collaboration.
U
Unknown1:00:53
I'm glad you leave randomly smashing together particles for the physicists. That's a good moment to close it. And also to plug that Kevin is about to give a lecture which covers many of these questions, i.e. what needs to be true to accelerate all parts of science and the tools to build for it. So I hope you'll come to that. But thank you all so very much.