Back
Kevin Weil
VP, OpenAI for Science, OpenAI

Kevin Weil - AI for Mathematical and Scientific Discovery - Bhaumik Public Lecture at CNSI

🎥 Mar 04, 2026 📺 Institute for Pure & Applied Mathematics (IPAM) ⏱ 62m 👁 236 views
Recorded 04 March 2026. Kevin Weil, VP of OpenAI for Science, presents "AI for Mathematical and Scientific Discovery." This Bhaumik Public Lecture at CNSI was part of IPAM's Accelerating Math and Theoretical Physics with AI Workshop. Learn more online at: https://www.ipam.ucla.edu/programs/sp...
Watch on YouTube

About Kevin Weil

Kevin Weil, then-VP of OpenAI for Science, discussed the company's efforts to hire practicing mathematicians, physicists, and biologists to develop frontier AI models. He described the models as "incredible," noting that they have progressed from achieving a 700 on the math SAT three years ago to "regularly solving open problems in math and physics and other scientific fields." Weil stated that OpenAI's mission is "not to win a Nobel Prize ourselves" but "to see a 100 scientists win a 100 Nobel prizes using our technology." He characterized AI as a "metal detector for hypothesis," saying it has read "substantially every paper across every field of science" and can generate more ideas than scientists can experiment with. Weil discussed OpenAI's role in the Department of Energy's Genesis Mission, calling it "one of the most exciting projects happening right now." He said the mission has a "huge amount of scientific data that is currently mostly unused" and that teaching AI models that science could "confer advantage on the US." He expressed particular interest in fusion energy, where AI could "iterate far faster on parameters using simulations and real experiments." Weil also addressed barriers to adoption, including the cost of compute, stating that scientists who could "do the most amazing things" with it often have "the least ability to pay." He acknowledged concerns about AI-generated "slop" in scientific publishing, comparing it to email spam and suggesting AI would ultimately be used to filter it out, while noting that "peer review will still be a thing."

Source: AI-verified profile updated from Kevin Weil's recent appearances. Browse all interviews →

Transcript (66 segments)
M
Moderator0:00
So welcome to the Bo public lecture. So right now there's a workshop, a one-day workshop organized at IPAM together with OpenAI and the BIC Institute on accelerating math and theoretical physics using AI. And we jumped at the opportunity to get the head of AI for Science here. And it's Kevin Weil. Oh, Kevin Weil. That is, sorry about that. And he has a math and physics degree from Harvard, a master's from Stanford. He did drop out from his PhD, but he had other things that he decided to work on, startups. He was head of product at Twitter and Instagram. I don't know if that's good, but he was also president of Planet Labs. If you don't know what that is, you should Google it. It's really cool. He built and launched more than 100 imaging satellites. And then he joined OpenAI about two years ago as chief product officer before recently starting OpenAI for Science, a group focused on accelerating scientific discovery with AI. And today he's going to tell us about AI for mathematical and scientific discovery. And he's going to tell us how to get superpowers.
K
Kevin Weil1:38
All right. Thank you, Z. Thank you, Monty. To the IPAM team for having us here. I got to say it's intimidating as a PhD dropout to be presenting here in front of all of you in this setting, this august group, but I'm excited to talk about AI for mathematical and scientific discovery and in particular to try and make the point that this isn't some dystopic view of the future where AI does all the work and we sit around collecting our universal basic income and writing poetry. Far from it. I think this is the most exciting time ever to be a scientist. The reach that every one of us has is broader because we have a collaborator in the AI that's extremely knowledgeable across basically every subfield and is also infinitely patient and we can go deeper in collaboration with AI. We can explore more routes of attack on any particular problem far more than we could just as people alone.
Now I'll tell you maybe a quick story from a conversation I had with a mathematician, a Fields Medal winning mathematician. So somebody at the very top of the field and he said, you know, I've written papers now for decades. So, I've got tens, hundreds of papers that I've written over the course of my career. And every single one of them, there were routes that I could have explored off of my papers that I knew would yield fruit, but I didn't have the confidence to explore them because they would have taken me outside of my domain of specialty. Now, with AI, I'm going back through all of my papers and following those paths that I couldn't take then. And if a Fields Medal winning mathematician is saying that he didn't have the confidence to explore these routes and now he does with AI, that just to me points to the value of these things.
So, but let's start by level setting on just how far we've come to get there. And in particular, let's go all the way back to three years ago. So far back that instead of an iPhone 17, you had an iPhone 14, the horror. And back then there was a model called GPT-4. And it blew my mind that GPT-4 could get a 700 on the math SAT. So GPT-4 could get a better score than 90% of high school students. You see an example of an SAT math problem over there normalizing a complex number.
Okay, so let's up the challenge. Instead of the SAT, let's use the AMC, which is a top high school math competition. Here, the problems are meaningfully harder. You can see the example on the right. And by the way, we're in mid-2024 now, so we're about a year further on. And rather than a 700 on the SAT, GPT-4's successor, GPT-4o, can only get 9% of the problems right. So, this is a much harder evaluation.
All right, let's go to the end of 2024. So, fast forward another six months. This is like what 15 months ago. This is when we released our first reasoning model which was called o1. So, we had a major jump. o1, as you can see, got 74% of the AMC questions right. And it's worth pausing for a second to explain what the difference is between o1 and all of the models that came before it. The major difference is that o1 could reason. Which is to say, unlike GPT-3 and 4 and unlike every other model at the time, which instantly gave you an answer to a question the minute you asked it, o1 would do what you or I would do if we were asked a difficult question. We wouldn't just spit out the first thing that came to mind. We'd pause. We'd think about the problem, maybe identify different ways to attack it. We'd find similar problems that we know how to solve. Maybe we'd look at reduced versions of the problem to try and simplify it and gain intuition.
o1 could use additional computing power, not in the pre-training of the model, but to think when it was answering the question. And just like you and me, there are a lot of questions that I can't give an answer to if you give me 30 seconds, but maybe if you give me 30 minutes or 30 hours or 30 days. So this ability to think in response to a question and not just at pre-training time but actually think when asked a question gives us a second dimension along which to scale the intelligence of AI models and it unlocked the ability for models to solve harder and harder problems which is what led us to this and to where we are today.
So now with reasoning unlocked the next model was called o3. Why was the next model after o1 called o3? We have some friends at Telefonica who operate a mobile network called O2. You can always count on OpenAI to find the most complicated naming scheme possible. But this is just 5 months later and o3 could get an 89% on the AMC, which is better than basically all high school students. Not quite all, but substantially all high school students. But, you know, okay, these are very smart high school students, but it's a high school group of high school students doing a high school math competition. What about a real top math competition? What about the International Mathematical Olympiad?
So now we're in July of last year, so you know, on the order of six, seven months ago, we had an internal model that was able to achieve a gold medal at the IMO, which means that its score was in the top five of all students across the world taking the IMO. And you fast forward another few months to this past December, so three months ago, and our models are now basically perfect on AMC. So, we've come a long way just in a few years from a model that couldn't even answer certain SAT math questions to a model that aces the AMC and gets a gold medal at the IMO. And we take this all in stride, but it's kind of a wild pace of progress, right? It's vastly faster than Moore's law. And Moore's law completely revolutionized the world over the last few decades.
So I actually have another story. Who's ridden a Waymo? Pretty good. So a lot of people have ridden in a Waymo. I don't know what your experience was. My experience when I rode in a Waymo for the first time, my first like 30 seconds in a Waymo, I was white knuckling the thing like, 'Oh my god, watch out for that bicycle.' And then after a couple minutes, your blood pressure goes down and you're like, 'Oh my god, I am living in the future. I am being driven around San Francisco by a robot. There's nobody in the driver's seat. This is amazing.' And that feeling lasts for like 5 or 10 minutes. And then after five or 10 minutes, I spent the rest of my first ride in a Waymo, bored, looking at my Slack and doing email and like, you know, when is this traffic going to... In the space of like 15 minutes, I went from scared out of my mind to blown away to like, okay, this is just part of my life now. Of course, robots drive me around the city. And I think there's a metaphor there for how quickly we adapt to these rapid shifts in the quality and capability of these models.
But speaking of quality and capability, everything I've talked about so far, this has all been contest math. And contest math is somewhat formulaic. You need creativity, but there's this selection of techniques and they'll get you pretty far once you kind of understand the set of techniques that you have on contest math. It's not like this is math research. We're not doing novel things. And this idea that LLMs would never actually do novel research was captured in this phrase stochastic parrots. And the idea was that LLM just probabilistically parrot back segments of the things that they've learned. And since all they do is sample from a distribution of stuff that they've learned in the past, well, they obviously can't do novel things.
So we felt differently and we could see it happening internally and externally with people using our most advanced models to do novel research. We wanted to show people. So in Q4 of this last year we had been seeing more and more of these examples. We got together a group of about 10 different external scientists, physicists, mathematicians, biologists, some material scientists. The idea was like not to break ground with novel research but to sort of put a stake in the ground to say where we were in the evolution of AI. So this paper is about 10 chapters beginning with more modest examples involving like AI powered literature search. It goes from there to showing how different scientists are using AI as a research partner and then finally at the end we have examples of AI actually helping solve open problems. And in some of these examples, these open problems were ones that the authors themselves had worked on for years prior to AI and hadn't been able to solve them. And then in tandem with AI, they were able to solve the problem.
One of these actually was an Erdős problem which is named after Paul Erdős who is this incredibly generative combinatorist and number theorist who left behind over a thousand open problems of varying difficulty. And these problems have been brought together under one roof now. They're known collectively as the Erdős problems. About 40% of them are solved now, which means there are about 700 of these Erdős problems still open. And as I said, they're of varying difficulty. So some are kind of minor forays beyond the frontier of human knowledge and others are important, well-known, and very consequential. So Máté and Mark on our team in tandem with AI solved one which at the time we were like oh my god we solved an Erdős problem. And then over the next month or two there was this huge rush from all over the world as other mathematicians used GPT-5.2 to find all sorts of low-hanging fruit among the Erdős problems and in very short order I think they solved like 20 or 30 of them either because they found solutions that were already sort of in the literature in places or in some cases just outright proved the answer. But these were all pretty low-hanging fruit, right? Maybe these were more bottlenecked on attention than actual intelligence as Terence Tao has said.
So we need the next more challenging test. We need a test that was designed from scratch to ensure that the problems are true research level problems. So now enter Frontiers Math which was now we're basically caught up to the present. This is February. Frontiers Math is a play on a baking term put together by 10 mathematicians and it could not have been better timed. So I'll quote from their paper. We present a diverse set of 10 research level math questions drawn from the mathematical fields of algebraic combinatorics, spectral graph theory, algebraic topology, stochastic analysis, symplectic geometry, representation theory, lattices and Lie groups, tensor analysis, and numeric linear algebra. Each of which came about naturally in the research process for one of the authors. Each question has been solved by the author of the question with a proof that's roughly five pages or less but the answers are not posted to the internet.
So they took those 10 problems which they had solutions to and they challenged the internet and basically they challenged the frontier AI companies to solve as many as they could. They gave a one-week deadline and they said we've posted our encrypted solutions. We're going to unencrypt them in a week. And they basically said go at it. Oh, and by the way, you need to solve these models without a human in the loop. So, it can't be progressive prompting with a human guiding it. It's got to be just give the question, see if you get the answer.
Now, of course, we chose to play. How could we not? There were a number of external folks who were betting that nobody would get more than two out of 10. We used actually an internal unreleased model to try and answer these questions and we believe that we were able to answer five out of 10. We actually thought that we did six initially but one was shown to be incorrect like almost immediately after we published. But our writeup, which is here, which by the way was also written by the model, includes not just the ones that we believe we got right, but also includes the ones that we're pretty sure were wrong because we felt it'd be interesting for the community to see the way that AI can be right and wrong when tackling problems at the frontier. But you hear me, by the way, saying we believe we believe we got five out of 10 right. Like, it's math, right? We did or we didn't get five out of 10 right. But this gets at the challenge of verifying AI proofs especially at scale especially at the frontier. Something I'll come back to later.
Okay. So now let's look at I've gone 10 slides into my presentation. We've traveled three years. We've gone from 2023 to now 2026. And in that three years, we've gone from a pretty good score on the SAT to being able to autonomously do five of 10 open research problems across a super wide spectrum of fields of math. I don't know whether there's a mathematician that could answer five of those 10 questions in a week. And that points at one of the novel and most interesting aspects of doing math with AI is there are certainly people probably in the audience here today who can do mathematics well beyond what GPT-5.2 can do today. But I'm not sure there are any, in fact I would wager strongly that there aren't any who know the sheer range, the breadth of research grade mathematics that GPT-5.2 knows, let alone the complement of other science it knows across physics and chemistry and biology and health and economics and material science. And that's one more reason why AI gives scientists superpowers.
All right, let's switch gears and talk physics. So here I'm not going to take you all the way back to GPT-4. We're going to start about a month ago. Now I was fortunate to do my undergrad in physics and math at Harvard. And while there I looked up to Andy Strominger. He's one of the top string theorists and particle physicists in the world. He has been for decades. And in particle physics, you think about how particles interact. You model it with what are called scattering amplitudes that give the likelihood of a certain set of particles interacting in a particular theory. We're fortunate to have, by the way, at least a couple of the people who've laid the groundwork for this theory in the audience today. Zvi Bern and Lance Dixon. So you can model these interactions with Feynman diagrams, an example over there on the right. And that gives you a perturbative expansion that allows you to calculate these quantities typically with huge amounts of pain and many many pages of calculations from grad students. Ask me how I know.
Increasingly though, there are examples where this hugely painful combinatoric expansion yields a messy result that upon aggregation miraculously cancels down to something nice. And when that happens, it tells you that there's some underlying simplicity often a symmetry that you've overlooked and you worked way harder than you needed to. So for years physicists had assumed that this particular amplitude where you have one negative helicity gluon and gluons are the particles that mediate the strong force which holds together quarks and protons and neutrons and so on. You have one negative helicity gluon coming in and you have n minus one going out with positive helicity. And you can show that when all of the helicities are identical that this quantity is zero. And it was thought to be zero when only one gluon had a different helicity. And so that the first non-zero terms were when you had two particles with identical helicities and n minus 2 of the other. And this has become so well accepted that this case of two particles of one helicity and n minus 2 of the other is called the maximally helicity violating amplitude or MHV amplitude. Right? It's maximal. This was like known for decades.
Except Andy Strominger believed differently. He thought that there was probably a region of phase space, this so-called half-collinear regime where the one helicity amplitude did not vanish and he wanted to prove it but the calculations as he did it were growing quite literally exponentially complex. So when you have just three particles, you have one of negative helicity and two of positive helicity, you get this expression that's just a nice single term. Although of course that term itself is shorthand for something more complex. So it's not perfect, but it's still pretty simple. Then you look at the four particle one, which turns into a sum of two products. And then the five particle one turns into a sum of eight different terms, each of which is a product of three factors. Okay, now let's go to n equals 6. And we're beyond what any human actually wants to deal with. And we're only at six particles. So you can make a further assumption and things simplify a little bit, but we're still going to grow exponentially complex as n grows.
So, Alex Lobser, who is on our team and is also a physicist at Vanderbilt, was a former student of Andy's, and we'd been looking for an excuse to invite Andy out and try to do some physics together. So, Alex and Andy, together with their collaborators, David and Alfredo, decided that this would be a great thing to work on together. So, we arranged a date and they headed out. On their way, like literally while they were on the plane, Alex started playing with GPT-5.2 Pro, trying to see if it could simplify this combinatorial explosion of terms. And it turns out GPT-5.2 Pro was able to predict a closed form for this expression for arbitrary n, but it couldn't prove it. And then we gave this problem to a more powerful internal model. And before Andy, David, and Alfredo had even gotten off the plane, it had proposed and it had proved the closed form of this expression for arbitrary n.
So instead of spending the week together at OpenAI actually doing the work, we spent the week verifying the work. There's that word again, by the way, verifying, which we'll come back to later, and then getting the paper ready. So we published and it's super exciting to see people's feedback. To me, the neatest thing in all the feedback is that the articles about the paper were like a little bit about the fact that AI did a non-trivial amount of the work, but they were more about the fact that this paper was actually just a meaningful result in particle physics. And again, I love this because I think AI gives every scientist superpowers. And when tools diffuse across an industry, the tools themselves fade into the background, right? Pretty soon we'll hardly ever mention when a result came about through interactions with AI because of course it did. Everybody will be operating this way.
One of the other physicists that I most looked up to in undergrad and grad school, Nima Arkani-Hamed, reminded us that finding simplicity in complexity, which is something that AI is great at, often points to new ideas and underlying structures. And Nathaniel, who's here with us today, gave the result perhaps the most powerful compliment that you can give to a piece of research, which is that it will inspire future developments and new publications. So, we're not talking about the fact that AI played a major role here. We're just talking about the progress itself. Together with AI, physicists are going to make a lot more progress. They're going to do it a lot more quickly. And that is something to be excited about.
And actually, hot off the presses, speaking of new publications and increasing iteration speed, Andy, Alex, Alfredo, and David already have a follow-up paper to the first. In this case, extending the gluon scattering result to gravitons, where the calculation is meaningfully more complex, but results in an equally beautiful final state. This time, one that points us towards quantum gravity. And that was put out this morning and I think will be on the archive tonight. So understanding how the theories of gravity and general relativity merge with quantum mechanics and quantum field theory is one of the great unsolved problems of the field. And it's really cool to see AI beginning to play a small role in helping us understand it.
Okay, we are going to switch fields one more time. This time we're going to talk about biology and the physical sciences. This was a result from about a month ago. And I'll quote from the abstract. We used an autonomous lab comprising a large language model, LLM, which was GPT-5.2 and a fully automated cloud laboratory to optimize the cost efficiency of cell-free protein synthesis. By conducting iterative optimization, the LLM-driven autonomous lab was able to achieve a 40% reduction in the specific cost, dollars per gram of protein, of cell-free protein synthesis relative to the state-of-the-art. This cost reduction was accompanied by a 27% increase in protein production. Iterative experimental design, experiment execution, data capture and analysis, data interpretation, and new hypothesis generation were all handled by the LLM-driven autonomous lab.
I want to show you what this looks like because I think it's the future of physical science. So you have an AI model trained in biology and given a goal, in this case which was to develop techniques for cell-free protein synthesis at high volume and low cost. It has as one of its tools an actual IRL robotic lab that it can use to evaluate its ideas quickly. So the AI model thinks, it designs different experimental concepts. It reasons through them. It tests them in silico as it reasons. And when it gets to a set of parameters it thinks could be valuable, it instructs the robotic lab to execute the experiment. The robotic arms run the experiment. And here you can see the sample that transits between different stations for different parts of the experiment and then at the end it submits the results back to the AI model. The model incorporates those results, uses them to refine its thinking, and then iterates again and comes back with an improved idea. So in this way, we ran 36,000 experiments dramatically faster than any human lab could ever do. And every part of this is scalable, right? You can give the AI more compute. You can assemble multiple robotic labs to parallelize the experiments and so on. So we don't need to be limited by postdocs pipetting things. Robots can do it and they can do it with a scale and precision far greater than humans. And the postdocs now having superpowers from AI can have more ideas. They can test them more quickly. So I really think this is a big part of the future of science.
Along these lines, I was talking to a professor named Pratyusha Sharma who's a linguist at MIT. His research is decoding whale language with AI. It's awesome. So, like, did you know sperm whales have vowels? They have complex social structures. They have a common language they use to speak to each other. And you can use AI to understand the patterns in their speech and sort of extract words and sounds that function like vowels in their language. His research is totally awesome. You should look it up. It's mind-blowing. But what he said, which has really stuck with me, is he said AI is a metal detector for hypotheses. As humans, we have tons of ideas. We have way more ideas than we have time to experiment with or investigate deeply. But when you have this advanced AI model at your side, you have this super smart thing that has read substantially every scientific paper written in the last n decades across every field. It's infinitely patient and it can be directed mercilessly to explore the pros and cons of any idea that you have or you can spawn 10 of them and explore 10 ideas in parallel and come back with the best ones. So, a metal detector for hypotheses, which is why I keep coming back to AI as a tool for scientists, one that gives scientists superpowers. The models are getting really, really, really good. They can do a lot of things you can't, but also you're better at a whole bunch of things than they are, and they're a power tool that you can guide with your experience, your intuition, your judgment, your taste.
So, I'm very optimistic about AI, as you can probably tell, but I didn't want to just talk about the stuff that works, because there's a bunch that still doesn't. AI needs to get a lot better, and I think it will, but what does it struggle with now from a research point of view? Well, the first that I've talked about a couple times is this idea of verification. The first proof set of problems that I was talking about is one example. When I said that we believe we've solved five out of 10 problems, why is it that we believe? Why isn't it we know? It's because it turns out it's very subtle to tell the difference between a correct proof of something new and an almost correct proof. The model can be overconfident. It can do things like refer to a lemma to solve a key part of the proof and then you need to go check the lemma whether it exists, whether it plays the exact role that it claims and so on. Or imagine applying AI to all 700 unsolved Erdős problems. Why not, right? I mean, this is a problem of scale. We can throw compute at it. We can try to solve all 700 unsolved Erdős problems at the same time. So, it turns out when you do that, the model will come back and say, well, you know, I couldn't solve 500 of them. I know I didn't get those right, but I think I've had to solve 200 of them. And you know, it turns out now you need an army of mathematicians to go check 200 solutions to 200 previously open problems, which takes time and is especially frustrating to the mathematicians because the majority of the problems actually aren't correct. Like AI isn't quite ready to solve 200 Erdős problems yet. And so really what you end up with is maybe it solves 20 or 30 of them, which is super cool, but you have 170 subtle errors that you have to go find. And so what this really means is the bottleneck is beginning to shift, right? It's no longer a problem of attention because we can direct compute towards whatever set of problems we want and we can see if AI models can make progress on them. But the problem is now one of verification where an AI model and a bunch of compute can propose a solution and we need to go figure out if it's correct. So in response to this a bunch of labs including us and a whole bunch of startups are looking at complementing proofs with formalization in Lean which is sort of a compiler for mathematical statements. This is also far from perfect today, but there's a huge amount of potential because if you could have the model propose 200 solutions to 200 Erdős problems and then automatically try and formalize each one and realize on its own that 170 of them weren't correct and 30 were or maybe it just needed a mathematician, a human mathematician to look at five of them. That's a very different world and you're able to go take that and solve lots more problems.
All right. Next is what I'll call unconventionality. So for most of what people use AI for, you kind of want AI to give the down-the-middle answer, right? Is 91 prime? Can you summarize this email in three bullet points? Who was the third emperor of the Holy Roman Empire? You don't want the AI to do low probability things when it answers those kinds of questions. You just want a quick, accurate answer. But when we go beyond IMO problems and
Sort of small theorems to truly original work. If we want to see AI models solve big open unsolved problems, every conventional angle of attack has already been explored by very smart humans. It's likely that many of these problems aren't going to fall without something highly original or unconventional. And we need to find more ways of rewarding the AI for trying these very low probability rollouts. I think this is going to be a key part of solving some of the hardest problems and it's something that we're working on.
And then last is invention. I've told the story now of the models going from SAT math to contest math to graduate math to solving open problems. But we haven't yet seen a model solve a major open problem, say at the scale of one of the Millennium Problems. I don't necessarily think it'll be that long, but even beyond a Millennium Problem is the notion of inventing a whole new field of science as part of the discovery. So, you know, think Grothendieck reinventing algebraic geometry with schemes. Think of the Langlands program proposing connections between what people thought were very different fields of mathematics. Think Einstein revolutionizing our understanding of spacetime with general relativity. These kinds of breakthroughs go way beyond solving an open problem. They uncover totally new areas of study and new vistas to explore.
AI is definitely not there yet. We don't have any examples of this, but I think one day we will be. We will have them.
And why do I think that? Why am I so confident that we're going to continue to see this kind of progress? It's because AI is progressing faster than any technology that we have ever seen in our lives. So this chart is put out by an organization called METR. It evaluates models on basically how long of a task, measured in like how long it would take humans to do it, that a model can consistently do. And as you can see, it's quite literally exponential.
Just in the time since I made this chart, Anthropic released a new model that is above GPT 5.2. And I'm pretty confident that in very short order, we're going to release a model that goes on top of that. So, this is the state of the world, right? Periodically, people talk about AI capabilities plateauing, and I'm here to tell you, I do not see that happening for the foreseeable future. Building AI is a surprisingly empirical science, and all the data we have points to continued growth in capabilities and intelligence.
We have models internally that are more capable than anything that we've launched publicly, and we're training models now that are already better than those. And I just do not think it will slow down.
One way I like to think about this, by the way, that really makes it sink in, you have to remember that the AI model you're using today is the worst AI model that you will ever use for the rest of your life. The AI model you're using today is the worst AI model you'll use for the rest of your life. Just remember that. But that's why I'm so excited about AI and science. I hope I've shown you today that today's AI models are already dramatically helpful for scientists. They accelerate our thinking, our testing, our calculation, our discovery. They've come so far in just 3 years that it's hard to imagine even where we'll be over the next three or four.
I think this year, 2026, will be for AI and science what 2025 was for AI and software engineering. So take yourself back to the beginning of 2025 and think about the state of software engineering. If you were using software agents to write most of your code, you were definitely an early adopter in the beginning of 2025. But then you fast forward 12 months, and 12 months later if you were not using software agents to write most of your code, you were falling behind. I think similar things are going to happen with AI and science in 2026. And those who adopt AI in their research are going to be able to accomplish dramatically more.
And if you go a few years out, I think it's very plausible that accelerated by AI, we're going to be doing the science of 2050 in 2030 instead. And I find that incredibly exciting for the world, for the lives we'll save through personalized medicine and curing disease, for better materials that lead to more abundant energy, for our understanding of the universe and the world around us. We can do the science of 2050 in 2030 instead. It is an incredibly exciting goal. It will take all of us, but it's a goal that I believe is within our reach. Thank you.
M
Moderator35:55
Well, thanks a lot for that very inspiring talk about the future. So we decided that the talk be shorter because we all suspect there will be many questions. So we have plenty of time for questions. So you're very welcome to ask questions. Who's first? We can start right. Yeah. The blue. Yeah.
It's okay. One then the other. Okay. Ask about the internal one.
K
Kevin Weil36:33
I mean, yeah, we well, there are tons of internal models because people are always, it's a, like I said, it's an incredibly empirical science. And so, if you're a researcher at OpenAI, what you do is build models. Most of them are not intended to be released. You're building models that test certain things. You're just basically running a ton of different ideas through and seeing if they show uplift in various different ways. And then over time we'll sort of assemble a whole bunch of those smaller experiments into a bigger thing and do a big run and that becomes an external model. And then from time to time we have models that we build to just sort of see how far we can max out capability in like particular areas, coding, math, science, things like that. So in a whole bunch of ways we have internal models and obviously, I mean our mission is about bringing AGI out to the world and to benefit all of humanity. We're not fulfilling our mission if we just have a bunch of models sitting on the shelf internally. So the goal is always to get them out to the public, but we have a lot of experimentation we do in the meantime.
M
Moderator37:42
These advanced internal models that can benefit humanity already. Are they outside of OpenAI as is?
K
Kevin Weil37:51
Well, they're internal now, but you know, fast forward a few months and you're going to start seeing more of them.
A
Audience Member38:10
Sometimes in some cases, yeah. In more of application and biological search, and that's one side. But I'm also curious about like how AI starts to interact with like people and patients when they have like immediate questions. So let me talk first. You mentioned about how like AI really helped in like routine research and those are more like a macro model where we can make the predictions and potentially get to more research on that. But I'm curious like if it'll keep the same level of accuracy and confidence in terms of like DNA or RNA when it goes down to like a lot smaller. You know, can AI really encapsulate the level of complexity that biology holds in very, very small?
K
Kevin Weil39:06
Yeah, I think it's one of the most interesting challenges right now on the bio side. People are trying to build a model of a virtual cell because if you had one then think of all the things that you could simulate in silico and you would be able to make progress a lot faster. To my knowledge there isn't a good virtual cell model right now but there are a lot of folks including us thinking about it.
A
Audience Member39:32
Like where I'm hearing that you're really optimistic about AI and onward, like really, how do you see like that virtual model to be out there and how nearly?
K
Kevin Weil39:54
A virtual cell model specifically or just like the kinds of models that were driving the experiments? Um, I don't know. It's a pretty big moonshot of an idea. I'm not sure it's like, you know, right around the corner. I think it'll take a little bit more time. But a big part of, I mean one of my beliefs about the value of AI and science is if you can shorten your iteration speed, you know, or shorten your loop, you get more, you can experiment more, you get more tries, you learn more quickly, that leads to quicker discovery. And so, you know, if you can develop something that even approximates a virtual cell, then instead of having to run experiments in labs, you can do more experiments in silico and then only run the most promising versions of those in a lab. And you know, you tighten the iteration loop and go faster. So, I don't think we're there yet, but it's one of the next big things for sure.
A
Audience Member41:00
Hi, thanks for sharing such wonderful news with us that all of, we are all mathematicians even though I'm not in the maths department and will soon become redundant. So that's good to know. I'm just kidding. Could you share or are there publications on how exactly are you training these models? Like you have a bunch of mathematicians conversing away or you're feeding more math books or what exactly are you doing?
Courses. I have a vested interest in the courses I teach. I can get the students to, you know, make better models, fine-tune with some of these strategies.
K
Kevin Weil41:44
Yeah. Unfortunately, we don't talk a lot about how we train models. So, that's one question I can't answer up here.
A
Audience Member41:53
But anything you can share?
K
Kevin Weil42:03
I will say if you, one of the things that we are realizing is the models are so good these days that to make them better at important topics like math and physics and biology it's no longer good enough to have, you know, even somebody like me, a PhD dropout, like I can't do what Alex can do in terms of teaching the model to be better at physics. You need someone like Alex who is at the frontier because the models are at the frontier. So it's already well past me. And you need to hire true experts in a field and you can see that we've been doing that across a variety of fields.
A
Audience Member42:46
I have a quick question about a lot of people are saying how like a big bottleneck for LLMs is that it can't simulate like the physical world and people are looking more into world models for like scientific research. Do you think that's something that like OpenAI would look forward towards and by like your 2030, 2030, for better research in science?
K
Kevin Weil43:08
Yeah, I mean, it's a really good question. You can think of Sora in some ways as a world model because you can't generate accurate video if you don't actually understand, you know, if it's not actually representing physics and other things appropriately, then your videos are going to look very wonky. So, in some ways, Sora gets you in a direction like that. Sora is a video model. We also care a lot about robotics, you know, both for science and for other reasons. And so, that's another, it's just one more reason why we care about world models. So, I agree with you. It's a big part of getting, a lot of the last few years has been AI in a digital space. I think over the next, you know, two to five years, we're going to really see AI in the real world flourish and you're going to need physical AI and you're going to need robots and you're going to need world models. So I do think they're a big part of the future.
A
Audience Member44:05
So like you were saying that for through innovation we're trying to create, you're trying to create models where not the most probabilistic answer is rewarded but reward systems are changed. And like I've been reading that most of the new development in AI is domain specific rather than broad now. So has most of AI research moved away from LLM towards new architecture? It's still just feeding more data?
K
Kevin Weil44:28
No, it's, I mean it's a mix of those things really. It's not like we haven't innovated on architecture at all internally but it's also, I mean these things are still LLMs at their core. You pre-train them and you know they go through a reinforcement learning phase and a bunch of other things. So no, I think, and we care a lot about getting transfer from one subject to the other as well. Like obviously we do need to do like say physics specific work to get the model to be better at physics but we also find that if you teach the model to be say really good at coding then it also gets more intelligent in other like technical and quantitative fields. So you want there to be some transfer and it kind of makes sense that there should be just thinking about, you know, humans and intelligence.
A
Audience Member45:16
But like now for instance Claude Code, it's creating sub-agents to do small tasks. So are the large scale LLMs these days also creating sub-agents that are accessing small fields? That's how most of the research questions are being tackled as well?
K
Kevin Weil45:30
I wouldn't say that's how most of the research questions are being tackled. Sometimes that's true. Sub-agents are nice also because you just don't pollute the context window and you can have an orchestrator model that, you know, I mean it's the same way if I give you a hard problem to attack, you would probably start by trying to break down the problem and solve it one piece at a time. And the sub-agent architecture is just one more way of doing that. So it's not that it like fundamentally changes everything. It's really just a good way of solving a hard problem by breaking it into pieces.
M
Moderator46:01
All right. Thank you.
A
Audience Member46:06
Thanks for the talk. So, one of the things that I personally find a little weird about AI is just how black boxy it is. You know, it's not really able to explain. For example, with that six-point gluon thing, right? I'm presuming it just like kind of spit out the closed form expression for, you know, n equals whatever. Is there effort at OpenAI to try to get the model to explain some of the logic of how it gets to that end answer and, you know, how's that going?
K
Kevin Weil46:44
Yeah, interpretability is a really hard problem. I think it's more or less unsolved at this point. We care about it because the better you understand it, the more you know, presumably the more you can also create outcomes that you want in terms of teaching the model to do certain things better. I mean, when you make the analogy to humans though, if I'm trying to tackle some hard problem and I have a flash of insight, do I necessarily, you know, like how interpretable are we? I'm not quite sure that we can live up to that ourselves as we think about things. So, you know, it's a, I don't know that we'll ever get a satisfactory answer, although I do think we're going to continue to get, you know, better and better at observability of these things.
M
Moderator47:30
Some hands. Okay. I think you've been waiting a while over there.
A
Audience Member47:35
Hi. I was curious about, you said how the role of like a scientist moves closer to like verification work of like an AI model. What do you do when the AI model is better at verifying model than the scientist?
K
Kevin Weil47:52
Well, you want to live in that world actually, right? Because then the model can be more productive. There are fewer bottlenecks. I think one of the major roles though will be, and we see this today, like if you take Alex again trying to solve physics problems with a model, he can do it better than I can. He has a better sense for what to prompt, how to prompt. He has better taste in physics problems than I do. And so there's a major role to play for experts. I actually think AI models, you know, they raise the floor certainly. They make everybody better if you, you know, they can be the best teachers in the world if you want a one-on-one education plan. You learn whatever you want faster than you ever could before. But they also raise the ceiling for experts. And I don't think that like, at least the point I was trying to get across was not like now all scientists need to go just be verifiers of hard problems. I actually would love scientists to get out of that world at least by and large and it's more about how do you, if these things give you superpowers, how do you direct the superpowers?
A
Audience Member49:07
That the more people ask questions, the more questions there are, the exponential curve.
Hi, my name is John. I'm doing my doctorate for computational sciences and he's 11. I wanted to get your perspective. Would you say that there's any particular sub-fields of mathematics or otherwise that these models are better or worse at or any?