Jakub Pachocki0:04
Thank you, Don. Hello everyone. Very excited to participate. I want to share a bit about where we at OpenAI see AI progress going. The big goal that we are working towards is automating research, automating discovery of new insights and development of new technologies. I believe this acceleration of technological progress is going to be the most important consequence of AI progress and of course it is going to become the core driver of AI progress as well in the future.
And there are still some technical challenges in the way. There's still a lot of research on alignment safety that we need to do. And I think there are also some big questions that are a bit less technical about what does it actually mean if we are able to succeed at this and what is the right way to use and govern such systems.
But let's talk about how the technical progress is going and I would like to highlight two results from the last few weeks that I have found very interesting. Both of them are the performance of our models in competitions. The first one is the AtCoder Heuristics World Finals. So this was a major international programming competition with the best contestants gathering in Tokyo to compete at solving a single very difficult problem over a 10-hour horizon.
And we have entered our system into the contest and we've seen the model was able to beat all but one person. And I thought it was a very cool result. Of course this is developed by a small team building on top of the more general AI machinery that we're building at OpenAI. And the model was able to find some interesting insights, it was able to optimize its code very well. But I think it's also very interesting to look at where the model fell short. And if you look at the one contestant that managed to defeat it, they actually figured out a quite novel approach to the problem that wasn't even in the model's search space. It didn't really look at that.
And I think this finding, this sort of new insight, is the thing that we're after. And so I think it's a very interesting thing to see that that is the thing that was missing there. And a very similar result happened in the International Math Olympiad this year where our model, similarly to DeepMind's model, was able to solve five problems perfectly and achieve a gold medal level performance but didn't make any progress on the sixth most difficult problem, which again required a fairly novel insight.
And so these are kind of similar results. We have these models that are able to beat most humans but they are missing this one big insight. And the IMO results I think is particularly interesting because this is something that's been a longstanding challenge in the field. I remember a few years ago we were talking about what are the AI milestones that have already been cleared, what are the AI milestones that we see in the future. A few years ago, the obvious ones, well, there was chess. There was Go, which was a long-standing challenge after chess. There was the ability to interact in natural language. And we were thinking of IMO as this great benchmark of is the model able to find these fairly general insights, is it able to perform this fairly general reasoning still in this very well constrained and measurable environment.
And one thing that wasn't clear at the time is should the benchmark be you got a gold medal performance at the IMO or you are able to solve six problems perfectly including problem six which tends to be a pretty difficult combinatorics problem. And so we were thinking oh maybe we should have two milestones. And at the time I really still didn't expect that we'll have this year where two labs independently say oh we have models that get a gold medal performance, they perfectly solve five problems but cannot touch the sixth one which happens to be a hard combinatorics problem. So I thought it was very interesting.
And I don't think that is accidental but still it makes you wonder what is next right because we are reaching this limit of these very clean-cut benchmarks, these very easily measurable milestones that we're used to for also measuring human performance. And we have to move to these fuzzier, longer horizon tasks where you actually develop insights and you develop things that have real world impact and actual value.
And part of this question as we go there is thinking about what is the mode in which the AI interacts with the environment right and I think that's very topical for today. Personally for a long time in the past I have been reluctant to push open research very quickly into extending tool use, interfaces and so forth. With my thinking being that while this is clearly very important, this will ultimately be needed, there's still so much progress that we can make in these more isolated environments and we need to pursue the kind of core intelligence of the model above other things. And it's not clear whether that was the right path obviously. And we did have a diversity of approach there. So for example, Jared Kaplan who now leads our RL team has been relentlessly pursuing this sort of capabilities for a very long time.
But in any case I think now we are reaching a sort of point of convergence where the models are smart enough that to even talk about to measure the progress on their capabilities, we need them to be able to more robustly interact with the world and to be able to tackle these longer horizon, more fuzzy tasks.
And I still want to be clear right, I think that this raw intellectual capability of the model is still going to remain at the core of our focus and I expect it to continue improving and maybe more rapidly than before. And maybe a few years ago we were thinking mostly about scaling pre-training, I think now we see there's also progress from reinforcement learning and new insights. But as we continue these axes of progress it is now very important to connect these models to the real world to have them pursue meaningful research results.
And if you kind of tread that line of thinking and think about what's coming after that, eventually it will become important for these systems to start interacting with the physical world and robotics will become extremely important. And another thing in this vein that's maybe a bit less obvious because we're already doing so much of it but obviously people are a huge part of the real world. And of course systems already interact with people quite a bit but I think especially for this reasoning, for these agentic AIs of today, there is still quite a lot of room to improve very meaningfully how they interact with us and I think there will be quite a bit of progress there.
A lot of our thinking on safety centers around this notion that we are building these agentic learning systems that are going to interact with the world, learning for a long time, and we have to think about how to keep them harmonious and beneficial right? And this starts from this innermost motivations of the system right? So thinking about value alignment and what is it really trying to accomplish? Going through can it follow instructions reliably? Going through robustness, I think challenges like prompt injection for example are going to be very important and challenging for the field. And all the way to physical and systemic constraints that we impose upon the model. So like what control do we have upon it without any reliance upon its behavior and can we monitor how it acts. And there's quite a lot of progress to make on these directions and we think about this quite a bit.
One thing I'd like to end with, that I think it's just interesting to think about. I think we have this kind of built-in notion that things that are truly interesting just have to take a lot of effort right? You cannot pull out just like very novel, interesting things that change your thinking out of thin air. And I think we've already seen AI systems bend this notion a little bit here and there. So for example, when you see AlphaGo discover a brilliant move that people haven't thought about, that feels absolutely magical right? It just feels like something that doesn't really make sense right? And just makes you feel a little differently about the world. For me personally, it was watching the Dota 2 bot play endlessly entertaining games, this sort of infinite TV with new strategies coming up.
And I think it's just interesting to consider what this will look like when we're talking about broad real world insights. And if you think about this world where new science just kind of falls out of GPUs, it's a very different, a very weird place to think about and I think that's worth reflecting on. Right, thank you for the time and looking forward to the panel.