About Jakub Pachocki
In a March 2026 podcast appearance, OpenAI Chief Scientist Jakub Pachocki discussed the company’s focus on continual learning, describing it as "the thing that we're building" and "what we're working toward." He addressed the use of math and physics benchmarks as proxies for general intelligence and noted that reinforcement learning is being extended beyond easily-verified domains toward longer-horizon tasks. Pachocki also expressed excitement about the "first proof challenge," a benchmark of unpublished problems from mathematicians and theoretical computer scientists, and recounted how an OpenAI model was prompted to solve those problems during a training run.
Pachocki stated that OpenAI believes its models are "capable enough to actually materially change the economy, change how things are done," and said the company feels "a lot of urgency about that." He also acknowledged challenges associated with automating intellectual work, including questions about jobs and wealth concentration, and said that "this requires real policy maker involvement."
Source: AI-verified profile updated from Jakub Pachocki's recent appearances.
Browse all interviews →
Transcript (115 segments)
J
Jakub Pachocki0:00
I definitely agree that continual learning is really the thing. It's really the thing that we're building. But I don't really think this is like a problem that's ignored and off the path of what we're doing currently. I think it is what we're working toward.
I
Interviewer0:09
What are like the other research areas within alignment that you're paying attention to or that you think are promising?
J
Jakub Pachocki0:13
A lot of the like longer-term challenge with alignment is about generalization. What are the values that the model falls back on?
I
Interviewer0:20
What are the things that you need to figure out to be able to really make models work well in some of these other spaces?
J
Jakub Pachocki0:25
I come back to this.
H
Host0:26
Jakub is the chief scientist of OpenAI. I think literally one of the most important people on the planet. And today on Unsupervised Learning, I got to ask him literally everything that I've been thinking about and I know a bunch of people in the ecosystem have too. We talked a lot about model progress, what's required to make long-running agents work, as well as the really interesting work OpenAI has done in the AI for science world and the progress he sees in that over the next years. We talked a lot about how companies should be thinking about model building in this moment, when they should be doing reinforcement learning, how they should be thinking about the evolution of harnesses and the impact that will have. We hit on a lot of his really interesting research, including the work he's done around alignment, the work that OpenAI broadly has done around math competitions. And we also talked about this focusing moment in OpenAI and what it means for the research organization and how he runs his team. Literally just such an awesome opportunity to talk to someone who is driving so much of the change that has revolutionized this space in the world. I hope folks enjoy this wide-ranging conversation as much as I did.
I feel like you are the perfect person to talk to about all the questions everyone has in the ecosystem. What's happening with model progress? A lot of companies are thinking about how they should be building things based on what's happening with the models. A lot of people at a societal level are thinking about the impact AI is going to have on science and broader society. And you've been at the forefront of the space for pretty much every generation of improvement these past years and so really excited to have you on the podcast.
J
Jakub Pachocki1:52
Happy to be here.
H
Host1:53
I think I'll start with one of the juiciest things you said which is, you know, four months ago I think you and the open team talked about aiming for a system with research level intern capabilities by September of this year. So coming up, I think that's what, six months from now. And then a more fully automated AI researcher by March 2028. And so I guess, you know, checking in four months later, how are you feeling about those timelines?
J
Jakub Pachocki2:16
Yeah, I think over the last months, the change that really happened is we've seen this explosive growth of coding tools.
J
Jakub Pachocki2:27
It's an understatement. Yeah, we've definitely like really kind of gone to a place in OpenAI where we use Codex for the majority of actual coding. And so I think for most people, the act of programming has changed quite a bit. So I definitely see this as a signal that something here is on track. The other kind of like very interesting update over the last few months to me has been the progress on the math research capabilities. Also the results we've kind of seen in physics and other fields. I think this kind of level of capability, this level of ability to provide insight when combined with ability to access infrastructure, ability to use maybe more compute at test time that's something that Codex is using currently, and very strong improvement in general intelligence which I also expect over the next couple of months. Yeah, it's something we're still very much planning for and very focused on.
H
Host3:29
And how do you know when you've gotten there? Like what's a workflow you might look to to say, hey, okay, I think we've got these research intern level capabilities?
J
Jakub Pachocki3:38
The way I would distinguish a research intern from a full automated researcher is the span of time that we would have it work mostly autonomously or the specificity of the task that has to be given. So I don't expect we'll have systems where you kind of just tell them, oh, like, you know, go improve your model capability, go solve alignment, and they will do it. Not this year. I think we might get there at some point, but I think for more specific technical ideas, like I have this particular idea how to improve the models, how to run this evaluation differently, I think we have the pieces that we mostly just need to put together.
H
Host4:18
Karpathy released, you know, a pretty viral version of using some of these models to improve some of his, you know, obviously way less complex models than what you guys are building here, but did that feel like generally in the spirit of some of what these tools might look like?
J
Jakub Pachocki4:36
Yeah, I think it's in the spirit. Yeah, I mean I expect it to look like a pretty continual evolution from kind of where Codex is now. I think towards a bit more autonomy, running for a longer time. But yeah, I think we'll see a lot of this sort of application. I think in general we'll see more autonomous and higher compute use of these models for different things.
H
Host4:58
You mentioned kind of like the math and physics side and obviously you've had these really impressive breakthroughs in math on some interesting different kinds of competition problems. Maybe, you know, I think for our listeners it intuitively makes sense how progress in coding directly translates to something like helping with AI research. How does math and physics progress also tie into this?
J
Jakub Pachocki5:21
The biggest role that focusing on these math benchmarks has played for us is as a general benchmark and a north star for how to improve this technology. Like math is very measurable, right? It's much easier to tell whether you've actually solved the math problem than whether you've even produced a good piece of software. And also it can get very hard, right? So you can have things where it's very definite whether you've solved them, but it can be arbitrarily pretty much hard to actually solve them. You know, I would say up until not too long ago, my perspective has been like, well, okay, our models are not maybe able to solve simple math problems. Okay, our models are able to solve simple enough problems but are not able to solve IMO level problems. So clearly there is just a gap in the intelligence of these models that is very measurable, very easy to run at. It's very clear what we need to do and this has been kind of our north star for reasoning models and so forth. Now of course that is changing quite a bit, right? And we have kind of reached these milestones that we've been working towards of IMO gold level, solving IMO problem six, and making progress into research level mathematics. And from this point, I think there still is utility in continuing to measure progress on this. I think there's also definitely transfer that you can get from getting better at mathematical reasoning to getting better at AI research. You know, a lot of our best researchers are mathematicians we're training or from other theoretical fields. But definitely we are very much changing how we think about these north stars and we are very focused on how the models, the next models that we're producing are actually useful in the real world, useful especially for research but also for other economically valuable activities and for other fields of science and especially maybe more applied sciences. And the reason for this shift is because we believe the models are now capable enough, not as smart as people always, but capable enough to actually materially change the economy, change how things are done. And so, yeah, we feel a lot of urgency about that.
H
Host7:43
In the early days, picking a domain like math that is so hard to solve, but then easy to verify whether you did it, like it's kind of the perfect place to get started. And I think code obviously shares a lot of attributes to that. You know, possible to check and verify and great for reinforcement learning. I think one question that a lot of people are thinking about is, okay, we've seen reinforcement learning work incredibly well in these domains where you can verify it rather easily. A lot of valuable tasks in the world, medicine, law, finance, you know, there's some level of the ability to do that, but it's certainly not to the same extent that math and code are. And so I think a lot of people are trying to figure out, you know, are we going to see similar improvements? You know, obviously code and math, the rates of improvement have been so astronomical and shocking.
J
Jakub Pachocki8:26
Yeah, I definitely expect so. I think an interesting duality that we think about a lot is for these more general tasks, for these tasks that are kind of harder to evaluate, they share a lot of commonalities with just longer horizon tasks, right? Because if you think about even a very well specified math or coding problem, again, if it's something that you need to work on for like a year, then even if it's very clear what the criteria of success are in the long term, like what to do on your first day of working on it is a pretty open-ended problem. And so I kind of believe these difficulties coincide and they're very clearly the next frontier for how these systems develop. And I think we've definitely seen very encouraging signs both on just our ability to scale RL on these more general domains. And I think also we can scale efforts that have a lot of promise.
H
Host9:24
In these other domains, it feels like one of the hardest things to know is just what was success in a task, right? And you can imagine there's going to be whatever the problems you are that are facing code or math that are short-term tasks and then longer-term tasks, feels it will be amplified in the space that is outside of those, right? Where a short-term legal task or medical task may be harder to run thousands of iterations on, right? And figure out was that done correctly and then those longer-term tasks like even harder. I'm curious like how you even conceptualize that research challenge, like what are the things that you need to figure out to be able to really make models work well in some of these other spaces.
J
Jakub Pachocki10:02
Yeah, I think I come back to this reality of like how do we make the models work for a very long time and how do we teach them to evaluate kind of partial progress.
H
Host10:12
Yeah. I mean I think if you look at even outside of RL, like where that sort of progress on longer horizons is coming from, right? Like I mean as the models kind of become more consistent from just pure supervision in pre-training, they gain some idea of like, you know, what does a good partial artifact here look like. And so I think even if we weren't scaling RL very meaningfully we would see an elongation of these horizons over time.
J
Jakub Pachocki10:40
Yeah, it's definitely a research challenge to figure out how to leverage these new ideas from RL and so forth to apply this to general domains. But I'm quite optimistic about that.
H
Host10:52
Yeah. And it's interesting. It sounds like part of your mental model is like the models themselves being able to check progress with some sort of cadence that is, you know, reliable enough from the outside at least. It's not totally clear if we've seen like generalization in RL yet. It feels like we clearly you seem to have some techniques that really optimize models around whatever we choose to focus on but it's like almost feels like an older school version of ML of like one thing at a time. Is that like, you know, I guess would you agree with that characterization and like how do you kind of see this current climate?
J
Jakub Pachocki11:21
Well, we are buying a lot of compute, right? Because we still believe a bit less and we believe more than ever to some degree. Yeah, we've seen new techniques and I think new ways to scale, but that is kind of the lens through which we've been viewing things. Yeah, I think there is a certain amount of complexity that we need to grapple with and kind of everyone needs to grapple with because, you know, we're no longer really purely building a brain in the sky that's completely isolated from the real world, right? Like if you actually, if you want this model to do medical research, if you want it to cure cancer at some point it needs to learn about the real world in a meaningful way, you know, maybe conduct some experiment and learn from its results. And for that you need to figure out how to actually connect it, right? And that is going to involve something that goes in the direction you described but I don't think that goes counter to actually finding and scaling the simple algorithms that we've been developing.
H
Host12:32
I feel like I talk to a lot of companies and one of the main questions everyone seems to be asking these days is like should we be doing our own reinforcement learning? Like take an open source model and we have some data on a task that people do. We have evals because we know our domain pretty well. Is this something that makes sense for us to do or should we just wait for the models to continue to get better at some of these things? You know, what advice would you give for the many builders that listen to the podcast as they think through the extent to which they invest on the reinforcement learning side?
J
Jakub Pachocki13:02
Reinforcement learning definitely can be a very data efficient way to really improve the model at some sort of task, right? There is a much more data efficient way of learning that we know, which is like learning in context, right? And this is maybe the most fundamental way that people teach these models, you just prompt them with examples, with instructions for what you want. I expect that learning is going to get much better over time. And so I think it definitely really matters that the models can adapt to your context. They can adapt to the kind of tasks you care about. So I think that will be very important. I'm not sure if replicating the current pipeline is going to be the right way to go about it. But yeah, it's definitely a problem that we're thinking about.
H
Host13:46
Yeah. So it's almost like, yeah, you still have to do the work, like you still should figure out what the evals are that matter, gather the data, the examples, but it may just turn out in the future you're far better off just feeding that into this context than trying to do anything on your own model.
J
Jakub Pachocki13:57
Yeah, I think that's quite plausible.
H
Host14:01
And I think that, you know, obviously people have seen the success of tools like Codex which I know you've obviously been a key part of and wondered like, hey, do we need to build our own harnesses or our own ways of using these things for our own domains whether it's legal or finance or healthcare or do we kind of just take the harnesses that the large models do and use them within the context that we have. Any thoughts around that?
J
Jakub Pachocki14:32
Like the implementation of the harness shouldn't really be a limitation for a very long time. I think we'll be able to get much more general harnesses that people can use for all sorts of other domains. I mean I think Codex is pretty good actually if you try using it for things beyond coding.
H
Host14:46
That's so interesting. Like a much more general harness being something that's almost adaptive to or just works across whatever the specific set of tools you have in your domain or specific set of things you want to expose to the model.
J
Jakub Pachocki14:59
Yeah. I mean I think and it's also worth thinking about like what is the ultimate interface that we want to interact with the model with. So the models give some UI hard for humans, right? They can build their own UIs. They can do things that people would find very time-consuming. But I definitely think there is also just a lot of space to enable the models to access the current interfaces that we use for people, right. So I think like we want to have AIs on Slack for example that are plugged into our context and are able to learn from it and able to realize these existing things, right. So definitely there is some meet in the middle here but definitely I believe long-term, by default the AI should meet you where you are and if not that would be because it has new abilities, not because it has limitations.
H
Host16:01
Yeah, it's an interesting point that basically today it feels like these harnesses are so bespoke to certain environments, but over time as you add more and more skills and tools and models can navigate across those effectively, it's like there'd just be a general way, like the way humans have, that makes a tremendous amount of sense. I guess I'm curious, you know, obviously I'm sure like every day you see kind of crazy stuff on the research side at this point, like what are the milestones that are still meaningful to you as you think about like it would be pretty crazy if I did a run one day and saw like X or Y, like what are the things you're paying most attention to?
J
Jakub Pachocki16:38
Yeah. I mean at this point it really is about research, right? Like is it about, can the model discover new things, can it execute on a longer horizon research problem.
H
Host16:54
It's almost like looking for some sort of insight that you're like, oh, someone on my team had come up with that, that would have been pretty intrigued by.
J
Jakub Pachocki17:00
Yeah, we've actually had some minor but I think quite impactful ideas come from even like GPT 5.2 Pro that we're using entirely. But you know, I think it's still very very small compared to where I expect it to be.
H
Host17:16
Yeah, I mean it seems like almost inevitably these models are going to get better. They will be used in research. They'll be used in science more generally. You're like one of the first people interacting directly with these models as like research partners almost at this stage. Anything you've learned around the right way to do that or do you think about like what a research organization, as these models continue to get better, might look like?
J
Jakub Pachocki17:34
Yeah, I think we're definitely kind of at a transition point where the short-term immediate quality of the model is about to be a quite determining factor for the pace of our research progress because the models are going to drive a lot of that. And so that definitely requires rewiring some intuitions about how to run a research organization. You know, normally you kind of try to not be too focused on immediate quality. You try to be much more focused on the longer term. I think we have a lot of very exciting stuff queued up that we are kind of working towards but I feel a lot of urgency to kind of yes, to actually execute on it and to actually use these advances in model intelligence to accelerate research on the AI and especially AI alignment.
H
Host18:24
Yeah, it's such a fascinating point because I've heard you talk before about running a research organization and I feel like in the past it was like giving people the space to pursue a lot of things that weren't directly, you know, hey, this is for a month or two months of progress, but it's like what are the ideas that are really going to drive things forward. But it makes total sense that we're in a time now where you're like, look, everything we do will be so much better if we just focus on this in the short term and make it better. It must be fascinating to navigate that and these maybe further off research ideas at the same time and running an organization.
J
Jakub Pachocki18:56
Yeah. Yeah. It's definitely something we spend a lot of time on with Mark nowadays. Yeah.
H
Host19:02
Right now you have a ton of compute as a company, but you obviously you have great scaling laws on the pre-training side, you have great scaling on the RL side, you have probably lots of experiments going on that have nothing to do with either of those vectors, but are like interesting new ways. How do you even think about allocating compute across all of this stuff?
J
Jakub Pachocki19:20
Yeah, it can get very complicated, right? Because there's so many things that we need to do. One thing we've been, one kind of discipline we've started keeping is we try to make sure we just explicitly budget a large chunk of our compute to the most scalable methods, to the things that we believe are the most responsible for driving general model intelligence. And you know, even if it's not the most efficient allocation of compute at all times because, you know, if you're allocating so much compute to one experiment or one set of experiments, there's so many things you can accelerate a little bit of that compute elsewhere. But I think it's easy to kind of with all the interesting and important things that we're doing, I think it'll be very easy to kind of paralyze all of it and not really end up doing the things that we believe are most important. You definitely want to understand the empirical evidence. You definitely want to make sure your evaluations are in order and the experimental rigor is there. And then you also want to apply some regularization based on like, okay, do we understand this method? Do we actually expect it will scale? Do we expect this is something you can actually build on in the future? Is this kind of a one-off? Right. And based on that determine the priority.
H
Host20:24
Yeah, it's so interesting. Probably find all the ways that you know you could improve things but they feel maybe like off a little bit to the side of where you think the overall arc of progress is and so you end up leaving some of these like low-hanging fruits to some extent because really the most important thing is finding the future direction and then the scaling within that and devoting compute toward that. Obviously the place where we talked about Codex a lot and the success of coding and it feels like last year was like the year of just incredible hill climbing on coding. I'm curious, you know, obviously Codex has been a super successful product in many ways, like Anthropic was kind of first to this market, you know, Claude Code was a dominant product there. What do you kind of, reflecting on that, I guess what do you make of the success Anthropic's had in this space?
J
Jakub Pachocki21:07
Yeah, I think it's a matter of really focusing your product direction on where you believe the next application of the technology is, right? And if you look at the prioritization we've had on our product, I mean we have been working on cutting products but they have kind of been a secondary thing compared to our main priorities. And the interesting thing is that is not very reflective of the priorities of the research organization within OpenAI. I think given that we've kind of had this explosive success of ChatGPT, you know, I think ChatGPT is going to evolve quite a bit but as it was in 23, right, is this particular product that's maybe not, you know, I think it's definitely quite aligned with our vision of where AI is going, but it's not really representative of everything that it enables. And so the majority of our work in research has been focused on that future thing. And I think increasingly it has decoupled from our short-term product strategies, right? Yeah. I'm very confident about the things we've been building and the things we are building on the research on the model intelligence side. You know, a lot of our reprioritization and increased focus on the product side is about actually getting to deploy them and the belief that actually they are the thing that really matters now.
H
Host22:38
Yeah. And now it feels like clearly the whole company priority is so locked in and focused on this and you've seen just incredible improvement in Codex in recent months. For all the developers that listen to the podcast, like if again it's almost like hard to comprehend what the world looks like as these models keep hill climbing on longer and longer tasks, like what do you think will look different in their lives or how will they be using Codex in three, six months? I realize three months and six months are very different timelines in this world, but take whichever, whatever in between point you'd like.
J
Jakub Pachocki23:08
I would expect just a gradual increase in just the level of autonomy you feel comfortable for the model, just the vagueness of description that can work with, the level of supervision it needs. I think we're not very far from models that can work autonomously for a couple days. Maybe use quite a bit more compute than they're using now and produce much higher quality artifacts on their own.
H
Host23:32
Do you have a gut instinct on like, you know, there's always been this question of like will the world, you know, do you need that software engineering skill set to supervise these models running for a few days or like, hey, does it turn out at some point of being able to run for a while, anybody can use coding agents and supervise them to some sort of output?
J
Jakub Pachocki23:49
I mean I think definitely for a lot of outputs you already don't need much experience, right? I think still the distinction I would draw between an intern here and really an autonomous researcher or software engineer would be that if you want to build something bigger, you probably still want to apply supervision, you still kind of want to have an overarching thing, you want to recognize what building blocks fit in and which don't. But yeah, I definitely expect that desired skill set to shift quite a bit over time towards this more general vision setting.
H
Host24:23
You know, I guess on the research side, I feel like there's been, maybe like a month ago, I feel like all anyone could talk about was continual learning and there's just, you know, it was in the Zeitgeist. There's all these new labs starting to go focus on continual learning. Some folks left OpenAI to go focus on that. I'm curious, I think part behind that is a belief that RL alone either won't get us there or will get us to some level of very inefficient scaling and it's kind of different than the way humans learn. I think even I've heard you say before that RL is still very different today than the way that humans learn. What's your take on that whole movement?
J
Jakub Pachocki25:02
Yeah, I am a little bit confused by it because in my mind the whole excitement that we've had, I mean even if you look at the titles of the GPT-3 paper, right, like it is that this class of models is actually capable of continual learning, right? It's capable of learning to learn in context, right? That has been really the driving force behind the excitement to scale these GPT models further. That has been the premise for why we really need to teach them with RL to learn in context more efficiently. And so I definitely agree that continual learning is really the thing, right? Like it's really the thing that we're building, but I don't really think this is a problem that's like, oh, you know, it's kind of ignored and off the path of what we're doing currently. I think it is what we're working towards.
H
Host25:54
Yeah. Like in your mind, this is like the single best path to get there is to continue to kind of scale the pre-training and RL.
J
Jakub Pachocki26:01
I think that is kind of how we've made the most progress on this problem so far and you know, I think there are, I think that there definitely are more ideas, more steps. I think also a lot of improvements that will just come from scale.
H
Host26:12
Yeah, and I guess like, you know, we have a lot of folks listening that maybe have been able to do a lot of simpler things with these models and then they try to do some of these more complex, you know, I don't know, call it 100-step or longer-term tasks and they're like, oh, the models don't work for this yet. And I think it's harder, you on the inside constantly feel this improvement but for them it feels like, hey, this is like night and day away from being able to do this much longer thing. How do you kind of articulate to them the set of things that need to be true for these much longer steps to happen? Is it around kind of checking in more often as you were talking about before or I feel like there's just this belief among the research community of like, oh, all of these tasks will be solved in the next year or two and then in the wild a lot of people maybe not totally grokking that improvement line that we've been seeing.
J
Jakub Pachocki26:58
Yeah. I mean I think a lot of that prediction comes from just looking at historical improvement lines, right? And but I think increasingly we can roughly see the shape here. I do think a lot of this is about just the models becoming intelligent enough to recognize whether they're making progress. I think some of this is like this very pragmatic work of like are the models actually, can they actually access all the context, all the files, all the infrastructure they need to do the work you want them to do. Which yeah, I remember like in the past when we were discussing the road map that we're taking with RL, you know, I definitely view like, okay, we just need to teach the model to kind of reason with its own tokens as kind of the priority and then of course we'll need it to use tools like the environment, you know, at some point we definitely need to teach it to see, right? At some point, we need to teach it to use a physical body, right? Like, but yeah, I mean, I think we're definitely well into the stage where it really needs to interact with the environment and it really needs to see and you know, someday soon we'll really care about robots, but yeah.
H
Host28:02
Yeah. I mean, it does feel like a lot of the times when I hear people complain about, oh, a model can't do X or Y, it's like literally just because you haven't fed, or connected it to systems or fed enough context into it. Actually, I do wonder if context was universally applicable and able to flow into these things. Like I feel like a lot of these problems would actually just be solved with today's models. You know, I want to talk about some of the AI for science stuff that you guys have been working on. And one thing in particular, you know, I feel like the coding stuff is something that everyone feels very viscerally, in every company they're using these tools and getting tons of productivity. You know, on the math side, not all of us competed in IMO competitions and necessarily have as much of an intuitive feel for some of these breakthroughs. And so one of them I know that was really interesting that you guys did is you used some compelling work around like first proof, right? And I think these are like very different problems than kind of traditional competition math. I wonder if you could just speak a little bit to that because I think it's just a space that our listeners might be less familiar with and kind of less familiar with understanding the implications of models being able to do pretty cool work here.
J
Jakub Pachocki28:58
Yeah, I mean, you know, I think yeah, I was very excited with the first proof challenge and you know, again, like I kind of, you particular one is kind of a
Benchmark, right? It's like a couple of respected mathematicians and theoretical computer scientists releasing problems that they believe are representative of their day-to-day work but haven't been published anywhere, so that we can really have our models take a crack. We were so excited about this challenge, but it was kind of dropped without any advanced warning with a week-long deadline to actually execute. We had a very exciting model training at the time, and so one of the people in charge of training, James Lee, kind of started prompting that model just by hand. And actually kind of seeing, oh okay, it's actually solving these problems, was really a fascinating thing to see. One of these problems actually is from a domain that I did my PhD in, and yeah, seeing the model kind of come up with these ideas which I would be quite proud to come up with in a week or two, seeing it come up with them in like an hour or so, that was a very weird feeling. I think in the past when I felt like that was when watching our data bot play just very interesting data games infinitely, and it feels like there's some sort of magic happening because interesting things should not be indefinite.
J
Jakub Pachocki30:32
And so seeing that happen for math, for something that I believe is actually quite representative of our work or a precursor to a lot of the work that we're doing and a lot of the work that really matters in the world, yeah, definitely really increased my feeling of urgency.
I
Interviewer30:49
One thing that's fascinating too is the idea that you're training these models and it's like you throw these problems in and nobody knows how good they will be at solving them. And I think it must just be fascinating to see something that you know so well and a space that you spend so much time in, and realizing hey, probably the previous generation of models wouldn't have been able to do that, and you wouldn't even have thought necessarily that this was the benchmark to do, but it's just generally showing the general purpose capabilities and improvements of the models.
J
Jakub Pachocki31:16
I mean, it is at a stage where we needed to seek out experts in the particular domains to be able to tell us whether these particular proofs are correct or not, but it's still much easier to tell whether you've actually made progress than for something like even coding, right? Because sure, competitive programming you can evaluate, but most programming is not competitive programming, and it's about are the abstractions right, are you handling all the cases.
I
Interviewer31:41
Yeah, I guess there was this maybe common criticism a year ago, and I don't know if it's as strident now, that these models are like pattern matchers, but you really want AI for science, we're not going to get new ideas or entirely novel things out of pattern matching. It feels like we continue to chip away at that narrative. Are we getting closer to fundamentally disproving that?
J
Jakub Pachocki32:02
I believe so, yeah. I mean, I think kind of on schedule we're starting to see minor advancements, right? Not huge things, a small idea here or there. I mean, maybe some bigger papers in collaboration with scientists. But, you know, was AlphaZero a pattern matcher? AlphaGo a pattern matcher? Our data bots, they did kind of come up with new strategies for the respective games.
J
Jakub Pachocki32:30
It's funny that there are counterexamples to it all the way back to 2016, 2017.
J
Jakub Pachocki32:33
And you can say, well, I guess you can always fall back on flaws in that, which I think is interesting. Like AlphaGo can be beaten with some strategy, our data bots could have been beaten with some strategy. I think there will be a lot of definitions for a while of these models. But I think also they are able to discover new things because they have a lot of these capabilities. And the way, you know, it's taken a couple years to go from very tiny game environments to much more general scientific research. It required kind of going through a decent approximation of all human knowledge in the meantime and learning all the human languages and so forth, but I think the basic principle is very similar.
I
Interviewer33:23
Yeah. You know, it's funny. I think when you guys had these first proof results, I remember the organizers said they were commenting on these AI solutions and they were like, this feels like 19th century mathematics of brute force, computation-heavy approaches rather than these elegant modern techniques. Which I'm not sure is a feature or bug of obviously the way these models work, but hearing that, does that concern you, excite you?
J
Jakub Pachocki33:47
It doesn't concern me. I mean, I think it's expected. I'm sure I thought for at least one of the problems that actually our model produced a pretty nice proof that was quite a bit shorter than the intended one. But I think in general you would expect these models, they can produce so much more reasoning in a short time than a person can, just in terms of raw number of tokens or thoughts. I don't expect that to be a long-term feature.
I
Interviewer34:13
It feels like there's so much momentum behind AI for science right now, and you mentioned obviously at some point you do have to connect these models to the physical world. You guys released some cool stuff with GKO and some of these other things you've been experimenting with. I'm sure you've thought a lot about AI for a bunch of different areas of science. As you've dug into some of this stuff, have you developed any intuition for as you think about three years from now, the spaces of science where you're like, oh, there's going to be crazy progress there versus the ones that might prove a little more resistant to immediate change? A tempting answer would be that it's really about what are the things that kind of require some manual work, where the models are not quite plugged into the ecosystem, or that the different laboratories will also evolve pretty quickly to adopt these new technologies within those STEM fields. Obviously, there's a question of is it an LLM with access to the physical world, or you've obviously had companies that have been started specifically around these domains, right? Like Isomorphic in biology, or Periodic in material sciences, or Physical Intelligence in robotics. What's your gut instinct on the extent to which it makes sense to pursue some of these things independently with different model architectures versus all within the context of one place?
J
Jakub Pachocki35:30
Yeah, I think it's kind of similar to my answer about the UI for Codex, which I would build around the capabilities of a technology and not around its limitations so much. So you definitely, if you have something that can suddenly design a huge amount of interesting chemical or biological experiments, yeah, it makes sense to build labs that enable that. I think if we did get to a place where the model is very capable of designing high-quality experiments, it also makes sense to have it work with humans in the loop. We shouldn't think of it as either you kind of automate it fully and you have this fun thing using some tools on the side. We will get to a world where it's just very natural to be collaborating with AI scientists that are working hard on a problem.
I
Interviewer36:16
Yeah, it's so interesting. It's almost like a different vision. It's like one world where this works is you just train a model to basically run these end-to-end tasks and be the automated biologist or chemist or whatever it is. And there's another one which is you're building really tools to both propose and run kind of work in tandem with a bunch of human researchers.
J
Jakub Pachocki36:37
I mean, I wouldn't necessarily categorize it as... of course there are tools in some sense, but I think we will get to a point where they're driving a lot of the design and ideation for the whole process. Yeah, with an LLM architecture, but just being able to figure out the right way, the right kinds of experiments to run and then actually design it. And yeah, when it comes to different architectures, for sure, natural language reasoning, the kind of things that we're prioritizing, that gives you a lot of generality. There are things that you kind of want to train a different model to model. I think even if you want to create a very good video model, I don't think large language models are the most efficient way to go about this, although they might result in the best model eventually. But I think it's similar for protein folding or other tasks of this kind.
I
Interviewer37:29
Yeah. So you think it makes sense to have some independent efforts around that, but obviously that will end up being paired with a core, really good researcher large language model that is helping drive a bunch of this stuff.
J
Jakub Pachocki37:40
Yeah. I want to also make sure just to talk about AI safety because I think that's an area that you've done a lot of really pioneering work on. And I'm not sure all our listeners will be familiar with, you actually did some really interesting work across the labs, right? And were focused on chain of thought monitoring. And so maybe to start, just tell us a little bit about that work and what you found.
I
Interviewer38:00
Yeah, so this is a realization that we had around the time we actually saw the first reasoning models of kind of the current crop. We realized that, okay, this works, and we were pretty, you know, we were thinking a lot about what this means. We kind of were like, okay, probably the world really changes over the next year or two or three. We were thinking what this means for safety and for our ability to kind of understand what these models are doing. And we realized that because of the way we train these models, because we don't supervise the reasoning process directly, it's not like ChatGPT is trained to kind of be polite and nice.
J
Jakub Pachocki38:43
And it always tells me I have great ideas.
I
Interviewer38:45
Yeah. Well, you know, that's a separate issue, right? But even assuming it's aligned exactly in the way we would want it to, which is definitely not, it's still kind of not going to be, there are just still some things it's not going to reveal about its motivations and time because maybe it would be unsafe or maybe it would be unkind, or maybe because it's not, maybe it's actually not aligned the way we think but it wants to hide that. And the way we train the reasoning models, the train of thought doesn't have any of that. It's not optimized to be in any particular way because it's just not directly trained. It's only trained in how it relates to producing a high-quality output. And we realized this is actually a very powerful paradigm for being able to interpret what the model is doing. It's actually not a very different idea from mechanistic interpretability, because in mechanistic, the idea is again you kind of have this model, you have these activations of the model that are not directly supervised to predict any label. They're kind of indirectly supervised, but the model kind of has never been trained with any sort of inspection of these activations. And so these activations might reveal something about its inner workings. But the big advantage of the chains of thought is that by default they are in English, and so it's so much easier to understand what is going on, especially as the concepts get more advanced. And the other interesting thing is, we were just talking about how probably how we believe in the future where these models work for a very long time, they work autonomously, and so there is much more of this reasoning. And so if this is a big axis of how the capability of these models increases, that our ability to supervise them will scale commensurately. Yeah, this really comes down to this principle though that you're not supposed to supervise the train of thought. And so this is actually something when we originally were releasing the preview model, we made this decision to hide the chains of thought.
Yeah, I remember.
J
Jakub Pachocki41:01
And for me that was the primary motivation. That was the reason, I didn't really even want to consider releasing it in different ways. There definitely was a bit of internal discussion about this, but the reason I felt very strongly that we should just hide it is because of this. Then there was this other concern that I didn't initially think about but I think was also very valid, of like, well, this model is going to be distilled to some extent, and that's definitely also been a big factor here. But I actually think that allowing the models some sort of private space... oh, and by the way, why do I think it's important that we don't show this chain of thought in product? If I'm saying the important thing is not to supervise them during training, well, I think if we did show it, if we established a paradigm where you just show these chains of thought in product, eventually you kind of have to train them, right? Like you'll have to train them for the same reasons you have to train whatever models you ship. And I just think that...
I
Interviewer42:02
We might not all want to know what the chain of thought our model has that gets to a response for.
J
Jakub Pachocki42:05
Right. I mean, I think it'll be useful to some extent, and we are trying to capture most of that value, either with chain of summaries, which I think are kind of a little bit of a stopgap. I think the longer-term solution here is having the model actually talk to you in real time, which the latest version of Codex kind of does, the latest version of the reasoning GP models kind of do, but I think that will get much better. But yeah, I think there's something very exciting here about just not having the training signal fight against us. And not... yes, because I think if you want to be able to understand what the model does in the long term, but you're scaling a method that is kind of going directly against that, you're probably not going to have a good time, right? That's the other side of the better lesson. And so this decoupling, I think, is an idea that gives me a lot of hope for our ability to at least understand how these models' motivations and generalization evolve as they get better, as they work for longer. I don't think it's a complete solution to AI alignment by a long shot. I think it's just another tool in our toolbox. But I am hopeful that building our toolbox with technical tools like this, we can actually continue chipping away at the fundamental problems here.
I
Interviewer43:27
Yeah, it seems like almost over the medium term, it's something that's going to be incredibly helpful. Probably not the catchall solution for long-term alignment.
J
Jakub Pachocki43:35
Yeah, I mean, I think it's a tool that can help us understand, I think it's actually very useful to build understanding of long-term alignment. For example, there has been this very exciting work from a planning collaboration with other labs on model scheming, where they investigate depending on what environment you put the model in, how you train it, is it prone to start having hidden objectives that it pursues, and what enables that. That whole line of work is chain of thought monitoring, right? Is this notion of, oh, you can actually inspect what the model's motivations are. And I think from that, that might take us in a completely different direction in terms of mitigations. Maybe the right way is changing the pre-training data of the model, or maybe it's something like the inoculation prompting from Anthropic. I think those are very interesting ideas, but I think having this ability to understand is very helpful to evaluate these.
I
Interviewer44:30
Yeah, it's almost foundational for any further area of research. What are the other research areas within alignment that you're paying attention to or that you think are promising areas to focus on?
J
Jakub Pachocki44:42
Yeah, I think a lot of the longer-term challenge with alignment is about generalization. We can train our models to do well, and or at least mostly to some extent, we can mostly kind of control their behavior in the things that are in distribution that we train for. But the things that are worrisome is what happens when a model is asked to do something very different, or it finds itself in a very different situation, or it's much smarter than it ever was before and it has all these capabilities. It's like we haven't really thought about how to train for. And so I think the study of this kind of longer-term value alignment is really a study of generalization, like what are the values that the model falls back on. One line of research I'm very excited about here, and something that we're investing in quite a bit, is understanding how that generalization falls back onto the pre-training data. And yeah, I think there's quite a lot there.
I
Interviewer45:53
I guess over the last six months, have your concerns around alignment increased, decreased? Like how do you, where are we kind of trending overall with this work?
J
Jakub Pachocki46:01
I will speak to the longer-term challenges of alignment, or what happens when you have very smart models. The way my thinking about the problem has evolved over the past few years has definitely kind of gone from, you know, oh, is this like very nebulous problem that is just very hard to even grapple with or define, to like, oh, you know, I think we can actually make progress at it by very concrete technical solutions and technical insights. And this is why we've really been viewing alignment as just a core part of research and really making sure that we are designing our reasoning models thinking about this, and we are conducting our alignment research with these reasoning models in mind and so forth. So I think my general belief that there's a research path here that actually gets us to an extremely happy world has increased quite a lot. At the same time, my timelines to very capable models have definitely decreased a lot. I think we're not that far. Again, I don't think these are models that are smarter than all of us, but I think these are models that are just very transformative. And so I'm quite optimistic we can keep a good grip on how we're doing on the alignment problem, how to roughly evaluate the risks of our models or the problems with them. But I do think we have to be, as an industry, really prepared to take trade-offs and possibly slow down development depending on what we see.
I
Interviewer47:37
It's already interesting to see a lot of this work happening across the major labs. The fact that you did this in collaboration with I think Anthropic and DeepMind, and it seems like, has that just come up organically, or is there a lot of alignment talk between the major players, given I guess the three of you are really at the forefront of all this?
J
Jakub Pachocki47:54
There's definitely some, I mean, there's definitely shared interest in these topics. Yeah.
I
Interviewer47:57
I want to shift a little bit to going inside OpenAI. I feel like no company probably in the world has been more interesting over the last two, three years, and particularly what it's like to run a research organization. We talked a little bit about this previously, but you talked before about how an important part of your job is giving researchers the comfort and space to almost be cave dwellers and think about what the models will look like in a few years. We were kind of alluding to it earlier. We're also in a time where it feels like there's just a massive competitive race, and everyone's going really gung-ho on these coding models. I'm wondering how do you actually operationalize this balance today, and anything you've kind of changed in your thinking overseeing this organization around the right way to do this?
J
Jakub Pachocki48:45
I focus on just high-quality experiments, recognizing are we actually making progress, being honest with ourselves, and promoting honesty about the results. I don't think that has changed. And even though our work will evolve a lot, I believe we still have quite a lot of work left to do. And so I don't think it's like, oh, we need to wrap up all our projects very quickly. So yeah, I don't think those fundamentals change. I think what does change is a level of urgency to really kind of bring some of these things that we think are most promising to fruition.
I
Interviewer49:19
And then obviously, there's been some very public internal moments at OpenAI over the years. You've been here for a long time. As you kind of reflect back, what were some of the difficult decisions that you guys made that maybe were like 51-49 that really defined the company, or any as you think back of the movie of the last seven, eight years of your life, the key moments that kind of stick out to you?
J
Jakub Pachocki49:41
Well, yeah. I mean, there's certainly a number of dramatic moments like this. I think the ways the company underwent the most change is not really these snap decisions, but more like shifts in how it operates. I would say OpenAI has gone through a couple phases. When I joined at the start of 2017, it very much felt like a very academic lab pursuing a lot of different ideas, not so much the scaling pill in practice. And I think that was the first big change with the data product, with GPT, we've kind of moved to, okay, we actually are going to have to buy big computers, we're actually going to have to scale things, we're going to have to develop the science of scaling, we'll have to develop the infrastructure for it. And so that kind of started the second phase of, okay, now we're scaling. We're still going to pursue a lot of these basic research ideas, but we are going to evaluate them for, are they scalable? Then yeah, then there was this interesting period I talked about earlier, where you kind of have ChatGPT as this big thing.
I
Interviewer50:52
Yeah, I mean, I thought it would look a little bit differently, right? I was actually surprised that text models are actually kind of the first thing. I thought we would be in a world where it's more the kind of video-style uses of generative AI are kind of the first big thing to take off, and we'll have to trade off pursuing the kind of longer-term text-based research. So yeah, but I think definitely we anticipated that this sort of tension would arise, where you have a thing that is kind of popular now, but you believe it's going to evolve quite a lot before you get to where you're going. And so I think that's kind of the phase we've been in for a while. And yeah, I think now we're like, well, yeah, I mean, we believe we are kind of starting to be in this phase where we're actually deploying AGI, or deploying models that are actually very economically transformative.
J
Jakub Pachocki51:53
No, it certainly seems that way.
I
Interviewer51:55
Well, I guess we always like to end interviews with a standard set of quickfire questions, which are basically me just stuffing all my overly broad questions I couldn't fit anywhere else. So if you'll shamelessly indulge me, I guess to kick it off, would love, what's one thing you've changed your mind on in the AI world in the last year?
J
Jakub Pachocki52:12
Yeah, I mean, I think it's really starting to reconcile this tension between the AI that you build ultimately is something that affects the world, but until you kind of get pretty close, it's a pretty theoretical thing that you're just kind of training and developing algorithms for. And so recognizing that, okay, now we actually need to make a lot of progress and focus on how we're actually deploying this technology in the wild. This is definitely something I've been thinking about a lot lately.
I
Interviewer52:50
Yeah, it's so interesting. Basically, outside of chat, it was almost more in the abstract or research hill climbing with some usage in the real world, and then in this last year we've obviously seen you primarily via coding agents just trickle in in a pretty massive way.
J
Jakub Pachocki53:06
Yeah, I believe it's kind of going in the same direction as the coding models, where it's actually going to be something very useful, it's going to be something that's a meaningful part of people's lives.
I
Interviewer53:20
When you say going in the same way, you mean just like executing longer-term tasks or more like the...
J
Jakub Pachocki53:24
I feel that's part of it, but also just coming to become a dependable, trustworthy assistant or companion.
I
Interviewer53:32
Yeah, it's amazing to watch the way younger people use it. I'd argue it's already pretty much there for the way a lot of folks in high school and college seem increasingly comfortable using it. I wouldn't be a shameless podcaster if I didn't ask a top researcher timelines for a few things. I think particularly interesting is the stuff outside of the core LM world. And so I think there's a lot of buzz around robotics these days. Do you have any intuition? I mean, obviously it's hard to pinpoint a moment robotics 'works', but I think whether it's finding scaling laws or finding some sort of ChatGPT-esque moment for robotics.
J
Jakub Pachocki54:05
Yeah. I mean, I definitely think there are very promising algorithmic ideas there that I believe are going to work that are not too dissimilar from the space of ideas. So I'm quite optimistic about timelines there. Although I do think they're longer than the kind of virtual AI.
I
Interviewer54:23
Obviously, I'm sure you think a lot about, because you're always thinking about the next frontier for what these models can do. Just the impact on society as a whole, as you think about this kind of pace of continued model improvement. What's maybe one thing that you think we're underthinking right now as a society in terms of the impact of these models?
J
Jakub Pachocki54:42
Yeah, I think getting to a point where so much intellectual work can be automated, I think comes with pretty big problems that I don't think have obvious solutions. One natural is a question of jobs and concentration of wealth, and I suspect this requires real policymaker involvement. I've heard some kind of optimistic takes on how this is resolved, but I think at a fundamental level it does seem like some things that used to be very valuable, used to cost a lot, and used to provide something, now can be done pretty cheaply. And in the long term it should be a good thing, but I think it can happen quite quickly. And there is a related question of, you really can, if you actually have an automated research laboratory, an automated company that can do so many things, it can be controlled by a very small number of people, right? It can do a lot. And this gets even more crazy when you have robots, but you don't need to have robots. And I think figuring out what does governance of such things look like, what are these organizations that are so powerful and yet maybe made of only a couple of people, how to think about these things, I think is a new question we have to grapple with as a society.
I
Interviewer56:03
When speaking of other new questions, one thing that's very top of mind for me, I recently had a kid and I've been thinking a lot about what is his life going to look like in 10 years. You're really close to this stuff. How has your work on AI changed the way you think about the way in which this next generation should be raised?
J
Jakub Pachocki56:21
A task for all of us is to build the AI, build a world in a way where at the end of the day humans have the agency, humans set the direction. And maybe a lot of the technical challenges that we cherish right now will become more of a pastime, that's something that we really kind of need to do in order to make progress, and the challenges will be more in figuring out what are the things that are important, what are the things we should go do. I think that will still be, in that world, people can end up with more things to do and definitely more exciting things to do. And I think you still want to have an understanding of technology, all the kind of basic education however you want to acquire it, for the sake of being able to think about these problems.
I
Interviewer57:15
Well, this has been fascinating, man. I really appreciate you sitting down and talking about so many different things. I want to make sure to leave the last word to you. Anything you want to point our listeners to, whether it's research you're doing or products you're excited about or really anything you'd like to plug, the floor is yours. I'm sure there's tons of threads people want to pull out of this conversation.
J
Jakub Pachocki57:35
I think the set of problems we just discussed, and also the questions around alignment, monitorability, I think those are growing to be very urgent challenges. And I don't think there are challenges only for AI researchers. I think there are challenges for policymakers, but also just things we have to think through as a society. And yeah, I'm happy to see some discourse starting to arise and I think we need more of it.
I
Interviewer58:05
Yeah. Well, I thought I could talk to you for hours more, but I'd be doing the world a great disservice by keeping you from your actual work of continuing to improve these models. Thank you so much for doing this. This was a ton of fun.
J
Jakub Pachocki58:14
Thank you.
I
Interviewer58:14
I'm Jacob Efron and this has been Unsupervised Learning, a podcast where I get to talk to the smartest people in AI and ask them tons of questions about what's happening with models and what it means for businesses and the world. As I hope is clear, I have a ton of fun doing this. It's a nights and weekends project in addition to my day job as an investor at Redpoint. But our ability to get these incredible guests on really comes from folks like you subscribing to the podcast, sharing it with friends. It's really what ultimately makes this whole thing work. And so, please consider doing that. And thank you so much for your support and listening. We'll see you next episode.