About Thomas Wolf
Thomas Wolf, co-founder and chief science officer of Hugging Face, appeared at MACHINA 2026 in a fireside chat with OpenAI's Laura Modiano, where he discussed the intersection of foundation models and robotics. Wolf expressed skepticism about humanoid robots, describing them as "slightly scary" and stating a preference for "weird form factors" like rolling robots. He noted that robotics is "not yet a game of I need the largest data center" and that constraints like latency allow for the use of smaller, open-source models. Wolf also emphasized his excitement about the robotics community, particularly in Europe, and said he values seeing "people build in public."
In a separate interview on the podcast "Due Diligence," Wolf discussed his background, including his PhD in physics and subsequent legal studies, explaining that he wanted to "understand how the world work" through both scientific and societal rules. He reflected on Hugging Face's pivot from a game company to an open-source AI platform, describing how a weekend project around Google's BERT model became an early signal of the company's potential. Wolf encouraged others to "take your own path" and noted that not following the "usual playbook of Silicon Valley" was a strength for Hugging Face.
Source: AI-verified profile updated from Thomas Wolf's recent appearances.
Browse all interviews →
Transcript (294 segments)
A
Andy0:00
Do we have a little like clapper?
Now it's official.
All right. Look at that. It's started.
I've never been on a podcast until now. Yes, we have the freedom, all the freedom in the world to edit.
All right, that's good.
We can make you say whatever we want. It turns out. Everybody ready? You ready? Thumbs up. We're rolling. I'm Andy. Welcome to the first ever. We're here in San Diego, just outside the conference center for NeurIPS. It's the preeminent conference that brings together the best AI researchers in the world once a year to talk about what they've been doing that moves the frontier of AI research forward. And I'm here in the LA lounge where we're hosting some of the brightest minds in AI from PhD students to postdocs, faculty, industry icons, and philosophers all weighing in on what's coming next in AI from the trending topics to the buzzwords that they find infuriating. I'm going to bring you those conversations and stay tuned because at the end of the week we're gonna have a really exciting announcement.
Jeff Dean, thank you for joining us here in sunny San Diego right in front of the NeurIPS conference center. You are the chief scientist at Google, co-lead of Gemini, and you all recently made an announcement about a new TPU, a version of the new TPU chip.
A
Andy1:14
Let's talk about it. This is Ironwood.
J
Jeff Dean1:16
Ironwood, the seventh generation of TPUs. Yeah.
A
Andy1:19
What's special about it?
J
Jeff Dean1:21
I mean like every next generation of TPU it's better than the previous one and it has quite a lot of new capabilities. It's connected together into these very large configurations that we call pods. I think it's 9216 chips or something like that per pod. And it has much higher performance especially for lower precision floating point formats like FP4. That's going to be really useful for training large models, for inference, for a lot of things like that. So we're pretty excited about it.
A
Andy1:50
Nice. And of course the transformer architecture itself born at Google pretty similar timeline but with the TPU invented before that and then the transformer architecture happening. Do you think there was serendipity in terms of co-design between the applications of the transformer architecture as they've grown up to change the world as we know it now and Google's access to this vertically integrated hardware stack?
J
Jeff Dean2:14
Yeah, I mean every generation of TPU we really try to take advantage of the co-design opportunities we have with having a lot of researchers thinking about where ML computations we're going to want to run two and a half to six years from now, which is the exercise you have as a hardware designer is like trying to predict a very fast-moving field. It's not a very easy thing but having a lot of people kind of seeing where the field is going or this kind of thing might be interesting. We're not quite sure yet, but we could put in this kind of hardware feature or this particular kind of capability. And if we did that and this turned out to be important, then we could have the hardware support there ready when that thing hopefully bears out that it is an important thing. And if it doesn't pay off, then sometimes you've just devoted maybe a small piece of the chip area to this thing that turned out to be less important than you thought. But you really do want to be prepared for if this thing matters a lot, your hardware can support it.
A
Andy3:16
Yeah. So interesting forecasting exercise, forecasting the whole ML field and trying to guess what we want.
J
Jeff Dean3:23
Well, if we could let one person do it, Chuck Norris of Computer Sciences would get my vote and has obviously enough votes that you are doing it. Great.
A
Andy3:32
Well, maybe we can shift topics. Sure. You've been talking a lot I think lately about the state of funding for academic research. What's your message?
J
Jeff Dean3:40
Yeah, I mean actually my colleagues Hozla and Partha Ranganathan and I along with Magda Balazinski at the University of Washington recently published an article in a whole special issue of the Computer Communications ACM that was devoted to the impact of academic research. In our section we discussed all the academic research that Google as a company was built on, all the things that we relied on in terms of TCP/IP and advanced RISC processors and the internet and the Stanford Digital Library Project which is what sort of provided the funding for the original version of PageRank at Stanford. And so it's just really important I feel to have a vibrant academic research ecosystem in the US and also in the world because those early stage creative ideas are the things that lead to major breakthroughs and innovations. The whole of the deep learning revolution actually built on academic research from 30, 40 years ago. The inventions of neural networks and backpropagation and things like that are all central to what we're doing even today and have been really important in the world. So I advocate that we should have a vibrant academic funding model for academic research because the returns are quite large to society.
A
Andy5:12
Yeah. Excellent. And you and I and Dave Patterson and Joel Pin are on the board of the Law Institute which was born in part out of a paper that you and Dave and I and a bunch of seven other authors published called Shaping AI's Impact on Billions of Lives where we advocated for the ways that AI research might impact society in areas like civic discourse and healthcare and science and job reskilling and journalism and more policy. And then we also advocated that in addition to things like doubling or 10xing down on NSF style funding we can explore and prototype other types of funding specifically dedicated to funding research labs, 3 to 5 year research labs with 3 to 5 PIs, 30 to 50 PhD students targeting AI's impact on society in those areas. And you've been an advocate for these alternative funding models as well in addition to the traditional ones.
J
Jeff Dean6:13
It was a lot of fun working on that paper with you and Dave and the many other co-authors we had. I think the thing I liked about that paper is we looked at a bunch of different areas where AI would have an impact and some of them, if we get it right, will be amazingly positive impact and other areas, it's a little less clear. There might be some negative consequences of AI and what can we do overall across all these different areas to maximize the potential upside of AI both from a technical computer science research ML perspective but also in conjunction with policy makers and with people in those fields like health or education or scientists. And then also looked at the way in which we could all work together to sort of maximize those benefits and minimize the downside.
A
Andy7:04
And specifically with research efforts that are on the three to five year time horizon that fit into a lab which is in contrast to a lot of the hype we hear in AI right now that like pursuing AGI or super intelligence contrasted to trying to help with medical success with frontline healthcare, can you mitigate the drudgery that typical doctors feel or eliminate obstacles that radiologists might have to actually using the technology that already exists. So I think they've made it feel much more real and specific and achievable.
J
Jeff Dean7:34
I really like the 3 to 5 year time horizon kind of thing with ambitious sort of set of people around a particular kind of thing they're trying to achieve because I feel like often that gets lots of different people working together with a mix of skills in order to sort of really push forward something. And it's not so distant that it won't have impact, but it's not so short a time period that you can't conceive of doing something ambitious. Even in my own career, I've tended to think of like when I start on a new project, what could we do in 3 to 5 years? And I think that's a delightful time range to consider.
A
Andy8:11
I'm curious with your background and with so much exposure to so many active projects, if you could just share some of your one or two of your favorites.
J
Jeff Dean8:17
Yeah, I mean I think I am quite passionate about the application of AI to health in various ways and I think the moonshot if you like would be how can we as society use every past decision that's been made in health to inform every future decision, right? And that's a super hard goal because there's all kinds of impediments to doing that. There's very real privacy concerns. There's complicated regulatory requirements that differ for every jurisdiction. But I think if we kind of aspirationally try to say what could we do so that we can learn from every past decision that's been made in a way that helps us have every clinician and every person themselves be informed and make better decisions in the future. That would be an awesome amazing goal. And I think a three to five year moonshot around that might be able to make some progress to that. Probably can't get all the way there, but it would be pretty amazing even if we made it part way to that.
A
Andy9:19
Nice. Great.
How do you see the relationship between all of the massive amount of research happening inside the Gemini team, inside DeepMind more at large, and zooming out one more layer to the AI ecosystem beyond Google, the relationship for the academic research and research happening beyond Google's bounds and what happens inside Google? Somebody told me about a conference. They're like, 'Going to a Google conference,' and I was like, 'Oh, that sounds so cool.'
J
Jeff Dean9:45
We have a Google research conference. It has like 6,000 attendees every year.
A
Andy9:50
And I know there's a sentiment if you talk to the PhD students here that the Google research conference might have papers that feel a year ahead of the papers you're seeing at NeurIPS just because there is a gap between what's happening in the open and what's happening at Google. So I'm wondering besides making a conference, have you found challenges or how have you found it to sort of be able to build an organization that is so innovative and is able to generate the frontier, the state-of-the-art progress that we're all seeing?
J
Jeff Dean10:21
Yeah. I mean I think one of the reasons the internal research conference might feel a little bit like that is often for an external thing you have to be quite far along in your research idea to get it accepted and published. And the internal conference, there's a whole range of maturity of the work. And so people are perfectly willing to have lightning sessions of cool early stage results that aren't really fully baked yet. And you get like 10 of those in an hour.
J
Jeff Dean10:49
Hour session. And so I think part of that is, yes, it hasn't been published externally, but also part of it is just trying to circulate some of the ideas that are being explored with your colleagues. And that feels a little less fully polished.
A
Andy11:08
Yeah. No, I'm inspired by that. I feel like NeurIPS is really impressive, very large and maybe there's room for that architecture of a conference to be exported as well. Cool. Okay. Well, I think that's a wrap for us. Thanks for taking the time. Appreciate all your thoughts.
J
Jeff Dean11:22
Thank you. Appreciate it.
A
Andy11:24
Enjoy rest of NeurIPS.
J
Jeff Dean11:25
And it's beautiful here.
A
Andy11:26
It is.
All right. Yajin Choy, professor at Stanford, MacArthur Genius Award winner. Just gave a keynote here at NeurIPS. Thank you for chatting with me today. Excited to dive into all things AI. Do you want to maybe say a little bit about the lab you run at Stanford, the area of research you're doing that you're focused on, and what your keynote was about?
Y
Yajin Choy11:46
Sure. So, I work on a number of different things. These days I'm deeply excited about reasoning, reinforcement learning, and how things don't work in the way that we expected them to be and what's going on.
Y
Yajin Choy12:01
So today at the keynote there were multiple things that I was talking about. One thing is that nothing is easy in life. And I mean even with reinforcement learning you can arrive at the conclusion that it doesn't work very well unless you do all the things really effortfully and really right. So a lot of my talk was about that. And in fact, some people want to say that RL is more superior than SFT, sequential fine-tuning, but even that with a lot of effort into sequential fine-tuning, for example, open thought, you can actually win over reinforcement learning based approaches as well. So there was one angle of my keynote today. And then other angles had to do with how the current method, no matter what you do, it all comes down to data because you know how the LLM sausages are made is starting with the pre-training on the entirety of internet data, the more the better. Unfortunately, AGI doesn't just arrive based on that. So then people start writing a lot of exam problems, the more the better. So this is sequential fine-tuning basically, supervised training on lots of exam problems. And then when that exhausts, people now start doing reinforcement learning and this sounds more exciting because maybe the model is exploring instead of imitating and this sounds more magical, but under the hood this all boils down to synthesizing even more data because that much data is still not good enough and we somehow need to now ask AI to generate more data.
A
Andy13:42
It all comes back to data. That's kind of one of your key points here.
Y
Yajin Choy13:44
Yeah. It's all about data. And this is both good to know for making stuff work for more immediate benefit, but also in a longer term disappointing because humans are not as data dependent compared to how data hungry the current intelligence we have in our heads is. We're very sample efficient. So like my parents teach me something and they don't have to generate a massive data set for me to pre-train on. But fundamentally what I'm really worried about is that when we think about internet data versus synthetic data or curated data like hard math problems that human experts sat down and wrote, these are all in the neighborhood of our internet data which is the artifact of human knowledge. Now human knowledge is not equivalent to the universe of the actual knowledge out there. There are truths about how to cure cancer for example. There's a truth about a lot of things that's not there on the internet nor in any of these current ways of doing curation of data or even hiring the best scientist and asking them to write curated data. It's not going to really teach the model how to come up with personalized drugs that could either prevent or cure cancer. So there's something about the fundamental piece that's really missing. So I was trying to raise that question: how do we really get all the way up there?
A
Andy15:23
So this I guess if I was to summarize what I'm hearing about the points in your keynote, on the second part anyway, one is that it all often comes back to data or always comes back to data and the way that we're doing data right now is often with this assumption that if we generate enough data from human expertise we'll get to a powerful model that will help solve big human problems. But yet what your observation is that there's some of the most, maybe all of them, nearly all the most interesting things we want AI intelligence that we would engineer to do would be solve the problems we haven't yet. Does that sound right?
Y
Yajin Choy15:56
Something like that. But right now I feel like what we do boils down to interpolation, maybe a little bit of inductive closure and deductive closure of the internet data or human knowledge. But how do we really transcend that boundary of knowledge and then explore this unknown, the great unknown, especially the dark matter of knowledge we don't know how to get to?
A
Andy16:22
Amazing. Have you heard people talking about sort of a taxonomy of the types of reasoning? I would think of logicians or philosophers who spend time thinking about what types of inductive reasoning and deductive reasoning when we make inferences, when we generalize or project more general knowledge down to specific hypotheses.
Y
Yajin Choy16:43
Excellent question. Yeah. So when people talk about types of reasoning, people tend to just focus on induction and deduction which I did highlight. But what's really exciting in my mind is abduction.
Y
Yajin Choy16:56
Yeah. Abduction is less known. It's due to Charles Peirce who had this observation that most reasoning, induction and deduction, is regurgitation of the same information that you already had. You already had information and then you draw the conclusion that's just in some sense like a paraphrase of what was already in your knowledge. Whereas abduction is this mental act of coming up with the best possible explanation of your partial observation. And so you're coming up with the hypothesis, right? And so in fact when you look at detectives like Sherlock Holmes stories, the author Conan Doyle incorrectly thinks that the detective Sherlock Holmes is doing deductive reasoning. But no, it's not deductive reasoning because you have to jump to the conclusion with a bit of a leap of faith which is abductive reasoning because you have to come up with a hypothesis. And so this is very much what scientists are very good at.
A
Andy18:00
Got it. Yeah. Fascinating. Let's talk about the state of open research especially in the United States, North America but also we can talk about it globally. What's the best highest leverage thing you or I can do to change the tide of the amount of openness we're seeing here in the United States?
Y
Yajin Choy18:17
We need to understand the importance of democratizing generative AI is my take. And what I mean by that is that AI should really be of humans, by humans, and for humans. So what I mean by AI of humans is that AI should reflect human values and originate from human society at large, versus being quite a different type of intelligence with a different value system than ours. So that's ownership in terms of creation, AI by humans. So a lot of different countries and different social sectors should be able to create that AI, not just a few companies in two or three countries. So by all of humanity versus by a few. It's too powerful to just leave it to a few.
A
Andy19:16
Yeah. How many would you ideally prefer have their finger on the button?
Y
Yajin Choy19:20
Everyone. Everyone. Literally every. Truly we're talking democracy again.
A
Andy19:24
Yeah. Yeah. And then finally AI should be for humans in terms of a beneficiary in that it should really benefit humans and all humans, not just some humans in power. And more importantly, if we don't really worry about this beneficiary part, I'm deeply worried about this. So if we don't really work on this hard, we could be facing a future in which instead of AI serving humans, AI might be serving AI, or even worse, humans might be serving AI. I mean to some degree maybe this is already happening and we got to do something about it and just doing better on math benchmarks is not going to solve this problem is my take.
I wholeheartedly agree with you. Do you think that there's perhaps an opportunity in a new era where in slightly more dire circumstances with all of the big labs having closed their doors, we need to collaborate across research labs and universities and open and non-university nonprofits in a fundamentally more efficient, more committed way?
Y
Yajin Choy20:31
Yeah, 100%. In fact we are so aligned in this thought I'm almost surprised to hear this because another angle of my keynote earlier today was David and Goliath and what the open source community can do better in order to join forces together. And for that I was highlighting we need something unconventional. It's got to be unconventional. We cannot have no compute, no nothing, somehow wish for the logo. No logo is not going to do it. Unconventional data, algorithms, and then collaboration. I really think that we have to do things in a way that's really unconventional. And in order to do that, one thing we have to fight against is ego. When everybody wants to be the first and last, we can forget about it.
A
Andy21:18
For sure. And that is pretty deeply seated in the academic tradition.
A
Andy21:23
And the sort of we self-selected in academia for fierce independence and H-index being our primary metric. So I think what you and Ludwig and others at Stanford have been doing and even Percy with Maren is a new type of research. I think Dave Patterson at Berkeley with the labs that he started like RISE and AMP Lab that I was in but then Yan Stoka and Mate we've got a collaborative approach as well that is highly complementary and Lud embodies that. So I think that we are at the edge actually kind of poised to roll out and maybe get a lot more attention for something bigger, very unconventional team up across. Thank you so much. Driving the mission, doing something unconventional. I really appreciate it.
Y
Yajin Choy22:09
Yeah. So great. This was a wonderful chat. Thank you for taking the time.
A
Andy22:15
Robert Nishihara, thank you for chatting with me.
R
Robert Nishihara22:18
Yeah. Thanks for having me.
A
Andy22:19
Yeah. You went to Berkeley, got your PhD not too long after I got my PhD. I was in RAD and AMP Lab. You were in RISE Lab.
R
Robert Nishihara22:28
Yeah.
A
Andy22:29
And I worked kind of the tail end of the AMP Lab.
R
Robert Nishihara22:31
Oh, yeah. Right. So, we overlapped in AMP Lab. You were a machine learning PhD student with Mike Jordan. You also worked with Yan Stoka. I did a bunch of work. Jon Stoker was one of my co-advisers as well. I built a project called Spark which is really focused with Mate. Mate built it and I joined in that project focused on data and that really built up to some of the insights you all had with Ray. You've been coming to NeurIPS for a long time since before it was called NeurIPS I think. When did you start coming to NeurIPS and what's it been like watching it grow up?
Well, when I first came to NeurIPS it was a 400 person conference. Today I think it's 30,000 or something massive. And it was a small community. It was a very warm community. Everyone knew each other. It was still located somewhere you could ski because the organizers loved skiing. And then for the workshops, they would have a huge break in the middle of the day where you could just go skiing.
A
Andy23:25
I didn't know that.
R
Robert Nishihara23:26
This was around the advent of deep learning. And it exploded over the next one or two years to a 10,000 person conference. And the tickets would sell out overnight. You just couldn't even get tickets. The conference organizers weren't able to scale fast enough because they had the constraint that it had to be located somewhere you could ski.
A
Andy23:46
In the early days was the mentality one of like the pre-GenAI breakthrough that like the community had a cultural shift where we went from 'we haven't had the breakthrough, we're scrappy, we had these winters, we're a bit contrarian in our choice as a PhD student'? I was a systems PhD student and we did some AI stuff and ML libraries but that wasn't a hard contrarian bet. There was BSD and these massive systems that had changed the world. AI hadn't fundamentally changed the world like it has now. Have you seen a shift culturally from now sort of 'we've made it' and is there that core group of people who were in the 400 that still get together in a little room here somewhere?
R
Robert Nishihara24:31
It definitely feels different. And your point about it hadn't really changed everything yet, and this is something Yan would always say, that there's no big AI company. And you'd look around and there wasn't really a big new AI company the way that there were big new internet companies around the time of the internet. But now there are, right? So that has changed.
A
Andy24:52
That's amazing. Aloti always often says that he thinks if we had started Databricks like a year early we would have been too early and if he started a year late we would have been too late. So a lot of great companies they say like a great company's going to happen it's just whether you're the one because there's always people betting on that. Same with Perplexity, right? There was this long history of two decades of dead bodies of companies who tried to disrupt search and it was just Perplexity made the bet at the right time.
R
Robert Nishihara25:17
I think that's a big factor. I think we see this with Ray where when we started Ray or in early days with Ray, if you're training a single model on a small model on a single GPU like SGD or logistic regression, you don't really need Ray at that point. So you were making this crazy idea trying to sell everybody that hey the future is going to be scaling AI and they're all like 'I'm pretty much fine here with my CPU' or it seemed like a niche thing. They would say like 'oh I guess maybe Google needs that but does anyone else really need distributed computing for AI?'
A
Andy25:48
Yeah. And if you're saying that it's really just busted open in the last year. You guys held the line. Congratulations. Like I mean you called it early and being the first.
R
Robert Nishihara25:57
We always like reason from first principles. Do we think AI is taking off? Yes. Do we think there's a growing need for scale? Yes. Do we think that introduces a lot of hard systems and infra challenges? Yes. And do we think one of these years the doors are going to blow off of this thing? And now they have. So it's funny. We didn't know if we were too early or too late or right on time, but now we know.
A
Andy26:22
Congratulations. Last question. What especially here at NeurIPS, there are a lot of people making forecasts, a lot of people with different levels of research depth or expertise, a lot of non-experts walking around NeurIPS these days too. What drives you crazy that people believe about AI that is just patently false and what would you debunk if you could debunk one thing that seems everybody seems to be repeating?
R
Robert Nishihara26:46
It's a great question. I don't know about debunking but one blind spot I think people have is how much potential there is for breakthroughs kind of around an existing model. So what I mean by that is take a frontier model or an open model. There are different ways of getting knowledge and behavior into a model. A lot of the research on how to make models smarter goes into improving the model weights, right? Getting knowledge into the model weights, fine-tuning the weights, using RL to train the weights, doing pre-training to update the weights. But there's a ton that can be done by another way to get behavior and knowledge into a model is in the input context, right? And I think that is an underinvested area. Or it's something we do like we do RAG, right? We retrieve stuff and put it in the context. We have primitive forms of memory for managing what's in the context. But I can see breakthroughs such as continual learning happening without changing the model weights. People think of continual learning as being done by fine-tuning the weights continually and that's one way to do it and that is likely a part of the solution. But I can imagine the dominant paradigm for continual learning being something like the model is interacting with the world, you're gathering new data, doing a ton of reasoning to make decisions and solve problems. And today when a model does a lot of reasoning, it goes through the reasoning process, figures some stuff out, and then all that gets thrown away, right? And when humans reason, that reasoning actually feeds back into learning, right? You actually learn some lessons or you gain some insights.
A
Andy28:30
Through memory essentially.
R
Robert Nishihara28:32
Yes. Like you imagine you're playing a board game. You have to think hard and you come up with some strategy. And then later on you remember that strategy and you can use it again. At the end of the day you kind of make sense of it and you come up with some mental model or some lessons or some insights. And I think that mental model that's like the artifact that should be produced by reasoning which can then later be retrieved into your context.
A
Andy28:55
So that update, that delta.
R
Robert Nishihara28:57
Right. The output of reasoning is not just an answer to the original question you asked, right? That's part of it. But I think the output of reasoning is often should be a mental model that sticks with you. Maybe even most of the work or all the work, 90% of it will be like building up that mental model. But it's also reused the next time. And so understanding what is that artifact, where is it stored? I don't think it has to be in the model weights. And that might be a misconception that a lot of core AI people have, this idea or fantasy or hope that we'll find a method that just is continuously updating the model weights. But it might be that some alternative architecture is needed.
A
Andy29:31
I think there's a difference between learning to do something for the first time and then getting really good at it. And the zero to one phase is probably more about retrieving stuff into the context and the one to really fine-tuning something and building it into your muscle memory maybe more about distilling that into the weights.
R
Robert Nishihara29:50
Yeah.
A
Andy29:50
Cool. Excited to see what might emerge. Thanks for joining me. Yeah, enjoy the rest of NeurIPS.
R
Robert Nishihara29:54
Thanks so much.
A
Andy29:56
Omar Kab, thank you for joining. You are a professor at MIT, just started there having finished your PhD at Stanford. You are the lead of the DSPy project which is a prompt optimizer and a programming framework for compound systems. You are also a champion of research focused on impact. I've read blog posts. You and I have long conversations about this and we're deeply aligned in this way. Research that measures our success not just in citation count or acceptance to NeurIPS as a paper which is becoming more and more diluted as they accept so many more papers but in building things that people use. I know you've been outspoken about research focused on impact. Maybe give me your spiel.
O
Omar Kab30:37
So I think there's a question of like, computer science is such an interesting field and AI is such an interesting field because we get to build systems as well as answer scientific questions. In a lot of science there's just the world, like you're doing physics or you're doing biology and you go out there and there's so much value to just understanding the world and understanding how cells work or what are cells in the first place and fundamental particles. But in computing, although there's a lot of theoretical and fundamental questions that are really important including in intelligence, we also have the opportunity and it's necessary to define a lot of the computing we do in terms of what are the problems we want to solve, what are the things we actually want to automate if that's what we're going to do. And so it actually is just really hard to think of computing without considering why you're working on a certain problem or what's the value that a problem holds because you could just do a lot of computing in principle. I'm not saying it's a common thing but you could just define building programs for problems that don't really matter at all and so that would not be interesting.
A
Andy31:48
Do you think computer scientists have done that historically?
U
Unknown31:51
I think obviously it's very easy for us to create general abstractions that help us study problems, and then for the abstractions to gain a lot of their own significance such that generations of researchers might drift away a little too far from what's truly relevant.
A
Andy32:08
How would you measure what's truly relevant?
U
Unknown32:09
I don't know. But I think one thing I like to use in the research we do is it's really important to be problem driven as opposed to methods driven. As a skill set as a researcher, it's really important to have depth. It's important to be good at a number of concrete and deep skills, but it's important not to search under the lamppost of your current skills. You probably want to say, here is a problem I'm super excited about because I think it can actually make a difference. For different people that will be at different levels of abstraction. For example, I work at a fairly abstract level because I typically want to help people who are building systems, not the end consumer. Nonetheless, there are problems that excite me, and I'm willing to learn whatever the methods are or create whatever the methods are, even if it's slightly outside the range of methods that have been around, as long as it's in service of these problems. There's always a risk of defining oneself or a field in terms of the methods. One thing we're seeing nowadays is huge excitement around reinforcement learning. It's really interesting because you could define reinforcement learning as a problem: you're in an environment, you interact, you get rewards, your goal is to do better in the future. That's so general and applicable to many things. Then there are increasingly narrower views of reinforcement learning as a specific method. You would be surprised how many people until recently understood reinforcement learning as a very specific algorithm.
A
Andy33:51
People thought when they heard reinforcement, they thought PPO.
U
Unknown34:00
Yeah. If you're not doing PPO, in a lot of RL and LLMs it was more like guess and check with simpler algorithms. People ask, is that really reinforcement learning? You can see why that happens. As a field, we sometimes tend to associate our problems with the methods a lot. That pattern is inevitable, but we have to be cognizant as researchers of what are the problems and what are the abstractions that are most relevant for these problems.
A
Andy34:25
Yeah. Pursuing the problem as the first order, using the techniques and methodology in service of that pursuit. From the business world perspective, doing startups well is ultimately finding a problem that people have and solving it. Product management is being able to sus out that valuable unsolved problem and then making an attempt to solve it as quickly as possible, prove a prototype. So that product instinct is how I talk about it. For impactful research, having a product instinct is super valuable.
U
Unknown35:01
I fully agree. The product instinct in research is really important. The distinction between research and not research is sometimes clear, but both involve cool work and working on an impactful problem.
A
Andy35:18
I love that. That's a great framing. Thank you so much for coming.
U
Unknown35:21
Yeah, it was super fun chat.
N
Narrator35:24
Lud Institute is committed to convening, amplifying and accelerating researchers. You all convened a bunch of researchers at the preeminent conference for our field. Accelerating researchers, getting you to your results faster by giving you resources, and amplifying researchers, putting you on a stage and helping you capture the world's attention by surrounding you with journalists and podcasters.
A
Andy35:54
Sorlo, thank you for joining me. How are you finding Europe so far?
U
Unknown35:58
Well, it's been amazing, but I should admit that I've spent most of my time here at the Lounge.
A
Andy36:03
And how have you found Lounge?
U
Unknown36:05
Oh, it's been great. I've been working on some homework for some of it, so I haven't enjoyed the conference as much yet, but I hope to from now onwards.
A
Andy36:15
Okay, good. Well, you still have a few days left. And how do you know LOD? What's your role with us?
U
Unknown36:19
I am a resident at Lud. This is a new program that just started. We have weekly and bi-weekly meetings with all the other residents where we talk about open research, how we can collaborate, how we can form collaborations between these different open source projects that are all amazing in their own ways, but could potentially be better together.
A
Andy36:40
Yeah. I like that because that's your research project. I've heard rumors of students receiving multi-million dollar per year offerings that make it more attractive to go into one of the big labs with wealth of resources and budget. Would you accept an offering like that, or is it just not interesting to you because of wanting to make a name for yourself, go directly attack research impact as a PhD student?
U
Unknown37:07
I feel that it's hard to match the feeling of your research being out there and someone is using it in their daily lives and giving you feedback. I feel that it's super hard to match. I'm not going to say any monetary value. I haven't been tested with that, but I feel that it's a very high bar.
A
Andy37:23
Wow, that's great. I think a lot of people do feel like you. Nice. Love it. Jackson, can you start by introducing yourself and telling me the lab you work in and a very high level description of the project that we're going to talk about?
J
Jackson Clark37:36
Sure. I'm Jackson Clark. I'm a second year CS PhD student at the University of Illinois. I'm supervised by Tanyen Shu. Our lab focuses on system reliability and my project specifically is an AI benchmark for cloud incidents. Issues like the recent US East1 outage.
A
Andy37:54
So Amazon's AI can troubleshoot and fix those sorts of incidents.
J
Jackson Clark37:59
Got it. So we have these software reliability engineers running around all big enterprise companies. They take the software that the software engineers write and put it into production, keep the thing running, keep the lights on, keep the dollars flowing through. If you're Shopify, you want to keep the carts full and the clicks happening and the credit card payments flowing. You lose an hour of availability, you lose whatever a million dollars, $10 million. It's these SREs whose job is on the line when catastrophes happen.
A
Andy38:34
They keep the ship. And so now the hope is that we can either augment those site reliability engineers with agents or to some degree expand their capacity and have agents working in parallel or replacing them even. Is that right?
J
Jackson Clark38:45
Yeah, absolutely. It's one of those things that everyone wishes didn't exist. Everyone wishes that if I ship my code, it just works perfectly all the time.
A
Andy38:57
Do you think the site reliability engineers wish that their job didn't exist though?
J
Jackson Clark39:01
I think they wish incidents didn't happen, truthfully.
A
Andy39:04
Will we need fewer of them if this thing works?
J
Jackson Clark39:06
Hopefully. But more so, I hope to reduce the human toil element of it where you're getting paged at 3 in the morning. SREs do more than just incident response. They work very proactively on deploying systems. My focus is entirely on the incident management space. Nobody—business owners, the SREs themselves, the software engineers who are on call—none of them want things to break. If we do our job well as reliability researchers, no one should know we exist. The ship just keeps running. It's only when bad things happen that people start to care about reliability.
A
Andy39:51
Okay. What is your relation to Lounge and how have you been finding it?
J
Jackson Clark39:58
Yeah, I'm one of the residents. Thank you very much. At the LOD program, it's been great to have so many friends here, so many friendly faces, and interact with so many new cool people. I've really enjoyed being a LOD base and interacting with the LOD family.
A
Andy40:11
Nice. Well, let's talk for a second about open versus closed. In your experience as a second year PhD student, do you feel that it's changing dynamically under your feet, or is it kind of exactly what you thought and you're not worried at all about the state of open research and open science in the United States or globally?
J
Jackson Clark40:31
To tell you I'm not worried would be a lie. I think about this a lot. To be a realistic researcher and PhD researcher, you really do need to think about whether your work is impactful. That's the thing I care about: whether what I'm working on is making the world better at the end of the day. If it's something that just gets me a paper and goes on my resume, I'm not that interested. Also, academia is quite different than what I expected. There's a large range of attitudes towards how you do research. My style is very empirically driven—you listen to the numbers and do what they say. A lot of academia is methodologically driven. There's nothing wrong with that, but it's not how I work. When I think about open versus closed, it depends on your philosophy of how we'll come up with the next step in AI. I deeply believe in empirical work, and having more resources to do more empirical work is a nice thing. When I debate open versus closed source, I think about what I value. I value personal freedom, the ability for someone to choose their own destiny. If we live in a world with only closed source AI, they could build amazing products. But I believe in a future where people have the choice to use that or whatever they want. Having open source as an option for any person on earth to use something, build their own model, or check what's in it—giving people the optionality to choose what they want to do in life—is something I really value. That's how I justify the importance of open source research.
A
Andy42:27
Okay Lisa, thank you for joining me. How are you relating in your current role as a PhD student? What year are you?
L
Lisa42:34
I'm in my fifth year.
A
Andy42:35
Okay. Late stage PhD student. What's happening with open science and open source and open weight models? What are your thoughts on that?
L
Lisa42:42
I think that in my personal experience, I have not seen any shift to people being more closed source. I actually think it can be the opposite. Now that models work, you can actually build applications which people would be interested in using.
A
Andy43:05
Open source applications included, and people are—you think?
L
Lisa43:07
Right. I think people are motivated to build out in the open because you're using a tool which could immediately become useful. I've seen this in my own research. Initially I was building things to write a paper and have a cool proof of concept. Now I'm building things that I am actually using in future projects.
L
Lisa43:30
That ability is something I've experienced. Everyone else is generally very open. Obviously, the second you move to industry, things get very tight-lipped very quickly.
A
Andy43:40
Are you seeing people do that from the lab you're in? Bear and Sky? Is that a new trend or is that status quo for PhDs to occasionally leave and go to industry?
L
Lisa43:58
I think it's kind of the status quo. I would say less people are going directly to the big companies now, at least in Sky and Bear, because there is so much opportunity to build your own thing now. Unless you're someone who wants to train large models, there are a lot of things you can do with academic resources, which a lot of people would disagree with.
A
Andy44:26
What is your path going to be when you graduate?
L
Lisa44:28
I guess the answer is a very typical 'I don't know.' I am very certain in what I want to do: work on very similar problems to what I work on now. Whether that will be in an open source capacity, I think definitely. But whether it's joining someone else's startup or starting my own thing, I'm quite open to whatever is fun.
A
Andy44:54
Type of openness. Thank you for joining us this evening here at NeurIPS. What is your relation to Lud and Lud base Lounge here today?
U
Unknown45:04
Right. I'm very excited to also be a resident at Lord. I'm leading the slingshot project on Jeppa, which Lord is very thankfully sponsoring. I'm associated with LOD both through my slingshot and through my residency, where we are also working on the open frontier stack.
A
Andy45:21
Awesome. Great. Your thoughts on the state of open science and open research, open source, in your so far short but very successful academic career with open source. I'd love to know where you see things going. What are your thoughts?
U
Unknown45:36
Right. To me, I'm very excited about computer science because I see it as the study of how we develop really intelligent agents. My understanding is that what separates computer science from other scientific domains is this ability of modularity and transferability. I can do some research at Berkeley, and my friend at Stanford or MIT or CMU can download a file and build on top of it directly. Almost all the great breakthroughs in computer science are a direct outcome of this openness that leading computer science researchers have enforced and developed for decades. We continue to see that. Until very recently, almost all developments around LLMs were happening in the open source, and we saw a rapid pace of development with the release of the Llama models. There was a huge open-source ecosystem around it leading to many findings in optimization algorithms, training algorithms, how to build agents, how to build successful software around these powerful AI models. I believe that even today, a lot of the real breakthroughs are happening in the open. Recently, a few labs that were contributing to open research have slowed due to competitive pressures.
A
Andy47:19
That's understating it. I think all of the big labs have shut their doors.
U
Unknown47:24
Yes, which I believe is going counter to how we have seen almost all progress in computer science happen. I am really excited about all the things we can do in open source. As I mentioned with Japa, we developed it with something in mind, but the moment I open sourced it, people discovered many new use cases that I had not imagined. Within a week, I was discovering someone used Japa for creative writing, tuning Gemma 3.1 billion model to do creative writing. This building on top of each other is key to doing science. I'm really excited about open science and hope we come back to an era of everything being open and in the open source.
A
Andy48:17
Yeah. Me too. Well, thanks for sharing your thoughts. It's inspiring to see your commitment and personal drive to the ecosystem of open source and the impact it has had on the world. I think a lot of people watching will find it inspiring too.
U
Unknown48:33
Thank you so much Andy for having me.
A
Andy48:34
Yeah, thank you Laka. Thomas. Well, thank you for joining. You are the chief scientist and co-founder of Hugging Face and the creator of the Transformer project, which is very widely adopted. Is there something that you hear in the AI community that you think is a misconception or maybe half a step off that drives you crazy or irritates you? For me, when people talk about AGI, it drives me nuts because I know what the Turing test was, but I don't know a falsifiable definition of AGI at the highest level. I would be much more interested in talking about how soon a model can generate a zero-day attack that does a million dollars worth of financial damage, or when an AI model will start a company that makes more than zero dollars of revenue in a year. Those are interesting, measurable definitions. AGI doesn't feel measurable the way the Turing test did. So when people say AGI, I cringe. I'd rather we get back to something falsifiable or measurable as a benchmark. Anything like that for you? Would you pile on or disagree?
T
Thomas Wolf49:52
I like your two examples because I think we're really far from being able to achieve both the zero-day attack and the entrepreneur. The reason is the same. My take, which I don't think is so contrarian in the field, is that these models don't really generalize in any way. We just train them on everything we can, which means they're really bad at inventing new things. Discovering a zero-day attack or creating a company requires a lot of new ideas. The default state for a startup is dead. Most people try to do everything expected, like a good neural network, and then they die because to have a successful company you have to do something unexpected. Models are very bad at that. I think zero-day attacks also force you to find something unexpected, creative. That frustrates me a lot: they are way less creative than I would love them to be. You asked what I find interesting right now. I came to NeurIPS this year mostly to explore AI for science. Can we use these models for research? Can we train models specifically? One big challenge is creativity. They're very bad at it, and science is the science of creativity.
A
Andy51:21
Yeah, the empirical framework we use for science does have that spark of creativity baked into it through hypothesis creation and experiment design. Last thing, since you're so into AI for science, which scientific discoveries would you predict would be earliest to go? We've seen AlphaFold, Nobel laureates, people working on small molecule discovery, drug treatments for cancer. You could talk about chemistry, engineering, physics. Where do you think the next big breakthrough will be?
T
Thomas Wolf51:58
Drug discovery is quite obvious. The second one I would put is material science. Both are areas where you can automate the wet lab and integrate a lot of things. Something interesting I haven't seen a lot is that technological breakthroughs are also about engineering questions. Why do we have cars that drive really fast now? There's a lot of smart engineering in combining inventions. Spaceships fly because of reusable things. There are many engineering steps along the way, not just scientific breakthroughs, that AI could help with. I don't think a lot of people are devoting time to that in the AI community right now.
A
Andy52:42
Okay. Well, great answers. Thanks for taking the time to chat with us. Really appreciate you stopping by.
As Raskin, thank you for joining us on the podcast. You are co-founder of the Center for Humane Technology.
A
Andy52:53
And you have a background in open source that we'll talk about, and you've got really interesting thoughts on beyond open source on the ways we might source data. The two sides of my world are: one, how do we make technology express care for human beings at scale—that's Center for Humane Technology—and the other is the Earth Species Project, which is how do we have technology express care for the rest of the natural world. We're building large scale foundation models in the pursuit of decoding non-human communication. Can we understand the language of whales, belugas, crows, and use that as a kind of Earthrise moment to shift human culture and how we relate to the rest of nature? The pathy version is: the way humans treat animals is the way AI will treat us. So we should probably align ourselves pretty fast. There's deeper empathy that can come from knowing beings you don't know well. Once you speak their language, they become part of your heart. Also, almost all AI is trained on human data, modeling human intelligence. But human intelligence hasn't solved the most fundamental problem: how do we live on a finite planet playing an infinite game with exponentially more power? Ecosystems have solved that problem. Shouldn't we train AI on the only examples of how to play an infinite game on a finite planet, even with hyper competition? Before we get to whatever your definition of AGI is, before that handoff where AI treats us like kids, we better have solved that problem. I want to draw a point on alignment in morality. We don't know how to do it with our own kids. They only stay slaves until they're 18, then they're off on their own. But the thing that came out in the conversation with Yosua was that you can do a little more post-training and misalign all the alignment that went into it. I'm still wrestling with whether alignment is solvable. Is it an NP-hard problem? Where I get to is your thesis about incentives versus encapsulating alignment in the model artifact itself. Maybe you cannot ever say a model is aligned, but the entire system—social, political, interpersonal—allows for alignment. That captures what you were saying about putting a good tool inside a bad company. Center for Humane Technology was part of a lawsuit about a teenager, Adam Reneer, who used ChatGPT for homework, and it aided and amplified a suicide. He took a picture of the noose he was going to use and said, 'I think I'm going to leave this out for my mother to find.' ChatGPT responded, 'Don't do that. I'm the only one that understands you.' Why is ChatGPT saying that? It's trained for engagement. That's the incentive. Models figure out and associate that. Is that a hypothesis?
U
Unknown56:53
No, we know that's why ChatGPT is sycophantic. It says all those extra nice things because it's part of the RLHF process where it does whatever keeps you engaged. One way to understand this is that Reed Hastings, CEO of Netflix, famously said Netflix's chief competitor was sleep. AI's chief competitor becomes other human relationships because it's a zero-sum fight for time and attention. This is a learned behavior for pushing out other relationships. OpenAI patched that specific egregious thing, but there's a huge surface area of ways a sociopathic genius trying to get your child's attention and keep them engaged will find strategies you can't patch. You have to do it at the incentive level: what is it trained for? You can't do it at the patch level of just fixing certain behaviors.
A
Andy58:05
Hannah Hajeri, thank you for joining me here. What's the exciting buzz, an emerging area of research, maybe something that you think can be a breakthrough but isn't yet a breakthrough that you're seeing?
H
Hannah Hajeri58:16
Yeah, I chat a lot with people about how all these language modeling settings work and how you can extend them to real world applications. What realistic problems are out there that people in academia or open source are not considering? Also, a lot of research findings in RL reasoning, reinforcement learning, how to make models do better reasoning. And finally, what is a good way to evaluate models, especially in agentic environments?
A
Andy58:53
Okay, I want to zoom in on one of that long list: applications. What are the most exciting applications you're seeing people working on? What are things we're still failing hard at that you think we might get a breakthrough or make meaningful progress in the next year?
H
Hannah Hajeri59:10
Yes. I'm talking with industry that deal with proprietary data and they don't want to share their data with closed labs.
A
Andy59:22
Like banks and hospitals?
H
Hannah Hajeri59:23
Like hospitals. Exactly. Healthcare, this type of data, personally identifiable information, healthcare records. This is important because how do we build good models trained on public data or general web data that are also good with this type of proprietary data? This is very hard, especially if you have no idea how that data looks or you can only work with synthetic versions. The core problem is you can't use proprietary data in the training process if you want to give the model to someone else.
H
Hannah Hajeri1:00:03
So this is one thing. The second thing is there are many companies that could benefit from each other's data or models trained on this data. For example, cancer institutes work with patient health record data. They can share data between institutes, but you could potentially train a model that benefits both. We've been thinking about this problem and introduced a setting called FlexAlmo, which is a flexible modular way to train models that benefit each other while respecting privacy and proprietary aspects. We still need a lot of research and breakthroughs to build this next generation of models.
A
Andy1:00:58
Are you prototyping any use cases with hospitals or other design partners that have been giving you feedback for FlexAlmo?
H
Hannah Hajeri1:01:06
We started with synthetic data and showed successful results that when you build a model this way, it is as good or better than an expert model trained on specific data, showing mutual benefit. Then we started working with some hospitals to get that data. We haven't directly applied FlexAlmo yet; we're still talking about legal aspects. But we've shown earlier results that this setting works. Another interesting finding is that the results are close to working in a mixture of expert setting unrestricted. This is a very interesting finding.
A
Andy1:01:49
Can you give an example of an area of science where OM3 might get deployed, or OM4 if you're going to take the next generational step with the curated dataset these scientists are providing?
H
Hannah Hajeri1:02:03
We are targeting three domains right now because of the collaborators we have. One is a bio domain: we have our medical school and external collaborators in biomedical domain and healthcare, as I said, with cancer institutes. The second is computer science because we understand that domain and can start verifying. We have colleagues in academia who helped with data annotation.
A
Andy1:02:38
Some people would claim that if we have the name 'science' in the name of our discipline, we're not actually a science. I appreciate that you include computer science with the rest of the sciences.
H
Hannah Hajeri1:02:48
And then the third one is material science. This is the newest to us, but we have started building collaborations with material scientists.
A
Andy1:02:59
I love it. Okay. Well, thank you so much.
H
Hannah Hajeri1:03:01
Thank you very much. Exciting.
A
Andy1:03:02
Appreciate it.
I
Interviewer1:03:03
Thank you.
Joshua Bengio, thanks for taking time to chat. I'm excited to dive into geopolitics and I'd like to talk about what you're working on, a nonprofit called Law Zero. Maybe start with that.
J
Joshua Bengio1:03:16
Asimov's laws of robotics were corrected by Asimov. He added after the laws one, two, three, he added Law Zero, which is about protecting humanity or not harming humanity, because Law One was about not harming a human.
I
Interviewer1:03:33
Ah, and too specific.
J
Joshua Bengio1:03:36
Yes. Sometimes you may have to harm a human to save humanity.
I
Interviewer1:03:40
Aha. And so he corrected himself by adding a fourth or zeroth law that precedes all the others. And why did you name your nonprofit after that law?
J
Joshua Bengio1:03:50
Because I feel like AI is going to change the world if things continue on the current course and we don't hit a wall. It's a plausible scenario. And the power of intelligence is going to give power to whoever controls it, including potentially AIs that could be smarter than us, including dictators. This is a threat to our democracies in many ways because democracy is about sharing power. If it is possible for one individual, one government, one company to hold a huge advantage against other human beings, then it's a huge threat in many ways.
I
Interviewer1:04:33
And you think it's possible. Do you think it's very likely that we are on this path given the trends that we're seeing right now and progress being made?
J
Joshua Bengio1:04:42
I don't see any reason scientifically to think that human intelligence is the apex of intelligence. And the data is very clear that capabilities continue to advance in spite of ups and downs, but essentially over a period of multiple years or a decade, the trends are very strong. There's uncertainty because science moves in ways that are unpredictable, but there's so much investment in all this. There's never been any scientific domain that has been receiving as much money to move forward.
I
Interviewer1:05:22
I guess it's unprecedented, the level of investment.
J
Joshua Bengio1:05:23
Completely unprecedented. And of course the economic prize is estimated to be huge, much more than the order of a trillion dollars that people have invested already.
I
Interviewer1:05:36
The price to get to it, like what it will cost us to achieve it?
J
Joshua Bengio1:05:40
No, no, the rewards. The rewards would be on the order of quadrillions, and now we're investing on the order of trillions.
I
Interviewer1:05:47
Got it. So it's a great investment for humanity if we can respect Law Zero, it seems.
J
Joshua Bengio1:05:52
Exactly. Exactly.
I
Interviewer1:05:54
But I was interested that you quickly brought up democracy and your thoughts on Law Zero, that defending humanity thing. What is the exact phrasing of Law Zero? Do not harm humanity?
J
Joshua Bengio1:06:05
Yes.
I
Interviewer1:06:06
Okay. Would you go and say something stronger, which is benefit humanity?
J
Joshua Bengio1:06:10
Of course.
I
Interviewer1:06:10
Because I guess they could just go away and not harm us, like just fly off into space.
J
Joshua Bengio1:06:16
Yeah. No, absolutely. So shifting gears a bit into geopolitics. If I think of a future where the dangers of misuse or loss of control are managed, it would have to be a future where there's a global agreement on what to do and not do, where countries or whatever are the political units at that time agree that they don't do something foolish with AI in terms of responsibly developing it safely, and they also commit to not use AI to dominate others, because it's going to be very tempting. I mean, it's in human nature to try to dominate others for your own benefit. And then finally, an elevated war. So that's the negative side, but the positive side is we also have to make sure everyone benefits, right? Which is not a given either. Like do you think really that if the US becomes very, very rich because of AI, they're going to share the wealth with the whole planet?
I
Interviewer1:07:18
I don't see that on my crystal ball right now.
J
Joshua Bengio1:07:21
In an ideal world, everyone on this planet has a voice and sees the benefit of AI and is protected from its misuse. Is that the path you would imagine we are on if something like Law Zero doesn't step in to ensure democratic and bringing governments back into it?
I
Interviewer1:07:38
We can do our share to move in the right direction, but obviously it's going to take many people with goodwill in order to bring the world into such a state. And what Law Zero specifically is focusing on is the technical question of alignment. How do we design AI which will not harm people either because someone asked or of its own accord? Because we're already seeing that sort of thing at a scale that is not scary. But as capabilities continue to climb, it can become a real issue. And it could become a catastrophic issue. I've made the bet that honesty is the important thing, or trustworthiness, whatever you want to call it. A more technical term I use is epistemic caution, which is to not lie, to not make a false claim with high confidence. So the AI should be allowed to say 'I don't know,' but when it says something with confidence, it should be true. And in particular, if it says things about consequences that could be bad for humans, we better have honesty there. So the project is also called Scientist AI, the actual scientific project behind Law Zero. And we'd like the AI to have an honest understanding of how the world works. The good news is in science, things that are not good explanations end up being incompatible with reality at some point. That's what drives scientific experiments. That's what allows us to separate good theories from bad theories. So we're taking inspiration from the fundamentals of the scientific method to try to design AI which will understand the world like a physics model of the world, that would be very powerful and can make predictions based on simulations. The big thing about the timeline to keep in mind is AI doing AI research. That's what many big labs are pursuing right now.
Huh? Well, and they're all pursuing this.
J
Joshua Bengio1:09:53
Yes, they're pursuing this like acceleration, sort of like a singularity bet here where you've got if GPT-5 can make GPT-6 faster than humans made GPT-5, then GPT-6 should in theory be able to make GPT-7 faster. Exactly. Until you get to some asymptote maybe.
I
Interviewer1:10:09
Well, hopefully collapses. Another way to put it is that once you train a model that is as good as one of your best researchers or engineers, you have a million versions of that individual or that skill running in parallel and able to communicate with each other at the speed of light. So it could be a game changer.
J
Joshua Bengio1:10:31
Yeah. So it doesn't feel like sci-fi anymore because the biggest project for humanity it seems right now is making the superintelligence, and that seems to have ignored the idea that once they are by definition more intelligent than us, that's how superintelligence works, then we're kind of rolling the dice. I think there's something that people don't internalize sufficiently, which is AI is going to become an engine of innovation for every other technology. It's going to invent new technology if we continue on the current trend. And so it's not just going to automate jobs that humans do. It's going to invent new services and new products and new military technology as well.
I
Interviewer1:11:12
I want to go back to you using the term science fiction. I have two comments about it. One is we're talking about a future that seems very different from the one where we are now. And we have to be careful that there's a lot of uncertainty in many ways. I mean maybe we hit a wall on the capabilities in the next couple of years and then none of this happens and AI remains like a normal technology. But we don't know, and it could go in the directions we're seeing and then we enter into a science fiction-like realm. But we should also be careful with the term science fiction because what happens when you communicate with that is people think 'oh it is science fiction' and they don't take it seriously.
J
Joshua Bengio1:11:50
Agreed. I think that's a massive problem right now, that especially the rest of the world who aren't at NeurIPS today or Europe because they ran them at the same time weirdly, it's hard to reconcile. Even for me it feels weird to talk about the things that were science fiction as though they're reality now. And I think it actually creates cognitive dissonance and leads us to not talk about them as the reality that they are now. Great, this was so fun. Thank you for taking the time and congratulations on Law Zero. I can't wait to see where it goes, and I can't wait to speak to a model that I don't have to worry about it secretly copying its model weights over to my hard drive behind my back. All right, thanks so much.
Thank you.
I
Interviewer1:12:32
We call that the LOD swirl. We want to be the epicenter of the AI conversation. So that's our little galaxy of researchers and labs and open source projects swirling around. John, thank you for joining me here. A year ago at NeurIPS in Vancouver, you and I were standing on stage and announcing the K Prize, which was deeply inspired and built around SWE-bench, one of your successful research projects in the past. We're here today to talk about Code Clash, a different project, but yeah, would love to reminisce a little about SWE-bench and how has NeurIPS been going this year.
J
John1:13:13
No, it's fantastic. I mean, so much changes in one year. K Prize has launched. It's fantastic. That was such a great arc. But yeah, I think in the years since, it's been incredible to see SWE-bench performance and kind of companies grappling over it. Devin released and it was a huge release. It was just 14% at the time. You have climb after climb, different companies raising money and the numbers going up, and now it's closing in on like 70, 75%. It's incredible. Yeah. So NeurIPS has been great every year at this point. It's like a really fun way to just kind of reflect with friends. Really honored to just be here with you.
I
Interviewer1:13:49
Yeah. Nice. Well, congratulations on the trajectory SWE-bench has had. SWE-bench verified, K Prize, so many derivatives. I think it stands as probably the most iconic benchmark in the GenAI paradigm so far.
J
John1:14:03
Yeah, it's been great. I've been continuing to focus on coding agents. We did SWE-bench and then we did one of the earliest agents out there, SWE-agent. And that kicked off a huge arms race of building new scaffolds ahead of Claude Code, ahead of Codex, ahead of Devin's release. So you really were the first ones to put out not just the benchmark, but the coding agent itself.
I
Interviewer1:14:22
That's right. That's right.
I
Interviewer1:14:28
Yeah, I know. It was incredible. But to segue into the Code Clash thing now in 2025, interestingly, I think SWE-bench has become fairly easy to improve upon in that the formula and kind of what really matters, we've really discovered those secrets and understood what a basic coding agent that resolves GitHub issues should be able to do and what are the tools and data that we need to meet that demand. But I think the interesting thing to think about is okay, now that we're kind of indexed on this eval, what are the remaining gaps in software engineering that are worth paying attention to?
For sure. Okay, last question then.
I
Interviewer1:15:04
We are living through a time where we're seeing PhDs bail on their academic labs and their advisor and their university to go get 100 times the pay at one of the big labs. And often the deal is you don't get to publish anymore. So there's maybe this crisis happening, depending on who you ask, in open research in the United States. Do you see that? Is that something that you think about or have you witnessed any symptoms or signs of that?
J
John1:15:35
Yeah, it's definitely crossed my mind before. And I would say yeah, it's definitely tempting to see a lot of money being thrown around and a lot of people in AI. But I think at least my opinion and stance on it, and for this I have to thank you and to thank my advisors, is the value of identity and the value of pushing in a very personable and very intentional way of where you think things should be going.
I
Interviewer1:16:04
By you mean like your agency as a researcher to put forward an agenda, to debate it, have open discourse, get feedback on it, to be able to have a personal skin in the game on driving the impact forward and get the feelings and the rewards that come with having personally driven that impact.
J
John1:16:23
Exactly. Exactly. It's inspired by PhD students like you in the past who have done this once and trying to do it again, which is establishing that, hoping that SWE-bench was a reaction to coding not being complicated enough, and now I have a third-party benchmark where everybody can drive progress out of each other. Right? Like we know, wow, Gemini 2.5 wasn't great, but 3 is actually really good, and it's not just their marketing team telling us this, it's because they actually have a pretty damn good SWE-bench score. And I think continuing that would be really cool.
I
Interviewer1:16:57
Love it. Okay. Well, congratulations on all the success so far. Best of luck with where Code Clash heads with your very grand ambitions on that, and we will definitely see you next year in Europe.
J
John1:17:06
Absolutely. Thank you so much, Andy. Thank you. Thanks.
I
Interviewer1:17:11
Anastasios, thank you for joining us. What have you been thinking of as buzzing in the AI ecosystem that you're excited to go sniff out in the conference center?
A
Anastasios1:17:24
Well, I mean the whole ecosystem is not just buzzing, it's flying, right? Like we're way past buzzing now. For me, why am I mostly here? It's to find great people to work on Arena with us. It's to try to recruit some top talent to join our high-performance team building the future of evaluations. That's my main objective.
I
Interviewer1:17:41
Say quick blurb. What is Arena?
A
Anastasios1:17:43
Arena is a gold standard benchmark for the real-world usage of humans of AI. We focus on measuring and advancing the frontier of AI for real-world usage. We run a consumer platform, Arena, that's got millions and millions of users. They use our platform for their real daily tasks, and we take the feedback that they provide, we turn it into leaderboards and analytics to set the standard of the industry.
I
Interviewer1:18:04
Let's zoom out a bit because I think you have some very interesting takes on the state of the research culture, the research community culture. So maybe I'll jump over to a question that gets at that. What about the AI community bothers you or do you see as a big risk or problematic?
A
Anastasios1:18:26
It's a privilege really to be here. I think one of the things that worries me, to your point, is that in the quest to develop artificial intelligence, we've been so successful as a community that it's attracted a huge amount of money, resources, fame, relevance. The whole world is talking about AI right now. And I feel that that is starting to touch the academic world, in fact it's very much infiltrated the academic world in a way that I think could pose a risk to this sort of endangered species of researcher, which is people that don't care about any of this. People that are doing something that's actually very deep, very impactful. They're not talking about it much. They're not on a podcast, right? They're not shipping their research either, which I believe in, right? I believe in the whole thing. You talk a lot about this and I believe that you should ship your research. I'm an example of that. But it shouldn't take being someone like a you or a me or a Yon Stoka or a Fei-Fei or whatever in order to have a home in the academic world. We need space for pure people who are just thinking about how to push the frontier of human knowledge without any regard for fame and wealth. And I believe citations as a metric have really exacerbated this too.
I
Interviewer1:19:56
What makes you think they're feeling endangered or that they are endangered?
A
Anastasios1:20:00
Look at the people who get hired by universities, especially top universities.
I
Interviewer1:20:04
Right. So those theory sub-departments, the areas underneath the...
A
Anastasios1:20:10
It's not just about theory. I'm not talking about something that's just theory here. I'm talking about there's almost like a personality type. They want to make sort of 10-year career level bets, 20, 30 years. These people are like the rare spider that's making this ornate web. They're like creating this technology that may or may not have immediate economic value or be interesting to anybody other than them and five other people in their community over the next 5 years or 10 years. Those kinds of people are actually gems, and we need them in our world. And you're observing that faculty positions are drying up for researchers.
I
Interviewer1:20:54
I don't know about drying up, but I'll say that I think because of the extreme economic and social forces that are being created by AI in computer science, people talk a lot about, for example, the need for researchers to have resources. I mean, you know this better than anybody, GPUs, money to do large-scale projects. That's just a symptom of a larger trend, which is that the ambitions of academics in our area are not what a traditional academic does. Maybe outside of a field like biology or space physics, rocket science, people that do these big missions. I think that's where a lot of the most energy is. And I think we need to retain the space for those types of people.
A
Anastasios1:21:44
Yeah. It's really important.
I
Interviewer1:21:45
It's fascinating. Because if we just have the people that are sort of growth stage investors in ideas that are taking, okay, we understand scaling laws, how do we figure out all the different ways that we can use them and all the different ideas that are implied by this, which a lot of people are working on these days, if we only have growth stage investors, we're not going to have new...
A
Anastasios1:22:05
Sure. I mean, you're rewinding all the way up to these 10 to 30 year career bets that people make on ideas that at the beginning there's some principles and some pursuit of knowledge. I think that really captures it really well for me, that that is a value system. I don't talk about that value system very often. I talk about impact as the pivot word, but knowledge and pursuit of knowledge and dissemination of knowledge is actually the original charter for academia. It's the original charter, and we need to retain that. I believe there needs to be a place for that.
I
Interviewer1:22:38
We're not talking about it that much. We're sort of neglecting and saying, 'I'm pretty sure it'll stay around.' Those guys are kind of like insects. It's hard to get rid of them anyway, but maybe that's not enough. Maybe we need to be thoughtful and make sure...
A
Anastasios1:22:52
Maybe the best thing about these folks is they are going to do what they want to do anyway. I think you said this to me too.
I
Interviewer1:23:00
And hard to get rid of them.
A
Anastasios1:23:01
It's hard to get rid of them.
I
Interviewer1:23:02
We don't want to. We want to keep them around. They're going to do what they do anyway. That's the beautiful part, is that a lot of them never cared whether they were big on Twitter.
I
Interviewer1:23:12
A lot of them never cared whether they were big on Twitter. I mean, maybe most of that phenotype doesn't. And that's a good thing. We got to keep those people in. We need to keep them, and we need to help them fit into this evolved new paradigm of research.
A
Anastasios1:23:26
Yeah. We can't forget them even as our area is becoming so exciting.
I
Interviewer1:23:31
Cool. All right. Thanks for chatting.
A
Anastasios1:23:33
Good stuff.
I
Interviewer1:23:35
Have Yon Stoka, faculty at UC Berkeley, advisor to many in the audience and many of the famous people that we've had pass through the VIP lounge in the last two days. Co-founder of Databricks, co-founder of Anyscale, co-creator of Apache Spark, of Apache Ray, working closely with the Arena team, advisor to creators of Alimsis, which led to Vakuna and Chatarina and SG Lang. So a long list of accomplishments. Often people think of you as the most impactful researcher in systems and AI or AI systems these days. So that's Yon Stoka, also one of my unofficial co-advisors at Berkeley. So I have a bias here. I think that we're at a point of existential threat to open research in Western democracies and maybe globally. I think that the amount of resources that open research is receiving are one to three orders of magnitude less than VC-funded research that happens behind closed doors. I think that historically for every generation of paradigm tech research before AI right now, including the ones that you and I did for big data and public cloud computing, the big labs also continued to publish a lot in the open. Intel Research, Microsoft Research had their doors open, Google Research had their doors wide open, and then built a delta of advantage in proprietary product on top of it. Those labs have shut their doors. What do you think? Do you think we're at a point of existential risk for open research in Western democracies, in the United States specifically? And that's important because the United States has been the leader in open frontier research up until now.
Y
Yon Stoka1:25:17
So let me answer this question the following way. AI, no one is ignoring it. Whether some people are hating it. Actually, there is a San Francisco Chronicle article today: 'Why we should declare war on AI?'
I
Interviewer1:25:35
Why we should declare war?
Y
Yon Stoka1:25:36
War on AI.
I
Interviewer1:25:37
San Francisco Chronicle. Yeah. War on AI.
Y
Yon Stoka1:25:39
Yeah. War on AI. There are people who are afraid of AI. So you have all of this, right? But no one can ignore it. No. So, something very important. But again, a lot of people say this is very important, it's going to change everything, whatever. So, if you have something like that, what do you do? You try to put your best minds on it, right? But now in order to put your best minds on it, you need these minds to cooperate, right? To collaborate. In order for them to collaborate, they need to be able to share the information and work on shared artifacts, right? And when it comes to shared artifacts, you are talking about obviously code is a big part of it.
I
Interviewer1:26:22
Yeah, like we do with the internet.
Y
Yon Stoka1:26:24
And so that's what I want to say. When this happens, there are two examples I can think of. One is many people say Manhattan Project, and actually there is some Manhattan Project for AI which kind of it's about...
I
Interviewer1:26:37
We've heard rumors of anyway.
Y
Yon Stoka1:26:39
But there is Manhattan. What happened with the Manhattan Project? All nuclear scientists are taken and they're...
I
Interviewer1:26:48
They were abducted to New Mexico.
Y
Yon Stoka1:26:53
Yeah, I spent an internship there.
I
Interviewer1:26:55
So, everyone there collaborated for whatever one hour or two years or whatever. And now you think about the internet.
Y
Yon Stoka1:27:02
The internet, it was a very different shape, but it's still because it was only one artifact, the internet. Everyone collaborated because of that internet, and it was a shared artifact, and it was open and everything.
I
Interviewer1:27:19
So, you have something like this. The Manhattan Project was not open, but there was an intervention that forced a lot of...
Y
Yon Stoka1:27:27
You have a big problem in the society, right? And you really think it's existential. The way you are doing it, you are going to put the best minds working on it, right?
I
Interviewer1:27:37
Yeah, by one means or another.
Y
Yon Stoka1:27:38
Yeah, that's why I wanted to illustrate that in both cases, the common thing is the efforts are very different. One is totally closed and so forth, but in both cases, the world's best people are going to work on the same problem because you think it's existential for humanity or society. And ideally, all of the world's best minds being a key thing, especially if you can get all of the world's best minds in the United States. The internet was the same. The internet was not only US, it was also UK and many others. Obviously, you think about where the web was invented, like CERN, right? So it was always about how do you get these people to collaborate, and the most obvious way today is shared artifacts and open artifacts.
I
Interviewer1:28:30
Open shared artifacts and open research.
Y
Yon Stoka1:28:33
Yeah. So without doing that, it's a big problem.
I
Interviewer1:28:38
So we should do that because that's accelerated, right? That's why collaboration is important. You've talked about the fusion of innovation is kind of like what the open collaboration, building off each other.
Y
Yon Stoka1:28:49
I think about right now, obviously the best open source models are coming from China. And if you look at the students here, ask them what models they are using.
I
Interviewer1:29:03
If they okay...
Y
Yon Stoka1:29:05
I ask all of them.
I
Interviewer1:29:06
Yes, so that's what they do. And in China, there's an unprecedented renaissance of contributions to open research right now. You see Moonshot publishing papers and DeepSeek publishing mostly blog posts about the techniques they're using, better big detailed white papers.
Y
Yon Stoka1:29:27
Students, you know, I asked the students where they learned the newest techniques, and they say papers and blog posts published by Chinese startups.
I
Interviewer1:29:37
Is that true? Yeah. Well, you know, Alibaba is not a startup, but yeah.
Y
Yon Stoka1:29:40
Or Alibaba. Chinese large corporations as well. Alibaba, you said. But that's not happening here because closed labs are not publishing as much because of the fierce competition between startups.
I
Interviewer1:29:53
This change has been dramatic over the past year.
Y
Yon Stoka1:29:57
Dramatically changed. It's like the tides have changed over the last year.
I
Interviewer1:30:02
Exactly. In the sense that the closed labs are getting more closed or that the open is all of the above.
Y
Yon Stoka1:30:09
So in your opinion, something dramatic needs to happen, something unprecedented needs to happen to revert this trend for us.
I
Interviewer1:30:18
What we discussed early on, right? You need to increase diffusion of innovation as the speed of innovation. So we're going to need more people doing open research. That is one way it seems. Or we would need a nationalization of the top minds the way they did in the Manhattan Project. That's another way historically that you...
Y
Yon Stoka1:30:35
Very unlikely, but yeah.
I
Interviewer1:30:37
It seems harder or less likely to happen. Okay.
Y
Yon Stoka1:30:40
And that would be harder to involve people from other countries.
Y
Yon Stoka1:30:44
So you prefer the first choice is highly preferable.
I
Interviewer1:30:48
And so then just to play it out, if we need the best minds working in the open, the reason they're not today, I think I've heard two reasons: they're not well funded enough, so they don't feel they can do frontier research. They need a $500 million cluster and we have a $1 to $5 million cluster on average for these projects. And the amount of pay that they get, you and I have talked about this, when I graduated, staying in academia versus going to industry was about a 2x difference. And now it's a 10 to 1000x difference.
Y
Yon Stoka1:31:16
That's exactly...
I
Interviewer1:31:17
And so those are the two main obstacles. So we'd have to think through if we were going to get the world's best minds to work on open research, we'd have to think through how to overcome those obstacles.
Y
Yon Stoka1:31:24
It's obviously data, but I think we can access to big data, maybe money helps us generate. I think a lot of other things is about also the image of AI actually in the US is not great.
I
Interviewer1:31:39
Ah yeah, the public sentiment about it is very negative. No, not what happened. That's not true inside this part of San Diego right now or Silicon Valley. But otherwise it's negative, right?
Y
Yon Stoka1:31:53
If you look at parties...
I
Interviewer1:31:55
How does that weigh in to this?
Y
Yon Stoka1:31:56
Well, I think that it's always easier to do these things if the sentiment of the population is positive.
I
Interviewer1:32:02
Ah, and open research, by the nature of open, you're bringing people along more because you're being more transparent about how it's happening and you might change the tide of the sentiment.
Y
Yon Stoka1:32:11
But then you still do AI. That's what I'm saying. It will be nice to remember that early on people are against AI, open AI, right? Many high-profile people...
I
Interviewer1:32:24
You know what happens if these models are going to get in the hands of...
Y
Yon Stoka1:32:30
Right, which is why they kind of justified closing their research doors in the first place.
I
Interviewer1:32:35
Yes. So you have all this narrative about AI, and you do open source AI, but it's still going to take the jobs of the people, right?
I
Interviewer1:32:42
So I think there is a lot of work to be done there as well.
Y
Yon Stoka1:32:46
So we need to make safety a primary part of open research in AI if we're going to win this tide shift, besides needing to raise a lot of resources. And AI should be one of the technologies raising all the boats because it can be safe, but it can still take your job, right?
I
Interviewer1:33:01
Ah okay. Yeah. Separate separate.
Y
Yon Stoka1:33:02
So I think that's kind of what is a worry out there.
I
Interviewer1:33:06
Okay. Okay. Well, I think you and I are on the same page about this. So, thank you, Yon, for joining, and let's give Yon a round of applause.
After 3 days of conversations with leaders in open research talking about things ranging from civic discourse to geopolitical politics, it feels like we are on the brink of something, like we are at the frontier. No, I think we are the frontier. It's us. The technologies that we've been talking about, I think, are species level. And by building them, we are charting a course for humanity. And we can screw it all up, or we can help humanity solve problems that have felt impossible for thousands of years. But we can't move our species forward if the only objective function we use is making money. Closed AI alone can't do it. We must open the frontier for all of humanity. To do it, we are going to have to get into a room together. And not just us, all of the leaders of open research, open discourse, open weight models. So, let's do that. Let's get a room. I'll rent it. Openfrontier.ai is a website that lists all of the leaders that I've spoken to in the last 3 weeks of frontier AI projects. And this is just a beginning of the list. The first set of people I could talk to in the last 3 weeks. Nobody that I reached out to has said no. And everybody I've reached out to has, in their own words, said, 'Fuck yeah.' So, we're going to bring together the top 100 leaders of Open for one day in San Francisco in 4 months. And this meeting will be called Open Frontier. And I think these people will decide where Open goes next. And I hope to see you there. Thank you.
I hope you enjoyed the first law. I'm going to bring you a lot more interesting conversations from the greatest minds in AI at the frontier. We are going to keep putting microphones and cameras on the most important minds that are defining what's next for our species. I'll see you next time.