Interviewer1:00:43
The subject of reasoning came up just a second ago. I think there's a perception that recently we made a lot of progress in reasoning. It's actually one of the main things that I think people are working on. We released a dataset recently called Sudoku Bench. I was actually quite happy to see it come up organically on your podcast a few weeks ago with Chris Moore. I wanted to tell you a little bit about this benchmark because I think I've been having a little bit of issue promoting it because it doesn't on the surface sound particularly interesting. Sudoku has a sort of feeling that it's already been solved. How interesting can a collection of Sudokus be for reasoning? We're not talking about normal Sudokus. We're talking about variant Sudokus. What variant Sudokus are are usually normal Sudokus, put the numbers one to nine in the row, the column, and the box, but then literally any additional rules on top of that. They're all handcrafted. They all have extremely different constraints. Constraints that actually require very strong natural language understanding. For example, there's one puzzle in the dataset where it tells you the constraints of the puzzle in natural language and then says, oh, by the way, one of the numbers in that description is wrong. You have to be able to meta reason about the rules themselves even before you start solving the puzzle. There are other puzzles where you have a maze overlaid on the Sudoku and the rat has to work out a way through the maze by following a path to the cheese. But then there are constraints on the path that it takes of what numbers and what they can be add up to. It's difficult to really describe how varied these variant Sudokus are. I think they're so varied that if anyone was actually able to beat our benchmark, they would necessarily have to have created an extremely powerful reasoning system. Right now, the best models get around 15%, but only on the very simplest and the very smallest Sudoku puzzles in the set. We're going to be putting out a blog post about GPT-5's performance, and it is a jump, but it's still completely unable to solve puzzles which humans can solve. What I really like about this dataset, and actually was the catalyst for me creating it in the first place, was that there was a quote by Andrej Karpathy saying, okay, so we have all this data from the internet, but what you really want if you wanted AGI, you wouldn't want all of the text that humans have ever created. You would actually want the thought traces in their head as they were creating the text. If you could actually learn from that, then you would get something really powerful. I thought to myself, well, that data must exist somewhere. My first thought was maybe philosophy, like there's a type of philosophy where you just write down your thoughts without thinking, like stream of consciousness. I thought maybe that could work. But then when I wasn't thinking about it and I was in my leisure time, I was watching a YouTube channel called Cracking the Cryptic, where these two British gentlemen will solve these extremely difficult Sudoku puzzles for you. Sometimes their videos are four hours long and they're professionals, this is their job. What was perfect, I realized, is they tell you in agonizing detail exactly what reasoning they used to solve those particular puzzles. We, with their permission, took all of their videos, which represents thousands of hours of very high quality human reasoning, like thought traces, and scraped them and made that available for imitation learning. We did try to do this internally. Turns out that I did a little bit too much of a good job of really creating a very difficult benchmark. We're still trying to get that stuff working and we'll publish it if we have some success. I want to really sell the fact that this reasoning benchmark really is different. Not only do you get something that's super grounded, you know exactly if it's right or wrong, so you can do RL to your heart's content, but you can't generalize very easily. Each puzzle is deliberately designed by hand to have a new and unique twist on the rules called a break-in that you have to understand. Right now, despite all the progress we've made, the current AI models can't take that leap. They can't find these break-ins. They'll fall back to, okay, I'll try five, I'll try six, I'll try seven. The reasoning becomes really boring and nothing like what you see in the transcripts that we've open-sourced from this YouTube channel. I just want to put the challenge out there that this is a really difficult benchmark and I think progress on this benchmark will really mean progress in AI generally.
Could you reflect a bit? After watching this Cracking the Cryptic YouTube channel, how diverse were the patterns? Chris was saying to me, oh, these guys go on Discord servers, they get these creative crazy ideas, and I'm obsessed. Maybe I'm just being idealistic, but I love this idea of there being a deductive closure of knowledge. That there's this big tree of reasoning and we're all in possession of different parts of the tree to different depths. The smarter and the more knowledgeable you are, the deeper down the tree you go. In this idealized form, there is one tree and all knowledge kind of originates or emanates from these abstract principles. We could in principle build reasoning engines that could just reason from first principles, and it might be computationally irreducible, so you have to perform all of the steps. It feels like because we're not in possession of the full tree, what we need to do is kind of fish around. We fish around to find Lego blocks. Oh, that's a good Lego block. I can apply that to this problem. Maybe that's just what we need to do in AI for the time being, is we need to just acquire as much of the tree as possible. But could we just do it all the way down?