Jonathan Ross28:44
Right now, when you go to use AI, it's a little it may feel somewhat fast, but that's because you're not used to using it much faster. Sort of like when you first use the internet, it felt fast compared to mailing things around. But when you got broadband, you realized, oh my gosh, like this is so much better. I'm never going to go back. The difference is broadband actually needed people to make their websites faster to make it usable. If the servers are slow, you don't get a benefit. So, it took a while to roll it out and make it good and get video that could stream and all that. The difference is you put these LPUs into a system and all of a sudden the generation of tokens gets faster. It's like getting broadband instantly on these existing models. And so now rather than having to wait a minute to get an answer you can get an answer in 10 seconds and that really starts to compound. So let me walk through an example of why it's not just speed but it's also quality.
The other thing that I did was I created Google TPU and at Google there was this time where I'd already moved over to Google X. I was no longer working on the TPU at that point. had already done it and someone else from the TPU team was at Google X came by, showed me this email from Deep Mind saying, 'Hey, we've got this competition. We think we're going to lose. There's a prize purse. Is your chip as fast as we've heard?' And we're like, 'Yes.' Like, it's like Ghostbusters, when someone asks you, 'Are you a god?' You say, 'Yes.' Someone asks you, 'Is your chip as fast as I've heard?' You're like, 'Yes.' So we reply back yes and they're like great competition's in 30 days. We're going to play the world champion and go and we played our test games and we lost. We need to win. So they had no choice but to port over to that TPU chip. So we did it and a bunch of interesting things happened. I think Do you know what an ELO score is? Yeah. Okay. So for those who don't know what an ELO score is, it's sort of your ranking in chess or go and something like a 200 point advantage is insurmountable. Like the probability of you winning is basically zero. Alpha Go running on GPUs had an ELO score of about 3,200. Lee Sedol was about 3,550. And I might be getting the first digit wrong. It might be 2,000 instead of 3,000, but it was more than 200. And then when put on LPUs, it actually went to like 3,900 or something ridiculous or 2,900, whatever the first digit was. It jumped dramatically. And so I think he didn't expect he was going to lose, but he lost badly, but it was the exact same model. So what changed was the ability to compute more made the results smarter. The way that these models work, thinking fast, thinking slow from Daniel Kahneman. Yeah, I read that book. What AI does is if I have 270 possible moves, which is what you have on a go board, the AI is going to rank those moves and say this is the best move, this is the next best, and so on. What happens is you virtually play that best move, and then you virtually play the counter move, and then you virtually play the next one. and then you see how the game unfolds. What you can also do is try that second best move. And occasionally that second best move when you play it out actually turns out to be the right move in this context. It's not the one that you would normally do, but in the context it's better. And you can see that as you play it out. In the second game, there was this famous move called move 37, which was creative. It was original. It actually wasn't completely original. It was a 1 in 10,000 game move. It had been in the cannon of games that we're trained on, but when we went back and played it on GPUs, it never found that move because it was too deep in the chain. So over time, these GPUs have gotten as good and better than TPUs. As the creator of the TPU, I have to admit GPUs are now better. This is the benefit of having an entire industry behind you and ecosystem and everything. But at the time TPUs had some novel innovations. Right? Now bring in the LPU and you can go deeper faster. You can search faster. And so you can actually make a model smarter by making it faster. And so the realization is and now that you've got this ability to reflect, think deeply and change the outcome based on your thinking, being able to think faster makes you think smarter. And so that's the advantage of pairing the LPU and the GPU.