Jonathan Ross0:08
So, I want to start out. Thank you very much for asking me to do this with you. It's a great honor. Let's get into what the partnership is and how this came about.
Yeah. So, and first of all, thanks for David coming. He normally doesn't do this kind of thing, so thank you very much. It all started early in 2025 when Nvidia released NVLink and was going to allow partners to connect to it. That's when Groq's COO, Sunny Madra, raise your hand. Yeah, that guy right there. He reached out to Jensen and said, "Hey, can we connect to this?" And Jensen said, "Sure, why not?" And so we got some GPUs and we didn't have NVLink at the time. And we just had Ethernet, but we got some GPUs and we started trying to get something working across GPUs and LPUs where we took different portions of the workload that ran on different chips or ran better on each different chip on those chips. It worked. We presented it to Jensen. 3 days later, Jensen called up and said, "Why don't we work more closely together?" Three weeks later, the deal was done. And one day after that, I was at NVIDIA working full-time and that was actually December 25th, Christmas. That's when I got my laptop and started working.
Perfect. So, can you tell us what the product is?
Yeah. So, let me start up with I can't see the slides, so I have no idea. Oh. Oh. So why don't we start off with me explaining it with words, then? All right. We're going to do the acappella. So, the best way to describe it is if you were building out a logistics network for the entire United States and I told you your two options were you could either use all 18 wheelers or just delivery vans. Which one would you pick? I mean, 18-wheelers are great for the long haul freighting or freight, you know, transport. The delivery vans are great for the cities, but you certainly don't want to use 18-wheelers in cities. The best answer is both. So the product is we actually take the decode layer of the LLMs and we take the sort of weights and the FFN portion. Oh, we can see it now.
I don't think they can see it though.
Okay. Great. Perfect. So, we're going to start off a little bit backwards here. So, what we do is we take the FFN layers and we put them on the LPUs and then we take the attention layers and we put them on the GPUs. So if you have about 40 decoder layers, there's going to be about 40 round trips. Now this is why we started talking with Nvidia when Nvidia offered NVLink to the ecosystem. So this is a really important point. Nvidia is all about growing its partners and making new partners and just building the rest of the ecosystem. So what I would recommend is if you have an idea where you can use NVLink, reach out. That's how this whole thing got started. But this actually allows us to sort of bend those curves that you saw in the presentation which is Reuben is great if you want lowest cost tokens it's your best option. It also bends the curve so that you can get faster tokens but you see that LPU or LPX that's brought in and it sort of flattens out the curve there and raises it up. Well, now when you put them together, you can go all the way to this ultra performance, which is just otherwise impossible to get thousands of tokens per second. That little arrow there that points off to the right goes quite far off to the right.
So, why is this important? Now, I'm going to go back to an analogy. So, let's talk about the economics here. So, when you buy energy, if you produce that energy from coal versus oil, it costs you about 7 times as much per BTU that you get out of oil versus coal. Now, people aren't dumb. They're not going to overspend on something. So, why are people willing to spend seven times as much per BTU from oil versus coal? There are things that you cannot do with coal. You can drive a railroad with coal. You cannot fly a plane with coal. You need velocity. If you are trying to code and you have to wait 10 minutes to get a result, you'll accept it if you have no alternative. But if someone else is able to get the results in 5 minutes, they can get twice as many iterations per day. That means they can get to what the customers want in half the time. And so that speed is very important. You're trying to get a report done. If you can get the report in half the time, one-tenth the time, it fundamentally changes what you can do.
So what does this mean for everything else?
Um, everything else is in...
You were saying earlier like one of the things we were talking about on the phone is like speed is really important especially when AI starts talking to other AI.
Yeah. Well, let's try and explain agentic AI really simply. So who here finds AI useful in getting tasks done? Okay. So, would you be surprised to find out that AI finds AI useful in getting tasks done? Okay. So, when AI invokes AI to solve a task, that's agentic AI. And there's all sorts of reasons why it does this rather than just doing everything itself. One is it can fork the context, right? So, there's a certain speed at which you can get those tokens out. And if I can fork and try 10 things at the same time, then I can try those 10 things. On the other hand, if I'm able to do something in one-tenth the time, I can change what I'm doing at each step based on what I just learned. So you get those feedback loops faster. In agentic AI, there have been some breakthroughs recently where 100,000 plus line browsers have been written by AI. But if you could do that instead of in days, if you could do that in a couple of hours, then you could start iterating on that design pretty fundamentally. So it's simply a way to do research, to explore, to and actually just stepping back for a moment, we actually just don't know. Bottom line is we don't know all the things that can be done with fast AI because we haven't had it yet, right? It's like when you first got oil, what did everyone do? They started running ships with oil. You didn't need oil to run a ship. It was cleaner. It was better, right? But what we knew was that ships, you know, being powered by energy was useful. We couldn't have imagined that we would have an internal combustion engine that could fly a plane. We don't know everything that's going to be done with speed. We all know that we want it. We all know that we want the answers faster. We're going to want to code faster. We're going to want our results faster from research reports, but there are going to be things that can be done. Think about having a conversational agent speaking to you where instead of when you ask it, you know, what is my bill? It goes, that's a very interesting question. Trying to buy time in order to generate an answer, what happens if it immediately answers? You're going to feel more engaged. You're going to feel like you're in flow and that conversation is going to feel much better.
So this is something I was actually talking to Jeff Bezos about because the way he describes AI you can go back in on YouTube and you see this TED talk he gave back in like I don't know 18 20 years ago and he was saying that there's these inventions of new technology there are these thin horizontal enabling layers and so he said like electricity is one he thought the internet was another one and the way he describes AI he talks about this thin horizontal enabling layer and what he told me and he said in the talk he's like listen when we invented electricity we weren't inventing electricity we were trying to put lighting in a house right but then there's no way after you did that, you could have possibly prevented all the different ways that electricity was going to be used. You just know it was going to go everywhere.
Yeah.
And so your point was AI is we know is we can't predict what is going to be created with AI and we certainly cannot predict what's going to be created with super fast AI.
Correct.
Okay. So yeah, going back to that point, when electricity was first invented and was put in houses and you started getting other appliances like washing machines, they didn't have prongs where you'd plug it into an outlet.
No, they put it in the...
You screw it into a light bulb.
Exactly. So you've seen this cuz he talks about that. Okay.
I don't know if the same source, but Yeah.
Yeah. But he talked about it on the...
People couldn't imagine that you have other uses besides lighting.
Yeah. But going back to this AI, it's not like the internet. It's not like telephones. Those are information age technologies. Information age technologies are about taking information, duplicating it, and distributing it, right? Some are better than others. But the internet is conceptually the same thing as the printing press. It's just faster. It's just cheaper. It's more convenient, but you're copying data and you're replicating it. AI is not copying data. It's creating answers that didn't exist before. And so AI is not an improvement like electricity. It's bigger. It's more foundational because we've never had the ability to offload creativity and this other kind of mental labor before.
Okay, this is perfect because we were talking about earlier how one of the biggest problems you had back at Groq was getting people to understand why speed matters. Yeah. And the conversation we were having earlier is like the difference between a company is trying to preserve what it has versus a company that's going to generate new value. And so speed, what will speed do to the second kind of company?
Well, there's two kinds of companies. There's value preservation and then there's value creation. If you're value preservation, you're going to use AI to reduce costs. If you are value creation, you're going to use AI to increase your revenues. And one of the most important things when you're increasing your revenues is that you're going to want to iterate more quickly. A lot of SaaS companies, a lot of web companies can push a new build once a week. What if you could do it once a day? What if you could do it once an hour? Going to an example of something that we did as a team. We were meeting with a customer and we got a customer request in that meeting and I sent a message to one of our engineers who using a coding agent implemented that feature and it was done I think less than an hour after the meeting was over. Now if we had more speed that could have been done before the meeting was over. Just imagine you meet with a customer they have a problem before the meeting is over it's solved. We just don't we're not able to wrap our heads around that because we can't do it yet. This is what that's going to enable.
Explain why people are always going to want to spend more money to go faster.
Think of it as competition. So, actually, let's take Nvidia. So, Nvidia went from us starting December 25th to us unveiling the LPX rack on stage today. Do you want to be half that speed? Everyone's going to want to be as fast as they possibly can. The best businesses move fast. And you're an expert in this. I'm going to ask you, is that true?
Yeah, of course.
Yeah. What business wants to move slower than everyone else? And especially when there's value creation, the faster you can create those new features, those new products, those new abilities, the faster you can bring that revenue in.
I think the important point you were making on the phone call we had earlier, it's like the product is already in production. Yeah, it's already being produced right now. And as Jensen announced, Q3 is when it's going to be available. It's probably one of the fastest ramps of a semiconductor in history.
Nvidia's legal department wanted us to insert the word probably there, just so you guys know. That's not a joke. Oh, I'm sorry. I'm not supposed to curse either. Um, real quick, why or not real quick, we have plenty of time. Um, why did you say that being the king of inference is a big deal at NVIDIA?
Well, everything started with training. We had to get models that worked. But the, you know, training scales with the number of researchers you have and inference scales with the number of users you have. Revenue comes from users, not researchers. The better your training is, the better your inference is, but the better your inference is, the more revenue you generate, the better your training gets. It's a virtuous cycle. If you are the best at inference, it means that you're going to be enabling that revenue for your customers and they're also going to buy more training hardware. But you can't be just one. There's a virtuous cycle.
How did can you explain in detail how you made it both faster and more cost-effective?
The way that we did that was so remember that Pareto curve here. Let's go to this one. So, what we're really showing here is throughput per megawatt. So, it goes back to what Jensen said. You're going to buy a gigawatt data center. You're going to fill it with hardware. And the question you're going to ask yourself is what are you going to run in that data center? How are you going to run it to earn as much revenue as you can or to get as much productivity as you can for your workforce? And what you see is the faster those tokens are produced, the fewer tokens you get per megawatt or per gigawatt. That's the trade-off. The LPU is really, really good at low latency. The GPU is really, really good at high throughput. But going back to the logistics network, if you deliver the long haul part with the 18-wheeler and then you hand it off to delivery vans to deliver the rest of it, you're going to get the most cost effective delivery network. Well, what you see here is that bend in the sort of curve. If you compare Blackwell to Reuben with LPX, you see that it doesn't just keep going down like it did. It bends up as you start going faster. That's because the LPU is particularly good at the high throughput matrix multiplies in the decode layer. So if you were to run something purely on GPU or purely on LPU, there would be different matrix multiplies that got different utilizations. If you ran everything on the LPU, you'd be underutilizing it on attention. If you run everything on the GPU, you underutilize it on the projection or the FFN layers. Putting them together, the utilization goes up for both. The speed also improves because the LPUs don't have to wait to fetch things from external memory. So now you actually get more throughput out of the same hardware while getting more latency or better latency. And it's not simply just faster is better. It's faster at what cost? It's faster at what capacity? The world right now doesn't have enough inference compute. It doesn't have enough energy to power enough inference compute. But if you can get more tokens out of the same amount of energy, then you can start to provide high performance tokens to everyone. One of the things at Groq that was great was we had very fast AI, but we didn't have the throughput to deliver it to those customers. We couldn't run anything large and customers always wanted to adopt us, but they couldn't. Now you mix in the GPU and you can do this at scale.
Can you tell the story that you told me earlier about using a model that's average at and how you were using your technology to actually have it solve math theorems?
So, one of the points here is that speed is intelligence. I mean, just think about it. The faster someone thinks probably the smarter they are, right? Same with AI. But why is that? There was a great example. So we were doing formal reasoning on the LPUs trying to prove correctness of portions of the circuitry and one of the folks at Groq did an experiment and they realized if they ran something on Opus, Anthropic's Opus at the time they got the answer in the smallest number of iterations but at the greatest cost. When they ran Qwen 32B, a good model, not a great model, but a good solid model, especially for its size, it would take more iterations, but it solved every single one of the math problems faster and cheaper than Opus did because it would just iterate until it got the answer. So on programming problems where you can get a debug an like you get a error on compile the tests run and you get an error you can actually iterate much more quickly with a less capable model. Now if you have a more capable model and it can run fast well that's even better.
You were doing that internally at Groq. Correct?
Yes.
What are some ways other people use this technology outside of Groq that's most surprised you?
Um, I would probably go back to voice again. So, voice was pretty heavily used.
Can you tell them why your girlfriend gets mad at voice?
So, my girlfriend uses ChatGPT and she mostly uses it in voice mode. And when she asks it a question, it starts off by answering that is a very good question. And then it'll go hm let me think about that. And she's like, "I just want the answer. Stop telling me all this." It's because it's slow.
You do that because you're adding filler words.
You're adding filler words. Yeah.
And those filler words are better than silence, but they're not much better.
If you if they might I don't know. I think I would prefer silence.
She certainly would. Um, so if you think about the most engaging conversations you've ever had, those are conversations where as you are stopping speaking, the other person is saying something that answers your question. They're not going in a different direction. They understood you. They really got you. They're aligned, but they're answering immediately. It feels like they're reading your mind. That's why speed is going to matter.
So if you will explain why how this applies to voice and like the other ways that you think it's going to apply to voice besides customer service like we use the example of calling call centers but like what else?
Coding. Some of the engineers here at NVIDIA, one of them on the LPU team has connected up I think to either Codex or Claude Code. I don't know which one, but they just have their phone connected to it and they just talk to it and have it make all the updates. So they're not typing anymore. They're just they use the computer to interact with the applications and stuff, but all of the direction of what the AI should do is voice. Our mutual friends Eric and Kareem from Ramp told me the same thing that they have a bunch of people engineers in their company. They're not typing anything.
Yeah.
Which is not going to be great for open floor plans.
Um if you you created technology, you invented technology, but if you were sitting in the audience, what are some ways that you think you would be using this if you were in their shoes?
I would probably focus on giving your premium customers speed. You certainly don't want to be slower than anyone else. You know in these charts we refer to this as a free tier at...
Go to the dollar.
Oh yeah.
Yeah. That one should just be up permanent display. Look at this.
I was asking you before we even came out here. Which plan would you use?
I always pay for time.
Yeah. Ultra.
So basically anything I buy.
Exactly. So slow pair private. There you go. So the thing is speed matters because you're another way to think of it is if you have engineers that aren't worth giving the ultra package to, are they good enough to hire? Like don't you want all your engineers to be so good that it's worth speeding them up? You should be giving the fastest speed tokens to all of your best engineers and then the other engineers are going to ask why don't I get the best tokens there.
You were mentioning the other two companies Anthropic and OpenAI are already doing this. Correct.
Yeah, they already offer plans where you can pay to get a speed up and it's I believe super linear. They charge more than linear for the speed up. So this is something that's already needed. I was using open-source models on Groq LPUs because even though like when I would ask a question, it normally wasn't perfect. I could immediately ask a follow-up and it would correct it versus having to wait five or 10 minutes for my answer on a better model. But if I could get a really good answer in a second, I'll take that every time.
What was this?
There's one more analogy here. Go.
So when we're looking at how much token capacity we're willing to give to engineers, like there are, you know, people spending $10,000 a month right now. It's on the high end. But that might be on the low end of where this is going. In fact, there's plenty of careers out there where people are responsible for things that are worth tens or hundreds of millions of dollars. Pilots, pilots fly airplanes that are worth tens of millions, actually hundreds of millions of dollars. And if you can generate the revenue, that makes sense. It probably won't be long until engineers are actually using millions of dollars of tokens per year and you're going to want to give them the fastest tokens possible because they're the best engineers.
I have a weird question that just popped to mind. Your name's Sunny, right? Okay. So, what was the because you're essentially encouraging them to use like find creative ways to use the latest technology, right?
It's a very old idea. Yeah.
From history of entrepreneurship. Andrew Carnegie in his autobiography would talk about the fact that like you know you have to invest heavily in technology the savings compound. It can give you it can be the difference between a profit and loss and give you a massive competitive advantage over your slower moving competitors which he applied to one of the most valuable industries in the world at the time during industrial revolution which is steel right.
Do you remember the breakfast we had in New York like two years ago? Yeah, it wasn't that long ago or a year and a half when I don't feels. Yeah, it does. Right.
This before you called Jensen and asked for access, right? That one idea of like, hey, this new thing came out. Let me figure out creative ways to use this new technology. What was the difference of value in your business? The state of Groq, I mean, Groq was doing okay, but certainly not. That was like life-changing. Correct.
So, this is really important. So Sunny proposed this to me and at first I was against it. Not because I didn't think it was a good idea. I just didn't know if it was going to work and there were a whole bunch of other things that were going to work.
But what was the state of the business at that time?
Well, this we were hitting some revenue targets, but they weren't nearly as big as the deal that was done.
Exactly. So it might have 5x 10x the value of the business.
Way more. But but I'm actually making another important point,
which is I hope you got paid.
Oh yeah, Sunny's okay. But but actually I'm making a more important point here which is we almost didn't do it. Whereas if we could have had AI at the time because AI wasn't at the state where it is today which is less than a year ago. So it's crazy, right? If we had AI where it's at today, it would have just tried the experiment. So the faster it get we would have just tried the experiment of seeing if we could get it to work. We'd have had AI go do it. Like I run experiments all day long.
That happened today.
Yeah.
Okay.
It would have been no question. The only question was opportunity cost. We had a finite amount of engineers. We were trying to deliver things for customers. Those customers told us how many dollars they were going to give us if we delivered. And so we were on a path and Sunny advocated for this and I said fine, take a small contingent and do it. It took longer than it would have, right? But now with AI, you can just run experiments. It's cheap relative to the lost opportunity.
So imagine so you advocated I would imagine advocated pretty hard. Imagine if you said no the difference in the valuation of the company is one-tenth 120th 150th.
I mean it wasn't that much of a difference but it was big.
Yeah. Substantial.
Yeah.
So I mean AI is going to allow you to run experiments. It's going to allow you to run them quickly. And if we run the experiment faster than someone else then we have the advantage. If you run the experiment someone else,
I guess this is the main point. It's like you're developing new technology to make other businesses more valuable. You're almost like your own proof of concept for that.
Yeah.
So I can't disagree with that.
Yeah. It's just like came to mind when well you want different example. Claude Code writes Claude Code.
Yeah.
Yeah. So Sunny just said Claude Code writes Claude Code.
And if you look at how fast it updates like more than once a day they're updating accelerate.
Yeah. saying one more than once a day it's updating.
I run experiments so quickly now on things that...
I just whispered in your ear. They were doing the introductions. I don't know if you want to talk about this or not. I was like, "What are you hacking on?"
Yeah. I'm not going to I was going to show you. No.
Okay. Okay. Show me after.
No. The reason I asked is cuz I just had Tobi Lutke on my new podcast. Yeah.
And we talked for like four straight hours. And what I like about him is like he's your favorite founder's favorite founder. So like people that don't like any other founders or any other products, they all talk about the fact they respect him. And I think he said this on the episode but he's like, you know, somebody's going to build the AI native version of Shopify and it's going to be me. And so like that's what he does at night when he's done working. Like he's trying to rebuild it as if from a complete blank sheet of paper. So I was curious like what you're doing.
I'm just doing AI coding just non-stop. It's all I'm doing right now.
You're like what's your schedule? Oh. Um, wake up at like 6:00 am. Um, spend the first couple of hours doing so I start off with an AI delivered daily briefing. Um, it tells me all the most important things from my email, my calendar, world events. It integrates Poly Market, Khi, everything, you know, market numbers, tries to give me an update of everything that's going on. Um takes news, integrates it, and then I read through that and I ask it follow-up. In fact, I told it to stop writing long descriptions. Just give me the title, and if I'm interested, I'll ask a follow-up question. And so, that's the first thing I do. And that takes me a little while to go through. It organizes all my emails for me. Um it drafts things for me. And then after I get past that, um I'll start coding a little bit. Um and then I'll start sending some emails to people. And then I come into the office and then when I'm in the office I'm in meetings just dreaming of getting back to coding and I was I wasn't coding pre-AI like AI coding I had stopped coding. I just didn't have the time for it. Now even while I'm in a meeting I'll have an idea and I'll type it. I'll run the experiment. I'll go back to the meeting.
How much sleep are you getting?
Oh six hours a night. Who do you like? A lot of people are coming here. They want to hear you speak. Right before we got on stage, he's like sometimes I wonder why anybody wants to hear me speak. I'm like, Jonathan, what are you crazy? Like you have this very unique lived experience. You have stuff in your head, very important, valuable information about the most important thing going on. Of course, people are going to hear you speak. Who do you go to learn like other things that you don't know about AI?
From I mean actually I just go on X and I hear things from Andrej Karpathy. I hear things from actually the daily brief usually has a bunch of new papers in it. So I'll just read whatever's in there. Um so AI I guess is teaching me about AI.
So there you go. Loops all the way down.
Yeah.
Um we're...
and also of course I listen to Founders Podcast a lot.
Of course.
I was just telling David I think I've listened to every single one that has made it to YouTube. Um, and a lot of the lessons for Groq were actually in there.
Oh, that's incredible.
Yeah.
How much of the purchase price are you going to give to me?
Well, I should have let you invest.
I think you did offer it. I said no. I'm pretty stupid with these kind of things. Um, we have eight more minutes. I don't have any other questions. I know you said no Q&A, but can I just like violate this and we take questions for eight minutes? And I'm sorry. I think I did two curse words. That's way less than I normally did. I see some nervousness over from the press relations, but yeah, let's do it.
What could possibly go wrong?
Yeah, there's a they have microphones set up. Why'd you say no questions?
All right. So, right now, as you guys have said, it's super linear to get fast mode and Claude or Codex. Do you think that's going to continue into the future?
I do. I think it will be super linear because there is an increased cost. What we're doing is we're improving the cost significantly, but there is an increased cost. And there's always a point on this chart, which goes way out way off of where this chart is that some people will want that others won't. Um but just imagine if there's some sort of emergency situation like a FEMA situation, you're going to want the fastest possible tokens responding to that and trying to action it. There's going to be use for the fastest possible token, but no matter what speed it's available here and it will be you'll get more of them per gigawatt or per megawatt deployed.
Did you guys kill the Reuben CPX?
Well, hey, one question per person. What are we doing?
Yeah, next time.
Can I go next?
Can you ask my question?
I just do LPU stuff.
Yeah. So, I think the LPX rack is cheaper than NVL72. So, why limit it to just the premium tier? Why not all workloads?
So, the LPX rack is actually very dense and has a lot more hardware in it. So, I don't actually agree that it's less costly. There's a lot more there's a lot of silicon in that rack or hero chips in that rack.
Hi, V from Microsoft slightly technical question but you spoke about the utilization of both GPUs and LPUs for different portions of the workload which totally makes sense but there is still the scale out latency cost you have to pay to transfer the activations. What secret sauce do you actually use to push towards the right hand side of the Pareto curve to hide that latency?
Well, one of Nvidia's best-kept secrets is we're really good at networking. Bought this company called Mellanox and have some nice gear there. So, we have gotten the latency down quite a bit on that. But we do have some improvements coming that will make it even better and so you'll see these tighten up a bit.
I actually have a question that you said earlier and I think it's important. Can you describe the speed at which Nvidia because from the outside you thought they were fast but now you're inside we were talking about this earlier just how fast they actually are.
Yeah. So at Groq we were a startup of 450 people and we moved incredibly quickly and most people thought it was crazy how quickly we could move. Um Nvidia's what are we over 40,000 people. Uh yeah okay getting thumbs up. Uh and we move just as fast no slower. No real bureaucracy, just things happen super quick. Um, someone's needed for something, they jump on it. The culture is very unique in that way. Uh, I've been at other large companies that didn't work this fast.
Google.
So, I was trying.
I said it, not me. Or, not him.
Oh, um, yeah, I'm going to get in trouble. Thanks.
No, you shouldn't invite me. I don't know what to tell you.
Go ahead. Um, regarding RSI or recursive self-improvement, how does fast AI help with that? I'm assuming since you can do so many fast experimentation also blows up RSI.
So I am not on the AI research side. So anything that I say is purely a bunch of speculative tokens here and you know some draft model or some verifier is going to have to come by and correct them. But yeah, that's a joke for a very small number of the audience. But the way to think about it is one of the differences between the way that we teach models and the way that human beings learn and why we're able to learn on so little data is we're curious and we're curious about anomalies. We see things that violate our expectations and then we double click on them. And so when a baby sees something that doesn't make sense, like if you try and trick them and make them think, you know, something's hovering in the air or something, they'll be like they'll look at it intently because they don't understand. That's a very innate thing. And the reason is whenever something violates your expectations, that's a very interesting thing to generate more data about and to train on. When we're training these models, we're just training them on a whole bunch of random things. We go, what's 1+1? What's 2*3? What's the second derivative of the square of the hyperbolic tangent? As if that should follow, but it's not ready to receive that information yet. It's too far out. When you're improving a model based on its interaction with the world, you're actually able to set the difficulty of the problems right at where it's ready to receive that training data. So what I would expect is if you were able to generate tokens faster in like do the rollouts faster in sort of reinforcement context at least you could potentially narrow the amount of cycles needed for training by targeting more carefully. This is just speculation like this is not an area I'm aware is being done too deeply but there are possibilities.
Makes sense. Thank you. Yeah, we got two and a half minutes.
Awesome. Thanks for doing this, Jonathan. This is really cool. I'm Nick from Midnight Capital on X. Um, so basically my question is, when we were listening to Jensen at CES recently, it sounded like the idea for Groq was at some point a little bit more narrow. Like I have the Meta Ray-Ban glasses and he was talking about how super low inference or super low latency inference would be useful for this type of application but it seemed like it was a more narrow thing. Today I think what we heard was a much more bigger roll out of Groq like a platform-wide thing where it's integrated into like the entire Reuben architecture. So I'm just wondering like was the interpretation that I had initially correct and then there was a change as far as how Groq was going to be integrated and if you can just talk a little bit about that. Thank you.
Well, Jensen likes to set expectations here and then over-deliver. So I think he's done that again. The expectations were that it would be used quite broadly but there was a lot of work done internally to verify that that could be the case before that was rolled out more broadly. At this point, you know, we're in production so more can be said.
Thank you.
Thank you. Um a quick question following up the question by the Microsoft gentleman on this Pareto curve in terms of a performance improvement. Can I guess right now the main bottleneck is the rack-to-rack latency and the bandwidth or in other words how do we get back to over 1.5 like say million tokens across the x-axis in the future what's the bottleneck I thought it's rack-to-rack.
You asked a question that only an LPU could answer in the amount of time left but I'll give a really good I'll give a high level so the number of bottlenecks so often times when doing computer architecture, you get focused on one bottleneck. Um, you know, I did a bunch of spreadsheet analysis on some of the more recent models and there was a point where a single layer of one of the open-source models with all the matmuls, one of the matmuls was bottlenecked on memory capacity, one was bottlenecked on memory throughput, one was bottlenecked on compute, one was bottlenecked on networking latency, one was bottlenecked on networking bandwidth. Literally everything you can imagine was bottlenecking one of them and if you improve that something else would be bottlenecked. You're right that if we do that it will speed things up but these are fairly balanced architectures where you have to improve a lot of things in order to get a noticeable improvement. But there's a lot of room for improvement going forward.