Back
Jonathan Ross
Chief Software Architect, Nvidia (formerly Groq CEO)

The Inference Revolution: Groq, Nvidia and the Future of AI

🎥 May 18, 2026 📺 Sohn Conference Foundation ⏱ 15m 👁 9206 views
At Sohn Investment Conference 2026, Jonathan Ross, Chief Software Architect, Nvidia, and John Yetimoglu, CIO, Infinitum, discuss the current landscape and future trajectory of AI.
Watch on YouTube

About Jonathan Ross

Jonathan Ross, founder of Groq and inventor of Google's TPU, has been discussing the company's near-collapse and subsequent partnership with Nvidia. He described a period when Groq was about three weeks from running out of money and considered layoffs, but instead implemented a salary-for-equity exchange called "Groq bonds," which he said 80% of employees participated in, with about half reducing their salaries to the statutory minimum. Ross stated that this move saved the company. He also discussed a $20 billion partnership with Nvidia, describing it as combining GPUs and LPUs for better performance, analogous to using both 18-wheelers and delivery vans in a logistics network. Ross has spoken about the future of AI compute, stating that "as long as there are unsolved problems in civilization, we will have a need for more compute" and that "there's no way to satiate the appetite for intelligence." He has advocated for revamping education to focus on asking questions rather than providing answers, saying "if you can come up with the right question, AI can go answer it for you." Ross also distinguished between intelligence and sentience, defining sentience as "your rate of improvement in your intelligence" and suggesting that AI is creating a feedback loop that accelerates societal intelligence.

Source: AI-verified profile updated from Jonathan Ross's recent appearances. Browse all interviews →

Transcript (30 segments)
U
Unknown0:01
Thank you to the Irish Open conference for having me back again this year. This is really a wonderful conference for a great cause and I'm honored to be invited back with Jonathan.
Jonathan, you guys might have seen, has been in the news a little bit lately. He's the chief software architect of Nvidia, founder and CEO of Groq, and inventor of Google's TPU. He's had four first-time right silicones. He's one of my dearest friends and someone I'd consider to be one of the greatest engineers of our time.
J
Jonathan Ross0:40
Well, thank you. And for those of you who don't know John, although probably most of you do, John's one of those rare hedge fund managers who both gets great returns and provides immense entertainment value. Watching you hold those shorts like you've got diamond hands. I don't know how you do it. Everyone else is like sweating for you.
U
Unknown0:59
Not everyone's supposed to know I'm like a crazy short seller, so...
J
Jonathan Ross1:04
Oops.
U
Unknown1:10
Crazy. Can't believe you just docs me like that.
So, you know, Jonathan called everything that's happening today like over 10 years ago. Everyone thought he was crazy. I was one of the lucky few that believed him back then. And, you know, up until a few years ago nobody even knew what inference was and so, I think that's a good place to start given it's now one of the most, probably the most important thing about AI today.
J
Jonathan Ross1:43
Well, for once I don't have to explain what inference is, that's nice.
Yeah, and so, you want me to get into the economics of what do you want?
U
Unknown1:53
Yeah, like, you know, one thing I think is very important is that, you know, portability across architectures is diminishing and porting is no longer really an engineering inconvenience but has evolved into more of an economic handicap. So, maybe we can talk through some of the tokenomics of, you know, inference.
J
Jonathan Ross2:14
Yeah, a good parallel here is we landed people on the moon about 60 years ago, roughly, and that was 60-year-old technology that got us there. What we're doing in AI is brand new technology. It's cutting edge even for today. We finally have enough compute to do these things with models. We've had the algorithm since the 1970s. So, what I'm seeing a lot of is people will look at one small part of the stack and think that that is the most crucial part, that's the bottleneck. And I don't think people understand just how much you have to build to make AI work. It's the chips, the packaging, the systems, the networking, the data centers, the servers, the racks, the power, all the stuff. But then also on top of that, it's both training and inference, which are two completely different problems. A little bit like, you know, getting to the moon and then landing on the moon were two very separate problems. And in terms of bottlenecks and all this, the focus I see a lot in investors are asking questions is what is the bottleneck? What's the next thing that's going to be constrained? What should I do next? And that's not how this industry works. Every time a bottleneck gets big enough, people solve it. So, as a person running a business, you're always going, 'Hey, what is my biggest problem and can I solve it?' And when you look at some of the components that are limited, when they're a problem but not a huge problem, people can charge a lot of money for it. But as they start to become a bigger problem, people start solving that problem. And so it's this constant shifting dynamic of where in the supply chain is the biggest problem, but don't become too big of a problem because then it'll get solved.
U
Unknown4:12
Yeah. So like memory is like the big theme today, right? It's the biggest bottleneck in AI. But so what you're saying is if memory continues to become a larger and larger bottleneck and continues to be extremely supply constrained that you think it's going to be solved?
J
Jonathan Ross4:30
Yeah, exactly. So memory used to be a commodity.
U
Unknown4:34
Correct. Yeah, it was the most commoditized segment of the semiconductor supply chain, pretty much.
J
Jonathan Ross4:40
Yeah, maybe a good way to explain this. There are two kind of goods that you can charge a lot for. One is a Veblen good and the other is a Giffen good. Veblen goods, the more you charge, the more desirable it becomes, like luxury goods. Giffen goods are a little bit different. A Giffen good is, an example is rice. So if you start to charge more for rice, some people can no longer afford steak, so they actually spend more money on the rice. So by raising the price, not for a luxury good, but by raising the price for a staple, you can actually increase the value. And that happens to some extent until someone goes, 'Let's stop eating rice and let's eat corn or something else, right?' And so there's this inverse sort of headwind where if memory is too expensive and if people don't build enough of it, people are going to solve that problem technologically.
U
Unknown5:34
So you think algorithmic efficiencies like, you know, for example, like Deep Seek's latest model release, the V4 release, it compressed KV cache by 90%. But the argument around that has been Jevons paradox, Jevons paradox, Jevons paradox.
J
Jonathan Ross5:50
Those engineers, them working on that problem is an opportunity cost. They could have been working on something else. If they had enough memory, they would have worked on something else. If it wasn't so expensive, they would have worked on something else. You make it the big enough problem, it's the tall poppy. As soon as it gets too tall, it gets chopped down.
U
Unknown6:03
Got it. Yeah. Cool. So...
J
Jonathan Ross6:15
Start building more memory fabs. That's the solution.
U
Unknown6:19
Yeah. So what do you think... We talk about diminishing returns on intelligence a lot, right? You and I, I mean, the way we prepared for this presentation is we literally just went through our conversations and thought, 'Oh, that was a good topic. That was a good topic.' And then we picked...
J
Jonathan Ross6:38
Long-standing disagreement between John and I.
U
Unknown6:41
Yeah. I think intelligence has diminishing returns and at a certain point the models get so smart above PhD level and human beings, you know, can't really understand the differences between one or the other. And then, you know, couple that with the closed-source versus open-source frontier landscape where open-source models are roughly 6 months behind the closed-weight model labs. So, you know, you can kind of make the assumption if you buy my argument that there's diminishing returns over time that the open-weight landscape will catch up to the closed-weight landscape. Now, there's a lot of nuances around that. So, I want to hear your thoughts, but like share your thoughts with us.
J
Jonathan Ross7:27
Well, I might surprise you a little. So, there are a lot of areas in the economy where if you produce more of something, it becomes less desirable. And then there's others where it becomes more desirable. And I would argue that intelligence, there's no way to satiate the appetite for intelligence. The more intelligence you get, the more intelligence you want. Now, I'll break it down. First of all, is intelligence going to plateau? Because that's important for decisions. And then second of all, not only would intelligence plateau, if it didn't plateau, would we get to enough of it where it'd be like we don't need any more? So, in terms of having enough intelligence, the first thing is as long as cancer isn't cured, as long as people still die of old age, as long as we don't have enough compute to run some of these AI models, we don't have enough intelligence. So, there's an economic incentive to keep building smarter and smarter machines to help us solve bigger and bigger problems. The second is competition. Everyone in this room probably could retire. You probably don't need to make more money, but you keep investing and you keep competing with each other. Why? It's competition. We just do it. We're humans. And competition isn't going away. And so, if my AI is less intelligent than your AI, and I can't tell the difference directly, I'm still going to be able to tell the difference in my returns based on which one I'm using. I'm going to want the better AI. And so, whenever I program, I actually use...
U
Unknown9:08
It's a gift to the Earth that Jonathan's programming again and writing code again. Just so, everyone should say thank you to him.
J
Jonathan Ross9:16
But I'm actually writing code which is being used for real things now again, thanks to AI, right? And I'll use multiple models. Each one is better at different things. But, you know, even though one model might be better than the other, I'll still use that model. And I think AI will know that this other AI is smarter, even if we can't tell the difference and will like... This is the whole agentic thing. So, why or what is agentic? Agentic is, do you get better productivity by using AI? Yeah. Well, so does AI. AI likes to use AI. So it calls to AI to do some task for it and return the result just like you do. It's just more of that. And the AI is going to recognize smarter AI and use that smarter AI.
U
Unknown10:06
There was actually a conversation we had outside with one of the hosts. I don't know if you remember about resumes.
J
Jonathan Ross10:12
Oh, yeah. Yeah. Yeah. So, I didn't know this. This is something I just learned. Someone did a study and showed that resumes generated from one LLM are preferred by that same LLM over the resumes from the other. Recruiters are now using LLMs to determine like who to interview. But you got to figure out which LLM the recruiter is using. So you should build one resume with Claude Opus 47 and one with ChatGPT and you'll have the highest probability of being selected basically.
So then the other question is, is intelligence going to saturate? Or are we just going to need more and more intelligence and this build out is going to make total sense? My argument for why it's not going to saturate is as follows, which is there are two components to intelligence and there's a great easily digestible book called Thinking Fast Thinking Slow that many of you have read by Daniel Kahneman that explains exactly what AI does. Thinking fast is the intuitive part. It's the, you're given a problem, do you have an answer immediately? Thinking slow is you iterate on it. So if you think about chess, speed chess is thinking fast, regular chess is thinking slow. You're evaluating multiple opportunities and recognizing better moves when you string moves together, right? And even though AI is coming from computers and so this is sort of a little bit hard for us to see, AI is very intuitive. It's actually better at being intuitive than we are. And that's because it's been trained on so much data. When Waymo is sending all their cars out, the amount of data that they get in a day is about, I don't know if it's now at the amount of experience that a human being gets in a lifetime of driving, but it's starting to approach that at least. When you're getting a lifetime of driving data in a day, you're seeing every possibility. You don't need to figure out how to deal with the fact that some I-beam is falling off the back of a freight truck and about to hit you. You've seen it happen two to three times and you know exactly what to do. The thing is the more these models produce data, the more they're able to just intuitively deal with the situation because they've already seen it, right? And that's intuition when you just have the answer. So, when you train these models, what you do now, they used to be trained on just data pulled from the world and that human beings were producing. And now what we do is we use the models to generate the data that they get trained on. So, you have a model at this level of capability and it produces data at these levels of capability and it's gotten good enough that it can tell what good is. It keeps this data, trains, moves up to here. Then it produces data of this quality, prunes it to here, trains, and goes up to here and just keeps moving up. And so, at this point, we're seeing these models improve at a pretty linear rate. And so, there's no reason to believe that they're not going to get any smarter. We may not recognize the difference between two really smart models, but one will be much smarter than the other. And that matters in the context of competition. Competition and solving big unsolved problems.
U
Unknown13:36
Yeah. Right. Do you want to talk about sentience?
J
Jonathan Ross13:41
Oh, jeez. Okay. I have a hobby. My hobby is to take words that people have used for centuries that don't have a good concrete meaning and try and ascribe a definition to it that helps me understand the world. And so I'm going to give you my definition of sentience. So, first of all, intelligence is your ability to make a prediction or influence an outcome to what you want to have happen. But it's stationary. It's like you have an amount of intelligence that is fixed in these models. But sentience is your rate of improvement in your intelligence. That makes sense because when we talk about sentience, we're talking about the ability to sort of self-reflect and get better and that's an important part. So, I say that intelligence is your capability and sentience is your rate of change. But rate of change doesn't have to be binary. You're not sentient or not sentient. It's how sentient are you? Are you linearly sentient? Are you asymptotically sentient? Right? You look at the world's best Go player and he asymptoted. He stopped getting better because he didn't have better players to play against. So, people often conflate LLMs with just being intelligent, but there's something else that language gives you beyond intelligence. It gives you the ability to transfer information. Going back to the example I gave with Waymo, if you have an entire civilization producing information, each participant in that civilization gets to benefit from that distillation of knowledge. And so while intelligence is a property of an organism or an individual, sentience is a property of a civilization. AI is producing more intelligence. It's getting smarter. You interact with it, you get smarter. You ask better questions. You're making the AI smarter. And so there's this feedback loop of sentience that's accelerating in our society. And AI is contributing to that. And so as that happens, I would expect our kids to get much smarter than we ever were, just like we're probably smarter than our parents were because we had the internet.