CEOInterviews.AI
Start App
Jensen Huang
Co-Founder, Chief Executive Officer, President & Director, NVIDIA

#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution

📅 Jul 08, 2026 MindVoice Production 152 MIN 100 SEGMENTS · 2 SPEAKERS
Jensen Huang is the co-founder and CEO of NVIDIA, the world's most valuable company and the engine powering the AI ...

What Jensen Huang said

Written from the verified transcript and checked against it. Every figure links to the moment it was said.

Jensen Huang discussed Nvidia's shift to extreme co-design, explaining that the problem no longer fits in one computer and requires distributing computation across the entire stack. He detailed the strategic decision to put CUDA on GeForce, which was an existential risk but built the install base that became Nvidia's core moat. Huang outlined four scaling laws (pre-training, post-training, test time, agentic) and emphasized compute as the ultimate blocker, mitigated by efficiency gains and supply chain partnerships. He proposed using idle grid power through contractual agreements and graceful degradation, praised Elon Musk's systems thinking, and described Nvidia's 'speed of light' philosophy. He discussed China's innovation driven by talent, open source, and competition, and TSMC's trust-based relationship. Huang argued AGI is achieved now, with AI agents creating viral companies, and predicted Nvidia's growth is inevitable, with $3 trillion revenue possible. He emphasized intelligence as a commodity and humanity as the higher value.

Key takeaways

  1. Nvidia's biggest moat is the CUDA install base, built over 20 years with 43,000 employees and millions of developers.
  2. Huang believes AGI is achieved now, as AI agents can create viral web services, but odds of building Nvidia are zero.
  3. Nvidia's growth to $3 trillion revenue is 'extremely likely and inevitable' due to generative computing and AI factories.
  4. Huang proposed using idle grid power (60% of peak) via contractual agreements and graceful degradation to solve energy constraints.
  5. Nvidia's 'speed of light' philosophy tests everything against physical limits, rejecting continuous improvement in favor of first-principles redesign.

Numbers and commitments

FigureWhat it refers toTypeAt
60 Direct staff size at Nvidia metric 9:53
$1.5 billion Nvidia's market cap during CUDA on GeForce decision metric 13:33
50 gigawatts Desired supercomputer footprint for data centers metric 50:52
1 gigawatt Power needed per week in supply chain for 50 GW of supercomputers metric 54:50
60% Typical grid utilization vs peak metric 53:50
200,000 GPUs Colossus supercomputer GPU count metric 58:45
1.2 quadrillion transistors Transistor count in Vera Rubin pod metric 1:04:14
60 exaflops Compute power of Vera Rubin pod metric 1:04:14
1.3 million components Component count in NVL72 rack metric 1:04:34
$1,000 per million tokens Potential price for high-intelligence tokens price 1:30:48

Chapters

  1. 0:00Extreme co-design and system architecture
  2. 13:33CUDA on GeForce strategic decision
  3. 28:51Scaling laws and compute blockers
  4. 50:52Supply chain and power grid solutions
  5. 58:45Elon Musk and systems engineering
  6. 1:06:22China's innovation ecosystem
  7. 1:11:52Open source AI and Nemotron
  8. 1:16:21TSMC relationship and culture
  9. 1:20:06Nvidia's moat and AI factories
  10. 2:01:17AGI timeline and job impact

Questions asked in this interview

12
  1. 0:00What is the hardest part of co-designing a system with that many components?
  2. 9:39What does that process look like of designing it all together?
  3. 22:50How are you able to make those decisions?
  4. 53:23What are your hopes for how to solve the energy problem?
  5. 58:00What was the third thing you were mentioning?
  6. 1:01:12Are there parallels in the NVIDIA extreme systems co-design approach that you see in the way Elon approaches systems engineering?
  7. 1:05:59What do you understand about how China is able to build so many incredible world-class companies, engineering teams, and a technology ecosystem that produces so many incredible products?
  8. 1:19:06There's a story that in 2013, TSMC's founder Morris Chang offered you the chance to become TSMC's CEO, and you said you already had a job. Is that true?
  9. 1:26:44What do you think about the space angle?
  10. 1:40:59What gives you strength given how many nations and peoples depend on you?
  11. 1:55:40What do you think of this drama?
  12. 2:03:31I can just launch an agent and make a lot of money?
Lex Freedman 0:00 ↗
The following is a conversation with Jensen Huang, CEO of Nvidia. After a brief mention of sponsors, I introduce the topic: You've propelled Nvidia into a new era, moving beyond chip scale design to rack scale design. Winning used to be about building the best GPU, but now you've expanded to extreme co-design of GPU, CPU, memory, networking, storage, power, cooling, software, the rack, the pod, and the data center. So let's talk about extreme co-design. What is the hardest part of co-designing a system with that many components?
Jensen Huang 7:12 ↗
Yeah, thanks for that question. The reason extreme co-design is necessary is because the problem no longer fits inside one computer. You want to go faster than the number of computers you add. So you have to break up the algorithm, refactor it, shard the pipeline, data, and model. When you distribute the problem, everything gets in the way due to Amdahl's law. Not only do you have to distribute computation, but also solve networking, switching, and distributed computing at scale. We have to bring every technology to bear, otherwise we scale up linearly or based on Moore's law, which has slowed.
Lex Freedman 9:35 ↗
How do you get them in a room together to figure out?
Jensen Huang 9:36 ↗
That's why my staff is so large?
Lex Freedman 9:39 ↗
What's the pro? Can you take me through the process of the specialists and the generalists? Like how do you put together the rack? What does that process look like of designing it all together?
Jensen Huang 9:53 ↗
The first question is what is extreme co-design? You're optimizing across the entire stack from architectures to chips to system software to algorithms to applications. The second is why: you want to distribute the workload to exceed the benefit of increasing the number of computers. The third is how: it's the miracle of this company. When designing a computer, you have to have an operating system of computers. When designing a company, you should think about what you want it to produce. My direct staff is 60 people. I don't do one-on-ones because it's impossible. We present a problem and all of us attack it. The company is doing extreme co-design all the time.
Lex Freedman 13:08 ↗
So, as you mentioned, Nvidia is this company that's adapting to the environment. At which point did the environment change and you began adapting from GPU for gaming to the deep learning revolution to now thinking of it as an AI factory?
Jensen Huang 13:33 ↗
I can reason through it systematically. We started as an accelerator company, but accelerators have a narrow application domain. The problem is market size dictates R&D capacity, which dictates influence. We had to find a way to become accelerated computing, struggling with the tension between specialization and generality. The first step was the programmable pixel shader, then FP32, which led to Cg and eventually CUDA. Putting CUDA on GeForce was a strategic decision that cost the company enormous profits, but it was necessary to build an installed base. That decision was existential: Nvidia's market cap went down to $1.5 billion, but we believed in it. GeForce took CUDA to everyone, and it became the foundation for deep learning.
Lex Freedman 16:44 ↗
Can you take me to that decision? Putting CUDA on GeForce, could not afford to do. Why boldly choose to do that anyway?
Jensen Huang 16:55 ↗
It was the first strategic decision that was an existential threat. The importance of install base is everything. GeForce was successful, selling millions of GPUs a year. We decided to put CUDA on GeForce and put it into every PC, even if customers didn't use it, to cultivate an installed base. We went to universities, wrote books, taught classes. It increased the cost of the GPU so much that it consumed all the company's gross profit dollars. Our market cap dropped, but we carried CUDA on GeForce. Nvidia is the house that GeForce built.
Lex Freedman 22:50 ↗
So Nvidia continues to make bold bets that predict and define the future. How are you able to make those decisions?
Jensen Huang 23:16 ↗
First, I'm informed by curiosity. I reason about the future and become convinced. I believe it in my mind and manifest that future. But I also shape the belief system of everyone around me step by step. I don't make sudden announcements. I lay down bricks over time so that when I announce something, everyone is already bought in. For example, I talk about initiatives for two and a half years before announcing. Leadership sometimes looks like leading from behind. By the time I declare something, it's obvious.
Lex Freedman 28:42 ↗
Are you still a believer in scaling laws? You've outlined four: pre-training, post-training, test time, and agentic scaling. What are the blockers?
Jensen Huang 28:51 ↗
We have more scaling laws now. Pre-training: we can use synthetic data. Post-training: inference is thinking, which is computationally intensive. Test time scaling: it's about reasoning, planning, search. Agentic scaling: agents spawn sub-agents, multiplying AI. The cycle feeds back. The ultimate blocker is compute, but we push extreme co-design to improve tokens per second per watt a million times over the last 10 years. Power is a concern, but we also improve energy efficiency. Supply chain is critical; I inform all CEOs of the dynamics so they can invest. For example, I convinced DRAM CEOs to invest in HBM memory, which became mainstream.
Total footprint of whatever data center you're going to build, let's say you would like to have 50 gigawatts of supercomputers running simultaneously, and it takes one week to manufacture that 50 gigawatts of supercomputers. Then each week in the supply chain, the supercomputers are going to need a gigawatt of power. So we're going to need the supply chain to increase the amount of power it has to build and test the supercomputers in the supply chain before I ship it.
Well, MVLink 72 literally builds supercomputers in the supply chain and ships them two or three tons at a time per rack. It used to be they came in parts and we used to assemble them inside the data center, but that's impossible now because MVLink72 is so dense. That's an example. I would have to go into the supply chain, meet my partners, and say, 'Guess what? Here's what we're going to do. This is the way we used to build our DGXs. We're going to build them this way. This is going to be so much better because we're going to need them for inference. The market for inference is coming, the inflection point is coming, it's going to be a big market.' So I first explain to them what's going on, why it's going to happen, and then I ask them to make several billion dollars of capital investments each. Because they trust me, I'm very respectful of them. I give them every opportunity to question me, I spend time to explain things, draw pictures, reason about it from first principles, and by the time I'm done, they know what to do.
Lex Freedman 52:37 ↗
So it's a lot about relationships and building a shared view of the future.
Jensen Huang 52:42 ↗
Yeah.
Lex Freedman 52:43 ↗
But do you worry about certain bottlenecks? I mean, what are the biggest bottlenecks in the supply chain? Are you worried about ASML EUV tooling? About the packaging, co-packaging of TSMC? About how fast it can scale? You're not only growing incredibly fast, you're accelerating your growth. It feels like everybody in the supply chain would have to scale up. Are you having conversations with them about how they can scale up faster? Do you worry about it?
Jensen Huang 53:15 ↗
No.
Lex Freedman 53:15 ↗
Okay.
Jensen Huang 53:16 ↗
Because I told them what I needed, they understood what I need. They told me what they're going to do, and I believe in what they're going to do.
Lex Freedman 53:23 ↗
Interesting. That's great to hear. So maybe if we can just linger on the power for a little bit. What are your hopes for how to solve the energy problem? One of the areas I'd love to talk about and just get the message out. Our power grid is designed for the worst case condition with some margin.
Jensen Huang 53:50 ↗
Well, 99% of the time we're nowhere near the worst case condition because the worst case is a few days in the winter, a few days in the summer, extreme weather. Most of the time we're nowhere near the worst case, probably running around 60% of peak. So 99% of the time our power grid has excess power sitting idle. But it has to be there just in case hospitals, infrastructure, airports need to be powered. So the question is whether we can go and help them understand and create contractual agreements, design computer architecture systems and data centers such that when they need maximum power for infrastructure in society, the data centers would get less.
But that's a very rare instance anyway. During that time, we either have backup generators for that little part, or we just have our computers shift the workload somewhere else, or run slower. We could degrade our performance, reduce power consumption, provide slightly longer latency response when somebody asks for an answer. I think that way of using computers, building data centers instead of expecting 100% uptime with very rigorous contracts, puts a lot of pressure on the grid. They'll have to increase from their maximum, but I just want to use their excess.
Lex Freedman 55:36 ↗
It's just sitting there. Yeah. So what's stopping that? Is it regulation? Is it bureaucracy?
Jensen Huang 55:45 ↗
I think it's a throughway problem. It starts with the end customer. The end customer puts requirements on the data centers that they can never not be available. They expect perfection. To deliver that perfection, you need a combination of backup generators and grid power to deliver on perfection. So everybody has to have six nines.
Lex Freedman 56:15 ↗
Well, I think first of all, we ought to have everybody understand that when the customer asks for these things, the data center operations team is disconnected from the CEO. I bet the CEO doesn't know this. I'm going to talk to all the CEOs. They're probably not paying attention to the contracts being signed. Everybody wants to sign the best contract, so they go to cloud service providers. I can see the contract negotiators now, negotiating multi-year contracts, both sides wanting the best contract. As a result, the CSPs then go to utilities and expect the six nines. So I think the first thing is to make sure all the customers, the CEOs, realize what they're asking for. The second thing is we have to build data centers that gracefully degrade.
Jensen Huang 57:23 ↗
Mhm.
Lex Freedman 57:23 ↗
We're just going to move our workload around. We'll make sure data is never lost, but we can reduce the computing rate and use less energy. The quality of service degrades a little for critical workloads, I shift that somewhere else right away. So whichever data center still has 100% uptime. How difficult of an engineering problem is smart dynamic allocation of power in the data center?
Jensen Huang 57:50 ↗
As soon as you can specify it, you can engineer it beautifully, so long as it obeys the laws of physics. On first principles, I think we're good.
Lex Freedman 58:00 ↗
What was the third thing you were mentioning? So the second thing is the data centers...
Jensen Huang 58:05 ↗
The third thing is we need the utilities to also recognize that this is an opportunity. Instead of saying it will take five years to increase grid capability, if you're willing to take power at this level of guarantee, I can make it available next month at this price. If utilities offered more segments of power delivery promises, then everybody will figure out what to do. There's way too much waste in the grid right now. We should go after it.
And instead of saying it will take me five years to increase my grid capability, if you're willing to take power at this level of guarantee, I can make it available next month at this price. So if utilities also offered more segments of power delivery promises, then I think everybody will figure out what to do. Yeah, but there's just way too much waste in the grid right now. We should go after it.
Lex Freedman 58:45 ↗
You've highly lauded Elon and xAI's accomplishment in Memphis in building Colossus supercomputer in record time, just four months. It's now at 200,000 GPUs and growing quickly. Is there something about his approach that's instructive to all data center creators? His approach to engineering, management of construction, everything.
Jensen Huang 59:17 ↗
First of all, Elon is deep in so many different topics, yet he's also a really good systems thinker. He's able to think through multiple disciplines. He obviously pushes things, questions everything: Is it necessary? Does it have to be done this way? Does it have to take this long? He has the ability to question everything down to the minimal amount that's necessary. You can't take anything else out, yet the necessary capabilities of the product retain. He's as minimalist as you could imagine at system scale. I also love that he is present at the point of action. He'll just go there and if there's a problem, he'll show me the problem. When you do all of that in combination, you overcome a lot of 'this is just the way we do it.' Everybody has a lot of excuses. The last thing is when you act personally with so much urgency, it causes everybody else to act with urgency. Every supplier has many customers and projects. He makes it his business that he's the top priority of everyone else's projects, and he does that by demonstrating it.
Lex Freedman 1:00:40 ↗
Yeah, I've been in a bunch of those meetings. It's fun to watch. Not enough people ask the question, 'Can this be done a lot faster? Why does it have to take this long?' And then that becomes an engineering question. When you get the ground truth, I remember hanging out with him, he was going through the entire process of how to plug in cables into a rack, working with an engineer on the ground. He's trying to understand what that process looks like so it can be less error-prone. Building up that intuition from every single task involved in putting together a data center, you immediately get a sense at the detailed scale and at the broad system scale of where the inefficiencies are. So you can make it more efficient, plus you have the big hammer of being able to say, 'Let's do it totally different.' And remove all possible blockers.
Jensen Huang 1:01:11 ↗
Yeah.
Lex Freedman 1:01:12 ↗
Are there parallels in the NVIDIA extreme systems co-design approach that you see in the way Elon approaches systems engineering?
Jensen Huang 1:01:20 ↗
Well, first of all, co-design is the ultimate systems engineering problem. We approach the work from that first principle. Another thing we do, a philosophy, a state of mind I started 30 years ago, is called the speed of light. Speed of light is shorthand for: what's the limit of what physics can do? Everything we do is compared against the speed of light: memory speed, math speed, power, cost, time, effort, number of people, manufacturing cycle time. When you think about latency versus throughput, cost versus throughput, cost versus capacity, you test against the speed of light for each constraint separately. Then when you consider them together, you have to make compromises because a system that achieves extremely low latency versus one that achieves very high throughput are architected fundamentally differently. You want to know the speed of light for a high-throughput system and for a low-latency system. Then with the total system, you can make trade-offs. I force everybody to think about first principles, the physical limits, before we do anything. We test everything against that. That's a good frame of mind. I don't love continuous improvement. You should engineer something from first principles at the speed of light, limited only by physics. After that, you improve it over time. But I don't like someone saying, 'It takes 74 days today, and we can do it in 72.' I'd rather strip it back to zero. First, explain why it's 74 days. Then think about what's possible today. If I build it from scratch, how long would it take? Often it might be six days. The rest of the gap could be well reasoned compromises and cost reductions. But at least you know what they are. Once you know six days is possible, the conversation from 74 to 6 is much more effective.
Lex Freedman 1:04:14 ↗
In such incredibly complex systems, is simplicity sometimes a good heuristic to reach for? I mean, the Vera Rubin pod you announced is incredible. Seven chips, seven chip types, five purpose-built rack types, 40 racks, 1.2 quadrillion transistors, nearly 20,000 NVIDIA dies, over 1,100 Rubin GPUs, 60 exaflops, 10 petabytes per second of scale bandwidth. That's just one pod.
Jensen Huang 1:04:32 ↗
That's just one pod.
Lex Freedman 1:04:34 ↗
That's just one pod. And the NVL72 rack alone is 1.3 million components, 1,300 chips, 4,000 pounds crammed into a single 19-inch wide rack. You'll probably crank out about 200 of these pods a week. The amount of different components is staggering. I suppose simplicity is impossible, but is that a metric you reach for in designing things?
Jensen Huang 1:04:39 ↗
The phrase I use most often is: we need things to be as complex as necessary but as simple as possible. The question is whether all that complexity is necessary. We ought to test that and challenge it. Everything else above that is gratuitous. But this is some of the most incredible engineering in history. These systems are truly marvels of engineering. It is the most complex computer the world has ever made. The engineering teams—I don't mean to make it a competition, but if it were an Olympics of engineering teams, TSMC and ASML do incredible work, but NVIDIA is giving them a run for their money. Incredible teams, gold medalists in every sport, all assembled right here.
Lex Freedman 1:05:59 ↗
And they have to work together and report directly to you. This is wonderful. You've recently traveled to China. So it's interesting to ask: China's been incredibly successful in building up its technology sector. What do you understand about how China is able to build so many incredible world-class companies, engineering teams, and a technology ecosystem that produces so many incredible products?
Jensen Huang 1:06:22 ↗
First, some facts: 50% of the world's AI researchers are Chinese plus or minus, mostly still in China. Their tech industry showed up at precisely the right time, during the mobile cloud era. Their way of contributing with software—this is a country with incredible science and math, well-educated kids. Their tech industry was created during the era of software, so they're very comfortable with modern software. China is not one giant economy; it has many provinces and cities with mayors all competing with each other. That's why there are so many EV companies, AI companies, every company you can imagine. They create some of them, and as a result, there's insane competition internally. What remains is an incredible company. They also have a social culture where it's family first, friends second, company third. The amount of conversation back and forth is essentially open source all the time. They contribute more to open source because they think, 'What are we protecting? My engineers' brothers are in that company, their friends are in that company, they're all schoolmates.' The schoolmate concept is like a brother for life. They share knowledge very quickly, so there's no sense keeping technology hidden. The open source community then amplifies and accelerates innovation. So you get rapid innovation from great talent, open source, and insane competition among companies. What emerges is incredible stuff. China is the fastest innovating country in the world today. Everything I've said is fundamental to how kids are raised: excellent education, parents wanting them to do well, a culture that values education. They showed up at the right time when technology is going exponential. Plus, culturally, it's pretty cool to be an engineer. It's a builder nation. Our country's leaders are mostly lawyers, while their leaders are mostly engineers who built the country out of poverty.
Lex Freedman 1:11:24 ↗
To take a small tangent, since you mentioned open source, I have to go to Perplexity, which you have been a fan of for a long time. I love it. Thank you for releasing open source Nemotron 3 Super, which you can also use inside Perplexity. It's a 120 billion parameter open weight model. What's your vision with open source? You mentioned China with DeepSeek, MiniMax, all these companies pushing the open source AI movement, and NVIDIA is leading the way with near state-of-the-art open source LLMs. What's your vision?
Jensen Huang 1:11:52 ↗
First, if we're going to be a great AI computing company, we have to understand how AI models are evolving. One thing I love about Nemotron 3 is it's not a pure transformer; it's transformer and SSM. We were early in developing conditional GANs and progressive GANs, which led to diffusion. Doing basic research in model architecture across different domains gives us visibility into what kind of computing systems will be needed for future models. That's part of our extreme co-design strategy. Second, we recognize that on one hand we want world-class models as proprietary products. On the other hand, we want AI to diffuse into every industry, country, researcher, and student. If everything is proprietary, it's hard to do research and innovate. So open source is fundamentally necessary for many industries to join the AI revolution. NVIDIA has the scale, skills, and motivation to build these models for as long as we live, so we ought to do that. We can open up and activate every industry, researcher, and country. A third reason is that AI is not just language. These AIs will use tools, models, and sub-agents trained on other modalities like biology, chemistry, physics, fluids, thermodynamics. Someone has to ensure that weather prediction, AI for biology, physical AI, etc., get pushed to the frontier. We don't build cars, but we want every car company to have access to great models. We don't discover drugs, but we want Lilly to have the world's best biology AI systems. So those three reasons—recognizing AI is broad, wanting to engage everyone, and co-design—drive our open source strategy.
Lex Freedman 1:15:34 ↗
Well, I have to say once again, thank you for truly open sourcing Nemotron 3. You open source the models, the weights, the data, and how you created it. It's pretty amazing. It's really incredible.
Jensen Huang 1:15:43 ↗
Yeah.

47 more exchanges in this transcript

Sign in free to read the rest of this interview. No card required.

Sign in to read the full transcript

Cite this transcript

APA, MLA, BibTeX
APA

Huang, J. (2026, July 8). #494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution [Interview transcript]. MindVoice Production. CEOInterviews.AI. https://ceointerviews.ai/interview/1070909/

MLA

Jensen Huang. "#494 – Jensen Huang: NVIDIA – The $4 Trillion Company & the AI Revolution." MindVoice Production, 8 Jul. 2026. Transcript, CEOInterviews.AI, https://ceointerviews.ai/interview/1070909/.

BibTeX
@misc{huang2026_1070909,
  author       = {Jensen Huang},
  title        = {#494 – Jensen Huang: NVIDIA – The $4 Trillion Company \& the AI Revolution},
  howpublished = {Interview transcript, MindVoice Production. CEOInterviews.AI},
  year         = {2026},
  month        = {jul},
  url          = {https://ceointerviews.ai/interview/1070909/},
  note         = {Speaker-attributed transcript with timestamps}
}