NUS120 Distinguished Speaker Series | Dr William Dally
ABOUT NUS120* The National University of Singapore celebrates its 120th anniversary in 2025, commemorating a legacy, forged ...
Chief Scientist & Senior Vice President of Research, Nvidia
Search every verified William Dally interview, podcast appearance, and on-the-record quote — each transcript cross-checked by AI and human review to confirm speaker identity. William Dally, Chief Scientist and Senior Vice President of Research at Nvidia, gave a keynote at GTC Taipei in May 2026 where he discussed the economics of AI factories, stating that tokens "are now profitable units of revenues" and that compute demand in Taiwan has "skyrocketed" as a result. He estimated the cost of a single gigawatt-level AI factory has risen from $30 billion to between $60 and $100 billion, and argued that "compute is revenues" and "performance per watt is your revenues," cautioning against choosing architecture solely on chip cost. Dally also said the number of software engineers is increasing, describing claims that AI reduces jobs as "complete nonsense," citing the productivity gain of $3 trillion worth of software engineer salary generating $9 trillion in output. In a June 2026 lecture at the National University of Singapore, Dally attributed the deep learning revolution to GPU hardware enabling algorithms and data that had existed for decades, and stated that progress remains "gated by how fast a GPU is." He contrasted Nvidia's product development, where "it has to work or we're going out of business," with Nvidia Research, where the ability to fail allows for innovations that can achieve "2x or 4x performance per unit energy on the next generation." He also referenced an earlier 2020 talk where he described a 317x increase in single-chip inference performance over eight years, a trend he termed "Huang's Law," and credited specialized tensor core instructions for allowing GPUs to achieve efficiency near that of dedicated hardware.
“AI really was created by three ingredients, the modern revolution in deep learning. Algorithms as sort of indicated by this picture from the AlexNet paper in 2012. If you look at AlexNet it's a comnet and it was trained with gradient descent and back propagation and all of those algorithms were around in the 1980s. So...”
“The single biggest thing that made a difference during this period of time was number representation. In Kepler, we had 32-bit floating point numbers for the bulk of the math. That's because the chip wasn't designed to do deep learning. It was designed to do graphics and high performance computing. And for graphics, we...”
“We want to be general but we can't be completely general but we have to necessarily specialize for example by deciding on certain ratios when we design a GPU we have to decide How much external memory bandwidth do we want to have? How much internal memory bandwidth do we want to have? How much arithmetic bandwidth do w...”
“I think deep learning is really affecting every aspect of the human experience and it's been enabled by hardware and its progress is gated by hardware. Remember the algorithms and the data were around. It wasn't until the GPUs came around to sort of be the spark that ignited that fuel, our mixture that lit off our curr...”
“I think academics have a huge advantage in that one thing I say about Nvidia research is we have a huge advantage over our product development people is that we can afford to fail. If you're working on product development, Nvidia, Intel used to start five CPU designs and one of them would actually ship so everybody goi...”
“30 million software developers representing about $3 trillion worth of GDP producing three, that's what they're paid. $3 trillion worth of salaries per year, which is generating economic growth for the rest of the industries. Say a hundred trillion dollars of the world's industries is impacted is generated by $3 billio...”
“The number of engineers, software engineers is actually increasing. People talk about AI reducing jobs. Complete nonsense. It's causing more software engineers to be hired. And the reason for that is very simple. If you can hire a software engineer and you could generate $9 trillion worth of productive work, why wouldn...”
“Tokens are now in extraordinary demand. Because if you could do this, you're going to want to produce more of it. And because tokens are now profitable units, tokens are now profitable units of revenues. because it is now profitable. The AI companies want to build a lot more tokens, generate a lot more tokens, build mo...”
“Each one of these at one gigawatt level started at 30 2030 billion dollars. It is at 5060 billion dollars and soon it will be 80 hundred billion dollars per gigawatt $100 billion into an AI factory. It must work the first time and it must work right away. The cost of capital is incredible.”
“Compute is revenues. Performance per watt is your revenues. Choosing the wrong architecture just because the chips are cheaper doesn't translate doesn't make sense. You need to make sure that your revenues per watt the more you buy the more you make.”
“If you know serial computers, single thread processors aren't getting any faster and you know if you want to add value to applications you need more performance then you need to have lots of threads in parallel and and to do that efficiently you really need to specialize for for a certain domain. And the gains to be ha...”
“When we started our Darwin project... if we took the existing best algorithms and through everything we knew about hardware to accelerate them we might get 4x... but we weren't very happy with 4x so we asked ourselves you know why is it limited to that... doing that turned it from being 4x better to being 15,000 times...”
“I think it's a really good one. Um so so as a leader of industrial research lab, you know, I I do one of my primary functions is to make sure that things that get done in in NVIDIA research result in improvements to NVIDIA products and not just publications at leading conferences because I think I'd have a real hard ti...”
“I'm I'm very worried about it and I guess I'm part of the problem. But um you know pe people react to incentives and and the incentives are all set up um you know even even you know both the selfish incentives and the selfless incentives are all set up to suck people into the industrial world. I mean the salaries are a...”
“Starting with a voltage generation and continuing with Turing and Ampere, we introduced specialized tensor core instructions for matrix multiply accumulate. By amortizing the instruction overhead across massive operations, our programmable GPUs operate with a negligible efficiency penalty compared to dedicated hardware...”
ABOUT NUS120* The National University of Singapore celebrates its 120th anniversary in 2025, commemorating a legacy, forged ...
Dr. Bill Dally is the Chief Scientist and Senior Vice President of Research at Nvidia, and a Professor of Computer Science at Stanford University. Dr. Dally has had a storied career with contributions to parallel computer architectures, interconnection networks, GPUs, accelerators and more. He has a history of designing innovative and experimental computing systems such as the MARS accelerator, the MOSSIM simulation engine, the J-Machine and M-machine, to name a few. He talks to us about computing innovation in the post-Moore era, domain-specific accelerators, and technology transfer in comput…
In this GTC China keynote, NVIDIA’s Chief Scientist Bill Dally reveals how accelerated hardware and software are solving the world’s most demanding computing problems. Discover the power of the Ampere A100—the world’s largest 7nm chip with 54 billion transistors—and its ability to exploit structured sparsity for a massive jump in deep learning performance. Key topics covered in this presentation: Huang’s Law: Learn about the staggering 317x increase in single-chip inference performance achieved over the last eight years. Energy Efficiency: Why the Selene supercomputer ranked #1 on the Green 50…
In this 60-minute wide-ranging discussion, NVIDIA Chief Scientist and GPU architect Bill Dally engages in a focused dialogue with ...
Biography: Dr. Bill Dally is Chief Scientist and Senior Vice President of Research at NVIDIA Corporation and an Adjunct Professor and former chair of Computer Science at Stanford University. Bill is currently working on developing hardware and software to accelerate demanding applications including machine learning, bioinformatics, and logical inference. He has a history of designing innovative and efficient experimental computing systems. While at Bell Labs Bill contributed to the BELLMAC32 microprocessor and designed the MARS hardware accelerator. At Caltech he designed the MOSSIM Simulation…
The Lectio Magistralis is the flagship event of the annual programme of the CSP-IAS – Institute for Advanced Study, the joint initiative promoted by the Italian Institute of Artificial Intelligence (AI4I) and the Compagnia di San Paolo Foundation to foster international dialogue among science, industry, and culture on artificial intelligence. Each year, the Lectio brings to Turin a leading global figure in research and innovation, offering a space for reflection on the questions and frontiers that define the evolution of AI.
... shows the evolution of of Nvidia hardware for running deep learning over um basically the last 12 years um and Blackwell is up ...
In this talk, Bill Dally, NVIDIA Chief Scientist and Senior Vice President of Research, discusses NVIDIA’s recent progress on deep learning hardware and reviews the origin of this technology with DARPA and DoE funded projects at MIT and Stanford. This supercut highlights the government-funded academic research that was critical both in developing key technologies and training the people needed to take these technologies to product.
По мере того как искусственный интеллект продолжает менять мир, пересечение глубокого обучения и ...
Sign in to search the full transcript archive, filter by topic, and access every quote from William Dally.
The summary and quote tags on this profile are produced with AI assistance from verified, first-person interview transcripts, then checked by our team to confirm the speaker's identity and the accuracy of every quote. See how we verify →