Back
Andrew Ng
Co-Founder & Chairman, Coursera

#73 – Andrew Ng: Deep Learning, Education, and Real-World AI

🎥 Jun 26, 2026 📺 MindVoice Production ⏱ 89m
Andrew Ng is one of the most impactful educators, researchers, innovators, and leaders in artificial intelligence and technology ...
Watch on YouTube

About Andrew Ng

Andrew Ng has recently been active in public discussions and product launches related to AI education and software development. He released courses aimed at teaching non-technical users how to build applications by describing ideas in words, and offered guidance on effective prompting for modern AI tools. At the AI Dev 26 x San Francisco conference in May 2026, he announced two new tools: Context Hub, designed to provide AI agents with up-to-date documentation to prevent the use of outdated APIs, and Code Dream, an interactive learning environment featuring AI-driven conversations and a browser-based terminal. He also participated in a fireside chat at Interrupt 26, where he discussed the rapid evolution of coding agents and the shifting bottlenecks in software development as AI accelerates coding speed. In interviews and talks, Ng offered perspectives on current trends in AI, stating that while the hype around the technology exceeded his expectations, he did not believe that a job apocalypse from AI is imminent. He argued that engineering roles would remain valuable and that as coding becomes easier, more people should learn to do it, not fewer. He also discussed the importance of rethinking entire workflows rather than making incremental efficiency gains, and highlighted the growing need for organizations to organize unstructured data for use by AI agents.

Source: AI-verified profile updated from Andrew Ng's recent appearances. Browse all interviews →

Transcript (150 segments)
L
Lex Fridman0:00
The following is a conversation with Andrew Ng, one of the most impactful educators, researchers, innovators, and leaders in artificial intelligence and the technology space in general. He co-founded Coursera and Google Brain, launched deeplearning.ai, Landing AI, and the AI Fund, and was the chief scientist at Baidu. As a Stanford professor and with Coursera and deeplearning.ai, he has helped educate and inspire millions of students, including me. This is the Artificial Intelligence Podcast. If you enjoy it, subscribe on YouTube, give it five stars on Apple Podcast, support it on Patreon, or simply connect with me on Twitter at Lex Fridman, spelled F-R-I-D-M-A-N. As usual, I'll do one or two minutes of ads now and never any ads in the middle that can break the flow of the conversation. I hope that works for you and doesn't hurt the listening experience. This show is presented by Cash App, the number one finance app in the App Store. When you get it, use code Lex Podcast. Cash App lets you send money to friends, buy Bitcoin, and invest in the stock market with as little as one dollar. Broker services are provided by Cash App Investing, a subsidiary of Square, a member SIPC. Since Cash App allows you to buy Bitcoin, let me mention that cryptocurrency in the context of the history of money is fascinating. I recommend The Ascent of Money as a great book on this history. Debits and credits on ledgers started over thirty thousand years ago. The US dollar was created over two hundred years ago and Bitcoin, the first decentralized cryptocurrency, released just over ten years ago. So given that history, cryptocurrency is still very much in its early days of development, but is still aiming to and just might redefine the nature of money. So again, if you get Cash App from the App Store or Google Play and use the code Lex Podcast, you'll get ten dollars. And Cash App will also donate ten dollars to FIRST, one of my favorite organizations that is helping to advance robotics and STEM education for young people around the world. And now here's my conversation with Andrew Ng.
The courses you taught on machine learning at Stanford and later on Coursera, that you co-founded, have educated and inspired millions of people. So let me ask you, what people or ideas inspired you to get into computer science and machine learning when you were young? When did you first fall in love with the field is another way to put it?
A
Andrew Ng3:01
Growing up in Hong Kong and Singapore, I started learning to code when I was five or six years old. At that time, I was learning the BASIC programming language and they would take these books and they'd tell you type this program into your computer. I'd type that program on my computer and as a result of all that typing, I would get to play these very simple shoot-em-up games that I had implemented on my little computer. So I thought it was fascinating as a young kid that I could write this code, which was really just copying code from a book into my computer, to then play these cool little video games. Another moment for me was when I was a teenager and my father, who's a doctor, was reading about expert systems and about neural networks. So he got me to read some of these books and I thought it was really cool you could write a computer that started to exhibit intelligence. Then I remember doing an internship while I was in high school in Singapore where I remember doing a lot of photocopying as an office assistant. And the highlight of my job was when I got to use the shredder. So the teenager in me remember thinking, boy, this is a lot of photocopying. If only we could write software, build a robot, something to automate this, maybe I could do something else. So I think a lot of my work since then has centered on the theme of automation. Even the way I think about machine learning today, we're very good at writing learning algorithms that can automate things that people can do. Or even launching the first MOOCs, massive open online courses, that later led to Coursera, I was trying to automate what could be automated in how I was teaching on campus.
L
Lex Fridman4:42
The process of education, tried to automate parts of that to make it more, to have more impact from a single teacher, single educator.
A
Andrew Ng4:52
Yeah. I felt, teaching at Stanford, I was teaching machine learning to about four hundred students a year at the time. And I found myself filming the exact same video every year, like telling the same jokes in the same room. And I thought, why am I doing this? Why don't we just take last year's video and then I can spend my time building a deeper relationship with students. So, that process of thinking through how to do that, that led to the first MOOCs that we launched.
L
Lex Fridman5:19
And then you have more time to write new jokes. Are there favorite memories from your early days at Stanford teaching thousands of people in person and then millions of people online?
A
Andrew Ng5:31
You know, teaching online, what not many people know was that a lot of those videos were shot between the hours of 10:00 p.m. and 3:00 a.m. A lot of times, we were launching the first Stanford course. We had already announced the course, about 100,000 people signed up. We just started to write the code and we had not yet actually filmed the video. So we had a lot of pressure, 100,000 people waiting for us to produce the content. So many Fridays, Saturdays, I would go out, have dinner with my friends, and then I would think, okay, do you want to go home now or do you want to go to the office to film videos? And the thought of being able to help 100,000 people potentially learn machine learning, fortunately, that made me think, okay, I want to go to my office, go to my tiny little recording studio. I would adjust my Logitech webcam, adjust my Wacom tablet, make sure my lapel mic was on, and then I would start recording often until 2:00 a.m. or 3:00 a.m. I think it's unfortunate that it doesn't show that it was recorded that late at night, but it was really inspiring, the thought that we could create content to help so many people learn about machine learning.
L
Lex Fridman6:45
How did that feel? The fact that you're probably somewhat alone, maybe a couple of friends, recording with a Logitech webcam and kind of going home alone at one or 2 a.m. at night and knowing that that's going to reach sort of thousands of people, eventually millions of people. What's that feeling like? I mean, is there a feeling of just satisfaction of pushing through?
A
Andrew Ng7:10
I think it's humbling, and I wasn't thinking about what I was feeling. I think one thing that I'm proud to say we got right from the early days was I told my whole team back then that the number one priority is to do what's best for learners, do what's best for students. And so when I went to the recording studio, the only thing on my mind was what can I say? How can I design my slides? What do I need to draw to make these concepts as clear as possible for learners. I think I've seen sometimes instructors, it's tempting to, hey, let's talk about my work, maybe if I teach you about my research someone will cite my papers a couple more times. And I think one of the things we got right launching the first few MOOCs and later building Coursera was putting in place that bedrock principle of let's just do what's best for learners and forget about everything else. And I think that as a guiding principle turned out to be really important to the rise of the MOOC movement.
L
Lex Fridman8:04
And the kind of learner you imagined in your mind is as broad as possible, as global as possible. So really try to reach as many people interested in machine learning and AI as possible.
A
Andrew Ng8:17
I really want to help anyone that had an interest in machine learning to break into the field. And I think sometimes I've actually had people ask me, hey, why are you spending so much time explaining gradient descent? And my answer was, if I look at what I think the learner needs and would benefit from, I felt that having a good understanding of the foundations, kind of back to the basics, would put them in a better state to then build on a long-term career. So I tried to consistently make decisions on that principle.
L
Lex Fridman8:48
So one of the things it actually revealed to the narrow AI community at the time, and to the world, is that the amount of people who are actually interested in AI is much larger than we imagined. By you teaching the class and how popular it became, it showed that wow, this isn't just a small community of people who go to NeurIPS, and it's much bigger. It's developers. It's people from all over the world. I mean, I'm Russian, so everybody in Russia is really interested. There's a huge number of programmers who are interested in machine learning. India, China, South America, everywhere, there's just millions of people who are interested in machine learning. So how big do you get a sense that the number of people is that are interested from your perspective?
A
Andrew Ng9:37
I think the number has grown over time. I think it's one of those things that maybe it feels like it came out of nowhere, but as an insider building it, it took years. It's one of those overnight successes that took years to get there. My first foray into this type of online education was when we were filming my Stanford class and sticking the videos on YouTube. We had uploaded a whole lot and so on. But basically the one hour fifteen minute video that we put on YouTube, and then we had four or five other versions of websites that I had built, most of which you would never have heard of because they reached small audiences, but that allowed me and my team to iterate, to learn what are the ideas that work and what doesn't. For example, one of the features I was really excited about and really proud of was building this website where multiple people could be logged into the website at the same time. So today if you go to a website, if you're logged in and then I want to log in, you need to log out if it's the same browser, same computer. But I thought, well, what if two people, say you and me, we're watching a video together in front of a computer. What if a website could have you type your name and password, have me type my name and password, and then now the computer knows both of us are watching together and it gives both of us credit for anything we do as a group. We piloted this feature in a high school in San Francisco, Sacred Heart Cathedral Prep. The teacher is great. And guess what? Zero people used this feature. It turns out people studying online, they want to watch the videos by themselves. You can play back, pause at your own speed, rather than in groups. So that was one example of a tiny lesson learned out of many that allowed us to hone in to the set of features.
L
Lex Fridman11:22
And it sounds like a brilliant feature. So I guess the lesson to take from that is there's something that looks amazing on paper and then nobody uses it, doesn't actually have the impact that you think it might have. And so yeah, I saw that you really went through a lot of different features and a lot of ideas and to arrive at the final, at Coursera, the final kind of powerful thing that showed the world that MOOCs can educate millions.
A
Andrew Ng11:49
And I think with the whole machine learning movement as well, I think it didn't come out of nowhere. Instead what happened was as more people learn about machine learning, they will tell their friends and their friends will see how it's applicable to their work, and then the community kept on growing. And I think we're still growing. I don't know in the future what percentage of all developers will be AI developers. I could easily see it being north of fifty percent. Because so many AI developers broadly construed, not just people doing the machine learning modeling, but the people building infrastructure, data pipelines, all the software surrounding the core machine learning model, maybe is even bigger. I feel like today almost every software engineer has some understanding of the cloud. I think in the future maybe we'll approach nearly one hundred percent of all developers being in some way an AI developer or at least having an appreciation of machine learning. And my hope is that there's this kind of effect that there's people who are not really interested in being a programmer, like biologists, chemists, and physicists, even mechanical engineers, all these disciplines that are now more and more sitting on large data sets.
L
Lex Fridman13:19
And here they didn't think they're interested in programming until they have this data set and they realize there's these set of machine learning tools that allow you to use the data set. So they actually become, they learn to program and they become new programmers. So not just because you've mentioned a larger percentage of developers become machine learning people, it seems like more and more the kinds of people who are becoming developers is also growing significantly.
A
Andrew Ng13:44
Yeah. I think once upon a time only a small part of humanity was literate, you could read and write, and maybe you thought maybe not everyone needs to learn to read and write. You just go listen to a few monks read to you. And maybe that was enough. Or maybe you just need a handful of authors to write the bestsellers and then no one else needs to write. But what we found was that by giving as many people, in some countries almost everyone, basic literacy, it dramatically enhanced human-to-human communications and we can now write for an audience of one, such as if I send you an email or you send me an email. I think in computing we're still in that phase where so few people know how to code that the coders mostly have to code for relatively large audiences. But if everyone or most people became developers at some level, similar to how most people in developed economies are somewhat literate, I would love to see the owners of a mom-and-pop store be able to write a little bit of code to customize the TV display for their special this week. I think it'll enhance human-to-computer communications, which is becoming more and more important today as well.
L
Lex Fridman14:56
So you think it's possible that machine learning becomes kind of similar to literacy, where, like you said, the owners of a mom-and-pop shop, basically everybody in all walks of life would have some degree of programming capability?
A
Andrew Ng15:13
I could see society getting there. There's one other interesting thing. If I go talk to the mom-and-pop store, if I talk to a lot of people in their daily professions, I previously didn't have a good story for why they should learn to code. But what I found with the rise of machine learning and data science is that I think the number of people with a concrete use for data science in their daily lives, in their jobs, may be even larger than the number of people with a concrete use for software engineering. For example, if you run a small mom-and-pop store, I think if you can analyze the data about your sales, your customers, there's actually real value there, maybe even more than traditional software engineering. So I find that for a lot of my friends in various professions, be it recruiters or accountants or people that work in factories, I feel if they were data scientists at some level, they could immediately use that in their work. So I think that data science and machine learning may be an even easier entree into the developer world for a lot of people than software engineering.
L
Lex Fridman16:20
That's interesting and I agree with that, but that's beautifully put. We live in a world where most courses and talks have slides, PowerPoint, Keynote, and yet you famously often still use a marker and a whiteboard. The simplicity of that is compelling and for me at least fun to watch. Thank you. So let me ask, why do you like using a marker and whiteboard even on the biggest of stages?
A
Andrew Ng16:46
I think it depends on the concepts you want to explain. For mathematical concepts, it's nice, you can build up the equation one piece at a time. And the whiteboard marker or the pen and stylus is a very easy way to build up an equation, build up a complex concept one piece at a time while you're talking about it, and sometimes that enhances understandability. The downside of writing is that it's slow, and so if you want a long sentence it's very hard to write that. So I think there are pros and cons, and sometimes I use slides and sometimes I use a whiteboard or a stylus. The slowness of a whiteboard is also its upside because it forces you to reduce everything to the basics.
L
Lex Fridman17:28
So some of your talks involve the whiteboard, you go very slowly and you really focus on the most simple principles, and that enforces a kind of minimalism of ideas that I think is surprising, at least for me, and is great for education. Like a great talk I think is not one that has a lot of content. A great talk is one that just clearly says a few simple ideas. And I think the whiteboard somehow enforces that. Peter Abbeel, who's now one of the top roboticists and reinforcement learning experts in the world, was your first PhD student. So I bring him up just because I kind of imagine this must have been an interesting time in your life. Do you have any favorite memories of working with Peter? Your first student in those uncertain times, especially before deep learning really blew up. Any favorite memories from those times?
A
Andrew Ng18:35
Yeah. I was really fortunate to have had Peter Abbeel as my first PhD student. And I think even my long-term professional success builds on early foundations or early work that Peter was so critical to. So I was really grateful to him for working with me. What not a lot of people know is just how hard research was and still is. Peter's PhD thesis was using reinforcement learning to fly helicopters. And so, actually even today the website heli.stanford.edu is still up. You can watch videos of us using reinforcement learning to make a helicopter fly upside down, fly loops, roses. It's cool.
L
Lex Fridman19:17
It's one of the most incredible robotics videos ever. So, people should watch it.
A
Andrew Ng19:21
Oh, yeah. Thanks.
L
Lex Fridman19:21
Inspiring. That's from like 2008 or seven or six, like that range.
A
Andrew Ng19:28
Something like that. Over 10 years old.
L
Lex Fridman19:30
That was really inspiring to a lot of people.
A
Andrew Ng19:32
What not many people see is how hard it was. So Peter and Adam Coates and Morgan Quigley and I were working on various versions of the helicopter and a lot of things did not work. For example, turns out one of the hardest problems we had was when the helicopter is flying around upside down doing stunts, how do you figure out the position? How do you localize a helicopter? So we wanted to try all sorts of things. Having one GPS unit doesn't work because you're flying upside down, the GPS unit is facing down so you can't see the satellites. So we experimented trying to have two GPS units, one facing up, one facing down. So if you flip over, that didn't work because the downward facing one couldn't synchronize if you're flipping quickly. Morgan Quigley was exploring this crazy complicated configuration of specialized hardware to interpret GPS signals, look into FPGA, completely insane. Spent about a year working on that. Didn't work. So I remember Peter, great guy, him and me, sitting down in my office looking at some of the latest things we had tried that didn't work and saying, you know, darn it, what now? Because we tried so many things and it just didn't work. In the end, what we did, and Adam Coates was crucial to this, was put cameras on the ground and use cameras on the ground to localize the helicopter, and that solved the localization problem, so that we could then focus on the reinforcement learning and inverse reinforcement learning techniques to then actually make the helicopter fly. And I'm reminded, when I was doing this work at Stanford around that time, there was a lot of reinforcement learning theoretical papers but not a lot of practical applications. So the autonomous helicopter work was one of the few practical applications of reinforcement learning at the time, which caused it to become pretty well known. I feel like we might have almost come full circle where today there's so much buzz, so much hype, so much excitement about reinforcement learning, but again we're hunting for more applications of all of these great ideas that the community's come up with.
L
Lex Fridman21:41
What was the drive, sort of in the face of the fact that most people are doing theoretical work? What motivated you in the uncertainty and the challenges to get the helicopter to do the applied work, to get the actual system to work? In the face of fear, uncertainty, the setbacks that you mentioned for localization.
A
Andrew Ng22:03
I like stuff that works.
L
Lex Fridman22:05
In the physical world. So it's back to the shredder.
A
Andrew Ng22:09
You know, I like theory, but when I work on theory myself, and this is personal taste, I'm not saying anyone else should do what I do, but when I work on theory, I personally enjoy it more if I feel that the work I do will influence people, have positive impact or help someone. I remember when many years ago, I was speaking with a mathematics professor and I just said, hey, why do you do what you do? And he said, he actually had stars in his eyes when he answered. This mathematician, not from Stanford, different university, he said I do what I do because it helps me to discover truth and beauty in the universe.
L
Lex Fridman22:54
He had stars in his eyes when he said that.
A
Andrew Ng22:55
And I thought that's great. I don't want to do that. I think it's great that someone does that, fully support the people that do it. A lot of respect for people for that. But I am more motivated when I can see a line to how the work that my teams and I are doing helps people. The world needs all sorts of people. I'm just one type. I don't think everyone should do things the same way as I do. But when I delve into either theory or practice, if I personally have conviction that there's a pathway to help people, I find that more satisfying to have that conviction.
L
Lex Fridman23:34
Yeah. You were a proponent of deep learning before it gained widespread acceptance. What did you see in this field that gave you confidence? What was your thinking process like in that first decade of the 2000s?
A
Andrew Ng23:51
Yeah, I can tell you the thing we got wrong and the thing we got right. The thing we really got wrong was the early importance of unsupervised learning. So early days of Google Brain, we put a lot of effort into unsupervised learning rather than supervised learning. And there was this argument, actually I think it was around 2005, after NeurIPS at that time called NIPS had ended, and Geoffrey Hinton and I were sitting in the cafeteria outside the conference, we had lunch, we were just chatting. And Geoff pulled up this napkin, he started sketching this argument on a napkin, it was very compelling. I'll repeat it. The human brain has about 100 trillion, so that's 10 to the 14, synaptic connections. You will live for about 10 to the 9 seconds, that's 30 years. So if each synaptic connection, each weight in your brain, has just a one-bit parameter, that's 10 to the 14 bits you need to learn in up to 10 to the 9 seconds of your life. So via this simple argument, which has a lot of problems, it's very simplified, that's 10 to the 5 bits per second you need to learn in your life. And I have a one-year-old daughter. I am not pointing out 10 to the 5 bits per second of labels to her. And I think I'm a very loving parent, but I'm just not going to do that. So from this very crude, definitely problematic argument, there's just no way that most of what we know is through supervised learning. But where you could get so many bits of information is from sucking in images, audio, just experiences in the world. And so that argument really convinced me that there's a lot of power to unsupervised learning. So that was the part that we actually maybe got wrong. I still think unsupervised learning is really important, but in the early days, 10-15 years ago, a lot of us thought that was the path forward.
L
Lex Fridman25:58
Oh, so you're saying that perhaps was the wrong intuition for the time.
A
Andrew Ng26:02
For the time, that was the part we got wrong. The part we got right was the importance of scale. So Adam Coates, another wonderful person, fortunate to have worked with him, he was in my group at Stanford at the time. And Adam had run these experiments at Stanford showing that the bigger we trained a learning algorithm, the better his performance. And it was based on that, there was a graph that Adam generated where lines going up and to the right, bigger, better performance. So it's really based on that chart that Adam generated that he gave me the conviction that we could scale these models way bigger than what we could on CPUs, which is what we had at Stanford, that we could get even better results. And it was really based on that one figure that Adam generated that gave me the conviction to go with Sebastian Thrun to pitch starting a project at Google, which became the Google Brain project.
L
Lex Fridman27:01
Google Brain, you go and found Google Brain, and there the intuition was scale will bring performance for the system. So we should chase larger and larger scale. And I think people don't realize how groundbreaking, it's simple but it's a groundbreaking idea, that bigger data sets will result in better performance.
A
Andrew Ng27:23
It was controversial at the time. Some of my well-meaning friends, senior people in the machine learning community, I won't name, but some of whom we know, came and were trying to give me friendly advice like, hey Andrew, why are you doing this? This is crazy. Look at the neural network architectures, look at these architectures people are building. You just want to go for scale? This is a bad career move. So my well-meaning friends were trying to talk me out of it. But I find that if you want to make a breakthrough, you sometimes have to have conviction and do something before it's popular, since that lets you have a bigger impact.
L
Lex Fridman28:00
Let me ask you just on a small tangent on that topic. I find myself arguing with people saying that greater scale, especially in the context of active learning, so very carefully selecting the data set but growing the scale of the data set, is going to lead to even further breakthroughs in deep learning. And there's currently pushback at that idea, that larger data sets are no longer that valuable, so you want to increase the efficiency of learning, you want to make better learning mechanisms. And I personally believe the just bigger data sets will still, with the same learning methods we have now, will result in better performance. What's your intuition at this time on these dual sides, do we need to come up with better architectures for learning, or can we just get bigger better data sets that will improve performance?
A
Andrew Ng28:55
I think both are important and it's also problem-dependent. So for a few data sets we may be approaching Bayes error rate, or approaching or surpassing human-level performance, and then there's that theoretical ceiling that we will never surpass Bayes error rate. But then I think there are plenty of problems where we're still quite far from either human-level performance or from Bayes error rate, and bigger data sets with neural networks, without further algorithm innovation, will be sufficient to take us further. But on the flip side, if we look at the recent breakthroughs using transformer networks or language models, it was a combination of novel architecture but also scale had a lot to do with it. If we look at what happened with GPT-2 and BERT, I think scale was a large part of the story.
L
Lex Fridman29:44
Yeah, that's not often talked about, is the scale of the data set it was trained on and the quality of the data set, because there's some, so it was like Reddit threads that had been upvoted highly, so there's already some weak supervision on a very large data set that people don't often talk about, right?
A
Andrew Ng30:04
I find that today we have maturing processes for managing code, things like Git, version control. It took us a long time to evolve the good processes. I remember when my friends and I were emailing each other C++ files in email. But then we had CVS, version control, Git, maybe something else in the future. We're very immature in terms of tools for managing data and think of how to clean data and how to solve very messy data problems. I think there's a lot of innovation there to be had still.
L
Lex Fridman30:36
I love the idea that you were versioning through email.
A
Andrew Ng30:39
I'll give you one example. When we work with manufacturing companies, it's not at all uncommon for there to be multiple labelers that disagree with each other. And so doing the work in visual inspection, we would take say a plastic pot and show it to one inspector and the inspector, sometimes very opinionated, they go, clearly that's a defect, this scratch, unacceptable, got to reject this part. Take the same part to a different inspector, different, very opinionated, clearly the scratch is small, it's fine, don't throw it away, you're going to make us lose money. And then sometimes you take the same plastic part, show it to the same inspector in the afternoon, I showed it in the morning, and very confidently in the morning they say clearly it's okay, in the afternoon equally confident, clearly this is a defect. And so what is an AI team supposed to do if sometimes even one person doesn't agree with himself or herself in the span of a day? So I think these are the types of very practical, very messy data problems that my teams wrestle with. In the case of large consumer internet companies where you have a billion users, you have a lot of data, you don't worry about it. Just take the average, it kind of works. But in the case of other industry settings, we don't have big data. It's just small data, very small data sets, maybe 100 defective parts or 100 examples of a defect. If you have only 100 examples, these little labeling errors, if 10 of your 100 labels are wrong, that actually is 10% of your data set, has a big impact. So, how do you clean this up? This is an example of the types of things that my teams, this is a Landing AI example, are wrestling with, to deal with small data, which comes up all the time once you're outside consumer internet.
L
Lex Fridman32:29
Yeah, that's fascinating. And so then you invest more effort and time in thinking about the actual labeling process, what are the labels, how are disagreements resolved, and all those kinds of pragmatic real-world problems. That's a fascinating space.
A
Andrew Ng32:45
I find it actually, when I'm teaching at Stanford, I increasingly encourage students at Stanford to try to find their own project for the end-of-term project, rather than just downloading someone else's nicely cleaned data set. It's actually much harder if you need to go find your own problem and find your own data set rather than go to one of the several good websites with clean scoped data sets that you could just work on.
L
Lex Fridman33:12
You're now running three efforts. The AI Fund, Landing AI, and deeplearning.ai. As you've said, the AI Fund is involved in creating new companies from scratch. Landing AI is involved in helping already established companies do AI, and deeplearning.ai is for education of everyone else, or of individuals interested in getting into the field and excelling in it. So, let's perhaps talk about each of these areas. First, deeplearning.ai. How, the basic question, how does a person interested in deep learning get started in the field?
A
Andrew Ng33:52
deeplearning.ai is working to create courses to help people break into AI. So my machine learning course that I taught through Stanford remains one of the most popular courses on Coursera to this day.
L
Lex Fridman34:07
It's probably one of the courses, sort of if I ask somebody how did you get into machine learning or how did you fall in love with machine learning or what got you interested, it always goes back to Andrew at some point. So you've influenced, the amount of people you've influenced is ridiculous. So for that, I'm sure I speak for a lot of people, a big thank you.
A
Andrew Ng34:24
No. Yeah. Thank you. You know, I was once reading a news article, I think it was Tech Review, and I'm going to mess up the statistic, but I remember reading an article that said something like one-third of all programmers are self-taught. I may have the number wrong, maybe it was two-thirds, but when I read that article, I thought this doesn't make sense. Everyone is self-taught. Because you teach yourself, I don't teach people. I just don't.
L
Lex Fridman34:50
That's well put. So, yeah. So how does one get started in deep learning and where does deeplearning.ai fit into that?
A
Andrew Ng34:58
So the deep learning specialization offered by deeplearning.ai is, I think, it was Coursera's top specialization. It might still be. So it's a very popular way for people to take that specialization to learn about everything from neural networks to how to tune your network, to what is a ConvNet, to what is an RNN or sequence model, or what is an attention model. And so the deep learning specialization steps everyone through those algorithms so you deeply understand it and can implement it and use it for whatever application.
L
Lex Fridman35:32
From the very beginning. So what would you say are the prerequisites for somebody to take the deep learning specialization in terms of maybe math or programming background?
A
Andrew Ng35:43
Yeah, you need to understand basic programming since there are programming exercises in Python. And the math prereq is quite basic. So no calculus is needed. If you know calculus, it's great, you get better intuitions. But we deliberately try to teach that specialization without requiring calculus. So I think high school math would be sufficient. If you know how to multiply two matrices, I think that's great.
L
Lex Fridman36:09
So a little basic linear algebra is great.
A
Andrew Ng36:12
Basic linear algebra, even very, very basic linear algebra, and some programming. I think that people that have done the machine learning course will find the deep learning specialization a bit easier, but it's also possible to jump into the deep learning specialization directly, but it'll be a little bit harder since we tend to go over faster concepts like how does gradient descent work and what is an objective function, which is covered more slowly in the machine learning course.
L
Lex Fridman36:36
Could you briefly mention some of the key concepts in deep learning that students should learn, that you envision them learning in the first few months, in the first year or so?
A
Andrew Ng36:45
So if you take the deep learning specialization, you learn the foundations of what is a neural network, how do you build up a neural network from a single logistic unit to a stack of layers, to different activation functions. You learn how to train the neural networks. One thing I'm very proud of in that specialization is we go through a lot of practical know-how of how to actually make these things work. So what are the differences between different optimization algorithms? What do you do if the algorithm overfits or how do you tell if the algorithm is overfitting? When do you collect more data? When should you not bother to collect more data? I find that even today, unfortunately, there are engineers that will spend six months trying to pursue a particular direction, such as collect more data because we heard more data is valuable. But sometimes you could run some tests and could have figured out six months earlier that for this particular problem, collecting more data isn't going to cut it. So just don't spend six months collecting more data, spend your time modifying the architecture or trying something else. So it goes through a lot of the practical know-how so that when you take the deep learning specialization, you have those skills to be very efficient in how you build these networks, to dive right in, to play with the network, to train it, to do the inference on a particular data set, to build the intuition about it.
L
Lex Fridman38:06
Without building it up too big to where you spend, like you said, six months building up your big project without building any intuition of a small aspect of the data that could already tell you everything you need to know about that data.
A
Andrew Ng38:23
Yes. And also the systematic frameworks of thinking for how to go about building practical machine learning. Maybe to make an analogy, when we learn to code, we have to learn the syntax of some programming language. But the equally important or maybe even more important part of coding is to understand how to string together these lines of code into coherent things. So when should you put something in a function call and when should you not? How do you think about abstraction? So those frameworks are what makes a programmer efficient, even more than understanding the syntax. I remember when I was in undergrad at Carnegie Mellon, one of my friends would debug their code by first trying to compile it, and then it was C++ code, and then every line that had a syntax error, they wanted to get rid of syntax errors as quickly as possible. So how do you do that? They would delete every single line of code with a syntax error. So really efficient for getting rid of syntax errors but horrible debugging practice. So I think we learn how to debug. And I think in machine learning, the way you debug a machine learning program is very different than the way you do binary search or use a debugger, trace through the code in traditional software engineering. So it's an evolving discipline, but I find that the people that are really good at debugging machine learning algorithms are easily ten times, maybe a hundred times faster at getting something to work.
L
Lex Fridman39:45
And the basic process of debugging is, so the bug in this case, why isn't this thing learning, improving, sort of going into the questions of overfitting and all those kinds of things, that's the logical space that the debugging is happening in with neural networks.
A
Andrew Ng40:04
Yeah. Often the question is why doesn't it work yet, or can I expect it to eventually work, and what are the things I could try. Change the architecture, more data, more regularization, different optimization algorithm, different types of data. So to answer those questions systematically so that you don't spend six months heading down a blind alley before someone comes and says why did you spend six months doing this.
L
Lex Fridman40:29
What concepts in deep learning do you think students struggle the most with, or is the biggest challenge for them, once they get over that hill, it hooks them and it inspires them and they really get it?
A
Andrew Ng40:45
Similar to learning mathematics, I think one of the challenges of deep learning is that there are a lot of concepts that build on top of each other. If you ask me what's hard about mathematics, I have a hard time pinpointing one thing. Is it addition, subtraction, is it a carry? Is it multiplication? There's just a lot of stuff. I think one of the challenges of learning math and of learning certain technical fields is that there are a lot of concepts, and if you miss a concept, then you're kind of missing the prerequisite for something that comes later. So in the deep learning specialization, we try to break down the concepts to maximize the odds of each component being understandable. So when you move on to the more advanced things, we hope you have enough intuitions from the earlier sections to then understand why we structure ConvNets in a certain way, and then eventually why we build RNNs and LSTMs or attention models in a certain way, building on top of the earlier concepts.
Actually I'm curious, you do a lot of teaching as well. Do you have a favorite 'this is the hard concept' moment in your teaching?
L
Lex Fridman41:55
Well, I don't think anyone's ever turned the interview on me. I'm glad you go first.
A
Andrew Ng42:03
I think that's a really good question. It's really hard to capture the moment when they struggle. I think you put it really eloquently. I do think there are moments that are like aha moments that really inspire people. I think for some reason reinforcement learning, especially deep reinforcement learning, is a really great way to really inspire people and get what the use of neural networks can do. Even though neural networks really are just a part of the deep learning framework, but it's a really nice way to paint the entirety of the picture of a neural network being able to learn from scratch, knowing nothing, and explore the world and pick up lessons. I find that a lot of the aha moments happen when you use deep RL to teach people about neural networks, which is counterintuitive.
L
Lex Fridman42:53
I find like a lot of the inspired sort of fire in people's passion, people's eyes, comes from the RL world. Do you find reinforcement learning to be a useful part of the teaching process or no?
A
Andrew Ng43:09
I still teach reinforcement learning in one of my Stanford classes. And my PhD thesis was on reinforcement learning. So, I clearly love the field. I find that if I'm trying to teach students the most useful techniques for them to use today, I end up shrinking the amount of time I talk about reinforcement learning. It's not what's working today. Now, our work changes so fast. Maybe this will be totally different in a couple years. But I think we need a couple more things for reinforcement learning to get there.
L
Lex Fridman43:37
To actually get there. Yeah.
A
Andrew Ng43:38
One of my teams is looking at reinforcement learning for some robotic control tasks. So I see the applications but if you look at it as a percentage of all of the impact of the types of things we do, at least today, outside of playing video games and a few other games, the scope actually, at NeurIPS a bunch of us were standing around saying, hey, what's your best example of an actual deployed reinforcement learning application? And among senior machine learning researchers, and again there are some emerging ones but there are not that many great examples.
L
Lex Fridman44:12
Well, I think you're absolutely right. The sad thing is there hasn't been a big application, impactful real-world application of reinforcement learning. I think its biggest impact to me has been in the toy domain, in the game domain, in the small example. That's what I mean for educational purpose. It seems to be a fun thing to explore neural networks with. But I think from your perspective, and I think that might be the best perspective, is if you're trying to educate with a simple example in order to illustrate how this can actually be grown to scale and have a real-world impact, then perhaps focusing on the fundamentals of supervised learning in the context of a simple data set, even like an MNIST data set, is the right path to take. I just, the amount of fun I've seen...
A
Andrew Ng45:03
People have great enthusiasm with reinforcement learning, but not so much in terms of applied impact on real-world settings. So it's a trade-off: how much impact you want to have versus how much fun you want to have.
Yeah, that's really cool. And I feel like the world actually needs all sorts. Even within machine learning, deep learning is so exciting, but AI teams shouldn't just use deep learning. I find that my teams use a portfolio of tools — maybe that's not the exciting thing to say, but some days we use a neural net, some days we use PCA. Actually the other day I was sitting down with my team looking at PCA residuals trying to figure out what's going on with PCA applied to a manufacturing problem. Some days we use a graphical model, some days we use a knowledge graph, which has tremendous industry impact but the chatter about knowledge graphs in academia is really thin compared to the actual raw impact. So I think reinforcement learning should be in that portfolio, and it's about balancing how much we teach all of these things. The world should have diverse skills. It would be sad if everyone just learned one narrow thing.
L
Lex Fridman46:09
Diverse skills help you discover the right tool for the job. What is the most beautiful, surprising, or inspiring idea in deep learning to you? Something that captivated your imagination. Is it the scale that could be achieved, the performance that could be achieved with scale, or are there other ideas?
A
Andrew Ng46:29
I think that if my only job was being an academic researcher, with an unlimited budget and I didn't have to worry about short-term impact and only focused on long-term impact, I'd probably spend all my time doing research on unsupervised learning. I still think unsupervised learning is a beautiful idea. At both this past NeurIPS and ICML, I was attending workshops or listening to various talks about self-supervised learning, which is one segment of unsupervised learning that I'm excited about. Let me describe the idea briefly.
L
Lex Fridman47:06
No, please.
A
Andrew Ng47:06
So here's an example of self-supervised learning. Let's say we grab a lot of unlabeled images off the internet — we have infinite amounts of this type of data. I'm going to take each image and rotate it by a random multiple of 90 degrees. Then I'm going to train a supervised neural network to predict the original orientation: has this been rotated 90, 180, 270, or 0 degrees? So you can generate an infinite amount of labeled data because you rotated the image, so you know the ground truth label. Various researchers have found that by taking unlabeled data and making up label datasets and training a large neural network on these tasks, you can then take the hidden layer representation and transfer it to a different task very powerfully. Learning word embeddings — where we take a sentence, delete a word, predict the missing word — is another example. And there's now this portfolio of techniques for generating these made-up tasks. Another one called jigsaw: you take an image, cut it up into a 3x3 grid, and have a neural network predict which of the nine factorial possible permutations it came from. Many groups including OpenAI, Facebook, Google Brain, and DeepMind are doing exciting work on this. Aaron van den Oord has great work on the CPC objective. I think this is a way to generate infinite labeled data and I find this a very exciting piece of unsupervised learning.
L
Lex Fridman48:51
So long term you think that's going to unlock a lot of power in machine learning systems, this kind of unsupervised learning?
A
Andrew Ng48:59
I don't think there's a whole enchilada. I think it's just a piece of it, and this one piece — self-supervised learning — is starting to get traction. We're very close to it being useful. Word embeddings are really useful. I think we're getting closer to this having significant real-world impact, maybe in computer vision and video. But I think there'll be other concepts around it too. Other unsupervised learning things I've worked on, I've been excited about — sparse coding and slow feature analysis. These are ideas that various of us were working on about a decade ago before we all got distracted by how well supervised learning was working.
L
Lex Fridman49:42
So we would return to the fundamentals of representation learning that really started this movement of deep learning.
A
Andrew Ng49:51
I think there's a lot more work that one could explore around this theme of ideas and other ideas to come up with better algorithms.
L
Lex Fridman50:00
So if we could return to maybe talk quickly about the specifics of deeplearning.ai, the deep learning specialization — perhaps how long does it take to complete the course, would you say?
A
Andrew Ng50:10
The official length of the deep learning specialization is I think 16 weeks, so about 4 months, but it's go at your own pace. There are people that finished it in less than a month by working more intensely. It really depends on the individual. When we created the deep learning specialization, we wanted to make it very accessible and very affordable. With Coursera and deeplearning.ai's education mission, one of the things that's really important to me is that if there's someone for whom paying anything is a financial hardship, then just apply for financial aid and get it for free.
L
Lex Fridman50:51
If you were to recommend a daily schedule for people in learning, whether it's through the deep learning specialization or just learning in the world of deep learning, what would you recommend? How do they go about it day-to-day, sort of specific advice about their journey in deep learning and machine learning?
A
Andrew Ng51:12
I think getting in the habit of learning is key, and that means regularity. For example, we send out our weekly newsletter, The Batch, every Wednesday, so people know it's coming. You can spend a little bit of time on Wednesday catching up on the latest news. For myself, I've picked up a habit of spending some time every Saturday and every Sunday reading or studying. I don't wake up on a Saturday and have to make a decision — do I feel like reading or studying today? It's just what I do. The fact is, a habit makes it easier. If someone can get into that habit, it's like brushing our teeth every morning. I don't think about it.
L
Lex Fridman52:03
But it's a habit that takes no cognitive load. This would be so much harder if we had to make a decision every morning.
A
Andrew Ng52:09
And actually that's the reason why I wear the same thing every day as well. It's just one less decision. I just get up and wear my shirt. But I think if you can get that habit, that consistency of studying, then it actually feels easier.
L
Lex Fridman52:23
So yeah, it's kind of amazing in my own life. I play guitar every day — I force myself to at least for 5 minutes. It's a ridiculously short period of time, but because I've gotten into that habit, it's incredible what you can accomplish in a year or two. You can become exceptionally good at certain aspects of a thing by just doing it every day for a very short period of time. It's kind of a miracle that that is how it works. It adds up over time.
A
Andrew Ng52:53
Yeah. And I think it's often not about the burst of sustained efforts and the all-nighters, because you can only do that a limited number of times. It's the sustained effort over a long time. Reading two research papers is a nice thing to do, but the power is not reading two research papers — it's reading two research papers a week for a year. Then you've read 100 papers and you actually learn a lot. So regularity and making learning a habit.
L
Lex Fridman53:21
Do you have general other study tips, particularly for deep learning, in people's process of learning? Any recommendations or tips as they learn?
A
Andrew Ng53:38
One thing I still do when I'm trying to study something really deeply is take handwritten notes. I know a lot of people take the deep learning courses during a commute where it may be more awkward to take notes, so it may not work for everyone. But when I'm taking courses on Coursera — and I still take some every now and then. The most recent one I took was a course on clinical trials because I was interested in that. I got out my little Moleskine notebook and was sitting at my desk, just taking down notes of what the instructor was saying. That act of taking notes, preferably handwritten, increases retention.
L
Lex Fridman54:17
So as you're watching the video, just kind of pausing maybe and then taking the basic insights down on paper.
A
Andrew Ng54:25
Yeah. There have been a few studies — if you search online you'll find some of them — that taking handwritten notes, because handwriting is slower, causes you to recode the knowledge in your own words more, and that process of recoding promotes long-term retention. This is as opposed to typing, which is fine — typing is better than nothing, and taking a class without taking notes is better than not taking any class at all. But comparing handwritten notes and typing, you can usually type faster than you can handwrite. So when people type, they're more likely to just transcribe verbatim what they heard, and that reduces the amount of recoding, which actually results in less long-term retention.
L
Lex Fridman55:16
There's something fundamentally different about handwriting. I wonder what that is. I wonder if it is as simple as just the time it takes to write is slower.
A
Andrew Ng55:21
Yeah. And because you can't write as many words, you have to take whatever they said and summarize it into fewer words. And that summarization process requires deeper processing of the meaning, which then results in better attention.
L
Lex Fridman55:35
That's fascinating.
A
Andrew Ng55:37
I've spent so much time studying pedagogy — this is actually one of my passions. I really love learning how to more efficiently help others learn. One of the things I do both when creating videos or when we write The Batch is I try to think: is one minute spent with us going to be a more efficient learning experience than one minute spent anywhere else? We really try to make it time-efficient for the learners because everyone's busy. When we're editing, I often tell my teams every word needs to fight for its life — if we can delete a word, let's just delete it and not waste the learner's time.
L
Lex Fridman56:17
Oh, that's so amazing that you think that way because there are millions of people that are impacted by your teaching and that one minute spent has a ripple effect right through years of time, which is just fascinating to think about. How does one make a career out of an interest in deep learning? Do you have advice for people? We just talked about the beginning early steps, but if you want to make it an entire life's journey, or at least a journey of a decade or two, how do you do it?
A
Andrew Ng56:46
The most important thing is to get started. In the early parts of a career, coursework like the deep learning specialization is a very efficient way to master the material. Instructors spend effort to make it time-efficient for you to learn new concepts. At Stanford, some of my PhD students want to jump into research right away, and I tend to say: in your first couple years as a PhD student, spend time taking courses, because it lays a foundation. It's fine if you're less productive in your first couple years — you'll be better off in the long term. Beyond a certain point, there's material that doesn't exist in courses because it's too cutting edge. After exhausting efficient coursework, most people need to go on to work on projects and continue learning by reading blog posts and research papers. Doing projects is really important, and it's important to start small and just do something. Building that tiny neural network, be it MNIST or Fashion MNIST or your own fun hobby project — that's how you gain the skills to do bigger and bigger projects. This is true at the individual level and also at the organizational level. For a company to become good at machine learning, sometimes the right thing is not to tackle the giant project but instead to do the small project that lets the organization learn and build up from there. Taking the first step and then taking small steps is the key.
L
Lex Fridman59:02
Should students pursue a PhD? You can have so much impact in machine learning without ever getting a PhD. So what are your thoughts? Should people go to grad school? Should people get a PhD?
A
Andrew Ng59:17
I think there are multiple good options, of which doing a PhD could be one. If someone's admitted to a top PhD program — MIT, Stanford, top schools — I think that's a very good experience. Or getting a job at a top organization on a top AI team is also a very good experience. There are some things you still need a PhD to do — if someone's aspiration is to be a professor at a top academic university, you just need a PhD. But to start a company, build a company, do great technical work — a PhD is a good experience but not the only path. I would look at the different options available and weigh the pros and cons.
L
Lex Fridman1:00:06
So just to linger on that a little longer, what final dreams and goals do you think people should have? What options should they explore? You can work in industry at a large company like Google, Facebook, or Baidu. You can also do more research-oriented groups like Google Research or Google Brain. You can be a professor in academia. Or you can build your own company, do a startup. Is there anything that stands out between those options, or are they all beautiful different journeys?
A
Andrew Ng1:00:50
I think the thing that affects your experience most is not whether you're in this company versus that company, or academia versus industry. The thing that affects your experience most is who are the people you're interacting with on a daily basis. Even in large companies, the experience of individuals in different teams is very different. What matters most is not the logo above the door when you walk into the giant building every day. What matters most is who are the 10 or 30 people you interact with every day. I tend to advise people: if you get a job offer, ask who is your manager, who are your peers, who are you actually going to talk to. We're all social creatures — we tend to become more like the people around us. If you're working with great people, you will learn faster. For small companies, you can figure out who you'd be working with quite quickly. If a company refuses to tell you who you'd work with — someone says 'Oh, join us, the rotation system, we'll figure it out' — I think that's a worrying answer because you may not actually get to a team with great peers and great people to work with.
L
Lex Fridman1:02:19
It's actually really profound advice that we don't consider too rigorously — the people around you. Really often, especially when you accomplish great things, it seems that great things are accomplished because of the people around you. So it's not about whether you learn this thing or that thing, or like you said, the logo that hangs up top. It's the people. That's fascinating. And it's such a hard search process of finding the right friends and somebody to get married with and that kind of thing. It's a very hard search. It's a people search problem.
A
Andrew Ng1:02:58
Yeah. But I think when someone interviews at a university or a research lab or a large corporation, it's good to insist on asking: who are the people, who is my manager? And if you refuse to tell me, I'm going to think maybe that's because you don't have a good answer. If something feels off with the people, then don't stick with it — that's a really important signal to consider. Actually, in my standard class CS230 as well as an ACM talk I gave, I did about an hour-long talk on career advice including the job search process, so you can find those videos online.
L
Lex Fridman1:03:40
Awesome, I'll point people to them. Beautiful. So the AI Fund helps AI startups get off the ground, or perhaps you can elaborate on all the fun things it's involved with. What's your advice on how does one build a successful AI startup?
A
Andrew Ng1:03:59
In Silicon Valley, a lot of startup failures come from building products that no one wanted. Cool technology, but who's going to use it? I tend to be very outcome-driven and customer-obsessed. Ultimately, we don't get to vote if we succeed or fail — only the customer gets a thumbs up or thumbs down in the long term.
L
Lex Fridman1:04:33
So as you build a startup, you have to constantly ask the question: will the customer give a thumbs up on this?
A
Andrew Ng1:04:41
I think so. Startups that are very customer-focused, customer-obsessed, deeply understand the customer and are eager to serve the customer, are more likely to succeed — with the proviso that I think all of us should only do things that we think create social good and move the world forward. I personally don't want to build addictive digital products just to sell a lot of ads. There are things that could be lucrative that I won't do. But if we can find ways to serve people in meaningful ways, those can be great things to do in an academic setting, a corporate setting, or a startup setting.
L
Lex Fridman1:05:19
So can you give me the idea of why you started the AI Fund?
A
Andrew Ng1:05:26
I remember when I was leading the AI group at Baidu, I had two parts of my job. One was to build an AI engine to support the existing businesses. The second was to systematically initiate new lines of business using the company's AI capabilities. The self-driving car team came out of my group, the smart speaker team — similar to Amazon Echo Alexa but we announced it before Amazon did — that came out of my group. I found that to be the most fun part of my job. So I wanted to build AI Fund as a startup studio to systematically create new startups from scratch with all the things we can now do with AI. The ability to build new teams to go after this rich space of opportunities is a very important mechanism to get these projects done that will move the world forward. I've been fortunate to build a few teams that had a meaningful positive impact, and I felt we might be able to do this in a more systematic, repeatable way. A startup studio is a relatively new concept — there are maybe dozens right now — but all of us are still trying to figure out how to systematically build companies with a high success rate.
L
Lex Fridman1:07:20
So a startup studio is a place and a mechanism for startups to go from zero to success, to try to develop a blueprint.
A
Andrew Ng1:07:30
It's actually a place for us to build startups from scratch. We often bring in founders and work with them, or maybe even have existing ideas that we match founders with, and then this launches hopefully into successful companies.
L
Lex Fridman1:07:48
So how close are you to figuring out a way to automate the process of starting from scratch and building a successful AI startup?
A
Andrew Ng1:07:56
I think we've been constantly improving and iterating on our processes — things like how many customer calls do we need to make to get customer validation, how do we make sure the technology can be built. Quite a lot of our businesses need cutting-edge machine learning algorithms. Even if it works in a research paper, taking it to production is really hard — there are a lot of issues for making these things work in real life that are not widely addressed in academia. How do we validate that this is actually doable? How do you build a team with specialized domain knowledge? I think we've been getting much better at giving entrepreneurs a high success rate, but the whole world is still in the early phases of figuring this out.
L
Lex Fridman1:08:52
But do you think there are some aspects of that process that are transferable from one startup to another?
A
Andrew Ng1:09:01
Yeah, very much so. Starting a company is a really lonely thing for most entrepreneurs. I've seen so many entrepreneurs not know how to make certain decisions — how do you do B2B sales, how do you market efficiently, or for a machine learning project, basic decisions can change the course of whether the product works or not. There are hundreds of decisions that entrepreneurs need to make, and making a mistake on a couple key decisions can have a huge impact on the fate of the company. A startup studio provides a support structure that makes starting a company much less lonely. When facing key decisions like hiring your first VP of engineering — what are the selection criteria, how do you source — having our ecosystem around the entrepreneurs to help at key moments hopefully significantly increases the success rate.
L
Lex Fridman1:10:19
So they have somebody to brainstorm with in these very difficult decision points.
A
Andrew Ng1:10:25
And also to help them recognize what they may not even realize is a key decision point.
L
Lex Fridman1:10:31
Right. That's the first and probably the most important part.
A
Andrew Ng1:10:35
Actually I can say one other thing. Building companies is one thing, but I feel like it's really important that we build companies that move the world forward. Within the AI Fund team, there was once an idea for a new company that, if it had succeeded, would have resulted in people watching a lot more videos in a certain narrow vertical. I looked at it, the business case was fine, the revenue case was fine, but I just said, I don't want to do this. I don't actually want to have a lot more people watch this type of video — it wasn't educational. So I killed the idea on the basis that I didn't think it would actually help people. Whether building companies or doing personal projects, it's up to each of us to figure out what's the difference we want to make in the world.
L
Lex Fridman1:11:31
With learning AI, you help already established companies grow their AI and machine learning efforts. How does a large company integrate machine learning into their efforts?
A
Andrew Ng1:11:42
AI is a general-purpose technology and I think it will transform every industry. Our community has already transformed, to a large extent, the software and internet sector. Most software and internet companies already have reasonable machine learning capabilities or are getting there. But when I look outside the software and internet sector — manufacturing, agriculture, healthcare, logistics, transportation — there are so many opportunities that very few people are working on. I think the next wave for AI is to transform all of those other industries. There was a McKinsey study estimating $13 trillion of global economic growth, and a lot of that impact will be outside the software and internet sector. We need more teams to work with these companies to help them adopt AI, and I think this will help drive global economic growth and make humanity more powerful.
L
Lex Fridman1:12:53
And like you said, the impact is there. So what are the best industries, the biggest industries where AI can help, perhaps outside the software or tech sector?
A
Andrew Ng1:13:01
Frankly, I think it's all of them. Some of the ones I'm spending a lot of time on are manufacturing, agriculture, and healthcare. For example, in manufacturing, we do a lot of work in visual inspection where today there are people standing around using the human eye to check if a part has a scratch or dent. We can use a camera to take a picture and use deep learning algorithms to check if it's defective, helping factories improve yield, quality, and throughput. The practical problems we run into are very different from what you might read about in most research papers. The data sets are really small, so we face small data problems. Factories keep changing the environment — the lights go on or off, and recently a bird flew through a factory and pooped on something, which changed things. Making algorithms robust to all the changes that happen in a factory is a lot of fun — we're on the cutting edge solving problems before many people are even aware there's a problem there.
L
Lex Fridman1:14:27
That's such a fascinating space, you're absolutely right. But what is the first step that a company should take? It's such a scary leap into this new world of going from the human eye inspecting to digitizing that process — having a camera, having an algorithm. What's the first step? What's the early journey that you recommend?
A
Andrew Ng1:14:48
I published a document called the AI Transformation Playbook that's online and taught briefly in the AI for Everyone course on Coursera about the long-term journey companies should take. But the first step is actually to start small. I've seen a lot more companies fail by starting too big than by starting too small. Take even Google — most people don't realize how hard and controversial it was in the early days. When I started Google Brain, it was controversial — people thought deep learning had been tried and didn't work. My first internal customer was the Google speech team, not the most lucrative or important project. But by starting small, my team helped the speech team build a more accurate speech recognition system, and this caused other teams to start having more faith in deep learning. My second internal customer was the Google Maps team, where we used computer vision to read house numbers from street view images. It was only after those two successes that I started a more serious conversation with the Google Ads team. The early small-scale projects help teams gain faith and also learn what these technologies do.
L
Lex Fridman1:16:54
Are there concrete challenges that companies face that you see as important for them to solve?
A
Andrew Ng1:17:01
I think building and deploying machine learning systems is hard. There's a huge gulf between something that works in a Jupyter notebook on your laptop versus something that runs in a production deployment setting in a factory or agriculture plant. A lot of teams underestimate the rest of the steps needed after getting something to work on a laptop. I've heard this exact same conversation between machine learning people and business people: the ML person says 'My algorithm does well on the test set, I didn't peek,' and the business person says 'Thank you very much but your algorithm sucks, it doesn't work.' There's a gulf between what it takes to do well on a test set on your hard drive versus what it takes in a deployment setting. Common problems include robustness and generalization — you deploy something in a factory, maybe they chop down a tree outside, the lighting changes, and the test set distribution shifts. In machine learning, especially in academia, we don't know how to deal with test set distributions that are dramatically different from the training set. Also, the ML model is maybe 5% or fewer of the lines of code relative to the entire software system you need to build.
L
Lex Fridman1:18:56
So good software engineering work is fundamental to building a successful machine learning system.
A
Andrew Ng1:19:04
Yes. And the software system needs to interface with people's workloads. Machine learning is automation on steroids. If we take one task out of many in a factory and automate it, it can be really valuable, but you may need to redesign a lot of other tasks around it. Say the ML algorithm says something is defective — what are you supposed to do? Throw it away? Get a human to double-check? Rework it? You need to redesign a lot of tasks around the thing you've automated. Planning for change management, making sure the software is consistent with the new workflow, and taking the time to explain to people what needs to happen. What Landing AI has become good at — learned by making mistakes and painful experiences — is working with our partners to think through all the things beyond just the ML model running in a Jupyter notebook, but building the entire system, managing the change process, and figuring out how to deploy this in a way that has actual impact.
L
Lex Fridman1:21:02
So you mentioned some people are interested in discovering mathematical beauty and truth in the universe, and you're interested in having a big positive impact in the world. So let me ask a —
A
Andrew Ng1:21:13
The two are not inconsistent.
L
Lex Fridman1:21:14
No, they're all together. I'm only half joking because you're probably interested a little bit in both. But let me ask a romanticized question. So much of your work and our discussion today has been on applied AI — maybe you can even call it narrow AI where the goal is to create systems that automate some specific process that adds value to the world. But there's another branch of AI starting with Alan Turing that kind of dreams of creating human-level or superhuman-level intelligence. Is this something you dream of as well? Do you think we human beings will ever build a human-level intelligence or superhuman-level intelligence system?
A
Andrew Ng1:21:54
I would love to get to AGI, and I think humanity will, but whether it takes 100 years or 500 or 5,000, I find hard to estimate.
L
Lex Fridman1:22:05
Do you have — some folks have worries about the different trajectories that path would take, even existential threats of an AGI system. Do you have such concerns, whether in the short term or the long term?
A
Andrew Ng1:22:19
I do worry about the long-term fate of humanity. I do worry about overpopulation on the planet Mars — just not today. I think there will be a day when maybe Mars will be polluted and there are children dying, and someone will look back at this video and say, 'How was Andrew so heartless? He didn't care about all these children dying on the planet Mars.' I apologize to the future viewer — I do care about the children, but I just don't know how to productively work on that today. I think the biggest problem with self-driving cars is not the trolley dilemma — self-driving cars will run into moral dilemmas roughly as often as we do when we drive. The biggest problem is when there's a big white truck across the road and the car should brake but it crashes into it. I think we need to solve that problem first. The problem with some of these discussions about AGI alignment, the paperclip problem, is that it's a huge distraction from the much harder problems we actually need to address today. Bias is a huge issue. I worry about wealth inequality — AI and the internet are causing an acceleration of concentration of power because we can centralize data and use AI to process it. Industry after industry, we've infected these other industries with win-and-take-most dynamics. How do we make sure the wealth is fairly shared? How do we help people whose jobs are displaced? Education is part of it, but there may be even more that we need to do. There are adverse uses of AI like deepfakes being used for various nefarious purposes. I worry about some teams making a lot of noise about problems in the distant future rather than focusing on some of the much harder problems.
L
Lex Fridman1:25:33
Yeah, that could overshadow the problems that we have already today that are exceptionally challenging — like those you said, and even the silly ones but the ones that have a huge impact, like the lighting variation outside your factory window that ultimately is what makes the difference between the Jupyter notebook and something that actually transforms an entire industry.
A
Andrew Ng1:25:54
Yeah. And then some companies where a regulator comes to you and says your product is messing things up — fixing it may have a revenue impact. Well, it's much more fun to talk to them about how you promise not to wipe out humanity than to face the actually really hard problems we face.
L
Lex Fridman1:26:13
So your life has been a great journey from teaching to research to entrepreneurship. Two questions: one, are there regrets, moments that if you went back you would do differently? And two, are there moments you're especially proud of? Moments that made you truly happy?
A
Andrew Ng1:26:32
You know, I've made so many mistakes. It feels like every time I discover something, I go, 'Why didn't I think of this 5 years earlier or even 10 years earlier?' Sometimes I read a book and I go, 'I wish I read this book 10 years ago — my life would have been so different.' That happened recently. I was thinking if only I read this book when we were starting up Coursera, it could have been so much better. But I discovered the book had not yet been written when we were starting Coursera, so that made me feel better. The process of discovery — we keep on finding out things that seem so obvious in hindsight, but it always takes us so much longer than I wish to figure it out.
L
Lex Fridman1:27:20
So on the second question, are there moments in your life that you're especially proud of, or that filled you with happiness and fulfillment?
A
Andrew Ng1:27:35
Two answers. One, my daughter — no matter how much time I spend with her, I just can't spend enough time with her. And two, helping other people. I think the meaning of life is helping others achieve whatever their dreams are, and also trying to move the world forward by making humanity more powerful as a whole. The times I've felt most happy and most proud were when someone else allowed me the good fortune of helping them a little bit on the path to their dreams.
L
Lex Fridman1:28:14
I think there's no better way to end it than talking about happiness and the meaning of life. Andrew, it's a huge honor. Me and millions of people thank you for all the work you've done. Thank you for talking today.
A
Andrew Ng1:28:24
No, thank you so much. Thanks.
L
Lex Fridman1:28:26
Thanks for listening to this conversation with Andrew Ng. And thank you to our presenting sponsor, Cash App. Download it, use code LexPodcast, you'll get $10 and $10 will go to FIRST, an organization that inspires and educates young minds to become science and technology innovators of tomorrow. If you enjoy this podcast, subscribe on YouTube, give it five stars on Apple Podcast, support it on Patreon, or simply connect with me on Twitter at Lex Fridman. And now, let me leave you with some words of wisdom from Andrew Ng: Ask yourself if what you're working on succeeds beyond your wildest dreams, would you have significantly helped other people? If not, then keep searching for something else to work on. Otherwise, you're not living up to your full potential. Thank you for listening, and hope to see you next time.