Back
Kai-fu Lee
CEO of Sinovation Ventures, Sinovation Ventures

Ep. 12 - Machine Learning and Speech Recognition with Kai-Fu Lee

🎥 Dec 01, 2025 📺 Stanford Digital Economy Lab ⏱ 39m 👁 328 views
Tom meets with Kai-Fu Lee, a pioneer in using machine learning to significantly advance speech recognition. Kai-Fu, former president of Google China and now Chairman of Sinovation Ventures and CEO of 01.AI, has led speech, machine learning and AI efforts at several top firms, and is now one of the top AI venture capitalists in China. Tom Mitchell is the Founders University Professor at Carnegie Mellon University. Produced by the Stanford Digital Economy Lab.
Watch on YouTube

About Kai-fu Lee

Kai-fu Lee, CEO of Sinovation Ventures and founder of 01.AI, commented on the AI landscape at the 2026 World AI Conference in Shanghai. He described the Kimi K3 model as "excellent" and said his company would use it in its products. Lee stated that 01.AI is preparing for an initial public offering in 2027, saying its "financial numbers are ready right now, but it's a matter of going through the process." He argued that U.S. export controls on GPUs have "failed" to contain China, attributing any GPU shortage to business frugality or insufficient supply rather than the controls themselves. Lee predicted that in the AI model race, the U.S. will make more money while China will have larger market share due to open-source distribution and enterprise demand for on-premise deployment. He cited Chinese government support and high consumer optimism—noting a poll showing 84% of Chinese view AI positively versus under 50% in the U.S.—as strengths, while acknowledging challenges in getting enterprises to pay substantial fees for AI. In earlier remarks from 2018, Lee discussed AI's impact on employment, stating that routine white-collar jobs are more vulnerable to automation than blue-collar roles due to the difficulty of robotics. He described the Chinese approach to tech markets as "do whatever it takes to win" in a winner-take-all environment, with companies aiming for domination of increasingly broad product categories. Lee also noted that AI technologies have become more accessible through open-source tools, advising a shift from a laboratory mindset to a business-oriented approach, except for those with genuine technological breakthroughs.

Source: AI-verified profile updated from Kai-fu Lee's recent appearances. Browse all interviews →

Transcript (19 segments)
T
Tom Mitchell0:00
Welcome to Machine Learning, How Did We Get Here? I'm Tom Mitchell, your podcast host. Today's episode is an interview with Kai-fu Lee, who as a PhD student in the 1980s showed how to apply machine learning to dramatically advance the field of speech recognition. After he graduated, Kai-fu went on to lead speech recognition at Apple, later became the founding director of Microsoft Research in China, later became president of Google China, and today is one of the top venture capitalists in China investing in artificial intelligence and machine learning. I hope you enjoy the conversation with Kai-fu.
I'm pleased to have with me today Kai-fu Lee, one of the pioneers in machine learning, speech recognition, artificial intelligence. Kai-fu, it's great to have you here.
K
Kai-fu Lee1:15
Thank you, Tom. Great to be here.
T
Tom Mitchell1:17
Let me kick it off toward the beginning. You're well known really as a person who brought machine learning and in particular hidden Markov models to speech recognition way back when you did your PhD thesis. Can you explain briefly what that was and why was it so important to the field to make that move?
K
Kai-fu Lee1:47
Sure. So I got just excited about AI when I entered college in Columbia University in '79 and I was mentored by a number of professors. And I just thought AI was really the final step for humans to understand ourselves and how our brain works and whether we can build something half as smart as us. And that was going to be my life's dream. And with one of my advisers, John Kender, who was a CMU graduate taught at Columbia, he recommended me to CMU and I was so lucky I got in. And then CMU had this matching process where I got to hear every professor and I was really blown away by Raj Reddy who described his vision of the future and with his confidence and passion I thought this was it. I considered a number of areas in machine learning but speech recognition appealed to me the most because at the time compute was still very very slow and expensive and I didn't think something as large as general natural language understanding or computer vision for that matter was really possible for me to build something that would be an important step forward. Speech recognition frankly just felt like a more tractable problem because it's converting speech to text in at the time a limited domain. And then I chose Dr. Raj Reddy to be my professor and he taught me all about speech recognition, signal processing. And then at the time the speech group run by Raj was in the process of getting a large DARPA grant. It was to create a speaker independent continuous speech recognition system for a domain and that at the time was already an unsolved problem but it didn't seem so intractable so I entered into it. The primary machine learning method at the time was expert systems so that was the way the DARPA project overall was structured. It was built around another amazing computer scientist, Victor Zue, who could read the spectrogram. So he could look at the visual representation and read out what the person was saying. So the belief was that the spectrogram contained enough information if a human could read it then so can a machine. So that was kind of the overall direction of the DARPA project. A collaboration by a number of universities with CMU and MIT being the leading universities. And during my studies I ran into another amazing person, Dr. Peter Brown. He was I think two years ahead of me, a grad student, but he was basically working mostly at IBM under the supervision of Fred Jelinek and Bob Mercer and they were the ones who really were the first to use hidden Markov models. But IBM was not very open in describing all the ways in which they did things. And also I think people just felt that wasn't the best approach. So actually the mainstream approach was expert systems. The old approach was dynamic programming. That was the known proven approach invented by a Japanese scientist named Dr. Itakura. That was the known approach with some tiny commercial success. And then the DARPA approach was the next one is expert systems. But I got introduced luckily through Peter and with support from Fred and Bob that he would mentor me. I think what they wanted to see was just to be nice to a young guy and maybe when he graduates he'll join IBM. I think that's what they had in mind. Peter obviously didn't tell me any of what IBM was doing. But he taught me about hidden Markov models. Hidden Markov models are basically a form of neural networks but simplified for time series. So it's basically left to right. It has the ability to do very fast decoding using the Viterbi algorithm which is also a more generalized form of dynamic programming or dynamic time warping. So I felt that was the one I want to try. So I learned from Peter. I built some simple programs and then one day I just got enough courage to go to Raj Reddy to say that I love you as my adviser. I love speech recognition but can I try something different using hidden Markov models which is stochastic machine learning. But I would provide a different way to try the DARPA problem. It might or might not work but it would just be me. And later one or two other students joined me but at the time it was just me and Raj said something really amazing which was I don't agree with you but I will support you and I think that is a level of magnanimity, generosity and a way of management that I've learned a lot from. That in a scientific exploration you cannot just say this is the way. You have to let people pursue what they're good at and what their passion is. So he not only let me do the speech recognition my way but he asked me what resources I needed and I said well I need a lot of data because Fred Jelinek was famous for the quote that there's no data like more data. So I became a strong believer in that and that is still true today. So Raj went to DARPA and said well we need to get the National Institute of Standards to collect a bunch more data and let everybody use it. Expert system approach needs it too. And I was also very very lucky that with the DARPA grant the speech group at CMU was probably one of the most compute-rich speech groups in the world. A lot of some workstations were bought. They were SPARCstations really the most efficient for CPUs, not GPUs, but still the most efficient at the time. And the people building expert systems just used it to label spectrograms and then the expert system wasn't very compute intensive. So I got to use all the leftover cycles from the SPARCstations and there were more cycles left over than used. So I had at my disposal about I think about 20 SPARCstations which was phenomenal compute power. So with Raj's support I had Peter as my mentor, hidden Markov models, Raj as my sponsor with his wisdom and strategy and encouragement and equally importantly data and compute and that's how it began.
T
Tom Mitchell9:53
That's amazing. That's a great story. So, at the time, I want to ask you, it seemed to me that speech recognition when we think about it today is very different from the state-of-the-art when you began your research. I recall at some point in the late 80s, it was possible kind of to recognize some isolated spoken words. But sentences like you and I are speaking right now would be really, really difficult. So can you just kind of summarize where the state-of-the-art was and what the impact of your work on hidden Markov models was?
K
Kai-fu Lee10:35
Sure. I think the state-of-the-art prior to my PhD work was IBM which built a so-called dictation system and which they also sold a little bit and also there was a company called Dragon also built by two CMU graduates, Jim and Janet Baker. And both systems pursued the dictation, their approach was the application and use case and people who would pay money are those who are unable to type. Namely people who were somehow disabled and then unable to type for whatever reason but clearly the technology could not handle the transcription like we are now getting used to. So they made two very very clever but very constraining requirements. One was each person has to train the speech recognizer. So when you bought the speech recognizer, you have to read a script. The system would say you know read the following sentence and you would read it and it would then take hours to train the hidden Markov model to recognize yours and only your speech and that was called speaker dependent speech recognition. But that wasn't good enough because continuous speech decoding was too large a problem even with Viterbi decoding. So they made another constraint which was you only could speak like this. Then what the system does is it takes the pause as a definitive breakpoint and recognize the next word so that your branching factor if your vocabulary was let's say 5,000 words the system only had to do okay I detected a pause the next word is going to be one of the next 5,000 given the previous word is recognized what might the probability be given both the acoustic evidence hidden Markov model recognition results but also language model which was what was the likelihood of the next word given the one or two previous words that were spoken and recognized. So that was one path. So it's called speaker dependent isolated dictation. That's one. The other that was also commercially sold was what was done by Bell Labs and because they were in the business of recognizing phone numbers. There was no hope of really recognizing anything complex because the telephone had a narrower bandwidth and also they could not do speaker dependent. If someone called up and the system said press one to ask about directions, press two to make a deposit, press three to be connected to operator. That got changed to press or say one to connect to sales, press or say two to get directions and so on. So they were able to add press or say. So they were able to make speaker independent but the system only did digits. So it had a vocabulary of 10. That was another way to constrain the problem. It could listen to anybody's speech. And the commercial systems generally also did isolated words even though there was beginnings of continuous digit recognition. So it was isolated digit recognition for speaker independent recognition and beginnings of speaker independent continuous digit recognition. That was what Bell Labs did and those were the state-of-the-art commercialized but with limited commercial benefit but nevertheless marginally useful to some people.
T
Tom Mitchell14:50
That's great. So then just briefly what was the impact of the research that you were doing around that time in your dissertation? What were you able to demonstrate and shortly after that?
K
Kai-fu Lee15:06
Yeah. So I also had to constrain the problem and the funding from DARPA came with a task. It was a resource allocation task. Basically something used in the government to ask about resources being used and applied and it basically had a context-free grammar. So it was continuous speech and also speaker independent that was the goal set out by DARPA with Raj's advice. So it was can expert systems deal with continuous speech speaker independent and large vocabulary. Large vocabulary being a thousand words, not quite dictation. And continuous speech being the way we speak, but only would work for things within a context-free grammar, which was a large branching factor, but not full dictation. And speaker independent was I think they had maybe around a hundred different speakers and it was important to prove speaker independent so that the speakers you test with must be different from the speakers you trained with. So those were the constraints set out and initially the accuracies were very low, 10 or 20% and then they got a little bit better but really never at a meaningful or useful level. I think they got maybe up to 50 or 60% word accuracy. So I decided I would basically take the same task they were doing because they had all the data and also they become a benchmark. So if I could beat that single-handedly that would be significant contribution and if I couldn't it would at least validate or invalidate an approach. And I came up with a number of ways that enhanced what I learned from IBM and Peter. At the time I don't think the IBM folks knew whether I would produce a great result or not because they were very much focused on speaker dependent. So it was meaningful to them to see if it would work but they didn't know I didn't know Raj didn't know but as I built a basic system it came out with about 70 or 80% accuracy. So that was already very very exciting. And on top of that I learned a lot from Victor Zue and his counterpart at CMU Ron Cole about the phonetic and acoustic aspects of speech. And I used some of their simple parts actually nothing complex. It was just that words are composed of phonemes which are sounds like b, a, t and these sounds these phonemes basically make up a word but some words might have multiple pronunciations like tomato tomato and furthermore phonemes don't exactly sound the same in different contexts so if I say bat and cat even though the a part sounds the same to our ears. But the B and C sound heavily affects the way a would be instantiated in the spectrogram. So that was one aspect of okay given that knowledge how can I build the hidden Markov model to deal with contextual influence across phonemes. That was one fundamental new idea that I had that other people did not have. And also I was lucky to have worked with a visiting researcher named Dr. Shikano from Japan and he was an expert in signal processing and he suggested Kai-fu we know that logarithmic representation of the spectrogram is much better for speech recognition so don't use the normal FFT spectrogram and he gave me his module which I plugged in which was called a mel-cepstral representation. So that also made a significant difference. And then lastly to build the language model I did not directly use the context-free grammar because that is very hard to, not hard but very time intensive to run with the Viterbi algorithm. So I thought why don't I just use bigrams and trigrams which is use the last two words to predict the next word and that was of course inherited from IBM. It's not a breakthrough for me but it was an important step. So and then there were a number of other small tweaks but it was a combination of Dr. Shikano's signal processing, my own ideas of how to improve the hidden Markov model structure to model the phonetics and acoustics, and using a statistical not a grammatical model to improve the speech recognition. So those were the things that I plugged in. And then I remember one Saturday I think I woke up and I saw that the system produced 96% accuracy and I was like blown away because before the best I could do was in the 80s and it turns out with these three main ideas plus a bunch of other tweaks I got up to 96%. And that was viewed as a big breakthrough at the time.
T
Tom Mitchell21:40
Yeah. That was and it was also in an era where people were beginning to as you say apply machine learning in different forms to Markov models, one form of it, to the speech problem and as you say collecting enough data that it was possible to train such a system was one of the key steps. So then over the years you've seen and you've actually been engaged in much of the advancement in speech recognition in a number of different organizations. Can you just kind of summarize the intervening years? How did we get to here in the year 2025 when we could have a conversation like we are having and it will be easy for us to get a speech-to-text transcript of this entire conversation when we're finished? How do we get to here from there?
K
Kai-fu Lee22:50
Yeah. So it was always my dream and it looks like that dream was accomplished with great advancements made by other people. So after my thesis the approach I took became kind of the standard. CMU licensed Sphinx to a number of companies and later open sourced it. I think that kind of put a performance level that enabled a lot of companies to build from that open source. And then people continue to tweak it using hidden Markov models and I think there were further improvements for another five years after my work. I worked on it at CMU as an assistant professor. I went to Apple and kept working on it and we built some prototypes and a lot of other people were building prototypes. We built a product, other people built products. But I think the improvements on hidden Markov models after five years after my thesis I would say became very slow and incremental because I think we squeezed the most we could out of this hidden Markov model technology and there were still improvements but very slow. I remember when I worked at Microsoft, Bill Gates would ask me, you know, when can we do full speech recognition and I would say we hope in five years and five years later he asked me again. I would say in five years and then I realized okay this is not going to get us there. Because we always draw a trajectory, right? People saw my big improvement in jump and said okay in five years we can do it and there were further improvements. Okay, not as much a jump but still improvement people were optimistic but then five years after that and five years after that the improvements became really incremental. So yes on the one hand there were more and more use cases and products and companies on the other hand we really saw the writing on the wall here the Markov models were going to get a bunch of applications but none of which made a lot of money for companies and I think the next big breakthrough was deep learning. And when Jeff Hinton first demonstrated deep learning broke all the benchmarks for computer vision using ImageNet Fei-Fei's database. Then the next milestone was speech recognition. So deep learning helped make another big jump. But deep learning was computationally much more expensive. So it created research results papers and benchmarks. So people knew this thing is going to be much better than hidden Markov models. It turns out the major thing is deep learning was always the ideas of using neural networks was always there but there just wasn't enough compute power to compute a large enough neural networks. Using one or two layers they were handily beat by hidden Markov models. But when you add a lot more layers and also get a lot more data not with an hour of speech but with a hundred or thousand hours of speech that really drove it up and then more applications proliferated and there were even cases where Microsoft some of my former colleagues demonstrated that Microsoft's deep learning system matched human performance. So that was considered a big feat. However, if you look under the covers, it was still a constrained environment that was fairly comparing human and machine, but it wasn't really able to deal with the full scope. So people were able to make the claim but we all know it wasn't there and also it was computationally very expensive. So really the next set of breakthroughs are what we now know today with transformer and all the technologies that come with large language models and there are a number of breakthroughs. Yan LeCun's proposal and also Andrew Ng who basically proposed to forget the language model and acoustic model all put it together make it end to end. So as with most large language model systems end to end works better. So we're just basically adding more data creating better structure and then let machine do all the understanding. So the work is somewhat analogous to what I did right add more data come up with an architecture of the machine learning algorithm and the model and get a lot of compute and train with it except this time I think deep learning plus transformer or call it large language model really did the trick. Also the idea of using bigrams and trigrams such as IBM and I did were sort of the most basic model of a dumb transformer, right? It was looking at context but really just the last two words. But now transformer can look you know as much as a million words and do it selectively using what's called the attention mechanism. So I think one could be basically reminiscent of the past and say well it's all flavors of similar things. It's all about model architecture, lots of data, lots of compute and good ways to deal with contexts. But honestly the ones that people came up with today are just so much better. And I think we're now looking at almost in all scenarios matching human performance and surely it will be beating human performance in the near future.
T
Tom Mitchell29:24
Do you think of the problem of speech recognition as solved?
K
Kai-fu Lee29:32
Technologically it is solved. I think there are still tiny issues like you know when it's really noisy can you do ambient and making it really cheap and run on a device but one could say these are just engineering.
T
Tom Mitchell29:55
That's amazing. And you know when you think about the 35 year roughly period that we're talking about from when you were first working on the problem till today. That's quite a set of advances. The other thing I wanted to get your opinion on I know that you have done much in your career far beyond speech recognition. You for example ran an entire research lab for Microsoft in China. Currently one of the main venture capital firms in China. So I want to get your take just on the whole field of AI, machine learning, how you see it developing. Maybe I break it into two questions. One is what are the biggest surprises to you in the past five years and the second question is in the next 5 years what do you think are some of the surprises that we might see?
K
Kai-fu Lee31:17
Yeah, I think the big surprise was that the transformer architecture could carry us as far as it did and that the scaling law would work as long as it did. And that we were able to finally get reinforcement learning to work after years. And I think each of the three I knew they would work because I believe in trajectories. So I was one of the early testers of GPT-2. So at the time it became clear to me this was going to have a huge exponential improvement given how fast compute would go and that this idea of transformer and attention was really brilliant. But I would not have predicted it scaled as far as it did even to today. But today we can all confidently predict it will scale further not necessarily through the scaling law as it was originally described for you know text and GPUs and data but in advanced functions. I would say all of the things that AI does today, computer vision, machine translation, and the super smart chat bots, ability to do deep research and many many others. Ability to program better than people. Ability to win at the top world math olympiad. These were all things I would have predicted would happen perhaps in my lifetime but not as quickly as they did. So I would say each one of the examples I gave, the speed at which it went from not quite as good as human but sort of useful to better than human and still growing really still really astounded me. I did not have that confidence that I think a lot of the young researchers had the likes of Google and OpenAI and I did and I think it was largely you know people think in trajectories. I was poisoned by a long trajectory that was very slow. So it took me a while to adapt to this. So I saw it coming eventually. It was clear. But I think the conviction that people like Ilya and others have are really inspirational and surprising. And now I think I and many other people are learning from the young generation. And I think the younger people, if you look at those super young people being hired by xAI and OpenAI and these people are in their late 20s and they grew up in the large language model era. So their mind is totally uncontaminated. Right? The way mine was, I was lucky to have been in machine learning uncontaminated by expert systems. But then people who were natively deep learning, they blew people like me away and now people who are natively generative AI and LLM blew away the deep learning people again. So I think the path forward will have so much more surprises for us and I wouldn't at all be surprised about AI making amazing scientific discoveries. We'll probably see that within a year. AI winning a Nobel Prize. Let's say AI making Nobel Prize level work. I would say that will surely happen in three or four years. But getting the Nobel Prize may take a little longer and once that happens you know advances to humanity bringing in healthier life and abundance of material goods to the world and letting us co-build the future and even watch AI build the future of humanity is completely predictable.
T
Tom Mitchell35:52
One final question. If you could speak with the incoming PhD student in computer science, AI, what kind of advice might you give them?
K
Kai-fu Lee36:08
I would actually give a very practical advice which professors may not like which is that the whole generative AI built on transformer the path forward must be built on a giant computer infrastructure which academic institutions don't have so if that's the path you want to work on and be a part of this breakthrough you need to work with a professor who has a partnership with a company that has the computing resources, then you've got the best of both worlds, right? The academic freedom to think and do things out of the box and the computing resources. Or if a student doesn't or cannot or chooses not to do that, then I would say go out of the box. There's probably something beyond Transformers, right? Don't go building on top of tweaking something Google or OpenAI already did because if you don't have the compute resources, you're not going to beat them. But you can do something they don't yet know how to do. Go build the next transformer. Go invent the next transformer or invent the next reinforcement learning. Those are riskier. Those can be tested on much smaller computing infrastructure. So I would say one of the two paths you know either get the large company's resources or go do something beyond the current state of the art. If we look historically once speech recognition or computer vision or autonomous vehicles or if we look further back search engines once they create enough commercial value the academic institutions alone can no longer compete. So either work with a company with the resources and the data or go invent the next thing. And I think the future of academia is still bright, but one has to be practical, right? Post Google, people who insisted on doing information retrieval just met a dead end. Early stages of autonomous vehicles were phenomenally exciting at CMU Stanford. But once Waymo and others and Tesla got into the game, building the next autonomous vehicle cannot be the research topic. Either you work with a Waymo or a Tesla or you do the next thing. So that's my suggestion.
T
Tom Mitchell38:44
Great advice. Kai-fu, thank you so much for sharing your time, your wisdom, your history. Much appreciated.
K
Kai-fu Lee38:53
Thank you, Tom.