About David Luan
David Luan, CEO and co-founder of Adept AI, has been discussing the company's focus on building AI agents for enterprise knowledge workers. He has described Adept's approach as training a foundation model that can translate natural language into actions on a computer, rather than competing directly with general-purpose LLM providers. Luan has stated that the company is "very enterprise focused" and is working on larger engagements where Adept agents can accelerate knowledge work. He has also emphasized that Adept controls both its own foundation model for agents and the product enterprises use, describing this as a bet on vertical integration.
In various interviews, Luan has shared his views on the broader AI landscape. He has argued that the business of training large base models is "quickly commoditizing" and that companies in that space will need to find alternate ways of making money beyond having better models. Luan has also commented on the importance of reliability for agent products, stating that if an agent makes operational errors a third of the time, people will stop using it. He has predicted that AGI is "really not super far away" but that it will not diffuse through society quickly due to other bottlenecks. Luan has also expressed concern about regulatory capture, saying that lawmakers' lack of understanding of the technology makes it easier for those with ulterior motives to shape policy.
Source: AI-verified profile updated from David Luan's recent appearances.
Browse all interviews →
Transcript (38 segments)
H
Heather Mack0:04
Hi everyone, welcome to Gray Matter, the podcast from Greylock where we share stories from company builders and business leaders. I'm Heather Mack, head of editorial at Greylock. Today we're rebroadcasting our episode featuring Greylock General Partner S.'s conversation with David Luan and Percy Liang. David is the co-founder and CEO of AI startup Adept, and Percy is a computer science and statistics professor at Stanford. While text and image-generating AI tools like ChatGPT and DALL-E are all the rage right now, Adept is developing tools that take things a step further by actually executing actions based on text commands. The company just raised 350 million in Series B funding to further its development of a tool that can be thought of as an AI teammate trained to use every software tool and API for every knowledge worker. Greylock contributed to the latest funding round, and the firm has been partnering with Adept since co-leading the company's Series A in 2022. In this interview, David, Percy, and S. discuss how advancements in large language models are paving the way for the next wave of AI. This interview took place during Greylock's Intelligent Future event in August 2022, the summit featured experts and entrepreneurs from some of today's leading artificial intelligence organizations. You can read a transcript of this interview on our website, and you can also watch the video on our YouTube channel. Both are linked in the show notes. And if you aren't already a subscriber to Gray Matter, you can sign up wherever you get your podcasts.
S
S.1:17
Okay, David, Percy, I'm excited about this. There's no doubt that large-scale models are topical for all of us here, and I'm really excited to have the two of you to discuss them with. For those of you in the audience who aren't familiar with these two gentlemen, Percy is the associate professor of computer science and statistics at Stanford, where among other things he's the director for the Center for Research on Foundation Models, and David is one of the co-founders and CEO of Adept, an ML research and product lab building general intelligence by enabling humans and computers to work together. And before Adept, David was at Google leading a lot of large models effort, and before that at OpenAI. And we're fortunate to get to partner with David and the team at Adept here at Greylock. Percy, David, thank you guys for being here and for doing this. So I want to start high level and just start with the state of the play. There's a lot of talk about large models and it's easy to forget that a lot of the recent breakthroughs and models that we're all familiar with, like DALL-E and GPT-3, are actually fairly recent. And so we're still in the early innings of these models running in production and delivering real concrete customer and user value. Maybe just give us the state of the play, David, starting with you. Like, where are we with large-scale models and what's the state of deployment of these models today?
D
David Luan2:35
Yeah, I think the stuff is just incredibly powerful, but I think we're still underestimating how much there is left to run on this stuff. It's still so incredibly early. Just let's take a look at a couple different axes, right. When we were training these models at Google, it became incredibly clear up front that you could basically take a lot of these hand-engineered machine learning models that people had been spending a lot of their time building, rip it out with this giant model, give it some fine data, and turn it into a smaller model again and serve it, and that would just end up outperforming all of these things that people had done in the past. And so the fact that they're able to improve existing things that companies are already using machine learning for, but also just how great it has been as a way to be able to create brand new AI products that couldn't exist before. It's fascinating to me to watch things like GitHub Copilot and Jasper and stuff like that just hit a nerve so fast and go from zero to hero in terms of adoption. I think we're just in the very early innings of seeing a lot more of that. So I think that's Axis one. I think Axis two too is that primarily what we're talking about so far has been language models, right? But there are so many other modalities, sources of human knowledge. What happens when it's not just about predicting the next token of text, but becomes about predicting all of those other different things? And we're going to end up in a world where a lot of humanity's knowledge is going to get encoded in various different foundation models for many different things, and that's going to be really powerful as well.
P
Percy Liang4:02
Yeah, I kind of want to highlight, I agree with everything that David said. I want to emphasize one distinction he made, which is already with all the applications out there, these foundation models can just lift all boats and just make all the numbers kind of go up. I think another thing which is even more exciting is that there's a huge sea of applications that we're not even maybe dreaming of because we're kind of stuck in this paradigm where what is ML? Well, you could get some data, you train on it, but with prompting and all these other zero-shot capabilities, I think you're going to see a lot more new types of applications. So I think we should be looking not just for how to make faster horses or faster cars, but new types of applications.
S
S.4:49
First, maybe to follow up on that, I totally agree, and I think it connects to David's point around something like Copilot. And the thing that's amazing to me about something like Copilot is both how new of an experience it is and how quickly it's taken off and gone into end-user adoption. What are some of the other areas that you're looking forward to and are excited about in terms of net new applications that become possible because of these large models?
P
Percy Liang5:09
Yeah, so I mean, maybe one general category you can think about is creation. So this includes code, text, proteins, videos, PowerPoint slides, anything that you can imagine humans kind of doing right now which could be a creative or sort of a more task-oriented activity. You could imagine these systems being very helpful, with you in the loop, taking you much farther and giving you many more ideas. So I think the space is quite broad, and underscoring the multimodal aspect of this which David touched on is really important. We shouldn't just think about language models and code models and image models. Think about things that you could do when you mix these together, creating different illustrated books or films or things like that. One thing that you have to deal with is long context dependence. Relatively, right now you're generating single images or text up to maybe 2,000 or 8,000 tokens depending on your model, but imagine generating full films. That's going to require pushing the technology further, but we have the data, and if we can harness that and scale up, then I think there's a lot of possibilities out there.
S
S.6:29
David, what would you add? I mean, at Adept, you guys spend a lot of time thinking about how to use these models to unlock new ways of collaborating with computers and software. I'm curious what some of the use cases you think about are.
D
David Luan6:43
So I think the thing that I'm most excited about right now is that all the creativity use cases I personally just highlighted are going to be extremely powerful. But I think what's fascinating about these models is if you ask these generative models to go do something for you in the real world, they kind of just pretend like they're doing something because they don't have a first-class sense of what actions are and what affordances are on your computer. So the thing that I'm really excited about in particular is how do we bridge this gap? How do we train a foundation model of all the actions that people take on a computer? And I think once you have that, you have this incredibly powerful base for being able to turn natural language into any sort of arbitrary complexity thing that you would then do on your machine.
S
S.7:23
So maybe if we take something like actuation as a key net new capability or we take longer contexts as an important net new capability, there's the form of the question is where do we still need to see key research blocks and where are the key areas of focus to actually make these products a reality?
P
Percy Liang7:46
I think there are maybe two sides of things. One is pushing up capabilities and one is making sure things are pushed up in a way that's robust and reliable and safe. The first one is in terms of scaling. If you think about video and the ability to scale to hundreds of thousands of sequence lengths, you're going to have to do something different. The Transformer architecture has gotten us surprisingly far, but you need to do something different there. And I think David mentioned this briefly, but I think these models are still in some ways chatterbots. They give you the illusion that there's something good going on, and in certain applications this is okay if there's another external validity check on things with the human in the loop. But I think there's a deep fundamental research question on how to make these models actually reliable. There's many strategies that people have tried, using reinforcement learning or more explanation-based or retrieval-augmented methods, but I feel like there's still something deeper missing. I think this is one thing I hope the academic community and researchers will work on to ensure that these foundation models have good and stable foundations as opposed to shaky ones.
D
David Luan9:14
Yeah, I agree with a lot of what Percy just said. I think I would just add that the default path that we're on is increasing scale and increasing data, and I think that will continue to lead to a lot of gains. But the question becomes how do we pull forward the future faster? I think there's a lot of different things that we should be thinking about. One is specifically on the data side. I don't think most people—I'm curious later on to understand from the audience how many people would agree that actually I think we're much more constrained on data than we think. Within the next couple of years, everyone's going to have, just to take language as an example, plus or minus 20% quality, similar number of tokens from web crawls, so then the question becomes where next? So I think that's a really important question. Another important question is what does true creativity mean? To me, true creativity means being able to discover new knowledge. And the new knowledge discovery process, at least for foundation models as we're training them today, we actually get better at training these models, but that just better models the training distribution. So I think giving these models the ability to go gather new information and try out things is also going to be really key. And finally, on the safety side, we have a lot more to invest in and a lot more questions we have to answer.
S
S.10:34
So let's get to safety in a moment. Continuing on data, because I think that is a really important topic here. David, at Adept, you all are thinking about how to build products that humans collaborate with, and I think one of the nice consequences of that becomes this data flywheel. Can you maybe add a little bit about how you're thinking about that and how you're approaching designing products that end users will work with?
D
David Luan11:00
Yeah, I think that it starts out with having a pretty crisp definition of what we want the end game to look like. And I think for us, what we want to be building is teammates and collaborators for people. A series of increasingly powerful software tools that help humans increase the level of abstraction at which they can interact with their machine. To use a different analogy, it doesn't replace the musician but it gives musicians synthesizers. That kind of analogy, except for doing things on your computer. Because that's where we want to go, I think what's really important to us is how do we solve these HCI problems where it really feels like you're working together with the machine at the same time, using that as an opportunity for us to be able to learn from how humans break down really complicated problems, how humans actually get things done. That may be part of things that are much more complicated than trajectories you might just be able to see on the internet.
P
Percy Liang11:51
Just add to that. I think the interaction piece is really interesting here because these models are in some ways the most interactive ML models we have. You have a playground, you type in a prompt, and you immediately get to play with a model, as opposed to the previous cycle where someone gathers some data, trains a model, and then you experience it from the user side. So the line between developer and user is actually getting blurred, which I think is a good thing because if you can connect these two up, then you get better experiences.
S
S.12:28
Is there anything interesting from the Stanford both the HCI perspective and the foundation models perspective on the research side that you all are working on around interaction?
P
Percy Liang12:40
Yeah, so one thing that we've been doing at Stanford is, as a part of a larger benchmarking effort, trying to understand the ways in which these models—what it means for humans to interact with these models. The classic way that people think about these models is you train them and then there's 100 benchmarks and you evaluate, taking the automation approach. But as we know, a lot of the potential here is in copilot or autocomplete experiences where there is a human in the loop. Humans, and I think Adept is also a good example of this. What does that mean? Should we be building our models differently if we know that humans are going to be in the picture, as opposed to doing full automation? That's interesting because maybe in some cases you want a model not just to be accurate, but you want it to be more interpretable or more reliable or understandable. And for creative applications, you may want a model to actually have a broader distribution of outputs. We're seeing some of this where what is good for actual interaction is not necessarily what's good just for the standard benchmarks. So that's really interesting.
S
S.13:59
How is that going to get resolved? In classical machine learning applications, there's some point of view on benchmarks and standards. There are different products out there that can actually measure these things around bias and auditing. As we massively blow up the scope around creativity, all of that kind of shifts. So how do you think this is going to resolve?
P
Percy Liang14:23
Yeah, so first order, scale definitely is helping, so we're safe on that. If you scale up the models, it lifts all boats. And then given a particular scale, you have a question of where you're investing your resources. I think what we want to do is develop effective surrogate metrics which you can actually evaluate that correlate well with human interaction. We don't really have a good handle on this quite yet. But having humans in the loop for an inner loop is also potentially problematic and hard and not reproducible. So you want something that's easy to evaluate but at the same time actually tracks what you care about.
S
S.15:09
I want to shift to building products and companies around large-scale models. David, maybe I'll start with you. There are people in the audience who are in the early stages of building these companies, and one fundamental question is, do you go build on top of an OpenAI API? Do you go build on something in the open source? Do you go build your own large model? How should a founder navigate making that decision?
D
David Luan15:27
I think probably the biggest question for people to ask right now, the root thing that is worth answering first, is: What is the loop you're going to run for your company to compound? Is it going to be oriented towards really deeply understanding a particular customer use case? Is it going to be oriented towards some sort of data flywheel that you're trying to build? Thinking about how that interfaces with the differentiation that you want to have as a business is going to be really key, because I don't think we want to live in a world where effectively these companies become sort of outsourced customer discovery engines, and then new Amazon Basic versions of these things come out over time. That would not be a particularly good world to live in. So figuring out what that compounding looks like is the most important first step. The other thing to think about here is just how many nines do you need? If you need a lot of nines of reliability, one thing that's really difficult is you lack all the affordances that you could possibly want if you are consuming this through an intermediary to get you to where you need to be with your customers. So because of those different reasons, you could end up choosing a very different point in space for how you ultimately consume these services.
P
Percy Liang16:44
Maybe just add one thing: the nice thing about having these APIs is it's extremely easy to get started and try something. You can sit down in an afternoon, punch in some data, and get a sense of the possibilities. In some cases, it's a lower bound on how well you can do, because you spend an afternoon, and if you invest more, if you fine-tune and build custom things, it can only get better. So I think that has opened up the challenge to even formulate what is the right problem to go after. Typically, you don't know, and you have to collect data and then train a model, and that loop becomes very expensive. But you could just sit down in an afternoon, try a few things, and maybe few-shot your way to something that's actually reasonable. That gets you into a different part of the space, and you can iterate much faster.
S
S.17:43
That makes a lot of sense. In terms of prototyping quickly and trying to take out product-market fit risk, one question becomes, and I'm curious for your take on this: if you start that way, how do you over time build durability into your product? Because I could make the argument that maybe you're just a thin layer on top of someone else's API. You can quickly de-risk product-market fit, but is there real durability in your layer of the stack?
P
Percy Liang18:05
Yeah, I think the transition out of API is a very discrete one in some sense. People also do like human wizard-of-Oz experiments. You put a human there and have the human do it, and then you see work out all the interface issues and whether this makes sense at all, and then you try to put something else, take the human out. Now you could put an API there and get a sense of what things are like. And then in some cases, maybe few-shot learning is actually for some things not that strong. If you have data, maybe a fine-tuned T5 model or something much smaller can actually be effective. I don't think the last thing on your mind should be to go pre-train a 500-billion-parameter model when you don't know what application you're building.
S
S.19:00
Maybe continuing on the theme of building on top of these models, despite the magical qualities of these things, there's still limitations. One of the limitations is falsehoods, and there are others that developers need to navigate as they think about building these applications. David, maybe starting with you, what do you think some of the key limitations are, and how do you guide people around navigating those?
D
David Luan19:19
That's a really good question. Falsehoods are definitely a very interesting thing to go talk about. These models love to be massive hallucination engines, so getting them to stick to script can be quite difficult. I think in the research community, we're all aware of a bunch of different techniques for improving that, from things like learning from human feedback to potentially augmenting these models with retrieval. I do think that on the topic of falsehoods in particular, this idea of packing all of the world's facts into the parameters of a model is pretty inefficient and somewhat wasteful, especially when some of those facts change over time, like who may be running a particular country in a particular moment. So I think it's pretty unlikely that's going to be the terminal state for a lot of these things. I'm really excited for a lot of the research that will happen to improve that. But I think the other part goes back to a question of practicality and HCI, which is that you have a sense of every year we're pushing fundamental advancements on these models, they get somewhat better on a wide variety of different tasks that already show receptivity to scale and more training data examples. But how do you surface this wave where the particular capabilities you're looking for from the model are good enough to be deployed, where you can learn how to get from there to the finish line, and how do you work around some of these limitations in the actual interface to these models such that it doesn't become a problem for your users? I think that's actually a really fascinating problem.
P
Percy Liang20:53
Yeah, I mean, these models are exciting, and the flip side is that they have a ton of weaknesses. Falsehoods, generating things that are not true, biases, stereotypes—basically all the good and bad and ugly of the internet gets put into them. I think it's actually much more nuanced than just 'let's remove all the incorrect facts and debias these models.' Their efforts on filtering can filter out offensive language, but then you might end up marginalizing certain populations. And what is truth? We like to think there's the truth, but there's actually a lot of text which is just opinions, and there's different viewpoints, and a lot of it's not even falsifiable. So you can't talk about truth without falsifiability. In a lot of cases, for example creative cases, you do want things which are a bit more edgy, like how you create fiction. If everything has to be true, there's no easy way to throw out all the bad stuff. So I think one good framework to think about is that there's no one way to make these models obviously much better. I think what you want is control and documentation. You want to be able to understand, given a model, what is it capable of, what should it be used for, and what it should not be used for. This is tricky because these models have a huge capability surface. You get a prompt, you can put any string and get any other string back. So what are the guarantees? I think as a community, we need to develop a better language for thinking about what are the possible inputs and outputs, what are the contracts like the things you have in traditional APIs and good old-fashioned software engineering. We need to import some of that so that downstream application developers can look at it and say, 'Oh yeah, okay, I'll use this model, not that one, for my particular use case.'
S
S.23:04
There's another side to this which, for lack of a better word, I'll use the word 'risk.' If I'm a product builder and I'm building these models and building different products, there's different levels of risk I might be willing to take in terms of what I'm willing to ship to end users and the level of guardrails I need in place before I'm willing to ship. How do you guys think about that and frameworks around that?
D
David Luan23:31
So I think there's some interesting perspectives here specifically around just how do you sequence out the different applications that you want to go after. I feel like one really nice property of these models is it's just so easy to go from zero to a hand-wavy 80% quality on a wide variety of tasks. For some of those tasks, that's all you need, and sometimes the iteration loop of having that 80% thing with humans is all you need to do to go run it over the finish line. I feel like right now, the biggest opportunity we all have is starting out by addressing things like that. But I think that over time, one of the things that Kevin said that I really liked is that there will be more of a standardized toolkit for how to erase some of the lower-hanging fruit risks related to generations that are inappropriate or models going off the rails in various different ways. I think there's another set of risks that are slightly longer term that are also really important to think about, and those are definitely much harder.
P
Percy Liang24:31
Yeah, to build on top of that, I think there's another category of risk which is adversaries in the world. Whenever you have a product that's getting enough traction, there's probably people who want to mess with you. One example is data poisoning. This side hasn't really been born out yet; there's some papers on it. But if you think about it, these models are trained on the entire web crawl, so anyone can go put up a web page or put something on GitHub, and that can enter the training data. And these things are actually pretty hard to detect. So from a security point of view, this is a huge gaping hole in your system. A determined attacker could probably figure out a way to screw over your system. And the other thing to think about is misuse. All powerful models like these are dual-use technologies. There's a lot of immense good that you can do with them, but they can also be used for fraud, disinformation, spam, and all the things that we already know exist, but now amplified. That's a scary thought.
S
S.25:46
Yeah, definitely a lot of asymmetric capabilities here. Just hooking up one of these giant code models to some RL agent to get into systems, there's so many things that become way easier for malicious actors to do as a result. Attack is so much easier than defense.
A
Audience Member26:11
This is a question for Percy. From an academic standpoint, what is an ideal wish list that you might have in ways that corporations and big companies that are building a lot of the ecosystem could help?
P
Percy Liang26:27
Yeah, I mean, I think one big thing is openness and transparency, which is something I think is really sorely missing today. If you look at the deep learning revolution, what's been great about it is that it's benefited from having an open ecosystem with toolkits like TensorFlow, PyTorch, datasets that are online, tutorials, and people can just download things, tinker, and play with it. It's much more accessible. And now we have models which are behind APIs with fees, and things have become such that you can't really tinker as much with it. So I think that. Also, a lot of organizations are addressing the same issues around safety and misuse, but there's not really an agreement on what the best practices are. So I think what would be useful is to develop community norms around what is safe or what are the best practices for mitigating some of these risks. In order to do that, there has to be some level of openness so that when a model is deployed, you get a sense of what these models are capable of, and benchmark them and document them in a way that the community knows how to respond, as opposed to 'Here is a thing you can play with. It might shoot you in the foot or not. Good luck.'
S
S.27:58
We have time for one more question. So maybe I'll ask a final question which goes back to creativity. I think one of the important things to inspire in all of us is what's going to be possible, the magic of these models. I think DALL-E was an important moment for people to see the type of creative generation that's possible. For each of you, what are some of the things that you think are going to be possible in the way we interact with these models in a few years that you're most excited about? Maybe starting with you, David.
D
David Luan28:28
I think language, as we were talking about earlier, is just the tip of the iceberg. We're already seeing amazing things just with language, but I think when we start seeing foundation models or even bundle together foundation models that are multimodal for every domain of human knowledge, every type of input to a system or to ourselves that we care about, I just think we're going to end up with some truly incredible outcomes with that. If I were to choose a personal thing that I'm not working on that I think would be really cool, I once talked to someone about this: what happens when you start having foundation models for robots? When you can take all of these demonstrations, all of the different trajectories of robots interfacing with the real world, and put them all into one particular system, and then have them have the same type of generality we've seen with language? I think that would be incredible.
P
Percy Liang29:24
Yeah, I agree with that and have definitely thoughts, but maybe I'll end with a different perspective, another example. If you think about it, we get excited about generating an image, but it's just an image. It's also something that you imagine an artist be able to do. And if you think about humans aren't changing that much; computers are. A year ago or two years ago, we weren't able to do that. Now we can. So if you extrapolate, now image is only just so small. Think about videos, or 3D scenes, or immersive experiences. Maybe with personas, you could imagine generating worlds in a sense. That's scary but also sort of exciting. There could be many possibilities there. But I think the bigness of things that you can create—if you think about these models as excellent creators of large objects now, and you think about what are big objects—well, there are environments in some sense. What if we could do that? What would that look like? What kind of applications could that unlock that would be interesting to think about?
S
S.30:49
Yeah, there's a commonality there which again connects back to multiple modalities and continuing to push scope. It's a really exciting glimpse of the future. Percy, David, thank you guys so much for doing this.
D
David Luan31:00
Thank you.
H
Heather Mack31:07
That concludes this episode of Gray Matter. Like what you hear? We encourage you to rate and review Gray Matter on your favorite podcast platform. We sincerely appreciate your feedback. You can also find all content on our website at greylock.com. I'm Heather Mack. Thanks for listening.