Back
Matei Zaharia
Cofounder and CTO, Databricks

Business AI in 2026: What’s Working, What’s Not, and What’s Coming (w/ Databricks CTO Matei Zaharia)

🎥 Jan 19, 2026 📺 Josue Bogran Channel ⏱ 40m 👁 504 views
‪@Databricks‬ Cofounder and CTO Matei Zaharia joins me to cut through the noise around AI adoption. We cover what’s working, what’s not, and how executives should approach AI in 2026, from ROI and workflow design to mindset shifts, model limitations, and yes, Agentic AI. Chapters 00:00 – Intro 01:12 – Are current LLMs good enough to deliver ROI? 02:33 – What happens if no new models are released? 04:54 – Can AI replace employees entirely? 05:07 – What can AI do reliably today? 06:53 – What’s are the biggest non-technical reasons AI efforts fail? 08:08 – How can a CEO or CTO tell if AI is a va...
Watch on YouTube

About Matei Zaharia

Matei Zaharia, cofounder and CTO of Databricks, introduced Omnigent, an open-source meta-harness for AI agents, at the company's Data + AI Summit in June 2025. He described Omnigent as a new layer above existing agent harnesses that aims to make agents interoperable, allowing sessions, policies, and skills to follow users regardless of where the agent or model runs. Zaharia said the project was released as open source because the company believes the meta-harness layer will benefit from being open in the same way data formats and sharing protocols have. He also demonstrated features including contextual security policies that vary permissions based on prior agent actions, cost controls, and an OS sandbox. In a separate appearance, Zaharia discussed Databricks' broader strategy, stating that many traditional software products will be rewritten under a paradigm of "get the data to be there and then let's slap some AGI on top." He attributed Databricks' differentiation to its commitment to open formats and its focus on AI. He also commented on the company's LTAP (Lakehouse Transactional and Analytical Processing) offering, describing it as providing the benefits of the HTAP (Hybrid Transactional/Analytical Processing) holy grail by making data available immediately for both reasoning and analytics workloads.

Source: AI-verified profile updated from Matei Zaharia's recent appearances. Browse all interviews →

Transcript (48 segments)
M
Matei Zaharia0:00
As I said, at minimum, LLMs are really great at basically search and approximate matching. So, at minimum, it'll help you surface that information and then you do something with it. Maybe in the extreme, it can actually, with high probability, just give you the right answer or give you a comprehensive report that tells you this is a little bit unclear or whatever it is. But people, you know, there's so much like people just see the sort of sci-fi version of AI, which, you know, they ask it to do things that it can't do.
I
Interviewer0:36
So, I'm here with Mate, who is one of the co-founders at Databricks, CTO, correct?
M
Matei Zaharia0:43
Mhm. Yeah.
I
Interviewer0:44
Like you have like you wear so many hats, I don't even know what your title is.
M
Matei Zaharia0:48
No, that's it. Yeah.
I
Interviewer0:49
Yeah. Well, thank you so very much for taking the time to meet with me and do this video with me. So, I'm very excited. So, lots of fun, very business-driven questions today with a few that are probably going to be a little bit more on the technical side, but let's start with the business ones, hard questions. Is the current generation of LLMs good enough to drive real ROI?
M
Matei Zaharia1:21
Yeah, it's a great question. Yeah, I think there's a lot of hype around it for sure. So clearly, you know, you'll get all kinds of answers. I do think it is useful for certain things and we're figuring out over time how to make it better for more of them. But not for everything. And I think one of the challenging things about it is that it's so easy to get a demo that works okay, you know, 50% of the time or works in a few examples, and then you think you solved the problem but actually getting it to be like highly reliable and diligent is not for sure. Like some things that it's good at and they're being used are any kind of search basically, trying to, you know, surface information, put it together, it helps you find things, but usually you have a human in the loop who looks at all the results. Coding agents, again, depends how you use them. It's not going to code a complete application for you that's like flawless and works, but it can help you navigate a lot of code, prototype things, do a small amount of code, things like that. And then there are others, you know, there sort of others that are exciting but maybe harder to get going.
I
Interviewer2:33
I'm curious, what do you think would happen to just the drive for AI if all of a sudden the Anthropics, the OpenAI, the Googles, the Metas stop making any new models? Do you think we would still have over time a net positive effect with what we have today?
M
Matei Zaharia2:57
Yeah, you know, I do think we can figure out how to use current models better. Actually, in my perspective, models haven't gotten that much better say in the past year compared to like the year before or the year before that. So they are in some sense slowing down and you could do similar things with today's models that you could like a year ago. Maybe it's cheaper. They've gotten faster and cheaper for the same quality, but that's sort of what it's like. So these models are not perfect. They make mistakes and they, you know, they can fall over, especially if you give them a lot of context for example, they'll just like stop working well. So we have to figure out, you know, what to build with them that's useful. So I would say a lot of the advances in actually applying these models in a way that's useful have been partly like user interface, like HCI advances. Like just as an example, everyone knows when ChatGPT came out, people were excited. But before that, OpenAI had something called InstructGPT which would follow instructions. The difference is you send that one message and it sent you one response, you couldn't send a follow-up message. So in some sense like ChatGPT was a user interface innovation where like when the thing misunderstands your question or makes a mistake, you can tell it like are you sure that's right or, you know, think about this and then it like follows up. So and that chat form factor worked really well and same thing with the coding agents, like even the initial like tab completion ones, right? You press tab, it shows a few options and then, you know, some of them are bad but like you pick the one that you like and continue. That's, you know, that's a UI that made it good for that particular thing. So I think we'll figure out, I mean both the techniques for like building agents but also how to, you know, get them to interact with people. And some, yeah, now what kind of things will it like, can it replace like all employees in like some company or some department? Probably not, at least not if the models stay at the current stage. But there are a bunch of things that it can accelerate and make better. Actually, the other thing I'll mention is there are things now that like people don't look at at all. For example, you just can't have humans like read vast numbers of documents that are posted every month in the financial markets or things like that and maybe you can have an LLM look at it. Even if it's not perfect, it will give you some summaries or some insights that are useful. That's actually one of the biggest use cases we're seeing is this sort of batch data analysis where like no one would look at this unstructured data. Now you do something with it and, you know, that it's not perfect but like you take that into account in what you do downstream.
I
Interviewer5:44
Yeah. That kind of ties into one of the things that I do tell people when they tell me like well what should we use AI for? And I tell them if there's something that anything would be better than nothing. You say, for example, let's say you have a thousand leads that come in and you cannot afford to have the staff to look through them to see what's qualified and what's not. If an LLM, an AI, a GenAI, you know, whatever you want to use can qualify all those leads and maybe 20% of them are wrong. That's a massive improvement for you.
M
Matei Zaharia6:25
Yeah. Yeah. And then this is also where I think the user interface things. Like for example, what if it gives you those and then you could in a chat you could send a follow-up like you could say, 'Oh, give me more that have this use case which I'm seeing popping up this week or whatever,' right? Like that would be an HCI sort of thing where suddenly like this is useful. You never as a user you don't have to look at all 10,000 leads, but you can ask some questions about them and get the ones you want.
I
Interviewer6:54
What is the biggest non-tech reason you see companies fail with AI?
M
Matei Zaharia7:01
One of them is like not being clear about like what outcome they want or maybe how to evaluate that outcome. That's one of them or like, you know, you build something but you can't tell how good it is. So, I'd say it's not necessarily a technical reason because if you knew how to evaluate, you might be able to build a better agent or whatever for the thing you want, but you just don't know like for example, you don't know that it's going wrong in a certain set of questions. And you just like deploy it and then it doesn't work and people don't like it. Yeah, I mean maybe even higher level is like the wrong expectation about what it's going to do. Although I think people are learning more now about what's possible and what isn't. But I would say okay if you do kind of know like what AI could do, what's practical to build today and you try to do something, you invest, it still goes wrong. Maybe it's because you didn't have a good plan for evaluating like either during development or even once it's deployed. Can you tell how it's doing?
I
Interviewer8:08
How can a CEO, a CTO, any executive actually tell if an AI project is a value driver or just an experimental cost center?
M
Matei Zaharia8:29
Yeah. You know, the way I think about that is you have to ask to close that loop of feedback from users or from whatever the use case is as quickly as possible. And I think it's easier than ever now to build prototypes of things. And with especially with generative AI like your prototype of an agent could just be a prompt or it could be you just tell it like go do this and maybe it's like really slow, maybe it takes a lot of steps but you can see, you know, does it have legs like is it likely to do this task or not. And then you can get people to try it and I would go even further and say like as an exec like you should try it, right? You should like sort of be critical, like, you know, try using it for a bit, see what happens. Obviously may not be possible for everything. But that's what I would do, having the tight feedback loop. If you're starting with saying well we need to first like staff up the team with like 20 people then we need to like pay to build a data set then we need to like spend three months investigating the framework we'll use to build it and like all that, you already invested a huge cost when you know maybe there's a very quick way to prototype to see is it possible, like what is it good at?
I
Interviewer9:46
Makes me think of like the Basecamp model for projects where it's like, hey, you're taking bets, you know, six-week bets, give it a try. At the end of the day, if after that six weeks, you see like, hey, this is not going to go anywhere or it's going to end up taking more and more time. Then just call your losses, but then you make other bets that perhaps, you know, if one out of five, you know, hit the right spot, then you might be able to make good ROI on that. I think about it.
M
Matei Zaharia10:17
Yeah. And another thing you can do here, it's like prototyping any software. Let's say there's a piece, you know, you think could be useful, but you don't have it yet. Like let's say if we collected or computed like a new data set in some way that would be useful. But you know, doesn't mean you need to build like a huge pipeline and a big like serving system and stuff for that. You can try running an agent where like you just handcode that information. You tell it, you know, here's this stuff and like see what it's going to do. And see if it would help to build that. We do that a lot here when we think about like what we want the assistant in the product to see. For example, like we can prototype stuff without fully building it out and then see, oh yeah, if we provided this type of, you know, tool to it or whatever, it would do better or it wouldn't do better. Yeah. So anything like that that can help you close the loop to get feedback, quick iteration, quick feedback loop. Yeah. Like sort of look for like what's a way you can tell like could it work, you know, at all and then if it can you invest in making that like more scalable, more solid, faster, whatever it is.
I
Interviewer11:30
Mind shift for how historically business has been done but, you know, you try to plan as much in advance, you gather up all your resources and you sign up for these big projects, right?
M
Matei Zaharia11:42
Yeah, yeah, I mean the other way you could do it, I think that the traditional way, like we kind of know if you know that okay if we invest this much time in a thing we can build it, then it makes sense to like perhaps go all the way ahead. But I think here with a lot of what people are trying to do with AI it's still pretty experimental, you don't know what's going to work. There might be a few things that are getting into that realm where like for example let's say you build some question-answering agent or something, you know, and you say well if I gave it this one, this other API to like, you know, look at customer whatever like transactions in real time or something, like could it answer questions better? Like, you know, there are small changes like that where you're like pretty sure that yeah if I build this it will do it and you can just add them in. But if it's a very unclear, like you haven't seen many people build this type of application, then I think you need to be in more of that experimentation mode.
I
Interviewer12:41
I know we've talked a lot about a little bit on the cautious side. Now, let's switch a little bit for a hot second to the art of what's possible. So, if again if an executive asked you which workflows in my company are most likely to benefit from Agentic AI, how would you help them figure it out? We kind of touched on it a little bit earlier, but I'm curious if you want to expand a little bit on that.
M
Matei Zaharia13:14
So obviously it's going to depend a lot on the company. I would say for sure anything where people need to like find information or even synthesize and summarize information, but there's a human in the loop working on stuff, I think it's very likely that you can help them do that and accelerate it. And that's I'd say at least like half the use cases we see. And even when I use stuff internally here, it tends to be that it's like, you know, I have some question about what's happening in the company and there's information across documents, tables, maybe many different data stores, maybe there's a little bit of analysis I want to run on it, and, you know, but it's all out there and like if I had to do it, I would just need to spend a lot of time finding each thing and putting it together. So that type of thing I think for sure you can do. And, you know, so like for me for example as CTO here, like I don't know what are examples of questions I asked. I mean one, you know, one example would be like for our Agent Bricks product, like what are the top, you know, concerns in like recent customer interviews? So it's about looking at all those customer interview notes or which customers, you know, talked about this, like use the latest version of like vector search or something, right? That's like joining a table with which has the actual events and like is the person using this or not with this unstructured data essentially in the document, like let's say they complain about latency but they didn't try the low latency version of like vector search or something, like that's kind of what I want to find. So these type of things are good but it can apply to anyone. Like for example we also have, you know, when there's any issue with a service here, like the engineering team is paged and normally as an engineer you get a big, you know, you get a whole bunch of dashboards you can look at and logs and things to figure out what is going on. So the team that runs those also built some little agents that, you know, while when the incident is happening the agent will go and run and sort of try to send you anything it found that looks unusual or you can ask it questions. So it's something like this that's good. That's the easiest one. There are cases where maybe you can automate more of a process if there are some standard decisions. I think as I said before one of the biggest ones is if you have like, you know, unstructured data like text or images, documents that like a person wouldn't even look at, for sure you can use AI and get a lot of value from that. And we've actually seen this with some companies that are, you know, who a big part of their business is essentially analyzing some data sets and they had very custom pipelines for it, you know, NLP models, regular expressions, whatever it is, entity matching, you name it. And some of them are switching to LLMs and finding that it's, you know, more flexible and you can get it to give good results. So, yeah, these are some of the ones. The things that I'd say the AI can do well is on its own, like diligently doing a task and getting lots of things right like for, you know, over a long duration and just giving you the right answer at the end. That's pretty tough. So it can do that very well. If you have a human in the loop who can verify stuff, then that's good or and who can even correct them by sending a follow-up message, then it works. There are some rare cases where you can have an automated verifier and that's where I think actually AI is most powerful today's AI. So for example like optimize this assembly code for like matrix multiplication and, you know, try to minimize the time it takes to run this benchmark while passing all the unit tests, like that's people have done that and they found like new, you know, new micro sort of algorithms and stuff for that using AI because it doesn't have to be right every time it just has to make like thousands of suggestions and then one of them might be better. Yeah. By the way, to capture some of these things if it's interesting, one of the things we're putting out is we have this sort of frontier AI benchmark which is not getting it to do like math olympiads or like, you know, programming contests or like weird obscure humanities last exam questions with obscure facts. We call it OfficeQA. All it is is you have a bunch of documents like PDFs from the US Treasury and you try to answer some questions that involve looking at a bunch of data like, you know, compare the interest rates in different years or whatever and tell me like what the, you know, peak years were in this time period or whatever. And so we specifically designed it so it's things that don't require like advanced math or world knowledge, finance, anything like that. They just require reading documents like a high school student could do them. But they do require diligence and, you know, like focusing on it and getting all the steps right. And this is one where like current agents and models get maybe like 30, 40% at best on this.
I
Interviewer18:37
When should a company not use LLMs or AI or Agentic AI even if they technically can?
M
Matei Zaharia18:48
I think that there's a couple of times when. One is if you want it to be very sort of deterministic and predictable. So you always, you know, get the same answer, that's pretty hard to do with LLMs and with AI today. And it's going to change anyway when you change your agent even if you get it to run that way today. And then I think the other one is for cost really, like you have to see if there is, you know, an easier way to do it, like both the monitoring cost of running it, latency, you know, yeah, maybe the other one would be if it's an adversarial setting where like the stakes are quite high of getting it wrong, you know, if someone can sort of trick the agent or whatever. So, because that's another like the weaknesses of LLMs today, you know, they sometimes get things wrong. They're not deterministic and they're all attackable through adversarial inputs. If someone really wants to get it to like, you know, say a bad word or whatever like or you can or worse. Yeah. Or worse. Yeah. So, yeah. And it is happening sometime that people prototype a thing with agents and then they realize actually I could have done something else. Yeah.
I
Interviewer20:02
Yeah. How are you finding that transition from a period of time where high levels of determinism, you wanted that 100% accuracy on everything, to now moving to that more probabilistic mindset which, which honestly I've seen a lot of businesses adopt that they're like hey they're okay because in reality day-to-day good enough sometimes is good enough for other but how do you find that as a technical person likely used to trying to make things as deterministic as possible. How do you find that?
M
Matei Zaharia20:36
It is a different mindset. I think it requires, you know, like you need to sort of think about evaluation a lot harder because normally with just like software engineering you write some tests and you know when they pass like it's likely to keep working and here the results might be a little bit noisy even on the test which you think it's going to get right and you really care about like the data distribution, you know, it's not like well strings that are like the empty string needs to work and like the string that's the maximum field size needs to work and then everything in between is going to work, like here it could be that for many inputs the thing works but if it's like, you know, if it's in French or something it doesn't work. So it's a much larger distribution of cases to try and you need sort of statistical look at it. But the other thing that I think many people don't really acknowledge is the applications we're using AI for, many of them are things we couldn't even specify or implement with standard software. It's kind of like for instance, let's say I want to make a great customer support agent. I mean to make it work great, I need to think a little bit about like psychology and like what is important for my business and so on. It's not like making a great sorting algorithm or something where like it sorts the list and, you know, it should be fast and it shouldn't use too much memory. There's a lot more context you need to think about and ultimately it needs to be sort of tested with people and see if it meets that outcome. So it's also a different kind of thing. It's almost like part of it is like, you know, kind of like imagine that the models actually were perfectly following instructions and stuff, doesn't mean everyone could build a customer support agent in the same way that a company might hire great people on their customer support team and still have a terrible experience because of the policies they tell those people to follow. So it's a different kind of application. So yeah, at the very least I think you need to take this more data-sciency, probabilistic sort of mindset about things. But even going beyond that like you need to think more about like what's this actually trying to accomplish and, you know, what type of like feedback, what cases do I worry about?
I
Interviewer22:57
Yeah. And the funny thing is I personally struggled a lot with that idea of like hey after so many years chasing after that one bug, that one number that keeps coming up because of XYZ reason, now we're okay with 10% mistake, it's okay. But then I also look back at even my times as an analyst whereas good enough. Was good enough. Like close, you know, being able to sort of hit the mark and being directionally correct was often very much of an okay thing.
M
Matei Zaharia23:34
Yeah. Yeah. I mean, there's definitely an argument that even the structured data you have is maybe incomplete or like wrong or whatever. So even then you need to step back and sort of, you know, like think a little bit critically about what is it giving you. But it's different, right? If you spend all your time on doing the computations with that data versus now trying to manage this other thing. And, you know, the final thing I'll say is I think again the user experience does matter a lot like can you, do you give users a way to do follow-up questions, you know, do you give them links to the sources, like can they dive in, these things matter. Actually funny story like with Genie which is our like, you know, basically like chat interface for BI, I was talking to our chief revenue officer and he was saying like hey, you know, customers want to trust Genie, like they don't know like can I trust Genie and I was asking well what's the error rate like what is it getting on what kind of question is it getting on and then he said that's not what they want, what they want is like when it gives me an answer can it show me all the steps and like let me follow up and so that as a user I can make up my mind that yeah, it computed the right thing. So, and again, it's basically it's an HCI. It's kind of like imagine you hired like the best data scientists in the world, right? You're the head of sales and then you ask like what's my churn forecast next quarter and they're like 50% of your customers are going to churn or 50.5%. And like they're like the smartest data scientists, I'm going to look for a new job. No, like no, you probably dig in and try to ask like, wait, how did you get that? What assumptions are you making? Maybe they're right, maybe they're not. But that's the kind of thing. So, that's why I'm saying like the user interface is a lot. Viewing this as just, you know, standard software, you fill in something in a form, you press a button, like a little answer pops up is too limiting I think for the kinds of both for the kinds of questions people really wanted to do and for the technology itself which is erroneous and not perfect.
I
Interviewer25:51
My concern when it comes to like the whole probabilistic approach is that it only takes one bad, one very bad miss for someone to be very pushed off. So like if I was from like the Databricks perspective, right? Building a software and there's this expectation that it's going to be very accurate, right? Especially like with like Genie that you can train it up to be better and better, but it's still not fully deterministic. So I think sometimes that is still a concern around financial documents, you know, financials and that is hard especially when you think any regulatory stuff and even Ali on stage in a Databricks Summit in 2024 he kind of called that out as one of the risks like hey you would still want the human to review this before you're submitting it as an official filing or something like that.
M
Matei Zaharia26:54
No, absolutely. But that's why I'm saying part of it is like what expectations you set for the AI product. And part of it is the interface like when it's showing me stuff, you know, is it giving me evidence? Is it saying, 'Well, there's actually two interpretations. I found two data sources that disagree' or whatever. And yeah, maybe people have set too many, well people have definitely set too many lofty expectations, like the AI companies are saying this is like PhD level, it's already like well past like a college grad. It's like it's going to do everything for you. But, you know, but clearly it isn't that. But, you know, you think about it, you know, like this way, right? Like if you're an executive at a financial firm and like you ask your data team to compute some result or whatever, it'll affect your trading strategy. Like they could also make a mistake. But you set things up so that you really minimize the chance of the wrong thing getting through. Like when you ask them, you kind of ask them follow-up questions. You look at the thing, they check with, you know, they have people review stuff. So, and that's kind of how it works. And you don't, and when you make a strong statement like I know we absolutely have to do this or that, like you're placing a lot of weight on it. So you check it even more. So part of it I think these tools can be extremely useful to help you make decisions. But they do need sort of the right interface. With Genie, one of the great things is that that data science team that would have had to know all the nuances and do all those things by hand can instead teach an agent and get a lot of the questions answered through that and then you only follow up when like it's really sort of hairy with them.
I
Interviewer28:45
Yeah. And yeah, the way that I've educated, you know, customers in conversations has been there's a risk things might be wrong no matter what you use with AI, but those mistakes could be acceptable. So, for example, I'd like to say, well, what if you have some solution that basically automatically ships out product, you know, the full order cycle and you accidentally ship out $2 million worth of product to the wrong place and that you lose out $2 million. Well, perhaps if you're saving $20 million a year and you're only losing, you know, $4 million from that, that might be an acceptable trade-off, but then you have to factor to the risk of like, well, is it going to make my customers upset? Am I going to lose my customers and all that? That's the way I think about it, too, because to your point is if you have the right expectations, then you won't be disappointed.
M
Matei Zaharia29:52
Yeah. Yeah. And exactly. And I also think people will adapt to these over time. But I think the best, like the upside, the really nice opportunity of these is for example in something like Genie, like basically it can transfer knowledge faster between people in the company essentially, like if someone on the data science team asked a bunch of questions and added them as examples or even if they didn't but just the fact that they asked them and they did like thumbs up on the answer or whatever and that means that the next person who asked it gets the right interpretation is good as much faster than you asking that person, you know, over like Slack or whatever. So there is, you know, I think at minimum as I said at minimum like LLMs are really great at basically search and approximate matching. So at minimum it'll help you surface that information and then you do something with it. Maybe in the extreme it can actually, with high probability, just give you the right answer or give you a comprehensive report that tells you this is a little bit unclear or whatever it is. People just see the sort of sci-fi version of AI which, you know, they ask it to do things that it can't do.
I
Interviewer31:09
What would you say to a team or leadership that hesitates adopting some AI capabilities?
M
Matei Zaharia31:22
Yeah, I mean I think it's important to look at that, you know, carefully and to take it seriously. And those exist for a reason. Now people have different interpretations maybe like some companies are very conservative, some are not. I would say at the very least like you should keep a pulse on what's possible, you know, what are you maybe missing out on by not trying to build those AI applications and also on the infrastructure you would need to use it in the future. Like as a simple example, well there's a few kinds of compliances, like one is we just don't want automated decision making for this thing, that one of course no, you know, any future AI like you probably still don't want it unless you change your mind about that. But then the other one is like oh we don't want to send data outside our network to like, you know, OpenAI or whatever or like we don't trust, you know, this model might have been trained on, you know, some kind of data that's wrong or whatever. So for those later ones, I think those are more surmountable in the future, whether it's through like open-source models that you host yourself, they're constantly getting better. And for like, you know, 90% of tasks that people do, they're like probably like pretty good, you know, already. Or even this like, oh, was it trained on like Reddit or something? I don't want that. Like there will be models that aren't trained on that or techniques to tune them to make them better. So for that stuff you just want to be ready in terms of like your infrastructure and business model. So for example I think like actually having solid data infrastructure is only going to help you in the future whether it's AI or other technologies that come out or simply the fact that like something can now be done digitally through APIs that couldn't be done before. You'll want to have accurate, timely, you know, data about what's happening in your company. And same thing with the business model like whatever, you know, data you have like maybe it becomes a competitive advantage in some form. In general like as any technology develops like data sort of becomes more useful probably, right?
I
Interviewer33:34
What should teams prepare to do in the next 12 months to really make the most of AI?
M
Matei Zaharia33:43
Any data that could be, you know, could be used in near real time by agents or like, you know, could power agents in some way by giving them like up-to-date information, getting that to be collected and computed and potentially served, you know, in near real time would be a big thing. Yeah.
I
Interviewer34:07
I know we've talked a lot about the risk and I appreciate it. By the way, I don't think there's many companies that sell AI that would be so willing, I think, to talk on the flip side of some of the risk and the trade-offs. So, I really appreciate that. But I also want to call out in a good way that you have Agent Bricks. A little bit ago, I wrote an article where I basically said, 'Hey, you have with LLMs, you have speed, quality, and cost.' Right now with LLM, you have the speed and the low cost. The quality obviously you need, there are different ways to overcome issues. But each one of those obviously has trade-offs that you might end up making things a little bit slower. You might end up making things a little bit more costly. And really the biggest cost being the time it takes really to figure out okay what should you be doing to counteract all the weaknesses and how do you even evaluate all the trade-offs. So with that as I was writing that article, I realized oh Agent Bricks actually abstracts a lot of those points. So I think it's a good point, good time for if you can share just basically given your awareness of the trade-offs, why Agent Bricks is a good solution to customers so they can actually go from we're afraid about using AI or we don't know where to start to actually beginning to get some wins that really they feel like they have the right speed, the right cost and more importantly the right quality and level of trust.
M
Matei Zaharia35:59
Yeah, totally. Yeah. I mean so the whole idea of Agent Bricks is there are some very common AI use cases out there like question answering about documents and tables, you know, structured data, like everyone has use cases for that or information extraction in batch. You got documents, you got text, you want to pull out structured information. So there's some very common use cases out there and for those use cases you as a company, you know, unless you are like primarily, you know, you want to be an AI development and algorithms company, you probably don't want to be developing your own pipeline. Like the last thing you want is you built like a RAG, you know, whatever like customer support bot but now you have to look at the latest like machine learning conference and there's like 100 papers on better ways to do RAG and you have to like evaluate all those and implement one and compare it with your previous one. So with Agent Bricks, it operates at this much higher level where like for these common use cases we give you a quick way of building a high quality agent for that task and then of tuning it. You can actually, you know, whether through prompting or just natural language feedback or examples you can tune it similar to how you tune Genie basically for your use case. And then our team does the research and development underneath and updates it. Like for example, Knowledge Assistant, the question answering agent that we have, since it came out we've changed the backend many times and we actually developed like, you know, some of our researchers are folks who worked on some of the key search products at Google for example. Like they revamped the way it does search and we revamped the way it does generation and citations and so on. And it just gets better and as a user if you deploy a thing on Knowledge Assistant, you know, it'll keep up with the quality and the cost of the best way of doing that. So that's what we're aiming to do there. It's sort of the next layer up above just LLMs and like vector stores and like the little building blocks you need to create an agent because, you know, you don't want to be doing that yourself. And of course you can only do it for kind of common things that many customers want. So we can create a team that does it and we know there'll be many customers using it but a lot of use cases are like that. The unstructured data use case, you know, parsing that, getting data out is hard. And we have, you know, a research team that basically built a large data set, trained models to do it, worked with a lot of customers, found like the top, you know, things that they're worried about, is constantly improving that. And that means a new customer can go in and use it and get better quality for like, you know, 10, 20 times lower price than like even the biggest LLMs for that task. And same with the Knowledge Assistant and Genie for text to SQL and so on. So it's a different point in the market. What we're finding is even the teams of like hardcore, you know, AI engineers who want to build agents, they actually like using Agent Bricks at least for sub-agents or tools or for a baseline. Like if I want my agent to consult, you know, documents and do RAG or consult some structured data, I don't want to build that piece from scratch. I can call into the Agent Bricks one and then I can work on whatever unique things my agent has, like it's going to, you know, control stuff inside my software or something that Databricks doesn't know about.
I
Interviewer39:37
And I know it's still in beta.
M
Matei Zaharia39:39
Mhm. Yep.
I
Interviewer39:40
Hopefully soon to, you know, to get out of there.
M
Matei Zaharia39:44
Okay. Got it. Yeah. I mean I think again for me, I think making that adoption easier going from like hey here's this technology but by the way you have to do a thousand things to be able to trust it to making it easy for people to have the right balance between building and still having some controls. We can always bring in like whatever is working best for that task.
I
Interviewer40:10
Yeah.
M
Matei Zaharia40:11
Yeah.
I
Interviewer40:11
That's awesome.
M
Matei Zaharia40:12
And you don't need to worry about it. You don't need to worry about like oh there's a new release I got to migrate my code just so my RAG thing keeps working so that future patches keep working.
I
Interviewer40:21
That's awesome. Well thank you so very much for your time, Mate.
M
Matei Zaharia40:24
Yeah, no worries, yeah.
I
Interviewer40:25
Thank you.