Back
Matei Zaharia
Cofounder and CTO, Databricks

From Lakehouse to Agent Cloud — Matei Zaharia and Reynold Xin of Databricks

🎥 Jun 12, 2025 📺 Latent Space ⏱ 70m 👁 7498 views
From open-sourcing the layer above coding agents to rethinking databases for the agent era, Databricks cofounders Matei Zaharia and Reynold Xin are pushing the company beyond the lakehouse into a full data-and-AI operating system. In this episode, Matei and Reynold join swyx after Data + AI Summit to unpack Omnigent, LTAP, Lakebase, agent security, open formats, Mosaic, and why databases may matter more than ever once AI agents start doing real work. We go deep on Omnigent: Databricks’ open-source meta-harness for combining, controlling, and sharing agents across Claude Code, Codex, Cursor, P...
Watch on YouTube

About Matei Zaharia

Matei Zaharia, cofounder and CTO of Databricks, introduced Omnigent, an open-source meta-harness for AI agents, at the company's Data + AI Summit in June 2025. He described Omnigent as a new layer above existing agent harnesses that aims to make agents interoperable, allowing sessions, policies, and skills to follow users regardless of where the agent or model runs. Zaharia said the project was released as open source because the company believes the meta-harness layer will benefit from being open in the same way data formats and sharing protocols have. He also demonstrated features including contextual security policies that vary permissions based on prior agent actions, cost controls, and an OS sandbox. In a separate appearance, Zaharia discussed Databricks' broader strategy, stating that many traditional software products will be rewritten under a paradigm of "get the data to be there and then let's slap some AGI on top." He attributed Databricks' differentiation to its commitment to open formats and its focus on AI. He also commented on the company's LTAP (Lakehouse Transactional and Analytical Processing) offering, describing it as providing the benefits of the HTAP (Hybrid Transactional/Analytical Processing) holy grail by making data available immediately for both reasoning and analytics workloads.

Source: AI-verified profile updated from Matei Zaharia's recent appearances. Browse all interviews →

Transcript (100 segments)
U
Unknown0:00
One of the advantages we have is that once you get the data in the right place, the AI models are becoming pretty good. The generic agents are fairly—I mean, Ali talked about AGI already here. They have pretty good reasoning capabilities. Actually, I think many of the traditional software will be sort of rewritten with this new paradigm, which is just get the data to be there and then let's slap some AGI on top. Magic will come out.
But without the right data, you can't really do that.
H
Host 10:28
Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering science and entertainment contents that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis, but fortunately enough of you actually subscribe to us to keep all this sustainable without ads and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you and it means absolutely everything to me and my team that works so hard to bring the content to you each and every week. If you do it, I promise you we'll never stop working to make this show even better. Now let's get into it.
Matei Zaharia from Databricks. Welcome to In Space.
M
Matei Zaharia1:20
Thanks for having us.
Yeah, thanks so much.
H
Host 11:22
Thanks for taking time out. You have your Databricks Data AI Summit going on. You were just telling me how the first summit that you guys ran was just 50 people.
M
Matei Zaharia1:30
Yeah, it was a little meet-up at Berkeley, I think. We put together—we did these tutorials and yeah, just teach people Spark.
H
Host 11:37
Yeah. You know, obviously now it's like—I think the headline number is like 100,000 people around the world, 30,000 in person. It's a crazy community. Well, I mean I just saw the keynote. Ali is just—did you know, was it obvious back then that Ali would be such a great CEO? Like he's a great presenter.
M
Matei Zaharia1:57
What do you think? I mean, I think among our group of founders, it was clear that I think he'd be the best at this. And yeah, it turned out great. I mean, he's ramped up on so many topics running a company. He would just go in and study it and talk to all the experts. Even if you can't hire the person, learn enough about finance and sales and whatever it was. And you know, go from there.
A
Ali Ghodsi2:23
Yeah. I mean, he's obviously very high IQ and has very high EQ, but Ali today is quite different from Ali from like 10 years ago. I think there's a lot of work that he put in to get to this point.
H
Host 22:34
Yeah. I mean, to me the most appealing thing about him is that he's funny. It's hard to make jokes about data, about serious topics and security and what have you.
A
Ali Ghodsi2:46
Oh, yeah. That's for sure.
H
Host 22:47
Yeah. So, you guys launched a whole bunch of things. I'll just name check briefly the stuff because we're not going to cover everything. Omnigents, your baby. El Tap, your baby. Your Dream Engine. We're also going to cover Genie, cover Customer League. You acquired Panther, Open Sharing, and there's Unity AI Gateway. A lot of these I think are things that you would expect a Databricks to do. It's like part of the road map. Everyone in your category has similar things. But I think probably the two of you are leading the two most unique and differentiated initiatives in the landscape. Maybe we'll start with Omnigents and then we'll go into it. A lot of people are exploring this sort of meta harness concept. What led you to it?
M
Matei Zaharia3:36
Yeah. There were actually a couple of converging lines, which I think is a good sign that you need something new. So, on the one hand, there's all the coding agent work internally. We have a really great dev infra team. They built something called Isaac that's basically a wrapper on Claude Code and Codex and lets you use them either on the web in sandboxes or just on your dev machine or on your laptop or whatever. And then they were adding all kinds of stuff there and we saw all the more advanced engineers building their own workflows with tons of agents and building their own UIs and stuff on top of that. And then the other one was that us building agents—we shipped this data science agent called Genie on the research team which I co-lead. We also build a lot of internal ones for various things and then we have all the customer ones and all of them were running into this thing of like, oh I need to switch model and harness and so on every few months. Plus the agent is completely useless if you can't share sessions with someone and have history and have search and all this layer on top of it for collaboration. I thought a bit about it from both contexts and at first people thought it was weird and like why are you doing coding agents and custom agents in the same thing but I said it's basically the same problems and you just want to build the stuff that lets you deliver the agent, maybe control it if you care about security, and make it portable across things. And then we prototyped some things as experiments. We said, yeah, actually we can make it work and then we sort of built it for real.
H
Host 25:20
I'm wondering if this kind of—let's call it architecture—maps to anything in your careers in the past. You know, like I always think about how a lot of things actually just tie back to operating systems.
M
Matei Zaharia5:32
A lot of operating systems tie back to databases, so or the other way around. So the thing I do think it ties a lot to is network protocols, you know, internet protocol. We also did stuff with data sharing also which is probably—most viewers probably won't know unless they—
A
Ali Ghodsi5:50
Open protocol is the term for sharing. Open sharing.
M
Matei Zaharia5:52
Open sharing. Yeah, so it's like you have a company, you maintain some kind of table—like let's say Walmart or something. They have the inventory and what's been sold in each store. And then you also have suppliers and they would love to produce more things and ship them exactly the moment you need them. So they would love real-time access to your table. So instead of sending emails around or Excel sheets or phone calls, why can't you share a view of that table in real-time with them? Then they query, they join it with their data and they decide what to send. So it's one of these things where you might ask—today since we can vibe code anything so fast, why do we even need to design protocols or APIs or software? Why can't you just vibe code things on demand? But actually for this type of interoperability where multiple parties that are moving at different speeds are building stuff and you still want some layer on top to coordinate, you do want to design it and build it. So it reminds me of that—agents talking to each other and users talking to agents and tools.
H
Host 26:56
Do we know of any other comments or alternative viewpoints?
A
Ali Ghodsi7:00
I think by the way, we had a debate on exactly the benefits of this. And I think around the time we decided to do this thing, I was telling Matei—it just happened to be there's a particular week that I was coding non-stop from the moment I woke up to the moment I went to bed. I was looking at my cloud sessions, my Codex sessions and one of the things particularly annoying was having to keep my laptop open. I was actually driving to a doctor's appointment and I remember because I wanted to make sure the whole thing continues working. By the way, it's so comforting to hear you say that because I'm like—I don't know if I'm a clown and I'm doing this or not. Honestly, I was driving and I was tethering my laptop to my phone, keeping it on the side, and whenever I hit a red light, I started looking at what's going on on my laptop. And I just felt that was ridiculous. It felt like we went back to the dark ages of programming. The productivity you gain from all this coding agent stuff is amazing, but—like, have you heard of cloud?
M
Matei Zaharia8:02
It was crazy to me. Was the thing we were working on the sandboxes or was this before that?
A
Ali Ghodsi8:06
It was the sandboxes. So you were—I was approaching from a very different angle. I wanted—hey, we're going to have cloud sandboxes that actually don't shut down. You can get one very quickly, but not just for running agentic sessions. It's actually also for running development. So I was personally building that that week and through building that, I ran into all these issues. And then I wrote actually a document for my case—here's my wish list of what the actual environment should do. And I think he actually ended up almost implementing every single one of them.
M
Matei Zaharia8:37
Yeah, I remember Reynold saying—because my first prototype of this had just chats with your agent—and he said, I have to be able to open a shell, like my own shell and list files and tail them and stuff. So I was—
A
Ali Ghodsi8:49
Is this an SSH into a mainframe?
M
Matei Zaharia8:51
Yeah, actually it has that.
A
Ali Ghodsi8:52
Tailing my logs. Yeah. And also, another thing I think I asked was—I still use Cursor for the sole purpose of rendering markdown files. Give me a way to see my markdown files and render them properly. I don't need a separate tool anymore. Yeah. I think you also built that in.
M
Matei Zaharia9:11
We did that. Yeah. We had a lot of engineers building their own vibe coding setup. But then the other thing they all said is—hey, I built something that's amazing for me, but no one else on the team can use that because I don't have a server to collaborate. And this is why we tried to set up Omnigents so you can have a server and have the security set up in there. So, you know, log in with Google or whatever and actually securely share stuff. And that's why we've seen a lot of other agents hit things like—people think they prototyped an awesome agent, but it's not allowed to connect to some really important data or whatever because of the security team. So, yeah.
H
Host 29:52
Yeah. At this point—so, for those watching along on YouTube, we're going to bring up an image of the structure here and we can talk through a little bit of the architecture. I think I just want to have people understand because when we're talking about software, it can be very abstract—and here's actually what we're talking about. You've worked out in open source this entire platform basically and there's a runner component and server component with a sort of uniform API that you figured out. Any other sort of element—and obviously you can plug in all these persistence layers and compute layers. This is a whole cloud. This is Agent Cloud.
M
Matei Zaharia10:25
Yeah, it's got these components to work with it. A lot of the action happens on the machine where you deploy your agent to. So, whatever you've got on there, you can run. But yeah, I think it's sort of the minimal thing you want to have to host collaborative agents and to have that server. And one of the reasons we open sourced it is anyone building agents, this gives them an app they can start with and customize, which we were seeing in Databricks too—someone would make a nice agent app and then other teams would ask, oh, can I just use yours for my agent?
A
Ali Ghodsi10:58
Yeah, I think we had like five or six different agentic frameworks built by every different team. They all do more or less the same thing.
M
Matei Zaharia11:04
Yeah, you need to basically—people want to take something that works and fork it, and you might as well have something open source.
A
Ali Ghodsi11:10
Yeah, which also was another question that's interesting for Databricks—what do you choose to open source, what do you choose to make proprietary? This goes back to Spark, right?
M
Matei Zaharia11:20
Yeah. One of the reasons to open source something is if you think it's a layer that will actually have some network effect. It'll benefit from many people collaborating on it. So, for example, with Spark—I don't know if you know—but when Spark came out, we also focused a lot on letting you have libraries on top. So they used to be different distributed computing engines for machine learning and graph computation. We said they should all be libraries that you can compose, and we made it super easy to add connectors to data sources too. And then we benefit because we don't have the time to write connectors to a thousand different databases and file formats. But we can just use the ones people make, and of course they benefit from joining this thing. So that's one way to think about it. Another way to think about it is—I can imagine our thing wasn't open. We had some kind of agent hosting thing, but it's not open, and then there's an open one. Which one's going to win in the long run? So here, because there is this benefit from people writing integrations, it'll be the open one. And then there are other things that you just can't even deliver as open source that are things the company does. Like, for example, how do you make sure your streaming jobs or your lake-based database doesn't lose all your data at night? Well, that requires an operational team that's going to sit there. There's no way—it has to be a service. So we want to make sure as a company we're really good at those infrastructure services, and then we're as open as we can in terms of what you build on top.
A
Ali Ghodsi12:55
I mean, speaking from the benefits, I think we're already seeing pull requests and ecosystem integration even though it was only released on Saturday.
M
Matei Zaharia13:03
Yeah, Saturday. Yeah, so someone—let's see what's going on. Yeah, you can look at the merge lines. I actually checked some metrics this morning about the—
A
Ali Ghodsi13:12
400 merges already?
M
Matei Zaharia13:14
Yeah. I think quite—I would guess around half are not from my team. But for example, someone added support for running it on Kubernetes. People added many cloud sandboxes. So this can launch a cloud sandbox and run your agent in there, which is great for sharing too because it's not on your laptop and someone's running code on there. So yeah, many startups have put those in and we expect to see more of them. We also have more agent harnesses already—Cursor, CLI, and Anti-Gravity also.
H
Host 213:47
Yeah. That's all beautiful. I feel like the last time this happened, there was the rise of the modern data stack. I don't know if it was that useful. I'm actually kind of curious in your postmortem. I think most people will agree that it is finally dead. But maybe this arises to a new modern AI stack that does the same thing. I don't know.
A
Ali Ghodsi14:07
I mean, I think the modern data stack was a pretty useful thing, probably even up until this day. I think—for the audience who don't actually understand the history—the modern data stack is effectively decomposed into: you need a layer to ingest the data in, you need a layer to transform your data, then all of this runs, and then you need a layer to maybe visualize your data. And all of this runs on some sort of data warehouse or later on as we're doing—data lakehouse. I think those concepts are all very powerful and very useful. It sort of enabled a lot of workloads. What people eventually run into is kind of a question of unification and consolidation. It's—hey, do you really need to chop all of this into different pieces and work with so many different vendors and platforms in order to get a very simple visualization done? Right? So I think over time everybody started realizing that customers are pushing us. We started to realize that, so we started building more and more capabilities and trying to consolidate. And at the end of the day now, customers don't have to worry about having to hook up five different systems in order to produce a chart. But I honestly think something like this is probably happening in how many different frameworks do you want to hook up together in order to do a very simple agent.
M
Matei Zaharia15:20
Just to be clear, I would say the core of this is this common API on top of all the harnesses. So the API is basically—you've got an agent session and you can send in a message or a file. Basically, that's what you can send in and then you get out these streams as it's streaming text or as it's doing tool calls. Or the other thing you can send in is you can tell it to cancel or turn. So that's the API. Now, the thing we did is we could get you that on top of Claude Code running in a terminal, Codex, Phi, OpenAI SDK, all that stuff. We map them all to that same interface. So that is something that you'd have to maintain yourself if you built your own agent orchestrator. And then whenever Claude changes its API, you got to tweak your thing or it's going to lose some messages. So that's the thing that's valuable to maintain. Then on top of that, we built a few apps. I think we built a pretty cool UI and stuff, but that's—and we built the security and control piece which I'm excited about. But it's that common interface. So we—it doesn't try to be a stack. And in fact, you could plug in your own UI on top of this server. That's one of the use cases we care a lot about because we want to use this in our own products.
H
Host 216:34
Yeah, it should be everywhere.
M
Matei Zaharia16:35
Yeah.
H
Host 216:36
I think one of those things that is really interesting to me is—well, first of all, I'll endeavor to do everything and not call it the modern AI stack because I think we have a name. But yes, one of the first people that told me about compute sandboxing was Nikita from Neon. Because a lot of people think about Neon as—well, it's serverless Postgres with the separation of compute and storage and instant branching and all those things. But actually, every database company is also a compute company. And so he was actually showing me his whole sandboxing solution. I don't think he ever launched it.
A
Ali Ghodsi17:10
So our sandbox solution—the reason we could have built it so quickly was because we realized if you just take the actual lakehouse architecture and remove the database from it. By the way, it was coming from you. Now there are some differences. For example, in the ones that support this particular workload, it's important to have local persistence. Because you want your state to persist. Your libraries—you don't have to install your library every time. Whereas the Neon architecture, because of the separation of storage from compute, you don't need persistent local disk. So there's some differences. But at the end of the day, yeah.
M
Matei Zaharia17:49
Yeah, so this is when you run like a coding sandbox. Like if I use the dev infra internally at Databricks, there's like many tens of gigabytes of data just for all the source code and artifacts and stuff that I built and I want that to come back next time. So but yeah.
H
Host 218:06
Before the show, we were talking about some statistics that might be surprising at the adoption. It could be internal, could be external, whatever comes to mind. Just to impress people the scale this is happening.
A
Ali Ghodsi18:15
So on the analytics side, I think we launch maybe 50 or 60 million virtual machines a day across our three clouds. So we're one of the biggest compute orchestrators out there. So that's for sure for CPU compute. And all of those process—I think—exabytes of data. I joked about depending on which time zone you are, typically before you have breakfast, Databricks would have processed exabytes of data already on that day. And on Neon it's actually pretty interesting—Kei, launching I think 13 million databases a day now.
M
Matei Zaharia18:48
Yeah, to me that was like a big—
A
Ali Ghodsi18:50
And that's just like—
H
Host 218:51
What do you mean?
A
Ali Ghodsi18:53
And then a lot of those were thanks to agents and branching experimentation. Because we made it so easy and so quick—and thanks a lot to Nikita's team—to launch databases. So it's changing the way people use databases.
H
Host 219:07
Yeah. Okay, we're going to go into more database talk in a bit, but I want to make sure we close up anything on Omnigents. You mentioned you're excited about the security and control side. A lot of companies are figuring that out right now, as well as the spend side. What have you found there?
M
Matei Zaharia19:24
Yeah, so I spent quite a bit of time talking to internal users, developers, security team, managers, and also lots of customers. And there's a few things. First of all, one thing that immediately became obvious for security—there's this tension between usability and security. And the way people do—a lot of coding agents today have very basic things like you can tell me which tool patterns are allowed or disallowed or whatever. It's like yes or no. But that puts you in a very tough spot. So, just as an example—should my agent be able to read some confidential documents? Or let's say should it be able to install new packages from NPM, which you know, maybe it's compromised? Yes or no. Maybe I want to allow it. Should my agent be able to publish stuff to the company website? Well, if I'm using it on the code on the website, yes. But should it be able to do both? So it can grab a confidential document and be prompt injected and leak it? Probably not. So the thing we decided we need is stateful or what we call contextual policies, where you keep track of the state of that session. It's not like is it allowed to push to the marketing site or not? But like, hey, if it did a risky thing—like it installed a one-day-old package from NPM, or it read a thousand confidential docs—then no. Then don't do it. Otherwise, maybe it's okay. That's one example of moving that trade-off. So it's both more secure and more useful by having a more powerful engine essentially. This requires tracking sessions. The other piece that was interesting there is—there are these very low-level events it's doing and you want some libraries on top that parse them. Like for example, we have an MCP server on Google Drive internally. It's got 60 API calls. Like how do I know which of those will share a document with stuff on the internet and which ones won't? It's annoying. So we designed in Omnigents the policy layer so that it's functions and you can have libraries. Like someone can make something that maps the low-level events to high-level ones and then you write a policy about the high-level things that came out. And that was related to the Panther—Panther will help with that. Panther is kind of a similar idea on the event processing side and it's Python-based versus a weird custom language. This is sort of more real-time. Those things are happening. Yeah. So yeah, but these are the cool things. I think the contextual or stateful part and then the way it can be libraries—and that was another reason to make it open source because others will write libraries and we and our customers can use them. And the final thing—because it's stateful, one of the states we track is how much you spent in that session. So I can—I've asked an agent to debug something and it spent $500 because it decided to read a lot of log files and burn a lot of tokens. But I can literally say, okay, launch a sub-agent to do this and cap it to spending $5. Like ask me for permission if it needs more. And because we're counting that within that session, it'll pop up and tell me—okay, you spent $5. Do you want to go on?
A
Ali Ghodsi22:40
So I'm more than context here. Matei spent the last five years—a lot of his time was architecting Unity Catalog at Databricks, which is the governance layer for data. And it's sort of combining expertise at that layer together with all the AI governance here.
M
Matei Zaharia22:48
That's right. Yeah. But I also spent a lot of time being annoyed by coding agents and getting bombs. And also as the CTO, I don't want to end up on the front page as—I installed some weird NPM package and leaked all the code. So I'm especially paranoid, but also I have very little time. So I don't want to sit there approving—do you want to run a 20-line bash script? Yes or no? So that's why I spend a lot of time figuring out—how can I make it as safe as possible and not annoying.
H
Host 223:24
Yeah. Is security a bigger concern than token maxing or token budgets? Which one is—
M
Matei Zaharia23:32
Oh, yeah, they're both there. I mean, I don't know. I guess it depends on the type of company you are. So I think some companies—the budget is limited and they really care about that.
A
Ali Ghodsi23:48
Or you can be Uber and still be concerned, you know.
M
Matei Zaharia23:50
Yeah, oh yeah, totally. For us, security is absolutely critical as a cloud provider. It's the most important thing. And token maxing, we're not so worried about it yet, but I've seen—like for example, I talked to some consulting companies. They have like 100,000 employees who are all coding for customers. If each of those spend like an extra $1,000 a month, that's not fun. We have like only a few thousand engineers.
H
Host 224:19
What's the policy in Databricks? Is it just unlimited or—
M
Matei Zaharia24:22
Unlimited, but we do—we use our own product to analyze the traces and stuff and we have a team that's looking to optimize and to see if anyone's doing something weird. And we actually had some really cool insights just from analyzing current traces—like which models are better at say Rust versus TypeScript or whatever. So yeah, at least in our code base.
H
Host 224:44
Yeah, amazing. Obviously, I have to ask the token maxing question. Obviously I think it's a key thing, but yes, security and control above that and figuring out a sane layer that you can have some autonomy, but not too much.
M
Matei Zaharia24:57
Yeah. Yeah, and we want to make it super easy. As an engineer, you should set the thing. So in Omnigents, you can ask your agent to set up a policy on itself to do this.
H
Host 225:05
If there's anything I should be showing—I don't see it on the GitHub, but you know, this is—
M
Matei Zaharia25:10
Put in the docs there. So you can look at it later. Just look in the docs on contextual policies if you want to see.
H
Host 225:17
I just like to show people—
M
Matei Zaharia25:20
Policies. Yeah, if you want to follow up on this, this is exactly where to look, right? Yeah. Yeah. And the story of these is—I just wrote—you know, I wrote a doc with like 10 ideas for things before as you were working on them. Well, that was like my wish list of things people asked. And I told the team—hey, can you do like at least five of these for the launch? And then they just got back with all of them. So, oh wow. So you can come up with more, but then some of them are just meant to be examples. Really, you can intercept any event the agent is making, and you can then either block or force it to ask the user or allow, and you can update state to keep track of stuff.
H
Host 226:00
Yeah, because you know, ultimately, I think of you as a systems designer—you let people plug in, right? That's the whole modus operandi of what you do.
M
Matei Zaharia26:07
Yeah. And we care a lot about composability. Like, can someone else write a library that others use? Which this is meant to—
H
Host 226:14
There's also a batteries-included philosophy here, probably very similar to how you did Spark, which is you could just start using it.
M
Matei Zaharia26:20
Yeah, that's right. It has to be good out of the box at certain things, and then you can build your own things on top. You know, we don't want to do—but in Spark, if you just want to read a table or do aggregation, it should be awesome out of the box.
H
Host 226:36
Yeah. People want to catch up on Omnigents, they should watch your keynote, they should go through the GitHub and the docs. If they wanted to contribute or build on this ecosystem, where would you call out as the most high-leverage places to get involved?
M
Matei Zaharia26:49
Yeah, do get involved in the Discord and GitHub. Our team is monitoring and some of the things people ask for we just built ourselves. Some of them, we're collaborating with them to build. And also tell us how you would like to use that because I think especially for developers—everyone wants it to work their own way and a really good developer tool—you have to hear the feedback on all the ways and figure out the abstractions and how to let people customize. So we love to hear—if you think, hey, I don't want it to work this way, tell us. We really just want to get that compatibility layer across agents and then let you do stuff on top.
H
Host 227:27
Yeah. In terms of the startup side—I'm a founder, I see an opportunity, I want to get in front of you. What's your request for like a startup that—I wish someone was working on this. You advise many startups too obviously.
M
Matei Zaharia27:39
Oh for a startup. I mean I do think—just as a company with a lot of engineers—anything that helps me make sense of how people are using coding agents and spend but also quality—or like you should add this skill or you should write this thing or your agents are really horrible at tasks involving this service or go spend time. That would be nice.
A
Ali Ghodsi28:14
Yeah, the closest I found is this team Get AI.
M
Matei Zaharia28:17
Mhm. Oh cool, yeah.
A
Ali Ghodsi28:18
They started with—we will just do code and human attribution—but they're basically building the analytics layer on top of that. I do think there are a bunch of—Artificial Analysis is obviously doing super well with their stuff. So there will be people—I think this is the domain of consultants first but then people will actually build software that is like—the management plane for coding agents.
M
Matei Zaharia28:44
Yeah, I think there'll be a lot of insights there. You have it in other areas.
H
Host 228:48
Okay, well, and then the other big thing is your dream engine.
If you want to tell the story of OLTP—and our background was—I'm going to make people listen to our Ankur Goyal episode where we talked about single store, HTAP, and all that history.
A
Ali Ghodsi29:06
Yeah, yeah. The OLTP idea is actually pretty simple. So, people have heard of Ankur's talk about HTAP—it's effectively the world of databases. Sorry, there's like maybe a lot of context that needs to be injected here. The world of databases—
H
Host 229:20
This is the database podcast that I'm forcing people to learn your databases, guys. You cannot vibe code with just markdown files.
A
Ali Ghodsi29:27
It's one of the most important fundamentals of systems technologies out there. But, the world of databases is effectively split into roughly two halves. There's what we call OLTP databases, which are transactional—think of your Postgres, your MySQL, your Oracle databases. And the other side is what we call analytics, and sometimes might refer to as OLAP. And the difference is on OLTP, you typically run some transaction on an event that looks up one specific row, we update that row, right? It's a very row-oriented data structure. And on analytics, you're trying to reason on the data—you're trying to compute, hey, what's my revenue per store? What's my—how's my website doing every day? And then you eventually want to probably end up running machine learning on it to predict—how will my sales be going in the future? They are very different architectures, and everybody starts with OLTP databases because every app, when it becomes serious enough—needs more than markdown files—you need to have a database. You don't want to lose your data. You want to have some transactional consistency. But, once you want to reason on the data, if you only have like 100 rows, it's probably okay to run it on your Postgres or your MySQL database. But, once you have more data and want to run more complicated analysis, the very analysis might crush your Postgres database. So you start getting data out of the—
M
Matei Zaharia30:49
Replication—replicate them into the analytic systems.
A
Ali Ghodsi30:53
Yeah, which for people—Elasticsearch is like a big—
M
Matei Zaharia30:56
Yeah, so some of them actually get into Elasticsearch for log analysis. A lot of our customers obviously get into Databricks to run more sophisticated things. And there's this term called CDC—
A
Ali Ghodsi31:07
CDC—change data capture. And what it does—it reads the bin log of the database. And if you don't understand what bin log is, fine. But it's a little delta of the data and then reconstructs based on the delta the state of the database on the analytic side. But CDC is a very painful thing. It's basically standard in the industry. Everybody uses it. But it ends up being—so I think many data engineers end up being woken up at like 3:00 AM because of some pipeline thing.
M
Matei Zaharia31:36
You know, my explanation is—everybody became a $5 billion company just doing CDC.
A
Ali Ghodsi31:40
Yeah, exactly. CDC is—it's one of the most boring but one of the most fundamental operations powering modern society. But it's so brittle that we joke that it should be called continuous data corruption because you might change your schema on your OLTP database and then the CDC pipeline fails to handle the schema change. And then everything goes out.
M
Matei Zaharia32:04
I mean there's all sorts of tricks that you can do—like you add in some versioning or whatever. But—
A
Ali Ghodsi32:08
Yeah, but it's in general very complicated. Like I think at my keynote I asked the audience—put up their hand if they love their CDC pipeline. Only like maybe two people put it up. So if single store—like about maybe a decade ago I think the industry had this idea—hey, what if I built a single database that can handle both workloads?
M
Matei Zaharia32:26
Which like by the way every database person ever has always dreamed about this.
A
Ali Ghodsi32:29
Yes, yes, this is the holy grail of database engineering. Why not build a single system that can do both of this? But it ends up just being a lot of compromises. One, I think one of the first issues is that—hey, Postgres has a massive ecosystem, right? You want to be using the tools that's built for Postgres. And Spark, for example, had a massive ecosystem. There's a lot of libraries you want to use. If you were to create now a new thing, you don't have an ecosystem. You tend to create a new smaller proprietary API and you're lacking both. And it's also very difficult to make it performance-wise to be sort of comparable on either side. So it ends up actually sucking up both. And our whole idea of LTab is kind of obviously a word play on the term HTAP—is that we think this is HTAP done right. HTAP wants to build a single engine for both. We think you can get 99% of what you need by unifying the storage. And just have a single storage layer. And once you have the single storage layer, if your Postgres databases are writing data in a column-oriented format, everything analytics can just go read that data directly without any delay. Right? There's no pipeline in between, so all the data would immediately be available for reasoning analytics. I think I was telling some customers earlier—hey, when we talked about this is going to be super useful for agents, I actually at first didn't really believe in it myself, even though we wrote that positioning. But then last night I was having dinner with an Australian customer and they actually told me—oh hey, one of the big issues we have is we have all these logs from our services and we see SLA dips and want to investigate. But then there's no way for those agents to even understand what's going on in the actual databases themselves. All we see is just product telemetry of the database and the services. You would actually make those agents 10 times more powerful if you understand, for example, who's actually placing those orders. What is happening? What exactly are they doing? So now I'm actually sold on our own message. I think it's really kind of—it gets you basically almost all of the benefits of the HTAP holy grail, which is hey, make the data available immediately for reasoning analytics.
M
Matei Zaharia34:40
Yeah, I think you know—in the way that humans are generally intelligent and want to have the ability and access to query anything. Even while they do the work, they also need history and they need context—and where else do they get context? That's an analytical workload.
A
Ali Ghodsi34:54
Exactly.
M
Matei Zaharia34:56
Yeah, and I remember when we had incidents with our databases and the engineer said—well, I can't just run a giant query on it to see what's going on because that's going to bring down the database and hurt it even more. Like that's the kind of stuff that this gets rid of because you spin up a whole separate fleet of machines that's doing the analytics. You're not overloading the main database that's still trying to serve stuff.
A
Ali Ghodsi35:17
Yeah.
H
Host 235:18
So, this has been a dream for a while. What had to get done in order to get to today? Like, you know—I feel like you have announced variants of this several times. But it wasn't as clear as LTab.