Back
Matei Zaharia
Cofounder and CTO, Databricks

Summit Live | Founder discussion: Matei on UC, Data Intelligence and AI Governance

🎥 Jul 17, 2025 📺 Databricks ⏱ 14m 👁 104 views
Matei is a legend of open source: he started the Apache Spark project in 2009, co-founded Databricks, and worked on other ...
Watch on YouTube

About Matei Zaharia

Matei Zaharia, cofounder and CTO of Databricks, introduced Omnigent, an open-source meta-harness for AI agents, at the company's Data + AI Summit in June 2025. He described Omnigent as a new layer above existing agent harnesses that aims to make agents interoperable, allowing sessions, policies, and skills to follow users regardless of where the agent or model runs. Zaharia said the project was released as open source because the company believes the meta-harness layer will benefit from being open in the same way data formats and sharing protocols have. He also demonstrated features including contextual security policies that vary permissions based on prior agent actions, cost controls, and an OS sandbox. In a separate appearance, Zaharia discussed Databricks' broader strategy, stating that many traditional software products will be rewritten under a paradigm of "get the data to be there and then let's slap some AGI on top." He attributed Databricks' differentiation to its commitment to open formats and its focus on AI. He also commented on the company's LTAP (Lakehouse Transactional and Analytical Processing) offering, describing it as providing the benefits of the HTAP (Hybrid Transactional/Analytical Processing) holy grail by making data available immediately for both reasoning and analytics workloads.

Source: AI-verified profile updated from Matei Zaharia's recent appearances. Browse all interviews →

Transcript (29 segments)
A
Ari Kaplan0:18
And welcome back. For those just joining in, I'm Ari Kaplan and I'm here with Holly and Mate. So welcome. We have over 55,000 people from all over the world, 140 countries. So we see you, we love you, we appreciate you and very excited to have you. If you are here for the pre-show, they asked, 'What am I most excited about?' And I'm like, this segment, this moment right now, it's Mate and we have tons of practitioners and you're great. So say hello and what you do here.
M
Matei Zaharia0:47
Yeah. Hi everyone. Great to be here. I'm co-founder and CTO of Databricks, working on a whole bunch of things throughout the platform, but lately a lot of Unity Catalog and AI stuff.
A
Ari Kaplan0:59
Yeah. So what does your organization do?
M
Matei Zaharia1:03
Well, I basically work with the product and R&D to actually build the Databricks product. A lot of what our co-founders do is make sure that everything fits together and is designed in a comprehensive way, because the whole premise of our platform and product strategy is a unified platform for everything you do with data and machine learning. So that requires designing everything so that it really can fit together and interoperate and remove those sources of friction you tend to have with very different systems.
A
Ari Kaplan1:34
So you focus on a huge amount of areas of the Databricks platform and lots of other things as well, which is incredibly impressive. But I think one of the things you've really been banging the drum on is governance for many years now. I think maybe it was three or four years ago, Unity Catalog was...
M
Matei Zaharia1:50
Yeah, four years ago we announced it.
A
Ari Kaplan1:52
So we've got four years of development of all the things that have gone on. What are you most proud that you've been able to either do from a technical perspective or from a customer enablement perspective?
M
Matei Zaharia2:01
Well, I'm excited about our customers. Basically, the majority of our customers, 97%, are now on Unity Catalog. We introduced this new governance model. It's different. It's hard to migrate a governance model. We thought it's a great one because it has a uniform way to think about files in your data lake, tables, machine learning models, dashboards, all that stuff. So we thought it's a great one, but it's still a lift to move people over. So it's great to see it actually adopted, and that required a lot of work from our engineering teams and field teams. And it's paying off. People are getting a lot of benefits out of it.
A
Ari Kaplan2:42
Yeah. And I was going to say, you debuted Unity Catalog and last year we had probably one of the keynote moments of the whole conference was you pushing this button to open source it. What was that like?
M
Matei Zaharia2:54
It was great. We always knew since we began it, we were actually discussing, 'Hey, should we make it open source from the beginning?' And we thought we want to get experience with it and iterate with our customers, make sure it's a great design, because once it's open source, it's hard to really change the fundamentals of how it works. So we were excited to bring it to that point and open source it. The reason it matters is the objective here is to really help you simplify governance across your company. It needs to be unified, and we believe it can't be unified unless it's open. It's got open APIs because everyone has so many different platforms and systems, legacy ones, new startups coming out. So it needs to have open standard APIs so they can work with it.
A
Ari Kaplan3:41
So it's something that's kind of fundamental to Databricks culture, being able to open source things, not just Unity Catalog, but we've had loads of announcements from this week. We had our open source summit on Monday. First of all, I'm very glad that this kind of culture of open source has stayed with the company, which doesn't necessarily stay with companies. Why do you think it's so ingrained in Databricks culture? Why has it stayed compared to other companies that start and then it maybe doesn't?
M
Matei Zaharia4:07
I think part of it is we always, when we started the company, there were these trends that are changing the way enterprise computing and data management is done. One of the trends was open source. We knew it's here, it's a very powerful force, it's going to be here. So we should learn how to work with it as a company and get all the benefits of it, because it's not going away as a movement and as a way to do collaborative development. And then the other part was the cloud. The interesting thing about the cloud, the thing that actually enabled the whole unified platform approach and things like Lakebase that we're launching, is in the cloud, all your storage, all your data you had in these siloed systems before, it's all bits sitting on the same hard drives in Amazon S3 or whatever. You don't even know where in the data center you are. Your database table for your online database and your machine learning model and your analytics data might all be on the same hard drive, you don't even know. And there's a super fast network. So in the cloud, it makes a lot of sense to have open interfaces and to be able to connect any engine to any data and work with it in a way that didn't make sense on premise. And as I said, enterprises understand the value of open. They're designing their architectures around it. We also think you can't have something truly unified and useful without that. So that's why we've been doubling down on it.
A
Ari Kaplan5:38
One thing I love hearing, we all love open source. The practitioners out there, if they could talk to you, they'd say thank you, Mate. But you open source everything, starting with Spark, Unity Catalog, MLflow, Delta, everything. Billions of downloads. What does it feel like to have such a huge impact on society?
M
Matei Zaharia6:00
Well, we definitely didn't expect any of these projects to grow so much. So it's been really awesome. I think we had a great team and we had the right thing at the right time, and then we learned a lot from early users. So it's awesome to be able to contribute to this and help so many folks be productive. Just even the people using things like Spark outside Databricks have done so many cool things with them, and seeing what our customers do is really great.
A
Ari Kaplan6:26
Could I ask you a bold question? You don't have to answer this. Open source applications of technologies that you've been involved in that aren't used at Databricks. What do you look at and think, 'Oh, that's really interesting' or 'That's really novel' that's not used at Databricks?
M
Matei Zaharia6:41
You can tell because we use so many things. We do front end development, backend development, machine learning, and so on. I think some of the things we're not doing is a lot of this edge stuff like Raspberry Pi and robotics. But when I see how creative people get with that, that's really cool to see. I wish I had time to do some things with that at home. But getting computing and AI into the physical world, I think will be really interesting.
A
Ari Kaplan7:10
You must have an abundance of free time. How often do you travel?
M
Matei Zaharia7:14
I travel pretty regularly, although I don't spend most of my time traveling because I also spend a lot of time just working with our teams to build things. But it varies from time to time.
A
Ari Kaplan7:27
You must have been very busy this week, not just speaking about things but also talking to customers. Of the stories that you've heard from people and the way they're using Databricks, what have been the trends and themes you've heard?
M
Matei Zaharia7:39
I think one really exciting trend is just the interest in Lakebase. The number of customers who say they're already doing a lot of work to try to publish data from the lakehouse as soon as it's computed into operational apps is really interesting. One customer told me they have an internal name for it that I thought was great: 'Lakeshore.' I thought, 'Can we steal that name?' I don't think we can, but it's such a common thing that they gave it a name. So the excitement around that was great, and as we announced it, we discovered so many use cases. I think the other one is I'm spending a lot of time with our AI team and our AI research team to figure out how to make it easier to build really high quality AI. People's feedback on how important evaluation is for AI and the ways they're doing it, and the ways they're plugging that into Databricks with our agent eval framework, is also great to see. We want to provide the best support for that because we think it's the key to make agents better.
A
Ari Kaplan8:42
So, but sorry, you mentioned Lakebase. Can you talk about Neon at all and what it was that you liked about their technology?
M
Matei Zaharia8:48
First of all, Neon is a fantastic product and we're continuing to run the product on it. Neon came in at the software engineer end of database users. You're building an app, you want the database, you want to be able to do modern development practices around it like CI/CD, branching, forking, and you want it to be serverless and pay as you go. So they made that super easy and it's great for that. Lakebase, we actually started at the analytics end because that's most of our users. So it's going to be very interesting to start from both ends. The Neon product has been growing extremely fast. For example, some of my past PhD students have startups and they were using Neon, and they told me it's such a good move because they love it. It's their favorite infrastructure provider. So a very high quality product, great team, great roots in the Postgres community. They have a lot of the Postgres committers that work on it, and they've been contributing things and adapting Postgres so that it can work better in this modern serverless and cloud stage that we're in. So these are the things that we liked, and I think they'll bring great product sense across much more than just Lakebase, across a lot of areas of Databricks to reach these apps.
A
Ari Kaplan10:10
I love with the Neon one of the things we learned is just how many AI are creating databases now. So what are your thoughts on AI creating things, ethically and effectively?
M
Matei Zaharia10:23
Well, there are two reasons. One is you have AI agents that need memory and you need to set it up, but the other reason is because of AI-based development and testing tools. The interesting thing is today, LLMs are not perfect, but they can do a lot of useful tasks. If your development environment, your code, and your database are something where you can easily generate a new version and test it, some of these tools do that. That's what's driving a lot of usage and pressure on development environments to be scriptable, forkable very fast, and testable very fast. So it seems like the use of that for development is only increasing, and people are much more productive. You ask the AI assistant to code something, it tries a lot of stuff maybe in parallel, it shows you what works, you give it feedback. So that's great to see. In fact, we're seeing that workflow also with Genie Deep Research, where it's a similar thing: you tell it to explore something, it does a whole bunch of queries in parallel and summarizes the results. Of course, you have to be careful as a developer using AI; you have to review the code. But people are learning. It's not that fundamentally different from when you did a web search and copied stuff off of Stack Overflow. The same kind of issues could happen where you copied in the wrong thing. But it really does speed up development and prototyping.
A
Ari Kaplan11:56
Absolutely. I hate to do it, but I'm going to do that annoying interviewer thing. I know you spent a lot of time working with Neon, getting that over the line. And now I'm going to ask, what's next for Lakebase?
M
Matei Zaharia12:06
What's next for Lakebase? Well, first of all, we're very excited to just ship it in public preview. It's been in private preview with hundreds of customers already. It's one of our largest private previews, and it's been great to see, but we're very curious to see what everyone else does with it. We've got a bunch of new functionality coming. The way it syncs with your lakehouse tables is going to be even faster and smoother, and some of the development features that Neon really excels at, we are pushing into Lakebase. I'm really excited in the long term of having everything else that's on Databricks available there, like Unity Catalog governance, ML, calling your machine learning models from Lakebase, metrics, all these things. So plenty of room to explore. The Neon team, the acquisition officially closed a few days ago. The team's been in the office getting ready, having discussions, and it's really awesome to see all the potential ways to make it better.
A
Ari Kaplan13:11
Okay, then. I think we've only got one minute left. Is there any one message that you wanted to leave with our audience? What would it be today?
M
Matei Zaharia13:17
I think overall, we're excited you're all here. We think it's still early days in terms of data and AI infrastructure. What is happening in the market today is just too complex, more than it needs to be. Our hypothesis has been you can have a unified platform. You can have unified governance, which is already paying a lot of dividends for our customers by making that simpler, but also unified storage and compute layers. The way you can now take a table in Databricks and ingest directly into it with zero bus, so it's also a message bus. You can then serve it through Lakebase or vector search. There's a lot of interesting technical work to do to make this one copy of the data a reality.
A
Ari Kaplan14:02
Thank you so much for your time today. I appreciate you've been very busy this week. We really, really appreciate it.