About Reynold Xin
Reynold Xin, cofounder and chief architect at Databricks, has been discussing the company's recent product announcements and his views on the future of data infrastructure. At the Data + AI Summit in June 2026, Xin introduced Lakehouse//RT, a real-time analytics engine powered by a new compute engine called Reyden. He demonstrated Reyden executing 8,000 queries at 6,000 queries per second with a tail latency of 37 milliseconds, and said that "none of existing systems can do this." Xin also described Lakebase branching, a copy-on-write clone of an entire database that he said can be created in under a second and costs roughly a penny before being thrown away. He argued that traditional databases "inhibit innovation" because they are slow and expensive to provision, and that infrastructure should enable "rapid experimentation at scale cheaply."
Xin has also discussed the impact of AI agents on database architecture. He stated that "agents are becoming the actual primary persona" for database usage, and that "99% of psycho value engineers these days don't write code manually and don't provision the database manually." He described Databricks' work on LTAP (lake transactional and analytical processing), which aims to unify transactional and analytical databases, and Omnigent, an open-source meta-harness for combining different coding agents. Xin said that "many of the traditional software will be sort of rewritten with this new paradigm, which is just get the data to be there and then let's slap some AGI on top."
Source: AI-verified profile updated from Reynold Xin's recent appearances.
Browse all interviews →
Transcript (23 segments)
I
Interviewer0:00
Reynold, I want to hear your take on the value of a handful of the capabilities that Databricks offers. Let's start with LakeFlow.
R
Reynold Xin0:09
The vision of LakeFlow is really to simplify data engineering. If you think about data, everything starts with data engineering. There's typically this persona, data engineer, that tries to get all the data into a central repository, which in our case would be the lakehouse. If you use a proprietary data warehouse, it could be the data warehouse. Then they need to transform this data through maybe what we call medallion architecture, from bronze to silver to gold. This is an extremely complicated, super costly process. LakeFlow aims to make that simple for data engineers. It has Flow Connect to help you get data in from a variety of different systems. There's workflows and there's also a Spark declarative pipeline to help you with the transformation. We're also working on more to make data engineering not just simpler for data engineers but also visible for non-data engineers.
I
Interviewer1:21
I'm very excited for LakeFlow Designer.
R
Reynold Xin1:23
Yeah, exactly. So they don't have to be blocked on having...
I
Interviewer1:26
As someone that before struggled with Delta Live Tables, for me, I'm liking where declarative pipelines is today. I'm excited for LakeFlow Designer even more.
R
Reynold Xin1:39
So then people don't have to block always on you.
I
Interviewer1:43
Exactly. Since we already mentioned a lake something, what about Lakebase?
R
Reynold Xin1:50
Yeah, Lakebase is our take on OLTP databases, which is the other side of the giant data estate that we haven't touched in the past. We want to make databases far easier to use and make them much easier to manage, and also make them great for this new persona that we call agents, in addition to humans. We want to actually make databases integrate much better with the lakehouse. When we talk to a lot of customers, they talk about how they have their OLTP database but it's in a different silo from their data infrastructure for analytics. But we always end up finding that the OLTP database is actually one of the most important, maybe the most important, data sources for the data lake. How do we actually combine them? We think we have everything it takes to actually solve this problem for everybody. That's why we're excited about Lakebase.
I
Interviewer2:47
AIBI dashboards.
R
Reynold Xin2:49
Yeah, ultimately, I think we have all these technical personas, data engineering teams, data science teams, but if I want to democratize data, I want to actually enable every individual, whether they consider themselves an analyst or not, to be able to get insights out of data. One of the most important tools, or toolkit, every enterprise has is their BI tool. AIBI is our take on BI. It's actually free because we actually think we can monetize enough just based on the level below that we don't have to worry about charging you with user-based pricing or some other seat-based pricing. It's completely free and it gives you not just what traditional dashboards can do. It's an independent dashboard product. It's gone from pretty crappy in the past.
I
Interviewer3:51
It keeps getting better.
R
Reynold Xin3:53
But we think what's really interesting about it is the AI part of it. BI dashboards are largely a static thing and they make a lot of sense. It's a great, very important thing. It will never go away. But I think with AI, we can now actually ask questions that are more general, whether it's based on a dashboard or just based on a bunch of semantics that are defined about your data. We think that's how we're going to disrupt BI tools in the past.
I
Interviewer4:25
No, honestly, I am personally very happy with how much dashboards has matured. I've shared this publicly because I asked him before. I once had asked a lead, like, hey, you have dashboards there, is there even any intention from Databricks to make dashboards good? This was almost two years ago and his response was basically like, hey, if we're going to offer it, we want it to be very good. So, personally, I'm very happy with how much the product has improved.
R
Reynold Xin4:58
To us, it's a long-term game. So, you shouldn't look at sort of what's on the truck today, but actually look at the velocity or the slope of that future.
I
Interviewer5:08
Shout out to Miranda and Chiao and Ken who took so much of my feedback and they probably got tired of my feedback.
R
Reynold Xin5:15
They love your feedback. Keep on coming.
I
Interviewer5:17
Delta sharing.
R
Reynold Xin5:18
I think Delta sharing is probably the first sort of open data sharing protocol that works at scale. I'm sure there are a lot of open source projects here and there, but the really interesting thing about Delta sharing is that in the past, we've seen among our customer base and just the broader community in general, people want to be able to share data either internally within the company or externally with their data providers and vendors and consumers. It creates this very important network, but every sharing protocol in the past was some sort of proprietary protocol. Which means now you have to implement specifically for a sharing protocol and a vendor. Delta sharing is just, when we decided to do it, it was like, hey, we are very proud of our open source roots and the way we build open standards. Let's build, why can't a data sharing protocol be open? So we just built one and it's taking off like crazy. Obviously, it benefits more than us because it's an open protocol. There are a lot of Delta shares that happen between different constituents that have nothing to do with Databricks. We think that's okay because we're creating more of an ecosystem here that we're an important part of. But then hopefully in the future, if you have a data sharing use case, you just implement the open protocol of Delta sharing once and you're done. You don't have to worry about, hey, do I implement this for this vendor and that for the other vendor and then keep doing it.
I
Interviewer7:01
I personally love, I've seen the value of it, especially within the... when I've worked with vendors of data, I think it really opens up, it makes things very frictionless, which I think is awesome. Last but not least, the darling of Databricks, Unity Catalog.
R
Reynold Xin7:28
Yeah, it's incredible. Honestly, when we first started doing Unity Catalog a few years ago, I didn't expect it to be this much of a hit. You can think of Unity Catalog from two angles. One is bottom up. It gives you a single pane of glass for all of your resources that are not just about data but also other things you care about alongside data like machine learning models, notebooks, your Lakebase databases. But also on the data side, it actually federates and can index not just Databricks catalog, can index Snowflake catalog, can index a Postgres instance, index everything. So it kind of provides a single pane of glass for everything in data and everything that's important in AI. And then on top of that, it gives all these governance capabilities to cover all of them. It gives them RBAC, standard row-based access control, column-based access control, and makes sure they're enforced properly regardless what the query engine would be. And honestly, probably more than half of Databricks engineering resources in the past few years went into making sure UC actually works at scale. And making sure we can enforce properly that governance there, which is a very complicated problem because Spark initially was not designed for that.
I
Interviewer9:03
Yeah, there's a massive amount of work that's gone into Spark itself.
R
Reynold Xin9:08
Which is, I think, now one of the main sort of competitive advantages we have. Because if you look at a lot of other platforms, they actually, first of all, they don't have a catalog to begin with in many cases, at least a catalog that can index this many things. And then often when they have a catalog, they can't actually enforce governance because they haven't done that heavy lift in engineering to actually make sure the governance there can be properly enforced, especially when they're usually defined in code. But yeah, it basically solves the enterprise governance problem, which is extremely valuable for enterprises.