Back
Ali Ghodsi
Cofounder and CEO, Databricks

Ali Ghodsi, Co-founder and CEO, Databricks kicks off Data + AI Summit 2026

🎥 Jun 23, 2026 📺 Databricks ⏱ 57m
Databricks Co-founder and CEO Ali Ghodsi opens Data + AI Summit 2026 and outlines how the Databricks platform solves the ...
Watch on YouTube

About Ali Ghodsi

Ali Ghodsi, cofounder and CEO of Databricks, has been speaking publicly about the company’s recent fundraising, product launches, and his views on the state of artificial intelligence. In July 2026, Ghodsi said on CNBC that Databricks was fundraising at a $188 billion valuation, driven by customer demand for open-source AI models such as Kimi and a resulting shortage of GPUs. He stated that the company’s Unity AI Gateway product, which allows customers to use both open-source and proprietary models, had seen explosive demand. At the company’s Data + AI Summit in June, Ghodsi announced the open-sourcing of a new project called Open Sharing, which he described as a superset of Delta Sharing that enables sharing of data, AI agents, skills, and models. He also noted that over 100,000 people had signed up for the event, making it the largest data and AI conference in the world. Ghodsi has repeatedly argued that artificial general intelligence (AGI) has already been achieved, stating that the current constraint on AI productivity is not intelligence but context. He said that AI models are already smart enough but lack the enterprise context—such as data, processes, and organizational knowledge—to be fully useful in the workplace. He described Databricks’ Lakehouse as a “system of record for agents” and highlighted new products including Customer Lake for marketing and Lake Watch for security. On the topic of an initial public offering, Ghodsi said he believes 2026 is a “terrible year” to go public due to macroeconomic uncertainty and the presence of several large IPOs, and that Databricks, which he said does not burn capital, can afford to wait for more stable conditions. He also noted that the company may raise additional private capital before going public.

Source: AI-verified profile updated from Ali Ghodsi's recent appearances. Browse all interviews →

Transcript (24 segments)
N
Narrator0:00
Data Bricks co-founder and CEO Ali Ghodsi.
A
Ali Ghodsi0:15
All right, super excited to be here. That was an awesome video showing customer use cases. This is the largest Data+AI Summit we've ever done, with over 100,000 sign-ups, 174 countries represented, and 3,139 attendees, making it the largest data and AI conference in the world. This conference is rooted in open source—we started with Apache Spark, which now has over three billion downloads a year, and we've added Delta Lake, Iceberg, MLflow, Unity catalog, and we've embraced Postgres. My co-founder Matei open-sourced Omniant this weekend. We'll hear from Greg Brockman, Satya Nadella, and many other leaders. Thanks to our partners and customers. Let me highlight three inspirational customers: Insulet's Omnipod helps people with diabetes using data and AI; Tonal personalizes workouts with 100 billion seconds of data; Merck trained Teddy, a transformer model for drug discovery that predicts gene regulatory networks. Now, is AGI here? Only about 5% of you think so, but frontier models can solve problems like computing reduced 12-dimensional spin bordism of the classifying space of Lie group G2—part of Humanity's Last Exam benchmark, where models score 50%. By the definition we used in 2009 at UC Berkeley's AMP Lab, AGI is already here by leaps and bounds. The problem is that AGI is not yet permeating our organizations. We're still using chatbots, not managing hundreds of autonomous agents. The key is to provide context from all your data and processes to these smart models. Getting that context is hard, and we need to solve four problems: context, control, cost, and choice—avoiding lock-in. Today's keynote focuses on these four.
Hey Ryan, how are you? Good. I want to embarrass you with a video from a bunch of years ago. This video actually had a very interesting title. It said...
R
Ryan Blue18:14
Talk to you guys about why you shouldn't care about Apache Iceberg. I mean, listen to it. I care deeply about it. And I sincerely mean that. You shouldn't care about Apache Iceberg.
A
Ali Ghodsi18:25
Okay, we should care about Apache Iceberg. Actually, I agree. I think we shouldn't also care about Delta Lake. But why did you say that?
R
Ryan Blue18:32
No one should have to care about formats. You should be able to use whatever tools you want with all of your data. Use the right tool for the job, which is why we built both formats in the beginning.
A
Ali Ghodsi18:45
Yeah. In fact, it's super funny. A few years ago, I would go to customer meetings and they would get into the nitty-gritty of how Iceberg was implemented versus Delta, asking about deletion vectors. I told them this is not where the industry should be going. So, how are we doing? Where are we? What's going on with the versions?
R
Ryan Blue19:05
Well, we just did the GA release of Iceberg v3 support, managed Iceberg tables in DBR. That is going really well. We're following that up later this year with the unified metadata layer built into Delta 5 and Iceberg v4. So we are very close to the full unification vision.
A
Ali Ghodsi19:27
So what's special about V3?
R
Ryan Blue19:29
V3 is the unified data layer. So you no longer have to rewrite any data files to share them across Delta and Iceberg tables.
A
Ali Ghodsi19:37
That's awesome. So basically the data that's laid out, whether Iceberg or Delta, looks identical now. Right.
R
Ryan Blue19:42
Exactly. Yes.
A
Ali Ghodsi19:43
Awesome. And then with V4, just a little bit of metadata left. When do we get that?
R
Ryan Blue19:49
We're aiming for later this year. I don't want to make too many promises, but I think we should finish the spec in Q4.
A
Ali Ghodsi19:58
Awesome. Super excited. Great job by Ryan and the whole Iceberg and Delta community. You shouldn't have to care about this. Can we stop talking about this now?
R
Ryan Blue20:07
Yeah, I got work to do if we're done.
A
Ali Ghodsi20:10
All right.
So it doesn't really matter if you're using database, Delta or Iceberg—they're the same now. We have LakeFlow with over 100 connectors to get data into the open lakehouse. ZeroBus is GA, Spark real-time mode enables tens of milliseconds latency, and LakeFlow Designer is a visual AI-powered tool. Our lakehouse has added 110 features for legacy data warehouse migration and AI functions. We also have Lakebase Postgres with autoscaling down to zero and branching for instant database clones, which agents love. Unity catalog is free and open source for governance of all data and AI assets. We're open sourcing Open Sharing to share data and AI assets across formats and on-premise. So that's how we solve context, control, cost, and choice.
About that—this is very cool. So we've got all these different things you can do in Unity Catalog. We have open sharing. But governance, which has always been part of Unity Catalog, is becoming very complicated. Organizations now have lots of agents, MCP servers, skills files, and models from frontier vendors. It's a quagmire. Three big problems: cost is skyrocketing with no visibility; no control over agents' data access, auditing, or identity; and lack of choice because models become obsolete in a month—Gemini, Opus, GPT-5, then Mythos or Fable, now canceled. So we're announcing Unity AI Gateway, a single pane of glass to manage and control all agents and AI spend.
Unity AI Gateway is part of Unity Catalog, which is open source, and it's also part of MLflow. It provides a single entry point for all AI—agents, models, harnesses. You can use your committed spend with Databricks to consume tokens from frontier models on any cloud. It gives you observability of spend with dashboards and allows you to set budgets down to the individual level, with alerts and rate limiting. It also enforces safety, compliance, and auditing for all AI assets. Now let me talk about the context layer. AI doesn't have an intelligence problem; it has a context problem. Current agents do a live random walk through your data, which is time-consuming, expensive, and suffers quality issues. So we're excited to announce Genie Ontology.
Genie Ontology builds a graph of your organization's most important knowledge in the background, connecting to not just the lakehouse but also Google Drive, SharePoint, email, and calendar. Our research team developed an algorithm called Onto Rank—like PageRank but for different asset types (code, docs, users, org charts). It constructs an ontology graph to feed context to agents, making them faster, cheaper, and higher quality. You can also bring your own semantics from partners like Atlan or existing BI tools. This feeds into three categories of agents: Genie 1, Genie Agents, and Genie Code. Genie 1 is a single interface for all business users to ask questions across all data sources, with skills, scheduling, and mobile access. Genie Agents lets you turn any conversation into a company-wide agent, deployable in Slack or Teams. Genie Code excels at data engineering and machine learning, helping write pipelines and train models.
We're also launching Genie Zero Ops, which automatically monitors your data pipelines and ML models at 2 am when something breaks. It investigates, experiments with fixes in a sandbox, and sends a notification for you to accept. No more middle-of-the-night calls. For developers, we've expanded Agent Bricks with sandboxes and agent memory, and we announced Omnigent—a harness of harnesses that lets different coding agents compete for better results. This all ties into the future of the software stack. The classic SaaS model with separate systems of record is breaking because vendors each have agents that don't talk to each other. The future is an agent system of record: all data in one open place, unified governance, cost management, and enterprise context. That's the Data and AI platform we've been describing. And on top, we have Databricks Apps to democratize access, with a marketplace where you can buy or build custom apps.
We're building some apps ourselves where data is intensive. First, Lakewatch—an agentic SIEM built on the security lakehouse. It collects all security data cheaply, then agents create detections, triage alerts, and hunt for threats automatically. We're also acquiring Panther Labs, a Python-based SIEM with hundreds of connectors and Pythonic detections, used by Anthropic, Coinbase, Plaid. Welcome, Jack Naglieri and team. Second, we're announcing Customer Lake—an agentic customer data platform built on the lakehouse. It has a profile agent for identity resolution using LLMs and a campaign agent for infinite personalized campaigns using distilled models, developed in partnership with vendors like Atlan and others.
Customer Lake enables one-to-one continuous campaigns. To wrap up: our platform gives you choice, governance, cost control, and context. With Unity Catalog, Unity AI Gateway, and Genie Ontology, you can get your data ready for AI, control costs and access, and deliver context to agents. Data and AI platform will help you do that. Thank you.