CEOInterviews.AI
Start App
Sundar Pichai
Chief Executive Officer & Director, Google

Sundar Pichai Opening Remarks | I/O 2026 Keynote

📅 May 22, 2026 Google and Sundar Pichai 42 MIN 1925 VIEWS 26 SEGMENTS · 4 SPEAKERS
Watch Sundar Pichai's opening remarks from the Google I/O 2026 keynote. Learn about Google’s vision for the future of AI, our latest progress with models and coding and how Google is unlocking the helpfulness of agents. 🚀🧠 Sundar sets the stage for our biggest announcements of the year, sharing how Google is making AI helpful for everyone. Watch the full Google I/O 2026 keynote here: https://youtube.com/live/wYSncx9zLIU Hear more from Sundar on his YouTube channel:    / @sundarpichai   0:24 - Keynote Intro & The Scale of AI 4:04 - Generative Search & Ask YouTube Updates 7:43 - Docs Live...

What Sundar Pichai said

Written from the verified transcript and checked against it. Every figure links to the moment it was said.

Sundar Pichai opened Google I/O 2026 by highlighting the company's AI-first pivot a decade ago and its full-stack approach. He reported token processing growth from 9.7 trillion monthly two years ago to 3.2 quadrillion today, with over 8.5 million developers building monthly and 375 customers each processing over one trillion tokens in the past year. Pichai noted 13 products with over a billion users, AI Overviews at 2.5 billion monthly users, AI Mode surpassing 1 billion, and the Gemini app reaching 900 million monthly active users. He announced Ask YouTube rolling out in the US this summer, Docs Live for Pro and Ultra subscribers, and projected 2026 CapEx of $180-190 billion. He introduced TPU v6p and v6e, Gemini 3.5 Flash with four times faster output, and Gemini Spark, a personal AI agent, with new Ultra pricing at $100 monthly and top-tier Ultra reduced to $200.

Key takeaways

  1. Gemini 3.5 Flash is four times faster than other frontier models and available today across products and APIs.
  2. 2026 CapEx expected to be approximately $180-190 billion, about six times the 2022 level of $31 billion.
  3. Gemini Spark, a personal AI agent, rolls out to trusted testers this week and as a beta for Ultra subscribers next week.
  4. New Ultra plan priced at $100 per month; top-tier Ultra reduced from $250 to $200 per month.
  5. Ask YouTube will roll out broadly in the US this summer; Docs Live for Pro and Ultra subscribers also this summer.

Numbers and commitments

FigureWhat it refers toTypeAt
$180 to $190 billion 2026 capital expenditure guidance guidance 9:47
3.2 quadrillion tokens processed per month across surfaces metric 2:27
8.5 million developers building with models monthly metric 3:21
19 billion tokens per minute processed by model APIs metric 3:21
375 customers each processing over one trillion tokens in past 12 months metric 3:21
13 products with over a billion users each metric 3:21
2.5 billion AI Overviews monthly users metric 4:04
1 billion AI Mode monthly users metric 4:04
900 million Gemini app monthly active users metric 4:49
$100 monthly price for new Ultra plan price 39:27

Chapters

  1. 0:00Opening and AI-first pivot
  2. 2:27Token growth and developer adoption
  3. 4:04Product usage and AI Overviews
  4. 4:49Gemini app growth and features
  5. 6:06Ask YouTube and Docs Live
  6. 9:47Infrastructure investment and TPUs
  7. 14:15World models and Gemini Omni
  8. 19:17SynthID and content credentials
  9. 21:36Gemini 3.5 Flash and Jules
  10. 32:05Gemini Spark and agentic era
Sundar Pichai 0:24 ↗
Hello Shoreline and hello to everyone watching from around the world. Excited to be back for this year's I/O. I think it's been an incredible year. All of the relentless shipping, the rapid advances in technology. It's been a period of progress. I've definitely felt it myself. It's been an intense year. Here's a look at what I've been up to.
OK, I wish this was my year. Actually, one of me plugging in the TPU is pretty accurate. So there's still more work to do before it's in space. We are working on that. On a serious note, it's been an extraordinary moment. It's been 10 years since we pivoted the company to be AI first. We knew how profound AI would be to advancing our mission and improving people's lives at scale. This is why we are taking a differentiated, full stack approach to AI innovation. From our custom silicon and secure foundation, to our world class research and models, to our products and platforms that reach billions of people. This approach enables us to iterate and innovate faster, and it's lighting up every part of the company.
What's really incredible is how people are using our AI. Students prepping for final exams with the Gemini app. Musicians and artists using generative AI models like Lyria and Veo as part of their creative flow. Developers coding and bringing their ideas to life. I've been using Gemini in a myriad of ways. Recently, I've been turning to Gemini to make sense of my parents' doctor visits. I'm sure many of you have done a version of that. These stories of how people are using AI are the best measure of progress to understand the scale at which people are adopting AI.
There's another great proxy: tokens. The fundamental units of data our models process, many representing a problem being solved. Two years ago, we were processing 9.7 trillion tokens a month across our surfaces, a huge number. Last year at I/O that grew to about 480 trillion tokens. And fast forward to today, that number has jumped seven times to 3.2 quadrillion tokens per month. Never imagined I'd say quadrillion in an I/O keynote, but here we are. Now, some out there might call this token maxing, and there's probably some truth to it. I still think it tells an important story about our products and how others are building as well, especially our developers.
Over 8.5 million of you are now building new apps and experiences with our models monthly, and our model APIs are now processing around 19 billion tokens per minute. Over the past 12 months, over 375 customers each processed more than one trillion tokens, representing incredible demand for AI across the industry. We are, of course, also seeing incredible demand across our products. We now have 13 products with over a billion users each. Five of those have more than three billion users.
Our Gemini models are a big reason more people are using our products and why they are using our products more. It all starts with search, which is bringing the benefits of generative AI to more people than any other product in the world. AI Overviews now has over 2.5 billion monthly users, and AI Mode has been a revelation. Our biggest upgrade to search ever. People love it. In just a year, it's already surpassed 1 billion monthly users. When people use our AI powered features in search, they use search more. I love how search has become less about individual queries and feels more like an ongoing conversation, giving you deeper insights and connecting you with the vastness of the web.
Another place we've been rapidly innovating is in the Gemini app. Last year at I/O, the Gemini app had 400 million monthly active users. Today we have surpassed 900 million. More than doubling in a year. In that same time, daily requests have grown over 7 times. It's incredible growth. We've been adding a lot of unique features like personal intelligence, which makes responses more customized and helpful. And to date, more than 50 billion images have been generated with our Imagen models. It was a breakout star this past year. I know you all have been having a lot of fun with it.
Beyond the Gemini app, we are also having much more natural conversations with Gemini directly inside many of our products. Recently, Maps got its biggest upgrade in a decade, including a new feature called Gemini in Maps. People are using it to ask more complex and much longer questions. Here's a real query from a parent: 'My kid just fell into the duck pond and the wedding starts in 30 minutes. Where can I walk and buy her a new dress?' I'd like to hear how that turned out.
We're also bringing this conversational AI to two more products. First, Ask YouTube. People come to YouTube every day to ask a lot of questions. There's a lot of great videos. Sometimes it's hard to know where to start. Ask YouTube entirely reimagines the experience. Say you want to teach your three-year-old how to ride a pedal bike, and they already know how to ride a balanced bike. Just ask YouTube. You'll see a couple of differences in results. The information is digestible and easy to navigate. You get an overview and helpful tips. You will see videos that best match your interests. So if you want to try a specific method of teaching, you can go deeper there. And best of all, it jumps right to the part of the video most relevant for you. Brings back memories of teaching my kids to ride. It remembers the context so you can follow up with questions like, 'Should I buy one with handbrakes or pedal brakes?' Making it an ongoing conversation. It even lays out the information in a table. So it's easy to compare. We are starting to test Ask YouTube now and it will roll out broadly in the US this summer.
So far, we have shown conversational text queries. There are a lot of times I want to get things done at the speed of my voice. There is much more possible today thanks to technical leaps in our audio models. A new feature called Docs Live takes this to another level. To create a doc with Gemini app before, you'd have to type up a really precise prompt. With Docs Live, you can just verbally brain dump whatever is on your mind and let Gemini do the rest. Let's see it in action with a demo from our product team. This is all in real time, not sped up.
In the future, you'll be able to create new docs and edit them directly, all with your voice. Docs Live is rolling out for Pro and Ultra subscribers this summer, and the same powerful voice capabilities will come to Gmail and Google Keep then too.
It's incredible to see the pace of innovation rolling out across our products. Supporting all of this at scale for our users, while also serving enterprises and developers around the world, requires massive investments in infrastructure. And we've been investing for today and for the future. In 2022, we were spending $31 billion annually in CapEx. This year, we expect that number to be about six times that, approximately $180 to $190 billion. A key part of this investment is our custom silicon. A decade ago, we announced our very first commercial Tensor processing unit, or TPU, on this I/O stage. Since then, we have transformed how the industry builds for AI. We recently announced our eighth generation of TPUs at Cloud Next. For the first time, we have taken a dual chip approach with specialized architectures for training and inference. TPU v6p and v6e. While they may look similar, they're actually pretty different. v6p is optimized for large scale pre-training, and it's nearly three times the raw computing power of our previous generation. We have taken a fundamentally different approach with our training infrastructure with Jupiter network pathways. Our training is no longer constrained by the limits of a single massive data center. Instead, we can now seamlessly distribute training across multiple sites, scaling across more than 100,000 TPUs globally. This gives us the ability to create the largest training cluster in the world for model builders. This means training larger, more capable models in weeks rather than months.
TPU v6e is designed for inference. We have dramatically improved speed at every step because if we learn anything in 27 years of working on search, it's that latency matters. To give you a live sense of what this speed feels like, here's a prompt on an upcoming Flash model. If it were running on v6e, I'll ask it to create a Chrome Dino game. Push submit. The response is generated in real time as you watch. Take a look at the tokens per second in the top right corner. The speed is pretty incredible, nearly 1,500 tokens per second. It almost took longer to write out the request, and the game is pretty fun too.
In addition to speed, we are also thinking about scaling sustainably. Both chips are more energy efficient, delivering up to two times better performance per watt. TPUs have been hard at work training for I/O this year. I'm told we have a behind the scenes look.
Our compute innovations enable our advances. There are three areas where I want to go deeper today to show you the progress in each: models, coding, and agents. Let's start with the exciting progress in world models. With world models, AI is moving from predicting text to simulating reality. Demis and the team at Google DeepMind have been working to push the boundaries of what these models can do. Let me invite Demis out to share more.
Demis Hassabis 14:58 ↗
Hi, everyone. It's really great to be here. Over the past year, AI capabilities have leaped forwards. We now have agents that can plan and act on our behalf. And artificial general intelligence is just a few years away. Today, I'm excited to share the progress we've made towards building AGI. Last year, I outlined our vision of extending Gemini's incredible multimodal capabilities to become a world model AI that can understand and simulate the world. This is a crucial aspect of achieving AGI, and it will be important for everything from building AI assistants to training robots. Now we're taking the next big step. I'm excited to announce Gemini Omni. Our new model that can create anything from any input. It combines Gemini's intelligence with the best of our generative media models for a new level of world understanding, multi-modality, and editing. Models like Veo, Imagen, and Genie are able to create extremely realistic videos, images, and interactive simulations. Although not perfect, they already demonstrate some impressive notions of intuitive physics. And with Omni, we've now made even more progress. It's a step change in simulating things like kinetic energy and gravity. Previous systems would have found these concepts difficult. Gemini's world knowledge and reasoning really shine in Omni. It can translate complex ideas into highly accurate videos. So for example, you can give it a simple prompt: 'Make a claymation explainer of protein folding,' and get this. Proteins start as chains of amino acids, they fold into patterns like the alpha helix and sections called beta sheets, forming a perfect 3-dimensional shape. But the initial generation is just the start. The creative process is rarely a single step. It's usually iterative, just like Imagen redefined image editing. Omni gives you a more natural way to edit video with conversational language. What's really cool is you can give it your own videos, for example, this selfie, and change reality in a really fun way. You can easily adjust the details and style or even add elements, and the whole scene morphs to reflect your new idea. A simple circle turns into a black hole, or an evening stroll comes to life. Anything becomes a canvas for creating entirely new realities. Let's take a look at what Omni can do.
But over time, Omni will be able to generate any output from any input. This was always our goal with Gemini and why we built it to be multimodal from the very start. It was a harder path, but the foundation is now paying off. Today we're launching the first model in the Omni family: Gemini Omni Flash. It's now available across our products and you'll hear more about this later. We're excited with the progress we're making, and we'll be able to share more about Omni Pro soon. We can't wait to see what you create. Back to you, Sundar.
Sundar Pichai 19:17 ↗
Thanks, Demis. This is huge progress. As generative AI gets better, so does the need for greater transparency. Research shows people can correctly identify high quality deepfake videos only about a quarter of the time. Three years ago, we launched SynthID, our watermark that is invisible to the naked eye. Since launch, it is now watermarked over 100 billion images and videos, along with 60,000 years of audio assets. Millions of people are using our SynthID detector in the Gemini app to verify AI generated content. We are now going a step further and adding content credentials verification across products. This will show if the origin of the content was AI or a camera, and if it's been edited with generative AI tools. In this example, Gemini can tell this photo was captured with the Pixel camera and then edited with Google Photos. We want more people to have easy access to these tools, so we are expanding both SynthID and content credentials verification to Search and Chrome. You can simply circle to search or right click in Chrome and ask, 'Was this generated with AI?' And you'll get a clear response along with other helpful context. For example, this image was making the rounds on social media last year. It's obviously fake. I don't eat hamburgers. It might not be as clear to everyone else. That's where these tools can be really useful. Of course, this only works at scale if more partners decide to watermark their own AI generated content. NVIDIA signed on to SynthID last year, and today I'm thrilled to announce that OpenAI, Kakao, and ElevenLabs are adopting SynthID too. It's great to see the cross-industry collaboration. We are looking forward to expanding to more partners and setting the standard of transparency for the AI era.
That's a look at the progress we are making with world models. Now let's talk about what's next for our Gemini 3 family. Gemini 3 launched a few months ago with a full family of models. It's been our most adopted series yet. We have loved seeing developers use Flash as their daily driver and build incredible experiences with Pro's deep reasoning and multimodal capabilities. We've been hard at work on improving these models, especially focused on agentic coding, long horizon tasks, and real world workflows. And today, I'm excited to introduce Gemini 3.5 Flash, our first in a series of models combining frontier intelligence with action. Two things I would highlight. First, when compared to 3.1 Pro, Flash is better across the board on almost all benchmarks. It's made huge progress in coding. And look at that extraordinary jump in GPQA, a benchmark that captures many real world, economically valuable tasks. Second, 3.5 Flash is a very capable model at the frontier and comparable to the best models, but much, much faster. You look at the intelligence versus output speed, it's in a whole league of its own in the top right quadrant. When looking at output tokens per second, it's four times faster than other frontier models, and it's an incredible delight to use. The new model has been a game changer for us internally at Google. We've been using 3.5 Flash with the reimagined version of our agent first development platform, Jules, and it's dramatically accelerated how we build. In March, we were processing half a trillion tokens a day internally for our developers. We've been doubling every few weeks, and now we are processing more than 3 trillion tokens a day. This scale has created a powerful feedback loop helping us improve 3.5. And of course, we are bringing it today to developers in Jules. Varun is going to share more.

4 more exchanges in this transcript

Sign in free to read the rest of this interview. No card required.

Sign in to read the full transcript

Cite this transcript

APA, MLA, BibTeX
APA

Pichai, S. (2026, May 22). Sundar Pichai Opening Remarks | I/O 2026 Keynote [Interview transcript]. Google and Sundar Pichai. CEOInterviews.AI. https://ceointerviews.ai/interview/928421/

MLA

Sundar Pichai. "Sundar Pichai Opening Remarks | I/O 2026 Keynote." Google and Sundar Pichai, 22 May. 2026. Transcript, CEOInterviews.AI, https://ceointerviews.ai/interview/928421/.

BibTeX
@misc{pichai2026_928421,
  author       = {Sundar Pichai},
  title        = {Sundar Pichai Opening Remarks | I/O 2026 Keynote},
  howpublished = {Interview transcript, Google and Sundar Pichai. CEOInterviews.AI},
  year         = {2026},
  month        = {may},
  url          = {https://ceointerviews.ai/interview/928421/},
  note         = {Speaker-attributed transcript with timestamps}
}