CEOInterviews.AI
Start App
Sundar Pichai
Chief Executive Officer & Director, Google

Google I/O 2025 keynote in 32 minutes

📅 May 20, 2025 The Verge 32 MIN 1544793 VIEWS 27 SEGMENTS · 2 SPEAKERS
Google just wrapped up its big keynote at I/O 2025. As expected, it was full of AI-related announcements, ranging from updates across Google’s image and video generation models to new features in Search and Gmail. We got a look at Project Astra’s latest prototype that will complete tasks on your behalf, Gemini 2.5 Pro updates, a new AI filmmaking app called Flow, Project Starline is now Google Beam. Plus, a look at Xreal’s new Project Aura prototype. Here’s everything you missed. 0:00 Intro 0:10 Ironwood 0:26 Google Beam/Project Starline 1:06 Speech Translation in Google Meet 1:55 Project Ast...
Sundar Pichai 0:00 ↗
Hello everyone. Good morning. I learned that today is the start of Gemini season. I'm not really sure what the big deal is. Every day is Gemini season here at Google. Our 7th generation TPU, Ironwood, it delivers 10x the performance over the previous generation and packs an incredible 42.5 exaflops of compute per pod. And it's coming to Google Cloud customers later this year.
Introducing Google Beam, a new AI-first video communications platform. An array of six cameras captures you from different angles. And with AI, we can merge these video streams together and render you on a 3D light field display. With near-perfect head tracking down to the millimeter and at 60 frames per second, all in real time. The result, a much more natural and deeply immersive conversational experience. In collaboration with HP, the first Google Beam devices will be available for early customers later this year.
We've been bringing underlying technology from Starline into Google Meet. That includes real-time speech translation to help break down language barriers. Let me turn on speech translation.
It's nice to finally talk to you. You're going to have a lot of fun and I think you're going to love visiting the city. The house is in a very nice neighborhood and overlooks the mountains. And today we are introducing this real-time speech translation directly in Google Meet. English and Spanish translation is now available for subscribers with more languages rolling out in the next few weeks and real-time translation will be coming to enterprises later this year.
We are starting to bring it to our products today. Gemini Live as Project Astra's camera and screen sharing capabilities so you can talk about anything you see. That's a pretty nice convertible. I think you might have mistaken the garbage truck for a convertible. Is there anything else I can help you with? What's this skinny building doing in my neighborhood? It's a street light, not a building. Why is this person following me wherever I walk? No one's following you. That's just your shadow. Gemini is pretty good at telling you when you're wrong. We are rolling this out to everyone on Android and iOS starting today.
We also have our research prototype, Project Mariner. It's an agent that can interact with the web and get stuff done. First, we are introducing multitasking and it can now oversee up to 10 simultaneous tasks. Second, it's using a feature called teach and repeat. This is where you can show it a task once and it learns a plan for similar tasks in the future. We are bringing Project Mariner's computer use capabilities to developers via the Gemini API and it will be available more broadly this summer.
Then there is the Model Context Protocol introduced by Anthropic so agents can access other services. And today we are excited to announce that our Gemini SDK is now compatible with MCP tools. And we are starting to bring agentic capabilities to Chrome, Search, and the Gemini app. Let me show you what we are excited about in the Gemini app. We call it agent mode. Normally, you'd have to spend a lot of time scrolling through endless listings. Using agent mode, the Gemini app goes to work behind the scenes. It finds listings from sites like Zillow that match your criteria and uses Project Mariner when needed to adjust very specific filters. If there's an apartment you want to check out, Gemini uses MCP to access the listings and even schedule a tour on your behalf. An experimental version of the agent mode in the Gemini app will be coming soon to subscribers.
Personalization will be really powerful. We are working to bring this to life with something we call personal context. With your permission, Gemini models can use relevant context across your Google apps in a way that is private, transparent, and fully under your control. You might be familiar with our AI-powered smart reply features. Now, imagine if those responses could sound like you. That's the idea behind personalized smart replies. Let's say my friend wrote to me looking for advice. Now, if I'm being honest, I would probably reply something short and unhelpful. Sorry, Felix. But with personalized smart replies, I can be a better friend. That's because Gemini can do almost all the work for me. Looking up my notes in Drive, scanning past emails for reservations, and finding my itinerary in Google Docs, Gemini matches my typical greetings from past emails, captures my tone, style, and favorite word choices, and then it automatically generates a reply. This will be available in Gmail this summer for subscribers.
Gemini Flash is our most efficient workhorse model. Today, I'm thrilled to announce that we're releasing an updated version of 2.5 Flash. The new Flash is better in nearly every dimension, improving across key benchmarks for reasoning, code, and long context. I'm excited to say that Flash will be generally available in early June with Pro soon after, but you can go try out the preview now in AI Studio, Vertex AI, and the Gemini app.
We are also introducing new previews for text-to-speech. These now have a first-of-its-kind multi-voice support for two voices built on native audio output. This means the model can converse in more expressive ways. It can capture the really subtle nuances of how we speak. It can even seamlessly switch to a whisper like this. This works in over 24 languages and it can even easily go between languages. So the model can begin speaking in English but then switch back all with the same voice. You can use this text-to-speech capability starting today in the Gemini API.
Second, we've strengthened protections against security threats like indirect prompt injections. So Gemini 2.5 is our most secure model yet. And in both 2.5 Pro and Flash, we're including thought summaries via the Gemini API and Vertex AI. Thought summaries take the model's raw thoughts and organize them into a clear format with headers, key details, and information about model actions like tool calls. Finally, we launched 2.5 Flash with Thinking Budgets to give you control over cost and latency versus quality. So, we're bringing Thinking Budgets to 2.5 Pro, which will roll out in the coming weeks along with our generally available model.
You've seen something like this before, right? Someone comes to you with a brilliant idea scratched on a napkin. I am going to add the image I just showed you of the sphere and I'm going to add in a prompt that asks 2.5 Pro to update my code based on the image. And I'm going to jump to another tab that I ran right before this keynote with the same prompt. And here's what Gemini generates. Whoa. We went from that rough sketch directly to code updating multiple of my files. And actually, you can see it thought for 37 seconds and you can see the changes it thought through and then the files it updated. But what if it talked? That's where Gemini's native audio comes in. That's a pangolin and its scales are made of keratin just like your fingernails. 2.5 Pro is available on your favorite IDE platforms and in Google products like Android Studio, Firebase Studio, Gemini Code Assist, and our asynchronous coding agent, Jules. Jules can tackle complex tasks in large code bases that used to take hours, like updating an older version of Node.js. It can plan the steps, modify files, and more in minutes. So today, I'm delighted to announce that Jules is now in public beta, so anyone can sign up at jewels.google.
Gemini Diffusion is a state-of-the-art experimental text diffusion model. The version of Gemini Diffusion we're releasing today generates five times faster than even 2.0 Flash Light, our fastest model so far, while matching its coding performance. So take this math example. Ready? Go. If you blinked, you missed it. Today, we're making 2.5 Pro even better by introducing a new mode we're calling Deep Think. It gets an impressive score on USAMO 2025, currently one of the hardest math benchmarks. We're taking a little bit of extra time to conduct more frontier safety evaluations and get further input from safety experts. As part of that, we're going to make it available to trusted testers via the Gemini API to get their feedback before making it widely available.
Gemini is already the best multimodal foundation model, but we're working hard to extend it to become what we call a world model. This is our ultimate vision for the Gemini app to transform it into a universal AI assistant. This starts with the capabilities we first explored last year in Project Astra, such as video understanding, screen sharing and memory. For example, we've upgraded voice output to be more natural with native audio. We've improved memory and added computer control. Let's take a look. I think I stripped this screw. Can you go on YouTube and find a video for how to fix that? Of course. I'm opening YouTube now. This looks like a good video. It seems like I need a spare tension screw. Can you call the nearest bike shop and see what they have in stock? Yep. Calling them now. I'll get back to you with what they have in stock. Hey, any updates on that call? Yep. I just got off the phone with the bike shop. They confirmed they have your tension screw in stock.
We've built AlphaProof that can solve Math Olympiad problems at the silver medal level. Co-Scientist that can collaborate with researchers, helping them develop and test novel hypotheses. And we've just released AlphaEvolve, which can discover new scientific knowledge and speed up AI training itself. In the life sciences, we've built AMIE, a research system that could help clinicians with medical diagnosis. AlphaFold 3, which can predict the structure and interactions of all of life's molecules. and Isomorphic Labs which builds on our AlphaFold work to revolutionize the drug discovery process with AI and will one day help to solve many global diseases.
As people use AI Overviews, we see they are happier with their results and they search more often. For those who want an end-to-end AI search experience, we are introducing an all-new AI Mode. It's a total reimagining of search. With more advanced reasoning, you can ask AI Mode longer and more complex queries. In fact, users have been asking much longer queries, two to three times the length of traditional searches. And I'm excited to share that AI Mode is coming to everyone in the US starting today. Soon, AI Mode will be able to make your responses even more helpful with personalized suggestions based on your past searches.
You can also opt in to connect other Google apps starting with Gmail. Now, this is always under your control, and you can choose to connect or disconnect at any time. Personal context is coming to AI Mode this summer. You already come to search today to really unpack a topic, but this brings it to a much deeper level. So much so that we're calling this deep search. It reasons across all those disparate pieces of information to create an expert level fully cited report in just minutes.
So I'll ask, show the batting average and on-base percentage for this season and last for notable players who currently use a torpedo bat. I get this helpful response, including this easy to read table. I can follow up and ask, 'How many home runs have these players hit this season?' Search figured out that the best way to present this information is a graph and it created it. It's like having my very own sports analyst right in search. Complex analysis and data visualization is coming this summer for sports and financial questions.
So, I'm excited to share that we're bringing Project Mariner's agentic capabilities into AI Mode. I'll say find two affordable tickets for this Saturday's Reds game in the lower level. Search kicks off a query fan out looking across several sites to analyze hundreds of potential ticket options. Search helps me skip a bunch of steps linking me right to finish checking out. We're taking the next big leap in multimodality by bringing Project Astra's live capabilities into AI Mode. We call this Search Live. It's like hopping on a video call with search.
We are introducing a new try-on feature that will help you virtually try on clothes. I really like this blue one. I click on this button to try it on. It asks me to upload a picture which takes me to my camera roll. I have many pictures here. I'm going to pick one that is full length and a clear view of me. We built a custom image generation model specifically trained for fashion. Wow. And it's back. The AI model is able to show how this material will fold and stretch and drape on people. And let's assume the price is now dropped. When that happens, I get a notification just like this. And if I want to buy, my checkout agent will add the right size and color to my cart. I can choose to review all my payment and shipping information or just let the agent just buy it for me with just one tap. Search securely buys it for me with Google Pay. And of course, all of this happens under my guidance. Our new visual shopping and agentic checkout features are rolling out in the coming months and you can start trying on looks in Labs beginning today.

1 more exchange in this transcript

Sign in free to read the rest of this interview. No card required.

Sign in to read the full transcript

Cite this transcript

APA, MLA, BibTeX
APA

Pichai, S. (2025, May 20). Google I/O 2025 keynote in 32 minutes [Interview transcript]. The Verge. CEOInterviews.AI. https://ceointerviews.ai/interview/626967/

MLA

Sundar Pichai. "Google I/O 2025 keynote in 32 minutes." The Verge, 20 May. 2025. Transcript, CEOInterviews.AI, https://ceointerviews.ai/interview/626967/.

BibTeX
@misc{pichai2025_626967,
  author       = {Sundar Pichai},
  title        = {Google I/O 2025 keynote in 32 minutes},
  howpublished = {Interview transcript, The Verge. CEOInterviews.AI},
  year         = {2025},
  month        = {may},
  url          = {https://ceointerviews.ai/interview/626967/},
  note         = {Speaker-attributed transcript with timestamps}
}