Back
Lisa Su
Chair, President & Chief Executive Officer, Advanced Micro Devices

AMD Advancing AI 2026: Lisa Su Full Keynote

🎥 Jul 23, 2026 📺 Yahoo Finance ⏱ 131m
AMD CEO Dr. Lisa Su takes the stage at the Advancing AI 2026 conference to outline the company's vision for the future of AI, ...
Watch on YouTube

About Lisa Su

Lisa Su, chair and CEO of AMD, spoke at the company's Advancing AI 2026 conference in San Francisco on July 22-23, 2026, where she announced new products and partnerships. Su stated that AI infrastructure demand is accelerating, not slowing, and described inference as the industry's largest growth driver. She said AMD expects the AI accelerator market to reach approximately $1.4 trillion by 2030, approaching the size of the entire semiconductor market today. Su also announced a $5 billion investment in Anthropic and discussed the Helios rack-scale architecture, which she said offers 10 to 15% more performance and up to 30% better tokens-per-dollar for customers compared to competitors. She reiterated her view that no single chip company will dominate the AI market, arguing that the world is heterogeneous and requires an ecosystem of different compute engines. In a commencement address at MIT on May 28, 2026, Su told graduates that "we may discover more in the next 10 years than we have in the last 30," but emphasized that "technology itself does not decide what the future looks like. The best people do." She advised graduates to "run towards the hardest problems" and said that "luck is not just being in the right place at the right time. It is taking the risk to work on something really hard." Su reflected on her decision to become CEO of AMD 12 years ago, describing it as her dream job, and noted that the company made a long-term bet that high-performance computing would be the most important technology of the future.

Source: AI-verified profile updated from Lisa Su's recent appearances. Browse all interviews →

Transcript (155 segments)
L
Lisa Su0:02
Good morning.
A
Audience0:03
Morning.
L
Lisa Su0:05
That's a pretty good morning. How about let's try one more. Good morning.
A
Audience0:09
Good morning.
L
Lisa Su0:11
And welcome to Advancing AI 2026. It's so great to be back here in San Francisco and to see so many friends, partners, and customers, and especially all the developers that are here with us today. And I want to say a big welcome to everyone who's joining us online from around the world. This is my absolute favorite event of the year. It's where we bring the entire AI ecosystem together to show what we've been building and where we're going next. And this year, this is our biggest show ever because we have so much to tell you. So, let's go ahead and get started. At AMD, our mission is to push the boundaries of high performance and AI computing to help solve the world's most important challenges. I often say AI is the most important technology of the last 50 years. The progress we're seeing is incredible. Every month, every few months, we see something new that was far beyond our imagination. You can already see the impact across every industry. In healthcare, AI is helping researchers identify new drug candidates faster. In science, we're solving problems that we didn't think were possible a few years ago. And across every industry, AI is changing the way we get work done. The most important thing is we're still in the very early innings of what's possible.
Now, just take a look at some of these charts. We look at these every year and you kind of see the rate and pace of AI growth. A couple of years ago, we were just starting. People were experimenting with chatbots and it was really cool. But if you look at today, we're seeing more than 35 quadrillion tokens consumed every single month. That's an increase of nearly 160 times in just two years. That curve is just getting steeper. You're going to hear a lot about that today. So just looking at where all that demand is, obviously training remains incredibly important and foundational. Over the last four or five years, we see the compute used for training advanced models continuing to increase by roughly 5x every year, and those models are getting better. We're seeing better reasoning, new capabilities, and a growing number of specialized models. One of the things I really believe in is there is no one perfect model. We're going to use a slew of different models depending on what task you're trying to solve and which industry you're in. But the bigger shift is actually in inference.
And we expected this. We expected that inference would grow faster than training. We can say that in 2026 for the first time the world is using more AI compute to run models than to train them. We think this year roughly 60% of global AI compute capacity will be used for inference. The reason is simple: as billions of people use AI every day, the workload shifts from building models to putting them to work. We're going to talk a lot today about agentic AI, which is accelerating the shift even faster. We see agents as the next big step for AI. This has really just started accelerating over the last five or six months. When we started with LLMs, they were great for answering questions, but to really get the power of AI, you want agents to be able to answer full questions. You give it a goal, it keeps working until it solves the problem. That means compute demand is growing at an incredible pace. We're seeing a step change because when you ask an agent to do something, it has dozens of steps, it has to reason, call tools, access data, and keep doing it over and over until it solves the problem. So you need lots of GPUs for reasoning, but importantly, you also need a lot of CPUs to orchestrate every step around it.
So when you look at what that means from a market standpoint, it's hard to put this market data up because every few months we're changing the perspective based on what our customers are saying. But what we're seeing is that the shift to agentic AI is growing the AI accelerator market very significantly. Last year at this event, we called the AI accelerator market at about $500 billion by 2028. At that time it felt like a very big number. But honestly, you look at it today and it looks conservative because AI demand is continuing to accelerate. Better models, more usage. More usage requires more compute, and more compute builds better models. So we now expect the AI accelerator market to reach about $1.4 trillion by 2030. What that means is by the end of the decade, the AI accelerator market will approach the size of the entire semiconductor market today. Although there will be many different types of accelerators, I'm a big believer that there's no one-size-fits-all. We do expect GPUs to make up the vast majority of that market because algorithms are still in their infancy and workloads continue to change, favoring programmability in the overall silicon ecosystem.
Now perhaps the most interesting part of the last six months has been that GPUs are only part of this story. Agentic AI is creating an entirely new growth vector for server CPUs. I've updated this number a lot over the last six to eight months, but we call it like we see it. We were early. We saw from our largest hyperscale customers that as AI inference goes up, they need more server CPUs. So we saw growth in the server CPU market. Last year we thought it might grow at 18 to 20% CAGR to an overall market size of $60 billion. But the rate and pace of agentic AI adoption is much faster than any of us thought. We're talking about agents going from millions to billions. That requires a tremendous amount of CPU infrastructure. Every customer conversation tells us that this buildout is just beginning. Based on what we're seeing today, we now expect the server CPU market to grow by over 50% to over $200 billion by 2030. That's starting from today's $25 billion market. So there's a lot of excitement about CPUs as well.
Now probably one of the things differentiating AMD is that this is not just a data center opportunity. As AI becomes a larger part of our daily lives, we want intelligence to run right where the work happens. Cloud is super important, but the devices we use every day are going to be extremely important. We're going to do a lot more at the edge as well with autonomous machines that sense and act in real time. We believe AI needs to be infused everywhere, and that's exactly what we're focused on at AMD. Putting all of that together, we expect the market for our high-performance and adaptive computing products to grow at roughly a 40% CAGR over the next several years, approaching a $2 trillion market by 2030. The only way to service a market like that is for us to work together as an ecosystem. No one company can solve it all. But this is the opportunity to bring the best and brightest together. We have never been in a better position to lead. We have the broadest product portfolio, the strongest roadmaps we've ever had, and the deepest partnerships with the companies building this future.
So a little bit about our strategy. We've been very consistent. Our multi-year strategy is built around three priorities. First, compute leadership: we're building the broadest set of compute engines in the industry so that we have the right compute for the right workload. Second, open platforms: we want everyone to come together in an open ecosystem. We believe an open ecosystem is essential to the future of AI. That's how we get the force multiplier of everyone coming together. That's open on hardware and open on software. That's why we're investing heavily in our ROCm software. AI has been a tremendous multiplier in the rate and pace of progress we're making in ROCm and software. It means that if we put all of this together, our customers can deploy on AMD hardware faster than ever. Third, powering AI everywhere: AI in enterprise, AI at the PC, AI in the physical world, and adding new capabilities across all our products. Today, you're going to see all of that in action. We're going to start with compute for the agentic era, show you a lot of hardware, talk about software, and our overall products. We're honored to have some special guests to help bring the technology to life. So, let's start with the data center.
Today, AMD EPYC runs on the most important workloads in the world. We power the largest cloud providers, the digital platforms that billions of people use every day, and the most important critical systems used by the largest businesses. More than 60% of the Fortune 100 use AMD EPYC, and that momentum is building. Last quarter, we reached a record 46% revenue share in the server market. We're continuing to see every major customer move more workloads to EPYC. On the GPU side, our Instinct adoption has also accelerated. We have a broad set of customers across the largest AI labs, cloud providers, leading AI startups, and national labs. Sovereign AI opportunities are being built on AMD Instinct, and we appreciate that opportunity.
Now, AI compute is becoming a lot more complicated. What we're seeing with frontier AI is that it's raising the bar for what infrastructure needs to do. It takes more than a single chip or a single server. You have to design the entire rack as one system. That requires leading CPUs, leading GPUs, high-speed networking that connects everything inside the rack and across the data center. And just as importantly, this is about making these systems super easy to use. They need to be easy to deploy, easy to service, and very reliable. That is exactly what we built with Helios. So today, I'm super excited to launch Helios, the industry's highest performance AI rack.
Helios is built to train and run the most demanding frontier models in the world at massive scale. I have a lot of show and tell for you. For those who know me, I love holding up chips. I call myself sometimes the Vanna White of chips. But with Helios there are lots of components. These are the chips inside the Helios rack: Instinct, EPYC, and Pensando chips for networking. Every one of these chips is part of the system. MI455 is the engine, the highest performance GPU in the industry. Venice which drives it is the world's fastest CPU cores. Pensando DPUs and Nyx connect them with leadership programmability and scale-out bandwidth. Let's look at how these fit into the system. This is the MI455 accelerator mounted on an enhanced accelerator module (EAM). It's a production module that goes into Helios, much more than the GPU package. The EAM integrates the GPU, memory, power delivery, high-speed interfaces, system management, and cold plates for liquid cooling into one compact serviceable module. Pretty cool, right? Specs: 320 billion transistors, built on TSMC's leading 2nm and 3nm process technology. It brings together nearly a decade of AMD chiplet innovation, combining 12 compute and IO chiplets with 432 GB of memory, all connected by leading industry 3D chip stacking packaging. It's the highest performance AI accelerator in the industry.
This is the CPU board inside the Helios compute tray. It's a high-speed 96-core EPYC processor with DDR5 memory and all the IO needed to feed the GPUs, all in a single motherboard. We also have our Selena DPU, a critical part of the networking infrastructure that allows us to deliver both front-end and scale-out and scale-up bandwidth. This is our Volcano AI NIC, which allows up to six Volcano NICs on a board, using the open Ultra Ethernet standard. Each Helios compute tray includes two of these Volcano boards for scale-out and scale-up networking. Inside the rack, we have an additional six dedicated networking trays that handle scale-up networking, connecting 75 GPUs together with UA link over Ethernet with silicon from our partners. That's what we mean by an open ecosystem. When you bring it all together, Helios is simply the best AI rack in the world.
Let's take a look at some numbers. It's a very competitive world. When you compare Helios to the competition, we're delivering 15% more compute, 50% more HBM4 memory capacity and bandwidth, and 50% more scale-out bandwidth. That means every Helios delivers more performance for the largest models, more capacity for longer context, and the bandwidth to scale across thousands of racks. Here is the production hardware that customers are deploying. Each tray weighs more than 160 lbs and stands less than 2 inches tall. The numbers are incredible: more than 18,000 CDNA 5 GPU compute units, over 4600 Zen 6 CPU cores, and 31 terabytes of HBM4 memory in a single rack. That's what it takes to run agentic AI at scale.
Today, I'm excited to announce that Helios is in full production. Shipments are on track to start at the end of the third quarter, ramping into the fourth quarter and the second half of the year. Customer demand for Helios is extremely strong. We're proud of the work done across the leading AI labs to adopt Helios. Let's start with some of our guests today. I'm excited to be joined by our newest Instinct partner and one of the leaders in frontier AI. Please welcome Anthropic co-founder and chief compute officer Tom Brown.
Hello Tom. It is so great to have you here and it is really such an honor to be working with Anthropic. We just had a big week this week.
T
Tom Brown19:51
We did. Yeah.
L
Lisa Su19:52
So, look, Anthropic has just had incredible success over the last few years. Claude has made such a big impact on the overall industry. Compute is such an important factor in progress. Can you talk a little bit about your compute strategy and how things have been evolving?
T
Tom Brown20:10
Yeah. Yeah. And first, thank you so much. A hard act to follow with Helios. Amazing. So with Anthropic, as you mentioned, the scale of growth in the industry is enormous. We've been working to make sure we have the best chips and can use the best chips for the best workloads.
L
Lisa Su20:37
Yeah. Look, I think you've been really leading in that area. I'm happy to say that I've always wanted to be part of your infrastructure. I remember the early times we talked as you were starting Anthropic. We had a big announcement this week: we're honored that you will be deploying up to 2 gigawatts of Helios. Why now? What was the key reason?
T
Tom Brown20:59
Yeah. I think the first thing was just Helios is an amazing machine. It's absolutely fantastic.
L
Lisa Su21:08
Do we like that? Does that work for us?
T
Tom Brown21:13
And then the thing we were thinking about originally was that bringing up a new hardware platform is a big effort. It's a huge thing. As part of that, we started our own evaluation of MI355. You guys generously got us a rack to start working. We expected a big process. Our actual experience was we had one engineer start doing it. They spun it up, connected it to Claude, asked it to bring up the machine, left it going over the weekend, and we ended up with a graph of the performance of our leading model on it just going up and up and up over the weekend. And so yeah, I think that's a testament to the open platform you've made, where anyone, human or AI, can now build real models on your platform.
L
Lisa Su22:15
Tom, I kid you not, when my team told me that you had an engineer working on MI355, I said, 'Hey, are we making sure we help them?' And they said, 'They say they don't need any help. They got it all covered.' I said, 'That's a wonderful story. Thank you for that.' Now, we're also super excited about the work we're doing together because one of the most valuable parts of this partnership is Claude and using Claude within AMD. We're big believers that AI is a force multiplier, and the more we have leading foundational models familiar with AMD architecture, the better we'll be able to service the overall market. So can you talk a little bit about how you see Claude evolving in these highly engineering-specific workloads?
T
Tom Brown23:02
Yeah. So I think this is a great place for collaboration. The bread and butter of Claude is doing software engineering, and now more and more we're seeing it help with workloads like design layout, which is adjacent to software engineering. I think that's where we can work together to make the next generation of chips even better.
L
Lisa Su23:30
Yeah. No, I think on the software side, the work we've seen with Claude and kernel development has been incredible. So, Tom, one of the things for our audience here: we like to talk about the present, but most of us are working on the future. This is the beginning of what we believe is a very strong multi-year partnership. Can you talk about what you're most excited about in the industry and what we can do more together to satisfy those key opportunities?
T
Tom Brown24:08
Yeah. That's a great question. I think working together for the scale-up is probably the biggest thing I'm most excited about. We see consistently that using more and better compute results in better models that add more value to downstream users. That's the biggest thing. Also, now that we're investing more with you, security is a place where we can work together. As machines get more complex, we can secure things at the chip layer, server level, rack, and whole network. That will help not just us but the entire industry stay safe.
L
Lisa Su24:56
No, I think you're completely right. We are super focused on how we can use AI to improve every aspect of our product, whether it's performance, power, software, or security as you said. So, Tom, huge thank you for the opportunity to work closely with you and your team. Our team loves working with you guys, and we look forward to everything we're going to do together.
T
Tom Brown25:21
Thank you so much.
L
Lisa Su25:22
Thank you.
All right. So, Tom talked about how important performance is. We have been laser focused on ensuring that the real-world performance of MI455 and Helios comes true. Let's take a look at some results. Here's an example running DeepSeek V4 Flash, a reasoning model. We're comparing MI455 to MI355. With every generation, we want huge leaps in performance. At lower concurrency, MI455 delivers 4x the throughput of MI355. But at the highest concurrency, Helios delivers up to 34 times more throughput than our prior generation. That means more users, much faster response, and much better efficiency at scale. Performance is just one aspect. The other is cost per token. Customers care about cost per token, and with every generation of Instinct, we're driving it down. With MI455, we're delivering up to 18 times more tokens per dollar, so customers can serve far more users with the same investment. And when you think about how that comes together in a rack, many data centers are limited by power. We tested Helios over a full range of workloads at fixed rack power compared to the competition, and we're seeing an average of 10 to 15% more performance across leading inference workloads. That translates to up to 30% more tokens per dollar than the competition. That's a pretty good value proposition.
Now I'm happy to say that customer demand for Helios is extremely strong. From the largest AI labs to hyperscale and neocloud providers, we are working across the entire ecosystem to enable Helios together with our OEM and ODM partners. One of our deepest and earliest partners deploying Helios is OpenAI. To talk about where AI is headed and the work we're doing together, please welcome OpenAI's head of infrastructure, Sachin Kati.
Thank you, Sachin. It's so great to have you here. Most importantly, we are so thankful for the partnership with OpenAI. We've been through so much on this journey together. You're really operating on the frontier of AI. Can you tell us a little bit about what you're seeing and what compute looks like for you?
S
Sachin Kati29:23
No, thanks. Great to be here. I was thinking about this event over the last two years. It feels like attendance is following scaling laws, doubling every year. But more seriously to your question, we've always predicted that models will evolve from chatbots to reasoners to agents to interns. That trajectory has kept up. We are seeing models become more capable, more agentic, able to do long-running tasks. All that means we need to keep scaling compute. The more compute we scale for training, the more capabilities emerge. The more compute we scale for distributing intelligence, the more people want to do with it. That's what we see with exploding token budgets. Internally at OpenAI, not just engineering but every aspect of the enterprise is beginning to use agents for every part of work. This is why we need platforms like yours that can scale with how quickly capabilities are scaling and how quickly the world is embracing these things.
L
Lisa Su30:47
Sachin, I don't think I've ever spoken to you where you haven't asked me for more compute. You're very consistent. Look, our teams have been working really closely together. It's been a true journey thinking about all the technology we've been doing. Last year we announced a landmark partnership where you're going to deploy up to 6 gigawatts of AMD infrastructure. You were one of the first to have MI455 racks a few months ago. Can you talk about the journey?
S
Sachin Kati31:23
It's been a phenomenal partnership. As you mentioned, we were first to bet on AMD, and we are thrilled with how that bet is turning out. We started on MI300 and expanded very quickly to MI355. It's been fun for our team to see how we could optimize the software stack, networking, and everything to deploy our models on AMD infrastructure. As Lisa mentioned, we just got hands on the Helios racks three months ago, and the collaboration has gone to the next level. Our engineers are working side by side with AMD engineers to optimize the software stack and run GPT-class workloads on Helios already. Phil will probably be on stage later talking about how productive and fast that collaboration is. We're really excited about the capabilities Helios is already showing us, and we expect to be deploying Helios at massive scale starting towards the end of this year and accelerating throughout 2027.
L
Lisa Su32:31
Like I said, I know more faster.
S
Sachin Kati32:33
We need it earlier.
L
Lisa Su32:37
So, Sachin, one of the most exciting parts of our collaboration is some of the research work. Thankful for the opportunity to bring our best engineers together with your researchers to talk about how AI should be implemented and used for future silicon and software. Can you talk a little bit about that work?
S
Sachin Kati33:01
Yeah. I think one of our key bets is about recursion: AI recursively improving the systems it needs to run on, and eventually AI doing the research itself. Today, expert engineers spend enormous time optimizing kernels, translating new models to run efficiently on different hardware, dealing with compilers, communications libraries, system configurations. What we've been super excited and encouraged by is the work we are doing with you, where AI is showing the potential to automate a significant portion of the work when anyone programs to an AMD GPU. It dramatically shortens the path from a new model coming up to efficient production deployment, and something we can much more easily tune to changing workload requirements. That's a big deal. Also, because AMD is built around an open software ecosystem, all these things we're inventing internally for our own use, we believe we can bring to the whole world and allow the whole world to use them for all AI models, not just ours.
L
Lisa Su34:26
I think super excited about that. It's really the tide that lifts all boats. As we work together, it helps our work together but also helps the overall ecosystem. So a little bit of your crystal ball looking forward: we all know that compute is super critical. What do you need from the next generation AI infrastructure, and what can we do as a key partner to help you accomplish that mission? If you haven't heard already, I need more compute, more, more, more. But at least he's consistent.
S
Sachin Kati35:02
I think the thing the whole industry is realizing is that it's a systems problem at a data center scale. We quickly internalized that AI was a rack-level problem, but now it's clear it's not just a rack-level problem; it's a rack in a data center. The whole data center is a system. We have to think about how we design the future based on what we know about our workloads and what's coming down the line, but co-design it with you from CPUs to GPUs to memory, networking, storage, power distribution, and cooling systems. These are all intertwined and have to be co-designed. One of the fun parts about working with you and your team is how easy and productive it is to give you insight into how we expect the world to change, and how quickly your team responds, turning those insights into better chips, systems, and software. That worked really well with MI400 and left-shifted a lot of engineering work. That's why we are very confident we can productionize those systems quickly. We're really looking forward to accelerating that with MI500 and beyond, and increasingly using AI to do that work for us rather than being limited by humans.
L
Lisa Su36:35
That's fantastic, Sachin. We really recognize that input is so valuable. Thank you again for the extraordinary partnership. The amount of work our joint teams are doing together, we couldn't be more excited about the opportunities in front of us. Thank you, Sachin.
S
Sachin Kati36:52
Thank you so much.
L
Lisa Su36:59
All right. So, now let's turn to the world of CPUs. There's been a lot of talk about CPUs recently, and I can say this has been the foundation of our AMD data center strategy for the longest time. We launched EPYC in 2017 and have been on a very clear mission: every generation, more performance, more efficiency, more capability for a wider range of workloads. Naples was our first generation, establishing the foundation. Rome and Milan changed the economics in the data center, adding core count and throughput. Then Turin and Genoa expanded performance, and Turin set a new bar for density, throughput, and total cost of ownership. That's the EPYC formula: a clear roadmap and consistent execution, with leadership that grows every generation. Fifth-gen EPYC Turin is the best server CPU in the world, with up to 192 cores and 384 threads, delivering leadership performance across cloud, enterprise, and HPC workloads. But compute never stands still. With agentic AI, it's creating a whole new class of infrastructure, and the CPU matters now more than ever. That's exactly what we built Venice for. There's been a lot of talk about agentic AI, but it's a very new field. Server computing is splitting across different workloads. First, GPU servers: the CPU's job is to drive the GPUs, all about speed, highest frequency cores, fastest IO, keeping GPUs fully fed. The biggest middle part is the largest growth: agent servers or agent sandboxes. Here you execute code, call tools, query data outside the model. The priority is density: highest performing cores per watt to run thousands of agents at once. Then we have traditional general-purpose servers for enterprise, running applications, databases, data services. With these different workloads, you need the right CPU for the right workload. EPYC is the only CPU portfolio that leads across all three.
Venice is our newest and most advanced server CPU family, designed for the agentic era. We're extending Turin's leadership across every metric: performance, efficiency, TCO for cloud, enterprise efficiency, HPC workloads. It starts with our all-new Zen 6 core, with higher IPC, higher frequency, delivering up to 1.8 times more performance than Turin. That's one of the largest generational gains in the history of EPYC. It's 203 billion transistors built on TSMC's newest 2nm process, and uses a next-generation chiplet that supports up to 512 threads per socket. This is a really important point. It gives us the highest compute density, but also allows us to double both memory and IO bandwidth. I have a few more chips to show you. Venice is not one chip, it's an entire family of chips. Our chiplet architecture is a huge advantage, letting us take the Zen 6 architecture and build a whole portfolio. Let's start with this guy: Venice HF. This is the highest performing CPU that drives AI host nodes. It has eight compute chiplets, each with 12 cores running at up to 5 GHz, with the IO and memory bandwidth to keep GPUs completely fed. This is the CPU I showed you before, shipping inside Helios. Now, this is its bigger brother: Venice 256 core. This is built for agent sandboxes. It also has eight compute chiplets, but each has 32 cores, scaling to 512 threads. It is the highest compute density in the industry. Customers get more agent capacity per watt, per dollar, per rack, bar none. And one more member of the family: Venice 128 core, optimized for enterprise and general-purpose servers. It has the same leadership performance per core, but in a lower cost, lower power design for enterprise applications, tuned to give customers the best performance per dollar. You can see the power of the family as they come together. Next year, we're adding a few other chips: Verona, the next generation for AI host nodes, with more power-efficient low-power memory and faster interconnect between CPU and GPU. Also for high-performance computing, Venice X, using our 3D V-Cache stacking technology, stacking memory chiplets right below compute chiplets to accelerate performance. One architecture, a CPU is not a CPU; there are many different types, and for us, one architecture spans all.
Dozens of different chips and that breadth is something that no other CPU vendor can match. And that's what makes sixth Gen Epic the best CPU in the data center.
Now, you're going to bear with me because I'm going to take you through some performance because the numbers are just incredible. So looking at AI workloads, again, you have different use cases and different workloads. Starting first with GPU servers. Many of the largest frontier models require you to move data between the CPU and the GPU. And in that scenario, Venice moves the data faster than the best competitive x86 processor, delivering up to 1.8 times more tokens per second. And when you look at the Agentic CPU servers and sandboxes, because of our density, Venice delivers more than twice the agents per watt for agent sandboxes. And for general purpose applications, in this case, we're standardizing on a 100 kilowatt rack, Venice delivers more than twice the performance per watt. So pretty incredible results.
And the gap is even wider if you look at ARM processors. So if you look at the leading ARM processors, when we're comparing at the chip level, Venice supports up to 2.8 times more agents per watt. And when you look at the rack level, Venice is delivering up to 3.3 times more performance per watt. And what that means is at the data center scale, you can just get a lot more agents in the same power envelope.
Now, there's been a lot of talk about what matters most for agentic AI with some saying that per core performance under load is the only thing that counts. But that's really only part of the story because agentic AI actually runs as a distributed platform and that is across databases and vector stores and orchestration and all of that has to happen at once and that requires not just raw speed but you absolutely need density efficiency and per core performance all together. And what I'm happy to say is that Venice leads across every one of them. So when you compare Venice against the highest performing ARM CPU from our competition, Epic delivers 20% higher per core performance.
I like that number. I like that number. And when we pair that with leadership core density, we're delivering 2.2 times more performance at the socket level. And so the main point is whether a customer needs per core or per socket performance, Venice is the leader. And there's one more point that a benchmark doesn't really capture because when you think about what agentic AI is, it's really like you're adding thousands of employees to your enterprise, to your traditional workflow, and that software enterprises have been running on for years.
The truth is that software runs on x86. So Venice runs all of it. And this is where x86 has an advantage. You must have the best performance. You must have the best density. But having that software compatibility is a big big plus. So customers scale up their agents on the stack that they already have.
And now let me just show you what that means at the rack level. So at the rack level, what we find is, density is super important. Everyone's trying to optimize their data center. Everyone's trying to optimize their power envelope. And with Venice, every major server OEM is offering a broad spectrum of racks for agentic AI. So what that means is we have the full spectrum from 25,000 cores to 50,000 cores and that's more agents per rack and that gives customers the choice. The key is every data center is different and with that flexibility allows you to pick the right cooling, the right space requirements for what you're trying to do in your data center. So there you have it, Venice the best server CPU in the industry.
And I'm very happy to say also today that Venice is in full production. Customer demand is incredible. It's the strongest we've ever seen for a new Epic generation.
We're seeing every major server OEM, every major cloud provider on track to begin rolling out in the fourth quarter as we start with making the broadest Epic launch we've ever had. So with that, I want to turn to my next guest who runs some of the largest and most advanced compute infrastructure in the world. And it's really been our privilege to partner with them as they've deployed multiple generations of both Epic and Instinct. To share more about our work together, please welcome to the stage Meta's head of infrastructure, Santo Dennard.
How are you? Wonderful. It is so great to have you here. Thank you for being here.
S
Santosh Dennard49:48
Good.
L
Lisa Su49:48
Santosh is a true friend. I have to say, it's been a tremendous opportunity over the last few years. My story about Santosh is the first time I sat in his office, he said to me, Lisa, we're going to deploy a lot of CPUs. Please make sure they work. And I said, I will. I will.
S
Santosh Dennard50:10
And you lifted it up. I have to say they won't care.
L
Lisa Su50:13
I try.
S
Santosh Dennard50:28
Sure. First of all, thank you. It's awesome. This is like a hardware geek's paradise. You come in, you're talking about transistors, you're talking about hardware, you're talking about cooling. This is sort of the jam. I like it. I think this is an audience that's receptive to it. I usually talk to audiences that have either not been to a data center or not really seen a chip ever. So this is refreshing. The thing I'll say is we are seeing demand go through the roof. When you look at inference, training, recommendation systems, that's how a newsfeed works. Content creation, the demand for that is just exponential right now. And when I think about the overall, Mark has this vision of delivering personal super intelligence to billions of people. So if you go to Facebook, Instagram, WhatsApp, or whatever surface, the idea is that we'll meet you there and deliver super personalized intelligence wherever you are. That's why we established this Meta Compute initiative, and because it's a top level thing, Mark himself oversees it. Now what that means is that we are now moving away from an area where we used to take servers and optimize them. This is what the conversations we were having many years ago: 'Hey, I'm deploying a CPU, I really need to maximize performance out of it.' It's different now. You talked about how we should be thinking about the whole system end to end. We are at a point where we have to look at the whole data center and think about it as one integrated system. Servers, hardware, networking, cooling, power, all of that is not optional. All of that has to work together. And this is where I think there's a huge opportunity for collaboration with partners like AMD, because you have to now co-design, co-create systems, not just deploy something that comes off the shelf. The other thing is flexibility. The thing I really liked about what you said is that this is such a big opportunity in the industry right now. All of us need to lean in and work on this together. It needs to be open. It needs to be heterogeneous. There's no one company, one partner that works for any one of us. We need to work on this together. And that's why I think you're one of our most important strategic partners.
L
Lisa Su52:47
Thank you Santo. And look, we completely agree with your philosophy. Now you are actually one of our broadest partners because if you think about our work on CPUs, GPUs, we developed the rack scale OCP systems together, you've deployed lots of CPUs, millions of CPUs, and you're a lead partner on Venice. Can you just talk a little bit about your CPU and infrastructure and some of our work together?
S
Santosh Dennard53:14
Sure. It's been many years. The story that Lisa was saying was many years ago. I think we went from Milan to Bergamo to Turin now to Venice. So it's at least the fourth generation. Long time partnership. Obviously, I actually think the world is changing in the sense that every year is a new world these days.
L
Lisa Su53:35
It's like every month.
S
Santosh Dennard53:37
Months exactly. So what ends up happening now is that it's not just a GPU game anymore. It's CPUs and GPUs. If anything, I think CPUs are becoming at least as important if not more. You're looking at a world where the workload is changing because while you have your workhorses in GPUs, at the end of the day, for agentic workloads you have to run your tools, you have to run your systems, and all the code that's being generated still needs to run somewhere. So I'm super excited to work on Venice. I think there's a lot to come there. And you have to think again, this is a theme that I assume will go throughout your presentation, that you have to start thinking about CPUs and GPUs as conjoined things. You hand off workloads depending on the workload, you employ the right hardware.
L
Lisa Su54:24
Yes. No, absolutely right. And we've gotten a lot of feedback from your technical team. I think that's what I really appreciate about the partnership with Meta. Now clearly you're deploying lots of GPUs too, a lot of accelerators, and you're also one of our deepest partners on the accelerator side starting with MI300 and then we had a very large strategic partnership announced around up to 6 gigawatts starting with MI450 and what we're doing is quite unique with Meta. So can you talk a little bit about the evolution and the MI450 plans?
S
Santosh Dennard55:01
Again, multi-year collaboration. I think we started with MI300, we did a bunch. It had a really nice memory capacity. I remember talking to you about that. And then there's a bunch of pipelining we did. It was important to deploy that at production scale, so get hands-on experience on both sides. 350 is the first time we deployed this on ranking and recommendation systems, so that's now getting into some good numbers. And 450 is, I think I'm super excited about it. We've got some of the racks and we are going to deploy it across the board because it gives us the opportunity to really collaborate. I think the big difference between the 300, 350, and 450 is that we have had an engineer sitting in the same room talking to each other, co-designing, co-creating, like power cooling. As I was saying, it's not just a single thing. And we are not shy in feedback.
L
Lisa Su55:55
They're not shy for sure.
S
Santosh Dennard55:57
But you're very deceptive. So this is true partnership. This is the point about co-designing and co-creating that I think, and as I said, that's a huge deployment coming. The more we deploy, the earlier we co-design, the better we are.
L
Lisa Su56:13
Yeah. Absolutely. Really appreciate the effort together on MI450. Now I'm going to ask you also about your crystal ball. So you look out into the future and you see what the next few years really means as you push forward in the AI industry. What are the biggest challenges and really opportunities for us? This is the industry ecosystem here. So what should we be focused on as an industry?
S
Santosh Dennard56:38
Listen, all of you know about the bottlenecks, you know about the power, the data centers, the silicon. All of these are choke points that the industry is sort of waiting on. But those are things I'm pretty confident we'll work our way through. At the end of the day, I think that demand is immense, people are responding. But the thing that I really want to make sure all of us realize is that we should not think about systems in isolation. As I was saying, CPUs and GPUs are one conjoint system. You should start thinking about it that way. We really need to co-design things early. Data centers take years to build. One of my favorite stories is Mark comes to me and says, 'Hey, I want a gigawatt worth of data centers.' Well, you should have talked to me two years ago. It takes time. It's the same with silicon. It just takes time. So we need to be starting and sitting down in a room co-designing today for what we need to deploy in 27 and 28. That I think is when we truly unlock the powers of the system. The systems are amazing, the specs are amazing, but think about how much more you can get out of it if you sit in a room and just co-design it today, and then in 28 we'll be having much better graphs out there.
L
Lisa Su57:52
Fantastic. Santo, I completely agree with you. I want to say again, thank you. We are so happy with the amazing partnership that we have across the board, and most importantly, with your engineering teams, and we truly are excited about what we're going to do in the future. Thank you so much.
S
Santosh Dennard58:07
Thank you. Thank you.
L
Lisa Su58:14
So, you heard Santos talk about just the diversity of workloads and the massive scale that you require with AI. Now I want to turn to another part of the inference market where the requirements are actually very different. So lots of workloads and as inference is moving into more products and services, we're actually seeing it segment into different workloads depending on what you're trying to do. Some of these are less interactive, so they're really tuned for maximum throughput or the lowest cost. Other applications require more balanced throughput, so you have to balance throughput and responsiveness. And then there's this new class of applications that actually want very fast results. You have extremely smart engineers and they don't like to wait. And that is where ultra low latency or every millisecond matters. Each one of these types of workloads requires a different type of compute. And one of the best ways to reach ultra low latency today is disaggregated inference. So if you think about the different pieces, both prefill and decode are two different jobs. Prefill tends to need more compute, decode tends to need more memory bandwidth. And with Instinct we have a very balanced machine so that we deliver both great compute and great memory bandwidth. But you can actually take this a step further if you know what workloads you're trying to run. You can actually let customers tune each of these pieces independently. And to address this market opportunity, we've been working with Cerebras. So to talk more about what we're building together, please welcome to the stage Cerebras co-founder and CEO Andrew Feldman.
Andrew, it's great to have you here. It's been an exciting few months for you. I know that for those who don't know exactly what you've been working on, can you talk a little bit about Cerebras and what problem have you been trying to solve?
A
Andrew Feldman1:00:18
Sure. Great to be here and great to be among people who love hardware. It's nice. Look, I'm thrilled to be here with you and to announce our cool new partnership at Cerebras. We build the world's largest and fastest chip. It's a full wafer and we package this into a system and systems into racks and then racks into clusters, and we deliver them both on premise and via the cloud. And you've had some of our customers up here already to date on the frontier labs. We serve customers like OpenAI and in the hyperscalers like AWS, in the small agentic and the coding space leaders like Cognition, and they do this because we're blisteringly fast.
L
Lisa Su1:01:14
There's no question, Andrew, that you have some tremendous innovation with what you've been working on and our teams have been collaborating really closely over the last few years to really have Helios together with the wafer scale engine for ultra low latency inference. Can you just kind of educate the audience a little bit? What are we trying to do and what does that mean for customers?
A
Andrew Feldman1:01:35
Sure. Right. I think what's happened and you described it previously is that AI has moved from being a novelty to being useful and then in some domains a necessity. And when something's a necessity, people want to use it and they want to use it quickly. And to serve this market, this segment of ultra low latency, we sought a partnership that could extend our capabilities and our footprint. And there was no better answer than AMD. We were already an AMD customer as we use AMD CPUs to surround our systems, and we were so excited at the opportunity to build a disaggregated solution that combined AMD CPUs, the cool new Helios rack, and the Cerebras wafer scale engine.
L
Lisa Su1:02:19
Yeah. Look, the technology is pretty cool. So just talk a little bit about how it comes together.
A
Andrew Feldman1:02:34
Sure. So as Lisa described, what you have with Instinct in the Helios rack is you have the leader in performance and memory capacity, and you marry that with our wafer scale engine which is the leader in SRAM and in memory bandwidth, and that combination allows us to deliver a solution that is unmatched in the industry.
L
Lisa Su1:02:59
So very exciting. I know that customers are excited about what we're going to do together as well. So, when are we going to have it in market?
A
Andrew Feldman1:03:07
Later this year and it will be first in the Cerebras cloud and it will be later at a store near you.
L
Lisa Su1:03:20
I think the key thing is we're going to give customers the choice to really put together the solution that they want. And I think we're very happy to be able to do this together with you. I think it's a huge step up in terms of what we can do for this ultra low latency segment.
A
Andrew Feldman1:03:35
I think that's right. I think until recently customers could have high throughput or they could have extraordinary speed. And by bringing together the Helios rack with the Cerebras wafer scale engine, we give you five times the throughput while continuing to deliver this extraordinary speed. It's really something amazing.
L
Lisa Su1:04:02
Fantastic. Well, look, Andrew, huge congratulations to what you and the Cerebras team have done. I mean, I think it's been incredible. I think you have been clear on what you were trying to accomplish and it's really nice to see not only it come together but us come together to offer a solution that will be very compelling to customers. So we can't wait to get this into customers hands.
A
Andrew Feldman1:04:22
Soon.
L
Lisa Su1:04:23
Thank you.
A
Andrew Feldman1:04:24
Good to see you. Thank you so much.
L
Lisa Su1:04:25
Thank you Andrew.
All right. So look, I think we've shown you a lot of hardware this morning. But as we all know, hardware is only part of the story. It's actually software that turns all of this into a platform that developers can really build on. So to share the progress we're making with Rockom, please welcome AMD senior vice president of AI Vomi Bopano to the stage.
V
Vomi Bopano1:04:59
Thank you Lisa. Good morning everyone. It's great to be back to talk about software. Just a few years ago, programming AMD GPUs required deep expertise and significant engineering effort. We've come a long way in a remarkably short period of time. Through our open source community collaboration and sustained investment in Rockom, our software stack, AMD Instinct GPUs are powering some of the most important and consequential workloads on the planet. But the biggest change is still ahead of us. AI is transforming how software is built and we are bringing that transformation to Rockom. Our software teams have been moving fast with relentless focus on developers. We've invested at every level of the stack in build and test infrastructure so we ship faster. Rockom releases now go out every six weeks, not every four months. We've expanded the work we do with our AI ecosystem partners and the result is more features, faster performance, and a significantly better out of the box experience. Our strategy is built on two core principles: partner deeply with the open-source ecosystem and build with the right layers of abstraction to enable developer productivity. Open source gives us velocity and scale, and abstraction makes developers more productive. Our deep commitment to open source has resonated with the community. From frameworks and compiler stacks to inference engines and model hubs, AMD is now becoming part of default enablement for the most important AI communities such as Hugging Face, PyTorch, Jax, VLM, and SG Lang. This is why new models get day zero support and why the world's most important AI workloads run on AMD today. That is a big shift. Inspired by engineers seeking productivity, abstractions have evolved to enable work at the right level of detail, hiding complexity when they want speed and exposing control when they need performance. From low-level programming in assembly and C to block programming abstractions like Triton to Pythonic frameworks, we've invested in key abstractions. Aer, our kernel library, and Atom, our serving engine, get you peak performance without writing the kernels yourself. Fly DSL is a brand new Pythonic domain specific language that gives you low-level control with the performance of hand-tuned C++. and Mori, our communications library, helps deliver leadership performance on the latest models like Mixtral.
But look, something very exciting is happening. Over the past year, I've seen something remarkable inside AMD. Our engineers are using AI models to generate GPU kernels, optimize code, debug issues, and improve performance. In some cases, these AI generated kernels are shockingly good, better than what we expected, sometimes better than the most manually tuned versions. The first time you see this, you're skeptical. You run more tests, you try to break it, you look for what went wrong, and then you realize this is real. What has happened in general software development is coming to significantly transform GPU programming. It will reduce the time to bring up workloads, make optimization automated, and remove barriers to broad adoption. We want to put that capability in the hands of every developer. So today I am excited to introduce Rockom.ai. It's an agentic AI platform that brings AI assisted GPU programming to developers. It starts with giving developers access to the power of the AMD ecosystem through the AI coding agents they already use. Whether you're using Cursor, Claude, Codex, Rockom.AI helps those agents understand AMD platforms, understand Rockom, and helps you build and optimize your workloads. We are making those popular coding agents into Rockom super users. You describe the workload you want to run, the performance target, and let the agents help you get there. For years, we have been building the open software foundation. Now we are adding an AI assisted layer that helps developers use that foundation faster. Let me show you what is inside Rockom.AI. It's built on the foundational capabilities of Rockom: core software stack, runtime, libraries, compilers, tools, framework integration. On top of that, we built a layer of AI assisted optimization called Hyperloom. Using Hyperloom, the system can analyze a workload, tune configurations, select and tune kernels, adjust parallelism strategies, and iterate towards performance goals. To give an example, I was talking to the lead developer last week. He told me they just pushed a suite of 14,000 models through Hyperloom, optimizing them and creating valuable insights. This would have been impossible even with a large team of engineers before. That's the power of Hyperloom. And finally, we provide an AI native interface for developers: AI skills that allow coding agents to understand Rockom natively. I am incredibly excited by what we have built. It will profoundly transform how developers interact with our platforms. Let me make this concrete with an example. Imagine you want to optimize Mixtral M3 with VLM on MI355s. Today you need the right VLM optimizations, model recipes, environment variables, and you may need to write new kernels. With Rockom.ai, you say 'optimize Mixtral M3 with Hyperloom.' The agent pulls a known good recipe, runs and profiles the workload, creates and tests many configurations, and iterates towards the target. You can watch it write a GPU kernel. It sees an opportunity for a more optimized group gem. The result is a 38% improvement in tokens per second. This is the experience we want. Simple to start.
That's pretty cool, huh? Simple to start and powerful to optimize. We've continued our relentless focus on performance. With every release, we've delivered significant gains on leading models like DeepSeek. Rockom.ai delivers 3.3 times speed up over Rockom 7. Innovations at every level: block scaled fused kernels, KV cache quantization, expert parallelism. But the bigger story is how AI agents profiled the workload, proposed new kernels, tested configurations, and validated results. That same focus is on training. On the same hardware, Rockom.ai improves training performance by an average of 2.4 times through optimized kernels like fused flash attention, efficient checkpointing, and advanced parallelism strategies. Our engineers have done amazing work to unlock the capabilities of MI455 through Rockom.AI. So I'm delighted to share some incredible results on hardware running real workloads. We are able to see the system deliver 20 terabytes of memory bandwidth and 20 petaflops of delivered compute. These are the highest demonstrated compute plus memory capabilities of any accelerator platform in the industry today. It's measured, not theoretical. Super exciting to share this on day one. Also thrilled to share that leading AI ecosystem partners who have had their hands on MI455 have had a delightful experience. From end to end PyTorch testing to Hugging Face model validation to VLM and SG Lang running well on MI455. We are incredibly encouraged by what the community is about to unlock with MI455. This matters because day zero readiness is one of our most important objectives, and that's what Rockom.ai enables.
So now let's look at Rockom.ai in action on Helios. This time we are using a Codex front end. Let's deploy DeepSeek V4 Pro, a 1.6 trillion parameter frontier model on Helios. Watch what happens. Rockom.AI pulls the right recipes, configures the runtime, implements the model graph onto hardware, and serves up the model. The model is up and running. Let's see it write a simple poem about advancing AI. Isn't that a pretty cool poem? Now there are few organizations shaping the frontier of AI as much as OpenAI. You heard Lisa and Sachin talk about our partnership. It's been truly wonderful collaborating closely with the OpenAI technical team to make rapid progress with Rockom. To talk about how they are advancing the frontier of software development and our collaboration, I'm delighted to welcome to stage Philip Tle from OpenAI.
Thank you for joining us here, Philip. So, just so you all know, Philip is the creator of Triton, one of the most important software innovations in modern AI infrastructure. It's a huge privilege having you here.
P
Philip Tle1:17:03
Thank you.
V
Vomi Bopano1:17:07
OpenAI has always pushed the limits of infrastructure. So, can you tell the team a little bit how your team develops GPU software and how technologies like Triton have evolved?
P
Philip Tle1:17:19
Yeah, for sure. Well, first of all, thanks so much for having me. I was here 2 years ago. It's my great pleasure to be back now. Look, our way of looking at things is there's no one way to fit all applications. We try to meet users where they are. So we've developed multiple solutions. For workflows where speed of iteration is more important, we've developed Triton as you mentioned. This is typically what researchers use. But more and more we've been in cases where we've had to hyperoptimize performance and Triton didn't give us the level of control we wanted. So we've developed Gluon for that, which is a lower level language. This way we've been able to optimize our kernels for AMD among other things. Our agents have become extremely good at both of them, and we've worked with AMD on both as well.
V
Vomi Bopano1:18:15
Yeah, and more and more what we're seeing is agents slowly taking over. They're getting very good at writing this kind of code. Amazing. So look, I think the work you have done with Triton and now Gluon has been really moving the ball forward for the entire industry. Now we began our work together with 300 and extended at 350, but now we're collaborating at a different scale with 450 and Helios. So would love to hear your thoughts on how that collaboration has been going and how has that experience been.
P
Philip Tle1:18:48
Yeah, super well. It feels like ages. I think we started years ago on hardware co-design when MI450 was in the very early prototyping stages. I think we were able to work really well together and it's so nice to see the outcome of that and the chip. More recently, we've been working very deeply on software enablement for MI450 both for our models and for the broader ecosystem. A lot of this development has been done open source. And yeah, we're collaborating on a full stack now, down to the very low-level LLVM code generation, so that all the nice hardware advances that came with MI450 can be properly leveraged in end-to-end applications.
V
Vomi Bopano1:19:36
Yeah, and you got your Helios racks. So tell us a little bit about that.
P
Philip Tle1:19:42
Yeah, it worked like we got the chip and very quickly a single engineer in a few days was able to confirm that everything worked. Obviously we have to optimize performance more, but we're very confident we can get in a very good place.
V
Vomi Bopano1:20:04
Awesome. That's awesome to hear. So look, one of the more amazing things we've done recently is teams have collaborated even more closely in terms of using AI to program AMD GPUs. So would love to hear your thoughts and anything you can share with the audience.
P
Philip Tle1:20:21
Yeah, no. I mean everyone in this room knows how much agents are taking over our jobs as kernel engineers. It's been amazing to see how much better they've gotten over just the past 6 months. They're extremely good at generating high quality GPU kernels in a way that were not before. It's been great collaborating with AMD on making sure that agents not only are capable enough, but also have all the right context they need to make good decisions and good optimizations.
V
Vomi Bopano1:21:02
That's great to hear. Yeah, I think the quality of code produced by AI assisted kernel generation has been truly remarkable through our collaboration. Very excited. Now as you look ahead, what crystal ball do you have for us for the future of AI software ecosystem and collaborations like ours?
P
Philip Tle1:21:22
Yeah, so there's two things that come to mind. One is just how good agents are getting. The other is how much the open approach to software of AMD is enabling not only for OpenAI but for the industry as a whole. I talked a little bit earlier about how we've been collaborating on LLVM code generation. LLVM is an open source project. The entire compiler stack from AMD is open source, and that has allowed us to make our agents really good at very low-level code generation down to scheduling instructions in assembly. This has led to very significant performance gains that I don't think we would have been able to achieve in a fully closed source stack.
V
Vomi Bopano1:22:22
That's one of the real big value propositions we have to offer, because everything we do is in the open and the fact that AI can actually pick that up is a big one. Thank you, Philip. We truly appreciate the partnership and we're really excited about what we're going to build together.
P
Philip Tle1:22:27
Yeah, my pleasure. Thank you so much for having me.
V
Vomi Bopano1:22:34
What you just heard is extremely important. When we combine better abstractions with AI assisted development and close hardware software co-design, we can create incredibly powerful capabilities. That's exactly aligned with our strategy. Now everything I have shown you so far has been about AI, but Rockom is also a great stack for developing high performance and scientific computing applications from frontier models to climate simulation. We use the same Rockom, same library, same tools, same AI assisted development, one open stack for every workload. When you look at our road map, we've retained that strong focus on HPC and scientific computing. Just like MI455 leads in AI compute with FP4 and memory performance, MI430X leads HPC compute with FP64 and memory performance. Two platforms purpose-built for very different classes of workload. The MI430X is built by leveraging our modular chiplet architecture to deliver a purpose-built variant with native FP64 hardware, giving customers leadership performance across AI and HPC in a single platform. This is hardware double precision, not emulation. For the scientific community, that distinction is everything. This product delivers 288 teraflops of FP64 compute, nearly nine times better than competition, and the same leadership memory capacity and bandwidth of MI455. The MI430X ships in the first half of 2027. This is why leaders in sovereign AI love this product. The first sovereign exascale AI factories in both US and Europe are being built on the MI430X at Oak Ridge National Labs and with Cineca in Europe. So look, as we look ahead, let me leave you with this thought. Every so often our industry goes through inflection points. We've felt it when neural networks first started recognizing cats and dogs, when transformers changed what was possible with AI. I believe we are experiencing another one right now, a real inflection in the ability of AI to program complex hardware systems. Building on the surface area from open-source software and leveraging abstracted interfaces, AI is making it dramatically easier to program AI systems. We cannot wait what you will build with it. Thank you.
Now AI is getting pervasively infused into all forms of computing. Enterprises are experiencing the same shift with their own unique requirements. So to tell us more about that transformation, please welcome to stage senior vice president and general manager of compute and enterprise AI Dan McNamara.
D
Dan McNamara1:25:55
Thank you, Vomi. And good morning everyone. It is a great pleasure for me to bring to you the third pillar of our strategy, which is powering AI everywhere. And the vision behind this has remained the same for several years: to deliver the right compute engine to the right workload and then continuously drive optimization points across our entire portfolio from CPUs, GPUs, networking, and FPGAs. So, building on what we've already shared, I'll cover how our strategy comes together across our enterprise portfolio and how the foundation we built with the CPU franchise has truly positioned us for the next era of AI across multiple deployment models. Lisa talked about the frontier model and hyperscaler landscape and the rapid and massive advances we've seen over the last 6 to 12 months. Those advances are driving the adoption of AI well beyond the cloud. As AI continues to grow, it will demand compute beyond the massive purpose-built data centers, extending into enterprise, personal, and physical AI. And each of these brings a new or unique set of requirements.
D
Dan1:27:17
Enterprises are usually constrained by power and cooling. Personal AI usually must operate within a fairly limited compute and memory footprint, and physical AI must perform reliably in real-world demanding environments. AMD is the only company delivering a complete portfolio that spans all of these model development models. So I just mentioned that our journey began in the data center with our CPU franchise, and I want to touch on it a bit. With each generation of EPYC, our strategy has been to listen to our customers, address their evolving needs, and deliver the highest performance, lowest total cost of ownership, and the fastest time to value. As Lisa mentioned earlier, we've introduced many industry firsts across our generations that our enterprise customers have actually asked for. That customer focus has brought us to where we are today with EPYC as the leader in enterprise computing. Over the last three years, there has been very strong adoption across the enterprise, with the world's largest companies moving more and more workloads to AMD. It's reflected in a number of places: first, public cloud where our VM consumption has grown at 76% CAGR over the last three years, and on deployments with our OEM and ODM partners, it is growing quite aggressively across key industries. But just as important, this position allows us to understand truly the diversity of the workloads that our enterprise customers run every day. Each workload places new demands on the CPU. Some require maximum thread count, some depend on per-core performance while others depend on memory bandwidth, cache efficiency, or IO performance. The really important point is that there is no single skew that solves every enterprise application. That is why our EPYC portfolio with Venice spans 8 to 256 cores with a range of power, frequencies, and IO capabilities, giving customers the flexibility to choose the right solution for their workload, and most importantly at the most optimized cost. So Lisa mentioned this earlier, and now we all know we covered Venice, and she showed how we have strong leadership across AI host nodes and agentic workflows. But I wanted to show you how Venice actually leads across the general-purpose workloads of the enterprise, these workloads that our customers rely on every day. As you can see, we have a minimum of 2.5 times the performance across our competition on these very important enterprise workloads. So now as we all know, agentic AI is really the next major transition for the enterprise, and it truly is a force multiplier for enterprises, and adoption is accelerating. In fact, the latest data shows that nearly every enterprise next year will have some form of agentic deployments. As enterprises move from pilots to production deployments, we talk to CIOs all the time, and they are looking at three major things: first, predictable infrastructure cost and token costs; second, making sure the data is secure; and lastly, and sometimes this is overlooked, integrating AI into their existing infrastructure. It's not all greenfield for the enterprise customer. So to solve these problems, we truly believe this is going to be a distributed deployment model that leverages frontier models, cloud services, on-prem infrastructure, and of course AI clients. As AI has moved from chatbots responding to prompts to agents doing real work, writing code, summarizing documents, and running complex flows, the requirements have evolved, and those are exactly the requirements we have been designing for for multiple years. So while AI will be deployed across all these environments, on-prem deployments in our enterprise customers have fairly unique constraints. First and foremost, power and cooling is a big challenge in space, and of course costs remain top of the list for CIOs trying to drive things forward with AI. And lastly, the full solutions need to be easy to deploy. That has been common for many years for the enterprise, but it is something we are very focused on. So to address these constraints, thank you, Laura. I am very happy to announce the launch of the Instinct MI 350P.
So today, every enterprise wants AI in the data center, but the challenge is how to add those capabilities without starting over. That is exactly what the 350P was designed to solve. It is an air-cooled GPU designed to fit within the power and cooling of enterprise servers today without a facilities upgrade, making it easy to bring LLM-scale inference into today's enterprise data center. At the same time, there is no compromise in capability. A single MI350P can support up to 260 billion parameters, allowing customers to run the majority of enterprise AI workloads on a single GPU. The economics are even more compelling: the 350P delivers more than four times the tokens per second per dollar than the competition. It is pretty cool. It actually turns the customer's existing data center into an AI data center. Now, we all know that benchmarks are one part of the equation. The other part is evaluating whether the advantage holds up across the actual workloads you run and how your business runs. So we tested a set of enterprise use cases using production models, and the 350P delivers much higher productivity than the competition, delivering up to two to five times faster tokens per second. So if you think about that, more tokens per second, more work completed, more users supported, all at a lower cost for the business.
So another area I wanted to talk about is our earliest customer, our own AMD IT team. They are a tough group to work with, even for us, but they put us through our paces. This has been our model from day one: we want to put our technology into our own data center first. So like many companies, we are looking at how to deliver AI services to our employee base while managing costs, protecting our data, and controlling the data. We focused on two use cases: first, autonomous threat detection, which is probably one of the faster-growing agentic applications we are seeing today; and second, a personalized AI assistant running on Open Co-pilot. Both of these route requests through the router and gateway, testing the complexity, sensitivity of the information, and service-level requirements. Some requests are sent to the frontier model, while others are routed to open-weights models running on EPYC and an MI350P. With intelligent routing, we reduced our token cost by 43% while delivering up to 3x faster response times for the workloads running locally. So we believe this is how enterprise AI will be deployed.
But we also know from our history with EPYC that enabling enterprise takes a lot more than great silicon. It requires a complete ecosystem of software and solutions. Over the last several years, we have truly grown and expanded our enterprise AI ecosystem. Today, we support numerous AI ISVs, provide zero-day support for all of the industry's leading open models, and deliver robust frameworks that help customers bring AI into production as fast as possible. It is all built on trusted platform, software, and OEM partnerships that enterprises rely on every single day.
Now, everything I just shared comes down to one goal: helping customers deploy AI where it creates the most value for them, whether in the cloud, on-premise, or at the edge. So to bring this to life, I would like to welcome Jeremy Le, Chief Technology Officer at AT&T. Hey crew, good to see you.
J
Jeremy Le1:37:06
Yeah, thanks for having me.
D
Dan1:37:08
Yeah, this is great. It is kind of a homecoming. I grew up in the East Bay, so it is good to be back.
Yeah, fantastic. We really appreciate you joining us today. So look, I think everyone knows AT&T is the largest communications company, one of the largest in the world, and I am sure half of the people out there probably have an AT&T network right now. But you know...
J
Jeremy Le1:37:30
Hopefully more than half, but yeah.
D
Dan1:37:33
But AI is touching way more than the network, and you and your team are doing a fair amount of work driving it. So can you share with us the opportunities you are working on and your experience to date?
J
Jeremy Le1:37:47
Sure, happy to. At AT&T, we are now burning about a trillion-plus tokens per month, and that number has gone up pretty dramatically, moving up double digits. So this is becoming very pervasive inside the enterprise. We have over a hundred gen AI models in production across our enterprise, stretching from things like customer care and fraud to things you would not necessarily think about, like where to place a cell tower or RAN for the most effective use. All of these things are happening across our enterprise at a scale that is not necessarily normal. We get over 300,000 calls a day; the transcription of those calls is an enormous workload, and driving insights out of those calls is another set of insights. It is a really interesting thing. Now we are beginning to rebuild entire workflows inside the company beyond just point use cases. So as we think about HR, finance, and different kinds of things, we are now identifying those entire workloads.
D
Dan1:39:14
Yeah, so it is you know we work closely with your team, and we have had many conversations. It is truly amazing the number of opportunities you are addressing with AI, a super broad range of use. So can you explain the key learnings in moving these to AI at enterprise scale and maybe a bit on what we have done together?
J
Jeremy Le1:39:41
Sure. I mean, I will start at a place that I do not think everybody starts from, which is the human. You need world-class teams to do this, and you need people that think in workflows. There is an enormous amount of what I call racing from stoplight to stoplight – folks going from zero to 100 and then hitting the next stage of the workflow and a red light because they haven't worked their way through. So I would start with people and the quality of the people you have across your enterprise. The second point is we are big believers in data sovereignty. Data sovereignty for us means not being tied to a specific chipset, model, or set of development tools. You need as an enterprise to manage your data; your data is your fuel for implementing AI. What we have done with AMD, which has been a great partner, is leveraging open-source models, leveraging different chipset sets, building our own models, and retraining other models. We are one of the few companies that has actually post-trained a model on AMD and then been able to drive equivalency in performance and accuracy with other more expensive chipsets. As we have done that, we have been able to drive down token cost through moneyballing across all these different models, leveraging that human capital in a way that has made a big difference. So increasingly, even though token consumption is going up, we are able to manage that underlying token cost at enterprise scale, which is not a small thing.
D
Dan1:41:26
Yeah, it is just a bunch of great work. Another key area that is groundbreaking is that you are the first telecom company to actually train a model on AMD, and you are driving fully open-source models in telecom. So tell us about that experience and why you felt it was needed.
J
Jeremy Le1:41:44
We felt the industry was moving down a path where folks were only going to use closed-source models to solve solutions. While closed-source models are an important ingredient, we did not want to lose sight of the fact that open-source models would be a very significant part of the solution. So we announced what we call 1.0, the open telco AI model, a bit over a year ago, and we have had over 18 million downloads since then. We are excited to announce today the launch of Telco 2.0, a more advanced set of models that we have trained on AMD, and we are making it available via open source across the broader industry. I think it will give folks an opportunity to see that you need to think about not only leveraging models but also training those models and how they work across your enterprise.
D
Dan1:42:38
So that is just great. We certainly appreciate the partnership. I really feel like this is germane to the conversation; it is exactly what we have been talking about today. I just want to thank you again for joining us and for the partnership.
J
Jeremy Le1:42:50
Thank you. You are a big part of it.
D
Dan1:42:52
Thank you.
Okay. So let me close with one last point. We have established ourselves as a leader in enterprise computing by listening to our customers and delivering the right solutions for their workloads. Today, we are extending that leadership to help customers accelerate their AI journey. This is our vision for powering AI everywhere. To expand a little bit beyond the data center into physical and client AI, it is my pleasure to introduce Mr. Jack Yun, who leads our Computing and Graphics group. Thank you.
J
Jack Yun1:43:41
Thank you, Dan. It is great to see so many friends, partners, and developers here in San Francisco. Today, we have seen how AMD is building the compute foundation for AI in the data center. But the full AI transformation requires two more ingredients: intelligence that becomes personal and intelligence that can understand and act in a physical world. Those are the two frontiers I want to explore with you here today. Let us begin with personal AI. Enterprise AI is already reshaping the data center, but agentic AI will extend far beyond it. Agents are the next great productivity multiplier, the jump from software that responds to software that completes work. A personal agent can double what one person can do; a team of agents can multiply that by 10. As agents gain more capabilities, they can grow exponentially. But that promise has a compute consequence. Unlike a single AI query, as an agent reasons, calls tools, interacts with other agents, and runs non-stop, agents become more capable and pervasive, so will token consumption. That means demand will require every layer of the computing ecosystem to scale together. Data centers remain the foundation for frontier AI in the most demanding workloads, but edge and client systems can extend that capacity, putting more intelligence closer to where data is created and where work gets done. The PC is already one of the most powerful compute engines ever placed in the hands of an individual, across billions of devices. Enormous computing capability is already deployed every day, yet its potential remains largely untapped. Putting that capacity to work changes the economics of computing. More workloads can run on systems that are already in place, reducing overall infrastructure costs and allowing data centers to focus on the largest, most demanding tasks. Local computing brings inherent advantages: your data stays with you, your applications remain responsive, and your system keeps working even when your network does not. Thanks to a new generation of smaller, more capable AI models, we are beginning to unlock that potential. The unlock is happening faster than anyone expected. Last August, GPQA needed 120 billion parameters to score 80 on GPQA. Just seven months later, Qwen 3.5 scored even higher with only 9 billion parameters, 13 times fewer parameters. This is not an incremental improvement; it is a dramatic shift in the compute required to deliver advanced intelligence. This is a trend, not an outlier. Smaller open models are closing the frontier gap at extraordinary speed. Qwen 3.6, a 27 billion parameter model, now outperforms leading frontier models from the previous generation on graduate-level reasoning at a fraction of the size and cost. That changes what is possible. Workloads once reserved for the cloud can now run on PC compute engines designed for personal AI. Let us map these model performance breakthroughs to AMD platforms. Ryzen AI 400 supports models up to 24 billion parameters, more than enough for Qwen 3.5 with substantial headroom. Ryzen AI Max extends that capacity to models as large as 200 billion parameters running natively on a personal system. This is not reduced AI compressed to fit on a PC; it is a new class of personal computing built to run powerful models where the work actually happens. This is exactly what led us to build Ryzen AI Halo. We wanted to give developers a platform powerful enough to run these models locally yet simple enough to use every day. With 120 gigabytes of unified memory and support for models up to 200 billion parameters, developers can build, test, and iterate locally on their desk. Ryzen AI Halo embodies the future of personal AI for everyone, shortening the distance between an idea and a working model. For developers, that means a dramatically more efficient and optimized path from experimentation to deployment. As Vont showed earlier, ROCm is making hardware programming radically more accessible. Developers can understand, modify, and optimize code for ROCm even without deep ROCm expertise. Better yet, the same open software foundation extends across all of AMD's hardware. Developers can use the frameworks and tools they already know, start locally on Ryzen AI Halo, and scale across the full AMD compute platform without rebuilding their work. We are not stopping there. Today, we are expanding our partnership with Hugging Face to bring the open AI ecosystem directly on Ryzen AI Halo. Powerful compute only creates value when developers can quickly access the right models, optimize them for the platform, and turn them into real applications. Together, AMD and Hugging Face will deliver Halo-optimized performance for open models and agentic workflows. How cool is that? Along with co-engineered libraries and toolkits designed to help developers spend less time configuring infrastructure and more time building, later this year every Ryzen AI Halo box will include a full year of Hugging Face Pro, putting the power of partnership in developers' hands from day one. Freedom and more power to the user. This is how we accelerate innovation in the AI era. We are already pushing that vision further with Gorgon Halo. Seeing is believing. Look at how small and beautiful this box is. We have increased unified memory from 128 to 192 gigabytes, the largest unified memory pool in its class, and increased model support from 200 billion to 300 billion parameters. Thank you, Laura. This is no longer limited to a single platform form factor. Together with the world's leading OEMs, AMD now powers the broadest portfolio of personal AI systems, from laptops and workstations to many PCs and developer platforms. That means developers, enterprises, and creators can choose the right system for the way they want to build and work. Personal AI is not a concept; it is a category, and it is available right now on AMD. We now have all the pieces. Models are efficient enough, hardware is powerful enough, the software stack is open and ready. Open harnesses like LM Studio can route workloads to the right model and right compute. The technology is here. The next challenge is taking it from one developer's desk to thousands of users across an enterprise. Cisco sees and believes in that future the same way we do. To take us deeper into how we turn that vision into enterprise-scale reality, please join me in welcoming to the stage G2 Patel, President and Chief Product Officer at Cisco.
G2, how you doing?
G
G2 Patel1:52:15
How are you?
J
Jack Yun1:52:16
You look great.
G
G2 Patel1:52:17
Thank you for joining me here today.
J
Jack Yun1:52:19
Congratulations. Thank you. Now G2, you and I talk about how we are at an inflection point. For the past few years, the industry has spent time just building intelligence. The next phase is how we deploy this at scale in enterprises. You spend so much of your time talking to CIOs and technology leaders. What are they telling you? What is changing?
G
G2 Patel1:52:39
Well, firstly, it is a really exciting time to be alive in tech. If you take a step back and see what is happening, there is a fundamental shift in the patterns of inferencing. In the chatbot era, inferencing was human-led, very spiky. You ask a question, you get an answer. What you are starting to see with agents is they are working 7 by 24, very consumptive on bandwidth and infrastructure. So you are starting to see a much more persistent pattern of demand signal for infrastructure. We usually joke internally: humans click, but agents swarm. So as you start to see this, you will have tremendous demand. There are two big concerns: every customer is concerned about security, and every customer is concerned about cost.
J
Jack Yun1:53:32
So what you are starting to see is inferencing not just limited to the data center. Inferencing is going to be distributed everywhere. There is a whole new class of computing emerging with this deskside computing that you just announced, where there will be inferencing on laptops or desktops where humans work, and deskside computers where agents go out and run their jobs on behalf of humans. That is what we are seeing.
G
G2 Patel1:54:00
And we are fully aligned, G2. That is exactly the shift we are seeing. If AI is moving closer to employees, compute has to move closer as well. The PC is no longer just a productivity device; it becomes an intelligence node in the enterprise. But deploying intelligence everywhere creates a new challenge for CIOs that you and I talk about so much. How do you operate thousands of these systems with the security, governance, and control enterprises require? How is AMD and Cisco solving this, and what are we building together?
Yeah, this is an exciting time. There are a few things that need to be done. You cannot just deploy deskside computers and hope everything works out. You need to make sure it is governed and managed. Number one, you have to have appropriate network bandwidth to satiate the needs of the agents you will run locally. So modernizing network infrastructure to keep up with the demand is very important. Number two, every customer is worried about token costs, so we need to contain the costs of an agent's token consumption when it starts to go up and down. Number three, we need to monitor agent behavior for safety and security so that if it does start to do things we do not want, we can provide runtime enforcement guardrails. And fourthly, overall safety and security from an apparatus perspective to get the whole thing humming.
J
Jack Yun1:55:39
Yes. G2, you and I think very alike. The opportunity is not just to bring AI to every device, but to enable enterprises to deploy, manage, and scale with confidence. Why don't we take a look at what we have been building the past few weeks and months?
G
G2 Patel1:55:53
Yeah, it is exciting. Basically, what we have done is a full stack. You folks have done with Halo: an isolated secure agent sandbox, intelligent routing, MCP core integrations. Above that, we have made sure you have the right level of security policy enforcement, observability for an end-to-end resilient infrastructure stack. Observability of how the apparatus and infrastructure are working, whether agent behavior is done correctly, and whether the token economics are within guidelines. Then we have a single unified management plane, a control plane that can look at every dimension and be managed within one environment. That is the full stack, and we are really excited about the partnership.
J
Jack Yun1:56:58
Us too. Our teams love working with your team, G2. Now, how is this all managed? How does the CIO deploy this at scale?
G
G2 Patel1:57:04
So let's take a look at what this looks like. We have Cisco Cloud Control, the overall management plane. If you have Halo devices on every desktop, you want to manage your entire estate, not just your Halo fleet but also what you have running in data centers and the cloud. You can provide full visibility for the whole estate. Then we also have the notion of monitoring token economics: is the agent behaving the way it should, or is it being overly consumptive? At which point you can isolate any individual Halo device to quarantine it.
J
Jack Yun1:58:00
I love this. Then we actually see the return on invested capital. It could pay for itself in three or six months, even quicker.
G
G2 Patel1:58:05
Exactly. So what you see here is you can tell how every single agent is performing and whether they are consuming too much token cost, and how to contain the cost.
J
Jack Yun1:58:18
No, I love this, G2. Now, everyone else is wondering when they can buy this? When is this open for everyone?
G
G2 Patel1:58:25
Well, the good news is you can buy the Halo already today. And what we are doing is providing this entire management apparatus that is currently in early availability for a select set of customers, but we will have it in general availability in the US in early fall. So we are excited to provide not just the compute but all of the apparatus needed to manage this, so you can have inference running in the cloud, in your data center privately, or on your desk side.
J
Jack Yun1:58:59
G2, we are so grateful for the partnership with you and Cisco. We are so excited.
G
G2 Patel1:59:03
Thank you so much. Congratulations. Take care.
J
Jack Yun1:59:10
What you just saw brings the vision to scale AI seamlessly and securely from personal devices all the way to the cloud. But in the age of AI, intelligence is where AI leaves the screen, enters the world, and acts. That frontier is physical AI, the ultimate application of agentic frameworks. In a digital world, an agent plans and completes a task. In a physical world, it must also sense, motion, navigate uncertainty, and protect people. Here, intelligence does not just produce an answer; it produces an action. When answers become actions, latency becomes safety, and reliability becomes trust. That is why the promise of physical AI is so incredibly powerful. It is not about replacing human potential; it is about amplifying it. Giving surgeons greater precision, keeping people out of harm's way, making farms more productive, supply chains more resilient, and factories more efficient. The best technology does not make people smaller; it gives them much greater reach. For more than 20 years, AMD has helped power the core capabilities that robotics depends on, from sensing and functional safety to deterministic control and precision motion. That foundation is already trusted by many of the companies defining the future of robotics. Now that foundation is coming alive in extraordinary ways. These are not machines following a script; they are beginning to perceive the world, reason through uncertainty, and act with purpose. The promise of physical AI is already in motion. The next era of robotics is not programmed; it is autonomous. While traditional robots execute instructions, autonomous robots are given an objective and determine how to achieve it, all at once and in real time. The robot must orchestrate intelligence, motion, and safety continuously, and every part must respond at the speed of the world around it. That demands an entirely new kind of robotics brain. To solve that, we rethought its architecture from the ground up. Today, I am so excited to introduce the AMD Korea AI system on module, powered by Ryzen AI Embedded X100. It brings CPU, GPU, NPU, and unified memory together in one compact open-standard architecture. This allows robots to process perception, AI reasoning, and do it all simultaneously in real time. The architecture is different, but the results are decisive. In third-party testing against Nvidia Jenson Thor, Korea delivers 3.4 times better real-time results, 2.3 times more concurrent agents, and 1.6 times more CPU capacity. That means faster reactions and greater headroom for the robot to take on more complex work, all without compromising control or safety. But to define the next era of robotics takes more than a powerful brain. Developers need a complete path from idea to a machine operating in the real world. Today, we are introducing AMD's Korea AI robot developer platform, the industry's first turnkey, open, fully integrated robotic system. Thank you, Laura. At the center is the Korea AI system on module paired with a robotics carrier card designed for sensing and connectivity. Built on ROCm and ROS 2, this platform gives developers everything they need to move from concept to prototype in days. The same module carries forward into production with no migration or redesign needed at the finish line. You can build faster and move from imagination to deployment with absolute confidence. This is the breadth only AMD can bring to robotics. From the first signal to the final motion, AMD Korea for the brain, Versal for the spine, Zynq for the joints, and Spartan for the sensors. Each layer purpose-built for its role, working together in one intelligent system, connecting perception and high-level reasoning to real-time control. The future will not be defined by one model, one machine, or one company. It will be built in the open across silicon, software, systems, and an ecosystem moving forward together from the cloud to the PC and now into the physical world. AMD continues to build a compute foundation for intelligence everywhere, with laser focus on the future, so our solutions can augment and expand what people are capable of achieving. The next era of AI will not just unfold behind a screen; it will unfold in the world around us. Now, to bring it all together and take us to what comes next, please welcome back Lisa to the stage.
Hi Lisa.
L
Lisa Su2:05:15
All right. Thank you, Jack. It has been a big day. We have covered a lot, from our data center products to enterprise AI to personal and physical AI, and the software that ties it all together. This is the strongest product portfolio in our history, and we believe it is the strongest portfolio in the industry. But it is really just the beginning, because what our customers count on is not just today; they count on us looking ahead and pushing the bleeding edge of technology. So I want to spend the last few minutes giving you a preview of what we are working on. Starting with CPUs. In 2028, we are going to introduce Florence. Florence brings the next-gen Zen 7 cores, leading-edge process technology, new AI compute extensions to ensure we have all the AI capability, and it supports the latest memory technologies. We are not stopping there; we are already deep in development of Ravenna, our eighth-generation EPYC family built on Zen 8, already well under development for 2030. To give you a flavor, Florence will extend our biggest advantage with EPYC: product breadth and having the right foundation and family for each workload. So you have Florence, Ferrara, Fidenza, built on a full family of CPUs optimized for different workloads, giving customers the right CPU for every workload. That is our commitment with EPYC. Moving to Instinct, we are continuing our cadence of delivering a new generation every year. MI500 brings next-generation HBM, a larger scale-up domain because scaling is everything, and introduces new copper and optical interconnects. With MI600, powered by our CDNA next architecture, it is already deep in development for 2028. To give you an idea of performance, every generation of Instinct has delivered a major step up. With MI455, we have talked about it delivering 35 times more inference throughput compared to MI355. But we are really excited about MI500. We see another opportunity to take a major step up and bend the curve again. MI500 will deliver the largest generational leap in the history of Instinct, putting us on track to deliver more than 2,000 times higher inference throughput in just four years. What I can tell you is we are working with a number of customers already on this design point. MI500 is super exciting, and the feedback we are getting is fantastic. Now, put the CPU, GPU, and networking roadmaps together, and you can see a complete cadence of our rack-scale systems. You should expect from AMD a new Helios system every year, delivering more performance, more capacity so we can run larger models and more agents at better total cost of ownership, giving our customers a complete predictable roadmap to plan and scale their infrastructure for years to come. Lots of exciting things in store over the next few years. So with that, let me wrap up this morning. I hope you saw a little bit about how we are looking at compute from every angle. Every aspect of AI compute, whether you are talking about Helios or MI455 or Venice for the world's largest AI systems, or MI430 and MI350 for sovereign and enterprise AI, you saw the passion in ROCm AI. It is so good to see how much has come through that platform in terms of making it much easier for developers to access the AMD platform. And Jack talked about Gorgon Halo and our new Korea AI platforms, really bringing together the leadership of the end-to-end AI story, including PCs and physical AI. Lots of new information, but I want to say a very special thank you to all of our partners who joined us today, because it is through those partnerships that we are able to do the most amazing things together. So if I leave you with a final thought: when I think about where AI is today, the biggest change is that we are no longer talking about what might be possible; we are actually seeing how AI can have real and significant impact across every industry and every part of our lives. I have spent my entire career in tech believing that high-performance computing can make the world an incredibly better place, and I have never believed that more than I do today. At AMD, we are focused on building the technology, the roadmaps, and the partnerships. There has never been a more exciting moment than today for our 30,000-plus engineers. This is really the next phase of AI, where we give AI the opportunity with all the tools and capabilities to bring meaningful real-world impact. I could not be more excited to build all of that together with you, our ecosystem. Thank you so much for joining us today.