Back
Lisa Su
Chair, President & Chief Executive Officer, Advanced Micro Devices

Advancing AI 2026 | Build What's Next with @AMD

🎥 Jul 23, 2026 📺 AMD ⏱ 133m 👁 42525 views
Join AMD and Dr. Lisa Su at #AdvancingAI on July 23rd at 9:30 AM PT to experience the latest in AI infrastructure, architecture, and development. Get up to speed on the latest AI compute innovations from AMD and our enterprise partners. And hear from industry leaders building what’s next in AI. Discover more: https://www.amd.com/en/corporate/even... 00:00 Introduction 00:02:16 Welcome & Opening Keynote: Dr. Lisa Su, Chair and CEO, AMD 00:21:31 Anthropic Partnership: Tom Brown, Co-founder & Chief Compute Officer, Anthropic 00:30:47 OpenAI Partnership: Sachin Katti, Head of Infrastructure, Op...
Watch on YouTube

About Lisa Su

Lisa Su, chair and CEO of AMD, spoke at the company's Advancing AI 2026 conference in San Francisco on July 22-23, 2026, where she announced new products and partnerships. Su stated that AI infrastructure demand is accelerating, not slowing, and described inference as the industry's largest growth driver. She said AMD expects the AI accelerator market to reach approximately $1.4 trillion by 2030, approaching the size of the entire semiconductor market today. Su also announced a $5 billion investment in Anthropic and discussed the Helios rack-scale architecture, which she said offers 10 to 15% more performance and up to 30% better tokens-per-dollar for customers compared to competitors. She reiterated her view that no single chip company will dominate the AI market, arguing that the world is heterogeneous and requires an ecosystem of different compute engines. In a commencement address at MIT on May 28, 2026, Su told graduates that "we may discover more in the next 10 years than we have in the last 30," but emphasized that "technology itself does not decide what the future looks like. The best people do." She advised graduates to "run towards the hardest problems" and said that "luck is not just being in the right place at the right time. It is taking the risk to work on something really hard." Su reflected on her decision to become CEO of AMD 12 years ago, describing it as her dream job, and noted that the company made a long-term bet that high-performance computing would be the most important technology of the future.

Source: AI-verified profile updated from Lisa Su's recent appearances. Browse all interviews →

Transcript (138 segments)
N
Narrator0:43
The next wave of AI is here. Evolving beyond thinking to acting. And it's bigger than anyone imagined. Except AMD. We recognize this agentic era would require the highest performing GPUs and CPUs working together. So we're the only ones who build both. It's how agentic AI moves from answers to action. Pushing inference compute to tremendous scale. Like an agent that frees thousands of employees from smaller tasks to focus on bigger things. A car that safely handles a new city like a local on its first day there. And an agent that pinpoints a disease in millions of cells, then designs a cure. We've built more AI computing power into a single data center than existed on Earth a decade ago. And now, we're giving you the tools. So together, we advance AI.
L
Lisa Su2:07
Good morning! That's a pretty good good morning. Good morning! Good morning! And welcome to Advancing AI 2026. It's so great to be back here in San Francisco and to see so many friends and partners and customers. And especially all the developers that are here with us today. And I want to say a big welcome to everyone who's joining us online from around the world. This is my absolute favorite event of the year. It's where we bring the entire AI ecosystem together to show what we've been building. And where we're going next. And this year, this is our biggest show ever. Because we have so much to tell you. So let's go ahead and get started.
At AMD, our mission is to push the boundaries of high performance in AI computing to help solve the world's most important challenges. And I often say, AI is the most important technology of the last 50 years. And frankly, the progress that we're seeing in the industry is just incredible. Like, every month, every few months, we see something new that was far beyond our imagination. And you can already see the impact across every industry. In healthcare, AI is helping researchers identify new drug candidates faster. In science, we're solving problems that we didn't think were possible a few years ago. And across every industry, AI is changing the way we get work done. And the most important thing is we're still in the very, very early innings of what's possible.
Now, just take a look at some of these charts. You know, we look at these every year. And you kind of see the rate and pace of AI growth. A couple of years ago, we were just starting. People were experimenting with chatbots. And, you know, it was really cool. But if you look at today, we're really seeing more than 35 quadrillion tokens consumed every single month. That's an increase of nearly 160 times in just two years. And that curve is actually just getting steeper. You're going to hear a lot about that today.
So, just looking at, you know, where is all that demand? I mean, obviously, training remains incredibly important and foundational. And over the last four or five years, you know, we see the compute that's used for training advanced models continuing to increase by, let's call it, roughly 5x every year. And those models are getting better. Like, we're all experiencing that. We're seeing better reasoning. We're seeing new capabilities. We're now seeing a growing number of specialized models. And one of the things that I really believe in is there is no one perfect model. I think we're all going to use a slew of different models, depending on what task you're trying to solve and which industry you're in. But the bigger shift is actually in inference. And we expected this. We expected that inference would grow faster than training. And we can say that in 2026, for the first time, the world is using more AI compute to run models than to train them. And we've seen that to a point where we think this year, roughly 60% of the global AI compute capacity will be used for inference.
And the reason for that is actually pretty simple. As billions of people every day use AI, the workload shifts from building models to putting them to work. And actually, we're going to talk a lot today about agentic AI, and that's accelerating the shift even faster. So if you think about all of this, we see agents as the next big step for AI. And the most interesting thing is this has really just started accelerating over the last, I would say, five or six months this year.
So when we started with LLMs, they were great to answer questions. But we found that to really get the power of AI, you want agents to be able to really answer full questions. So you give it a goal, it keeps working until it's solved the problem. And what that means for us in the compute industry, it means that compute demand is growing at an incredible pace. We're actually seeing a step change in compute demand. Because when you ask the agent to do something, it actually has dozens of steps. And it has to reason, and it has to call tools, and it has to access data. And it has to keep doing it over and over until it's solved the problem. And so you need lots of GPUs to do all that reasoning. But importantly, you also need a lot of CPUs to orchestrate every step around it.
So when you look at just what that means from a market standpoint, it's really hard to put this market data up. Because I feel like every few months, we're changing the perspective based on what our customers are saying. But what we're seeing is that the shift to agentic AI is certainly growing the AI accelerator market very significantly. It was actually last year at this event that we called the AI accelerator market at about 500 billion by 2028. And at that time, it actually felt like a very big number. Didn't you think it was a big number? But honestly, you look at it today, and it looks conservative. Because the AI demand is just continuing to accelerate. And you can see it, right? Better models, more usage, more usage requires more compute. And more compute actually builds better models. And so we're now expecting that by 2030, the AI accelerator market is going to reach about 1.4 trillion by 2030. So what do you think about that? Is that a big number?
I mean, what that means is by the end of the decade, the AI accelerator market is going to approach the size of the entire semiconductor market today. And although there will be many different types of accelerators, you know, I'm a big believer in there's no one size fits all as it comes to chips. We do expect that GPUs are going to make up the vast majority of that market because the algorithms are still very much in their infancy, and we're still continuing to see the workloads change, and that favors programmability in the overall silicon ecosystem.
Now, perhaps the most interesting part of the last six months has been, you know, GPUs are only part of this story. What we're actually seeing is agentic AI is creating an entirely new growth vector for server CPUs. Now, I've updated this number a lot also over the last, you know, six to eight months. But look, we call it like we see it, and we were early. I mean, we saw from our largest hyperscale customers, and you're going to see some of them here today, who said, look, as AI inference is going up, we need more server CPUs. And so we saw growth in the server CPU market, and we said, you know, last year we thought it might grow, let's call it 18 to 20% CAGR, to an overall market size of 60 billion.
But frankly, what we're seeing is that the rate and pace of agentic AI adoption is much, much faster than any of us thought. So we're talking about agents going from millions to billions, like we're able to increase all of our productivity, and that requires a tremendous amount of CPU infrastructure. And every customer conversation is telling us that this build-out is just beginning. So based on what we're seeing today, and we're going to talk more about how agentic AI is evolving with CPUs, we now expect the server CPU market to grow by over 50% to over 200 billion by 2030. So that's starting from today, a $25 billion market going to over 200 billion. So there's a lot of excitement about CPUs as well.
Now, probably one of the things that is differentiating AMD in how we think about the market is this is not just a data center opportunity. As AI becomes a larger part of our daily lives, we want the intelligence to run right where the work happens. And that means cloud is super important, but the devices that we use every day are going to be extremely important. And that means that we're going to do a lot more at the edge as well with autonomous machines that sense and act in real time. And we really believe that you need AI to be infused everywhere. And that's exactly what we're focused on at AMD.
So putting all of that together, we actually expect that the market for our high-performance and adaptive computing products to grow at roughly a 40% CAGR over the next several years, approaching a $2 trillion market by 2030. And frankly, the only way you're going to be able to service a market like that is for us to work together as an ecosystem. There's no one company that can solve it all, but this is the opportunity to bring the best and brightest together. And we have never been in a better position to lead. We have the broadest product portfolio. We have the strongest roadmaps we've ever had. And the thing that I'm most proud of is we have the deepest partnerships with the companies that are building this future.
So a little bit about our strategy. I think we've been very consistent. We've talked about our multi-year strategy, and it's really built around three priorities. The first is just compute leadership. We're building the broadest set of compute engines in the industry so that we have the right compute for the right workload.
Second, it's going to be about open platforms. We want everyone to come together in an open ecosystem. We believe an open ecosystem is essential to the future of AI, and that's how we get the force multiplier of everyone coming together. And that's open on hardware, in terms of hardware standards, as well as open on software. And that's why we're investing so heavily in our ROCm software. And you're going to see that AI has been a tremendous multiplier in this great pace of progress that we're making in ROCm and in software. And it really means that if we put all of this together, we can have our customers deploy on AMD hardware faster than ever.
And the third piece is we're going to get a chance to talk to you about powering AI everywhere. So that's AI in enterprise. That's AI at the PC. That's AI in the physical world. And that's adding new capabilities across all of our products. So today you're going to see all of that in action. We're going to start with compute for the agentic era. I'm going to show you a lot of hardware. And then we're going to talk about some software and then our overall products from a platform standpoint. And we're honored to have some really special guests who are going to join us to really help bring the technology to life. So let's start with the data center.
Today, AMD EPYC runs on the most important workloads in the world. We power the largest cloud providers. We power the digital platforms that billions of people use every day and the most important critical systems that are used by the largest businesses, including more than 60% of the Fortune 100 are using AMD EPYC. And that momentum is just building. Last quarter, we reached a record 46% revenue share in the server market and we're continuing to see every major customer move more workloads to EPYC.
And on the GPU side, our Instinct adoption has also accelerated. We have a broad set of customers across the largest AI labs, cloud providers, leading AI startups, as well as some of the national labs and sovereign AI opportunities are being built on AMD Instinct. And we really appreciate that opportunity.
Now, AI compute has become a lot more complicated. And what we're seeing with Frontier AI is it's really raising the bar for what infrastructure needs to do. It takes more than a single chip or a single server. You actually have to design the entire rack as one system. And that requires leading CPUs, that requires leading GPUs, that requires high-speed networking that connects everything inside the rack and across the data center. And just as importantly, this is about making these systems super easy to use. And so they need to be easy to deploy, easy to service, and very reliable. And that is exactly what we built with Helios.
So today, I'm super excited to launch Helios, the industry's highest performance AI rack. Now, Helios is built to train and run the most demanding Frontier models in the world at massive scale. And I have a lot of show and tell for you. So for those of you who know me, you know that I love holding up chips. I call myself sometimes the Vanna White of chips. But it turns out with Helios, there are lots and lots of components. And so you can see, these are the chips inside the Helios rack. These are the Instinct... Thank you.
So this is Instinct, EPYC, and then the Pensando chips for networking that are inside each rack. And every one of these chips is part of the system. So MI455 is the engine. It's the highest performance GPU in the industry. Venice, which drives it, is the world's fastest CPU cores. And Pensando DPUs and NICs actually connect them with leadership programmability and scale-out bandwidth. So I have a few more props to show you. Let's kind of take a look at how these things fit into the overall system.
So first, this is now the MI455 accelerator, which is mounted on what we call an Enhanced Accelerator Module, or EAM. It's a production module that goes into Helios, and it is much more than the GPU package. The EAM actually integrates the GPU, memory, power delivery, high-speed interfaces, system management, cold plates for liquid cooling, all into one compact, serviceable module. Pretty cool, right?
Now, if I just give you some of the specs, we're talking about 320 billion transistors. This is built with TSMC's leading 2-nanometer and 3-nanometer process technology. And most importantly, it brings together nearly a decade of AMD chiplet innovation. So it allows us to combine 12 compute and I/O chiplets with 432 gigabytes of memory, all connected by leading industry 3D chip stack packaging. And I can say for sure, it's the highest performance AI accelerator in the industry.
Okay, so this is the CPU board that's inside the Helios compute tray. And I'm going to talk a lot about CPUs later. But just to give you a high-level view, it's high-speed, 96-core EPYC processor, DDR5 memory, all the I/O that's needed to feed the GPUs, all in a single motherboard. And we also have our Salina DPU. This is a critical part of the networking infrastructure, and it really allows us to deliver both front-end as well as the scale-out and scale-across bandwidth overall. And this is our Vulcano AI NIC. And you can see this actually allows us to really have up to six Vulcano NICs on a board. And it uses the open ultra Ethernet standard. So each Helios compute tray includes two of these Vulcano boards, and that allows us to have the scale-out and the scale-up networking overall.
So you put all that inside the rack, and we have an additional six dedicated networking trays that handle the scale-up networking. And those are all connecting 75 GPUs together with UA-Link over Ethernet with silicon from our partners. And that's what we mean by an open ecosystem. So when you bring this all together, I want to say Helios is simply the best AI rack in the world.
Now, let's take a look at some numbers. Clearly, it's a very competitive world out there. So when you compare Helios to the competition, we're delivering 15% more compute, 50% more HBM4 memory capacity and memory bandwidth, and 50% more scale-out bandwidth. And what that means is that every Helios can deliver more performance for the largest models, more capacity for longer context, and the bandwidth to scale across thousands of racks.
And now we can kind of come over here and take a look at another very large piece of hardware. This is now the production hardware that customers are deploying, and you can see more of that as you go through the exhibit center. And each one of these weighs more than 160 pounds and stands less than two inches tall. That's just an extraordinary amount of technology when you look at what's in one of these trays. And when you bring these things together, the numbers are just incredible. More than 18,000 CDNA5 GPU compute units, over 4,600 Zen6 CPU cores, and 31 terabytes of HBM4 memory, all in a single rack. That's what it takes to run agentic AI at scale.
So today I'm excited to announce that Helios is in full production. We have shipments on track to start at the end of the third quarter and ramping into the fourth quarter in the second half of the year. And I can tell you customer demand for Helios is extremely strong. And we're extremely proud of the work that we have done across the leading AI labs to adopt Helios. And so let's start with some of our guests. Today I'm excited to be joined by our newest Instinct partner and one of the leaders in Frontier AI. To share more about what we're doing together, please welcome Anthropic co-founder and chief compute officer Tom Brown.
Hello, Tom. It is so great to have you here. And it is really such an honor to be working with Anthropic. We had a big week this week.
T
Tom Brown21:56
We did, yeah.
L
Lisa Su21:57
So look, Anthropic has just had incredible success over the last few years. You know, Claude has made such a big impact on the overall industry. Compute is such an important factor in the progress. Can you just talk a little bit about your compute strategy and just how things have been evolving?
T
Tom Brown22:15
Yeah. Yeah. And first, thank you so much. A hard act to follow with Helios.
L
Lisa Su22:19
Amazing.
T
Tom Brown22:20
So with Anthropic, as you mentioned, the scale that we're growing just as an industry is enormous. And so we have been working to make sure that we have the best, like different chips, and we can use the best chips for the best workloads.
L
Lisa Su22:37
Yeah. And look, I think you've been really leading in that area. And, you know, I'm happy to say that I've always wanted to be as part of your infrastructure. I remember the early times we talked as you were starting Anthropic. But we had a big announcement this week. We're honored that you will be deploying up to 2 gigawatts of Helios. Why now? What was the key reason?
T
Tom Brown23:02
Yeah. So I think that the first thing was just Helios is an amazing machine. It's absolutely fantastic. Do we like that? Does that work for us? And then I think the thing that we were thinking about originally was whenever we're bringing up a new hardware platform, it's a big effort. It's like a huge thing. And so as part of the... as we were thinking about this, we started doing our own evaluation of MI355. You guys generously got us a rack to start working. And we expected this to be kind of a big process. Our actual experience was we had one engineer start doing it. They spun up... they connected it to Claude, asked him, hey, bring up this machine, left it going over the weekend. And we ended up with a graph of the actual performance of our leading model on it just going up and up and up over the weekend. And so, yeah. I think that that's a testament to the open platform that you guys have made where anyone like human or AI can now build real models on your platform.
L
Lisa Su24:20
Tom, I kid you not. When my team told me that you had an engineer working on MI355, I said, hey, are we making sure we help them? And they're like, they say they don't need any help. They got it all covered. And I was like, hey, that's a wonderful story. Thank you for that. Now, we're also super excited about the work we're doing together because one of the most valuable parts of this partnership is Claude and using Claude within AMD. We're big believers in the fact that AI is this force multiplier and the more we're able to have the leading foundational models really familiar with the AMD architecture, the better we're going to be able to service the overall market. So can you just talk a little bit about how you're seeing Claude really evolve in these types of highly engineering-specific workloads?
T
Tom Brown25:04
Yeah, so I think this is a great place for collaboration where the bread and butter of Claude is doing software engineering. And now more and more we're seeing it help out with the type of workloads that you guys are doing also, like design, layout, which is not quite the normal software engineering, but is adjacent. And I do think that that is the place where we can work together to make the next generation of chips even better.
L
Lisa Su25:34
Yeah, I think, and also on the software side, I think the work that we've seen with Claude and the kernel development has been just incredible. So, Tom, one of the things for our audience here, we like to talk about the present, but it's actually most of us are working on the future. So this is really the beginning of what we believe is a very strong multi-year partnership. Can you talk a little bit about, number one, what are you most excited about in the industry? And then number two, like, what can we do more together to really make sure that we're satisfying those key opportunities?
T
Tom Brown26:13
Yeah, that's a great question. So I think the working together for the scale-up seems like it's probably the biggest thing that I'm most excited for, where it does seem like we see consistently using more better compute results in better models that then add more value to all the folks using it downstream. So I think that's the biggest thing. And then I think one thing now that we're investing more and more with you guys, too, and I think we can continue to, is security, where that's a place where now, as the machines are getting more complex, we can work together to secure things at the chip layer, the server level, the rack, the whole network, and that'll help out, like, not just us, but the entire industry stay safe.
L
Lisa Su27:01
No, I think you're completely right. We are super focused on the idea of, you know, how can we use AI to improve every aspect of our product, whether it's performance, power, software, and security, as you said. So really, Tom, huge thank you for the opportunity to work very closely with you and your team. Our team loves working with you guys, and we look forward to everything that we're going to do together. Thank you so much.
T
Tom Brown27:26
Thank you.
L
Lisa Su27:27
Thank you.
Alright, so look, Tom talked about how important performance is, and we have been laser focused on ensuring that the real-world performance of MI455 and Helios really comes true when we look at, you know, system results. So let's take a look at some of those results. So here's an example. We're running DeepSeek V4 Flash, one of the latest reasoning models, and we are actually comparing MI455 to MI355. And, you know, the key thing is with every generation, we want to make huge leaps in performance. So we're seeing that at lower concurrency that MI455 delivers 4x the throughput of MI355. But clearly what you see with, you know, large users across the board, that you want to increase the number of simultaneous users, you want to increase the amount of work that you can do at the same time. And so at the highest concurrency, we're seeing Helios deliver up to 34x more throughput than our prior generation. And what that means is more users, much faster response, and much better efficiency at scale.
And performance is just one aspect of it, right? The other aspect and, you know, what I should have mentioned when Tom was on stage is every one of us in the enterprise is seeing our AI budgets go up every single month. Is that right? Are you seeing some of that? And so customers really care about cost per token. And with every generation of Instinct, we're driving the cost per token down. And that's with more memory, more bandwidth, and much more compute. And we're taking another major step with MI455, delivering up to 18x more tokens per dollar, so customers can serve far more users with the same investment.
And as you think about how that comes together in a rack, many of our data centers right now are really limited by power. So power is kind of the maximum limiter. And so we tested Helios over a full range of workloads where you have a fixed rack power compared to the competition. And what we're seeing is across the highest throughput workloads, the most interactive applications, all of these leading inference workloads, we're seeing Helios an average of 10-15% more performance than the competition. And that comes with all of the system work that we've done in the overall system.
And how that translates to a customer is cost. What we expect is the Helios rack delivers more performance, and it also delivers up to 30% more tokens per dollar than the competition. So I think that's a pretty good overall value proposition.
Now, I'm happy to say that the customer demand for Helios is extremely strong. From the largest AI labs to hyperscale and NeoCloud providers, we are working across the entire ecosystem to enable Helios together with our OEM and ODM partners. And one of our deepest and earliest partners deploying Helios is OpenAI. To talk about where AI is headed and the work we're doing together, please welcome to the stage OpenAI's Head of Infrastructure, Sachin Katti.
S
Sachin Katti31:28
Thanks, great to be here. I was thinking about this event over the last two years. It feels like the attendance at this event is following scaling laws. It's been doubling every year over the last two years. But more seriously, to your question, we've always predicted that models will evolve where they will start as chatbots, then they become reasoners, then they become agents, then they become interns. And the trajectory has kept up with that. So we are all seeing that models are becoming a lot more capable, more agentic, and able to do very long-running tasks. Let them go and they just take care of it. And so all that really means is we need to keep scaling compute. The more compute we scale and throw at training, the more capabilities emerge, the more compute we scale and provide it to distributing intelligence to the world, the more people want to do with it. And this is what we see with exploding token budgets. We've seen internally, for example, in OpenAI, that not just engineering, but every aspect of the enterprise is beginning to use agents for every part of the work. This is why we need platforms like yours that can scale with how quickly the capabilities of these models are scaling, as well as how quickly the whole world is beginning to embrace what these things can do.
L
Lisa Su32:52
Well, Sachin, I don't think I've ever spoken to you where you haven't asked me for more compute. So he's actually very, very consistent. Look, our teams have been working really closely together. I think it's been a true journey when we think about just all of the technology that we've been doing. Last year, we announced a really landmark partnership where you're going to deploy up to six gigawatts of AMD infrastructure. You're one of the first, I think you're the first to actually have MI455 racks a few months ago. So can you talk a little bit about the journey?
S
Sachin Katti33:28
No, it's been a phenomenal partnership. As you mentioned, we were first to bet on AMD, and we are really thrilled with how that bet is turning out for us. It got started on MI300 and expanded very quickly to 355. And it's been really fun for our team to see how it was possible to optimize the software stack, the networking and everything, and deploy our models running on AMD infrastructure. So, as Lisa mentioned, we just got hands on the Helios racks three months ago, and the collaboration has gone to the next level. Our engineers are working side by side with AMD engineers to optimize the software stack and run GPT class workloads on Helios already. Phil will probably be on stage later talking about how productive and fast that collaboration is. So we are really excited about the capabilities Helios is already showing us, and we expect that we'll be deploying Helios at massive scale starting towards the end of this year and then accelerating throughout 2027. Like I said, I know, more faster. We need it earlier.
L
Lisa Su34:42
So, Sachin, the other thing is, one of the most exciting parts of our collaboration is actually some of the research work. So very thankful for the opportunity to spend really bringing our best engineers together with your researchers to talk about how AI should be really implemented and used for future both silicon and software. So can you talk a little bit about that work?
S
Sachin Katti35:06
Yeah, I think one of our key bets, and I think for a lot of us in AI, is about recursion, right? AI recursively improving the systems it needs to run on and eventually, obviously, AI doing the research itself, right? Today, expert engineers spend an enormous amount of time optimizing kernels, figuring out how to translate new models to run them efficiently on different hardware, dealing with all the usual things you think about, compilers, communications, libraries, system configurations. What we've been super excited and encouraged by the work we are doing with you is now AI showing the potential and demonstrating it in production to automate a significant portion of the work when anyone programs to an AMD GPU. So it dramatically has shortened the path from a new model coming up to an efficient production deployment and something that we can much more easily tune to changing workload requirements, right? And that's a big deal. So one of the other things that's super exciting for us is because AMD is built around an open software ecosystem, all these things that we are now inventing internally for our own use, we believe can, we can also bring it to the whole world and allow the whole world to use it for all of the AI models out there, not just our models.
L
Lisa Su36:33
I think super excited about that. I think the idea is that it's really the tide that lifts all boats, right? As we work, it helps our work together, but it really helps the overall ecosystem. So a little bit of your crystal ball. Looking forward, I think we all know that compute is super critical. What do you need from the next generation AI infrastructure? And what can we do as really a key partner to help you accomplish that mission?
S
Sachin Katti36:58
If you haven't heard already, I need more compute more quickly. At least he's consistent. I think the thing that the whole industry is realizing is it's a systems problem at a data center scale. So I think we obviously very quickly internalized that AI was a rack level problem, but now it is very clear that it's not just a rack level problem, it's a rack in a data center. So the whole data center is a system. And we have to think about how we design the future based on what we know about our workloads, what's coming down the line, but co-design it with you from CPUs to GPUs to memory, networking, storage, power distribution in a data center, and the cooling systems that go with it. These are all intertwined and they all have to be co-designed. And one of the really fun parts about what we do with you and your team is how easy and productive it is for us to work with you and give you insight into how we expect the world to change in the future from what our workload needs are, what our models are, and how quickly your team responds in turning those insights into better chips, better systems, better software. So that's what worked really well with MI400 and left shifted a lot of the engineering work that we needed to do to bring this up online quickly. And that's why we are very confident we can productionize those systems very quickly. And so really looking forward to accelerating that with MI500 and beyond and increasingly using AI to actually do that work for us rather than being limited by humans.
L
Lisa Su38:39
That's fantastic, Sachin. And look, we really do recognize that input is so, so valuable. So I have to say thank you again for just the extraordinary partnership, the amount of work that our joint teams are doing together, and we couldn't be more excited about the opportunities in front of us. Thank you, Sachin.
S
Sachin Katti38:57
Thank you so much.
L
Lisa Su39:03
Alright. So now let's turn to the world of CPUs. There's been a lot of talk about CPUs recently, and I can say that this has been the foundation of our AMD data center strategy for the longest time. We launched EPYC in 2017, and we've been on a very, very clear mission. Every generation, more performance, more efficiency, more capability for a wider range of workloads. Now, Naples was our first generation. It established the foundation, but Rome and Milan really changed the economics in the data center. We did more than what people expected in terms of adding core count and adding throughput. And then Turin set... Genoa expanded our performance, and then Turin set a new bar for density, throughput, and total cost of ownership. That's really been the EPYC formula. A clear roadmap and very consistent execution, and leadership that actually grows with every generation. This is why, if you look today, 5th Gen EPYC Turin is the best server CPU in the world. With up to 192 cores and 384 threads, Turin delivers leadership performance across a broad range of cloud, enterprise, and HPC workloads. But the thing about the data center is compute never stands still. With agentic AI, it's actually creating a whole new class of infrastructure, and the CPU matters now more than ever. And that's exactly what we built Venice for.
So, there's been a lot of talk about agentic AI, but it's actually a very new field. And what we see is that server computing is actually splitting across a couple of different workloads. So, the first section is GPU servers, and that's fairly well known. That's where the CPU's job is basically to drive the GPUs. And so, the job there is all about speed. We want the highest frequency cores, we want the fastest IO, and we want to keep the GPUs fully fed. Now, the biggest middle part, this is the largest growth, is actually in agent servers, or what we call agent sandboxes. And this is actually a whole new class of workload. Here, you execute code, you actually call a bunch of tools, you actually query data that's outside the model, you have a whole bunch of things around it. And in this place, the priority is actually density. So, we want the highest performing cores per watt to run thousands of agents at once. And then we have our traditional general purpose servers that run the enterprise. And here, you're going to have a diversity of workloads. It's all about efficiency, it's running the applications, it's running databases, it's running data services. And all of that is what we traditionally say is general purpose server. And so, with these different workloads, you actually need the right CPU for the right workload. And EPYC is the only CPU portfolio that leads across all three.
So, Venice is our newest and most advanced server CPU family, and it's actually designed for the agentic era. We're extending Turin's leadership across every single metric. That is performance, that's efficiency, that's TCO for cloud, that's enterprise efficiency, that's HPC workloads. And Venice starts with our all new Zen 6 core. We have higher IPC, higher frequency, and that delivers up to 1.8 times more performance than Turin. That is one of the largest generational gains. I mean, we have been doing this for six generations, but this is one of the largest generational gains in the history of EPYC. It's 203 billion transistors built on TSMC's newest two nanometer process, and it uses a next generation chiplet that supports up to 512 threads per socket. And this is actually a really important point. I'm going to show you why in a few minutes. This gives us the highest compute density, but it also allows us to double both the memory and the IO bandwidth. So I have a few more chips to show you, if that's all right. And what's clear is Venice is not one chip. Venice is actually an entire family of chips. And this is where our chiplet architecture becomes a huge, huge advantage. It lets us take the Zen 6 architecture and build a whole portfolio around it.
So let's start with this guy. This is Venice HF. This is the highest performing CPU that drives AI host nodes. It has eight compute chiplets, each with 12 cores, running at up to 5 gigahertz with the IO and memory bandwidth to keep the GPUs completely fed. And this is the CPU that I showed you before shipping inside Helios. Now, this is its bigger brother. It's Venice 256 core. This is the chip that's built for agentic sandboxes. And it also has eight compute chiplets, but each of the chiplets now has 32 cores. So it scales all the way to 512 threads. And it is the highest compute density in the industry. Customers get more agent capacity per watt, per dollar, per rack, bar none. Okay, and then one more member of the family. This one is Venice 128 core. And this guy is optimized for enterprise and general purpose servers. This is the same leadership performance per core, but it's in a lower cost, lower power design for enterprise applications. And this one is really tuned to give customers the best performance per dollar across a wide range of configurations. So you can see really the power of the family as they come together. And in addition with these guys, I can tell you that next year we're adding a few other chips to the family. We're adding Verano, which is the next generation for AI host nodes. And what that adds is a more power efficient, low power memory, as well as faster interconnect between the CPU and the GPU. And then we're also adding for our high performance computing, our technical computing, Venice X, which is using our 3D vCache stacking technology. And we're really the first to stack memory chiplets right below the compute chiplets to accelerate performance, even for those types of workloads. So you can see one architecture. A CPU is not a CPU. There are many different types of CPUs. And for us, one architecture spans dozens of different chips. And that breadth is something that no other CPU vendor can match. And that's what makes 6th Gen EPYC the best CPU in the data center.
Now you're going to bear with me because I'm going to take you through some performance because the numbers are just incredible. So looking at AI workloads, right? Again, you have different use cases and different workloads, starting first with GPU servers. Many of the largest frontier models require you to move data between the CPU and the GPU. And in that scenario, Venice moves the data faster than the best competitive x86 processor, delivering up to 1.8 times more tokens per second. And when you look at the agentic CPU servers and sandboxes, because of our density, Venice delivers more than twice the agents per watt for agent sandboxes. And for general purpose applications, in this case, we're standardizing on a 100 kilowatt rack. Venice delivers more than twice the performance per watt. So pretty incredible results. And the gap is even wider if you look at ARM processors. So if you look at the leading ARM processors, when we're comparing at the chip level, Venice supports up to 2.8 times more agents per watt. And when you look at the rack level, Venice is delivering up to 3.3 times more performance per watt. And what that means is at the data center scale, you can just get a lot more agents in the same power envelope.
Now, there's been a lot of talk about what matters most for agentic AI, with some saying that per-core performance under load is the only thing that counts. But that's really only part of the story, because agentic AI actually runs as a distributed platform, and that is across databases and vector stores and orchestration. And all of that has to happen at once. And that requires not just raw speed, but you absolutely need density, efficiency, and per-core performance all together. And what I'm happy to say is that Venice leads across every one of them. So when you compare Venice against the highest performing ARM CPU from our competition, EPYC delivers 20% higher per-core performance. And when we... I like that number. I like that number. And when we pair that with leadership core density, we're delivering 2.2 times more performance at the socket level. And so the main point is, whether a customer needs per-core or per-socket performance, Venice is the leader. And there's one more point that a benchmark doesn't really capture, because when you think about what agentic AI is, it's really like you're adding thousands of employees to your enterprise, to your traditional workflow. And that software, enterprises have been running on for years. The truth is, that software runs on x86. So Venice runs all of it. And this is where x86 has an advantage. You must have the best performance. You must have the best density. But having that software compatibility is a big, big plus. So customers scale up their agents on the stack that they already have.
And now, let me just show you what that means at the rack level. So at the rack level, what we find is density is super important. Everyone's trying to optimize their data center. Everyone's trying to optimize their power envelope. And with Venice, every major server OEM is offering a broad spectrum of racks for agentic AI. So what that means is we have the full spectrum from 25,000 cores to 50,000 cores. And that's more agents per rack. And that gives customers the choice. The key is, every data center is different. And with that, the flexibility allows you to pick the right cooling, the right space requirements for what you're trying to do in your data center. So there you have it. Venice, the best server CPU in the industry. And I'm very happy to say also today that Venice is in full production. Customer demand is incredible. It's the strongest we've ever seen for a new EPYC generation. We're seeing every major server OEM, every major cloud provider on track to begin rolling out in the fourth quarter as we start with making the broadest EPYC launch we've ever had. So with that, I want to turn to my next guest, who runs some of the largest and most advanced compute infrastructure in the world. And it's really been our privilege to partner with them as they've deployed multiple generations of both EPYC and Instinct. To share more about our work together, please welcome to the stage Meta's Head of Infrastructure, Santosh Janardhan.
S
Santosh Janardhan51:50
It is so great to have you here. Thank you for being here. Santosh is a true friend. I have to say it's been a tremendous opportunity over the last few years. My story about Santosh is the first time I sat in his office. He said to me, Lisa, we're going to deploy a lot of CPUs. Please make sure they work. And I said, I will, I will. And you lift it up. I try. I try. But look, Santosh, Meta operates, you operate some of the most advanced infrastructure in the world. What's happening in the data center world? What's happening in your world? Tell us a little bit about the architecture and what you're working on.
Sure. First of all, thank you. It's awesome. This is like a hardware geek's paradise, right? You come in, you're talking about transistors, you're talking about hardware, talking about cooling. This is sort of the gem. I like it. And I think there's an audience that's receptive to it. I usually talk to audiences. They have usually either not been to a data center or not really seen a chip ever. So this is refreshing, I have to say. The thing I'll say about... Listen, we are seeing demand go through the roof. When you look at inference, training, recommendation system, that's how a newsfeed works. Content creation. The demand for that is just exponential right now. And when I think about sort of the overall... Mark has this vision of delivering personal super intelligence to billions of people. So if you go to sort of Facebook or Instagram or WhatsApp or whatever surface you go to, the idea is that we'll meet you there and we'll deliver super personalized intelligence right wherever you are. And that's why we established this Meta compute initiative and because it's a top level thing, Mark himself oversees. Now what it means is that we are now moving away from an area where we used to take servers, optimize it. This is what the conversation we were having many years ago, that, hey, I'm deploying a CPU, I really need to maximize performance out of it. It's different now. You talked about how we should be thinking about the whole system end to end. We are at a point where we have to look about the whole data center and think about it as one integrated system. Servers, hardware, networking, cooling, power, all of that is not optional. All of that has to work together. And this is where I think there's a huge opportunity for collaboration with partners like AMD because you have to now co-design, co-create systems, not just deploy something that comes off the shelf. The other thing is, by the way, flexibility. The thing I really like about what you said is that this is such a big opportunity in the industry right now. All of us need to lean in and work on this together. So it needs to be open. It needs to be heterogeneous. And I think that's not one company, one partner that works for any one of us. We need to sort of work on this together. And that's why I think you're one of our most important strategic partners.
L
Lisa Su54:53
Thank you, Santosh. And look, we completely agree with your philosophy. Now, you are actually one of our broadest partners because if you think about our work on CPUs, GPUs, you know, we developed the Rackscale OCP systems together. You've deployed lots and lots of CPUs, so millions of CPUs. And you're a lead partner on Venice. Can you just talk a little bit about sort of your CPU and infrastructure and some of our work together?
S
Santosh Janardhan55:19
Sure. It's been many years. And the story that Lisa was saying was many years ago. I think we went from Milan to Bergamo to Turin, now to Venice. So it's at least the fourth generation. So a long-time partnership, obviously. I actually think the world is changing in the sense that it used... The world is changing. Every year is a new world these days. It's like every month. Months, exactly. So what ends up happening now is that it's not just a GPU game anymore. It's CPUs and GPUs. If anything, I think CPUs are becoming at least as important, if not more. You're looking at a world where sort of the workload is changing because while you have your workhorses and GPUs, at the end of the day, there's agentic workloads. You have to run your tools. You have to run your systems. And all the code that's been generated still needs to run somewhere, right? So I'm super excited to work on Venice. I think there's a lot to come there. And you have to think again. This is a theme that I'm assuming will go throughout the presentation that you have to start thinking about CPUs and GPUs as conjoined things. You hand off workloads. Depending on the workload, you employ the right hardware.
L
Lisa Su56:30
Yes, absolutely right. And we've gotten a lot of feedback from your technical team. I think that's what I really appreciate about the partnership with Meta. Now, clearly, you're deploying lots of GPUs too, a lot of accelerators. And you're also one of our deepest partners on the accelerator side, starting with MI300. And then we had a very large strategic partnership announced around up to 6 gigawatts, starting with MI450. And what we're doing is quite unique with Meta. So can you talk a little bit about the evolution and the MI450 plans?
S
Santosh Janardhan57:06
Again, multi-year collaboration. I think we started with MI300. We did a bunch. It had really nice memory capacity. I remember talking to you about that, right? And then there's a bunch of pipelining we did. It was important to just have deployed that at production scale, sort of get hands-on experience on both sides. 350 is the first time we deployed this on our ranking and recommendation systems. So that's now getting into sort of some good numbers. And 450 is, I think, I'm super excited about it. We just got some of the racks. And we are going to deploy it across the board because it gives us the opportunity to go and sort of really collaborate. I think the big difference between the 300, 350, and 450 is that we have had engineers sitting in the same room talking to each other, co-designing, co-creating, like power cooling. Like I was saying, it's not just a single thing. And we are not shy in feedback. They're not shy, for sure. But you're very receptive. So this is true partnership, right? This is the point about co-designing and co-creating that, I think. And like I was saying, that's a huge sort of deployment coming in. The more we deploy, the earlier we co-design, the better we are.
L
Lisa Su58:19
Yeah, absolutely. Really, really appreciate the effort together on MI450. Now I want to ask you also about your crystal ball. So you look out into the future and you see what the next few years really means as you push forward in the AI industry. What are the biggest challenges and really opportunities for us? This is the industry ecosystem here. So what should we be focused on as an industry?
S
Santosh Janardhan58:42
Listen, all of you know about the bottlenecks. You know about the power, the data centers, the silicon. All of these, I think, are choke points that I think the industry is waiting on. But those are things I'm pretty confident we'll sort of work our way through. At the end of the day, I think the demand is immense. People are responding. But the thing that I really want to make sure all of us realize is that we should not think about systems in isolation. Like I was saying, CPUs and GPUs are one conjoined system. We should start thinking about it that way. We really need to co-design things early. See, data centers take years to build. One of my favorite stories is Mark comes to me and says, hey, I want a gigawatt worth of data centers. Well, you should have talked to me two years ago. It takes time. It's the same with silicon, right? It just takes time. So we need to be starting and sitting down in a room co-designing today for what we need to deploy in 27 and 28. That, I think, is when we truly unlock all sort of powers of the system. The systems are amazing. The specs are amazing, right? But think about how much more you can get out of it if you sit in a room and just co-design it today. And in 28, we'll be having much better graphs out there.
L
Lisa Su59:57
Fantastic. Santosh, I completely agree with you. I want to say again, thank you. We are so, so happy with the amazing partnership that we have across the board and, most importantly, with your engineering teams. And we truly are excited about what we're going to do in the future. Thank you so much.
So you heard Santosh talk about just the diversity of workloads and really the massive scale that you require with AI. Now I want to turn to another part of the inference market where the requirements are actually very different. So lots and lots of workloads. And as inference is moving into more products and services, we're actually seeing it segment into different workloads depending on what you're trying to do. Some of these are, let's call it, less interactive. So they're really tuned for, let's call it, maximum throughput or the lowest cost. Other of these applications require more balanced throughput, so you have to balance throughput and responsiveness. And then there's this new class of applications that actually want very, very fast results. You have extremely smart engineers and they don't like to wait. And that is where ultra-low latency or every millisecond actually matters. Each one of these types of workloads requires a different type of compute. And one of the best ways to reach ultra-low latency today is disaggregated inference. So if you think about the different pieces, both prefill and decode are two different jobs. So prefill tends to need more compute. Decode tends to need more memory bandwidth. And with Instinct, we have a very balanced machine so that we deliver both great compute and great memory bandwidth. But you can actually take this a step further. If you know what workloads you're trying to run, you can actually let customers tune each of these pieces independently. And to address this market opportunity, we've been working with Cerebras. So to talk more about what we're building together, please welcome to the stage Cerebras co-founder and CEO Andrew Feldman.
A
Andrew Feldman1:02:10
Andrew, it's great to have you here. It's been an exciting few months for you. I know that for those who don't know exactly what you've been working on, can you talk a little bit about Cerebras and what problem you've been trying to solve?
Sure. Great to be here and great to be sort of among people who love hardware. Yes, it's nice. Look, I'm thrilled to be here with you and to announce our cool new partnership. At Cerebras, we build the world's largest and fastest chip. It's a full wafer. And we package this into a system into racks and then racks into clusters. And we deliver them both on-premise and via the cloud. And you've had some of our customers up here already today. On the Frontier Labs, we serve customers like OpenAI and in the hyperscalers like AWS, in the small agentic and the coding space, leaders like Cognition. And they do this because we're blisteringly fast.
L
Lisa Su1:03:19
There's no question, Andrew, that you have some tremendous innovation with what you've been working on. And our teams have been collaborating really closely over the last few years to really have Helios together with the wafer scale engine for ultra low latency inference. Can you just kind of just educate the audience a little bit? What are we trying to do and what does that mean for customers?
A
Andrew Feldman1:03:39
Sure, I think what's happened, and you described it previously, is that sort of AI has moved from being a novelty to being useful and then in some domains a necessity. And when something is a necessity, people want to use it and they want to use it quickly. And to serve this market, this segment of ultra low latency, we sought a partnership that could extend our capabilities and our footprint. And there was no better answer than AMD. We were already an AMD customer as we use AMD CPUs to surround our systems. And we were so excited at the opportunity to build a disaggregated solution that combined AMD CPUs, the cool new Helios rack and the cerebrus wafer scale engine.
L
Lisa Su1:04:33
Yeah, look, the technology is pretty cool. So just talk a little bit about how it comes together.
A
Andrew Feldman1:04:40
Sure. So as Lisa described, what you have with Instinct and the Helios rack is that you have the leader in performance and memory capacity. And you marry that with our wafer scale engine, which is the leader in sort of in SRAM and in memory bandwidth. And that combination allows us to deliver a solution that is unmatched in the industry. Very, very exciting.
L
Lisa Su1:05:06
I know that customers are excited about what we're going to do together as well. So when are we going to have it in market?
A
Andrew Feldman1:05:12
Later this year. It will be first in the cerebrus cloud and it will be later at a store near you.
L
Lisa Su1:05:25
I think the key thing is we're going to give customers a choice to really put together what is the solution that they want.
A
Andrew Feldman1:05:32
And I think we're very, very happy to be able to do this together with you. So I think it's a huge step up in terms of what we can do for this ultra low latency segment.
L
Lisa Su1:05:40
I think that's right.
A
Andrew Feldman1:05:41
I think until recently, customers sort of, they could have sort of high throughput or they could have extraordinary speed. And by bringing together the Helios rack with the cerebrus wafer scale engine, we give you five times the throughput while continuing to deliver this extraordinary speed. It's really something amazing.
L
Lisa Su1:06:07
Fantastic. Well, look, Andrew, huge congratulations to what you and the Cerebras team have done. I mean, I think it's been incredible. I think you've been clear on what you're trying to accomplish. And it's really nice to see not only it come together, but us come together to offer a solution that will be very, very compelling to customers. So we can't wait to get this into customers' hands.
A
Andrew Feldman1:06:27
Soon. Thank you.
L
Lisa Su1:06:29
Good to see you.
A
Andrew Feldman1:06:29
Thank you so much.
L
Lisa Su1:06:30
Thank you, Andrew. Alright. So look, I think we've shown you a lot of hardware this morning. But as we all know, hardware is only part of the story. It's actually software that turns all of this into a platform that developers can really build on. So to share the progress we're making with ROCm, please welcome AMD Senior Vice President of AI, Vamsi Boppana, to the stage.
V
Vamsi Boppana1:07:03
Thank you, Lisa. Good morning, everyone. Good morning. It's great to be back to talk about software. Just a few years ago, programming AMD GPUs required deep expertise and significant engineering effort. We've come a long, long way in a remarkably short period of time. Through our open source community collaboration and sustained investment in ROCm, our software stack, AMD Instinct GPUs are powering some of the most important and consequential workloads on the planet. But the biggest change is still ahead of us. AI is transforming how software is getting built, and we are bringing that transformation to ROCm. And that's what I'm excited to share with you guys today. Our software teams have been moving fast with relentless focus on developers. We've invested at every level of the stack in build and test infrastructure so we ship faster. ROCm releases now go out every six weeks, not every four months. We've continued to expand the work we do with our AI ecosystem partners. And the result? More features, faster performance, and a significantly better out-of-the-box experience. Our strategy that has gotten us here has been remarkably consistent. It's built on two core principles. Partner deeply with the open source ecosystem and build with the right layers of abstraction to enable developer productivity. Open source gives us velocity and scale, and abstraction makes developers more productive. Our deep, deep commitment to open source has resonated with the community that has truly embraced AMD. From frameworks and compiler stacks to inference engines and model hubs, AMD is now becoming part of default enablement for the most important AI communities, such as Hugging Face, PyTorch, JAX, vLLM, and SGLang. This is why new models get Day Zero support and why the world's most important AI workloads run on AMD today. That is a big shift. Inspired by engineers seeking productivity, abstractions have evolved to enable work at the right level of detail. Hiding complexity when they want speed, and exposing control when they need performance. From low-level programming in assembly and C to block programming abstractions like Triton to Pythonic frameworks, we've invested in key abstractions. Aiter, our kernel library, and Atom, our serving engine, get you peak performance without writing the kernels yourself. FlyDSL is a brand new Pythonic domain-specific language that gives you low-level control with the performance of hand-tuned C++. And Mori, our communications library, helps deliver leadership performance on the latest models, like Minimax. But look, something very, very exciting is happening. Over the past year, I've seen something remarkable inside AMD. Our engineers are using AI models to generate GPU kernels, optimize code, debug issues, and improve performance. And in some cases, these AI-generated kernels are shockingly good. Better than what we expected. Sometimes better than the most manually-tuned versions. The first time you see this, you're a little bit skeptical. You run more tests, you try to break it, you look for what went wrong, and then you realize this is real. What has happened in general software development is coming to significantly transform GPU programming. It's going to reduce the time it takes to bring up workloads, make optimization automated, and it will remove any remaining barriers to broad adoption. We want to put that capability in the hands of every developer.
So today, I am excited to introduce ROCm.ai. It's an agentic AI platform that brings the capabilities of AI-assisted GPU programming to developers. It starts with the simple idea of giving developers access to the power of an AMD ecosystem through the AI coding agents they already use. Whether you're using Cursor, Claude, Codex, rocm.ai helps those agents understand AMD platforms, understand ROCm, and helps you build and optimize your workloads. In other words, we are making those popular coding agents into ROCm super users. You should be able to describe the workload you want to run, the performance target you want to hit, and then let the agents help you get there. For years, we have been building the open software foundation. Now, we are adding an AI-assisted layer that helps developers use that foundation faster and more effectively. So let me show you what is inside ROCm.ai. It's built on the foundational capabilities of ROCm, our core software stack, runtime, libraries, compilers, tools, framework integration, all of the core software that makes the stack work. On top of that, we built a layer of AI-assisted optimization called Hyperloom. Using Hyperloom, the system can analyze the workload, tune configurations, select and tune kernels, adjust parallelism strategies, and iterate towards performance goals. To give you an example of the power of Hyperloom, I was talking to the lead developer last week. He told me they just pushed through a suite of 14,000 models through Hyperloom, optimizing them and creating valuable insights. This would have been impossible to imagine, even with a large team of engineers before. That's the power of Hyperloom. And finally, we provide an AI-native interface for our developers, AI skills that allow coding agents to understand ROCm natively. I'm incredibly excited by what we have built with rocm.ai. It's going to profoundly transform the way developers interact with our platforms. So let me make this concrete for you with an example. Imagine you are a developer who wants to optimize Minimax M3 with vLLM on MI355s. Today, there are many things you need to get right. You need the right vLLM optimizations, the right model recipes, the right environment variables. You may need to write new kernels or optimize kernels, and then you still need to tune all of this. With rocm.ai, this interaction becomes much simpler. So let's see it. You can say, optimize Minimax M3 with Hyperloom. Enter. And behind the scenes, the agent does what an expert would do. Pulls a known good recipe, runs and profiles the workload. It creates and tests a bunch of kernel and runtime configurations, and it keeps iterating towards target. You can actually watch on screen here rocm.ai write a GPU kernel. It sees an opportunity to write a more optimized MOE group gem. And the result? rocm.ai delivers a 38% more tokens per second improvement on this specific example. This is the experience we want for developers. Simple to start. That's pretty cool, huh?
Simple, simple to start and powerful to optimize. We've continued our relentless focus and performance. With every release, we've delivered significant performance gains. On leading models like DeepSeek, rocm.ai delivers 3.3 times speed up over ROCm 7. This has been made possible by innovations at every level of the stack. Techniques such as the use of block-scaled fused MOE kernels, KV cache quantization, export parallelism. But look, the bigger story is how AI agents actually profiled the workload, proposed new kernels, tested configurations, and actually validated all of the results. That exact same focus is also on delivering performance gains for training. On the same hardware, rocm.ai improves training performance by an average of 2.4 times. Through optimized kernels such as fused flash attention, more efficient checkpointing, and advanced parallelism strategies. Our engineers have done amazing work to unlock the capabilities of MI455 through rocm.ai. So I'm delighted to share some incredible results on hardware. On running real workloads, we are able to see the system deliver 20 terabytes of memory bandwidth and 20 petaflops of delivered compute. These are the highest demonstrated compute plus memory capabilities of any accelerator platform in the industry today. It's not theoretical, there's measured. Super exciting that we were able to share this on day one. I'm also thrilled to share that the leading AI ecosystem partners who have had their hands on MI455 have had a delightful experience. You can see here what they've had to say. From end-to-end PyTorch testing, to ensuring Hugging Face models are getting validated, to vLLM and SGLang running well on MI455. We are incredibly encouraged and excited by what the community is about to unlock with MI455. This matters because day zero readiness is one of the most important objectives we strive for, and that's what rocm.ai enables. So now let's look at rocm.ai in action on Helios. This time we are going to use a Codex front end. Let's deploy DeepSeek V4 Pro, a 1.6 trillion parameter frontier model on Helios. Now watch what happens. rocm.ai pulls the right recipes, it configures the runtime, it implements the model graph onto hardware, and finally serves up the model. There you go. The model is actually up and running. And let's see the model write a simple poem now about advancing AI. Let's hit enter. Isn't that a pretty cool poem? This is the experience we want to give our users. The developers should not have to manually reason through every layer of model architecture, kernel selection, COM strategy, and rack level deployment. AI is transforming this experience, and that is a very, very big shift for GPU hardware and software.
Now, there are few organizations shaping the frontier of AI as much as OpenAI. You heard Lisa and Sachin talk about our partnership. It's been truly wonderful collaborating closely with the OpenAI technical team to make rapid progress with ROCm. To talk about how they are advancing the frontier of software development and our collaboration, I'm delighted to welcome to stage Philippe Tillet from OpenAI.
Thank you for joining us here, Philippe. So just so you all know, Philippe is the creator of Triton, one of the most important software innovations in modern AI infrastructure. It's a huge privilege having you here.
P
Philippe Tillet1:19:08
Thank you. For sure. Well, first of all, thanks so much for having me. I was here two years ago. It's my great pleasure to be back now. A way of looking at things is there's no one way to fit all applications, and we really try to meet users where they are. So for that reason, we've developed multiple solutions. So for workflows where speed of iterations is more important, we've developed Triton, as you mentioned. So this is typically what researchers will use. But more and more, we've been in a case where we've had to hyper-optimize performance, and Triton didn't give us the level of control that we wanted. So we've developed Glion for that, which is a lower-level language. And this way, we've been really able to optimize our kernels for AMD, among other things. Our agents have become extremely good at both of them, and we've worked with AMD on both of them as well. More and more, what we're seeing is agents slowly taking over. They're getting just very good, as you just mentioned, at writing this kind of code.
V
Vamsi Boppana1:20:31
Amazing. So look, I think the work you have done with Triton and now Glion, it's been really, really moving the ball forward for the entire industry. Now, we began our work together with 300 and extended at 350, but now we're collaborating at a different scale with 450 and Helios. So I would love to hear your thoughts on how that collaboration has been going, and how has that experience been?
P
Philippe Tillet1:20:53
Yeah, super well. It feels like ages. I think we started years ago on hardware co-design. You know, when MI450 was in the very early prototyping stages, I think we were able to work really well together, and it's so nice to see the outcome of that and the chip. And more recently, we've been working very, very deeply on software enablement for MI450, both for our models, but also for the broader ecosystem. So a lot of this development has been done open source. And yeah, we're collaborating on the full stack now, down to the very low-level LLVM code generation, so that all the nice hardware advances that came with MI450 can be properly leveraged in end-to-end applications.
V
Vamsi Boppana1:21:42
Yeah, and you got your Helios racks, so tell us a little bit about that.
P
Philippe Tillet1:21:46
Yeah, it worked. We got the chip, and very quickly, a single engineer in a few days was able to confirm that everything worked. Obviously, we have to optimize performance more. But we're very confident we can get in a very, very good place.
V
Vamsi Boppana1:22:09
Awesome. That's awesome to hear. So look, one of the more amazing things that we've done recently is our teams have collaborated even more closely in terms of using AI to program AMD GPUs. So we'd love to hear your thoughts and anything you can share with the audience.
P
Philippe Tillet1:22:25
Yeah, no, I mean, everyone in this room knows how much agents are taking over our jobs, in a sense, as kernel engineers. And it's been amazing just to see how much better they've got over just the past six months. And just now, they're extremely good at generating high-quality GPU kernels in a way that were not before. And it's been very great collaborating with AMD on just making sure that agents not only are capable enough, but also have all the right context they need to be able to make good decisions and good optimizations.
V
Vamsi Boppana1:23:07
It's great to hear. I think the quality of code that's getting produced by AI-assisted kernel generation has been truly remarkable through our collaboration. Very excited. Now, as you look ahead, what crystal ball do you have for us for the future of AI software ecosystem and collaborations like ours?
P
Philippe Tillet1:23:28
Yeah, so there's two things that really come to mind. One, as we keep talking, is just how good agents are getting. And the other one is, I think, how much the open approach to software of AMD is enabling, not only for OpenAI, but I think for the industry as a whole. I talked a little bit earlier how we've been collaborating on LLVM co-generation. LLVM is an open source project. The entire compiler stack from AMD is open source. And that has allowed us to make our agents really good at very, very low-level co-generation, down to scheduling instructions in assembly. And this has led to very, very significant performance gain that I don't think we'd have been able to achieve in a fully closed source stack.
V
Vamsi Boppana1:24:15
Yeah, that's one of the real sort of big value propositions that we have to offer, because everything that we do is in the open, and the fact that AI can actually pick that up is a big one. Thank you, Philippe.
P
Philippe Tillet1:24:31
Yeah, my pleasure.
V
Vamsi Boppana1:24:39
What you just heard is extremely important. When we combine better abstractions with AI-assisted development and closed hardware software co-design, we can create some incredibly powerful capabilities, and that's exactly aligned with our strategy. Now, everything I have shown you so far has been about AI, but ROCm is also a great stack for developing high-performance and scientific computing applications. From frontier models to climate simulation, we use the same ROCm, the same library, same tools, same AI-assisted development, one open stack for every single workload. And in fact, when you look at our roadmap, we have retained that very strong focus on HPC and scientific computing. Just like MI455 leads in AI compute with FP4 and memory performance, MI430X leads HPC compute with FP64 and memory performance, two platforms purpose-built for two very different classes of workload. The MI430X is built by leveraging our modular chiplet architecture to deliver a purpose-built variant with native FP64 hardware, giving customers leadership performance across AI and HPC in a single platform. And this is hardware double precision, not emulation. For the scientific community, that distinction is everything. This product delivers 288 teraflops of FP64 compute, nearly nine times better than competition. And the same leadership, memory capacity, and bandwidth of MI455. The MI430X ships in the first half of 2027. This is why leaders in sovereign AI love this product. The first sovereign exascale AI factories in both US and Europe are being built on the MI430X. Discovery at Oak Ridge National Labs and Alice Recoque with Cines and Genci in Europe. So look, as we look ahead, let me leave you with this thought. Every so often, our industry goes through inflection points. We've all felt it when neural networks first started recognizing cats and dogs, when transformers changed what was possible with AI. I believe we are experiencing another one right now, a real inflection in the ability of AI to program complex hardware systems. Building on the surface area that's been exposed from open source software and leveraging the power of abstracted interfaces, AI is making it dramatically easier to program AI systems. We cannot wait what you will build with it. Thank you. Now, AI is getting pervasively infused into all forms of computing. Enterprises are experiencing the same shift with their own unique requirements. So to tell us more about that transformation, please welcome to stage Senior Vice President and General Manager of Compute and Enterprise AI, Dan McNamara.
D
Dan McNamara1:28:00
Thank you, Vamsi, and good morning, everyone. It is a great pleasure for me to bring to you the third pillar of our strategy, which is powering AI everywhere. And the vision behind this has remained the same for several years for us, to deliver the right compute engine to the right workload, and then continuously drive optimization points across our entire portfolio, from CPUs, GPUs, networking, and FPGAs. So building on what we've already shared, I'll cover how our strategy comes together across our enterprise portfolio and how the foundation we built with the CPU franchise has truly positioned us for the next era of AI across multiple deployment models. So Lisa talked about the frontier model and hyperscale landscape and just the rapid and massive advances we've seen over the last 6 to 12 months. And those advances are driving the adoption of AI well beyond the cloud. And as AI continues to grow, it will demand compute beyond the massive purpose-built data centers that we all know and love, extending into enterprise, personal, and physical AI. And each of these brings a new or unique set of requirements. Enterprises are usually constrained by power and cooling. Personal AI usually must operate within a fairly limited compute and memory footprint. And physical AI must perform reliably in real-world demanding environments. And AMD is the only company delivering a complete portfolio that spans all of these development models. So I just mentioned that our journey began in the data center with our CPU franchise, and I want to touch on it a bit. With each generation of EPYC, our strategy has been to listen to our customers, address their evolving needs, and deliver the highest performance, lowest total cost of ownership, and the fastest time to value. As Lisa mentioned earlier, we've introduced many industry-firsts across our generations that our enterprise customer has actually asked for. And that customer focus has brought us to where we are today, with EPYC as the leader in enterprise computing. And over the last three years, there's been very, very strong adoption across the enterprise, with the world's largest companies moving more and more workloads to AMD. And it's reflected in a number of places. First, public cloud, where our VM consumption has grown at 76% CAGR over the last three years. And on-prem deployments with our OEM and ODM partners is growing quite aggressively across key industries. But just as important, this position allows us to understand truly the diversity of the workloads that our enterprise customers run every day. And each workload faces new demands on the CPU. Some require maximum thread count, some depend on per-core performance, while others depend on memory bandwidth, cache efficiency, or I/O performance. The really important point is that there is no single SKU that solves every enterprise application. And that's why our EPYC portfolio with Venice spans 8 to 256 cores, with a range of power, frequencies, and I/O capabilities, giving customers the flexibility to choose the right solution for their workload, and most importantly, at the most optimized cost. So, Lisa mentioned this earlier, and now we all know, we covered Venice, and she showed how we have strong leadership across AI host nodes and the agentic workflows. But I wanted to show you how Venice actually leads across the general purpose workloads of the enterprise and these workloads that are running our customers' line every day. As you can see, we have a minimum of 2.5x the performance across our competition, across these very, very important enterprise workloads. So now, as we all know, agentic AI is really the next major transition for the enterprise. And it truly is a force multiplier for enterprises, and adoption is accelerating. In fact, the latest data here shows that nearly every enterprise next year will have some form of agentic deployments. And as enterprises move from pilots to production deployments, we've talked to CIOs all the time, and they are looking at three major things. First, predictable infrastructure cost and token cost. Secondly, making sure the data is secure. And then lastly, and sometimes this is overlooked, right, integrating AI into their existing infrastructure. It's not all greenfield out there for the enterprise customer. So to solve these problems, we truly believe that this is going to be a distributed deployment model that leverages frontier models, cloud services, on-prem infrastructure, and of course, AI clients. And as AI has moved from chatbots responding to prompts to agents doing real work, writing code, summarizing documents, and running complex flows, the requirements have evolved also. And those are exactly the requirements we've been designing for for multiple years. So while AI will be deployed across all these environments, the on-prem deployments in our enterprise customers have fairly unique constraints. First and foremost, power/cooling. Really big challenge in space. And of course, costs remain top of the list for CIOs that are trying to drive things forward in the AI world. And then lastly, the full solutions they need to be easy to deploy. That's been common for many, many years for the enterprise, but that's something we're very, very focused on. So to address these constraints, thank you, Laura, I'm very happy to announce the launch of the Instinct MI350P. So today, every enterprise wants AI in the data center, but the challenge is how do you add those capabilities without starting over? That's exactly what the 350P was designed to solve. It's an air-cooled GPU designed to fit within the power and cooling of enterprise servers today without a facilities upgrade. Making it easy to bring LLM scale inference into today's enterprise data center. And at the same time, there's no compromise in capability. A single MI350P can support up to 260 billion parameters, allowing customers to run the majority of enterprise AI workloads on a single GPU. And the economics are actually even better and more compelling. The 350P delivers more than four times the tokens per second per dollar than the competition. It's pretty cool. And it actually turns the customer's existing data center into an AI data center. Now, we all know that benchmarks are one part of the equation, right? The other part is really evaluating, does the advantage hold up across the actual workload you run and how your business runs? So we tested a set of workloads, enterprise use cases, using production models. And the 350P delivers much higher productivity than the competition, delivering up to two to five times faster tokens per second than the competition. So if you think about that, that's more tokens per second, more work completed, more users supported, all at a lower cost for the business. So another area that I wanted to talk about is, our earliest customer is our own AMD IT team. And they're a tough group to work with, even for us. But they put us through our paces. And really, this has been our model from day one is, we want to put our technology into our own data center first. So like many companies, many enterprises, we're looking at how to deliver AI service to our employee base, but also manage costs and protect our data, manage and control the data. So we focused on two use cases that we wanted to try out. And first was autonomous threat detection. And that is probably one of the faster-growing agentic applications that we're seeing today. And the second was a personalized AI assistant running on OpenCloth. So both of these run route requests through the router, and the gateway sort of tests the complexity, the sensitivity of the information, and the service level requirements. So some requests are sent to the Frontier model, while others are routed to OpenWeights models running on EPYC and an MI350P. So with intelligent routing, we reduced our token costs by 43%, while delivering up to 3x faster response times for the workloads running locally. So we believe this is how enterprise AI will be deployed. But we also know that, from our history here with EPYC, that enabling enterprise takes a whole lot more than great silicon. It requires a complete ecosystem of software and solutions. And over the last several years, we've truly grown significantly and expanded our enterprise AI ecosystem. Today, we support numerous AI ISVs, provide zero-day support for all of the industry's leading open models, and deliver robust frameworks that help customers bring AI into production as fast as possible. And it's all built on trusted platform software and OEM partnerships that enterprises rely on every single day. Now, everything I just shared comes down to one goal for us, and that's helping customers deploy AI where it creates the most value for them, whether it's in the cloud, on-premise, or at the edge. So to bring this to life, I'd like to welcome Jeremy Legg, Chief Technology Officer at AT&T.
J
Jeremy Legg1:39:09
Good to see you. Thanks for having me. Yeah, this is great. It's kind of a homecoming. I grew up in the East Bay, so it's good to be back.
D
Dan McNamara1:39:19
Yeah, fantastic. We really appreciate you joining us today. So look, I think everyone knows AT&T as the largest communications company, one of the largest in the world, and I'm sure half of them out there probably have AT&T network right now. Hopefully more than half. But AI is touching way more than the network, and you and your team are really doing a fair amount of work driving it. So can you share with us sort of the opportunities you approach, you're working on, and your experience to date?
J
Jeremy Legg1:39:52
Sure, happy to. At AT&T, we're now burning about a trillion-plus tokens per month, and that number has gone up pretty dramatically. It's moving up double digits, and so this is now becoming very pervasive inside of the enterprise. We've got over 100 Gen A models in production across our enterprise as well, and those are stretching from things that I think a lot of folks are doing in customer care, things that we're doing in fraud, but also things that you wouldn't necessarily think about, like where's the best place to place a cell tower or a RAN that is the most effective use of it inside of a network. So all of these things are happening across our enterprise, but they're happening at a scale that is not necessarily normal. We get over 300,000 calls a day. The transcription of those calls is an enormous workload, and then driving the insights out of those calls that we then bring back to the business is another set of insights. So it's really an interesting thing, and now we're beginning to actually rebuild entire workflows inside of the company beyond just the point use of a use case. So as we think about HR, as we think about finance, and we think about different kinds of things that each of those organizations do, cash forecasting, things that we do on onboarding employees as examples, we're now agentifying those entire workloads.
D
Dan McNamara1:41:20
Yeah, so we work closely with your team, and we've had many conversations. It's truly amazing the number of actual opportunities you're addressing with AI, so a super broad range of use. So can you just explain sort of what are the key learnings in moving these to AI at enterprise scale, and then maybe a bit on what we've done together?
J
Jeremy Legg1:41:45
Sure. I mean, I'll start at a place that I don't think everybody starts from here, which is the human. You need world-class teams in order to do this, and you need people that think in workflows. There's an enormous amount of what I call racing from stoplight to stoplight. Folks are going from 0 to 100, and then they hit the next stage of the workflow, and then there's a red light, and they haven't worked their way all the way through that. So I would start with people and the quality of the people that you really have across your enterprise in order to do these things. The second point is we are big believers in data sovereignty, and data sovereignty for us means not being tied to a specific chipset, not being tied to a specific model or a specific set of development tools. You need, as an enterprise, to manage your data. Your data is your fuel in terms of how you implement AI across that broader enterprise. What we've done with AMD, and AMD's been such a great partner in this, is leveraging open source models, leveraging different sets of chipsets, building our own models, training other models. We're one of the few companies in the world that's actually post-trained a model on AMD and then been able to drive equivalency in terms of performance and accuracy with other more expensive chipsets in the marketplace. So as we have done that, we've been able to drive down token costs, essentially through money-balling across all of these different models and leveraging that human capital in a way that's made a big difference to our company. And so increasingly, even though token consumption is going up, we're able to manage that underlying token cost at enterprise scale, which is not a small thing.
D
Dan McNamara1:43:31
Yeah, it's a bunch of great work. Another key area that's sort of groundbreaking, which is exciting also, is you're the first telecom company to actually train a model on AMD. And also, you're driving fully open source models in telecom. So tell us about that experience and why you felt like that was needed.
J
Jeremy Legg1:43:49
We felt like the industry was moving down a path where folks were only going to use closed source models to solve solutions. And while closed source models are an important ingredient in that, we didn't want to lose sight of the fact that open source models were going to be a very significant part of the solution. So we announced what we call OTEL 1.0. That's the open telco AI model a bit over a year ago or so now. And we had over 18 million downloads of that model since then. We're excited to announce today the launch of OTEL 2.0, which is a more advanced set of models that we have trained on AMD. And we're now making that available via open source across the broader industry. And it, I think, is going to give folks an opportunity to see that not only do you need to think about leveraging models, but you also need to think about training those models and how they work across your enterprise.
D
Dan McNamara1:44:44
So it's just great. We certainly appreciate the partnership. And, you know, we really feel like this is germane to the conversation. It's exactly what we've been talking about today. And I just want to thank you again for joining us and the partnership. Thank you.
J
Jeremy Legg1:44:56
You're a big part of it. Thank you.
D
Dan McNamara1:45:02
Okay, so let me close with one last point. We've established ourselves as a leader in enterprise computing by listening to our customers and delivering the right solutions for their workloads. Today, we're extending that leadership to help customers accelerate their AI journey. This is our vision for powering AI everywhere. And to expand a little bit beyond the data center into physical and client AI, it's my pleasure to introduce Mr. Jack Huynh, who leads our computing and graphics group. Thank you.
J
Jack Huynh1:45:45
Thank you, Dan. It's great to see so many friends, partners, and developers here in San Francisco. Today, we have seen how AMD is building the compute foundation for AI in a data center. But the full AI transformation requires two more ingredients. Intelligence that becomes personal, and intelligence that can understand and act in a physical world. Those are the two frontiers I want to explore with you here today. Let's begin with personal AI. Enterprise AI is already reshaping the data center, but agentic AI will extend far beyond it. Agents are the next great productivity multiplier. The jump from software that responds to software that completes work. A personal agent can double what one person can do. A team of agents can multiply that by 10. And as agents gain more capabilities, the opportunity can grow exponentially. But that promise has a compute consequence. Unlike a single AI query, as agent reasons, calls tools, and coordinates to other agents, and run nonstop, agents become more capable and pervasive. So will token consumption, meaning that demand will require every layer of the computing ecosystem to scale together. Data centers remain the foundation for frontier AI in the most demanding workloads. But edge-inclined systems can extend that capacity, putting more intelligence closer to where data is created and where the work gets done. The PC is already one of the most powerful compute engines ever placed in the hands of an individual. Across billions of devices, enormous computing capability is already deployed every single day, yet its potential remains largely untapped. Putting that capacity to work changed the economics of computing. More workloads can run on systems that are already in place, reducing overall infrastructure costs and allowing data centers to focus on the largest, most demanding task. And local computing brings inherent advantages. Your data stays with you, your applications remain responsive, and your system keeps working, even when your network does not. And thanks to a new generation of smaller, more capable AI models, we're beginning to unlock that potential. And the unlock is happening faster than anyone expected. Last August, GPT OSS needed 120 billion parameters to score 80 on GPQA. Just seven months later, Qwen 3.5 scored even higher, with only 9 billion. Better performance, with 13 times fewer parameters. This is not an incremental improvement. It's a dramatic shift in the compute required to deliver advanced intelligence. And this is a trend, not an outlier. Smaller, open models are closing the frontier gap at extraordinary speed. Qwen 3.6, 27 billion parameters now outperforms leading frontier models from the previous generation on graduate-level reasoning, at a fraction of the size and cost. That changes what is possible. Workloads once reserved for the cloud can now run on powerful compute engines designed for personal AI compute. Let's map these model performance breakthroughs to any platforms. Ryzen AI 400 supports models up to 24 billion parameters, more than enough for Qwen 3.5, 9 billion, with substantial headroom. Ryzen AI Max extends that capacity to models as large as 200 billion parameters, running natively on a personal system. This has not reduced AI, compressed to fit on a PC. It is a new class of personal computing built to run powerful models where the work actually happens. And this is exactly what led us to build Ryzen AI Halo. We wanted to give developers a platform powerful enough to run these models locally, yet simple enough to use every day. With 120 gigabytes of unified memory and support for models up to 200 billion parameters, developers can build, test, and iterate locally directly on their desk. Ryzen AI Halo embodies the future of personal AI for everyone, a future that shortens the distance between an idea and a working model. For developers, that means a dramatically more efficient and optimized path from experimentation to deployment. And as Vamsi showed earlier, rocm.ai is making hardware programming radically more accessible. Developers can understand, modify, and optimize code for ROCm, even without deep ROCm expertise. And better yet, the same open software foundation extends across all of AMD's hardware. Developers can use the frameworks and tools they already know, start locally on Ryzen AI Halo, and scale across the full AMD compute platform without rebuilding their work. And we're not stopping there. Today, we're expanding our partnership with Hugging Face to bring the open AI ecosystem directly on Ryzen AI Halo. Powerful compute only creates value when developers can quickly access the right models, optimize them for the platform, and turn them into real applications. Together, AMD and Hugging Face will deliver Halo optimized performance for open models and agentic workflows. How cool is that? And along with co-engineered libraries and toolkits designed to help developers spend less time configuring infrastructure and more time building. And later this year, every Ryzen AI Halo box will include a full year of Hugging Face Pro, putting the power of partnership in developers' hands from day one. Freedom and more power to the user. This is how we accelerate innovation in the AI era. And we're already pushing that vision further with Gorgon Halo. And seeing is believing, look at how small and beautiful this box is. We've increased unified memory from 128 to 192 gigabytes. The largest unified memory pool in this class. And increased model support from 200 billion to 300 billion parameters. Thank you, Laura. And this is no longer limited to a single platform form factor. Together, with the world's leading OEMs, AMD now powers the broadest portfolio of personal AI systems. From laptops and workstations to mini PCs and developer platforms. That means developers, enterprises, and creators can choose the right system for the way they want to build and work. Personal AI is not a concept. It is a category and is available right now on AMD. We now have all the pieces. Models are efficient enough. Hardware is powerful enough. Software stack is open and ready. And open harnesses like Lemonade Server can route workloads to the right model and right compute. The technology is here. The next challenge is taking it from one developer's desk to thousands of users across an enterprise. Cisco sees and believes in that future the same way we do. To take us deeper into how we turn that vision into enterprise scale reality, please join me in welcoming to the stage Jeetu Patel, President and Chief Product Officer at Cisco.
J
Jeetu Patel1:54:19
Jeetu, how are you doing? How are you? You look great. Thank you for joining me here today. Congratulations. Thank you, thank you.
J
Jack Huynh1:54:26
Now Jeetu, you and I talk how we're at an inflection point. And for the past few years, industry has spent time just building intelligence. The next phase is how we deploy this technology at scale enterprises. You spent so much of your time talking to CIOs, technology leaders. What are they telling you? What's changing?
J
Jeetu Patel1:54:44
Well, firstly, it's a really exciting time to be alive in tech. And if you just take a step back right now and see what's happening, there's actually a fundamental shift that's happening in the patterns of inferencing. So if you think about the patterns of inferencing that existed during the chatbot era that was human-led, it was very, very spiky. You ask a question, you get an answer. What you're starting to see with agents is they're working 7x24. They're very consumptive on bandwidth and on infrastructure. So you're starting to see a much more persistent pattern of kind of demand signal for infrastructure. You know, we usually joke around internally, humans click, but agents swarm. And so as you start to see this, what you're going to have is a tremendous amount of and there's two big concerns that are also there, which is one is every customer is concerned about security and every customer is concerned about cost. So what you're starting to see happen is inferencing is not just going to be limited to the data center. Inferencing is going to be distributed everywhere and there's a whole new class of computing that's emerging with this desk-side computing that you just announced where there's going to be inferencing that you're going to have a laptop or a desktop where humans work, and they might have a desk-side computer where agents are kind of going out and running their jobs on behalf of humans. So that's what we're seeing, essentially.
J
Jack Huynh1:56:05
We're fully aligned Jeetu and that's exactly the shift we're seeing. If AI is moving closer to employees, compute has to move closer as well. The PC's knowledge as a productivity device becomes an intelligence node in the enterprise but deploying intelligence everywhere creates a new trend for CIOs that you and I talk about so much. How do you operate thousands of these systems with the security, governance and control enterprises require? Jeetu, how is AMD and Cisco solving this and what are we building together?
J
Jeetu Patel1:56:36
Yeah, so this is an exciting time because what you're starting to see is there's a few things that need to really be done. You can't just go out and deploy your desk-side computers and hope that everything works out in the enterprise. You need to make sure that it's governed, it's managed and so there's a few things that are really important. Number one, you have to have an appropriate amount of network bandwidth to satiate the needs of the agents that you are going to run locally. So this notion of network infrastructure that can get modernized to keep up with the needs that you're going to go out and generate is going to be very important. That's number one. Number two, like I said, every customer is really worried about token costs and so we need to make sure that we can actually contain the costs of an agent's consumption of tokens when it actually starts to go awry. Up and down when needed. Exactly. Number three then is how do we monitor agent behavior for safety and security so that if it does start to do things that we don't want it to do, we can in runtime provide enforcement guardrails. And then fourthly, it's just the overall safety and security that you need to have from an apparatus perspective to get this entire thing humming in the way that you want it to hum.
J
Jack Huynh1:57:44
Yes, it's great how you and I think very alike. So I mean the opportunity is not just bringing every device as you and I talk. How do we enable enterprise to deploy, manage the skill with confidence? Why don't we take a look at what we've been building the past few weeks and months.
J
Jeetu Patel1:58:00
Yeah, it's exciting. So basically what we've done is there's a full stack that you can start to think about. So what you folks have done with Halo is you've got an isolated secure agent sandbox. You've got intelligent routing. You've got an MCP kind of core set of integrations. And then what we've done above that is we've made sure that you have the right level of security policy enforcement. We've also got the right level of observability so that it's an end-to-end resilient infrastructure stack. So it's observability of how is your apparatus and the infrastructure working. Is the agent behavior being done in the right way? Can you observe the agent behavior? And are the tokenomics within the guidelines that you want them to be? And then what we have is a single unified management plane, a control plane that can look at every single dimension that you have and be managed within one environment. And so that's kind of the full stack, and we are really, really excited about the partnership.
J
Jack Huynh1:59:03
Us too, our teams love working with your team Jeetu. How is this all managed? How does the CAO deploy this at scale?
J
Jeetu Patel1:59:09
Yeah, so let's take a look at what this looks like, right? Because what we would have is this is Cisco Cloud Control, which is the overall management plane. So if you happen to have Halo devices on every desktop, which it might actually get to, what you want to make sure that you do is you are able to go out and manage your entire estate of not just your Halo fleet, but also what you have running on the data centers, as well as what you have running in the cloud. So you can provide full visibility for that entirety of the estate. And then once you've got that, what we also have is this notion of how do we go out and monitor tokenomics? And so is the agent behaving the way that it needs to behave, or is the agent actually going out and being overly consumptive? At which point, you should be able to isolate any individual Halo device so that you can make sure that you can quarantine it.
J
Jack Huynh2:00:05
No, I love this. Then we can actually see the return on invested capital. It could pay for itself in three or six months even quicker.
J
Jeetu Patel2:00:10
Exactly. So what you see over here basically is you can tell how every single agent is performing and whether or not they're consuming way too much kind of token costs. And if they are, how do you actually contain the costs as you're going through it?
J
Jack Huynh2:00:23
No, I love this Jeetu. Now, everyone else is wondering when can they buy this? When is this open for everyone?
J
Jeetu Patel2:00:30
Well, the good news is you can buy the Halo already today. Yes. And then what we're doing is we're making sure that we provide this entire management apparatus that's currently in early availability for a select set of customers. But we'll have it in general availability in the US in early fall. So we're excited to make sure that we can provide not just the compute, but all of the apparatus that's needed to go out and manage this so that you can have inference running in the cloud, you can have inference running in your data center privately, or you can have inference running on your desk side.
J
Jack Huynh2:01:05
Jeetu, we're so grateful to be part of you and Cisco. We're so excited. Thank you so much. Congratulations. Take care, Jeetu.
J
Jeetu Patel2:01:10
Thank you.
J
Jack Huynh2:01:15
What you just saw brings a vision of scaling AI seamlessly and securely for personal devices all the way to the cloud. But in the age of AI, intelligence will be pervasive and it will be everywhere. The next frontier is where AI leaves the screen, enters the world and acts. That frontier is physical AI, the ultimate application of agentic frameworks. In the digital world, an agent plans and completes a task. In the physical world, it must also sense motion, navigate uncertainty, and protect people. Here, intelligence does not just produce an answer, it produces an action. When answers become actions, latency becomes safety and reliability becomes trust. That is why the promise of physical AI is so incredibly powerful. It is not about replacing human potential. It is about amplifying it, giving surgeons greater precision, keeping people out of harm's way, making farms more productive, supply chains more resilient, and factories more efficient. The best technology does not make people smaller. It gives them much greater reach. For more than 20 years, AMD has helped power the core capabilities that robotics depend on, from sensing and functional safety to deterministic control and precision motion. That foundation is already trusted by millions of companies to find the future of robotics. And now, that foundation is coming alive in extraordinary ways. These are not machines following a script. They are beginning to perceive the world, reason through uncertainty, and act with purpose. The promise of physical AI, already in motion. The next era of robotics is not programmed. It is autonomous. While traditional robots execute instructions, autonomous robots are given an objective. And from there, they determine how to achieve it. And they do this all at once, and all in real time. The robot must orchestrate intelligence, motion, and safety continuously. And every part must respond at the speed of the world around it. That demands an entirely new kind of robotics brain. And to solve that, we rethought its architecture from the ground up. Today, I'm so excited to introduce the AMD Kria AI System On Module, powered by Ryzen AI Embedded X100. It brings CPU, GPU, NPU, and unified memory together in one compact, open standard architecture. This allows robots to process perception, AI reasoning, and do it all simultaneously and in real time. The architecture is different, but the results are decisive. In third party testing against NVIDIA Jetson Thor, Kria delivers 3.4 times better real-time results, 2.3 times more concurrent agents, and 1.6 times more CPU capacity. That means faster reactions, and greater headroom for the robot to take on more complex work. And all this without compromising control or safety. But to find the next era of robotics takes more than a powerful brain. Developers need a complete path from an idea to a machine operating in the real world. Today, we are introducing AMD's Kria AI Robotics Developer Platform. The industry's first turnkey open, fully integrated robotic system. Thank you, Laura, so much. At the center is the Kria AI System on module, paired with a robotics carrier card designed for sensing and connectivity. Built on ROCm and ROS2, this platform gives developers everything they need to move from concept to prototype in days. And the same module carries forward into production, with no migration or redesign needed at the finish line. You can build faster and move from imagination to deployment with absolute confidence. This is the breadth only AMD can bring to autonomous robotics from the first signal to the final motion. AMD Kria for the brain, Versal for the spine, Zynq for the joints, and Spartan for the sensors. Each layer purpose built for its role, working together in one intelligent system, connecting perception and high-level reasoning to real-time control. The future will not be defined by one model, one machine, or one company. It will be built in the open, across silicon, software, systems, and an ecosystem moving forward together. From the cloud to the PC and now into the physical world, AMD continues to build a compute foundation for intelligence everywhere. We'll laser focused on the future so our solution can augment and expand what people are capable of achieving. The next era of AI will not just unfold behind the screen. It will unfold in the world around us. And now, to bring it all together and take us to what comes next, please welcome back Lisa to the stage.
Hi, Lisa.
L
Lisa Su2:07:20
Alright. Thank you, Jack. And it's been a big day. We've covered a lot, from our data center products, to enterprise AI, to personal and physical AI, and the software that ties it all together. This is the strongest product portfolio in our history. And we believe it's the strongest portfolio in the industry. But it's really just the beginning, because what our customers count on is not just today. They really count on us looking ahead and pushing the bleeding edge of technology. So I want to spend the last few minutes giving you a preview of what we're working on. Starting with CPUs. In 2028, we're going to introduce Florence. Florence brings the next Gen Zen 7 cores. It's leading-edge process technology. It's a new set of AI compute extensions to really ensure that we have all of the AI capability. And it supports the latest memory technologies. And we're not stopping there. We're already deep in development of Ravenna, our eighth-generation EPYC family built on Zen 8. And that family is already well under development for 2030. Now, to give you a little bit of a flavor, Florence is again going to extend our biggest advantage with EPYC. It's all about product breath and having the right foundation and also the right family for each workload. So you have Florence, you have Ferrara, you have Fidenza. And it's really built on a full family of CPUs that is optimized for these different workloads. That's giving the customers the right CPU for every workload. And that is our commitment with EPYC. Now, moving to Instinct. We are continuing our cadence of delivering a new generation every single year. MI500 brings next-generation HBM. It brings a larger scale-up domain because scaling is everything. And it introduces new copper and optical interconnects. And with MI600, it's powered by our CDNA Next architecture, and it's already deep in development for 2028. Now, to give you an idea of what we see from a performance standpoint, every generation of Instinct has delivered a major big step up in performance. And with MI455, we've talked a lot about it, delivering 35 times more inference throughput compared to MI355. But we're really, really excited about MI500. What we see is another opportunity to take a major step up and actually bend the curve again. So MI500 will deliver the largest generational leap in the industry of Instinct, putting us on track to deliver more than 2,000 times higher inference throughput in just four years. Now, what I can tell you is we are working with a number of our customers already on this design point. MI500 is super exciting, and the feedback we're getting from customers is just fantastic. Now, put the CPU, GPU, and the networking roadmaps together, and you can now see a complete cadence of our rack scale systems. So you should expect from AMD a new Helios system every year delivering more performance, more capacity so that we can run larger models, more agents at better total cost of ownership. And that gives our customers the complete, predictable roadmap to plan and scale their infrastructure for years to come. So lots of exciting things in store over the next few years. So with that, let me wrap this up for this morning. Today, I hope you saw a little bit about how we're really looking at that compute from every angle. So every aspect of AI compute, whether you're talking about Helios or MI455 or Venice for the world's largest AI systems, or you're talking about MI430 and MI350 for Sovereign and Enterprise AI. You saw the passion in Vamsi with ROCm AI. It's so good to see how much has come through that platform in terms of making developers make it much, much easier for developers to access the AMD platform. And then Jack talked about Gorgon Halo and our new Kria AI platforms really bringing together the leadership of the end-to-end AI story, including PCs and physical AI. So lots of new information, but I want to say a very, very special thank you to all of our partners who joined us today because it is really through those partnerships that we're able to do the most amazing things together. So if I just leave you with a final thought, when I think about where AI is today, the biggest change that we see is we're no longer talking about what might be possible. We're actually seeing how AI can have real and significant impact across every industry and every part of our lives, our personal lives. And I have to say that I've spent my entire career in tech believing that high-performance computing can make the world an incredibly better place. And I have never, ever believed that more than I do today. At AMD, what we're focused on is building the technology, the roadmaps, and the partnerships, and there's never been a more exciting moment than today for our 30,000 plus engineers. This is really the next phase of AI. This is where we make AI, give the AI the opportunity with all the tools and all the capabilities to really bring meaningful real-world impact. And I can tell you I could not be more excited to build all of that together with you, our ecosystem. Thank you so much for joining us today.