Hi everybody, welcome to this special cube power panel with supporting advancing AI. My name is Dave Volante and the topic today is the architecture of uniting CPU and GPU for enterprise AI. I have a little setup and then we're going to introduce the guests. So let me tell you about the angle that we're going to take today. AI infrastructure is very often discussed as though it were just exclusively a decision around GPUs. But in production, you have CPUs, accelerators, memory, networking, storage, cooling, software, and rack integration. All of these have to work as one system in balance. So this becomes even more important as trusted enterprise systems of record converge with AI-driven and agentic applications, the probabilistic nature of these. Today we're going to examine how AMD and Super Micro are combining roadmaps and engineering to translate Venice into the H15 family and broader rack-scale infrastructure, and how that impacts performance, efficiency, and time to value. Venice is the code name for AMD's sixth-generation EPYC processor, the workhorse for heavy-duty cloud, data center, and agentic AI workloads. The H designation is reserved for the AMD-powered Super Micro data center servers. With me to discuss that is Derek Dicker, corporate vice president for enterprise business group at AMD. Derek, thank you for hosting us here at your EVC.
Thank you for coming and for having us. Looking forward to the conversation.
And Vic Malyala, who's the chief business officer at Super Micro. Vic, good to see you.
Okay, so let's start with the partnership. What is the unique role that each of your companies plays in this partnership? How would you describe that? Why don't we start with you, Derek?
First off, the only way that somebody can build technology that's required today to address all the workloads that you described is to have incredibly close partnerships between the companies in the chain delivering products to the end. With respect to Super Micro, an incredible relationship since the beginning and inception of EPYC. I think Vic, we've had five generations of product launched on time at launch together.
Absolutely. The only way that happens is by not just having a relationship between the two of us, but all the way down to the engineers thinking through what's coming next, aligning our requirements, designing products together, co-innovating together. It's been a fantastic relationship so far.
How about from your perspective, Vic? I mean, I think EPYC goes back to probably the 2017-2018 timeframe, but Genoa was probably 2022 timeframe, and really been on a clear cadence since then. What's your perspective on the partnership?
I think the most important thing is it's two engineering-centric companies that are collaborating and co-developing platforms and solutions together. AMD is focusing on the silicon. Our role as Super Micro is to see how best to design a platform or a solution that takes advantage of every one of those features, making it available for customers. So basically making it real at a server level for the customers to use. It goes very deep from the very beginning. From the first generation when AMD introduced EPYC, it was all about maximum number of cores, maximum memory you could pack in a system, and it had 128 lanes of PCI Express, which was unheard of. It started with that and we never looked back. Every generation there is tremendous improvement, generation over generation, features and functionality get added at the silicon level, and to support that from Super Micro at a system level, now at a rack scale level or a data center level, we are working towards getting the right platforms out together.
I want to get into the deep engineering that you guys do, but let's start with the demand. Everybody's excited about AI. What can you share, Derek, in terms of what you're seeing for demand? And then Vic, maybe you can take it closer to the end customer, because we really want to understand what's driving this next generation enterprise infrastructure. What is demand? I mean, obviously demand is through the roof, but how would you describe it?
I would attack it from a couple of different perspectives. First, every single enterprise customer is being asked by their management and their board how they are going to address AI and utilize it in their infrastructure and day-to-day activities. That spurs two different sets of activities. First, inside data centers, most are bursting at the seams with legacy technology, three to four years back server infrastructure. To put AI into that infrastructure, they have to create room. So there's quite a bit of activity we see before the agentic world, wrapped around modernization and data center modernization. That's one level of demand. When you modernize, you give yourself the opportunity to compress the footprint, reduce power, but that's in the pursuit of freeing up room for AI infrastructure. The next piece is how enterprises are deploying enterprise AI and what they are building out for that, driving the next layer of demand. The only way to service those is by having an incredibly tight partnership with companies like Super Micro to build the next generation of infrastructure fueling a large amount of CPU demand as well as GPU demand.
Your perspective on demand. Common knowledge, I think you talk to anyone in the industry, whether it's CPU or GPU or anything that connects all these things, everyone is saying they're sold out for the next several quarters. Take a look at the demand, it's insatiable. In the case of AI, it all started with training clusters, massive training clusters being built. But in the last 8 to 10 months, we see the demand coming up a lot more with inferencing, specifically agentic AI. The RAG models were there, GenAI models came up, and now we're talking about all these agentic AI workloads that are completely transforming the requirement for all customers. Traditional workloads are great, but in addition, people are trying to adopt AI into their workflows to improve efficiency and productivity for their companies. That is driving the entire decision-making process. The idea is to take something they feel comfortable with and use it to enhance their workloads to get better business value. That is driving the demand, along with efficiencies. As people start deploying these, they are deployed at scale, not just one or two, but clusters of them. Now people need to figure out how to optimize their cost function and total cost of ownership, bringing out the efficiencies. So a lot of these things are happening and that's driving the demand.
When you think back to the cloud era, now the new legacy, you had public cloud workloads and on-prem workloads. That was the state-of-the-art hybrid. They really weren't working together, maybe some bursting. It seems that's changing. You've got a lot of activity in the public cloud, we had the Neo Cloud panel which people should check out. And then you've got organizations saying, 'Hey, we want to bring the AI to the data that's on prem. We're not going to move that all into the cloud.' So you're seeing this true hybrid situation. So my question is, where do they put all this stuff? They're bursting at the seams. They've got to look to companies like yours to consolidate the existing infrastructure. What are you seeing in terms of the ability to consolidate servers down from where they are today to a better ratio?
Well, maybe just one data point we can offer is in the partnership between our companies. As we've gone out and done case study-related work within customers, we've built up a consolidation playbook. Within that, what we reflect is based on using a Turin, a fifth-generation EPYC processor, as the baseline inside a two-socket server today. You can take three to four-year-old technology on legacy infrastructure and replace eight of those legacy servers with a single Turin-based server. That equates to a substantially smaller footprint, 70% less power, 71% improved total cost of ownership, freeing up space for AI workloads. That's one baseline fact. We're seeing different levels of consolidation depending on what customers are looking for, but it's very real and happening on a repeated basis, driving some of the demand you talked about. The other piece is that on the AI side, we're experiencing more of a desire for people to put historically token-related workloads out into the cloud, looking at how they run on open-source models on top of x86 servers. That's the next level we expect to see building out.
That's a good point. Customers, Vic, they want to be their own token generators, don't they?
Absolutely. I was trying to see where to start. But what is happening is it's expensive. Anything with AI is an additional expense on top of whatever infrastructure spend people already have. Initially, the easiest way is to go to a public cloud service provider and take these tokens and start using them. Everyone is happy until they see the bill. Then all of a sudden, what do they do? There are definite benefits, so people are trying to figure out how to bring it on prem. That means they need to retrofit data centers. So what happens? If you take a look at the amount of power available per rack, that is being revised as they want to bring AI infrastructure in. Once that is taken care of, they can bring very high-density compute right next to it. So we are seeing that we can make the best use of available power and space by bringing both AI infrastructure, which is GPUs, as well as compute. We see people are more receptive to the idea of using liquid cooling. When they bring liquid cooling, they can increase the density like anything. We're talking about 40,000 plus cores per rack for CPU, not GPUs. When was the last time you heard about that kind of density at a rack level for CPU?
Right. Right. Okay, let's get into it. Let's explain the whole CPU and GPU and why it's an 'and' story and not an 'or' story. The business news will cover it. Lisa, I think, was pretty explicit on the last earnings calls. She said it used to be eight to one, it's now four to one, we think it's soon going to be two to one, maybe even one to one, maybe even flip. What's driving that? When you let agents loose in the organization, they're accessing different tools, different documents, spending time reasoning, and that requires a lot of compute. But help us understand specifically what's going on under the hood.
Yeah, I'll take a first crack at it. It would be great for Vic to pile on. If you look at the work being done in agentic workflows themselves, it's a lot of leveraging existing applications that are in place, but it's orchestration occurring across them. So you have preservation of the applications we use and love today, getting connected, but there needs to be orchestration across all of those, tool calls, and a bunch of work on the data itself. All of these things fit really well to CPU architecture, particularly x86 architecture. Those combination of activities is driving the demand we see on x86 CPUs in the agentic era. It's a non-trivial explosion in consumption and use we're seeing coming over the horizon.
I think to add to that, GPUs are the most expensive part and we want to keep the GPUs busy doing the actual work. Anything that can be done to support the GPUs and make sure they are fed properly is where the compute comes in. With agentic workloads, a lot of these things can be done on a CPU far better than on an expensive GPU. That's number one. Second, so many small batches are coming in with these agents running. With more processor cores packed in, like 256 cores per CPU, with a lot of memory and IO bandwidth, a lot of stuff can be done. I was reading about the total latency impact of running these things in CPU architecture. You're talking about 50 to 80% of the latencies of these agents can be addressed from a CPU point of view. So not just for the premise that CPUs were used in the past, now we're talking about a lot of agentic AI workloads, whether it's database lookups, tool callings, all the orchestration. All this is happening on the CPU. CPU counts are increasing considerably in inference-specific workloads for enterprise customers. Coupled with properly balancing with the GPU architecture, that is what is going to provide the right value for customers.
Am I correct? Am I correct? It's not just a CPU problem. You've got networking, storage, IO. It's a balanced system to keep that GPU busy and get the most efficient use out of the system.
I think that's the reason the CPU to GPU balance comes into picture. You take a look at the latest generation processor. Now you're talking about PCIe Gen 6, 16 channels of memory with 8,000 plus MHz memory. So now you're talking about how much memory bandwidth you can get, what kind of IO you can get because of faster PCI Express, and how many GPUs to a specific CPU we are mapping. If that memory and storage is not sufficient, what other storage you can add to create a more balanced compute, GPU, storage, and network with very fast interconnect, like InfiniBand for example.
I would add to that the beauty of the relationship between the two companies is that we can sit down and look at mutual experiences within customers and translate that to what the requirements are at a system level. Then between the two of us, leveraging their system design expertise as well as our ability to design silicon in a flexible manner, we can address these workloads. We now have GPU-centric technology, CPU-centric technology, networking technology, and rack-scale architectures that we are developing jointly together.
It leverages the knowledge of the end customers, maps to what the art of the possible is, but also leaves room for innovation. That's the beauty of the relationship between the companies. In my discussions with AMD folks, listening to Lisa on various calls, Vic was talking about the preponderance of training, then we saw reasoning sort of save the AI day for a while, and now we're into agentic. Did you see that coming? My point is AMD always has said inference is going to be the larger opportunity. You would agree with that, but you had to make a bet there and things changed so fast. Did you see that coming? How did you see around that corner? And what's happening next is really what I want to know.
Yeah, that's totally fair. First off, the beauty of the company's evolution, if you look at the overall share of gain and the place AMD plays in the market, the company has done an amazing job working with all the leaders in the industry. You see the announcements out there with OpenAI, Meta, and a number of enterprise customers. Those discussions afford us the ability to see through their eyes where the biggest business opportunities are and where the potential business outcomes are. Couple that with what Super Micro sees in their engagements and discussions, and it starts to put together a picture of what the future may be. Go back in time and trace when we started talking about agentic. It was back last year as we were having our Financial Analyst Day. We were talking about the impact on the CPU market, what our CPU technology could look like in the future. All those pieces together mapped really well on top of agentic workflows. So, did we see around corners perfectly? No, of course not. But did we have footprints in the sand that we were following as we developed solutions? Yes. And was that coupled by work we were doing with Super Micro to understand the complexity of the market out in time? Absolutely.
So what were the signals you were seeing at the time and what are you seeing now?
One of the things that working with AMD, as they define the processor SKUs, in a given thermal design you can pack fewer cores with higher frequency or more cores with less frequency, how memory mapping should happen, how MCM designs fit into workloads. That's something we always look at as we design our platforms. We would not know whether people would be focusing on these systems for using accelerators, storage, network, or high-performance compute. But what we do know is by giving those options in the most optimized manner and being able to build these systems not just at a system level but conceptualize at a rack level, how we bring a cluster and deliver it, that's the main emphasis. It becomes an easier transition for us as we start looking into the AI-centric workloads. What kind of GPUs you are using, how many GPUs, what it takes to keep the GPU busy, and what new workloads you want to run. Traditional workloads are not going anywhere. They are still running, but the same systems can be tuned a little further in terms of system configuration to run all these new agentic and inference workloads. That's where the difference is. No one has a crystal ball. If that were the case, people would have talked about agentic-optimized CPUs five years ago. Why is it only being heard now? Because things are changing. We just need to be ready with the right tools to deliver the solution at the right time. I would add one other thing. The magic of a relationship like this is where we can communicate together on things we see coming, but be able to pivot, execute, and make modifications. The beauty is that Super Micro has a great set of technologies that allows them to build different types of systems within our silicon architecture. Our engineers have thought forward to what we may want to modulate on. That's where the chiplet architecture comes in really well at a CPU level. We can adjust solutions out in time.
I mean think about enterprise workloads, right? People want security and sovereignty. CPU architecture already has trusted execution in it. That gives an enterprise workload the security aspect. It's not an afterthought, it's there. It becomes the right tool available for this kind of core of the architecture from the beginning. But to your point about seeing out into the future, there are certain things you can look at, follow themes, build in capability. You're not exactly sure how it's going to be used, but you provision it sufficiently, work with system providers to make sure it can be taken advantage of, and deliver it to market when it's ready. And like you said, the chiplet architecture is a benefit; it gives you the ability to flex.
What is materially new about Venice? And I'm curious, is it specifically designed for this CPU and GPU agentic world, or was it just kind of luck of the draw? And how does that translate into H15 deployability for customers? So that's where I want to go.
I'll talk a little bit about the silicon side. Venice itself was created with a number of different applications in mind, but really it was the next generation, a sixth generation in the EPYC family. By the way, the first five were delivered on time with quality, with Super Micro standing there right with us at the launch. That only happens through incredible collaboration. Venice itself, the way I would describe it, is the highest performant solution that will be on the market in the x86 side. Generation over generation, we'll see a 70% performance increase in the product. We're featuring PCIe Gen 6, time to market with an ecosystem we're building around it again in partnership with Super Micro. A lot of the use cases originally envisioned were how to marry flavors of this product, from 256 cores down to the base of the stack from a core count perspective, but for GPU-centric applications, how do we design a system leveraging Venice in such a way that when you build something like Helios, which we haven't talked about yet, you get an optimized solution out of it? So a lot of work was done envisioning what the application sets would be.
I think you hit the nail right on the head. When we started looking at Venice, one of the things that we started looking at is what are the kind of environments that it is going to be deployed in besides the traditional way. For example, we have standard 1U, 2U pizza boxes with a single or dual socket that's been there like our cloud DC hyper these kind of platforms. Then we extended that into a dense compute like FlexTwin, which is liquid cooled platforms and a 2U4 four-node kind of thing. And then we have integrated switch into that for those who wanted to have a dense compute but in a smaller form factor like a SuperBlade. And we went, the kind of product line extends one after another. But then we also have seen people are using either air or liquid because things are being now dealt at a rack scale, especially in a larger deployment arena. So what we did was, what is the power delivery going to be in these kind of systems? Traditionally, you have like a 60 amp or 100 amp power drops coming in, and then you put PDUs and then you put all these systems together. That's all great, but now we started looking at what is that we can do with an OCP enhanced or OCP inspired designs where like OV3 rack with power delivery using the DC bus bar. We have kind of transformed the product line not just for the traditional ones but also to take advantage of these things. So things are being considered at a rack scale or a data center scale. What is that we need to do in a given operating environment and how do we basically pack as much punch in a smaller footprint? You kind of figure out where it would be needed and how do we get the most optimal solution there. So that's the thought process and we've been working with AMD on how best to put these things together so that it will be relevant for customers.