About Vik Malyala
Vik Malyala, Senior Vice President, Managing Director and President of the EMEA Division at Supermicro, has been discussing the company's AI infrastructure portfolio and the importance of balancing CPU and GPU architectures for enterprise AI workloads. In interviews and presentations, Malyala emphasized that AI infrastructure should not be viewed solely as a decision around GPUs, stating that CPUs, accelerators, memory, networking, storage, cooling, software, and rack integration must work as one system. He noted that with the rise of agentic AI, workloads on the CPU can impact 50 to 90 percent of latency, and called it "a criminal waste not to have the best computation" for orchestration. Malyala also highlighted the rapid pace of technology refresh, which he said is now less than a year, while typical data center projects take 12 to 24 months, creating a fundamental challenge where infrastructure can become outdated before deployment.
At COMPUTEX 2026, Malyala showcased Supermicro's next-generation platforms powered by NVIDIA, AMD, Intel, and ARM processors, including the Vera Rubin, HGX B300, and AMD Helios systems. He described the AMD Helios as a "beast" and a "massive 7,000 lb machine." Malyala also discussed the company's data center building block strategy, which includes liquid cooling, rack-scale systems, and infrastructure management software. He referenced estimates of $1.5 trillion in IT infrastructure spending, calling it a "big number," and noted that customers are increasingly exploring on-premises deployment after seeing the costs of public cloud services. Malyala added that Supermicro is working with partners such as AMD to co-design systems that can be quickly taken from proof of concept into production.
Source: AI-verified profile updated from Vik Malyala's recent appearances.
Browse all interviews →
Transcript (59 segments)
D
Dave Volante0:05
Hi everybody, welcome to this special cube power panel with supporting advancing AI. My name is Dave Volante and the topic today is the architecture of uniting CPU and GPU for enterprise AI. I have a little setup and then we're going to introduce the guests. So let me tell you about the angle that we're going to take today. AI infrastructure is very often discussed as though it were just exclusively a decision around GPUs. But in production, you have CPUs, accelerators, memory, networking, storage, cooling, software, and rack integration. All of these have to work as one system in balance. So this becomes even more important as trusted enterprise systems of record converge with AI-driven and agentic applications, the probabilistic nature of these. Today we're going to examine how AMD and Super Micro are combining roadmaps and engineering to translate Venice into the H15 family and broader rack-scale infrastructure, and how that impacts performance, efficiency, and time to value. Venice is the code name for AMD's sixth-generation EPYC processor, the workhorse for heavy-duty cloud, data center, and agentic AI workloads. The H designation is reserved for the AMD-powered Super Micro data center servers. With me to discuss that is Derek Dicker, corporate vice president for enterprise business group at AMD. Derek, thank you for hosting us here at your EVC.
D
Derek Dicker1:32
Thank you for coming and for having us. Looking forward to the conversation.
D
Dave Volante1:35
And Vic Malyala, who's the chief business officer at Super Micro. Vic, good to see you.
V
Vik Malyala1:40
Same here. Thank you.
D
Dave Volante1:41
Okay, so let's start with the partnership. What is the unique role that each of your companies plays in this partnership? How would you describe that? Why don't we start with you, Derek?
D
Derek Dicker1:51
First off, the only way that somebody can build technology that's required today to address all the workloads that you described is to have incredibly close partnerships between the companies in the chain delivering products to the end. With respect to Super Micro, an incredible relationship since the beginning and inception of EPYC. I think Vic, we've had five generations of product launched on time at launch together.
V
Vik Malyala2:17
Absolutely. The only way that happens is by not just having a relationship between the two of us, but all the way down to the engineers thinking through what's coming next, aligning our requirements, designing products together, co-innovating together. It's been a fantastic relationship so far.
D
Dave Volante2:33
How about from your perspective, Vic? I mean, I think EPYC goes back to probably the 2017-2018 timeframe, but Genoa was probably 2022 timeframe, and really been on a clear cadence since then. What's your perspective on the partnership?
V
Vik Malyala2:47
I think the most important thing is it's two engineering-centric companies that are collaborating and co-developing platforms and solutions together. AMD is focusing on the silicon. Our role as Super Micro is to see how best to design a platform or a solution that takes advantage of every one of those features, making it available for customers. So basically making it real at a server level for the customers to use. It goes very deep from the very beginning. From the first generation when AMD introduced EPYC, it was all about maximum number of cores, maximum memory you could pack in a system, and it had 128 lanes of PCI Express, which was unheard of. It started with that and we never looked back. Every generation there is tremendous improvement, generation over generation, features and functionality get added at the silicon level, and to support that from Super Micro at a system level, now at a rack scale level or a data center level, we are working towards getting the right platforms out together.
D
Dave Volante3:55
I want to get into the deep engineering that you guys do, but let's start with the demand. Everybody's excited about AI. What can you share, Derek, in terms of what you're seeing for demand? And then Vic, maybe you can take it closer to the end customer, because we really want to understand what's driving this next generation enterprise infrastructure. What is demand? I mean, obviously demand is through the roof, but how would you describe it?
D
Derek Dicker4:20
I would attack it from a couple of different perspectives. First, every single enterprise customer is being asked by their management and their board how they are going to address AI and utilize it in their infrastructure and day-to-day activities. That spurs two different sets of activities. First, inside data centers, most are bursting at the seams with legacy technology, three to four years back server infrastructure. To put AI into that infrastructure, they have to create room. So there's quite a bit of activity we see before the agentic world, wrapped around modernization and data center modernization. That's one level of demand. When you modernize, you give yourself the opportunity to compress the footprint, reduce power, but that's in the pursuit of freeing up room for AI infrastructure. The next piece is how enterprises are deploying enterprise AI and what they are building out for that, driving the next layer of demand. The only way to service those is by having an incredibly tight partnership with companies like Super Micro to build the next generation of infrastructure fueling a large amount of CPU demand as well as GPU demand.
V
Vik Malyala6:18
Your perspective on demand. Common knowledge, I think you talk to anyone in the industry, whether it's CPU or GPU or anything that connects all these things, everyone is saying they're sold out for the next several quarters. Take a look at the demand, it's insatiable. In the case of AI, it all started with training clusters, massive training clusters being built. But in the last 8 to 10 months, we see the demand coming up a lot more with inferencing, specifically agentic AI. The RAG models were there, GenAI models came up, and now we're talking about all these agentic AI workloads that are completely transforming the requirement for all customers. Traditional workloads are great, but in addition, people are trying to adopt AI into their workflows to improve efficiency and productivity for their companies. That is driving the entire decision-making process. The idea is to take something they feel comfortable with and use it to enhance their workloads to get better business value. That is driving the demand, along with efficiencies. As people start deploying these, they are deployed at scale, not just one or two, but clusters of them. Now people need to figure out how to optimize their cost function and total cost of ownership, bringing out the efficiencies. So a lot of these things are happening and that's driving the demand.
D
Dave Volante8:10
When you think back to the cloud era, now the new legacy, you had public cloud workloads and on-prem workloads. That was the state-of-the-art hybrid. They really weren't working together, maybe some bursting. It seems that's changing. You've got a lot of activity in the public cloud, we had the Neo Cloud panel which people should check out. And then you've got organizations saying, 'Hey, we want to bring the AI to the data that's on prem. We're not going to move that all into the cloud.' So you're seeing this true hybrid situation. So my question is, where do they put all this stuff? They're bursting at the seams. They've got to look to companies like yours to consolidate the existing infrastructure. What are you seeing in terms of the ability to consolidate servers down from where they are today to a better ratio?
D
Derek Dicker9:03
Well, maybe just one data point we can offer is in the partnership between our companies. As we've gone out and done case study-related work within customers, we've built up a consolidation playbook. Within that, what we reflect is based on using a Turin, a fifth-generation EPYC processor, as the baseline inside a two-socket server today. You can take three to four-year-old technology on legacy infrastructure and replace eight of those legacy servers with a single Turin-based server. That equates to a substantially smaller footprint, 70% less power, 71% improved total cost of ownership, freeing up space for AI workloads. That's one baseline fact. We're seeing different levels of consolidation depending on what customers are looking for, but it's very real and happening on a repeated basis, driving some of the demand you talked about. The other piece is that on the AI side, we're experiencing more of a desire for people to put historically token-related workloads out into the cloud, looking at how they run on open-source models on top of x86 servers. That's the next level we expect to see building out.
D
Dave Volante10:30
That's a good point. Customers, Vic, they want to be their own token generators, don't they?
V
Vik Malyala10:33
Absolutely. I was trying to see where to start. But what is happening is it's expensive. Anything with AI is an additional expense on top of whatever infrastructure spend people already have. Initially, the easiest way is to go to a public cloud service provider and take these tokens and start using them. Everyone is happy until they see the bill. Then all of a sudden, what do they do? There are definite benefits, so people are trying to figure out how to bring it on prem. That means they need to retrofit data centers. So what happens? If you take a look at the amount of power available per rack, that is being revised as they want to bring AI infrastructure in. Once that is taken care of, they can bring very high-density compute right next to it. So we are seeing that we can make the best use of available power and space by bringing both AI infrastructure, which is GPUs, as well as compute. We see people are more receptive to the idea of using liquid cooling. When they bring liquid cooling, they can increase the density like anything. We're talking about 40,000 plus cores per rack for CPU, not GPUs. When was the last time you heard about that kind of density at a rack level for CPU?
D
Dave Volante12:06
Right. Right. Okay, let's get into it. Let's explain the whole CPU and GPU and why it's an 'and' story and not an 'or' story. The business news will cover it. Lisa, I think, was pretty explicit on the last earnings calls. She said it used to be eight to one, it's now four to one, we think it's soon going to be two to one, maybe even one to one, maybe even flip. What's driving that? When you let agents loose in the organization, they're accessing different tools, different documents, spending time reasoning, and that requires a lot of compute. But help us understand specifically what's going on under the hood.
D
Derek Dicker12:48
Yeah, I'll take a first crack at it. It would be great for Vic to pile on. If you look at the work being done in agentic workflows themselves, it's a lot of leveraging existing applications that are in place, but it's orchestration occurring across them. So you have preservation of the applications we use and love today, getting connected, but there needs to be orchestration across all of those, tool calls, and a bunch of work on the data itself. All of these things fit really well to CPU architecture, particularly x86 architecture. Those combination of activities is driving the demand we see on x86 CPUs in the agentic era. It's a non-trivial explosion in consumption and use we're seeing coming over the horizon.
V
Vik Malyala13:38
I think to add to that, GPUs are the most expensive part and we want to keep the GPUs busy doing the actual work. Anything that can be done to support the GPUs and make sure they are fed properly is where the compute comes in. With agentic workloads, a lot of these things can be done on a CPU far better than on an expensive GPU. That's number one. Second, so many small batches are coming in with these agents running. With more processor cores packed in, like 256 cores per CPU, with a lot of memory and IO bandwidth, a lot of stuff can be done. I was reading about the total latency impact of running these things in CPU architecture. You're talking about 50 to 80% of the latencies of these agents can be addressed from a CPU point of view. So not just for the premise that CPUs were used in the past, now we're talking about a lot of agentic AI workloads, whether it's database lookups, tool callings, all the orchestration. All this is happening on the CPU. CPU counts are increasing considerably in inference-specific workloads for enterprise customers. Coupled with properly balancing with the GPU architecture, that is what is going to provide the right value for customers.
D
Dave Volante15:06
Am I correct? Am I correct? It's not just a CPU problem. You've got networking, storage, IO. It's a balanced system to keep that GPU busy and get the most efficient use out of the system.
V
Vik Malyala15:18
I think that's the reason the CPU to GPU balance comes into picture. You take a look at the latest generation processor. Now you're talking about PCIe Gen 6, 16 channels of memory with 8,000 plus MHz memory. So now you're talking about how much memory bandwidth you can get, what kind of IO you can get because of faster PCI Express, and how many GPUs to a specific CPU we are mapping. If that memory and storage is not sufficient, what other storage you can add to create a more balanced compute, GPU, storage, and network with very fast interconnect, like InfiniBand for example.
D
Derek Dicker16:05
I would add to that the beauty of the relationship between the two companies is that we can sit down and look at mutual experiences within customers and translate that to what the requirements are at a system level. Then between the two of us, leveraging their system design expertise as well as our ability to design silicon in a flexible manner, we can address these workloads. We now have GPU-centric technology, CPU-centric technology, networking technology, and rack-scale architectures that we are developing jointly together.
D
Dave Volante16:36
It leverages the knowledge of the end customers, maps to what the art of the possible is, but also leaves room for innovation. That's the beauty of the relationship between the companies. In my discussions with AMD folks, listening to Lisa on various calls, Vic was talking about the preponderance of training, then we saw reasoning sort of save the AI day for a while, and now we're into agentic. Did you see that coming? My point is AMD always has said inference is going to be the larger opportunity. You would agree with that, but you had to make a bet there and things changed so fast. Did you see that coming? How did you see around that corner? And what's happening next is really what I want to know.
D
Derek Dicker17:23
Yeah, that's totally fair. First off, the beauty of the company's evolution, if you look at the overall share of gain and the place AMD plays in the market, the company has done an amazing job working with all the leaders in the industry. You see the announcements out there with OpenAI, Meta, and a number of enterprise customers. Those discussions afford us the ability to see through their eyes where the biggest business opportunities are and where the potential business outcomes are. Couple that with what Super Micro sees in their engagements and discussions, and it starts to put together a picture of what the future may be. Go back in time and trace when we started talking about agentic. It was back last year as we were having our Financial Analyst Day. We were talking about the impact on the CPU market, what our CPU technology could look like in the future. All those pieces together mapped really well on top of agentic workflows. So, did we see around corners perfectly? No, of course not. But did we have footprints in the sand that we were following as we developed solutions? Yes. And was that coupled by work we were doing with Super Micro to understand the complexity of the market out in time? Absolutely.
D
Dave Volante18:36
So what were the signals you were seeing at the time and what are you seeing now?
V
Vik Malyala18:40
One of the things that working with AMD, as they define the processor SKUs, in a given thermal design you can pack fewer cores with higher frequency or more cores with less frequency, how memory mapping should happen, how MCM designs fit into workloads. That's something we always look at as we design our platforms. We would not know whether people would be focusing on these systems for using accelerators, storage, network, or high-performance compute. But what we do know is by giving those options in the most optimized manner and being able to build these systems not just at a system level but conceptualize at a rack level, how we bring a cluster and deliver it, that's the main emphasis. It becomes an easier transition for us as we start looking into the AI-centric workloads. What kind of GPUs you are using, how many GPUs, what it takes to keep the GPU busy, and what new workloads you want to run. Traditional workloads are not going anywhere. They are still running, but the same systems can be tuned a little further in terms of system configuration to run all these new agentic and inference workloads. That's where the difference is. No one has a crystal ball. If that were the case, people would have talked about agentic-optimized CPUs five years ago. Why is it only being heard now? Because things are changing. We just need to be ready with the right tools to deliver the solution at the right time. I would add one other thing. The magic of a relationship like this is where we can communicate together on things we see coming, but be able to pivot, execute, and make modifications. The beauty is that Super Micro has a great set of technologies that allows them to build different types of systems within our silicon architecture. Our engineers have thought forward to what we may want to modulate on. That's where the chiplet architecture comes in really well at a CPU level. We can adjust solutions out in time.
D
Dave Volante21:01
I mean think about enterprise workloads, right? People want security and sovereignty. CPU architecture already has trusted execution in it. That gives an enterprise workload the security aspect. It's not an afterthought, it's there. It becomes the right tool available for this kind of core of the architecture from the beginning. But to your point about seeing out into the future, there are certain things you can look at, follow themes, build in capability. You're not exactly sure how it's going to be used, but you provision it sufficiently, work with system providers to make sure it can be taken advantage of, and deliver it to market when it's ready. And like you said, the chiplet architecture is a benefit; it gives you the ability to flex.
What is materially new about Venice? And I'm curious, is it specifically designed for this CPU and GPU agentic world, or was it just kind of luck of the draw? And how does that translate into H15 deployability for customers? So that's where I want to go.
D
Derek Dicker22:09
I'll talk a little bit about the silicon side. Venice itself was created with a number of different applications in mind, but really it was the next generation, a sixth generation in the EPYC family. By the way, the first five were delivered on time with quality, with Super Micro standing there right with us at the launch. That only happens through incredible collaboration. Venice itself, the way I would describe it, is the highest performant solution that will be on the market in the x86 side. Generation over generation, we'll see a 70% performance increase in the product. We're featuring PCIe Gen 6, time to market with an ecosystem we're building around it again in partnership with Super Micro. A lot of the use cases originally envisioned were how to marry flavors of this product, from 256 cores down to the base of the stack from a core count perspective, but for GPU-centric applications, how do we design a system leveraging Venice in such a way that when you build something like Helios, which we haven't talked about yet, you get an optimized solution out of it? So a lot of work was done envisioning what the application sets would be.
V
Vik Malyala23:29
I think you hit the nail right on the head. When we started looking at Venice, one of the things that we started looking at is what are the kind of environments that it is going to be deployed in besides the traditional way. For example, we have standard 1U, 2U pizza boxes with a single or dual socket that's been there like our cloud DC hyper these kind of platforms. Then we extended that into a dense compute like FlexTwin, which is liquid cooled platforms and a 2U4 four-node kind of thing. And then we have integrated switch into that for those who wanted to have a dense compute but in a smaller form factor like a SuperBlade. And we went, the kind of product line extends one after another. But then we also have seen people are using either air or liquid because things are being now dealt at a rack scale, especially in a larger deployment arena. So what we did was, what is the power delivery going to be in these kind of systems? Traditionally, you have like a 60 amp or 100 amp power drops coming in, and then you put PDUs and then you put all these systems together. That's all great, but now we started looking at what is that we can do with an OCP enhanced or OCP inspired designs where like OV3 rack with power delivery using the DC bus bar. We have kind of transformed the product line not just for the traditional ones but also to take advantage of these things. So things are being considered at a rack scale or a data center scale. What is that we need to do in a given operating environment and how do we basically pack as much punch in a smaller footprint? You kind of figure out where it would be needed and how do we get the most optimal solution there. So that's the thought process and we've been working with AMD on how best to put these things together so that it will be relevant for customers.
D
Dave Volante25:32
And you mentioned OCP, so you're probably with your liquid cool, you're probably running warmer water right or warmer liquid, and so that affects the environmentals. When you think about advances in things like cores, memory, IO driving efficiency, what does that mean for customer outcomes?
V
Vik Malyala25:52
Well, so to start with, your workloads are changing. Some of these could be requiring more memory. Some of them could be using, take advantage of faster memory. Some of them could be requiring very fast interconnect as systems cannot just handle all the workload in a single system and it needs to be looked at multiple systems clustered together. So the idea behind this is to give that option for the customers and see how different applications will behave and make sure that we are ready for that. Let's get into the joint engineering because a lot has to happen between when you get to high yields and an OEM ready to a deployment. So what is the engineering that occurs between your two companies in between that cycle?
D
Derek Dicker26:46
I think you almost have to go back even before that. You have to go back to where we're talking about what are the products that we want to build and what are the end outcomes at a customer that we want to enable, and that begins with the concept of the actual silicon that's being designed. So the relationship that we have is one where we come in and we spend time going through what are the concepts that we're contemplating building for next generation, looking at what are the specs associated with the silicon itself and, most importantly, what's driving the features that we're putting inside the silicon. That's kind of the beginning of the dialogue. It then goes beyond that to once we instantiate a design and put a design team on it and they're down a path. When we get silicon back, as we're prepping to get silicon back, we're working very closely together on the system design complexities that exist. We may have to make adjustments along the way based on feedback as they're trying to build a system based on the silicon that we're delivering. Post that, you end up getting into validation and that's an incredibly tight relationship where if issues are being found, how do we get a very tight relationship for tight turnaround because of course we're trying to move very quickly to launch product. And ultimately, there's validation work that they're doing testing our silicon and their applications with ISVs on top. All that compatibility work needs to be done in coordination between our two companies also. So it goes all the way out to at the point of launch and then even beyond launch, Dave, we're spending time talking through issues that a customer may find out in the field. And so it's from the day that we start thinking about what we might want to build to the day we sunset the product on both sides, their servers and infrastructure and our products.
D
Dave Volante28:17
And you have to optimize for performance, you have to optimize for reliability, the environmentals, obviously cost comes in. It sounds like you have to make a lot of tradeoffs and then find that perfect balance. Maybe you could describe that process.
V
Vik Malyala28:33
So I think what we are seeing is like silicon is getting hotter and we need to figure out a way to optimize the thermal design. So we need to pick the right type of form factor and the solutions associated with it could be air cooled or could be liquid cooled because not everyone has liquid cooled in their data centers. So we need to have platforms that can accommodate and provide the highest performance using air cooled and what does that system look like? So thermal design becomes very important for that. And the second part of it is the interoperability. So obviously people are not just going to just use that for a specific workload. Some put accelerators in it. These accelerators could be inferencing or network optimization or any of those things from a security point of view. All these things need to be validated within those designs. And the networking options whether people want to use Ethernet or InfiniBand or whatnot, how best you actually put it together, what kind of storage you want to be supporting in that because now it's not just about the capacities but it's about the form factors too, E1S, E3S, and you're talking about U.2 to U3, and if somebody wants to use Compute Express Link as a part of these devices, and if somebody wants to try UALink as a part of these systems, all these things we start looking at and say, okay, is AMD going to support it and if so what needs to be done, what are the BIOS optimizations that need to come up with, and what kind of firmware updates that we need to make the system run operate in a certain way. All these things are the ones that we work hand in hand with AMD. And the designs of the systems are also a function of what we discuss with AMD and also what we discuss with customers. If a customer says, 'I'm looking for a single socket and I'm looking for one DIMM per channel,' that's quite a bit different from 'I'm looking for the most dense kind of compute doing blah blah blah.' I think the idea again here is talking to customers and getting a pulse of that and looking at what kind of technology that we need to make sure that we'll be validating because we don't know whether it's going to be popular or not but we still need to make sure that the platform is designed, otherwise we'll be late for the market. We want to be as a launch partner, we want to bring as many products, the broadest product portfolio that we can bring, air or liquid, rack scale or a single node. All these things is what we are looking at to bring, and in order to do that we need to have every bit of the design checked twice. What Supermicro does and AMD engineering comes in and like, 'Hey, if you do this, is the signal integrity going to be okay? Is it going to work on this? What needs to be done from our AGESA code for you to do a code drop for your BIOS?' and things like that. A lot to be done and a lot is being done prior to launch and after launch. I think we're going to continue to refine to get the best tuned solutions for our customer.
D
Dave Volante31:34
Because you're obviously optimizing for time to market, so a lot of that pre-work is critical in terms of being able to hit that time to market for customers. What does it mean for a customer when you can hit that time to market benchmark? I mean, what does it mean? What's the value for them?
V
Vik Malyala31:47
Time to value. I think the whole idea is that you want to take systems that are fully validated and to the extent that one can, we run synthetic benchmarks and we run ISV qualifications. Ultimately, it's the customer that is going to take it and use their applications, their workloads to go do something with it. We want to minimize that pain by bringing, like, okay, time to value or time to online, make the systems, bring them to the customer as quickly as possible so that they can have a time to market advantage and they can see the benefits quickly.
D
Dave Volante32:22
What changes, Derek, from your standpoint when we move from the server to the rack as sort of the key unit? I'm interested in where Helios fits into this equation. Like you said, we haven't talked about it much, but what changes?
D
Derek Dicker32:33
I would maybe take a step back and look at what have the customers been telling us mutually over time. One of the single biggest themes that we get on a regular basis is open, deep interest in having choice in what customers want to be able to buy. And so a lot of the strategy that you'll hear all of AMD talking about is a criticality of open, whether it's open source software that we deliver, whether it's open standards that we're a part of. The whole idea is to build an ecosystem of players that can add innovation at the appropriate levels. And so that's something that we're very much involved in and driving ourselves. And Vic mentioned it with the Ultra Ethernet Consortium, UALink, all the work we do around PCIe just developing standards as a part of that work. All of that tucked together essentially translates into systems that customers can depend upon but also get different levels of innovation from different players in the market. So Helios, to your point, is a rack scale architecture that AMD is developing. We're doing it in concert with a number of different players in the market at different levels of the stack. It's essentially a large-scale GPU design system. 72 GPUs paired with 18 Venice CPUs, all stitched together with networking from AMD. Also, our Volcano NICs coming together, and it's really designed to be validated as a complete system and then done in partnership with Super Micro and others in the industry to go deliver instantiations of this design.
V
Vik Malyala34:01
I call it like A plus A plus A. Like, we want to use the AMD EPYC with Instinct as GPUs as well as the connectivity coming from either Polar or Nyx or you're talking about Volcano going forward. Why that matters is because I just want to make sure that interoperability is fully and thoroughly done, and it's easiest when it is under a single roof. So AMD has in optimizing on their end and we validate from our front. We still give the flexibility for customers to choose their own architecture for the overall solution. But the idea again here is to validate all these things together and make it readily available for the customers to deploy. Again going back to our initial thoughts, if you're putting Helios, that's fantastic, but what is the right size of compute that you want to put in front of it and what is the connectivity that we need to be putting, what network topology that we need to bring in order to get the maximum throughput from this one? So you can throw the most expensive equipment and you can actually connect them with the fastest fabric and then you can give all the power and cooling you want and you can get the best performance, but CFOs are not going to like it. The idea here is again giving a balanced architecture that is well orchestrated and how we can provide the value to customers, and that's what we are trying to do, especially when working with AMD on this compute plus GPU together to provide an optimal solution whether it's training or inferencing.
D
Dave Volante35:36
So when the server is the sort of unit that you're bringing in versus the rack, the customers have to think about that differently, don't they? I mean, everything you just mentioned, they have to think, they have to make those trade-offs, and then these things are heavy, they can't just, sometimes they can maybe, but they make sure that the infrastructure around that can support it. What do the customers have to think about in that transition from server to rack?
V
Vik Malyala36:00
So the good thing is that I think dense compute is getting more and more common mainly because of how the industry is trending. If we take a look at our existing MI355 type of systems, Instinct 355 series, we are able to put like 8 to 10 systems in a rack getting to like 150 kilowatts per rack currently that people are already using in the data centers, so it's not going to be a total surprise to start with. But at the same time, when you're deploying it at scale, often times a traditional data center may not even fly, and we are getting strong customer demand on these platforms and they know what this platform is going to look like, and based on that they're actually retrofitting their data center or building a brand new data center. And since we are getting that pulse, what we are also able to do is like how we can optimize these compute platforms to go along with it so it doesn't look like an oddball. So we want to make sure that the compute systems are also thought about at a rack scale, think about the power delivery, the cooling, and everything so that way you don't have to start creating parallel environments with different dynamics in it. So that's something that we are working with the customers and customers are quite open and happy about it because now we are really thinking in lines of what they are trying to do and taking their feedback and start putting in our platform designs. So is it difficult? Yes. And is something that needs to be done? Absolutely. Are we doing it? Absolutely.
D
Dave Volante37:36
Liquid cooling is state of the art. It's going to give you the best performance, the best price performance, the lowest cost per token, the best price per token, cost per token per watt. But most data centers today and most enterprises aren't designed for liquid cooling. Do you think that in the fullness of time, and use an Andy Jasse statement, that that will change? Will organizations because of the attractiveness of what they can get out of liquid cooling rethink their data center infrastructure, or will they go to the cloud and the neo clouds? How do you see it?
D
Derek Dicker38:13
I think that the reality is any transition of this nature takes time. There are going to be some customers that value the business outcome and have the capital to deploy to go create greenfield infrastructure to be able to deploy liquid cooling, and that's happening today. We see that in mass in a number of different places. Internal to existing infrastructure where people can't retrofit, maybe it's going to take some time before they get to a point where they're able to either afford it or go into a retrofit of their existing data centers. But I think it's undeniable at least the value, as you suggested, that comes associated with liquid cooling. The reality is for us between our two companies, we're going to offer both and help monitor the transition over time so that we can give the customers what they need when they need it in time. That's how I simplistically look at it.
V
Vik Malyala39:01
Yeah. I mean, I'm not a golfer, but you know you don't take a putter for a driver. So it's the idea again here is to give the optimal solution for the customer. So if people are doing small scale inferencing workloads, I mean the Instinct MI350, 355 and also the PCIe version the MI350P, these things are going to just fit fine in an aircooled environment for customers. If they want to go with liquid cooling in a small scale, they can still do the MI355 that can do that. And as for people looking for the best in class and the beast, if you want to have it, Helios is there. So the idea is whether it's air cooled or liquid cooled, rack scale or at a node level, could be like 1 GPU to 8 GPU type of systems, we want to give them the option, and that's what the industry wants and that's what we are trying to do so that we don't want to slow down the innovation, we want to accelerate together.
D
Dave Volante39:58
That's exactly right. Okay. I want to ask you guys about the, you have a lot of experience obviously with general purpose infrastructure, x86, AMD has just done a remarkable job over the last decade. What happens to the traditional workloads that are running out there? The trillions of dollars of investments that are there, does that get absorbed into the functions that are running on that legacy software estate and the infrastructure around it, the offloads to storage and networking and all those functions in x86 land, do they get absorbed into a new AI infrastructure? How do you see that playing out?
V
Vik Malyala40:43
Where is that crystal ball? [laughter] I always...
D
Dave Volante40:45
If you have it, it could share.
V
Vik Malyala40:47
But you're right. I think it's a very good point because traditional workloads they are going to evolve. People are trying to figure out a way to optimize that to the extent possible because now all of a sudden, let's take an example of like two generations ago systems where if you want to run an application that requires a certain memory footprint, it may not even fit in a given system. So now they need to look at how it's going to run across four or five systems put together and connect with whatever the network at that time. So now look at 16 channels of memory, 256 cores in a single system put in a dual socket, you are able to get like what, 512 cores or maybe double that number of threads and whatnot. The entire architecture can be redefined in terms of how the software is going to run. So the idea again here is if people are running certain things, they can do at least the same if not multiple times better traditional workloads, but at the same time when they are doing that, they want to take advantage of what is happening with AI. Every company is looking to see what is that they can do to improve the decision-making, how do they run agentic workloads, and how do they improve the customer experience? Think of customer support calls or an internal knowledge base, database queries and whatnot. Every one of these things will be done far better either you use a traditional workload as is or you optimize it further using some agentic workloads. Both of them can be supported with the same system. So it's prudent for anyone to take a look at what they are trying to do and how they can take advantage of these new systems that are bringing the value to them.
D
Dave Volante42:30
And I think you would agree that obviously a startup is going to write code that is AI native, that could create some interesting competitive dynamics in industries with existing software companies, and those existing many of those existing software companies will likely refactor their code for AI. And it's a transition. It's not an either/or necessarily, but when we hit these waves, you tend to see these disruptions occur, new companies come out of the woodworks that we've never heard of before. And it is a crystal ball question. I don't know if you have anything to add to this conversation.
D
Derek Dicker43:10
I would just agree with what Vic said. I think there's going to be a transition and every single enterprise customer is going to have a decision to make in terms of how quickly they want to migrate and the way that they want to migrate. The good news is that existing applications today can be coordinated into an agentic workflow with some level of additive technology. Then for those who want to go and really invest even more heavily and have the capital to do so, reimagining their entire operation is entirely possible too. I think the ones that are going to have the most difficult time are those who choose not to do anything and wait and see. I think that's the most important thing that we see right now.
D
Dave Volante43:44
Yeah, I think that's a key point. You're not going to run your agentic enterprise on legacy infrastructure. It's just that's a failure mode. You shouldn't even try.
V
Vik Malyala43:52
Right. I think going back, I was thinking through and one of the things that we have seen is especially the enterprise customers or the people who have already data centers, they are struggling to see how to ensure that they are able to take advantage of the latest gear that is coming in. What Supermicro is doing is to figure out a way to lessen the pain. We have started investing in so-called data center building block solutions or DCBS, where we are going to help them to retrofit their data centers by bringing liquid cooling into that, and could be like water towers and things like that, or could be power shelves that need to be put in this or change the power delivery to the systems. This way, one doesn't have to go and build a brand new data center, but can take advantage of all the latest within the existing data center by making certain modifications. If someone wants to start brand new, we can certainly help them on that front as well. But to just kind of wrap the thoughts on that one, if the customers want to take advantage of the latest, they are not alone. We can actually certainly help them on a system level as well as the deployment level.
D
Dave Volante45:03
All right, let's close with a single action item for IT decision makers. Derek, you can start and then Vic, you bring us home.
D
Derek Dicker45:13
I think probably the most important thing that I would offer up is it's not a scenario where anyone can afford to wait. I think the infrastructure is changing so fast. The capabilities that are available are changing so fast. And really the most important thing that one can do is understand where they want to head from a workload perspective, put infrastructure in place, and we'll partner with folks like Super Micro and ourselves to the extent that there's an interest. We spend a lot of our time working closely with end customers in partnership with Super Micro to help them evaluate options, to conduct POCs, to take advantage of some of the transitionary technologies and offerings that Super Micro has, and to put new infrastructure in place to take advantage of this exciting time that's coming ahead of us.
V
Vik Malyala45:57
Balanced platforms and architecture is the key. What we are doing together is whether you want to use the standard compute, you want to use accelerated compute, how to put these things together to make sure that it meets or exceeds someone's expectations. And in order to do that, one has to do a proof of concept quickly. We can give as much information possible that we have already done together. But ultimately, it boils down to customers using it as a POC and very quickly roll it into production. And that's what I think people have to start with.
D
Dave Volante46:31
Guys, it's been really interesting to learn more about the partnership over these three panels that we've had going deeper into some of the engineering work that you're doing, the roadmap. I can't thank you enough for spending some time with us.
V
Vik Malyala46:42
Thank you so much for the opportunity. It's been great. Thank you.
D
Dave Volante46:45
Yeah, you're very welcome. And thank you for watching. I'll just leave you with this, to Derek's point. Don't wait. You got to get on that AI curve and your legacy infrastructure is not going to support your future agentic infrastructure. You need to think about that and plan for the future. This is Dave Volante for the Cube. Check out the panels that we're doing here on the Cube.net. This is in support of Advancing AI where the Cube is there as well. Check out Silicon Angle for all the news. And thank you for watching.