CEOInterviews.AI
Start App
Vik Malyala
Senior VP, MD & President of the EMEA Division, Supermicro

Vik Malyala, Supermicro | AMD Advancing AI 2026

📅 Jul 22, 2026 SiliconANGLE theCUBE 23 MIN 3394 VIEWS 54 SEGMENTS · 3 SPEAKERS
Vik Malyala, Supermicro | AMD Advancing AI 2026 00:00 - Intro 00:02 - Navigating AI Leadership and Industry Transformation 04:08 - AMD's Role in AI and Partnering Strategies 06:41 - Demand for GPUs and Broad AI Adoption 08:43 - Configurability and Scalability in AI Systems 11:57 - Co-design and Engineering with AMD 15:41 - Optimizing Helios: Balancing Demand and Overcoming Data Center Limits 18:19 - Strategies for Sustainable Market Growth and Innovative User-Centric Design

What Vik Malyala said

Written from the verified transcript and checked against it. Every figure links to the moment it was said.

Vik Malyala, Senior VP and President of Supermicro's EMEA division, discussed the rapid evolution of AI infrastructure at AMD's Advancing AI event. He noted a shift from RAG models to agentic AI, which is driving demand for both GPUs and CPUs, with CPU workloads impacting 50-90% depending on the source. He highlighted supply chain constraints, especially in memory, and how Supermicro works with partners like Micron, Samsung, and Hynix to secure allocations. Malyala compared AMD's Instinct GPU adoption to its EPYC CPU trajectory, noting ecosystem development is key. He mentioned strong demand for MI350, MI355, and Helios, and that Supermicro is designing platforms with DC bus bar power delivery to support Helios. He dismissed bubble concerns, citing strong pent-up demand, and emphasized the importance of reducing token costs for enterprise adoption.

Key takeaways

  1. Agentic AI is driving significant CPU demand, with 50-90% of workload impact on CPUs, making CPU architecture critical.
  2. Supermicro is designing most server platforms with DC bus bar power delivery to align with AMD's Helios architecture.
  3. Demand for AMD MI350 and MI355 GPUs is strong, extending into Helios, with 12,000 people signing up for the event.
  4. Supply chain constraints mean systems can no longer ship in a week or 10 days; longer lead times are expected.
  5. Malyala sees no bubble, citing strong pent-up demand and a growing TAM, not a speculative top.

Numbers and commitments

FigureWhat it refers toTypeAt
50 to 90% CPU workload impact in agentic AI metric 10:13
12,000 people signed up for AMD Advancing AI event metric 4:08
600W power for future AMD GPUs metric 13:29
40 cents cost per million tokens for enterprise adoption price 20:55
$40 cost per million tokens that CFOs would reject price 20:55

Chapters

  1. 0:00AI infrastructure demand and agentic AI
  2. 2:15Supply chain and memory constraints
  3. 4:08AMD's evolution and Instinct adoption
  4. 6:29Supermicro's backlog and margins
  5. 8:43Configurable platforms for diverse workloads
  6. 10:13CPU role in agentic AI
  7. 13:01Co-designing with AMD for future GPUs
  8. 15:41Helios demand and data center readiness
  9. 18:29Bubble debate and market outlook
  10. 20:55Token costs and enterprise adoption

Questions asked in this interview

9
  1. 2:15Prices go up in memory, but people want solutions now, right?
  2. 6:29Is it as simple as just demand is so far outstripping supply?
  3. 8:07Who cares about token cost and frontier models are going to dominate everything?
  4. 8:37How are you and AMD engineering for these rapid changes and seeing around corners and how do you accommodate this change?
  5. 11:51So it's criminal to not have the best CPU architecture to orchestrate all of that, right?
  6. 13:01How do you guys partner now in the AI era and where is it the same and where is it different?
  7. 15:41What's the fundamental customer value prop?
  8. 19:05You don't think we're in a bubble?
  9. 19:11The market's going like this, isn't it?
John Furry 0:06 ↗
Welcome back. We're the cube's live coverage here in San Francisco, California. AMD's advancing AI event. All the leaders are here, industry participants. It's a free event. So, a lot of practitioners and technologists here checking out the rack scale systems, checking out all the new gear and networking, servers, everything's here to power the next generation. I'm John Furry, host of the cube with Dave Volante, my co-host Nick Mielis here. He's the chief business officer of Super Micro cube alumni. Recently I interviewed — we interviewed — long time no see.
Dave Volante 0:31 ↗
Super Micro, Dave just interviewed him. You guys had a big storage summit with the cube.
John Furry 0:37 ↗
Great to see you.
Vik Malyala 0:38 ↗
Thank you for having me as always. It's a pleasure to talk to you. It's lots happened since supercomputing when we were riffing on the neo cloud growth AI infrastructure a lot of content we put publish out together but the big thing is just the massive buildout the demand curve Dave was mentioning before we came on camera and just the nature of these AI factories. They're kind of taking shape as a new architecture. It's not the rack and stack days anymore, although there's still a lot of that going on, but it's a whole different computing paradigm. Explain the current situation. Where are we today? November. We just had a lot's happened since November. I'm telling you, whatever we think is the pace at which the industry is moving, three months later, it is proving us wrong. What I have seen is that last time in November when we were talking about there was hardly much discussion about agentic AI, it was all about rag models and people are trying to develop these different applications but then in a snap it all changed. Now we are talking about heavy use of agentic AI across the board and we have a strong demand that is coming up not just on a GPU but also on the CPU because of that. As for the GPUs are concerned, now enterprises are starting to pick up. It's not just all about the training clusters but also the application stack that is being used on them which is basically making the enterprises adopt that. So what I see is that it's a very broad scale adoption of AI infrastructure which is propelling this massive demand.
John Furry 2:15 ↗
And you guys have such a great history. I always kind of flex the Super Micro success over the many generations. You know, supply chain and right now that is the number one thing people care about. Prices go up in memory, but people want solutions now, right?
Dave Volante 2:31 ↗
Talk about the impact of what that's done and also how that's changed, how you guys think about building these large AI factories and these systems.
Vik Malyala 2:38 ↗
It's a complex equation. I mean, no, there's no ifs and buts about it. If you take a look at the likes of AMD and Nvidia for example, for the rack scale solutions they are bringing the predictability in the supply as well as pricing by them working directly with the memory manufacturers on the supply side at least and in most of the cases some level of pricing also. But as for the standard HGX platforms and everything else is concerned, this is going to be a very tricky market for quite some time to come and luckily for Super Micro, we've been working with pretty much every one of the industry leaders, whether it is memory or whether it's flash, think of like Micron, Samsung, Hynix, Solid, and every one of them for the longest time. As we build these platforms and solutions, we give them enough visibility into where the systems are, knowing who the customers that we are supporting and what problem we are trying to solve, which sometimes helps actually in getting an allocation at the right time. But no matter how much planning that we have, we know for a fact the supply is a lot more constrained than ever before and we continue to work with them to see how best we can address and most importantly we set expectations to customers that it's no longer possible to have a system ship in like a week or 10 days. It takes a longer time and by taking the orders in by having a proper forecast it's also helping us to plan better.
This event, I won't say coming out party, but we've been following AMD. They've been in the x86, we've seen that growth but I kind of been following, we can see them hiding the ball a little bit. They had FPGA GPUs coming. What is the new AMD? You talk about the relationship you have with AMD and how the new and improved AMD, just the now the AI high systems side of AMD starting to show, starting to see the results. What's it? Explain what it means to the market and what people should know about it. I'll give a parallel story here. Back in 2015, 2016, when I was working with Dr. Releases on the first generation of the EPIC platform, what happened was the hardware was fantastic. I mean you have many cores and the PCI Express lanes and everything, but the software and other ecosystem wasn't quite ready. So while it's adopted in a very specific marketplace, whether it is HPC, whether it is hyperscalers, the general adoption took an extra two cycles after Milan, then the next step like Genova, Bergamo, and all these things. So what I have seen at every generation, more people started adopting, the ecosystem started expanding, and to the point where they actually have the market leadership in x86. The same situation I see happening with respect to Instinct also. Initial units, hardware was fantastic, but the ecosystem is the one that needs to be developing. I was told that some 12,000 people signed up for this event, which is something that no one would have imagined two years ago if you ask.
In terms of the software developers, what I am seeing is that as people start figuring out a way to use the GPUs, then the adoption of the GPUs beyond the training clusters is going to happen and we have seen a good demand buildup especially around the MI 350 and 355 and the demand is kind of extending into the Helios. There's a lot of people asking about it and people are trying to understand. Mind you, it's not just about the technology. People need to figure out a way to host them in their data center. It's not going to be easy to just run a Helio.
John Furry 6:19 ↗
It's like the GPU is the initiation. Wow, I love this value. Like, wait, I need more compute. So, the discovery on the ecosystem becomes kind of its own progression.
Vik Malyala 6:28 ↗
It is. It is.
Dave Volante 6:29 ↗
So, you guys hinted that you're going to that your margins are up. Your backlog is I think 60 billion was the number you published. Is it as simple as just demand is so far outstripping supply?
Vik Malyala 6:41 ↗
I mean we have done something right. Ultimately we need to take care of our customers' demands and as customers become successful, the business starts to grow and that's precisely what happened in this case. We have a growing set of customers and wider adoption of the platforms. Think of an enterprise, think of banking and financial sector as well as the traditional compute and GPU environments. These are the ones that are actually creating the demand for us. As far as the margins are concerned, it's a mix of various things. You're talking about data center building blocks that we are bringing to expedite the deployment phase of these GPU clusters, part of it service organization coming up to make sure that the things are going to be up and running and be able to stay up and running at a high percentage of time, and the relationship that we have with ISVs for the storage and what we are bringing in terms of the storage solutions to customers. So a combination of all these things is driving both the adoption as well as the impact on the margin. Margin is something that's going to fluctuate especially because of how the whole supply chain dynamics are working today, but we are certainly emphasizing on value and how we can actually help customers to bring some predictability into their equation.
John Furry 8:07 ↗
Yeah, the Super Micro storage summit. I was there. It was a great event. Some really good content. You were talking before about how things change so fast. We actually earlier today we came John and I came up with a list. It's like, oh yeah, rag based chatbots are going to add huge value. You got token cost. Who cares about token cost and frontier models are going to dominate everything? You know, these other small models don't mean anything.
Dave Volante 8:37 ↗
How are you and AMD engineering for these rapid changes and seeing around corners and how do you accommodate this change?
Vik Malyala 8:43 ↗
So none of us have the crystal ball but what we can do is to bring the configurability into the platform design and be able to size it differently. One of the things that we have done is if you were to take a look at the GPUs itself, you have the MI 350 and MI355 that is the HGX platforms predominantly being used especially with AMD and then we also have the inferencing point of view which supports different accelerators. We have several accelerators that actually based on AMD as a compute platform and they have these accelerators as a PCIe and AMD themselves are coming up with MI350P which is also a PCIe based accelerator. Why I'm saying it is that for the guys who want like you know think of like the Rolls-Royce, you have these platforms at the rack scale and it's fantastic but for the ones who are trying to use for different applications they may not need that in order to have the right ROI for what they are trying to do. They can go with partially populated systems like for example I can have a system that supports up to eight of these PCIe GPUs and people can actually start with two or four or scale up to eight depending on what they are trying to do. If it goes beyond eight and if they are trying to scale up then we have the rack scale solution and multiple racks we can connect and create a cluster. So there's no rocket science per se but what we are trying to do is to give the option to the customers on what might work for them.
The second part of what we are seeing is you mentioned you started with the rag models, that one is a single loop. You ask a question, the model is going to run it, spit out the answers and you're good to go, whether it's right or wrong and how much of it is hallucination depends on the data it's trained on. But this agentic is a completely different spin. You're talking about it needs to think, it needs to reason and it needs to go into multiple loops. It needs to memorize things and it spawns off so many more sub agents who are doing all their work. What is happening with this is there is a significant demand for the compute, not necessarily the GPUs. The GPUs are doing a fantastic job. Think of the rag models, when it gets to that point it has matrix multiplications it can do very quickly, it can spit out the answers. That's fantastic. But on the compute side, how do you make sure that the right type of queries are going to the GPU and the rest of it is handling on this? In the rag model all it is doing is just sequencing, managing the pipeline, orchestrating done. But with respect to agentic AI, now we're talking about the workload that's happening on the CPU which is impacting 50 to 90% depending on where you read. So if the impact is that much that will be handled by the CPU then it's a criminal waste not to have the best computation.
John Furry 11:38 ↗
And the cost and the cost financially, the overhead and the technical costs are massive because...
Dave Volante 11:45 ↗
You got to feed the GPUs. You got to do routing so there's a whole another orchestration layer here.
John Furry 11:51 ↗
But finish that thought. So it's criminal to not have the best CPU architecture to orchestrate all of that, right?
Vik Malyala 11:56 ↗
Right, right. And that's where I think you know now you have the CPU, you have the GPU, you're connecting bringing AMD into the equation, having the polar cards now, and then the Pensando part of the equation. So what we are trying to do is depending on what customers are trying to do, we can size it right. So have a PCIe based accelerators or the standard HGX type of platforms, how much precompute clusters that you need and how much that you need it for agentic AI, what is the kind of bandwidth that is needed between these compute as well as the GPU platforms, how do you pair it with the appropriate storage so that when things go past the size of available memory and storage available within the system, how it's going to expand into the storage for the KV cache allocation. So the entire thought process of Super Micro as a building block is absolutely making sense right now in this kind of scenario because as things get more complex, ultimately what matters is how are we bringing the right value to customers and sizing it right and making it run most optimal way.
Dave Volante 13:01 ↗
Talk about the AMD relationship when it comes to code designing and engineering. A lot's going on to these rack scale systems. They're essentially super servers as one and then you have distributed computing also. Other areas they're going to have footprint, the nice cards are going to actually target workloads. They all have to talk to each other. How do you guys partner now in the AI era and where is it the same and where is it different?
Vik Malyala 13:29 ↗
So take an example of the PCIe based cards. It started with 75W, 150W, 225W, 300W, now 600W. So AMD gives us a visibility: hey, my GPUs in future is going to be air cooled but I need a platform that supports 600W, or I'm going to have 600W but I would like to bring the maximum density in which case I will say like well if that's the case let's see if that design can accommodate liquid cooling so I can actually put those many cards. The other part of it is let's say AMD silicon for the GPUs is based on PCI Express Gen 5 or Gen 6. Based on that I can say well I can actually put in the Venice platform with the PCI Express Gen 6 so you have much higher throughput that is going to be supported in the platform. So what we typically look at in this scenario is what are the different technologies coming from them, what is the time frame, and when I bring a platform, how am I going to make it run in the most optimal manner for that? Does it require liquid cooling? Does it need to be air cooling? What is the form factor? What could be the power delivery? And we look at where that product could be adopted. Like I mentioned, when people start adopting Helios, let's say deployment starting whenever it is, we will be ready with AMD to have our compute platforms to go along with it. The public information is that AMD Helios is going to be running with the ORV form factor with the power delivery using the DC bus bar. The entire server product portfolio we have transformed, not entire but most of them, like Hyper, Cloud DC, the Twin platform, the Flex Twin, all of them we have started designing with the power delivery using the DC bus bar. Why? Because as they go into the data center they don't need to figure out like okay now I need to have standard power supplies, I need to have a 63 amps or 100 amp power drops and the PDUs. Helios doesn't need it, DC bus bar, so I'll just do the same DC bus bar on the compute also. That way you don't have to go and rediscover. This is how we are looking at what would be needed for the customers in a given time frame and what kind of platform that I need to develop or platforms I need to develop in order for us to get that rolling quickly.
John Furry 15:41 ↗
So where do you see Helios fitting that rack scale system? What's where do you see the demand for that? What type of workloads is it going to support? What's the fundamental customer value prop?
Vik Malyala 15:51 ↗
So at this point, people if they're doing the primary inferencing workloads, HGX platforms are doing a phenomenal job. But when it comes to frontier models, that's where I see the immediate demand for Helios because the models are getting trickier. Trickier because it's no longer just about how many trillions of parameters, it's all about the optimization that they're bringing into that. And when it comes to inferencing, I think it was like a year ago when we talk about DeepSeek and everyone like 'oh my god the world is coming to an end' and after that everyone has multiplied their numbers by like three or four times.
Dave Volante 16:28 ↗
It's happening again now. Models of innovation, they're talking about the point is all of it is going to drive the democratization. So whenever people figure a way to run things more efficiently, it's not a negative thing. It basically makes it easy for others to engineer it. It's called engineering buying opportunity.
John Furry 16:43 ↗
Well, no. This is where I like the compute direction because when that gets smaller, faster, cheaper. That sounds familiar. Moore's law. I mean, so you're starting to see a kind of a dynamic where the engineering is going to enable things to run cost effectively. That should help the enterprise for sure.
Vik Malyala 16:58 ↗
For sure.

25 more exchanges in this transcript

Sign in free to read the rest of this interview. No card required.

Sign in to read the full transcript

Cite this transcript

APA, MLA, BibTeX
APA

Malyala, V. (2026, July 22). Vik Malyala, Supermicro | AMD Advancing AI 2026 [Interview transcript]. SiliconANGLE theCUBE. CEOInterviews.AI. https://ceointerviews.ai/interview/1115190/

MLA

Vik Malyala. "Vik Malyala, Supermicro | AMD Advancing AI 2026." SiliconANGLE theCUBE, 22 Jul. 2026. Transcript, CEOInterviews.AI, https://ceointerviews.ai/interview/1115190/.

BibTeX
@misc{malyala2026_1115190,
  author       = {Vik Malyala},
  title        = {Vik Malyala, Supermicro | AMD Advancing AI 2026},
  howpublished = {Interview transcript, SiliconANGLE theCUBE. CEOInterviews.AI},
  year         = {2026},
  month        = {jul},
  url          = {https://ceointerviews.ai/interview/1115190/},
  note         = {Speaker-attributed transcript with timestamps}
}