Back
Vik Malyala
Senior VP, MD & President of the EMEA Division, Supermicro

Vik Malyala, Supermicro | SC25

🎥 Nov 20, 2025 📺 SiliconANGLE theCUBE ⏱ 24m 👁 9587 views
In this interview from SC25, Vik Malyala from Supermicro joins theCUBE's John Furrier and Jackie McGuire to discuss the ...
Watch on YouTube

About Vik Malyala

Vik Malyala, Senior Vice President, Managing Director and President of the EMEA Division at Supermicro, has been discussing the company's AI infrastructure portfolio and the importance of balancing CPU and GPU architectures for enterprise AI workloads. In interviews and presentations, Malyala emphasized that AI infrastructure should not be viewed solely as a decision around GPUs, stating that CPUs, accelerators, memory, networking, storage, cooling, software, and rack integration must work as one system. He noted that with the rise of agentic AI, workloads on the CPU can impact 50 to 90 percent of latency, and called it "a criminal waste not to have the best computation" for orchestration. Malyala also highlighted the rapid pace of technology refresh, which he said is now less than a year, while typical data center projects take 12 to 24 months, creating a fundamental challenge where infrastructure can become outdated before deployment. At COMPUTEX 2026, Malyala showcased Supermicro's next-generation platforms powered by NVIDIA, AMD, Intel, and ARM processors, including the Vera Rubin, HGX B300, and AMD Helios systems. He described the AMD Helios as a "beast" and a "massive 7,000 lb machine." Malyala also discussed the company's data center building block strategy, which includes liquid cooling, rack-scale systems, and infrastructure management software. He referenced estimates of $1.5 trillion in IT infrastructure spending, calling it a "big number," and noted that customers are increasingly exploring on-premises deployment after seeing the costs of public cloud services. Malyala added that Supermicro is working with partners such as AMD to co-design systems that can be quickly taken from proof of concept into production.

Source: AI-verified profile updated from Vik Malyala's recent appearances. Browse all interviews →

Transcript (46 segments)
J
John Furrier0:02
Welcome back everyone to Supercomputing live stream from theCUBE. I'm John Furrier hosting with Jackie McGuire from theCUBE Research. We're breaking down all the AI infrastructure innovations from HPC now full on AI cloud. This is day two of three. Vik is here. He's the Managing Director and President of EMEA, Senior Vice President of Technology and AI at Supermicro, a company that's been instrumental in providing a lot of the compute and gear over many decades. Now, the center of the AI factory conversation. Vik, thank you for coming on theCUBE to wrap up day two.
V
Vik Malyala0:40
Thank you for having me here.
J
John Furrier0:42
I forgot to ask you. So you're SVP of Technology and AI at Supermicro. I know you guys have been doing a ton of work, storage, networking and obviously the compute side. I think my startup, 16 when I started this company, I actually bought a Supermicro box and I had to host it and then the cloud came and they're still buying. So you're at the center of the AI revolution, but the game is still the same.
V
Vik Malyala1:06
Yeah, game is still the same, getting better because the nerd nature of humans is not going anywhere. So people still like to build and experiment with things, so we still continue to support that. That's a huge customer base for Supermicro. But at the same time, a lot of things are going into cloud, but there's also repatriation from cloud happening because of cost and other reasons. Cost is how it started, but cloud is increasing. The total pie is exploding, so don't think it's going away, but people are repatriating because of cost, data sovereignty reasons, and because they want the flexibility to choose whichever platform or technology they want to adopt.
J
John Furrier1:54
The systems game is back. If you look at all the AI factory conversations around AI infrastructure, it's essentially a system architecture. It has all the elements, and the density is increasing. This is the core competency of Supermicro. How are you guys looking at that right now in terms of what's state-of-the-art for you guys as enterprises? The data and the compute, that's the name of the game. The on-prem, those are the crown jewels seeing a lot of build out. Of course, the neoclouds are emerging fast, and then you have the neoclouds and the enterprise, two hot sectors of build-out right now.
V
Vik Malyala2:30
The good thing is that even when Supermicro was 30 years ago making subsystems, the largest exciting part for Supermicro in larger deployments is high performance compute. Because of that, we are used to being part of larger deployments with performance being the key and density being a part of it, because in order to reduce the total cost of networking overhead, nodes need to be next to each other. That's how the denser platforms came about. In fact, we worked with Intel and we were the first to introduce the twin platform, which is pretty much a de facto standard now in the compute world. Second, as we look into AI deployments, one of the things we leveraged is that we develop everything in-house, which helped us adapt to changes very quickly and bring platforms that are more relevant for different verticals, whether enterprise, cloud, or HPC. Now, when you go into AI factories, the way I see it is we are the factory that is enabling the AI factories, because you need to bring the infrastructure together in a certain way very quickly and efficiently, and be able to roll in these systems along with the software with NVIDIA and get them to customers.
J
John Furrier3:51
We are going to do a big thing on inside the AI factory content series, because there are a lot of answers around what's in the factory: the memory, the storage. We actually did a showcase this year with Supermicro, and people might not know this, but you guys have a huge ecosystem of partners.
V
Vik Malyala4:09
Absolutely.
J
John Furrier4:14
Open systems has been a great wave that you guys have ridden as a company, but now you've got open ecosystems, co-design. You have partners that are critical to that. So talk about the open and the partner network that you guys have, because it's pretty huge.
V
Vik Malyala4:28
It's really a fundamental shift in business and the kind of collaboration over competition.
J
John Furrier4:33
I think it is. It's all the collaborative effort. It takes a village to get things done.
V
Vik Malyala4:37
So from a standards point of view, initially we started with all the 19-inch standard rack designs, which are open standards, and we started adopting the best practices in the industry. For example, we started incorporating OpenBMC and OpenBIOS, which gives people the flexibility to modify to make it work for them. Then we went with the OCP NIC, Open Compute ORV3 based GPU accelerator modules. When we started incorporating that, we also looked at the rack level as people deploy at scale, when density increases. Take an example of NVIDIA GB300 NVL72, which a lot of people are interested in. This is nothing but an ORV3 rack. So we started adopting that as well. So two-fold: whether it's conventional standards like 19-inch racks with IPMI as a management interface and Intel or AMD as standards, that's one thing, but we're also adopting quite a bit with the OCP community to incorporate those changes into platform design. It's about giving customers a choice. That's on the hardware front. On the software side, on an accelerated computer like GPU platforms, the most expensive thing in the cluster is GPUs. We want to keep them busy, which means the network with the lowest latency, highest throughput, and highly performant storage. We come up with many high-performance storage arrays, but storage array is hardware. We have at least 15 different ISVs we work with on storage, like VAST Vector, DDN, Qumulo, OSNexus.
J
John Furrier6:28
The who's who.
V
Vik Malyala6:30
Exactly. Whoever is focusing on high performance, high resiliency, we are working with them. For networking, we work very closely with Arista, Juniper, Cisco switches, or we also have Supermicro switches based on Broadcom silicon with SONiC enterprise OS.
J
John Furrier6:53
Like Tomahawk?
V
Vik Malyala6:55
Standards, standards, standards to the extent possible. If it is open standards or industry standards, we want to incorporate them into our designs.
J
John Furrier7:00
How are you bringing all those subsystems closer together with the density? Because you guys have been living the storage compute network game for generations, but now it's large-scale clusters, in some cases the data center itself is a supercomputer. So you have to bring together all those three.
V
Vik Malyala7:19
It's a very complex equation. I'm actually super excited to be part of it. Take an example of a GB300. A single rack goes in excess of $4 million. If someone were to deploy it in a data center and the liquid cooling is not ready, the power delivery is not ready, the cooling is not there, then that entire infrastructure sits unused. This is something we are looking to address. Because of our building block approach, we took it to the next level and say think of the data center as a building block. In your parlance, it's like a massive supercomputer, and everything in it needs to work, and the telemetry has to work for people to deploy and use it. So what we are doing is taking a system sitting in a rack, a liquid cooling distribution unit, and the distributor, plumbing into the facility water and water towers outside, and bringing all of them into a single pane of glass to manage it, based on open standards. We work with different ecosystem partners like CoolIT or Motivair, and we have our own offerings validated on that. At every stage, we want different building blocks so when you put it all together, it works seamlessly.
J
John Furrier8:55
I love this for technology. In the later part of the 90s, 2000s, even the 2010s, we had this whole walled garden economy where venture capital and private equity poured gas on the flame of intellectual property, don't work with other people, nothing is interoperable. It's really cool to me that as millennials start to come into buying power in the enterprise, they have a really strong preference for collaboration over competition. So it's nice as an end user to see the walls come down, this interoperability, the embrace of open standards, because I want to be able to choose best of breed. Or if I have an older data center and I want to start bringing in new components, I may not be able to bring in a full factory, I may need to do it piece by piece. I think this is one of the best things that can happen in technology: we have to grow so fast that companies don't have a choice but to interoperate, because it kept a lot of innovation from happening through at least the 2010s.
V
Vik Malyala10:01
Look at it, we are at Supercompute. This is a bunch of engineers and developers, whether hardware, software, or AI-centric applications. People are talking about agentic AI workloads. When all these people work together and innovate quicker, that is what we want. That only happens when there are some standards. Standards make it easily available for people to develop on. It's going to be a win-win for the industry and everyone.
J
John Furrier10:30
Standards really make the builders innovative and creative. That's one of the things coming out of the AI side too, the creativity. You mentioned GPU utilization. The theme we've been hearing throughout the week is GPU starvation: this is the most expensive resource, but you've got memory and storage, so let's rethink that system architecture to make those GPUs utilized. We're hearing some of the big neoclouds are clocking 95% plus utilization, while enterprises are around 65%. So there's a lot of idle GPU. Engineers are thinking, 'Hey, I could do things like offload.'
V
Vik Malyala11:15
There are so many companies coming up and working on solving this problem. It started with training workloads, which beat up the systems as much as possible as long as the network and storage supported them. But as inferencing workloads come in, it's going to be spotty. The other thing is heterogeneity. We talk about AMD, NVIDIA, Intel, and others coming on inferencing applications. When people start adopting all this, companies are popping up to abstract that layer and guide the application to the appropriate accelerator or GPU. It's all about how we keep GPUs busy, improve productivity, and bring total cost of ownership down.
J
John Furrier12:07
It's load balancing on steroids. It's funny because GPU came about as graphics processing unit, but it's actually just compute. NVIDIA gets the name GPU and it's got the math behind it.
V
Vik Malyala12:22
Because it started off that way. They led the industry and continue to do so.
J
John Furrier12:26
It's a high performing AI processor, but as inference becomes nuanced or situational, whether at the edge or in a workflow, you can use other compute and architect things around it. Maybe the workload will use a small language model. The big guys are doing training, but you can do reinforced learning on other models and then do inference on the fly.
V
Vik Malyala12:52
The best thing about agentic AI is that we're not trying to use an LLM to solve all the world's problems, because they're so inefficient.
J
John Furrier13:00
Yeah, but you can figure it out.
V
Vik Malyala13:01
I came from a chip development background and used to spend so much time in validation. I was talking to a company called ChipAgents, and looking at how they use agents, it reduces time by 60-70% and does a phenomenal job. Incorporating these things and the amount of innovation happening requires huge compute power, which translates to how we make deployment easy, run it more efficiently, and innovate.
J
John Furrier13:32
I think DeepSeek was really instructive. It came out around November 2023, and through 2024 there was speculation about what it means. But it taught us that you can configure other chips to get performance levels with open source software, and it's not a one-trick pony with just GPU. You have other mechanisms. So you're in this business of builders. Builders can now configure.
V
Vik Malyala14:02
Right. I see that innovation is happening. We have several companies working with us, taking our hardware and popping up their accelerators. For example, FuriosaAI is one, Rebellions is another, and Qualcomm and many others. Whoever is making accelerators needs systems that provide the best value and performance so they can innovate. From our point of view, while we lead with NVIDIA based on customer preferences and requirements, there are many other options coming up, and we are very much in the game to help as many ecosystem partners.
J
John Furrier14:41
I have to ask you. We've been talking about HPC and AI coming together, kind of one thing now, even though HPC has a storied history at this show. What's the difference between HPC workloads and AI workloads as they blend together? What's your view?
V
Vik Malyala14:59
There are a few different ways to slice and dice. First, from an infrastructure point of view, AI clusters took all the best practices from HPC and took them to the next level because of the need. Second, today much of the AI workloads are about training models, which is FP32 or FP16, and inferencing is going towards FP8 and FP4. Whereas the scientific community is still dealing with floating point 64-bit operations. So from an HPC point of view, there's still a lot of focus on Intel and AMD later generation CPUs, but we also have GB200s which are fantastic for HPC applications. When you look at newer ones primarily focusing on AI, the scientific workload shifts sideways. HPC historically was designed to scale across many nodes, while AI is trying to pack as much punch in a single node. So both scale-up and scale-out are becoming very important. HPC historically is more scale-out than scale-up.
J
John Furrier16:27
So AI has got a little bit of both. I have to ask you because you're involved in the EMEA side of the business, Middle East and Africa, as well as SVP of Technology and AI. Sovereign clouds are coming, AI clouds. You already see neoclouds. Talk about the sovereignty piece, because sovereign AI is being talked about a lot. There are different versions of that depending on where you sit. What is sovereign AI? These AI factories continue, you're going to have core factories being built out, and edge factories going to be hyper-converged at the edge, and micro... This is distributed computing. So if that happens, we're going to see a lot of clouds. Sovereign has to play into that.
V
Vik Malyala17:08
The way I see it is that the entire AI we are planning to use is to make the user experience better, whether it's language, health, shopping, or anything. When you're dealing with people's data, and at a country level, it becomes a lot more political than ever before. Earlier, data was there but people weren't doing much with it. Now, people are very sensitive about what data is used for, including national security. So data sovereignty is becoming important. To provide localized services, you need models optimized for what you're doing. That's why certain large language models are adopted and retrained with local data, with guardrails, and they stay local so no one else has access. This is not a fad; it's absolutely needed and increasing day by day.
J
John Furrier18:09
Vik, you mentioned the countries. One of the things I'm learning from this show is that the national angle is a new metric that's superseding or in some cases at parity with classic nerd metrics. How do we optimize efficiency? But now you have this other qualitative aspect. It has nothing to do with money, though it does at the end of the day, but decisions are being made at large scale over national cloud. Even in the US we see NVIDIA doing what they're doing, and Europe is huge.
V
Vik Malyala18:43
Europe is also doing the Euronet. Europe is really going all in on open source with this huge project called the Euronet, focused on open source technologies, democratizing the technologies powering AI and cloud computing.
J
John Furrier19:00
It is certainly happening, but the difficult part is this. The cost of infrastructure is so expensive, the amount of power needed, the data centers that need to go along with it, the whole thing needs to come together, and investments need to happen from governments. In Europe, some countries are going forward with those investments and many are not. With a fixed budget, what do you prioritize? And no matter whether you like open source or not, you will be using it, but the priority will be getting the work done quicker, then comes the rest. For example, one of our key customers, I was talking to him about infrastructure and efficiencies, how to build liquid cooling systems. His comment was, 'I'm sold on physics, but can you give me the liquid cool systems one day prior to air cool?' I said no. Then he said, 'I want the air cool to go with it. When you're ready to provide liquid cooling at the same time as air, I'll go with it.'
Supply chain is going to drive so much.
V
Vik Malyala20:19
Business is driving it. We've been able to deliver a lot of capacity fairly quickly, but I think our ability to physically continue to deliver capacity at that scale is going to slow down. You can only build so many data centers so fast. So people will take things faster rather than better quality right now.
J
John Furrier20:42
You may believe it or not, but in the last two and a half days, I've had at least 10 meetings with data center providers. People who have power are trying to build data centers, or people with power and data centers are trying to refine the infrastructure to bring equipment home. They're trying to find out who to work with. Customers with dire need are trying to figure out which data centers to work with. These two groups are trying to figure out who the ISVs and end users are. The equation is getting quite complex, and deployment needs are increasing tremendously, so whatever we can do to speed up.
It's a supply chain on one hand, but also a stakeholder evolution because now you have different stakeholders aligning around resources like energy. I've got my real estate, but the grid is not good here, but I've got space. So new dimensional physical aspects.
V
Vik Malyala21:45
Perfectly stated. Five years ago, if I went to a data center and asked what they were doing, they'd say, 'Why are you asking me?' Now, if I don't ask, they won't talk to me because they know I don't know what I'm talking about. If someone says they're deploying a thousand node cluster, the first question I ask is, 'Where? Do you have the data center ready? Show me the data center contract. Who is the end user? Where is the money coming from? How are you planning to manage it?' It's everything.
J
John Furrier22:17
The biggest thing I keep hammering on is even if you have the real estate, you need the people. You need people who can pour the foundation, pull cables, service generators and chillers. That's one of the biggest impediments. If you're trying to build a data center in Kentucky...
V
Vik Malyala22:33
IT is getting sexy. That's all I've got.
J
John Furrier22:35
Yeah, we've loved it for years. Vik, it's great to have you on. I'm super glad you stopped by to wrap up our day two.
V
Vik Malyala22:40
Thank you.