Vik Malyala2:57
Good afternoon. Exciting times that we are in. I have been here so many times to attend Computex, and this time it is completely different. It is filled with so much technology, mostly around AI, mostly around how we bring the infrastructure together, and how things can be handled better. We are no exception. The fact of the matter is last year our CEO presented the thought process and how things can be accelerated by bringing a building block approach into infrastructure. First things first, if I take a look at what Supermicro has been doing, we are the largest and the fastest in terms of the AI infrastructure for an OEM. And this primarily happened because of the growth that happened with respect to AI. We are one of the only companies that is totally vertically integrated in order for us to conceptualize the design of the systems as a building block. That way we can expand and be able to bring the latest and greatest technology quickly in a more operable and scalable model. As a company, we have the US as a headquarters, but we have a significant presence in Taiwan. We have about half of our strength here in Taiwan for engineering, manufacturing, operations, logistics, and everything in between. That is because here we have a lot of partner ecosystem that we closely collaborate with, which is essential for us to leverage that incredible talent to bring the technology out to the customers. So, I sincerely appreciate all the technology and the partners that are actually contributing for our success here.
The growth also poses challenges. What we have also seen is that as we bring these products and start deploying, the scale is changing. It is something that people are trying to use at the edge, probably a rack or two, versus a frontier model or a larger data center that is deploying thousands or tens of thousands or even hundreds of thousands of systems. What ends up happening is there's a severe shortage of data centers. Why would it matter? Data centers have been there for a long time, and data centers are going to be there for a long time. But the difference is back in the days when people are looking at 10 to 15 kilowatts per rack, network probably a gig or 10 gig if you are lucky. And some data centers they are talking about 100 to 200 megawatts. That was the case. Nowadays when somebody talks about having a 100 megawatt data center, it doesn't really move the needle much because of the sheer amount of compute that is needed to take care of the insatiable appetite we have for AI. They're talking about around 1 and a half trillion dollars of IT infrastructure spend. 1 and a half trillion dollars, which basically means it could be a conservative number, it could be an aggressive number because I have seen numbers going three or four trillion also, but the fact is that it's a big number that we all are exposed to. So it's something that we cannot ignore. That's number one. Based on public information on how many data centers already kicked off, we're talking about roughly 190 gigawatts of power that is needed in order for us to support this 1 and a half trillion dollars worth of spend. It's not easy. 190 gigawatts. The other aspect of it is we are bringing technology very quickly and we are deploying it. And ultimately people are going to use it to do something with their applications. So the whole concept of AI factory and the tokens and how the tokens are used is the one that is driving it. Unfortunately, the range is quite large. We're talking about some 40 cents to some 25 dollars per million tokens, which is going to hurt because in order for many people to use it, AI needs to be cheaper. It needs to be democratized. It's a function of the infrastructure, the technology that is being used, the software that is being used, how the overall thing is going to come together. Great. We can actually speed up and start deploying these data centers quickly and start deploying the models quickly. Problem is, you take a look at our friends like Nvidia GB300 to Vera Rubin, for example, or in the case of AMD you have MI350 series to Helios that is coming very soon. And they're available in HGX platform as well as a rack scale. When you combine these things together, the technology refresh is less than a year, less than a year. But a typical data center takes 12 to 24 months if you are lucky, assuming that nothing else is blocking it. This is where the problem fundamentally lies. By the time you initiate a data center project and you finish it, the technology changes are changing which may or already going to impact this data center in a way that it will be outdated. This is the problem. That is part of the reason why the industry leaders like Jensen and Lisa Su are telling this is how my platform is going to be two years down the road, so the data centers can actually start preparing for that kind of thing. So, what are the things that are impacting it?
So, obviously we need to have enough energy. And with the energy doesn't mean that a land and then some gas pipe or something. So, we're talking about either from the grid or from using fuel cells. It could be renewable energies. All of these things need to come together for the data centers to run. And one has to be able to have sufficient power in a predictable manner in order for them to move forward with. Land. When you're talking about smaller to larger data centers, smaller ones I have talked to many customers, they were talking about 5 MW, 6 MW kind of data centers distributed. But if you're talking about a giga scale, where are we going to find that contiguous chunk of land that is going to be able to host these data centers? Supply chain is going to be another headache. And the time it takes, as I mentioned, last but not least, in terms of the cooling, because data centers are going to run hotter and hotter with the technology that we are putting in them, how are we going to make them run efficiently? These are the very basic principles, right? Then, you take a look at a data center. How do you define it? For some of them it could be a single system that is being used to provide tokens for whatever they are doing. So, small enterprises, for example. Versus many of these things put together in a rack or a pod or a super pod, that is what many of the neo clouds are using today. But when you take a look at a larger deployment, let's say, if people are looking for tens of thousands of systems, it could be a data hall. But when multiple data halls need to be stitched together in order to form the entire campus. This is how the spread of these data centers is today. And every one of them is technically an AI factory. And how do we make sure all these things are going to operate? That is another tricky part that one has to worry about. The other aspect of it is Nvidia is a de facto standard for pretty much anything with training and inferencing today, but others are not too far behind. You have AMD coming up with the AMD Instinct GPUs. We have Intel coming up with their own GPUs, you have heard probably from Li Bu the other day. Arm has Arm Neoverse CPU primarily used for AI specific acceleration, and there are at least half a dozen, if not more, accelerator companies that are coming up primarily for inferencing, but overall, when you take a look at it, we have several of these technologies that are coming in. So, you cannot design just for one technology or one particular product because it's going to change. So, you need to be ready for all these things as well. Like this is not enough, the other thing that one has to be concerned about is sovereign AI because one has to worry about how do we make sure whatever data that is getting trained on is staying within a given region. There's nothing new in that. We have done sovereign primarily for data in the past, but now AI applications are also going to be part of it. Every one of them, depending on the budget, depending on the region, depending on what kind of technology they want to use, those data centers are going to look different as well. And cooling is becoming an incredibly important thing. And cooling happens to be important because as we bring denser and denser solutions together, air cooling is not going to be sufficient. And liquid cooling is coming into the picture. That needs to be at the silicon level versus at a rack level where you are having CDUs and PDUs and whatnot. Then, you're also looking at what happens if the CDU needs to be in row, in rack, or in the data center. And as you look at it even further out, you're talking about the chillers that are outside. Could be dry or wet or anything in between. These are all the things that need to be looked at as well based on, again, where it's going to be located.
Power delivery is becoming another incredibly difficult thing for people to handle, mainly because there is not enough power available. Even, let's say you have a hydroelectric plant, you need to update your grid in order for you to get the electricity to where you have enough land available. Because otherwise you have to take your data center towards wherever the power source is, which sometimes happens – you bring gas lines to that and put your own generators and whatnot. So, these are the type of things from a data center perspective – the power coming into the data center, but once you get into the data center, you have several things, right? You're talking about the UPS systems. You're talking about the battery backup units within a rack, and you are also talking about super caps. And lately, if you have seen people who are handling the large language models at larger scale, they're struggling with yet another thing, which is the spikes. In order to handle the spikes, you need to have a larger battery pack sitting outside. Each of them having their own strengths and limitations in terms of capacity, latency, and how quickly it responds to a request and whatnot. And within a data center, networking used to be always an afterthought. With AI, that is actually taking the forefront, especially as you start handling larger and larger models. The east-west traffic, as we call it, in order to handle that, we have InfiniBand or RoCE that is primarily being used as a fabric and whatnot. But unless you know exactly how the performance is going to tune for a size of the cluster, it's going to be difficult to go and take either or. This is also another reason where people struggle. You know, do I need to go for rail optimized? Do I need to go for a factory approach because I'm coming from an HPC background and things like that? That's yet another variant that people are playing with. But having that understanding and having a way to modularize it is something that's extremely important within the data center as we look into that. So, with all these complexities, what is it that one can do? A data center building block.
Just as I mentioned about a system when we are designing, we are looking at what are all the different pieces that are put together to make the system – like a power supply, the fans, the add-on cards, the CPU, GPU, whatever it may be. Similarly, if you take a look at the entire data center, now a data center is nothing but one single large computer. And within that you have interconnects, you have CPUs, GPUs, storage, all these things together. So, the idea what we are looking at is by making it modular, by making it scalable, by making it predictable, one can accelerate the deployment of these data centers. Easier said than done, but it is something that is needed in order for us to speed things up. Otherwise, any single element missing in action, you're losing the entire project and everything gets delayed. So, why? If you were to just basically compare a traditional data center build versus what it would be if you were looking at the data center building block, it's pretty trivial. Time to market is one of the main advantages of going with a data center building block. And just to kind of shed some light on it, it basically means that for every piece of these different elements, which I will go through in a sec, we will have one or more options. And we want to make sure that all these things are going to operate well by interoperability testing. And depending on which one has what lead time and what is the size of the data center, where it's going to be deployed, we can plug these things together and in theory, it should work seamlessly because these are all going to scale as a building block.
From Supermicro's point of view, we started looking at it from the system level. So, given the fact that we have an extensive product portfolio that is based on either Intel or AMD, now Nvidia and even Arm as a CPU, right? A single socket, dual socket, different form factors, optimized for storage, optimized for compute and whatnot. So, in any data center, the very basic element is that we need to have this compute and we need to have the network and we need to have storage and whatnot to be put together. If you were to look at one step above that, then things will be slightly different. We do need to take this into consideration because if the systems that we are making cannot be accommodated in a data center, then all the effort, no matter how good your technology is, it is useless. Liquid cooling, as I mentioned, is an integral part of what we are trying to do. So, the liquid, which is often either PG 25 or more innovations are coming, including even from Supermicro, where we are bringing new solutions that have much higher impedance, which actually helps in case there is a leak or something. Then, we also need to be looking at the rear door heat exchangers, mainly because not everyone has the very clear hot and cold aisle containment. We need to be looking at that. And in-rack or in-row cooling is something that we also need to be looking at for the different types of racks, how the CDUs need to be done. And for those data centers that do not have liquid cooling plumbing, then you still want to be able to use this infrastructure. So, we have a liquid-to-air type of solution, which actually helps remote data centers to be enabled for this. And you need to be able to also handle the overall infrastructure to be put together in a cohesive manner. So which includes rack integration and whatnot. One needs to be looking at heterogeneous or homogeneous compute and how we put together within the same thing. And as we step outside the data center, we're talking about either liquid chillers based on water or dry coolers. Last but not least, as I mentioned, for the energy to be stored in a battery pack, which is becoming extremely important for very large scale – think of frontier models to be run for training especially. This is what we call a battery energy storage system or BESS, right? So these are only a subset of what we're talking about obviously, but the idea is to continue to enhance and expand this portfolio, to validate and eventually be able to help speed up these things. And when you stitch all these different elements together, how do we know it's going to work? And let's say if everything is working, if something were to fail, then how do you know what is failing? Software becomes an important factor by taking all the telemetry and bringing it into a single pane of glass for people to look at. And for those larger organizations that already have a tool available, then we expose that through APIs for people to look at. None of this is rocket science, but somebody has to do it. So the idea again here is for us, we are moving out of the server only mode, server to racks to data centers. So that way when we engage with customers, we are able to understand what their pain points are and how we are able to address that.
Case in point, one of the largest and first deployments that we have done is actually with xAI Colossus, the 100,000 plus GPUs. The entire project got completed in 122 days. In a typical data center, if someone were to start from scratch and deploy, it's going to take probably 2 to 3 years for that size. But I cannot take the full credit for it, other than the fact that we worked closely with them in order to bring the AI infrastructure. But we were able to bring the liquid cooling platforms very quickly. And because of the building block approach, we were able to develop that and bring it at a rack scale and roll these racks into the data center as they start doing the project planning to make these things happen. This is how we were able to quickly deliver that kind of large scale. And since then, we have done a lot more. Case in point, once we understand the requirements from the customers and what needs to be done with respect to the data center design, then we can see what are the different elements that we can help with through this building block approach and work with our partner ecosystem to bring the right pieces of the puzzle together and be able to accelerate the deployment. Besides that, I also want to bring one of our partners on stage to talk about what they are doing as a part of our data center building block solutions to bring value to our customer base. With that, S.H. Lee, if you can get on the stage, please.
Thank you for being here.