About Neil Macdonald
At HPE Discover 2024, Neil MacDonald, Executive Vice President of Compute, HPC & AI at HPE, discussed the company's partnership with NVIDIA, announced as "NVIDIA AI Computing by HPE." MacDonald described the collaboration as a strategic relationship focused on accelerating enterprise deployment of generative AI. He stated that the technology is intended to change business processes by enabling systems to generate responses based on a company's historical data, such as past proposals, making sales and pursuit teams more productive. MacDonald said that enterprises that do not adapt to generative AI risk becoming obsolete.
MacDonald also addressed the technical requirements for enterprise AI, noting that power efficiency is critical as systems scale and that 100% liquid cooling will be necessary for handling multi-kilowatt power levels. He emphasized that the key to enterprise AI success is integrating compute, data, and models at the system level, rather than focusing on individual servers or accelerators. The partnership includes joint go-to-market campaigns and equipping customer briefing centers with the combined technology, with a detailed plan reviewed regularly by HPE CEO Antonio Neri and NVIDIA CEO Jensen Huang.
Source: AI-verified profile updated from Neil Macdonald's recent appearances.
Browse all interviews →
Transcript (28 segments)
J
John0:10
2024, I'm John for host CU with Dave Volante, my co-host co-founder. It's our 14th year covering HP and HP Discover. Got two great guests here unpacking the big news with Nvidia and HP, the strategic relationship, a deep relationship. Neil McDonald, Executive Vice President HPC and AI for HP, and Bob Petty, Vice President and General Manager of Enterprise Platforms at Nvidia. Gentlemen, thank you for coming on. I think we had a great preamble that wasn't on camera, but I think we had some good content coming out. Thanks for coming on.
N
Neil Macdonald0:39
Glad to be here.
J
John0:41
So I got to say, this is unprecedented times. The world's being changed by generative AI, the new platform shift. It's different tokens, neural networks, vectors. These are the terms that Jensen's talking about. Nims, whole another architectures here. And you guys announced the partnership between Nvidia and HP.
N
Neil Macdonald1:10
The announcement that we made this week is NVIDIA AI Computing by HP, and our focus is really accelerating the next wave of deployment of generative AI. We have this technology that's going to change the world and impact all kinds of processes inside businesses, because fundamentally what it's about is productivity. We believe that every enterprise is going to become, to some extent, driven by generative AI, or they're going to become obsolete because they're going to be beaten by the competitors who do so. The challenge is that as great as all of this amazing innovation and technology is right now, enterprises have different needs. There are lots of solutions out there that are really reference model stacks before they can ever get to the experimentation on use cases and capture the competitive benefits of productivity. So the wonderful thing that we're engaged on with our great partners here at Nvidia is delivering a co-developed portfolio of products and solutions that enable enterprises to accelerate that journey, because we take care of that for them. They can take these technologies, deploy them in a few clicks in less than 30 seconds, and get going on their journey of discovery and transformation leveraging generative AI.
J
John2:42
Bob, talk about the Nvidia piece, because we've been covering you guys for a while. And in your conference, the AI systems were front and center. I mean, it's a data center in a box, it's a huge monster machine. But the appetite for generative AI is not just in the broad bigger market, it goes down into the enterprise.
B
Bob Petty3:11
It will be gathered and obtained, and they require a different model. You know, there's data privacy things to consider, there's GDPR things that could be considered, data governance and the like. As you move from hyperscaler technology into the enterprise, you'll still do fine-tuning. And so our GPU and our CPU-GPU assortment covers, and what we announced here as part of the HP Private Cloud for AI, is a slew of product from small, medium, large, and extra large to address these different use cases. There'll be some training done, there'll be fine-tuning done where you take an existing LLM or small LLM and you fine-tune it for your particular industry or your particular corporation. Probably most importantly is where you get into retrieval augmented generation, AI co-pilots for code assist. But the Hopper systems and the Grace Hopper systems that we've announced here with HP are state-of-the-art. We'll be time to market with HP when we're shipping Blackwell next year. Jensen's already talked about Rubin in 2026, and we're looking forward to the next couple years with our partners here at HP. One-year cycle. You know, John points out, and John's the only person I know who times Jensen at events. Other than GTC where he spoke for hours, he was 25 minutes at the Sphere the other day, which is a new record, double roughly what you see at other shows. But the thing about Jensen is he always brings something new to these shows. We watch all his videos and it's like, wow, he's always got a new take.
J
John5:10
And this time, how does that fit into what you guys announced?
N
Neil Macdonald5:14
Well, any form of generative AI relies on data, and enterprise data is on-prem. It's on-prem for reasons of confidentiality or protecting IP. It's on-prem for regulatory reasons. It's on-prem for any number of concerns. And so any generative AI endeavor that's really going to transform business processes in the enterprise has a natural gravity to on-prem for part of that architecture. And as you know, we've long held that the world is going to be hybrid, and it's been borne out, and nowhere is it being borne out more than in generative AI. So the challenge then becomes, how do you make that easy for the enterprise to deal with? And that's where the three stacks come in. In the market today, what we've done is co-developed a turnkey integrated solution that integrates that compute stack, including the storage that's needed and the ability to ingest data across a disparate enterprise into gen AI initiatives leveraging our data fabrics. But it's more than that. You need to be able to observe this infrastructure, you need to be able to curate and lifecycle manage it, so we've done all of that too. But that's only the compute stack. When you think about the data that feeds into it, that's also part of what we've delivered. And most critically for the enterprise, it's about the models. So Bob, maybe you can add a little bit about the model stack and how we've integrated it.
B
Bob Petty7:10
GPU, but an integrated approach that looks at latency to data, latency between CPU and GPU. Neil covered that greatly. The data stack partnership, GreenLake for storage, all the great storage that we have integrated, Nemo Retriever. You know, the slowest horse in the race is the one that's going to determine the length of the race. And if that's a query that someone has and they need timely feedback, so knowing where the data is, making sure that it's guardrail protected, that's part of the data stack. All of that is built into NVIDIA AI Enterprise that they've incorporated into their stack. And then the model stack is really based around NVIDIA Inference Microservices. These essentially are the runtime for models. We have a slew of models covering a variety of enterprise use cases across multiple industries, which is why we're able to take multiple data sources and predict, and that's where more intelligence is extracted. So these NIMs, these containerized models, can come from anywhere. People can download them forever from wherever they are, but we have a nightly NIM AI Factory at Nvidia that we build, continue to build these models, tune, test, optimize. We have replicated all of the servers and the systems that are part of NVIDIA AI Computing by HP. We test them out there, so our enterprise customers, our collective customers, know they're getting the latest state-of-the-art models. And so models are the key. They need to get to the data and they need to run on the infrastructure that's going to let them truly run at performance for time to insight, time to value.
J
John9:10
There are other model opportunities out there, open source and whatnot, so that's important to point out. I want to bring this up because all the people we talk to in our research and our conversations, yeah, they use the proprietary and/or large open source model for things, some confidence, some query, some retrieval. But when you talk about their data, they want a simple and easy way to stand up infrastructure, kind of like the old cloud days. I want to do a SaaS app, I don't need to, I just spin up a credit card and I get some services. They want the same thing for their data because that's their crown jewel. Exactly. And so take us through what you guys see there, because that is different. It might be smaller, it might be different sizes, but that's a key part of their IP. And the workloads are end to end. So you know a couple things: I know my data and I know my workload. How does that factor into what you're doing?
N
Neil Macdonald10:10
To get the results you want in terms of quality, because it doesn't have the specific intelligence about your business or your industry. This is why for enterprises to be successful with generative AI, they have to pursue either fine-tuning a model based on their own data, which they're going to do on-prem, or leveraging retrieval augmented generation, where the generation process is leveraging that data from inside the enterprise. But to do that, enterprise data is very fragmented. And so that's why we've integrated into Private Cloud AI our data fabric that enables you to pull that together from those disparate sources in order to underpin that work. But that tuning or RAG or the development of small specialized models is not just about data from the internet. You want your data, your SAP system, Oracle system, your Salesforce system. You want all of your information to be a part of that model. And some of that is done via RAG, but a lot of it is done to an industry-specific model that is developed. We see this today in climate, in geospatial large language models, and financial large language models. And when I say large, I'm not talking about trillions of parameters, but billions or hundreds of billions of parameters. Those by industry you'll see that, and we are seeing that happen. To your point, that's the transformation. So I take that industry-specific one, then I make it enterprise-specific.
B
Bob Petty12:10
Big transformation, and that's part of what NVIDIA AI Enterprise, HP AI Essentials, and the NIMs are meant to do: help companies go from these truly large language models down to models that they're going to work with on a regular basis to improve business. You mentioned RAG, obviously retrieval augmented generation is what it means for the folks who don't know what it means. But that's a lot of text. There's also tokens can be used for multimodal, so you got images, PDFs, X-rays, I mean a lot of DNA, non-text. So that's an important point. But again, that's part of the new system.
J
John12:44
Which brings up my next question: can you guys highlight or illuminate the difference between old school data processing, which I could do in the cloud? By the way, cloud's great for data processing. Put data in the cloud and I can...
B
Bob Petty13:10
This gets back to this, I don't want to call it a soundbite, but it's really important. Instructional computing means you, the programmers, have an idea of what they're trying to find and they go through the code in an order of if-then-else to see what's there with the hopeful outcome of coming up with something important. And that's served us well today. Intentional computing goes well beyond that, right? And so it's not so much about how I'm going to crunch data, it's about what I'm trying to do. And AI learns that based on the questions that you're probing the model with, based on the data that you're exposing it to. It will not only answer your questions, but it understands what your second question is going to be. And we turn it on to a different modality and we're shocked. And so that's the true power: not just getting the answers that you want, but getting the answers that maybe you should be asking. Knowing that your intended goal is to save lives, improve the climate, improve the supply chain, increase sales, whatever it may be. And you think about an enterprise example: you got a B2B business, you're responding to requests for quotes, requests for proposals. In an instructional world, the way we would run all of our IT in the past, you'd try to come up with a perfect algorithm to write all of that. In an intentional world, you leverage HP Private Cloud AI to train your sales teams and your response teams, your pursuit teams, vastly more productive. So you can cover more opportunities. You've gotten the best practices out of what's worked for you in the past, the deals you've won, the deals you've lost, and you've factored that intelligence in a systematic way into the business. So that's transformational. If you're competing with another organization who's trying to do this all by hand, you're going to be faster to respond, you're going to be better at leveraging your best practice, you're going to win more. And that's the power of transformation. The challenge is how do you enable enterprises to do that without having to become gurus in the compute stack, the data stack, and the model stack, because the architecture is so different from traditional enterprise IT. And that's why we're so excited about what we're doing together with Nvidia.
J
John16:12
The reason why this is so important is the one-year cadence rhythm. I think Jensen calls it, right? H100s are becoming a little bit more plentiful, and then can't wait to get Blackwell. We already see Rubin announced. What does that mean from your standpoint in terms of managing the life cycles, which used to be two, maybe even extended tape-outs were stretching out to two to three, sometimes even four years? Now it's like this one-year rhythm. What does that mean for you?
N
Neil Macdonald16:36
Well, the first thing it means is aligning the amazing engineering teams at HP rigorously with Nvidia's cadence. And we're thrilled to have announced our first Gen 12 servers here in the HP ProLiant line, which are based on Nvidia Superchips. So what you can expect to see moving forward is incredibly tight alignment on that cadence. For the enterprise, we've thought of it like a server, it's a box. The architecture is not about the individual nodes anymore, it's not even about the individual accelerators anymore. It's about the whole system architecture that brings together all of the capabilities in that compute stack, the data stack, and the model stack that you need to be effective with generative AI. So another shift here, both for our partnership and our collaboration, but also for our enterprise customers, is thinking at that system level in a much more prominent, cohesive way.
B
Bob Petty17:47
Yeah, and I think what will tie these architectures together in a way that customers don't feel like they're going to be missing out is CUDA compatibility. You know, we're testing out all of the software, not just our models and CUDA, but also all the HP software. And from a lifecycle management standpoint, one of the reasons we partner with HP, besides their broad reach, now they've got financial options to make it easy for people to transition. You know, some people might want Rubin today. You always want the latest thing, but do you want to improve your business today or do you want to wait till 2026? And so it's important that they know the code's going to run, the code will run, it's compatible, will continue to optimize, and they know that there are financial options to not only expand. I doubt that they'll replace those Hopper systems in the room in time frame; they'll just add Rubin.
J
John18:54
Well, we had Moore's law, now we have Jensen's law. I actually stole that in Taiwan, our chart. It wasn't really our chart, you did your own version. It was a great chart. The reason I asked, and we understand, thank you. We understand the importance of systems, okay? But it does come back to that core, that GPU. We need bigger GPUs. And when you look at Blackwell and what you did with NVLink, and to keep that lead, everybody says, okay, the most valuable company in the world, they're going to get all this competition. How long is that lead? When you look at Blackwell and the things you did, I'd love to ask you if you can answer. Most people from Nvidia are super smart, smart guys, as we say in Boston. So maybe you can answer this, maybe you can't. But we noticed reading about it, learning about it, that you had a little trick in there. You dialed down the floating point to FP4, and then dials things back up. So you have a lot of flexibility there. What can you tell us about what's coming on that roadmap, without obviously giving away too much?
B
Bob Petty20:21
Yeah, I'm not going to give away too much. But you know, the level, whether you need INT4, INT8, FP8, FP16, FP32, depends on the workload. Our R&D group is doing dramatic results. We're getting what used to require FP16 now runs in INT4 at the same level of accuracy. This is important because now I can fit more data, larger models into a frame buffer size that I need to run that model. So I may not need the hyperscale level of compute. You need to reserve space for FP64, which we always had, or FP32 calculations we always have. When you get into the tensor world, the AI world, you want to try to squeeze as far as you can. If we could get to an INT1, an INT level, I'm not saying we would, but if we could, we would. Why wouldn't you? If we could make the same prediction in the same or faster time frame with less precision, using up less memory, you're going to do that. And we'll continue to innovate that way. Will there be some things that hope that some of that floating point was there? Yes, but they'll still run faster. It's just the AI will run 100 times faster.
N
Neil Macdonald22:11
And so, Neil, from our partners at Nvidia, we're thrilled to be there time to market with every one of these transitions. The key though is if we languish in the detail of exactly what's in a piece of silicon, we're having the wrong conversation with enterprises. Because we shouldn't be burdening customers with having to understand the whole technical detail of the stack. That's the power of our focus and our approach. They're just going to have access to more performant, more capable systems that lets them deploy a larger number of generative AI use cases inside the enterprise. We don't want them thinking about INT4. You got to extract that, but we still want to know.
B
Bob Petty22:53
But I think the other thing that we need to talk about though is that not just the architecture changing, but the number of bytes required, we're taking power up. It's just the way it's going to work. We need bigger and bigger power systems. You saw that from Hopper to Blackwell. Let's assume that that will continue. That doesn't mean that we're going to run out of power, because while we may be taking power up 2x or 3x, if the performance is going up well beyond that, it still is a more efficient system. So if we can improve Llama 3 70B 5x with our NIMs running on a system today at a certain power level, when that power goes up and you get that 5x improvement, it still is going to be saving you money. The more you buy, the more you save. So in the same power envelope, you're going to get more done. There's only so much space and so many data centers. And that's one of the main reasons we partner with HP on this announcement: they're experts in liquid cooling, whether it's 70% liquid cooled or direct liquid cooled. So we're confident in taking the power up because it's still going to be a profitable formula for our customers. They're still going to save money, and especially with the liquid cooling piece, they'll save even more from a power standpoint. You guys can reuse that heat that's pulled out and heat buildings in winters and the like. So we're confident that people say GPUs are power hungry. I think customers are information hungry, intelligence hungry, and so we're feeding that.
N
Neil Macdonald25:11
We have all of the experience from decades of Cray building 100% liquid cooled systems. You can see them here on the show floor at Discover with the Venado system at Los Alamos, which Jensen and Antonio inaugurated just a few short weeks ago. So we're really thrilled that all of that experience is coming to bear right at the heart of this enterprise generative AI journey. Dare I say democratization of supercomputing? But I won't say it. I just said it.
J
John25:37
Well, I think that last conversation about getting into the details proves that the engineering's different, but by design has to be different. We're in a different world. You guys take care of that on behalf of the customer. So I have to ask, in our last few minutes, on the go-to-market together, this is a significant part of the relationship. What's the compensation? Who's getting paid? You guys holding hands together as a joint solution? What's the go-to-market?
N
Neil Macdonald26:16
So the first thing that we're doing is we're training and certifying our entire enterprise salesforce on Nvidia Technologies as part of NVIDIA AI Computing by HP, so that we can go bring that benefit of Private Cloud AI to all of our enterprise customers. But we're not stopping there. We're also enabling certification and training in our channel community, which is such a core set of partners for both of us in expanding reach and being able to bring these competitive benefits in a democratized way to the broad sweep of enterprises. And then of course we're working with Nvidia in the field.
B
Bob Petty26:49
Yeah, we start with the product group. So my product group is actually just focused on the enterprise and just focused on the OEMs because of that. Jensen sent this morning, and he said, 'I want a tiger team pulled together.' He sent it to our head of sales, myself, a couple of my peers, and said, 'Let's get a tiger team together and figure out the next 90 days. I want to know every go-to-market activity.' And so he's a genius at pulling that together and holding us accountable. There'll be joint go-to-market campaigns that we do together. We'll be making sure that their customer briefing centers are fully equipped with our technology. Our sales force, which is much, much smaller than their sales force, we'll be working with them on probably 1000 accounts. But it'll be a detailed plan that Jensen wants to view on a regular basis, and I'm sure with Antonio.
J
John27:54
That's great. Being in the valley for 25 years, I can tell you, a person who fits that. Congratulations on the relationship.
N
Neil Macdonald28:14
Well, we're excited to be working with the best partner in the space and excited to bring that value to our customers.
J
John28:20
And we're timing the CEOs on keynotes. We see how they do around the track. You know, they're tech athletes. We want to see their speed date. We're going to track it all. Right, you watch the Cube here. Bring all the Nvidia and HP and a really monumental partnership. We're going to keep tracking here on the Cube. We're bringing all you the data here. We got our own NIM. It's called the Cube Data. Stay with us for more after this.