About Mark Papermaster
Mark Papermaster, Chief Technology Officer and Executive Vice President of Technology & Engineering at AMD, spoke at the RAISE Summit 2026 in July, discussing the company's strategy for AI compute and system-level optimization. He described a shift toward agentic AI workflows, stating that AMD is using such workflows to "shave months off of our chip design schedule" and bring new features to market in "weeks and days." Papermaster attributed AMD's competitive position to a culture he described as "a scrappy underdog, a fighter" combined with "a culture of execution." He also said that "the days of the homogeneous data center" are over, and that enterprises can no longer rely on a single vendor for computing engines.
Papermaster highlighted AMD's acquisition of ZT Systems and its focus on optimizing entire clusters of CPU, GPU, and networking hardware at the rack level. He noted that the company is working to make AI more economical for enterprises by offering solutions that span from cloud data centers to edge devices, including embedded neural processors in PCs. He also previewed AMD's upcoming "Advancing AI" event, where he said the company would provide details on its Helios rack and new 2-nanometer "Venice" CPUs.
Source: AI-verified profile updated from Mark Papermaster's recent appearances.
Browse all interviews →
Transcript (61 segments)
J
John Furrier0:00
The connection, Silicon Valley and Wall Street. I'm John Furrier, host of the Cube here with Dave Vellante, my co-host.
Welcome to the Rave Summit coverage, brought to you by Solidigm, with support from Tentative D Matrix and Ardentum.
Welcome back to the Cube's live stream in Paris, France for the Rave Summit. I'm John Furrier, host of the Cube. Day one kicking off day zero side car event yesterday with Machine as well. Physical AI and robotics, which is going to lead into hot year in robotics. But that's all part of the power and the computing and the software efficiencies coming from all these new integrated systems. This next guest is a legend in the industry, the conscious engineering brain for AMD, 40-year history, Mark Papermaster, CTO and EVP with AMD. Thanks for coming on the Cube. Good to see you again.
M
Mark Papermaster0:52
John, great to see you.
J
John Furrier0:52
Cube alumni back. I said you're the conscious brain of AMD because the engineering mindset of AMD has been one of the hallmarks of the company. And there's so much you guys are doing. It's so broad and in the spectrum of what capabilities, but performance has been a very huge deal. So thank you for what you do. We appreciate more horsepower, go faster. What's your current take right now because right now the demand is so high for performance, but it's not the way it was last decade, last generation. It's transitioning to a new architecture, new requirements. We're paving new roads, we're pouring new concrete in AI infrastructure. What's your take on all this?
M
Mark Papermaster1:35
Well, I love your analogy. We're paving new roads. John, so true. Because look, the need for high performance, that's actually not new. So I'm four decades plus in the industry, so that means four decades plus driving more and more performance and performance at efficiency across the computer infrastructure. So that's been my whole working career. But it is a new road now because generative AI is changing everything. It's such a diverse workload, John, that is driving us at AMD and certainly the whole industry to think, design, architect, and implement much more holistically than we ever have. And what I mean by that is the workloads are so complex because people are looking at what they do end to end. They're looking at whole processes, not just one bespoke task. And that means you need different computing engines, and they need to work together. And they need to work together at scale. I mean, we're talking across massive clusters of racks.
J
John Furrier2:37
I love that analogy of like just focusing here, now we're looking at a larger scale because at the end of the day, it's all a system. So, if you have a system mindset, what's changed the most for you as you evaluate efficiency and some of the constraints and managing through those because when you look at the broader workloads that are running on a gigantic, there are use cases, there are new requirements. What are some of those system elements that need to be in place as first principles? That's different than building a motherboard and putting software on it, accessing it. Now we got and then putting it into a data center, then putting it in the cloud. What's different now from a system standpoint from a first principles?
M
Mark Papermaster3:18
You know, when you look at data center, the movement of the data needs to be so efficient because it's always been moving data from memory into that beast, that processor core, moving it to IO into that processor core as efficiently as you can. But, again, what I said a moment ago, it's heterogeneous now. So now you have to think about that data movement and optimizing that performance across CPU, across GPU, and across the network. And all of that to a memory complex, you need low latency to memory. You need to get out to storage. And so, what's changed is not the local optimization on each of the elements. Still have to do all of that. But you have to, as you said, at a system level, really optimize. And that's what we've changed at AMD. We've become system optimizers. Still optimizing every component, but then equally looking at how you optimize how each of the pieces come together. And in that vein, we've expanded our portfolio. So look at what we've done over the last few years, right? We added networking capability. We acquired Xilinx. We acquired Pensando. And then we acquired ZT Systems that brought a whole rack design capability. So for the data center design, we're optimizing at the rack level, hardware and software.
J
John Furrier4:42
I have to ask you this question because it comes up a lot. Okay, I got this large-scale system. I want to put this AI factory or refinery into the system. Great. I plug it in. I put it into a data center. Great. But whoa, I got all this other stuff. So backwards compatibility is a big discussion that's come up a lot on the queue, but also backwards compatibility with existing workloads. So I have a lot of AMD, but I might want to run at the edge. I'm going to have distributed computing architecture. Now I'm looking at a cost envelope and a performance envelope. So I might look at like I might get some performance here, but if I look at the big picture, the performance cost ratio changes. A lot of these infrastructures have a lot of AMD, a lot of x86.
M
Mark Papermaster5:23
That's right.
J
John Furrier5:24
What's your vision on that? Because this has come up a lot.
M
Mark Papermaster5:26
Yes. It's coming up in you know, you say a lot. I'll say for me, every enterprise conversation, John, I'm having now, this comes up. Because enterprises have decades they've been running x86. They're not going to move that install base. And what we've done at AMD, you know, since we launched the new Zen processor back in 2017. We're now on our sixth generation. So at our Advancing AI event on July 22nd and 23rd, we're rolling out this new generation. You know, it continues the kind of leadership x86 CPU, but it's designed in such a way that it matches what I just described a moment ago. It's optimized for standalone x86 traditional workloads. But what about when you're running heterogeneous workloads that you need a GPU, but then you still have your traditional x86 apps? Helios has that sixth generation epic processor built in. And so the very control of the Helios rack with our new MI255 that we're again, we'll be sharing more details in a couple weeks. We'll be there, by the way. We're streaming live from your event.
J
John Furrier6:35
So we'll be there. Let's get into the price performance because you're what you're getting into is obviously your efficiency your career does that. Performance is obviously great, but the cost also is important. Price performance has always been a benchmark. Let's map that to the current environment. So x86 factors into the cost. What about when you look at things like, okay, agentic workloads spanning to the edge and inference needs that are different in scope? So I might say, 'What's the weather like in Palo Alto?'
M
Mark Papermaster7:01
Of course.
J
John Furrier7:02
That's different than saying, 'Hey, solve my high-end math problem and write a like this new thesis and solve these big problems.' They have different levels of differentiated services. So I'd love to have that route in the right areas. How does that impact you guys because that's coming up a lot, especially policies policy-based resources is becoming a big
M
Mark Papermaster7:24
It is.
J
John Furrier7:25
cost factor.
M
Mark Papermaster7:27
Yeah, no, you're spot-on. And that's what we've architected in our stack. So our AI stack's open. It's called ROCm. It's the same ROCm that we run on those massive clusters. When you're running AI at the edge, you're running AI locally on a PC like you have in front of you there, it's also ROCm. And so if you think about what we're doing, we're making it very easy to port. Now what you have to do is you have to orchestrate your biggest task, a massive inference model that's running on a foundation model, that's running in the cloud or a big data center. But most enterprises, that's very expensive if you run everything there. So they're exactly on the point you made, John. They're looking to run that more economically and often at the edge it has to be done locally because you need real-time response. And so we've done that for not only our CPU and GPU, but the embedded neural processors that we have on the PCs and also in the embedded edge.
J
John Furrier8:25
flexibility
M
Mark Papermaster8:26
More flexibility.
J
John Furrier8:27
All right. So I want to bring something else up. It's not related to the question I'm going to ask, but it's a way to tease out the question. So if you look at what's going on in China, they get constrained GPUs.
M
Mark Papermaster8:36
Mhm.
J
John Furrier8:36
And so they're smart engineers too out there. And they do is they actually figure out how to run better, faster
M
Mark Papermaster8:41
They cluster it together.
J
John Furrier8:43
All entrepreneurial engineers will figure out how to work around constraints. So their constraint I don't have the big GPUs. So what do they do? They engineer better solutions to run on non-GPUs.
M
Mark Papermaster8:52
Mhm.
J
John Furrier8:53
I think that's going to be something that's
M
Mark Papermaster8:54
or even just lower capability GPUs.
J
John Furrier8:58
So I think that's a tell sign to how the engineering entrepreneurial thinking's going to happen. The constraints are going to be in memory. We're going to see some constraints around cost and deployment savvy IQ. I think this is a topic. What's your thoughts on this engineering around constraints as they will come up because we're still at the beginning of a new era. What are some of those thoughts around what will come because an edge box might not have a big footprint than a big monster data center.
M
Mark Papermaster9:29
Right.
J
John Furrier9:29
So, but they'll talk to each other.
M
Mark Papermaster9:30
Right.
J
John Furrier9:31
But tokens will be running involved.
M
Mark Papermaster9:32
Yeah, exactly. But that's where modularity comes into play. We've designed our portfolio. We have the broadest portfolio in industry because of those acquisitions I described and because of our history across data center, PCs, and embedded. And so the problem you described that's exactly been our focus. Leverage the broad portfolio and enable it. So constraints are the mother of invention. And what we've done at take our big MI Instinct GPUs. Those run in massive clusters. We power the most powerful supercomputers in the world. But we now scale that down into the new MI 350P as a PCIE card you can plug right into an X86 server. So that was necessity that drove that. Our customer said, 'Look, I can't run all my workloads on these massive clusters. You know, I'm a mid-sized enterprise. What do you got for me?' And so we do have that solution. And when we take that then across to the neural processor right into the PC, right into the edge. So it is in fact that customer demand driving invention. And what we've done is created a modularity into our portfolio to enable it.
J
John Furrier10:50
I love having you on the Cube with all your experience and your leadership and knowledge. So I have to ask you. It's like having a consultant sitting at the table with me. Because I'm seeing this is a question I'm seeing. You guys are building as much capability to manufacture semiconductors and chips and software as fast as possible. The demand curve is undeniable. People see the demand curve.
M
Mark Papermaster11:09
Yeah, so you're doing your job. And the speed of building is a goal.
J
John Furrier11:13
Mhm.
Outside of your customers are trying to build large-scale infrastructure, housing, data centers, footprint as fast as possible. So there's a demand there. And so people are trying to go faster, your customers. And there's this new model emerging. When we just talked to Cerebras, he's got a chip, big chip. I talked to a hot startup, Argentum. They're just trying to build. So you have this kind of general contractor mindset. The demand is so big, they can stockpile the demand and just get the best supply possible. So there's a lot of builders.
M
Mark Papermaster11:46
Mhm.
J
John Furrier11:46
And they're building large-scale systems.
M
Mark Papermaster11:49
Yes.
J
John Furrier11:49
Fast. And they want the best product.
M
Mark Papermaster12:05
Sure.
J
John Furrier12:06
The at large scale. And they got to move fast. What's your advice to them?
M
Mark Papermaster12:10
Well, my advice to all of the customers is really understand the problem you're solving. You know, if you look at, if you need a wafer scale, you might need you might be using that because you need real low latency. You might have vibe coding you're doing where you just want that immediate response. You might be most interested in total cost of ownership, just running most efficiently. And more than likely these days you have a mixture of workflows. You need multiple optimized systems. So John, the days of the homogeneous data center or homogeneous computer or CIO could just say, 'I'm going to just populate my computer with one engine, you know, one vendor.' That's gone.
J
John Furrier12:54
Yeah, and I saw an example where someone was I won't say the name, but they were kind of complaining to me or crying to me about Yeah, we built spent all this money for training.
M
Mark Papermaster13:02
Yeah.
M
Mark Papermaster13:03
Now it's all moving in front of you.
J
John Furrier13:04
We did that already. It's not a lot of Well, there's still demand for training, but not at the scale of inference.
M
Mark Papermaster13:09
Right.
J
John Furrier13:09
And inference is where the token economy is going to live.
M
Mark Papermaster13:12
Right. No, you need both and we support both, training and inference. The MI455 cluster we're coming out it is a massive training cluster, but the demand for inference we saw was taking off. So we optimized configurations that really provide efficient inferencing. And if you needed a density, that's the big cluster. But then again, with our modularity, you know, you can run that on premise, you can run it air-cooled.
J
John Furrier13:39
Modularity and the X86, I think the X86 is going to be a big factor for what I call inference engines that'll be deployed. Okay, what's the coolest thing you're working on now? You got the big event this month. And going to have a lot of stuff, but can you share what's exciting you these days? I mean, a lot of I mean, it's a great time to be an engineer. I mean, this I mean I pinch myself sometimes.
M
Mark Papermaster14:01
Well, what I have to tell you is the last, I'll call it, 6-7 months what we've seen is agentic AI is transforming how we, even at AMD, how we're running engineering. And it's going to transform every enterprise. So we're looking at every process we have. We're saying, 'Look, we can create agentic workflows. We can tag these together. We're shaving months off of our chip design schedule. We're going to be able to get to market faster.' And our story applies to the kind of chip design we do, but it's going to be repeated at every enterprise. It's causing everybody to rethink how they do what they do.
J
John Furrier14:40
Yeah. I mean, I think that everyone wants you to go faster. We love what you do. We're paving new roads. We're pouring new concrete. I think the AI infrastructure is just beginning. I mean, if you look at the 5G buildout alone, that was close to a trillion. If you look at all the telecom, trillions and trillions of dollars. The upside for AI infrastructure is incredible. And again, it's a shift. The shift has happened.
M
Mark Papermaster15:01
The shift is happening right now.
And again, we saw this coming. We have the highest density of CPUs, the most efficient CPUs, the higher performance CPUs, and they're X86. They run your applications.
J
John Furrier15:15
[laughter]
M
Mark Papermaster15:15
And we complement that with the GPUs, our network infrastructure, and our rack optimizations through PCs, through embedded, like John, we could not be more excited. We feel like we intercepted the right products, optimized products, right as this wave is completely taking off.
J
John Furrier15:32
Well, we appreciate you. Thanks for coming on the Cube. Thanks for spending the time on your busy schedule. CTO and EVP at AMD, efficiency, price performance, but the new architecture, the new workloads are setting a new architecture for the infrastructure and then the applications that are named, we're hearing it on stage, lovable, replit, all the applications, coding, you name it. And we didn't even get to sovereignty yet, but lot more of that coming here. I'm John Furrier, your host of the Cube. Thanks for watching.