Back
Marc Whitten
Chief Product & Technology Officer of Create, Unity Software

Unlocking the Power of AI in Gaming with Unity’s Marc Whitten

🎥 Aug 14, 2023 📺 ARKInvest
Hosts Andrew Kim and Nicholas Grous venture into the world of artificial intelligence (AI) in the gaming sector with Marc Whitten, the President of Create at Unity. They discuss AI's role in creating vivid, immersive 3D experiences. Marc gives them a behind-the-scenes look at the complexities and opportunities involved in deploying AI for content creation in games. They also explore Unity's mission to make game design accessible to all. The conversation doesn't stop at gaming; they also cast a wider look at how real-time 3D technology is making waves across various industries. So join Andrew,...
Watch on YouTube

About Marc Whitten

Marc Whitten, President of Create at Unity, addressed the company's pricing changes and community backlash in a September 2023 Q&A. He described the revised pricing model, which includes a 2.5% revenue share and seat license options, as a "better program" that resulted from community feedback. Whitten stated that the changes were intended to establish a "sustainable business" for Unity and that the company was committed to transparency, including making terms visible on GitHub. He acknowledged that "trust is easy to lose and hard to earn" and said Unity would seek ongoing feedback through surveys and direct engagement. In an August 2023 podcast, Whitten discussed Unity's AI initiatives, including Unity Muse, a platform for integrating AI into real-time 3D creation. He described AI's potential to reduce development time from "months to minutes" and to enable more dynamic, interactive NPCs and environments. Whitten also highlighted Unity's collaboration with Apple on spatial computing and VisionOS, aiming to create cross-platform 3D experiences. He noted that AI's expansion into industries such as manufacturing and medical would reshape how complex systems are designed and simulated.

Source: AI-verified profile updated from Marc Whitten's recent appearances. Browse all interviews →

Transcript (34 segments)
N
Narrator0:12
This show offers an intellectual discussion on technologically enabled disruption, because investing in innovation starts with understanding it. To learn more, visit arc-invest.com.
Arc Invest is a registered investment advisor focused on investing in disruptive innovation. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. It does not constitute, either explicitly or implicitly, any provision of services or products by Arc. All statements made regarding companies or securities are strictly beliefs and points of view held by Arc or podcast guests and are not endorsements or recommendations by Arc to buy, sell, or hold any security. Clients of Arc Investment Management may maintain positions in the securities discussed.
N
Nick Gruse1:10
Internet and fintech, and I'm joined by Nick Gruse, associate portfolio manager. Today we have the great privilege of speaking with Marc Whitten, president of Create at Unity. Hi Marc, thanks so much for your time today.
M
Marc Whitten1:25
Yeah, for sure. Happy to be here.
N
Nick Gruse1:27
And for our listeners, we think it'd be great if you can just start with an introduction and how you ended up where you are today.
M
Marc Whitten1:33
So I'm Marc Whitten, and as you said, I'm the president of Create at Unity. What Create is is all the tools and services that we provide for creators to build in real-time 3D, whether that's the Unity engine and editor, whether it's services like Unity Gaming Services, whether it's tools like Weta Digital, so that people can... I spent the majority of my career at Microsoft, where I was one of the first employees on Xbox and spent 15 years building the original Xbox, Xbox 360, Xbox Live, and Xbox One on the platform side, sort of running a big chunk of those businesses. So I spent a lot of time around games, a lot of time around 3D and developers and creators. And as far as what brought me to the company, I knew our CEO quite well from my time in Xbox, and we had about a 20-second conversation on the mission of Unity, which is we believe the world is better with more creators in it. And this opportunity that they saw was obvious to me, something that I care super, super deeply about. But then as you got closer to...
N
Nick Gruse3:11
And obviously Unity has been at the forefront of AI. In the past couple of months, that's obviously on the minds of many, many people. AI everywhere, people talking about AI a lot. And you guys talked about Muse and Sentis as the primary AI products. For our listeners, I think it'd be great if you can just provide us with an introduction. Let's start with Muse. What does it do today and what use cases do you aim to support in the future?
M
Marc Whitten3:47
Yeah, well let me give you a top-line introduction to both because they're related, and then we can talk about Muse in a little bit more detail. Because frankly, it starts with the belief that every aspect of a real-time 3D experience is going to be touched by AI. At create time, creators using it deeply inside of their workflow, but also at runtime, actual inference happening on every pixel being displayed in every frame of a game or other real-time 3D experience. And so as we started thinking about that, it became clear to us that what we needed to deliver was a platform approach for each half of that. So how do we enable great AI workflows at create time to make creators more productive, to give them more capabilities, and make it easier for more creators to create? And then also, how do we unlock the real power of AI at runtime?
N
Nick Gruse5:10
When you say we're at a certain point today in this journey of AI being a part of every aspect of a game, which of the two segments poses the most challenge? And also, where are we in that journey? AI has had this huge run-up in the last six to eight months. It feels like it's starting to seep into every sector, gaming being one of those. But there's so much innovation happening so fast, it almost feels like we've made so much progress. But I'm curious, where are we in this journey and where are the challenges ahead?
M
Marc Whitten5:50
Those are great questions. And it's funny, I was sitting there trying to say, well, I don't know which one is actually the bigger challenge because they're kind of different. A lot of the energy and investment is being spent on the create side. When people talk about generative AI today, they generally are talking about how to use AI tools to create assets or help with coding or whatever. And I think when we talk about Unity Muse, people say, 'Oh yeah, that's exactly what I expect Unity would be doing.' When we talk about Unity Sentis, they say, 'Oh my god, I didn't even think about this being the implication of artificial intelligence as it applies to games and other real-time 3D experiences.' As for where we are, the first thing I'd say is AI has been used in a lot of ways. We've used AI inside of Unity for a lot of things. But the availability of large language models and generative models really brought generative AI to the forefront. But in another sense, especially on the create time side, this is kind of a continuum from manual placing of pixels through procedural tools through full inference-based creative tools. It's a little more analog than digital, going between those states. A good example of a non-AI tool today at Unity that serves a lot of the same purpose is something like SpeedTree. I like to joke that most video games are just SpeedTree. They're not picking from a class of trees; they're defining the environment they're trying to fill with vegetation and allowing the smarts of SpeedTree to create variation through procedural generation, so you walk through a forest that looks real and varied. That's procedural, though, not an AI tool. Now we're moving more into using generative and inference-based techniques, but they've been heavily used for many years by artists, programmers, and developers in real-time 3D.
N
Nick Gruse9:11
You mentioned text to 3D models and procedural generation. But with respect to more complex, dynamic, malleable assets like a human, you can't just generate a holistic human; you want it to move the way a human would. What are the puts and takes with training such a complex AI model that has to reflect how it should behave in any given environment?
M
Marc Whitten9:43
Totally makes sense. There are several challenges there, and there's not one system of training that solves all problems. In fact, part of the reason we think of Muse as a platform is that we want to allow creators to define what they're trying to accomplish and then allow a variety of AI or procedural techniques to take over and help get them further along. In large, complex environments, there are a couple of big problems we're spending time on. First, there's a lot of research going on in text to 3D, but one problem is that models can start to generate a 3D mesh that's essentially 'triangle soup'—an undifferentiated set of triangles where you can't tell where the ground ends and the chest begins. So we're spending a lot of time on how to think about problems like that. The other thing is you're going to see a lot of expert-trained AI systems for specific characteristics. You mentioned humans, which is a great example. We spent a lot of time, and in fact acquired a company called Ziva Dynamics, that's really focused on ML and AI-based training to solve this problem. The old way of creating a high-quality character was to hand-create it, then hand-rig every point that's potentially animatable, and then hand-move them through every scene. With Ziva, we've trained an AI model that understands a huge number of data points and tens of thousands of different emotional states around a character's facial animation and deformation. Then we run it through an advanced deformation solver so you can get instantaneous inference of what the animation should be, versus having to hand-rig it. Something that might take a team of artists three months can suddenly take a server about 10 minutes to get you back something you can use in real-time and interactively. So it's not pre-canned. The challenge is how do you create orchestration in a platform layer so that we can plug in all the expert AI systems for a particular area.
N
Nick Gruse13:12
Wow, and you mentioned a really interesting stat there: three months to 10 minutes. I want to zoom out a bit and talk about productivity gains at large. When you look at the video gaming space, AAA games take years to develop. What do you think are the areas to focus on when you talk about productivity? And are you hearing stories from customers already saying they're rocking and rolling with all these services?
M
Marc Whitten14:10
We're rolling out some of these tools, learning with our customers, and we'll continue to move them forward. But I'll step back a bit. The truth is we have a crisis of content creation. It's not possible to make the amount of 3D content and assets that video games and other real-time 3D experiences demand. When everybody was so hyped about things like the metaverse, one thing many of us were talking about is that it's going to be a bunch of really empty worlds because there's not enough content. In video games, especially larger ones, if you look at the announce video of a game, it often looks amazing, but then the game is delayed or becomes less ambitious or fails to reach its goals. There's a desire for much more content, so we have to solve the productivity problem. The interesting point about the three months to 10 minutes is it's actually zero to 10 minutes, because everybody looks at the three months and says we can't afford it, so we don't do it. Or we use a much worse model. So we're able to change the quality bar and give them productivity gains. On the asset side, we see this in two ways. Number one, we want to make sure these tools enable a broader class of creators to have access to technology and become creators. We believe the world is better with more creators in it. Number two, for existing creators, this isn't about 'generate some art for me'; it's about 'give me a generation and then all the knobs and nodes I need to turn that into something truly special for the game I'm trying to build.' In the end, I think it will mean significantly more high-quality content and significantly more games. I actually think it's going to mean more net creators and be expansive for the industry, but also allow people to make bigger bets on the types of games they want to try because they can get there a lot faster in prototyping to understand the challenges.
N
Nick Gruse16:53
Wow, yeah, that's really fascinating. I think there's a concern that people will have to learn new tools, but I actually think it will be expansive of the types of things people are trying to accomplish, especially in real-time 3D. So you've iterated more content and more high-quality content, but looking at the other side of the equation on the consumer front, given that people have limited time and resources, how do you think the market will evolve as it may look more competitive for developers to get their game in front of consumers?
M
Marc Whitten17:48
Yeah, gaming is already extraordinarily competitive to find and keep your audience. And it's also one of the longest uses of AI that we've had. If you look at the ad monetization in Unity, the ad business is really about how do you help people find their customers and keep those customers engaged, and then monetize their game so they can be successful. That's all based on using AI models and neural networks to better target, so you find the specific player that would most like your game and is most likely to stick around. So AI will play a continuous role in doing that. However, for live games or live ops games, it's not about the initial bits you shipped, but about building a community and maintaining the game to keep it engaging and robust so you hold on to players for a long period. I'm a big Marvel Snap player, and every time a patch comes out, I wonder how I'm going to change my deck. The core reason they're doing that is to keep the game engaging for existing players and drive new players in. One of the big opportunities in AI at runtime is to create truly differentiated experiences. With Sentis, you can see a little bit of this, like an NPC that you can truly have a conversation with, shaping the game experience in a way that's very different. If you think of Skyrim, one of the best games ever, you'll go talk to someone and they'll say, 'I used to be an adventurer like you, but then I took an arrow to the knee.' It's a total meme because millions of gamers have heard it dozens of times. You can't create content for every guard to say something different every time. But when you switch that out where an NPC can have a unique conversation with you, it's going to feel a lot more alive than anything we have today. That's an example of NPCs, but it plays out in a lot of the systems inside your games.
N
Nick Gruse21:25
Just following up on intelligent NPCs, how far are we from a fully interactive Skyrim? And how big do the models have to be to run that locally?
M
Marc Whitten21:42
It's a really good question. There's an economic question and a capability question. I don't think we're very far away. In fact, you can see lots of examples, including people hacking Skyrim to do this. But the problem is that inference can be very expensive. It's one thing if it's at create time because it's a relatively small number of people for a defined period. If instead you have to pay every time you talk to an NPC for a cloud inference times the number of concurrent users, and you have a successful game, the cost could scale in a way that's not manageable. So part of what we're trying to do with Sentis is to make it easier to use neural network-based features at runtime with deterministic cost. The other problem is running it locally. The tech stack for most neural networks and the languages behind them, whether TensorFlow or PyTorch, don't exist broadly on all devices. So you'd have to say only people on this particular device can play the game, which is too small a market. What we're trying to do with Sentis is make it so you can take any standard model, assuming it fits in local memory, and we'll make sure it works anywhere the Unity runtime runs. We'll transpile the PyTorch into whatever we need locally to run on the GPU or CPU for the best performance. So suddenly, you can have small models that work for most people, and maybe you go to the cloud for a boss or head wizard that needs deeper inference. But the other thing is, it's not just about those types of features. I believe we're going to enter a golden age of inference-based techniques used at runtime, and many of those will be very small models. There's one hard truth in game development: you have 16 milliseconds to get your frame done. You have to decide how many of those 16 milliseconds you can spare for a particular technique because you have to do everything for the game—the whole simulation, input, everything—in every 16-millisecond turn. Inference means you can take some of those problems and turn them into inference instead of direct calculation. We're hearing that people are super interested in having models that do lighting, for example, instead of calculating the physics. Think about ray tracing, which is a very high-performance physics calculation to generate how a photon gets to your eyeball. Many types of techniques can be turned into inference, making it easier to deliver across every platform.
N
Nick Gruse26:14
Gotcha. And do you think as we enter this golden age of AI, does this accelerate growth in memory? If you think about the iPhone's unified memory, it's been growing for the past decade but slowed down in the mid-2010s. That presents a bottleneck for more performant models running locally. Is that a meaningful constraint in the longer term?
M
Marc Whitten26:51
I think on core RAM, it probably will end up helping drive additional innovation. You use a GPU for a lot of the silicon, and they're becoming more tailor-made to run inference versus a render pass. That will continue, and it will be on every SOC. Everybody's going to be thinking about their inference engine embedded in silicon, and that will drive a cycle about what that changes in terms of hardware cache, RAM size, or other things to get the most out of it, just like GPU evolution has driven core GPU evolution for 3D graphics for the last 20 years.
N
Nick Gruse27:51
Marc, I want to change gears here a bit. You mentioned the Apple Vision Pro. What can you tell us about the collaboration with Apple? When can developers expect to get their hands on it? Is it already out in the market, or will it be coming later this year or early next year?
M
Marc Whitten28:17
Yeah, we've collaborated with Apple for a very long period of time across many things. You might remember that when they originally launched their M1 silicon, Unity was there—not just the runtime capable of running on it, but the full editor already ready to go on M1 silicon. And obviously, we've worked very closely on iOS and iPhone for many years; it's a huge part of the gaming market. So when we first heard what they were thinking about doing, we were really excited to see where they were going to go. They've really gone super deep into how to make the platform work. Our goal is to make sure as much of Unity just works for you as you target whatever platform matters, so we want to make it as easy as possible for you to build your game and then target platform X, Y, or Z. That's standard: make sure it's easy to use Unity on day one of the launch. But the other thing that I think is very unique about Apple's approach to spatial computing and the Vision OS is shared mode in the experiences. That meant we needed to work quite closely with Apple to think about how to render, how to build an experience that is shared between multiple people. That's different from what we've done on any other platform, and I think it creates an opportunity. You can imagine a crazy game or experience I've built, and next to it is another experience, like a video screen where you and I are chatting, and maybe I have something off to the side that I'm just slightly paying attention to. That's cool, and you can start to imagine those experiences being able to interact with each other over time. Whenever there's this type of paradigm shift, game creators and developers do amazing things. So yeah, we're super excited about it.
N
Nick Gruse31:15
Yeah, it sounds amazing. I'm so excited for this collaboration. I have a follow-up question that ties back to what we were originally discussing with Muse. Do you think as Muse scales, it will be possible for developers to deploy on a device like Vision Pro even if they didn't design for that platform originally?
M
Marc Whitten31:36
Well, that's always our goal, to make that as easy as possible, regardless of whether using AI or not. We'll do everything we can to make it easy. I think there are two parts to that, and AI likely makes part of it easier. There's the technical part: is my game capable of running on a particular platform? But then there's the design part. Think about a PC game or application; it's extremely information-dense because you have a large screen. One of the things people learned when they first started doing game streaming of console games on a phone is that it's really cool to play Assassin's Creed on a phone, but there's a lot of text—your screen is covered with text because it was designed for a 50-inch display, and suddenly it's on a five-inch screen. Twin-stick is different than mouse and keyboard, which is different than touch. Those aren't just about translating between one versus the other; they often go deeply into how you design the gameplay loop to be fun. You can translate between mouse and keyboard and a controller, but you'll still have to fundamentally understand the type of game or experience you're trying to build and how a user is going to interact with it.
N
Nick Gruse33:19
That makes sense. Very excited to see what gets built for the Apple Vision Pro.
M
Marc Whitten33:23
Yeah, me too. It's going to be really fun to watch. I think your point on shared experiences for the Vision Pro is really interesting. At least from what I gathered from the Vision Pro demo and the Quest Pro, the Quest is taking a very different approach to gaming compared to the Vision Pro, with 2D in a spatial environment versus full immersive 3D. Do you think those two experiences will converge over time, or will they remain distinct, and what does that mean for developers developing for either device?
What's exciting about Unity is that we end up working with a lot of different platforms, and for us, we're always looking at how that creates more opportunity for creators. We think all those spaces will get filled. If there's a high-quality device and a company building really cool stuff for it, game creators will come in and fill out interesting use cases. So whether you're talking about a device like the Quest, Vision Pro, Xbox, or a PC, all of them are going to have really robust ecosystems, and people are going to explore them. Now, if we talk about Vision Pro particularly, it's an extraordinarily powerful device—whether it's the sensors, the display quality, or the power of their M silicon to drive amazing results. So I think you're going to see some very interesting stuff. What will be exciting is when people start to use that kind of palette of toys they have. I always think of it as how many crayons do they have in their crayon box to start building something. They're going to do something that none of us think the device is actually built for. One of my favorite first examples on the iPhone when they launched the App Store was from a company called Smule, who built an app called Ocarina, which essentially made the iPhone a flute. You blow into the microphone and use the touch screen. Developers weren't thinking about that, but suddenly they came out and did it, and you could see everybody at Apple go 'oh,' and every other game creator go 'oh.' That's how the space gets broader and broader. I've been doing platforms for a long time, and I always like when people ask me what I think people are going to do with this. I can tell you that whatever idea I have is low on the creativity scale, because the beauty of being a platform creator is you unleash the creativity of so many brilliant people. People do amazing things. There were people on the original PlayStation One, which lasted a long time, who were still finding new things to do with it at the end.
N
Nick Gruse37:10
Now with the Apple Vision Pro and the Unity collaboration, I think we'd be remiss not to mention that the beta just opened, right?
M
Marc Whitten37:19
That's right. We just started our first closed beta. We've had extraordinary interest since the WWDC announcement, continuing with people signing up, both in games and non-games. On Wednesday of this week, we started putting out our first closed beta, and our goal is to expand that as soon as possible to make it easy for everybody interested to target it. It's really exciting to get started with our first creators on it.
N
Nick Gruse37:50
Very exciting. Everything that Unity and Apple are doing in this collaboration... We want to end with just one last question. We've talked a lot about gaming, but what about non-gaming use cases? How do you see AI being used in different industries?
M
Marc Whitten38:10
One of the main reasons I joined was because of the massive opportunity in the continued growth of real-time 3D in non-gaming use cases, in industrial use cases, and digital twins. I believe deeply that essentially every industry is going to be rebuilt around real-time 3D for a variety of really important reasons. We are 3D and real-time people. So whether it's medical, manufacturing, or the way I interact with my house or devices, the same techniques that were pioneered in gaming and continue to be extremely robust in gaming are as useful, and we expand through all the different types of use cases. AI is really critical inside of it. You go through that content creation problem we talked about; it's even more so outside of gaming, where there may not be the same concentration of artists and 3D technical creators in a particular industry or enterprise. So we're working very hard to make it easy for everyone to pick up and use. Frankly, it's the fastest-growing part; it continues to see extremely robust growth, whether in automotive, manufacturing, or retail. It's a really exciting opportunity in a place where Unity plays specifically well. We're not trying to change the fact that every industry has a bunch of 3D data. A manufacturing company builds their stuff in 3D CAD tools; they have tons of 3D data. We're not trying to say you shouldn't use that super high-end CAD tool you've always used. What we're trying to do is free that 3D data, because today it's captive on high-end desktops in the R&D department of one part of the company. If you want to use that data for marketing, sales, support, or any other thing, the way it works today is you wait to get a real object and take photos of it. There's no reason for that. You should be able to have the marketing department access the same 3D data, transformed to be real-time and interactive, so it can run on every device and get throughout the enterprise and to all the customers and partners of a particular company. That's our goal: to take the 3D they already have and truly free it so it adds value across everything they're trying to do, from core sales to support to training.
N
Nick Gruse42:11
Well, thank you so much for your time, Marc. And for all our developer listeners out there, make sure to check out Muse, Sentis, and the Unity VisionOS open beta. Also, make sure to check out ARK's blog on generative AI's impact on gaming. Thank you.
M
Marc Whitten42:29
Thank you.
N
Narrator42:37
ARK believes that the information presented is accurate and was obtained from sources that ARK believes to be reliable. However, ARK does not guarantee the accuracy or completeness of any information, and such information may be subject to change without notice from ARK. Historical results are not indications of future results. Certain of the statements contained in this podcast may be statements of future expectations and other forward-looking statements.