Everyone is focused on LLMs. ChatGPT, Claude, large language models. But you have raised $1 billion to build something different. Large world models. Make the case for us. What bet are you making that others aren't?
Right. So, this is my co-founded startup, World Labs, and we are all in on spatial intelligence. And the path to spatial intelligence is building a large world model. So what is the case for us? The case for us is a 500-million-year story. Animal intelligence starts with seeing and moving in the physical world. Evolution began with us as animals, knowing what the world is, knowing who we are, and how to move around it, interact with it. And much of life, human life, human work life, human private life, has a lot to do with perceiving, understanding, reasoning, and interacting with the world, including imaginary worlds of creativity, and productivity, as virtual worlds. So unlocking that capability in machines, unlocking the capability of generating any 3D or 4D worlds, unlocking the capability of reasoning within any world, unlocking the capability of teaching agents, or robots, or assisting humans to interact with the world, is what spatial intelligence is about. And that is what we are focusing on. So what can world models ultimately do that LLMs will never be able to? Can words put out fires? Can words cook an omelet? I think there is so much, right? So we, for example, creativity. People design. People, whether we are designing interior spaces, we are designing machines, we are designing homes, we are designing stories. So much of that is beyond words. We also use agents, whether we use agents in virtual worlds, whether it is for entertainment, like gaming, or for more serious industrial applications, whether it is digital twins, design, inspection, or optimization, or many kinds of optimization tasks. Or we build robots to help us do many things, from putting out fires, to helping in healthcare scenarios, to manufacturing. All of these are downstream applications of unlocking spatial intelligence and building world models.
So what do you think the ChatGPT moment for world models will be like? How will we know this has arrived?
Yeah, that is a great question, Emily, because chat is such a consumer behavior that the ChatGPT moment tends to describe a viral, public consumer moment of getting very close to what AI can do. In the world of world models, the kind of spatial intelligence we are trying to unlock... I am still trying to figure out whether there is a corresponding consumer moment, because the kinds of applications we are talking about tend to go first to professionals: professional creators, professional designers, professional developers, professional researchers and engineers, who will use it for robotics, industrial design, and all of that. So maybe we will not necessarily have a consumer moment. But maybe we will, and you know, I would love to design my home in a much easier way, and just change the color of the curtains with a click.
All right, that sounds pretty cool. So in the last six months, Yann LeCun left Meta to work on world models, Google shipped Project Genie, NVIDIA has its own world model, Cosmos. NVIDIA is also one of your investors. What do you have that they don't? And which competitors worry you the most?
Yeah. So first of all, we started World Labs in 2024. I still remember when we were out talking about world models and spatial intelligence, it was just a year after ChatGPT. People were still totally talking about LLMs. So we really had a head start in understanding that this is going to be the next frontier of AI. I am very excited by that. So what do they have that we don't? First of all, I think we have an incredible team. We have the conviction. They don't have the godmother, that is for sure. But the world is big. And I think this is just like LLMs. I think there will be many companies doing incredible work in world models. Just 24 hours ago, we kind of got fed up that the term world model has become so confusing and is being used in so many different ways that we actually put out a blog post explaining what a functional taxonomy of world models is, instead of mushing everything together. And the way I see it right now, there are three ways of referring to world models when it comes to spatial intelligence. One is what I call a renderer. That is when the model puts beautiful pixels on the screen, mostly like a video generation model. And the consumer is mostly human eyeballs. While the model commits to beautiful pixels on the screen, it doesn't necessarily commit to physics, dynamics, and geometric correctness. Because that is just for human visual consumption, not necessarily for computation or other tasks. Then another kind of world model is what we call a planner, which is more for machines, more for robots, where it outputs, whatever the input is, the state of the world or the action, it outputs the correct action to take next. And you see that kind of world model a lot in robotics applications. And you hear it in that context. The third kind, which I think is the linchpin of the three, is a simulator. It is consumed by humans as well as machines. It tries to respect the structure, the physics, and the dynamics of the world, and truly simulate the 3D and 4D information of the world, as well as the semantic information. And the simulator could become a renderer. The simulator could become a planner. But this layer is a huge critical path, in my opinion, to unlock spatial intelligence. And that is what World Labs is working on.
All of this rolls up into robotics. So I want to get your take on the field, and humanoids in particular. Funding for humanoids hit $6 billion, but, you know, they still cannot load my dishwasher as fast as I can. They still cannot go get my Amazon packages. Will world models and World Labs close the gap between hype and reality?
That is a loaded question, Emily. First of all, that is my job, yes. I get it. First of all, robotics is going to be one of the most important revolutions in human industrialization. $6 billion is too small, right? If you look at self-driving car investment, if you look at language model investment, it took way more than $6 billion. I am not saying we now... I think it will take time to invest. And hopefully it will not just take hype, but take thoughtfulness to invest in the right efforts. For example, unlocking world modeling, spatial intelligence, and the simulation layer — all of this is part of that important effort. Are we going to close the gap? I do believe World Labs is working on one of the most critical technologies in spatial physical intelligence. And obviously, that is the hope.
You have been more measured on AI safety, skeptical of the doom narrative, but also of heavy-handed regulation. When you look across the industry, where do you see real safety work versus safety theater? Is anyone getting it right?