CEOInterviews.AI
Start App
Fei-fei Li
Co-Founder and Chief Executive Officer; incoming Executive Vice President and Chief Scientist, AMD, World Labs

Beyond ChatGPT? Fei-Fei Li Bets on Large World Models to Redefine AI

📅 Jun 07, 2026 New SciTech 新科技 18 MIN 416 VIEWS 18 SEGMENTS · 2 SPEAKERS
Everyone is focused on large language models — ChatGPT, Claude, Gemini, and the future of text-based AI. But Fei-Fei Li is making a different bet. Through World Labs, she is building Large World Models: AI systems designed not just to understand words, but to understand space, objects, motion, interaction, and the physical world itself. Her argument is simple but powerful: Can words put out a fire? Can words cook an omelet? If AI is going to power robots, virtual worlds, 3D creation, healthcare, industry, and real-world agents, it must go beyond language. It must learn to see, reason, act, and...

What Fei-fei Li said

Written from the verified transcript and checked against it. Every figure links to the moment it was said.

Fei-fei Li, co-founder of World Labs, argues that the next frontier of AI is spatial intelligence, built through large world models, not just LLMs. She explains that World Labs, started in 2024, focuses on a 'simulator' type of world model that respects physics and 3D/4D structure, unlike 'renderers' (video generation) or 'planners' (robotics). Li says robotics investment of $6 billion is too small compared to self-driving or language models, and that world models are critical to closing the gap. On safety, she calls for scientifically grounded guardrails, citing real AI use in healthcare, and criticizes hype and 'theater.' She acknowledges mixed student sentiment amid an 'AI hate wave,' urging thoughtful discourse. Li insists AI must change K-16 education, and she declines to engage with the term AGI, focusing instead on building technology that improves lives. She hopes to ship a spatial intelligence model this year.

Key takeaways

  1. World Labs is building a 'simulator' world model that respects physics and 3D/4D structure, unlike renderers or planners.
  2. Li says $6 billion in humanoid funding is too small; world models are critical to closing the gap between hype and reality.
  3. Li says AI must change K-16 education, and kids should not be scared of AI but should lead it.

Numbers and commitments

FigureWhat it refers toTypeAt
$1 billion raised for World Labs metric 0:00
2024 year World Labs was started timeline 4:50
$6 billion funding for humanoids metric 7:48

Chapters

  1. 0:00Case for large world models
  2. 3:20ChatGPT moment for world models
  3. 4:50Competition and world model taxonomy
  4. 7:48Robotics and humanoids funding
  5. 9:39AI safety and real-world use
  6. 12:38Student sentiment and AI hate wave
  7. 15:17AI in education
  8. 17:08AGI and shipping plans

Questions asked in this interview

9
  1. 0:00What bet are you making that others aren't?
  2. 3:13How will we know this has arrived?
  3. 4:32And which competitors worry you the most?
  4. 7:48Will world models and World Labs close the gap between hype and reality?
  5. 9:19When you look across the industry, where do you see real safety work versus safety theater? Is anyone getting it right?
  6. 12:13And if they are scared, are the fears justified?
  7. 15:09How do you think AI will change learning in the college experience?
  8. 16:46Are they wrong, or is the disagreement about what we are calling the goal?
  9. 18:01What is the one thing you will have shipped this year that we will be talking about next year?
Emily 0:00 ↗
Everyone is focused on LLMs. ChatGPT, Claude, large language models. But you have raised $1 billion to build something different. Large world models. Make the case for us. What bet are you making that others aren't?
Fei-Fei Li 0:19 ↗
Right. So, this is my co-founded startup, World Labs, and we are all in on spatial intelligence. And the path to spatial intelligence is building a large world model. So what is the case for us? The case for us is a 500-million-year story. Animal intelligence starts with seeing and moving in the physical world. Evolution began with us as animals, knowing what the world is, knowing who we are, and how to move around it, interact with it. And much of life, human life, human work life, human private life, has a lot to do with perceiving, understanding, reasoning, and interacting with the world, including imaginary worlds of creativity, and productivity, as virtual worlds. So unlocking that capability in machines, unlocking the capability of generating any 3D or 4D worlds, unlocking the capability of reasoning within any world, unlocking the capability of teaching agents, or robots, or assisting humans to interact with the world, is what spatial intelligence is about. And that is what we are focusing on. So what can world models ultimately do that LLMs will never be able to? Can words put out fires? Can words cook an omelet? I think there is so much, right? So we, for example, creativity. People design. People, whether we are designing interior spaces, we are designing machines, we are designing homes, we are designing stories. So much of that is beyond words. We also use agents, whether we use agents in virtual worlds, whether it is for entertainment, like gaming, or for more serious industrial applications, whether it is digital twins, design, inspection, or optimization, or many kinds of optimization tasks. Or we build robots to help us do many things, from putting out fires, to helping in healthcare scenarios, to manufacturing. All of these are downstream applications of unlocking spatial intelligence and building world models.
Emily 3:13 ↗
So what do you think the ChatGPT moment for world models will be like? How will we know this has arrived?
Fei-Fei Li 3:20 ↗
Yeah, that is a great question, Emily, because chat is such a consumer behavior that the ChatGPT moment tends to describe a viral, public consumer moment of getting very close to what AI can do. In the world of world models, the kind of spatial intelligence we are trying to unlock... I am still trying to figure out whether there is a corresponding consumer moment, because the kinds of applications we are talking about tend to go first to professionals: professional creators, professional designers, professional developers, professional researchers and engineers, who will use it for robotics, industrial design, and all of that. So maybe we will not necessarily have a consumer moment. But maybe we will, and you know, I would love to design my home in a much easier way, and just change the color of the curtains with a click.
Emily 4:32 ↗
All right, that sounds pretty cool. So in the last six months, Yann LeCun left Meta to work on world models, Google shipped Project Genie, NVIDIA has its own world model, Cosmos. NVIDIA is also one of your investors. What do you have that they don't? And which competitors worry you the most?
Fei-Fei Li 4:50 ↗
Yeah. So first of all, we started World Labs in 2024. I still remember when we were out talking about world models and spatial intelligence, it was just a year after ChatGPT. People were still totally talking about LLMs. So we really had a head start in understanding that this is going to be the next frontier of AI. I am very excited by that. So what do they have that we don't? First of all, I think we have an incredible team. We have the conviction. They don't have the godmother, that is for sure. But the world is big. And I think this is just like LLMs. I think there will be many companies doing incredible work in world models. Just 24 hours ago, we kind of got fed up that the term world model has become so confusing and is being used in so many different ways that we actually put out a blog post explaining what a functional taxonomy of world models is, instead of mushing everything together. And the way I see it right now, there are three ways of referring to world models when it comes to spatial intelligence. One is what I call a renderer. That is when the model puts beautiful pixels on the screen, mostly like a video generation model. And the consumer is mostly human eyeballs. While the model commits to beautiful pixels on the screen, it doesn't necessarily commit to physics, dynamics, and geometric correctness. Because that is just for human visual consumption, not necessarily for computation or other tasks. Then another kind of world model is what we call a planner, which is more for machines, more for robots, where it outputs, whatever the input is, the state of the world or the action, it outputs the correct action to take next. And you see that kind of world model a lot in robotics applications. And you hear it in that context. The third kind, which I think is the linchpin of the three, is a simulator. It is consumed by humans as well as machines. It tries to respect the structure, the physics, and the dynamics of the world, and truly simulate the 3D and 4D information of the world, as well as the semantic information. And the simulator could become a renderer. The simulator could become a planner. But this layer is a huge critical path, in my opinion, to unlock spatial intelligence. And that is what World Labs is working on.
Emily 7:48 ↗
All of this rolls up into robotics. So I want to get your take on the field, and humanoids in particular. Funding for humanoids hit $6 billion, but, you know, they still cannot load my dishwasher as fast as I can. They still cannot go get my Amazon packages. Will world models and World Labs close the gap between hype and reality?
Fei-Fei Li 8:09 ↗
That is a loaded question, Emily. First of all, that is my job, yes. I get it. First of all, robotics is going to be one of the most important revolutions in human industrialization. $6 billion is too small, right? If you look at self-driving car investment, if you look at language model investment, it took way more than $6 billion. I am not saying we now... I think it will take time to invest. And hopefully it will not just take hype, but take thoughtfulness to invest in the right efforts. For example, unlocking world modeling, spatial intelligence, and the simulation layer — all of this is part of that important effort. Are we going to close the gap? I do believe World Labs is working on one of the most critical technologies in spatial physical intelligence. And obviously, that is the hope.
Emily 9:19 ↗
You have been more measured on AI safety, skeptical of the doom narrative, but also of heavy-handed regulation. When you look across the industry, where do you see real safety work versus safety theater? Is anyone getting it right?

9 more exchanges in this transcript

Sign in free to read the rest of this interview. No card required.

Sign in to read the full transcript

Cite this transcript

APA, MLA, BibTeX
APA

Li, F. (2026, June 7). Beyond ChatGPT? Fei-Fei Li Bets on Large World Models to Redefine AI [Interview transcript]. New SciTech 新科技. CEOInterviews.AI. https://ceointerviews.ai/interview/1112168/

MLA

Fei-fei Li. "Beyond ChatGPT? Fei-Fei Li Bets on Large World Models to Redefine AI." New SciTech 新科技, 7 Jun. 2026. Transcript, CEOInterviews.AI, https://ceointerviews.ai/interview/1112168/.

BibTeX
@misc{li2026_1112168,
  author       = {Fei-fei Li},
  title        = {Beyond ChatGPT? Fei-Fei Li Bets on Large World Models to Redefine AI},
  howpublished = {Interview transcript, New SciTech 新科技. CEOInterviews.AI},
  year         = {2026},
  month        = {jun},
  url          = {https://ceointerviews.ai/interview/1112168/},
  note         = {Speaker-attributed transcript with timestamps}
}