Welcome everybody to a very special live stream today. Our live stream is called Robots on Their Best Behavior. It's a fireside chat with Dr. Fei-fei Li and Dr. Jim Fan from Nvidia. Very excited. You'll see them both with us. Now, we're not going to do any news at the beginning of this live stream. I'll cover news at the end after their discussion. But today, we would love for you to join us for an in-depth discussion of BEHAVIOR, a large-scale benchmark and challenge for advancing embodied AI. The session is going to outline scientific vision and research motivations behind BEHAVIOR, highlight how it links perception, reasoning, and action in real world household settings, and explain the design of the BEHAVIOR challenge. It's also going to cover evaluation methodologies, the difference between standard and privileged information tracks, and the role of simulation in advancing robotics research. It's an opportunity to understand how BEHAVIOR unifies efforts across academia and industry to build robust human-centered intelligent agents. First, let's introduce our special guests and then I'm going to disappear into the background. But Jim, thank you so much for being part of this live stream today. And I understand you are a PhD student of our special guest, Dr. Fei, is that correct?
Yes. Yes, totally. What a great honor.
So, tell everybody what your role is at NVIDIA and what's the purpose for us doing this today?
Yeah. Hi everyone. My name is Jim Fan. I am the director of robotics at NVIDIA and currently my team and I are spearheading Project GR00T, which is to build the robot foundation models for the ecosystem. And today, I'm very honored to have Fei here. I had a great honor to be her student and then even a greater honor to be the first student who live streams with Fei. So you know, at Stanford with Fei, those years completely reshaped me as a young scientist. I also witnessed the lab's transition from computer vision to embodied vision robotics and Fei's guidance has always been a north star for me. So thank you so much Fei for making me a better researcher, a better thinker and also a better live streamer.
I don't think Jim you learned live streaming from me but as you kindly say it's an honor for you but the truth is it's an honor for professors and advisors to work with such brilliant and wonderful students and you're absolutely one of them.
That's so nice. Well Dr. Fei, please introduce yourself and your background a little bit and then I'm going to hand over the microphone to Jim so we can start the discussion off.
Yeah, my name is Fei-fei Li. I'm a professor of computer science at Stanford as well as the founding co-director of Stanford Human-Centered AI Institute and right now I'm also a co-founder CEO of World Labs, which is a startup that is focusing on spatial intelligence empowering from virtual worlds to embodied AI worlds.
Amazing. Okay. Well, I can't wait. I'm going to be watching with everybody else. So, have a great discussion. This live stream is going to last about a half hour and then like I said, I'll cover the news at the end. Have a great time.
Awesome. I will jump right in. So, among the countless stellar achievements of yours, ImageNet is the one that stands out to most people. Looking back, I think ImageNet didn't just benchmark vision models. It really benchmarked an entire era in the history of AI. And when you first launched ImageNet, did you ever imagine it would completely reshape AI?
Thank you Jim for opening up with ImageNet. No. I did not imagine I would reshape anything or not. I imagined what my curiosity would take me. And at that time my curiosity was really an intertwined dual focus. One is what is a fundamental problem to define in visual intelligence. So that unlocking that problem would largely unlock visual intelligence. It's not gonna, of course visual intelligence is very rich but there are some fundamental problems and ImageNet defined object recognition especially classification as one of the most fundamental problem. It was a north star problem. The second equally important problem I was going after was a more technical problem of what's the role of big data in machine learning and it was as you know ImageNet together with neural network especially convolutional neural network as the first paper out of ImageNet and Nvidia's GPUs defined the beginning of deep learning but before that neural network was not clear for everybody as a defining methodology. So, was big data. People were not understanding the role of big data. But we had a conjecture that big data would change machine learning, would change statistical learning in a fundamental way. And so that's why ImageNet was developed with this dual purpose of a technical bet on big data as well as a north star of visual intelligence task which is object recognition.
That is amazing and I think the world needs more ImageNets. So today our main highlight is BEHAVIOR, a new robotics benchmark that your team has built. So could you share with the audience here a bit more about the BEHAVIOR project and its connection and perhaps inspiration from ImageNet?
Yeah, I think well, okay, in a nutshell, we just announced to the world the BEHAVIOR 1K or BEHAVIOR 1000 challenge, which is a comprehensive simulation benchmark and training environment for embodied AI and robotics. And the focus of the 1,000 tasks are the everyday household long horizon, which we will define later, long horizon tasks that robots can help people with. And BEHAVIOR is intended to give an open-source environment for researchers across the world and across the embodied AI discipline to use to train their robotic learning algorithms and also give an opportunity for the community to benchmark against a standardized list of tasks. So BEHAVIOR is inspired by the fundamental problems of robotics and embodied AI and the problems we when it began you know almost more than five years ago we were seeing the lack of standardization in robotic learning. We were seeing very anecdotal choices of tasks that is very hard to compare from one paper to another and also obviously we are also seeing that it's very hard to find training data right we faced that problem in the pre-ImageNet days for computer vision but we are still facing that in robotics so those are the pain points in robotic learning and research. So we were inspired by these and wanted to build BEHAVIOR. I remember Jim, you were part of the lab when we started BEHAVIOR and I even had conversations with you. You even though you were focusing on your own dissertation and finishing the PhD at that time, but you were very excited and gave me and the team a lot of encouraging words about that.
Yes. I was there when BEHAVIOR was first launched and I can see how big of an initiative it grew into right like how many PhDs dedicated their work to it and then making it really great. So it actually is super exciting for me and to see it now reaching like a thousand tasks and becoming this new competition that I hope the community can rally behind. Yeah, I don't know if you want me to get into a little more details about what BEHAVIOR is consisted of the three core components. Should I?
Yes. You have mentioned many things about BEHAVIOR. It's a super complex project. Would love to unpack a bit. Perhaps we can start with you mentioned a thousand activities. How did you select that thousand activities in the first place?
That's a great question. That did somehow bring me back some memories of ImageNet, right? Like the taxonomy of a problem domain is itself science and of course I spend a lot of time explaining how we selected the 22,000 categories for ImageNet. For robotics and embodied AI, it was even more challenging to be honest. Robotics is not just semantics. You know, objects can at least be captured more or less by word or word phrases of semantics, but robotics is trying to imitate human behavior that's accomplishing tasks. And many many human behaviors of accomplishing tasks is not defined by words so easily. For example, I was just told that my microphone is scratching my zipper, so I have to use my hand to hold this. What is a word for that? What is a category that you can assign a label for this very behavior of my two fingers pinching this cable and lifting it away from my zipper, right? That's how challenging robotics is. But we have to start somewhere. And what we have seen is many of the tasks are very short horizon. For example, grab a block and put it on a table or we call them more punctual tasks or open a drawer or pick and place, pick something up and place it somewhere. These are very important tasks, don't get me wrong, but BEHAVIOR was intended to be bold and ambitious and to really encourage ourselves and the community to go after truly complex human world relevant tasks. And the way we pick them is to look at how humans spend their time in every day. And how do you find where humans spend their time? It turns out governments, whether it's America or other governments have surveys. The American Labor Bureau has a survey called American Time Usage. And that survey breaks people's daily life as well as daily work into a hierarchy of thousands of tasks. We actually went through that survey very carefully, created a candidate list of more than 2,000 tasks. And then there's a very important question. Which tasks should we encourage the robotic community to build towards? And who is it to say? Is it Jim to tell us or is it Fei or is it someone else? We have smart students at Stanford but how do we do this that's respectful of human lives and human desire. So we applied a human-centered approach and did a survey across more than a thousand people over different demography and asked them a very fundamental question. What would you like the robot to help you with?
That's very important. It's a very human-centered way of approaching this problem. Instead of saying, what do you want robot to do that could potentially replace humans, we want robots to help our daily life and productivity. So, the tasks ranges from opening Christmas gifts, playing squash, cleaning pets, all the way to wiping the floor, making bed. And as you can guess, very rarely people want robots to open their Christmas gift. Even though it's probably I'm sure your robot can already do that, but that's just such a human behavior that we would love to keep for ourselves. It's the...
We're not referring to the robot.
Exactly. There was even a task of asking a robot to pick a ring. I mean, I really hope nobody does that. But we learned, for example, the top three tasks that people just unanimously want robots help are cleaning bathrooms or washing floors and cleaning after a wild party. I'm sure some of them are parents.
Vote for all three, which is I badly want that to happen as soon as possible.
How many wild parties do you have, Fei?
Well, now that we one day we'll have a robot, you can have more wild parties, but...
Anyway, so the thousand BEHAVIOR tasks were the top ranked by humans who want robots help.
Absolutely. And regarding robot task, so unlike static data sets for LLMs and computer vision, the task they cannot exist outside of an interactive environment or simulation. So Fei, could you share a bit about how you design these interactive environments and about the simulation framework that you choose?
Well, this brings us to the collaboration with Nvidia. I remember I was talking to Bill Dally, your chief scientist and then a number of scientists on the Nvidia robotics team and there was a common shared desire let's try to do something that push robotic learning forward and for simulating robotic behavior you need to simulate environments that are interactable. So there are two very important things. One is an interactive environment of objects. So BEHAVIOR ended up gathering 50 really high-fidelity fully interactive scenes that across multiple spaces from offices to restaurants to apartments. And within these scenes we have around 10,000 objects that we have accumulated that are realistic that are household scale. It can be articulated. Some of them are deformable and these are what we call assets as you know and for academic or open source research scale this is very very large scale. Of course, I think privately this scale can go a lot larger. But in contrast, most of the papers you see in robotic learning are just one or two much smaller desktop or countertop environments with dozens or not even interactive objects. So, we're taking this to two three orders of magnitude higher. And of course, interactive environments and objects need to obey the laws of physics and need to have collisions, need to drop with gravity, need to be heated and have higher temperature when there's fire and need to get wet and all that. And that requires a simulation engine. And that was our collaboration with the Omniverse and Isaac Sim within Omniverse. And that was great for us because now the BEHAVIOR simulation engine which we call OmniGibson as an acknowledgement of our collaboration supports rigid body physics as well as deformable objects like cloth and fabric and fluid interactions like making smoothies and complex object states like heating, cooling, cutting and all that thanks to this collaboration.
Great. Yeah. With so many objects and scenes, I think this is almost unprecedented scale of a benchmark for robotics in simulation and now that we have...