Back
Vicki Cheung
Cofounder, OpenAI

A Kubernetes Ghost Story by Vicki Cheung

🎥 Oct 19, 2020 📺 Camp Cloud Native ⏱ 3m 👁 58 views
Watch on YouTube

About Vicki Cheung

Vicki Cheung, a cofounder of OpenAI, has discussed the challenges and lessons learned from managing cloud-native infrastructure for machine learning research. In a 2020 talk, she recounted an incident where an intern accidentally launched a job that consumed all available compute resources, which she estimated would have cost roughly a million dollars to complete. She emphasized the importance of setting reasonable limits and implementing cost monitoring to prevent such overuse, noting that the infrastructure had been made "too easy to scale." In a 2017 keynote, Cheung described OpenAI's use of Kubernetes to manage a cluster spanning multiple cloud providers and on-premises hardware, scaling up to 4,000 nodes. She explained that the infrastructure was designed to support large batch jobs, some running for weeks, and to allow researchers to scale experiments from one core to 10,000 cores as needed. Cheung highlighted the value of open-source software and Kubernetes' flexibility, which enabled the team to build custom tools while shielding researchers from underlying complexity.

Source: AI-verified profile updated from Vicki Cheung's recent appearances. Browse all interviews →

Transcript (3 segments)
V
Vicki Cheung0:00
Hi, this is Vicki and I want to tell a cautionary tale of how cloud native infrastructure can be super scalable. So this is a story from a number of years ago now. My team had the setup where we had a Kubernetes cluster for running machine learning experiments, and we had set it up so that I had the ability to overflow to other regions based on capacity or availability limitations. So it was actually able to burst to quite a bit of scale, and this is also that machine learning researchers have the ability to launch experiments and burst whenever they need to, and it was designed with usability in mind.
So one day, an alert that was basically like we're out of capacity on the cloud, and I was kind of confused because I thought that we had pretty good access to capacity at the time. And so, you know, going to poke in through the cluster, I realized that basically one job was taking over our entire infrastructure. And so this happened to be in summer during intern season, so it was already sort of like a busier time than normal. And so what we discovered was that our intern was so excited about, you know, how easy it is to scale all this infrastructure that he had launched this job that was just going to consume all the compute that he could.
And so yeah, after talking to the intern, we sort of did a little bit of back-of-napkin math and realized that had he been able to run this to completion, it would probably cost on the order of a million dollars to finish running the experiment, and also consume all of the available infrastructure. So after that, we stopped the experiments and we realized that we had made our infrastructure too easy to scale. So I guess the moral of the story is, you know, it's great that cloud native infrastructure is super easy to scale and also quite usable, but always put in reasonable limits, and also cost monitoring is super important.