Back
Jeff Hammerbacher
Cofounder, Cloudera

Centerstone Research Institute - The Role of Big Data in Health Care

🎥 Jun 13, 2017 📺 Centerstone Health ⏱ 45m 👁 111 views
Interview between Former CEO of Centerstone Research Institute Tom Doub and Cloudera Chief Scientist Jeff Hammerbacher ...
Watch on YouTube

About Jeff Hammerbacher

Jeff Hammerbacher, cofounder of Cloudera and an assistant professor at the Icahn School of Medicine at Mount Sinai, has focused his recent work on applying data science to biomedical research, particularly cancer immunotherapy. In a 2020 talk, he described a relation extraction project on biomedical literature aimed at understanding T cell function and differentiation, noting that immune checkpoint blockade is forecast to generate over $100 billion in sales by 2024 and that over 2,000 clinical trials for such therapies are active. He has emphasized the importance of open science, stating that his lab’s pipeline, epiD, is open source under an Apache 2.0 license and that all development occurs publicly on GitHub to allow others to reperform analyses. Hammerbacher has also spoken about the challenges of translating high-throughput web experimentation to healthcare, arguing that the field needs to conceive of healthcare delivery as a high-frequency, low-cost touchpoint with patients to enable rapid learning. He has criticized the tendency of institutions to outsource data infrastructure to large tech companies, calling such use cases “just marketing” from a Silicon Valley perspective. In earlier talks, he discussed the ethical implications of data collection, stating that “the best minds of my generation are thinking about how to make people click ads” and that decisions about what to measure involve “implicit political, moral, and ethical choices.”

Source: AI-verified profile updated from Jeff Hammerbacher's recent appearances. Browse all interviews →

Transcript (28 segments)
I
Interviewer0:09
I think Jeff for responding and appreciate your being here today. So just a little bit about his background: prior to co-founding Cloudera—and I'll let you share a little bit about Cloudera and what you do there—but he conceived and built the data team at Facebook. I think you were one of the first 50 or 100 employees there, really early on, and built the infrastructure that became the backend for that system. And of course when you talk about data, we don't have big data in this industry—that's an overused term—but we have a lot to learn. In fact, many of the things we've done have been looking at what happens outside of healthcare, what ideas and innovations can we bring back to healthcare and behavioral healthcare. So really coming from large-scale data analysis, he's got a great perspective and can help us there. Currently, in addition to his work with Cloudera, he is very engaged with Mount Sinai in New York, working with them on getting more out of their data. So he's very focused on healthcare and the opportunities data presents, and how science is transforming itself from just a hypothesis-driven research, randomized controlled trial model to a more iterative model where you can gain knowledge from the data in an appropriate and considerate way, then feed that back into science to improve practice. So with that, we're going to have a conversation for a few minutes and then open it up to questions with the group. I'm very excited. Jeff, you'd like to come up?
J
Jeff Hammerbacher2:13
Thanks again for being here today. I think it's probably helpful to start at the beginning. I'd love to hear from you about what led you to Cloudera and what has led you to your current work with Mount Sinai.
Absolutely, sure. I'll try and keep it slightly condensed. I was at Facebook for about two and a half years, and as you mentioned, we built from scratch the infrastructure for storing the data generated by the website for offline analysis. We grew that data store from zero to well over a petabyte of data over the course of a few years, primarily with open source software running on commodity hardware, at much lower cost than we would have been able to do with commercial software. Through that process, I became quite familiar with the enterprise software vendors and startup companies attempting to tackle scalable analytical data management infrastructure. I really felt that what we built at Facebook, in conjunction with a few other organizations in an open source project called Apache Hadoop, would be broadly applicable. So with Cloudera, the idea was to take that software and commercialize it, similar to how Red Hat did with Linux, to bring it to a much broader audience than just consumer web properties. That said, my undergrad degrees are in mathematics, and my background is really in doing data analysis, not building infrastructure software. After about four years of building Cloudera, I became interested in being an end user of the software again. Through a fairly circuitous path, I also serve on the board of a nonprofit called Sage Bionetworks, which is attempting to build globally coherent datasets about disease for modeling and drug discovery, making them openly available. Another board member, Eric Schadt, was hired by Mount Sinai to build an institute applying scalable analytical data management to hospital data. Through talking with him, we carved out a way for me to have two employers: I'm still employed by Cloudera as Chief Scientist, and also at Mount Sinai as an Assistant Professor. I split my time between coasts—at Cloudera I'm tool-building, at Mount Sinai I'm tool-using, and each informs the other.
I
Interviewer5:21
How are you finding the work at Mount Sinai? What questions are you interested in, and where do you see that leading?
J
Jeff Hammerbacher5:28
Early on, there's a tremendous amount of domain knowledge I need to build, so I don't plan on authoring my own scientific hypotheses in the first year or two. The questions I'm addressing are the ones that seem most important to the people in the Department of Genetics, where I find myself. I've been paying attention to projects with real headway that cross multiple labs, participating in their meetings to understand their work and see if the software tools I'm familiar with could help, or if I can build new tools to speed their research. The first two areas that have emerged as cross-lab collaborative problems are personalized cancer therapy—where we do molecular profiling of a patient's tumor to identify novel therapeutic regimens based on data analysis—and hospital-acquired infection. Mount Sinai just merged with Continuum to form the Mount Sinai Health Network, now one of the largest nonprofit medical centers in the US. With seven campuses, about one in 20 people will acquire a hospital-acquired infection during a stay. So better quantifying and characterizing that is a focus. However, my long-term interest is diseases of the central nervous system and mental health. We have leading researchers like Pamela Sklar and her lab, including Sean Purcell and Menachem Fromer, who have done high-profile work identifying genetics of mental disorders. They've found similar copy number variations present in both schizophrenia and bipolar disorder, suggesting a genetic similarity. So I pay attention to their work but haven't yet been able to contribute.
I
Interviewer8:13
When you think about Cloudera taking this data infrastructure and making it a cloud-based service, how does that matter for smaller organizations? How can they leverage findings that occur at scale?
J
Jeff Hammerbacher8:49
At Cloudera we initially built a cloud service hoping it would be our vector of disruption, but the vast majority of customers run our software on servers they own. The migration to cloud for scalable data management and analysis has been slower than the press suggests. But the point is valid: how can small or medium businesses take advantage of this trend? There's an expertise gap—a small business doesn't necessarily want to hire world-class IT staff. Cloud architectures can help them avoid building expertise in non-differentiating areas. Jeff Bezos talks about undifferentiated heavy lifting—no reason every business should have to do that. But there were ways to do that before public cloud, like using hosting providers. The more important trend is when it makes sense for a small business to analyze very large datasets with commodity hardware and open-source software, whether in a cloud data center or one they manage. For the cluster I'm building at Mount Sinai, all in, I'm paying under $180 per terabyte, so a petabyte of data costs less than $200,000—roughly the cost of one fully loaded employee. That's a huge amount of data for a small cost. So the kinds of problems addressed by small businesses expand. Companies like Tumblr had under 30 employees at acquisition and served hundreds of millions of customers. Instagram and WhatsApp had even fewer. Small businesses can now think about what they could do with a petabyte of data, but it takes time to change processes. In my work with infection control at Mount Sinai, we were discussing tracking patients and infections, and the department head suggested starting with one bacteria to avoid overwhelming me. I laughed, because the software can handle billions of interactions daily. We need practitioners and domain experts to evolve their understanding of what's possible, to start asking questions that can only be answered with a petabyte of data.
I
Interviewer12:02
That takes me to thinking about what that means for the generation of knowledge. In healthcare, traditional notions of science focus on the randomized controlled trial as the gold standard. There are new opportunities from data volume, but it's also dangerous, with more spurious findings than ever. How do you balance the scientific method with discovery? How do you see that model evolving?
J
Jeff Hammerbacher12:42
It's evolving. There was a really interesting set of lectures by Sir Michael... Rawlins? a few years ago, where he articulated the biases about evidence in the medical field. We all have qualitative notions of how good different evidence is, but he wanted to quantify that. With RCTs, you hit issues around external validity—the people enrolled in trials often don't match the general population—and post-market surveillance. After an RCT passes, you let your guard down, so you rely on organizations like Kaiser Permanente to find that Vioxx is killing people. We need to rethink scientific evidence to be broader than just RCTs. In the consumer web, we run very low-cost, high-throughput controlled trials—hundreds of thousands a day—so bringing that lighter-weight, rapid iteration from hypothesis to conclusion would be valuable in healthcare. Also, we need robust causal inference in observational studies, with methods emerging from labs like Gary King's at Harvard. The reality is most data we analyze is observational. It's a challenge for statistics and for providers to work with those results. There's no silver bullet—it's not like there's something just over the horizon. We need to be open to evolving what is a historical artifact. RCTs weren't written in stone; they emerged over a century and are now treated almost with religious veneration. In reality, we need to treat them as just one way of learning about the world.
I
Interviewer15:37
As we have new data volumes, methods, and tools to learn from observational data, especially in healthcare, how do you see that intersecting with practice? There are dangers—algorithms can go awry, like flash crashes in the stock market, but in healthcare, a wrong dose suggestion could kill someone. How do we prevent that?
J
Jeff Hammerbacher16:20
Humans give wrong doses quite regularly too. When building models from data, you're usually creating predictions or understanding—trying to evolve a mental model of the domain. For predictions, you can often remove humans entirely. At Medtronic, they're moving to closed-loop models for insulin delivery and pacemakers—no human in the loop. For clinical decision support, a human mediates between algorithm recommendations and treatment decisions. To prevent mistakes like in finance, you need to stay close to the thing you create. If you generate predictions and don't participate in their deployment and evaluation, you'll do much worse. There's a nice talk by Brett Victor about creators needing to be close to what they create—you have to watch it perform to iterate. My experience on Wall Street was that models drifted away from actual trading practices. I'd work on a differential equation, then talk to a trader who had a mental model of a market that wasn't in my model. So staying close to the actual act in the physical world prevents dangerous mismatches between theory and reality.
I
Interviewer19:50
One exciting development is the explosion of data from personal sensors, phones, apps. Most of that sits in silos, but walls are coming down, opening doors to new ways of understanding. In mental health, we have relatively unsophisticated diagnostics, but now with this data, we might be able to tell from biometric devices if someone's depression is worsening. How do you see tackling that problem?
J
Jeff Hammerbacher21:33
It usually takes one visionary organization that imagines value in integrating these datasets. That organization has to build the data model, write software to transform raw data, perform analyses to extract signal, and demonstrate value by altering a business process. Once that cycle is performed, other organizations copy it, and it spreads. There's no shortcut. At Mount Sinai, touring our genomics core facility, the director has been running facilities like this since the 90s. He talked about how things evolved slowly—better measurement technology, better data management, building unified datasets, and tools for inference. Each cycle gets tighter until it becomes automated. The early stages are uncertain because there may be no value, but someone has to take the risk and allocate capital to evaluate it.
I
Interviewer23:50
Our work through Knowledge Network aims to leverage data to advance mental health and substance abuse treatment. From your perspective, where would you like to see that go? What opportunities do we have?
J
Jeff Hammerbacher24:22
Through my work at Sage Bionetworks, which is relevant to what you're doing, Sage started as a set of assets from Merck. They wanted researchers from industry and academia to collaborate in one place with all datasets and lineage tracking. But scientists have habits that make centralized work difficult. Side projects included allowing scientists to cite datasets rather than papers, getting funding for dataset citations, and giving data producers exclusive access for a period. Also, consenting patients was a challenge—we worked on standardizing consent forms so patients could be recontacted and their data used beyond the original purpose. Logistical issues were not technical. When it came to analysis, we found that traditional biomedical expertise isn't necessarily valuable for building predictive models. We joined with the Dream Foundation, which runs contests for predictive modeling, similar to the Netflix Prize. Sage now runs contests to bring predictive modelers to the data. So building datasets isn't enough—you need contests to draw modelers. Sage is on the knowledge creation side; you guys are both creating knowledge and trying to use it to alter practice. On that front, I have fewer ideas, but I think if you try to alter the practice of well-educated, credentialed care providers, there's tremendous inertia. But if you work with community clinics and less credentialed providers, they're more adaptable. That seems promising. Also, at Cloudera we invested early in training and certification; there's a surprising amount of budget for continuing education. Creating real certification programs might shift practice. For example, a 'certified care provider' credential could have value.
I
Interviewer29:52
You mentioned how in the web environment, rapid hypothesis testing is done. How does that map to healthcare? Blue vs. red web page is easy, but how does that apply?
J
Jeff Hammerbacher30:37
The biggest problem is the frequency of exposure to the healthcare system. In the consumer web, we have lightweight contacts with individuals, allowing high-throughput experiments. Healthcare is low-frequency, high-cost interaction. Healthcare is nothing more than a time series of interventions in a dynamical system. One way to intervene is to wait for the patient to visit an office and make an aggressive intervention. A better way would be to sample frequently, find low-cost ways to measure the system, and take lighter interventions. Until we conceive of healthcare as high-frequency, low-cost touch points, we won't do high-throughput experiments. We need to rethink healthcare delivery before we can bring in those tools.
I
Interviewer32:05
We've thought a lot about the next phase of healthcare—payment reform, pay-for-value, large-scale data, mobile sensors. How do you see redesigning the system? It's a big question.
J
Jeff Hammerbacher32:52
My wife runs a full-service seed investment fund for digital health called Rock Health. She worked at Apple in the App Store for healthcare and medical verticals, frustrated by the quality of solutions on ubiquitous computing platforms. She's bringing Silicon Valley design and technical talent to healthcare. Mobile devices and tablets are powerful sensors and vectors for touch points between healthcare and patients. Figuring out how to evolve from decades-old delivery models to a world with these devices is key. The challenge is recasting treatment. I spoke with someone about home care devices for seniors—they provide high-frequency measurements available to clinicians, a preview of what we'll all have. The hard part is whether we're headed for a direct-to-consumer future where companies like Nike dominate healthcare via sensor platforms that perturb behavior, or whether the traditional system evolves to subsume those devices. I'm biased toward the regulated industry, but I don't know. That tension will play out over the next 5-10 years.
I
Interviewer36:00
I want to thank you again. Let's open it up for questions.
A
Audience Member36:37
Did you hear the question? I'm talking about whether at Mount Sinai you see situations where patients collect massive amounts of data that is visually displayed during the patient-provider encounter, and is that making a difference? Is the review and interpretation of information together changing the interaction?
J
Jeff Hammerbacher37:10
I personally haven't sat with a physician and patient and observed that interaction, but in the Department of Genetics and Genomics Sciences at Mount Sinai, we have projects working closely with physicians on joint review of genomic data integrated with clinical data. They're building web interfaces that patient and physician look at together to understand the genome—a clinical version of 23andMe. So yes, empowering collaboration is something we're thinking about. Whether traditional parts of the hospital are rethinking that, I'm not privy to.
That's a great question. If there were a quantum leap, it's probably going to come from the microbiome people—studying bacterial composition of your body. Ninety percent of cells in your body belong to bacteria, and there are remarkable links between the gut and mental illness. The biggest quantum leap I can imagine from my uninformed perspective is something that completely changes how we think about disease in our lifetimes. Genomics work would be more of an evolution, not a leap. The biggest thing everyone is waiting for is biomarkers for mental illness—being able to compute a number and align it to a treatment regimen rather than just talking. One thing that drew me to Mount Sinai is that the Dean of the medical school, Dennis Charney, is a psychiatrist. He's done interesting work on ketamine for depression and finding people robust to trauma. He's invested heavily in neuroimaging and asked me to look at intelligent ways to control for covariates and do matching in observational studies. Some op-eds claim observational studies shouldn't even be considered, but I think there's value in both RCTs and causal inference from observational studies. The tools for causal inference have improved remarkably since the 1970s. We need to update curricula to include those tools from quantitative social science and epidemiology, so we don't only think about RCTs as a way to objective reality.
I
Interviewer42:14
From a time standpoint, we'll take one more question.
A
Audience Member42:29
I'm curious about tools and applications used to directly engage with the patient and collect data. Are there special challenges regarding privacy, consent, and how do you work that?
J
Jeff Hammerbacher43:00
This isn't my area of expertise, but I've sat next to experts. We need to move toward machine-readable representation of consent. Out of our work at Sage, we evolved both the consent form and the representation of what was consented. A nonprofit called Genome Bridge from the Broad Institute is trying to build a centralized dataset around genomics, and they find when they get data from partners, there's no metadata about how patients were consented—so they don't know what they're allowed to do with it. We need a standard for representing consent in machine-readable fashion attached to every data point, exchangeable between partners. In terms of educating people about data use, I'm not a medical ethicist. But the problems I see are number one, giving everyone a standard consent form that's friendly for research—allowing recontact and downstream analysis—and number two, a machine-readable representation of consent that travels with the data for sharing.
I
Interviewer45:10
Thank you very much for taking the time to come down and share your experience, and for thinking about what it might mean for our work in mental health.
J
Jeff Hammerbacher45:25
Thank you so much.