Back
Jeff Hammerbacher
Cofounder, Cloudera

Jeff Hammerbacher: Extracting T cell function and differentiation characteristics

🎥 Jan 20, 2020 📺 Ai2 ⏱ 57m 👁 692 views
Many promising cancer immunotherapy treatment protocols rely on efficient and increasingly extensive methods for manipulating human immune cells. T cells are a frequent target of the laboratory and clinical research driving the development of such protocols as they are most often the effector of the cytotoxic activity that makes these treatments so potent. However, the cytokine signaling network that drives the differentiation and function of such cells is complex and difficult to replicate on a large scale in model biological systems. Abridged versions of these networks have been established...
Watch on YouTube

About Jeff Hammerbacher

Jeff Hammerbacher, cofounder of Cloudera and an assistant professor at the Icahn School of Medicine at Mount Sinai, has focused his recent work on applying data science to biomedical research, particularly cancer immunotherapy. In a 2020 talk, he described a relation extraction project on biomedical literature aimed at understanding T cell function and differentiation, noting that immune checkpoint blockade is forecast to generate over $100 billion in sales by 2024 and that over 2,000 clinical trials for such therapies are active. He has emphasized the importance of open science, stating that his lab’s pipeline, epiD, is open source under an Apache 2.0 license and that all development occurs publicly on GitHub to allow others to reperform analyses. Hammerbacher has also spoken about the challenges of translating high-throughput web experimentation to healthcare, arguing that the field needs to conceive of healthcare delivery as a high-frequency, low-cost touchpoint with patients to enable rapid learning. He has criticized the tendency of institutions to outsource data infrastructure to large tech companies, calling such use cases “just marketing” from a Silicon Valley perspective. In earlier talks, he discussed the ethical implications of data collection, stating that “the best minds of my generation are thinking about how to make people click ads” and that decisions about what to measure involve “implicit political, moral, and ethical choices.”

Source: AI-verified profile updated from Jeff Hammerbacher's recent appearances. Browse all interviews →

Transcript (46 segments)
J
Jeff Hammerbacher0:00
I'm ostensibly going to tell you about a relation extraction project we've been doing on the biomedical research literature, but it's not novel enough, I think, to be interesting to everyone in this audience. So I'm going to take a circuitous route to describing the problem to motivate it and hopefully make the talk more interesting and introduce some features of biomedical research that maybe you don't get to see every day in your work.
To motivate this, I'm going to pull a couple of stories out of my far-far past. So after graduating from Harvard a year after Brian, for reasons that are pre-2005 so we're not going to go into them in too much detail, I worked in Wall Street a little bit, but then I went to Facebook. Subsequently started a company, Locally Terra, and actually I used to come out to Seattle pretty regularly because for 10 years I was on the board of Sage Bionetworks, which was a nonprofit bioinformatics that was spun out of Merck. Merck bought a company called Rosetta Informatics, it was Seattle-based. At some point Merck decided they didn't want the Seattle office. The Seattle folks wanted to keep doing the work that they were doing, and so they negotiated with Merck to be able to do that work in a pre-competitive open-source fashion. And they reached out to me for advice on open-source and scalable data management. So that was a fun way to learn about biomedicine, and that organization is still going well. I just live further from Seattle so the board meetings didn't make as much sense.
So I thought that I went and dug into my old slides from 2008, and we actually did some text mining at Facebook. So we had a product that was an advertiser, so the advertising product at that time was called Sponsored Groups, which most people probably don't remember. But you could pay to sponsor a group on Facebook and you get all these analytics about how people are interacting with your group. Victoria's Secret was a very popular one that I can recall right offhand. Oh, I think Red Bull had like a popular sponsored group. So yeah, we put a lot of time into this. And so Facebook Lexicon was this like text analytics product. I think the most interesting thing about it for our work was that it was actually the use case that drove significant Hadoop adoption inside of Facebook. So we ended up building out a 600-node Hadoop cluster to be able to process all of the text that was being generated on Facebook.
And also amusingly, I tried to get an intern that summer to do computer vision over our image collection, and I really should have saved the email from Zuck in which he told me he can't possibly imagine what commercial value there would be in doing computer vision on Facebook's image collection. So I got denied an intern who's actually now the lead product manager on TPUs. So he was the product manager on TensorFlow at Google for a while, now he's managing their GPUs. So yeah, took a while to get religion at Facebook on machine learning.
So yeah, while we were building that, and so you can see we did kind of like state-of-the-art text mining in 2008 at the time. We did sentiment analysis, and I actually remember talking to a lot of leading researchers at the time. They're like, 'Oh yeah, this is an insoluble problem, you know, there's basically like the Bayes error is like 35% and like we're never going to get above that, so stop trying to work on it.' And I actually like, I had to Google, I didn't even realize there was an Indiana Jones movie that came out in 2008. I was like, why did I put this example of Indiana Jones? So I don't know if people, it was the second highest-grossing movie of 2008. So yeah, that was why we showed the example that people didn't like Indiana Jones relative to Iron Man in 2008.
But in any case, I also, when I was trying to like meet some of the best people working in this field, I still just remember a very cool talk from a guy named David Blei who was working on topic modeling. And he'd gotten access to the historical archives of Nature and Science, and so he's doing a really cool thing showing topic modeling across like, you know, 150 years of scientific literature. So I guess that was a bit of foreshadowing of what we would eventually work on. And the reason I found his research pretty interesting, I mean, talking with him, was one of the problems we wanted to work on was new topic discovery. So that was a hard text mining problem in 2008 was saying, you know, how can we, so a lot of the topic models at the time were like fixing the number of topics, and David was working on like some statistical methods for allowing the number of topics to vary. And so we would, it was very interesting to see if we could discover when like a new topic would emerge.
And actually like, I went back and it's still up, you can go to the Cloudera blog, and I wrote the first blog post I ever wrote for Cloudera was showing how to use NLTK within the context of MapReduce to do text mining. So this is to say that ten years ago I was thinking about this quite a bit. And then I kind of felt like we hit the limits of what text mining was capable of doing, and I just kind of put it on a shelf. But it's been percolating for a long time.
So in any case, after Cloudera, I went into biomedicine. And so one of my fellow board members at Sage was a guy named Eric Schadt, and he was hired to run the genetics department at Mount Sinai in New York City. So Eric, I was kind of thinking about doing something. I'd kind of been doing, you know, I went to Facebook, I was using data, I ended up building data infrastructure so that I could use data more effectively. Cloudera, we commercialized a lot of that data infrastructure, but at some point I wanted to return to like using data. I didn't really perceive myself to be a person who built data infrastructure, it's just that the data infrastructure at the time didn't work. So eventually I was thinking about what I wanted to do, and I found biomedicine to be compelling. It's hard to get bored there. I spent, you know, the more, there are more papers in biomedicine that I read and go like, 'Wow, how does that work?' than any other kind of domain in which. So I just wanted like a place to apply data science, and so I ended up starting a lab in New York City.
So I hired up three people. One of these guys, Arun Ahuja, worked with Doug Downey, and so I did a reference call for Arun with Doug, so I met him a long time ago. And then another guy, Alex Rubenstein, I robbed him of the opportunity to make tremendous amounts of money. He finished his PhD in 2014 at NYU where he was writing a dialect of Python, a compiler for a dialect of Python that would be compiled on GPUs, where he was helping his good friends win ImageNet. So one of his good friends was in Yann LeCun's lab, and then ended up starting a company called Clarifai. So Alex was basically sitting next to him like making his ImageNet code run faster. So instead of going into deep learning commercially, he decided to come and work in biomedicine.
So yeah, so we didn't really have a strong thesis about what we were going to be addressing the work on. We were in the Department of Genetics, obviously, so genomics and genetics looked like interesting things to work on, but the application domain, we didn't know if it was going to be clinical, if it was going to be research. We worked on a variety of problems early on. We were software-only, didn't have any kind of wet lab capacity. So we looked at, we had a big cohort of irritable bowel disorder patients that Mount Sinai was working on in conjunction with J&J, which has one of the larger IBD franchises. So we had done like deep genetic sequencing and all kinds of microbiome data generation, and so we were looking for kind of biomarkers of disease progression in IBD patients.
We randomly wrote a paper on how like all of the predictive models for predicting adverse events in patients who are undergoing surgery for cerebral arteriovenous malformations, which are effectively just tangles of blood vessels. And this was actually kind of like, you can use off-the-shelf scikit-learn to outperform every existing, so it was kind of like we were setting like a lower bound for what machine learning should do in medicine with that paper, which I feel like is probably not done enough. So we were competing against all of the proposed scores that existed at the time for predicting adverse events during this, and this is still I think our second most cited paper from the lab.
We worked on C. diff infection, which was actually a null result, we didn't get a chance to publish it because that's how biomedicine works. But I thought I found it very interesting. So C. diff was a rapidly growing cause of death amongst people age 65 and above. It's a kind of hospital-acquired infection and it can be very fatal, it dehydrates people. So we were looking at C. diff infection. So we took patients that were admitted to Mount Sinai who were diagnosed with C. diff while they were in the hospital, and we could take microbiome samples and sequence the microbe and look at the genome sequence of every C. diff bug that we saw. And so we had like 50 or so patients, and the theory was that C. diff was being transmitted within the hospital because when these people were admitted they didn't have C. diff infection. The reality was we found 50 different genomes of C. diff. So it looked like there's actually a large community reservoir of C. diff bugs, and there's something about the hospital environment that induces the infectious state of that bug. So that was a really interesting project to work on.
We actually got to some like RFID real-time tracking of like every instrument and doctor in the hospital to be able to try and trace infections, but then it turned out there were no infections to trace. And then ultimately we settled on this work in cancer immunotherapy that I'll tell you more about. And we worked on a neoantigen vaccine. So cancer is a disease of the genome. When the reason why cells start growing out of control and forming tumors is they accumulate mutations in their DNA that aren't present in the other cells around them. Some of these mutations may in fact be immunogenic, they may alter the proteins that are expressed within that cancerous cell in such a way that the immune system recognizes them and determines that that cell should be eliminated. And these mutation-derived tumor antigens are referred to as neoantigens largely in the literature.
And at the time we got to work with a principal investigator named Eina Bhardwaj who was taking an approach in which she was giving a therapeutic vaccination. So she would take a tumor sample, take a blood sample from the patient, identify mutations that were in the tumor not in the blood, and then she would find the most immunogenic of those mutations and construct a vaccine which contained the short amino acid fragment containing that mutation, or say the top 10 mutations in that tumor that are most immunogenic. And then you would vaccinate the patient in the hopes of generating an immune response that would clear the tumor. So I say this, that she wanted to do that, she did not have the capacity to do it. So I looked at it and I said, 'Oh, this is kind of a cool computational problem.' And so we actually wrote a pipeline which is used in now three different Phase 1 clinical trials. So our software is actually, you know, taking data generated from the sample of these patients that are in the clinical trial and then formulating the vaccine. We actually send the request off to a peptide synthesis vendor who then sends it back to Mount Sinai, and that is what the patient receives. So that was a really fun project to work on.
Several different companies have started up to commercialize this strategy. Neon Therapeutics is the one that we work closely with. So our software runs in pretty much all of those for-profit versions, and many of them have gone public and they're starting to read out clinical trial results now. So that was a big part of what we worked on. But then we grew up to be about 16. We wrote a lot of open-source software on GitHub, Sledgehammer Lab, probably way too much actually. A lot of it's in OCaml, which I think people looked at as funny for doing that, but it was fun.
And ultimately we ended up working on a project, my wife and I, of having a baby. And I just want to share this because I miss him, I get to go home tomorrow. But this is our guy Bear, and that's my dad. And my dad lives in Charleston, and we wanted more room to hang out with Bear. So I ended up winding down the lab in New York and started up a small lab in Charleston, South Carolina. Different focus this time, much tighter focus. We knew we wanted to work in cancer immunotherapy, we wanted to primarily be a wet lab, not really do much software, and we really wanted to focus on the biology of T cells. So T cells are the effector cells of the cellular immune system, they're the ones that are actually directly killing cancer cells. And so we really want to understand the biology of T cells and how they interact with pretty solid tumors. So that's our lab, it's pretty big, there's a lot more room in Charleston than in New York City.
So yeah, so that's the context in which we're doing this work. Now I'm going to tell you two different scientific problems that brought us to doing text mining to try and solve them. So one is the problem of finding biomarkers for response to immune checkpoint blockade. This is something that people spend a lot of money on. So I try and convince people that work in like AI or cloud computing that what they're doing is actually far less commercially interesting than immune checkpoint blockade, because the scope of what is happening, I mean, checkpoint blockade is outrageous. So you can see the projected sales of just this one category of cancer therapeutic that did not exist prior to 2011 in terms of the ability to prescribe for a patient.
So the very first immune checkpoint blockade agent, so this is kind of a category of drugs, like these drugs all work in a similar fashion, they all target interestingly T cells, not tumors. So you're not actually treating cancer, you're treating the immune system, which then itself will treat cancer. So the very first immune checkpoint blockade agent was approved in 2011. The most successful immune checkpoint blockade agent, Keytruda, was approved in 2014. So I know it as pembrolizumab because I worked with it before it was approved, and it didn't have the name Keytruda until it was approved. And most significantly, Keytruda was approved for the first time for a non-tissue-of-origin-specific indication in cancer in 2017. So this means you can come to your doctor and no matter, like often oncologists are divided by the tissue of origin that they treat. So generally when you get a cancer drug approved, it's approved for ovarian cancer or breast cancer or prostate cancer. So Keytruda was the first approval for a biomarker that could span many different tissue-of-origin types in solid tumors. And so in this case it was for mismatch repair deficient tumors.
So all of your cells accumulate mutations all the time. Luckily you have proteins in your cells that are repairing those mutations, but some people are born with issues or their tumors develop issues in their mismatch repair machinery. So if you've heard of BRCA as a gene for like breast or ovarian cancer, that's a common mismatch repair deficiency. And so there are other forms of that. And so for anyone who has a tumor which has a mismatch repair deficiency, Keytruda is indicated for treatment. That's a very broad population, and so that's what's driving it. So it's forecast to become the number one selling drug across all indications, not just cancer, by 2024. And in that year it's expected to generate about 17 billion dollars in revenue. I think it's going to do about 10, I believe, this year alone. So 17 is not a huge crazy forecast for it. And then this entire category of immune checkpoint blockade is forecast to do over a hundred billion dollars in sales in 2024. So this is an incredibly large commercial space.
Another way to conceive of that is there are over 2,000 active and rolling clinical trials for immune checkpoint blockade right now. This number is as of late 2018, it's I'm sure higher now. And then the total number of patients that are expected to enroll in those clinical trials is over 400,000 once they all meet capacities, are multi-year later-phase clinical trials for the most part. So this is really a massive, massive scientific enterprise right now that's completely revolutionizing how we treat cancer.
So I'll give you a quick understanding of how it works. I mentioned that it treats the T cell. So the pointer disappears, but in any case, if you can see all the way down there, the immunotherapy, one of the checkpoints, so an immune checkpoint is a protein which is expressed on the surface of a T cell, which when it is bound by its ligand will enhance or diminish the T cell's activity against the cancer cell. And so the first checkpoint protein against which a drug was approved is called CTLA-4. And so CTLA-4, you can see on the left of this figure, it's co-expressed on a T cell alongside of the T cell receptor. So the T cell receptor is what determines the specificity, that's what determines the T cell whether or not it's going to kill a particular cell. But if CTLA-4 is expressed and it binds its cognate ligand, then it is going to actually cause the T cell to inactivate and not kill the cells. So it's a very important, your immune system has a lot of ways of turning off the T cell response because autoimmunity is a very scary thing. So we need this. So anti-CTLA-4, ipilimumab, or Yervoy, still sounds weird to me, I know it as ipi, binds to CTLA-4 to cover it up. So when the T cell goes to do its work against a tumor, it doesn't have that off switch anymore, it's being bound by the antibody. So that's how immune checkpoint blockade at a high level works.
So this is the figure that caused me to work on cancer immunotherapy. So it's a low-res clinical trial result figure. So these are very common figures that you'll see if you read clinical trials. So the way to read it is you start with 100% of the patients in the clinical trial are alive at time T equals zero. As you go along the x-axis, that is time, and as you go down the y-axis, that is death. So the y-axis is the percent of surviving patients in the clinical trial. So I, you know, this is very motivating when you're, the data you're working with is human lives and survival. So first of all, that really caught my attention. But what's most interesting here is that, so for the first three months, immune checkpoint blockade does not appear to do anything. So the two lines that you pay attention to here are the purple one and the grey one. I won't explain what the red one is, but basically the purple one is somebody who got a checkpoint blockade, that's the arm of the clinical trial in which they were treated with checkpoint blockade, and then the grey is the arm of the clinical trial in which they were treated with standard of care for melanoma. So no difference between standard of care and immune checkpoint blockade in the first three months.
In the first two years after the first three months, you start to see some separation, but this is not uncommon. This is the kind of separation you see with a lot of different anti-cancer drugs. So a lot of anti-cancer drugs work for a few months or sometimes even the first year or two, but then you start to see survival curves collapse. You start to see the tumor develops resistance, maybe the patient develops another form of cancer as a result of the treatment. But in any case, you generally see survival curves collapse. But what's magical about immune checkpoint blockade is what's happening from years two onwards. And so this figure is quite old, this figure actually extends beyond 10 years now, and this plateau has roughly held for 10 years. So what has happened here is this is a durable remission, or as we might say, a cure. So these people, about 20% of the original clinical trial population for the Phase 3 clinical trial for ipilimumab, they're surviving, they just keep surviving. So they've effectively been cured of their cancer.
And so melanoma, this is what happens on standard of care with melanoma, this is crazy. So if you talk to a melanoma doctor from prior to 2011, you know, it was just, well, this is advanced metastatic melanoma, so it's a subset of melanoma, but effectively this just goes to zero. Like you had below 5% five-year survival in advanced melanoma, and ipilimumab lifted that above 20%. So this is a really remarkable result. And this durable remission is really unique to immunotherapy. There are a very small subset of other forms of cancer drugs that can deliver this, but really this is kind of a signature of immunotherapy because you are treating the immune system, and the immune system itself is adaptive. That's the theory. We still don't understand fully why we get durable remissions, but so this is really remarkable.
But the problem with this chart is it's only 20%. So the question that everyone, the literal hundred-billion-dollar question that everyone's been asking is how can we distinguish the 20% that respond from the 80% that don't respond, and what can we do to shift this curve upwards? How can we take the 80% that don't respond and make them into responders? So we worked with a lot of the pharmas running these clinical trials, we got their datasets, and we worked on this problem of biomarker discovery for checkpoint blockade. And one of the earliest biomarkers that was proposed was a very simple one. It said, how many mutations do you have in your tumor? If you have greater than 100 mutations, you're actually much more likely to respond. So this is that curve, the blue, and then if you have less than 100 mutations, that's the red curve. So that was a very interesting and early signal. And this is generally known as the lottery ticket hypothesis, where every mutation you get is a lottery ticket for whether you're going to have a protein that's immunogenic and targeted by your immune system. It's a controversial hypothesis, but that's generally how people explain this chart. And so this kind of kicked off, so this is known in the field as tumor mutation burden, or TMB. It's a proposed biomarker. It's quite related to the mismatch repair deficiency that we talked about earlier, in that mismatch repair deficiency leads to an increased tumor mutation burden. So we're kind of coherent, a nice post facto story that we can all convince ourselves is true.
But so our group, we tried to kind of bring more rigorous statistical methodology in and open reproducible science into the identification of immune checkpoint blockade. And so now we're finally getting to the point where I tell you about the scientific problem that we wanted to use text mining to help solve. So this is, so when we're thinking about biomarkers for immune checkpoint blockade, a lot of the proposed biomarkers have to do with the composition of what's called the tumor microenvironment, or the TME. So a solid tumor, when it forms, it forms for years. So if you look at, you know, the Japanese survivors of the atomic bomb dropping, and they didn't get cancer for tens of years afterwards. So cancer is a really slow-growing disease, it accumulates over time. And so these tumors, the solid tumors that grow, they've been growing for often years if not decades, and they have radically transformed the tissue microenvironment around them to better support their existence. And they've done a lot of things to manipulate that microenvironment to make it hard to kill them and to make it easy for them to grow.
And when doctors and researchers talk about the tumor microenvironment, they talk about it in terms of its constituent components. And probably the most important component of an immune microenvironment is a cell type that lives within it. So if you look at a lot of these labels on this picture here, you see things like tumor cells, macrophages, T effector cells, myeloid-derived suppressor cells or MDSCs, antigen-presenting cells or APCs, T regulatory cells. So these are all cell types. So people have been working to understand tumor microenvironments in terms of the cell types that live within them. And then there is a posited mechanism by which those cell types cause this tumor microenvironment to be, generally people say cold or hot. If you have a cold microenvironment, that's one in which there's not a lot of immune activity and you're probably not going to do well on immune checkpoint blockade. If you have a hot microenvironment, it's one where you have a lot of effector immune activity and we expect you to perform well on immune checkpoint blockade.
So decomposing the tumor microenvironment into the individual cell types and counting the number of them is a big sub-problem within the larger problem of biomarker discovery for checkpoint blockade. And one of the ways in which you decompose the cell types present in the tumor microenvironment, and likely the most popular, is you take a tissue biopsy, you create a slurry from the cells in that biopsy, you extract the DNA or the RNA, generally the RNA, from that slurry, and then you send that through a sequencing device. I was up until a few years ago, now it's actually quite common to do individual single-cell sequencing, but we were primarily working with these bulk RNA sequencing datasets. And so what people do is they take individual cell types that they believe to be pure samples of just a single cell type, and they take that pure sample of say dendritic cells, DCs, and they take a pure sample of dendritic cells, they get all the RNA out of that sample, they run it through RNA-seq sequencing, and then they count up. So RNA sequencing, the dataset is basically you have like 20,000 or so genes. So imagine as a matrix where it's like sample on the rows and on the columns you have a gene, and then in the cell is some intensity measure of like how much of that gene is expressed in that cell type.
So for each purified sample of individual, say, ontologically real single cell type, we're going to generate an expression signature effectively. Like we can do that for like say ten samples and just average them, and then that could give us what we believe to be sort of the expression signature for that cell type. So what we expect to see this much of gene 1, we expect to see this much of gene 2, etc. If you want to get fancier, you might want to not just do averaging, you may want to get spread and other forms of statistics about it. But that's effectively how people, so you now have, so if we have 29 immune subsets, we now have 29 different signatures of that cell type. And now if you hand me the RNA-seq data from a tumor biopsy, and the RNA sequencing that was performed in that tumor biopsy, I can take these 29 different cell type signatures, and there must be some kind of like linear composition of those 29 signatures that could cause the data that I'm seeing for the bulk RNA-seq for the tumor microenvironment. And so that problem of trying to figure out what are the weights that we should put on the individual cell types, what is the amount of each individual cell type that we are seeing in our bulk tumor microenvironment RNA-seq data, is generally known by the problem of deconvolution.
So we want to deconvolve the tumor microenvironment into counts of the number of cells of different subsets, in this case 29, how many of each of those 29 subsets is present in this particular tumor microenvironment. And now that becomes a set of features about an individual patient that can be fed into a biomarker discovery algorithm. And so our group worked on this, we did some fun stuff with hierarchical models and Stan to infer these mixtures. You might think that it's an ill-posed problem, and it kind of is. So sometimes you have to bring inside information to try and get at it. But this was a fun problem to work on.
But so this is the first scientific problem that led us to start thinking about text mining, which was one way to define cell identity is through its expression signature as discussed here. But where, how did we pick these 29 immune subsets to then purify and generate expression signatures that we can then deconvolve against? These 29 subsets are folk knowledge that's encoded in the scientific publications of immunologists. So the 29 subsets that were gathered here, it was maybe they had a flow cytometry marker panel that they could use to isolate them, and that's why they chose those 29. But how could you principally choose the N cell types that you would like to deconvolve from the tumor microenvironment? That's the one phrasing of the scientific problem that we were interested in solving.
The second scientific context for making use of this, for doing this text mining exercise, is adaptive cell therapy. So immune checkpoint blockade, big story, very commercially successful. Adaptive cell therapy, probably more scientifically interesting. So immune checkpoint blockade, like I said, it delivers 20 to 30% durable remission rates in solid tumors, which is exciting, particularly when those durable remission rates were below 5% previously. But adaptive cell therapy generates much higher durable remission rates, but for a much smaller set of tumor types, and in fact primarily just B-cell lymphomas for now are the tumors that we can use adaptive cell therapy to cure. But the remission rates are amazing, they're 80 plus percent. And so adaptive cell therapy was approved by the FDA very recently, and it is yet to really make significant alteration of clinical practice because it's very difficult to deploy. What you do is you have to pull white blood cells, or T cells, out of the cancer patient. For now, potentially there's some world in which we could actually pull these out of healthy subjects. So you pull T cells out of the cancer patient, and then you do some form of genetic engineering on those T cells to make them better at killing the tumor. In this case, for the clinically approved, for the FDA-approved adaptive cell therapy, what they do is they take what is effectively a receptor which is specific for a surface protein...
On B cells called CD19, they stick a CD19 receptor—they just glue it to the outside of a T cell—and then they put all of the machinery for the T cell firing and killing of cells below the signaling apparatus of that receptor. This allows it... So what we do when we do adoptive cell therapy for B-cell lymphomas today is we're not actually treating cancer; we're killing all the B cells in your body. We're killing everything that's CD19, and it just turns out that you can replenish your B cells, so this is an effective strategy. The reason why it hasn't worked for other forms of cancer is we generally just can't get rid of, for example, all of your skin cells—that would be a bad move. However, it's generating really remarkable remission.
So there's a big question in adoptive cell therapy: which T cells should we use to actually do the genome engineering on and then send back into the patient? There's been a tremendous amount of analysis on what subsets of the T cells are most fit for the long-term clearance and keeping out of the tumor. So people have proposed... This is within the context of tumor micro-environment deconvolution. We care about many immune cell types in addition to T cells. In this context, we're really just taking the T cell branch of the cell family tree and just trying to get a high-resolution view: what are the subsets of T cells in this case based on their functional performance in adoptive cell therapy? So this is pulled from a paper where they've assigned a certain set of T cells—naive, stem cell memory, and central memory—they say these are more effective cells to start with for adoptive cell transfer. And then they say effector memory and effector T cells, those are worse. So fine, how do you define those cell types? In this case, they give you surface proteins that you can use to define them, but they also tell you that where do those cells come from? Well, they come from—they get different forms of signal strength from the antigen-presenting cells. So these are different ways of defining what different subsets of T cells are that are different from their expression profile. And these are actively big—they're huge debates. They're like scientists that won't talk to each other at conferences because they believe that this cell belongs here on this chart and not here. Like, I genuinely know two people who won't talk to each other because they believe that to be true.
And I also just kind of wanted to show, like, I really like the intersection of geometry and probability theory, and someone has kind of snuck differential geometry into the world of single-cell genomics now. So there's this guy who posited that something existed called the Waddington landscape, which said we start with a pluripotent cell—in this case, it's a naive T cell—and we say that this cell has an opportunity to become many different kinds of cell. And as it rolls down this hill, it's kind of losing potential energy; it's losing its ability to become other kinds of cells. So it's making these fate choices that cause it to kind of close off paths of differentiation to itself. So if it chooses to become a T helper 2 cell, then it probably can't ever become a helper 1 cell. That's one of the things that the Waddington landscape posits. So this is another form of cell identity that is actively discussed: are there kind of points of no return in cell identity? And so we should be able to use that to inform our deconvolution exercises as well.
So last thing I'll talk... So that's the second scientific problem: what's the best T cell subset for adoptive cellular therapy? And that begs the question of what T cell subsets exist. So we wanted to survey literature and see what T cell subsets exist, which ones have kind of solid footing amongst immunologists and are frequently referenced and have identities that seem stable, and then which are kind of emerging T cell subsets that may be only a subset of labs are talking about, maybe its identity is unstable. And so last thing I want to talk about is how are we gonna figure out which T cell subset is better at adoptive cellular transfer without having to, like, stick it into a human? Well, right now you stick it into a mouse model of cancer, and that'll take you three months to figure out if you have a better T cell for ACT. So the lab next to us at MUSC, they are working on ACT and they have ideas about better T cells for ACT. And when they have a new idea, they buy some mice. There's a giant—the fourth floor of our building is just like a giant mouse... or whatever you call it. They raise the mice up, they inject some cancer cells into them, they let the tumors grow, and then once the tumors reach a certain size, they inject the T cells into that. Or they perform this whole dance of getting the T cells out of the mouse, genome editing the T cells, putting them back into the mouse, and then they measure the tumor growth or shrinkage for a few weeks. And that whole process takes about three months, and it also involves, you know, taking animal life, which ethically I don't love.
So one of the other things we work on in our lab, which is also kind of a hot area of research, are these three-dimensional model systems, often called organoids. In our case, we call them spheroids because a tumor isn't necessarily an organ. And so this is an actual picture taken by one of the microscopes in our lab. So you can actually—you can take cancer cells and you grow them in certain conditions. In this case, it's an incredibly simple method called the hanging drop method. So you literally just put dots on the bottom of one of these dishes, and then you seed each one of those dots with a few cancer cells, and then you just flip it over so they're all kind of like hanging from what is now the ceiling, informally the bottom of the dish, and you let them sit like this for a day and they start to form a spheroid. And so then after 24 hours, you harvest the spheroids and you put them into... this is just a microfluidic chip in this case. The interior of this microfluidic chip contains a gel which is intended to mimic the extracellular matrix in which cells are embedded in the body. And so you put your organoid in here—you can barely see them, there's one, there's one, there's another—and then these are channels into which you can flow various liquids. In our case, we flow T cells into this. And so the T cells will come in... well, let the organoids grow out until they're like big enough to mimic the structure of a tumor, and then we'll put in the T cells that we want to test, and those T cells will infiltrate that gel and they'll try and kill the organoids. Over there are some Z stacks, you know, like basically going down the Z axis, and so you can look and see what these organoids look like. So they truly are like three-dimensional models, and this is very important because solid tumors—I mentioned that they grow over the course of years or decades—and they really make use of the three-dimensional structure to make it more of an obstacle course for T cells to get in. So there's a lot of things that happen in three dimensions that don't happen in two dimensions, and it's one of the reasons why things that work in a petri dish don't work in humans. So this is an attempt to get a more representative model.
So this is a little video of T cells trying to kill cancer cells from our lab. And so what we're trying to do—a separate project that we worked on—is actually use computer vision to automatically quantify. So today, the people who do organoids to study ACT, they take these pictures and then there's a grad student somewhere counting the green cells versus the red cells. The red cells are the dead ones, the green cells are the live ones. So we've done some computer vision work to automate the counting of, like, live versus dead. So this gives us a higher rate of iteration on determining whether we've found better T cells for killing tumors. So, T cell relation extraction... made it 40 minutes. So this is the actual text mining problem.
So now that you understand that we're going to different problems—biomarker discovery for checkpoint blockade and better T cell subsets for adoptive cellular transfer—both of which have at their base a requirement that we understand what subsets of T cells exist and how are they defined in the literature. This is a very common diagram you'll see in papers. So you'll say this is related to that Waddington landscape that I talked about. So naive T cell is kind of this like progenitor T cell state. So that's like the T cells are circulating through your lymph and blood system right now; we believe most of them to be in this naive state. And then they can become many different other kinds of T cells. How do they become? So this is a T helper 1 cell, or Th1. How do you become a Th1 cell? Well, we have posited that if you are exposed to these cytokines—a cytokine is kind of a protein, it's a signaling protein—so IL-12, IL-18, or interferon gamma. If you are put into a petri dish, if you're a naive T cell and I surround you with these cytokines and I wait, you're gonna become a Th1 cell. That's one theory. Another way to describe it is within the genome, within the nucleus of every cell, the genome, there are certain proteins which are called transcription factors which can turn on the expression of certain proteins. And so there are certain transcription factors which are believed to be kind of gating transcription factors, where once you've turned on the expression of that transcription factor, you have kind of unlocked this transcriptional program that's going to send you down a pathway to becoming a specific cell type. So T-bet and STAT1 are transcription factors that are believed to bias naive cells towards becoming a Th1 cell. So if you see those being expressed inside of a cell, you think this cell is on the way or is a Th1. And then the third way of thinking about T cell identity is through what cytokines does it secrete. So these cytokines that cause the transformation of the cell, there's kind of a feedback process in which the T cells themselves have secretion profiles. When stimulated, they'll send out into their near paracrine environment, they'll send out different kinds of proteins. And so in this case, interferon gamma is both a cytokine which induces Th1 cells, but it's also a cytokine produced by Th1 cells. So obviously you now have a positive feedback cycle there.
So we went looking to see these three kinds of relations: can we extract these from literature? Can we find inducing cytokines for a specific cell type? So we look for sentences which have cell type and a cytokine in them, and we try and find examples in which that cytokine is being referred to as one that induces that particular cell type. We look for inducing transcription factors—this is a different class of proteins, comes from a different ontology—and then we look for secreted cytokines. So these are... and these are, I think, BRAT annotation diagrams, or that comes out of spaCy. Yeah. So yeah, our poor first author has spent more time in BRAT than he'd like. So the corpus that we're using, we built a 10,000 document corpus just by querying PubMed, and 7,000 of those 10,000 had full text available; 3,000 just had title and abstract. We went and ran kind of a simple named entity recognition pipeline over them. This is where spaCy enters the picture. So we use one of the four kind of like specific named entity taggers that spaCy made. It's JNLPBA, which is an acronym that I frequently fail to expand correctly or even state correctly, so I had to read it there. But in any case, JNLPBA helps us identify proteins. Well, we had to define our own patterns for cell types; there wasn't really a good cell type named entity recognizer out there. And so we ended up with 160,000 sentences—or 160,000 cytokine either secretion or inducing and transcription factor candidate sentences. Sorry. And we went and we did some manual annotation on these sentences, which derive from about 90 of those documents. So just sitting in BRAT and labeling them and seeing if that was right. So that ended up giving us about 280 cell-cytokine relations and about 120 cell-transcription factor relations.
These are some examples that have been cherry-picked, because that's best practice when showing that natural language processing can work. So yeah, so you can see, like, secreted cytokines: IL-17 is mainly produced by Th17 cells, so that's a pretty good example, I think. The established role of IL-2 in T regs, one may not be as obvious to me, but yeah. So these are some of the sentences that we're staring at and trying to pull relations out of. And so these are some charts that we created. This one has a typo in it, so this presentation I don't love, but the idea here is that this direction tells you the intensity of that cytokine for that cell type in terms of induction relations, and this direction tells you in terms of secretion relations. These aren't actually opposed relations; it was just one way of getting two figures on that. We haven't... this is still a kind of mid-stage project, so we haven't optimized for this. And so in this case, you can start to see: so IL-12 and Th1 in terms of induction feels like a very well-established scientific fact. But then things start to get more interesting in the middle here. So IL-23A, I'm actually kind of interested in this one. What's that doing? Like, some people have said it, but clearly not a lot. IL-27, IL-6 is actually a cytokine which has many functions, and so that's interesting to see that show up there. So we're kind of using this to do kind of hypothesis generation to figure out what do a bunch of scientists agree on. They agree that interferon gamma is secreted by Th1 cells for sure, but they may not agree that IL-4 is secreted by Th1 cells. So then that dictates a potential experiment that we can run.
So we also looked and tried to see: are there cytokines which are more likely to be inducers versus cytokines that are more likely to be secreted? This is just an analysis that sometimes people put in a textbook, so we're just reproducing a figure that is built from other forms of data. We're just... this is kind of a version of the figure which is text mining. So surprisingly, most of these cytokines kind of live on the y equals x axis; that is, there's no real bias between inducing and secretion. Like, if a cytokine is often secreted, then it is probably also involved in the induction of T cell fate choice, which is, I guess, a seemingly reasonable response. To find transcription factors, so we were looking at for each cell type here. So like for T regulatory cells, this is kind of a really famous example. So people have won very big prizes for this discovery that Foxp3 is a transcription factor which induces this regulatory T cell. And so it's good to see that our data support that conclusion. But as you go down here, you can see there are certain cell types, like these T helper 9 cells, which are less well-defined, for which the intensity of the speculation on the transcription factors that kind of gate the identity of that cell is lower. So there's probably someone out there who's thinking like, if I can only find the transcription factor that gates the identity of a Th9 cell, then I'll be famous. And the same goes for inducing cytokines. So transcription factors, cytokines. So yeah, we spent a lot of time kind of staring at this, and we've actually used this to inform some experimental design that we do. So for example, one of the things that we do is there are these panels that measure cytokines secreted by individual cells. And so when we think about which cytokines do we want to measure, we've actually used this chart to determine what we're going to measure in the dish.
So for relation extraction, we knew that we were in the domain of learning from limited labeled data, and we wanted to play around with some of these new ways of doing learning from limited labeled data. So I know Chris Ré from a past life; we hung out at 2007 VLDB in Vienna when he was working on probabilistic databases. That was an interest of mine at the time, and I think he was a grad student up here at UW. So Chris Ré is now a PI who won like a MacArthur grant and sold a company to Apple and things like that. So his group has a kind of line of research that is, I think, many of you are likely familiar with, called Snorkel, and they refer to it as like data programming. And it's this idea where instead of labeling your data, you should try and define these sort of loosely correct labeling functions. They're kind of rules where you could take as input a sentence, a candidate sentence, and you could spit out a noisy label. And then you just like hand a big laundry list of labeling functions to some kind of adjudicator, and there's a generative model built that determines how to put weight on those different labeling functions using things like how much the different labeling functions overlap, etc. It's probably not a great description of Snorkel, but in any case, in our hands, it's a Python library that we're using to write functions that map from sentences to labels. And when we look to see how our relation extraction performance improved because of labeling functions, what we realized is we didn't have enough labeled data to determine this. So we were doing these labeling functions, they were looking good on the validation set that we had held out, and then we had held out a test set, and then we went to see if that performance transferred onto the test set. We weren't really seeing that performance transfer. So we looked at it, and we only had like a few dozen examples in our test set. So we now need to, you know, once more into the breach, go label some data. But we do have a lot of opinions about labeling function writing. So Eric, the first author on this work, actually went out to the most recent Snorkel workshop, and we spent a lot of time talking with them about the fact that, like... so this was our process for labeling function writing. I don't know, has anybody in the room written a labeling function? Okay, so it's not ready for primetime, I think is what I would say.
So what we did, our methodology was we took those like 90 or so candidate papers which contain candidate sentences, we took a hundred or so candidate sentences, and we manually labeled them. And we just kind of like took notes, really. Oh yeah, like what did we use in this sentence that allowed us to make that decision? And then we had a laundry list of many, many rules. We wrote the actual Python code to code up like 50 or so of those rules. In the process, we realized a lot of these rules were kind of like slight variations on each other, and we could figure out a way... and a lot of them were just using like regular expressions to look for things like, you know, 'this is induced by that' or like some synonym of 'induce'. So we ended up building for ourselves our own thesaurus of like synonyms for 'secrete' and 'induce'. And then another thing that we did that we felt kind of clever about is that we realized like secretion and induction are like relatively opposed, so an example of one is probably not an example of the other. So we started using labeling functions from opposing tasks as negative examples. There's a part of me that feels like the Snorkel sort of adjudicator should somehow figure that out, but this actually did lift performance on the validation set, so we feel like it was a good idea. And then we started throwing in like little dumb heuristics, like how far the entities from each other in the sentence, does an entity co-occur between the two? And these also started lifting performance a little bit. And then we spent a lot of time on the phone and in person with Alex Ratner, the kind of lead author around a lot of historical papers, just sort of looking at... he's doing sort of like manual inspection of our labeling function metrics and telling us like, okay, you've done a good job here. Like, so there's clearly some kind of like augury that is still required for doing labeling functions, and we still are unable to tell you whether that's worth doing. But what I can tell you is that like there needs to be better support for... there needs to be some kind of like intermediate kind of labeling function IDE or instead of libraries that facilitate writing labeling functions for like common tasks. So I don't know if that's something that you all have encountered, but that's an area that we're hoping... like, I told Eric, I feel like the world... so I used Spark very early on, Apache Spark, and it just didn't work. And I didn't... I would go to conferences and people would be giving these like glowing talks about how great Apache Spark was, and in my head I'm thinking like, I feel like I'm being gaslit. Like, there's like all I get is like debug spew from Apache Spark when I try and run it. And so I thought it was like important, so we started giving talks about like how Spark just didn't work, and then we got all these emails of people being like, oh my god, it doesn't work for me either. So I kind of feel like Snorkel is in like a similar place, where like Eric was at a workshop at the AAAI conference and he gave a talk on this work, the state of the work at the time, and he was like, I kind of felt like it wasn't interesting people because like everybody's kind of like already mentally mapped like, Snorkel, okay, that's a good idea, we've all moved on. But it just feels like the actual ergonomics, the experience of using it, is actually kind of far from its impact on the literature and how people are treating it when they cite it. And so we're just trying to like... so I like finding these things where you have a technology which looks like it's obviously useful, people are very excited about it, but the reality of actually like using it in your hands is quite painful. And starting to kind of like pave the way, because it's likely there's signal here, it's likely that what Snorkel is doing is a good idea, but it's also likely that there's a gap between the research community and trying to use it in industry in production. And I like trying to kind of like interpose myself in that place.
So what we've done recently is we've worked on named entity linking. So biology, one of the nice things about biology is that it has these very rich, well-curated ontologies for almost every entity type you can think of. So for cell types, there's the Cell Ontology; for proteins, there's the Protein Ontology. Both of those are OBO Foundry ontologies. So this is kind of a consortium of ontologies that have all agreed to use the same top-level ontology, BFO. There's actually a really good book by the guy that formulated BFO, Barry Smith. I ended up like reading like horse roll, you know, to understand what was going on in BFO, because like he's like a philosopher, so it's fun reads. But in any case, this is what it looks like. So this is like a Cell Ontology. So this is a mature T cell, and it has mature gamma-delta T cell, alpha-beta T cell, and memory T cell as children. And so what we want to do is we want to take these entity mentions that we're identifying and we want to map them back—you all know the problem setting of entity linking. And so these also contain like synonyms and descriptions, so we can use that as some kind of entity representation. And so for cytokines and transcription factors, there actually doesn't exist a well-curated ontology. The cytokines and transcription factors are proteins, so we could use the Protein Ontology, but there's no like annotation within the Protein Ontology which says this is a cytokine and this isn't. So people have built these kind of controlled vocabularies or spreadsheets, if you will, of all the cytokines and transcription factors that exist in the world. So we're just like importing those as CSVs and linking against them. So we tried to use the NormCo, so it was like a best paper at AKBC, NormCo. So for this problem, they were trying to do disease normalization. So they just kind of the obvious thing: they took the surface form of the entity mention, they tokenized it, and then they just did a word embedding of each token, and then they just like added those vectors—maybe they averaged it, maybe they multiplied it, I can't remember the exact form of how they mashed them all together. And so we did that, and then this is a UMAP embedding of those vectors that come from the surface forms of the entities. And so we started seeing this looked pretty reasonable, but we started... so Eric, in his diligent way, went through and started looking at like the different clusters and started looking at, starting to like categorize the tokens that were used to kind of like distinguish different clusters. This is always a fun thing. And so like you look in there and you can see things. So this is maybe the most legible part of the slide—this is a heavyweight slide, this is from an internal presentation. So the way that you distinguish which kind of T cell you're looking at, you often use either a name or a list of markers, or both. So by a name, we mean something like Th17 that distinguishes a subset of T cells versus just saying T cells. So we look to see when a name was used versus just T cells with adjectives. And then we also looked when these markers were listed. So these are surface markers that are expected to be present on a cell, and so it's often you'll distinguish the cells you're working with as CD4 positive, CD25 negative, Foxp3 positive T cells. So I happen to know that that's a kind of T reg, but they didn't say T reg there. So we realized when we were working on this and just using the NormCo strategy for embedding these surface forms of entities and trying to map them back to Cell Ontology surface forms, that the tokenizer wasn't doing a good job of actually generating the tokens we needed, because it was just taking these things like that string 'CD4 positive CD25 negative Foxp3 positive' and it wasn't tokenizing that correctly in a way that like a scientist would tokenize it. So we took... so these are long strings that actually exist in text: 'CD4 bright', 'CD4 are low', 'CD4 RO- 4-1BB-'. So the right way to tokenize this string is you've got like two tokens here: you've got CD4 is the protein name, and then it's expression level. So we went and built like a custom tokenizer for the surface forms of entities that understood this in the way that an immunologist would understand it, and we were able to generate what we believe to be much better embeddings. And this is kind of interesting data on how often a marker is used in a positive, neutral, or negative sense for a particular cell type. And so this is also a place where we generate hypotheses. So for example, why is CD27 used both equally, nearly equally, as often as a positive and a negative marker of these naive cells? It's a subset of T cells. That's a... this is a pretty odd thing to see turn up. So writing these... so basically along the way at named entity linking, we didn't fiddle with neural networks; we ended up fiddling with tokenization. So that was kind of a lesson from this.
All right, so future work: we need to get more labeled data, it's gonna be annoying. We want to add new terms to the Cell Ontology. So our named entity linking, we're gonna have a set of entity mentions that map nicely to the existing Cell Ontology, but our knowledge of T cells is far greater than the knowledge of T cells in the Cell Ontology. So we want to figure out a way to expand their ontology. So I've been looking at like Jure, who I know from the databases world, also happens to work in the problem of named entity linking. So that was kind of fun to see that name. So he had a paper recently called HiExPan, which is this like hierarchical set expansion strategy for like taking a seed ontology and then expanding it with entities that you see. So sometimes entities map back to existing terms, but sometimes they actually belong somewhere in that tree. So figuring out like where that entity mention lives in that tree and then adding it. And then the hard part for us, and we've tried to do this, is determining when you've got a new entity and when you've got a synonym for an entity. So that's been a kind of interesting problem for us. So we want to kind of do this in a more rigorous, formal way. Negation detection, we haven't really tackled yet. There's this NegBio paper from a young guy at the NIH that does a lot of this stuff, who... that looks pretty good. We want to incorporate document context, so where the assertion occurs within a scientific paper is gonna determine how much we believe that to be like knowledge versus speculation. So there's a project called Section Tagger that can just tell you like what section of a paper you live in. Adding coreference resolution seems to have really helped for this SciERC project, which is trying to do entity extraction, relation extraction, and coreference resolution in a unified model. And it's some folks that you saw. And then we want to start making use of the dependency parse in labeling functions. So some other people who worked on this problem have shown that the dependency parse tree is actually pretty useful. You know, we want to work with Alex to get better kind of support for like the intermediate task of writing labeling functions. I'm gonna try out some ideas in data augmentation. The Snorkel team, when we were out there, was talking a lot about this. It's obviously been very interesting for computer vision. So we've been looking at different ways of doing data augmentation for... so there's like back-translation if you're working on like machine translation, but like what does data augmentation look like in the context of this problem? And then most importantly, we want to start making predictions and validating them in experiments. So one thing we want to play around with is taking the relations that we've extracted, treating that as like a knowledge graph, and then doing like link prediction on that knowledge graph using like knowledge graph embedding and things like that. So yeah, so that's the project. Landed at 30 seconds spare. Happy to take questions.