Back
Taha M.s.
Chief Science & Technology Officer, GE HealthCare Technologies Inc

#AIMI23 | Keynote Talk and Fireside Chat - Taha Kass-Haut

🎥 Jun 20, 2023 📺 StanfordAIMI ⏱ 44m
The 2023 AIMI Symposium was a hybrid conference presented by the Stanford Center for Artificial Intelligence in Medicine and Imaging (AIMI Center) on June 8, 2023. Website: https://aimi.stanford.edu Twitter:   / stanfordaimi   #StanfordAIMI #AIMI23 // Speaker: Taha Kass-Haut, Chief Technology Officer, GE HealthCare
Watch on YouTube

About Taha M.s.

Taha Kass-Hout, Chief Science & Technology Officer at GE HealthCare, has discussed the company's focus on precision diagnostics and artificial intelligence. In a January 2023 interview, he described GE HealthCare's spin-off as a "catalyst moment" and outlined the company's D3 Precision Care strategy, which emphasizes smart devices, digital products, and personalization of diagnostics and therapeutics. He stated that the company aims to build an open framework for innovators and focus on disease states to help healthcare providers manage populations and costs. He also noted that patient privacy and cybersecurity are top priorities. In a June 2023 keynote at the Stanford AIMI Symposium, Kass-Hout discussed the potential of generative AI and foundation models in healthcare. He said that large models can help structure medical data, automate workflows, and improve diagnostic accuracy, citing examples such as AI that reduces organ contouring time from three hours to under 20 minutes. He described AI as an assistive tool to automate manual tasks rather than replace clinicians, and emphasized the need for interpretable and explainable AI, rigorous validation, and partnerships with cloud providers to scale healthcare AI.

Source: AI-verified profile updated from Taha M.s.'s recent appearances. Browse all interviews →

Transcript (26 segments)
I
Interviewer0:10
Look forward to just getting started with the meat of our program. So I'd like to introduce our keynote speaker, who is Dr. Taha Kass-Hout. He is the Chief Technology Officer at GE Healthcare. GE has been a strong partner of the AIMI Center since its inception, with collaborative research projects that involve more than 40 people here at Stanford and at GE. Dr. Kass-Hout holds an MD and an MS in Biostatistics from the University of Texas. He did his clinical cardiology training at Harvard's Beth Israel Deaconess Medical Center. Before he was at GE for five years, he served as Vice President, Distinguished Engineer, and Chief Medical Officer at Amazon Web Services. He led Amazon Health AI strategy and efforts, including Amazon Precision Medicine initiative and the widely accessible OpenFDA that enables researchers and the public to search and analyze adverse event data. So an incredibly interesting career journey and background, and we are so delighted to have him here today to speak to us as our keynote speaker. Thank you, Dr. Kass-Hout.
T
Taha Kass-Hout1:34
Thank you. Thank you so much for having me today. Really awesome to be with all of you here, and also building on the great collaboration that we've had with Stanford over the years. Let me just start really quick, just oriented about today's topics that I'd like to cover. Accelerate the work that we're doing towards workflow optimization and increase that by 10x folds by reducing administrative burden, and also improve diagnostic accuracy and align also with some of the work that we're doing with regulatory agencies. Advanced machine learning, deep learning, has been at the core of product development at GE Healthcare now for more than a decade. We leverage deep learning for a variety of tasks that are related to object detection, segmentation, image quality improvement, as well as reconstruction, just to name a few. One of the key innovations from last year included a program called Air X for the brain. It can eliminate rescan wastage as well as scan inefficiencies by ensuring that both small and large organs within the brain are aligned consistently. Sono CNS is another one that helps in similar tasks, for example during post-3D acquisition by automatic detection of four standard views in our ultrasound release on Vscan ultrasound images. In addition to support of the auto segmentation and consistency, this AI feature also can take on average about 20 or so minutes just to do that, so the Sono CNS workflow really helps optimize. Another key collaboration we've done also with Stanford is just this week with the 510(k) cleared Sonic DL, which is an MRI feature to help with the scanning of the heart by taking accurate measurements as well as structure of the heart in one single breath hold, something that has been fairly troublesome for patients being in an MRI machine, trying to hold their breaths several times, sometimes several minutes, just to get the accurate measurement.
Also, I mentioned ultrasound. We have deep learning that supports medical practitioners who are less experienced in sonography, trying to extend the care beyond, especially with the shortage of ultrasound sonographers, and ultimately around how you can label organs automatically, such as the upper right quadrant, the liver, gallbladder, and right kidney for abdominal scans to improve quality reports. At GE Healthcare for the last 20 years now, we've been in the care pathway for radiation therapy. Here I show the steps involved in the RT workflow. We are automating the steps with deep learning, which can clearly benefit the radiation oncologists and also the dosimetrists in finding the right treatment for the right patient. To overcome these challenges, we recently received FDA clearance for anatomy delineation using deep learning out of segmentation. This is a post-processing step application designed to automatically generate contours of these organs at risk. It can reduce that time from hours to 15-20 minutes. It's trained with one deep learning localization algorithm and about 15 deep learning segmentation algorithms designed to contour about 40 different organs and sub-organs. The DL localization model analyzes all the input CT images, segments those organs, delineates the boundaries, and incrementally processes those in the sub-volumes of the CT. From an interoperability perspective, the DICOM images can be ingested directly into the radiotherapy planning software for further processing.
I started the discussion with Air X for the brain, which automatically detects anatomy and prescribes slices. At GE Healthcare, we started with the brain as the first anatomy, which took us about probably two years, almost half of that just curating the images and then spending the second half just training the model, deploying the model, and validating the model, and then through the regulatory clearance. For each of the anatomies, even with some optimization using what we've learned from the brain, it took us over a year just to develop and deploy each one of those models. This motivated us to think really differently. As you can see, translating that to the knee, the spine, and the prostate, each one is a separate journey on its own, and each one of these models requires its own training. This brings us to foundation models. There is a passage in Ernest Hemingway's novel The Sun Also Rises where a character named Mike was asked how he went bankrupt, and he answered, 'Gradually, then suddenly.' Technological advancements happen much the same way, where small changes accumulate and then suddenly the world is in a different place. This is what we've seen just in the last few months with large foundation models.
Generative AI models or large foundation models are how we can create and visualize new content, communicate, and even work efficiently. They will likely impact societies from business development to medicine, from education to research, and even art, from coding software. At a high level, these are models trained on a large amount of data in a self-supervised manner without any manual labels. They can generalize to tasks and data distributions beyond what they have already seen. They are capable of creating new content, combining inputs from various examples, just like how humans make reasoning from few examples. We have four million installs touching 1 billion patients, pretty much every modality you can imagine—imaging, ultrasound, even monitoring. So for example, we had to develop separate models for each one of the Air X anatomies, restarting the model development step from scratch every time. Despite recent progress in the field of medical AI, most existing models are fairly narrow, single-task systems that require large quantities of highly curated, high-quality data to train, and these models cannot easily be reused for new clinical contexts.
It would be advantageous to first leverage all the acquired healthcare data to learn the overall healthcare-specific domain, then generative data to learn the AI domain, and then fill out any gaps with knowledge-based synthetic data. Leveraging such large-scale generative models, called foundation models, we can really and very quickly build domain adaptation, usually with much less labeled data, as has been the case for us in building our existing models, which is always a challenge in healthcare to find that labeled data or just to curate that information. As we build this capability, we believe we can support a 10x growth in the amount of breadth and depth of the models that we can deploy.
Let's talk about the foundation model where it's been a recent development as it pertains to the medical field. Large language models, for the most part, are specific to text, with very limited applications right now in medical imaging specifically, so computer vision to a lesser extent. Still, it is very exciting to see some of the recent work, like the Segment Anything Model or SAM. The goal for the builders of these models was for image segmentation. They are fairly powerful AI models that can segment any object in an image or video. They trained on a dataset of 11 million images and almost 1.1 billion masks, which is the largest segmentation dataset to date. This dataset covers a wide array of objects such as animals, plants, vehicles, furniture. The ability to segment any object in the image is at a bar that's never been seen before, thanks to its generalizability and the diversity it was trained on. Recently we've seen MedSAM, which explored further fine-tuning on medical datasets. However, SAM has performed poorer in various other scenarios, such as segmentation of brain tumors or, in our case, the delineation between white matter and gray matter in the brain. MedSAM found that performance can be significantly improved after fine-tuning. So that's the latest iteration there. Then there is a new direction that came out of the Google Romulus paper. They use a combination of two steps: first, supervised representation learning on large-scale datasets of labeled natural images pulled from ImageNet or JFT, and then the second step involves intermediate self-supervised learning, which does not require any labels at all and instead trains a model to learn medical data representations on its own. This has demonstrated two abilities: improvement in zero-shot generalization to out-of-distribution settings with high accuracy, and a significant reduction in the need to annotate data. It has shown an improvement in in-distribution performance of up to 11.5% relative to strongly supervised baselines in diagnostic accuracy. So there is absolutely quite a bit of promise in this approach.
The third really exciting model we have seen recently is Med-PaLM and Med-PaLM 2 specifically, which is another medical foundation model built on top of 540 billion parameters that uses instruction tuning, medical data fine-tuning, as well as better prompting strategies. This model has achieved a new state-of-the-art for the MedMCQA dataset, and in axes relevant to clinical utility, it has shown high performance in factuality, medical reasoning capability, and low likelihood of harm, but has also shown a higher level of empathy compared to clinicians. This can be a significant step to help physicians also have access to the most relevant information as a search engine or even a chatbot in clinical practice. Let me share with you a few use cases or applications in healthcare that we are also working on. One example is when you look at radiology reports. This space has accelerated over the past decade with almost 5.5 billion studies generated every year around the globe, by a shrinking number of radiologists. It can be fairly difficult to structure this information, and it takes many of our customers weeks or even months just to structure and analyze that data to derive insights. Here's how Gen AI can really help structure this information and relieve the burden on overstretched healthcare workers. Gen AI can help radiologists create a combined report from all the prior studies, a summary in layman's terms for the patient, or a summary report for the surgeon or specialist who commissioned the imaging study, all from the same set of information, and even help the radiologists generate a rich report with all the measurements in hyperlink format with bi-directionality between text and imaging. This can happen in a fairly interactive, self-supervised way where radiologists can intervene where they need to or edit the output. In this case, we see for example a nodule in the left upper lobe periphery, size 5 millimeters in 2013, 6 millimeters in 2014. Instead of having separate models for each of these tasks as is the case today, here we have one large model serving multiple purposes for different kinds of personas, seamlessly integrating multimodal data, both structured and unstructured, interacting bi-directionally between images and text, and easily integrating with radiologist and clinician workflows.
Another example is large language models to track care pathways and recommend next steps in a treatment based on medical guidelines. Guidelines drive most of our diagnostic decisions, and it can be fairly complex to optimize patient care based on these guidelines. These can be fairly complex documents that define care pathways, including many alternatives that depend on patient characteristics, their history, and where they are in the care pathway. When followed consistently, these care pathways can potentially improve patient outcomes and quality of care substantially. These are areas in diagnostics and screening where large models can really help with these guidelines. Here's an example where Dr. Clifford Saper of Rheumatology in the U.S. uses ChatGPT to deal with insurance denials. Usually he sits down, writes a letter, puts references, and spends about two or three hours just going through that. Now he can have letters automatically drafted, even replies to patient messages. We've seen recently Microsoft and OpenAI partnering with Epic about how they can communicate directly with patients using Gen AI for general questions, instead of needing a data scientist to query specific data, then be able to draft the reply automatically to patients as they interact with their care. Another example is a virtual assistant to help patients manage their health. Patients can have a dialogue with their digital health history, tracking metrics like activity, energy, burnout, and sleep. You can imagine a future where you can actually track all your data in a privacy-first manner, understand unique needs, and offer more tailored advice based on your personal health journey. Ultimately, you can ask questions like, 'My leg is cramping, what would you recommend is causing it?' or 'I'm training for a half marathon, can you create a training plan for me?' or 'What was my last cholesterol level? Is it too high or too low? What can I do about it?' and that sort of interaction.
In the pursuit of ethical AI, especially in healthcare, it is really crucial to prioritize data privacy, minimize bias, ensure rigorous validation, and uphold ethical considerations and regulatory involvement. The FDA is looking to partner with med tech and big tech companies on regulating AI. If we look specifically at responsible AI, the great thing is that over the last few years we have made big strides in understanding the challenges and gaps in this space, and things have definitely improved a lot, but the problem is nowhere near solved. Interpretable and explainable AI becomes particularly important in breaking the black box, especially in the form of readily usable tools and diagnostics that we can use to improve patient care. A holistic approach is needed to bring all parties together and collaborate across all sectors and stakeholders to develop an environment that ensures the responsible use of these technologies. They are wonderful technologies if put to the right use, because they will drive efficiency, improve efficacy, protect privacy, promote equity, and close gaps in care, but also support data and interoperability of a future healthcare system. Together, we can truly build a more flexible health system and imagine a future where healthcare has no limits.
I
Interviewer28:08
Thank you so much. That was very interesting—a compelling vision for the future of AI. We've been getting a couple of questions from Slido, and I understand there is one in the room as well. I guess the first question that we heard about is: replace versus assist medical professionals? You showed us a lot of examples there. Can you just give us a perspective of how GE thinks about that question of replacing versus assisting caregivers?
T
Taha Kass-Hout28:46
Yeah, I know this is a really great question, and we get that a lot. We look at AI truly as an assistive tool. It's not to replace the physician; it's to help them get hours of their time back, like the linear organs example. If you have an assistive tool that can help you do that in less than 20 minutes, accurately, consistently across patients and the same patients, reproducible across different modalities and planes, why not leverage that? Almost like how you think about it, as a cardiologist, you think of a lot of assistive tools that can help you. Imagine what the stethoscope has done to the field of medicine over the last 250 years. We think about it the same way. This is an intelligent tool in a toolbox. The more you use it, the more you give it feedback, the more it adapts to your needs. But it's also very important that you do that in a way that you know how the output and inferences were made in the first place.
I
Interviewer30:11
Great, thank you. Was there a question in the room at some point here? I think there was—no? You're good. Wonderful. You know, I also want to make sure: you have such an interesting background, having gone from academics to government and then to industry. Let's start with the academics to industry transition. Can you tell us a little bit about how you made that decision? We have this vast AI ecosystem across all those institutions, but how did you think about that choice?
T
Taha Kass-Hout30:50
Yeah, I mean, before I joined GE Healthcare, I worked for about five and a half years at Amazon Web Services. I do have a Master of Science in Biostatistics and Machine Learning, and it's really near and dear to my heart how we can leverage data and population health data to really drive a lot of the decisions that we make. But it's fairly important, honestly, those lines are getting more and more blurred right now. Innovation can come from anywhere, and if we really want to take it to scale, oftentimes we look at massive inflection points that can happen anywhere in industry to be able to translate that to our domain. In this case, almost all large models were born in the cloud. Machine learning at scale is born in the cloud, and you need that high computation power, but also the tools to handle billions of products and try to personalize one item for you, and search in real, messy data. If you look at our health data, 97% of our data in healthcare is unstructured—medical images, reports, genetic mutations mentioned somewhere, EKG traces, etc. So there is a lot to be borrowed from those industries to bring that to our field. Back in 2009 during the H1N1 pandemic, the question was how to scale public health efforts. I needed a technology platform to help me do that, and in 2009, not a single health department could scale at that massive scale, not even the largest data center in the U.S. at CDC. I partnered with AWS back in 2009 to work directly with hospitals and labs to cover about 320 million population in the US. So, you can imagine that scale, which they built for consumer products and startups, was a great partnership to leverage all these tools to serve that purpose.
I
Interviewer33:57
Thank you. Then from government to industry, you moved to industry. How do you think about that problem of access to health data? Because it's really being produced primarily by healthcare delivery organizations and academic medical centers, so how do you think about the need for data when you're developing these innovations like you showed today?
T
Taha Kass-Hout34:31
Absolutely. This is really where you lean on a lot of partnerships. Even at GE Healthcare, we work very collaboratively with institutions like Stanford about how we can bring the technology innovation in MRI and CT and bring our breadth and depth about how these can be brought to scale with advancements in Gen AI or deep learning. One of the things is that unlike images of dogs and cats on the internet where you can build amazing models, or access to the largest Wikipedia in the world and build these really amazing large language models, healthcare data is fairly private, and as I mentioned earlier, a large volume of it is unstructured and not curated. So finding high-quality data is a challenge. I believe that the innovations we have seen with Gen AI are going to really help that quite a bit, because it can take us much of the way, reducing the time we spend today on curating and labeling a lot of this information.
I
Interviewer36:10
You talked about guidelines and that sort of thing, and I think one of the challenges with large language models is that they are trained on data that is maybe up to a certain date, and they give citations, but they hallucinate. There's no way to check how up-to-date it is. How do you keep up to date, for example, with guidelines that change? Several years ago the screening age for colon cancer dropped to 45, and recently we lowered mammograms from 50 to 40. Are you incorporating on a daily, weekly, or monthly basis new information so that you can give up-to-date answers?
T
Taha Kass-Hout36:47
From our world in medical imaging, how we are deploying these algorithms, for example Air Brain, Air X, Sonic DL, they are diagnostic and require regulatory clearance. So we are not updating them on a daily basis. However, the example I mentioned about guidelines is where the power of these technologies can help you summarize information in a most concise way. But if you remove the human from the loop, or the expert from the loop, or the content is not updated, then the data starts decaying and these models won't perform as well in the wild. So benchmarking and having feedback in the loop are super, super important for deploying these models in the wild. You want to be able to say, 'This is based on data that's two years old,' absolutely, because that's not being updated.
I
Interviewer38:10
And that lack of referability—where did you get this answer from? It's back to your early comment about breaking the black box.
T
Taha Kass-Hout38:18
Very, very important. By understanding what went into the model and what should have gone into the model in the first place—bias is one aspect, but there are many other aspects. By breaking the black box, you can start understanding that. I have spent much of my career working really hard on this particular problem. Back in the '90s, you built a neural net with six layers, and now that's a joke. Now we're talking about hundreds of thousands of layers in between, so how do you really know what went into the model and what came out? It's a very interesting area.
I
Interviewer39:11
And that is what are the benchmarks, and I imagine you're thinking about this because I know in the market, customers of AI algorithms are asking the question, 'How do I know this is going to work for my patients?' How do you help answer that question?
T
Taha Kass-Hout39:28
Yeah. Again, at GE Healthcare, we work very, very closely with physicians, clinicians, radiologists, and technicians who are constantly giving us feedback on how the algorithm is working and how we can refine and fine-tune it. As you deploy the models beyond their FDA clearance or CE mark, it's now in the hands of tens of millions of patients. So we work closely with regulatory agencies like the FDA, especially in the world of Gen AI, to determine the best ways to build, deploy, and benchmark these models. When I was at the FDA with PrecisionFDA and Rob Califf, we established PrecisionFDA because nothing existed in the world to benchmark next-generation sequencing technologies and ensure your algorithms are calling a variant accurately and reproducibly. If you get an F1 score of 0.7, how do you know if that's good or bad? You don't want to kill innovation because in some areas 0.7 is amazing, but in others, if everyone else is getting 0.99, then you should continue working. We need to continue working together as a community and educate the regulatory agencies about how best to benchmark this field.
I
Interviewer41:18
Thank you. Time for one more?
A
Audience Member41:24
Yes, first of all, I'm a clinical radiologist working in the largest healthcare system in the United States, the VA Healthcare System. I wanted to share with everyone the challenges we have as clinical clinicians getting AI adopted. Some of us feel that the AI tools we've implemented in our practice make us better radiologists, but there are others who do not even want to try to implement it. So what are your thoughts about the challenges in the real world of getting clinicians comfortable with using these tools?
T
Taha Kass-Hout42:02
Yeah, I think it's very, very real, especially if you're talking about pixel-based AI or clinical decision support. You can't change culture overnight. But things that we are seeing that are really driving adoption are in areas like advanced visualizations and 3D post-processing, which take a lot of time. Automating those is of great value to clinicians, especially radiologists. Automated measurements, reconstructing 3D from 2D images, and doing accurate measurements with an adjudicated process where you can give feedback—it's like a Tinder-like experience where you actually feel you're part of designing the system. Then you lean on clinicians to design these systems in the first place, rather than just pushing the technology out. This is why our advanced visualization has propagated around the world for the last three decades at GE Healthcare—because we were working backward from those clinicians. For our machine learning in the last decade, especially since COVID, there has been almost a hockey stick of adoption because it solved a pain point and was a heavy lift, rather than just a nice add-on. It comes down to solving real problems.
I
Interviewer44:00
Fascinating. Thank you for the conversation. Thank you. We have a 10-minute break, and we'll be back right here with the next session in 10 minutes. Thank you all.