Taha Kass-Hout1:34
Thank you. Thank you so much for having me today. Really awesome to be with all of you here, and also building on the great collaboration that we've had with Stanford over the years. Let me just start really quick, just oriented about today's topics that I'd like to cover. Accelerate the work that we're doing towards workflow optimization and increase that by 10x folds by reducing administrative burden, and also improve diagnostic accuracy and align also with some of the work that we're doing with regulatory agencies. Advanced machine learning, deep learning, has been at the core of product development at GE Healthcare now for more than a decade. We leverage deep learning for a variety of tasks that are related to object detection, segmentation, image quality improvement, as well as reconstruction, just to name a few. One of the key innovations from last year included a program called Air X for the brain. It can eliminate rescan wastage as well as scan inefficiencies by ensuring that both small and large organs within the brain are aligned consistently. Sono CNS is another one that helps in similar tasks, for example during post-3D acquisition by automatic detection of four standard views in our ultrasound release on Vscan ultrasound images. In addition to support of the auto segmentation and consistency, this AI feature also can take on average about 20 or so minutes just to do that, so the Sono CNS workflow really helps optimize. Another key collaboration we've done also with Stanford is just this week with the 510(k) cleared Sonic DL, which is an MRI feature to help with the scanning of the heart by taking accurate measurements as well as structure of the heart in one single breath hold, something that has been fairly troublesome for patients being in an MRI machine, trying to hold their breaths several times, sometimes several minutes, just to get the accurate measurement.
Also, I mentioned ultrasound. We have deep learning that supports medical practitioners who are less experienced in sonography, trying to extend the care beyond, especially with the shortage of ultrasound sonographers, and ultimately around how you can label organs automatically, such as the upper right quadrant, the liver, gallbladder, and right kidney for abdominal scans to improve quality reports. At GE Healthcare for the last 20 years now, we've been in the care pathway for radiation therapy. Here I show the steps involved in the RT workflow. We are automating the steps with deep learning, which can clearly benefit the radiation oncologists and also the dosimetrists in finding the right treatment for the right patient. To overcome these challenges, we recently received FDA clearance for anatomy delineation using deep learning out of segmentation. This is a post-processing step application designed to automatically generate contours of these organs at risk. It can reduce that time from hours to 15-20 minutes. It's trained with one deep learning localization algorithm and about 15 deep learning segmentation algorithms designed to contour about 40 different organs and sub-organs. The DL localization model analyzes all the input CT images, segments those organs, delineates the boundaries, and incrementally processes those in the sub-volumes of the CT. From an interoperability perspective, the DICOM images can be ingested directly into the radiotherapy planning software for further processing.
I started the discussion with Air X for the brain, which automatically detects anatomy and prescribes slices. At GE Healthcare, we started with the brain as the first anatomy, which took us about probably two years, almost half of that just curating the images and then spending the second half just training the model, deploying the model, and validating the model, and then through the regulatory clearance. For each of the anatomies, even with some optimization using what we've learned from the brain, it took us over a year just to develop and deploy each one of those models. This motivated us to think really differently. As you can see, translating that to the knee, the spine, and the prostate, each one is a separate journey on its own, and each one of these models requires its own training. This brings us to foundation models. There is a passage in Ernest Hemingway's novel The Sun Also Rises where a character named Mike was asked how he went bankrupt, and he answered, 'Gradually, then suddenly.' Technological advancements happen much the same way, where small changes accumulate and then suddenly the world is in a different place. This is what we've seen just in the last few months with large foundation models.
Generative AI models or large foundation models are how we can create and visualize new content, communicate, and even work efficiently. They will likely impact societies from business development to medicine, from education to research, and even art, from coding software. At a high level, these are models trained on a large amount of data in a self-supervised manner without any manual labels. They can generalize to tasks and data distributions beyond what they have already seen. They are capable of creating new content, combining inputs from various examples, just like how humans make reasoning from few examples. We have four million installs touching 1 billion patients, pretty much every modality you can imagine—imaging, ultrasound, even monitoring. So for example, we had to develop separate models for each one of the Air X anatomies, restarting the model development step from scratch every time. Despite recent progress in the field of medical AI, most existing models are fairly narrow, single-task systems that require large quantities of highly curated, high-quality data to train, and these models cannot easily be reused for new clinical contexts.
It would be advantageous to first leverage all the acquired healthcare data to learn the overall healthcare-specific domain, then generative data to learn the AI domain, and then fill out any gaps with knowledge-based synthetic data. Leveraging such large-scale generative models, called foundation models, we can really and very quickly build domain adaptation, usually with much less labeled data, as has been the case for us in building our existing models, which is always a challenge in healthcare to find that labeled data or just to curate that information. As we build this capability, we believe we can support a 10x growth in the amount of breadth and depth of the models that we can deploy.
Let's talk about the foundation model where it's been a recent development as it pertains to the medical field. Large language models, for the most part, are specific to text, with very limited applications right now in medical imaging specifically, so computer vision to a lesser extent. Still, it is very exciting to see some of the recent work, like the Segment Anything Model or SAM. The goal for the builders of these models was for image segmentation. They are fairly powerful AI models that can segment any object in an image or video. They trained on a dataset of 11 million images and almost 1.1 billion masks, which is the largest segmentation dataset to date. This dataset covers a wide array of objects such as animals, plants, vehicles, furniture. The ability to segment any object in the image is at a bar that's never been seen before, thanks to its generalizability and the diversity it was trained on. Recently we've seen MedSAM, which explored further fine-tuning on medical datasets. However, SAM has performed poorer in various other scenarios, such as segmentation of brain tumors or, in our case, the delineation between white matter and gray matter in the brain. MedSAM found that performance can be significantly improved after fine-tuning. So that's the latest iteration there. Then there is a new direction that came out of the Google Romulus paper. They use a combination of two steps: first, supervised representation learning on large-scale datasets of labeled natural images pulled from ImageNet or JFT, and then the second step involves intermediate self-supervised learning, which does not require any labels at all and instead trains a model to learn medical data representations on its own. This has demonstrated two abilities: improvement in zero-shot generalization to out-of-distribution settings with high accuracy, and a significant reduction in the need to annotate data. It has shown an improvement in in-distribution performance of up to 11.5% relative to strongly supervised baselines in diagnostic accuracy. So there is absolutely quite a bit of promise in this approach.
The third really exciting model we have seen recently is Med-PaLM and Med-PaLM 2 specifically, which is another medical foundation model built on top of 540 billion parameters that uses instruction tuning, medical data fine-tuning, as well as better prompting strategies. This model has achieved a new state-of-the-art for the MedMCQA dataset, and in axes relevant to clinical utility, it has shown high performance in factuality, medical reasoning capability, and low likelihood of harm, but has also shown a higher level of empathy compared to clinicians. This can be a significant step to help physicians also have access to the most relevant information as a search engine or even a chatbot in clinical practice. Let me share with you a few use cases or applications in healthcare that we are also working on. One example is when you look at radiology reports. This space has accelerated over the past decade with almost 5.5 billion studies generated every year around the globe, by a shrinking number of radiologists. It can be fairly difficult to structure this information, and it takes many of our customers weeks or even months just to structure and analyze that data to derive insights. Here's how Gen AI can really help structure this information and relieve the burden on overstretched healthcare workers. Gen AI can help radiologists create a combined report from all the prior studies, a summary in layman's terms for the patient, or a summary report for the surgeon or specialist who commissioned the imaging study, all from the same set of information, and even help the radiologists generate a rich report with all the measurements in hyperlink format with bi-directionality between text and imaging. This can happen in a fairly interactive, self-supervised way where radiologists can intervene where they need to or edit the output. In this case, we see for example a nodule in the left upper lobe periphery, size 5 millimeters in 2013, 6 millimeters in 2014. Instead of having separate models for each of these tasks as is the case today, here we have one large model serving multiple purposes for different kinds of personas, seamlessly integrating multimodal data, both structured and unstructured, interacting bi-directionally between images and text, and easily integrating with radiologist and clinician workflows.
Another example is large language models to track care pathways and recommend next steps in a treatment based on medical guidelines. Guidelines drive most of our diagnostic decisions, and it can be fairly complex to optimize patient care based on these guidelines. These can be fairly complex documents that define care pathways, including many alternatives that depend on patient characteristics, their history, and where they are in the care pathway. When followed consistently, these care pathways can potentially improve patient outcomes and quality of care substantially. These are areas in diagnostics and screening where large models can really help with these guidelines. Here's an example where Dr. Clifford Saper of Rheumatology in the U.S. uses ChatGPT to deal with insurance denials. Usually he sits down, writes a letter, puts references, and spends about two or three hours just going through that. Now he can have letters automatically drafted, even replies to patient messages. We've seen recently Microsoft and OpenAI partnering with Epic about how they can communicate directly with patients using Gen AI for general questions, instead of needing a data scientist to query specific data, then be able to draft the reply automatically to patients as they interact with their care. Another example is a virtual assistant to help patients manage their health. Patients can have a dialogue with their digital health history, tracking metrics like activity, energy, burnout, and sleep. You can imagine a future where you can actually track all your data in a privacy-first manner, understand unique needs, and offer more tailored advice based on your personal health journey. Ultimately, you can ask questions like, 'My leg is cramping, what would you recommend is causing it?' or 'I'm training for a half marathon, can you create a training plan for me?' or 'What was my last cholesterol level? Is it too high or too low? What can I do about it?' and that sort of interaction.
In the pursuit of ethical AI, especially in healthcare, it is really crucial to prioritize data privacy, minimize bias, ensure rigorous validation, and uphold ethical considerations and regulatory involvement. The FDA is looking to partner with med tech and big tech companies on regulating AI. If we look specifically at responsible AI, the great thing is that over the last few years we have made big strides in understanding the challenges and gaps in this space, and things have definitely improved a lot, but the problem is nowhere near solved. Interpretable and explainable AI becomes particularly important in breaking the black box, especially in the form of readily usable tools and diagnostics that we can use to improve patient care. A holistic approach is needed to bring all parties together and collaborate across all sectors and stakeholders to develop an environment that ensures the responsible use of these technologies. They are wonderful technologies if put to the right use, because they will drive efficiency, improve efficacy, protect privacy, promote equity, and close gaps in care, but also support data and interoperability of a future healthcare system. Together, we can truly build a more flexible health system and imagine a future where healthcare has no limits.