Back
Amit Paka
Cofounder, Fiddler AI

Model Performance Monitoring and Why You Need it Yesterday // Amit Paka // MLOps Coffee Sessions #42

🎥 Jun 01, 2021 📺 MLOpscommunity ⏱ 66m
Coffee Sessions #42 with Amit Paka of Fiddler AI, Model Performance Monitoring. //Abstract Machine Learning accelerates business growth but is prone to performance degradation due to its high reliance on data. Moreover, MLOps is often fragmented in many organizations, causing frictions to debug models in production. With new rules from the EU that focus on trust and transparency, it’s becoming more important to keep track of model performance. But how? We propose a new framework, a centralized ML Model Performance Management powered by Explainable AI. Learn more about how you can stay complia...
Watch on YouTube

About Amit Paka

Amit Paka, cofounder of Fiddler AI, discussed model performance monitoring and the proposed EU artificial intelligence regulation during a June 2021 appearance on the MLOps Coffee Sessions podcast. Paka stated that before founding Fiddler, while leading the product team on shopping apps at Samsung, he observed challenges that data teams faced in operationalizing models, including difficulty running A/B tests. He said the lack of transparency in model decisioning "played a key role and it came up again and again," and that the Fiddler team aligned on a mission of "helping teams build trust with AI." Paka described the proposed EU regulation as aiming to help teams build "human-centric and trustful AI," classifying applications into categories including banned "unacceptable" uses and "high-risk" applications such as self-driving cars or credit scoring that would face new oversight. He argued that for high-risk applications, the law would require sufficient transparency for users to understand and control how models work, and that monitoring, record-keeping, and fairness validation are what the regulation is pushing toward. Paka also introduced the concept of "Model Performance Management" (MPM), which he described as a centralized framework powered by explainable AI that involves measuring, validating, monitoring, and analyzing model behavior across the lifecycle, with the ability to feed new data representations back into training.

Source: AI-verified profile updated from Amit Paka's recent appearances. Browse all interviews →

Transcript (33 segments)
H
Host0:10
because they have sponsored this session. They have decided they're actually one of the first companies to sponsor the community. And so for me, it is something that is truly humbling to see people that are so invested and so interested in the community that they would like to sponsor, they volunteer for sponsoring these initiatives that we're doing, and they want to make sure that the community continues to grow and it continues to be the place for having conversations around ML ops today. We're talking with Amit Paka from Fiddler AI, and we're going to be talking about all kinds of things. We've got a lot on the agenda. We're just machine learning and not just monitoring for a few metrics after the model is out, but a real in-depth deep dive on monitoring. If you've listened to any other podcast or meetup from us, you probably know that monitoring is one of my favorite subjects. And so today we get to dive into it real deep. And as usual these days, I'm also joined by Vishnu. So welcome to the both of you. It's great to be here. And again, Amit, just before we start, thank you to Fiddler for sponsoring this and sponsoring the community. It's really awesome to see and yeah, it's heartwarming. You guys are awesome.
A
Amit Paka1:52
Thanks so much to Major. I think it's been amazing to watch the like engaged communities on ML ops, and it's so insightful to be part of these conversations, right, to drive all these learnings. Given it's still a wild west with ML ops, it's like kind of communities like these are where standards get defined, workflows get defined, people help each other to make their lives easier in this wild wild west. So we are really happy that you and the team have actually picked up this kind of tough job on maintaining this kind of community. It's a lot of work, and we're happy to actually play a small part of it.
H
Host2:51
That's awesome. Yeah, for us it's a labor of love. ML ops space and even so much more traction. It was where you would have to, I think a year ago you would have to really search and know who to follow on LinkedIn or on Substack to get any information about ML ops. But now that's completely different. I think every single person in my news feed on LinkedIn is talking about how what is ML ops or ML ops this or ML ops that. So it's definitely gaining some traction. But that might also be because I'm in an echo chamber on my LinkedIn and I just follow people who talk about ML ops and all that good stuff. All right, so let's actually definitely yeah, the market is actually seeing the demand and they're following up with a lot of supply. So very exciting space right now.
V
Vishnu4:22
Yep, yeah it is very true. I mean there's a lot of crazy stuff happening with ML ops.
H
Host4:33
So let's jump into what we came to talk about today. I would love to hear a little bit of background about you before we get into the regulation and the monitoring for ML. And maybe it would be a great start just to know how you came to be where you're at. You're the Chief Product Officer at Fiddler, if I'm not mistaken?
A
Amit Paka5:11
And so we were using machine learning models for product recommendations. And so if you went to a Samsung mobile app and you looked at the list of products, those were recommended by a machine learning model. And I was observing a lot of challenges that the data team was facing just operationalizing these models, right? So just starting off with managing versions of these models, things that we take for granted on Git, was very hard for them. Doing an A/B test on these models was incredibly hard. The way the team would do an A/B test is they would run a previous model for a week and then they would stop running the model, they would run a new model for the next week. Based on the output of these models, one of the things I was asked to do is I was testing the hypothesis that Samsung users would be open to buying a Samsung product via a chatbot. So if you speak to a Samsung chatbot using Facebook Messenger, you'd go through this conversation script and then you'd buy it. I was testing that hypothesis. To actually test the hypothesis, we ran an ad campaign, a million dollar ad campaign. The target of this ad campaign, 300,000 users, was called from 21 million users in the database that came from a machine learning model. And of course, as a product manager, you want to understand why those 300,000 from the 21 million. Was that the wrong audience that we went after, or was it something else? So the lack of transparency in the decisioning of these models played a key role, and it came up again and again. And so I connected with Krishna, who's the founder CEO, who I had gone to grad school with and overlapped at Microsoft. And he was facing a similar issue at Facebook. He was working on the news feed and he joined after the US elections, and there was a lot of angst about potential bias in the news feed. And it's very hard to dispute that given it's a complex ensemble. And so he worked on trying to help internal teams learn why you're seeing a news story using explainability, and that was also a challenge. And it really keeps us going. So that led to the genesis of the company. And it's been great and almost validating to see how our worldview has, close to two and a half years back when we started, is becoming more mainstream. The trust in AI is becoming not just a nice to have but a must-have, especially in light of this new EU regulation where Margrethe Vestager comes out and says that trust is a must-have. That really is one of the things.
H
Host9:12
Oh yeah, go ahead.
A
Amit Paka9:14
Yeah, I was just going to add that at Fiddler, I look at ensuring that our products stay true to the mission of not just giving you visibility into model behavior at train time, but also post-deployment, and kind of weaving that into an experience that's very unique to ML ops, given machine learning models are unique probabilistic entities that need constant new iterations given their actual performance can decay constantly.
H
Host10:12
Regulation, I should say. There's a lot of people that are very quick to correct me that it has not been enacted yet and it is at the very infancy still. So there's lots that can change, and we're looking at something that is probably realistically three years out, potentially longer, potentially a little shorter. But that being said, you probably shouldn't wait three years before you start getting all of your ducks in a row. It's a lot that is being proposed, and what I've heard from different sources that are very close to this that are in Brussels is that it should be looked at. The proposed EU regulation... Oh, I'm so bad with the acronyms, especially when they make things like GDPR and all that. But there's like, I can't remember the acronyms, but I'll figure it out and put in the show notes later. The three different proposed regulations that they have: there's one on AI specifically, there's one on data, and then there's one on the sharing of the data. And all of that good stuff. I'll figure that out, I'll come back to it in a minute. But we should probably mention what we were talking about before we hit record, which was this idea of what is the new regulation, the high risk, the three different sectors that they've put together.
A
Amit Paka12:10
Yeah, yeah, yeah, great point. It's a pretty, I just wanted to start off by saying, Nikki, this is something we were discussing. It's a pretty detailed proposed regulation. It's around 80 pages, and I am surprised how much time the team spent on crafting this. It's pretty well written. The goal that the EU seems to be going for with this proposed regulation is they really want to go after helping teams build human-centric and trustful AI. That's typically how proposed regulations get voted on, and then it does take a couple of years for it to go into effect, and there might be some changes that will happen from now until the one that gets voted on and accepted, and there will be a couple of years before it really goes into effect, just like GDPR did. But to summarize what the proposed regulation is doing with trust at its core, it first classifies AIs into four categories. It says the first category is unacceptable risk applications that are banned. These are applications like behavior manipulation, for example. They're banned. The second one, and now we're getting into high risk applications, and we're going to talk about what this oversight is. The third one is limited risk applications, and these are the likes of chatbots, and there are some transparency requirements for these limited risk. For example, it says if you're talking to a chatbot, you should know that you're talking to AI, really simple guidelines. And the last one is minimal risk application, and this they believe will be the majority of AI applications, and there's no proposed intervention here. An example of this is a spam filter. Limited minimal risk application that is high risk. And so high risk applications are really ones that have implications on public safety, services, there are actual opportunities. Examples of these are credit scoring, loans, recruitment, self-driving cars, robotic surgery. So these are high-risk applications. So the first thing is, do you fall in this high-risk application? And the next thing is, what does high-risk application mean? What are these new guidelines? And so what the proposed regulation stipulates is that you can go back to it. There are requirements on transparency, knowing how this model was built, how it operates, and included in this transparency is fairness and bias implications on certain protected attributes like race and so on. And human oversight. It uses human oversight a little bit differently from transparency. The ability for a human to come in and control this. Of course, the human needs transparency, but they need the ability to actually control it. And finally, it says that once you know how this model is working, it's transparent, high quality data you need to have. It was trained well, but it ran into live data that completely makes it a hazard. That should not happen. You should be able to monitor it. And finally, there's robustness aspects, accuracy aspects, and security aspects to get this operational visibility. So that's the high level of what the proposed regulation means, and we can deep dive into what it means in the context of transparency and what it means in the context of operational visibility.
H
Host17:45
Yeah, I think there's something to note that we were talking about before about the idea of how they specifically called out with AI and that's not a problem at all, and you're like, wait a minute, the hypocrisy here is ridiculous. So hopefully in some of these new iterations, it will come into effect. But just as a side note, I went and I Googled what the damn names of these other acts are that I couldn't remember. And there's the Digital Services Act and there's the Digital Marketing Act, the DSA and the DMA. And they should be looked at as like the trifecta of all of these together really are trying to help Europe as a society go forward in a way that they feel is human-centric. About as far as the military goes and why they chose not to single out the military.
A
Amit Paka19:11
My sense is there's a little bit of an unknown there. They probably have not really run into a lot of military applications, so they don't know what they don't know, and they probably want to keep that door open. But it feels like there would be potential pushback on that, and there might be language in other iterations of this regulation that might... but it feels like they wanted to have that option, not knowing how they're going to deploy, to just make sure that it doesn't hold them back in situations of emergencies.
H
Host20:26
Yeah, that makes sense too. And one person I interviewed before said it's very hard to get regulation passed on current dangerous weapons, you know, something like chemical weapons or landmines, and getting those banned or getting treaties signed around those which we know what the effects are. It's even harder to get things passed when...
A
Amit Paka21:12
What's been really interesting is in the US, the DoD has actually come out with more focus on trust in AI. So there has been a team formed called JAIC, and that team is now looking at how do you bring in transparency and trustworthiness to AI. So the US has kind of absorbed the implications of black box AI, especially in the context of DoD. So I think the DoD had a principles document that they kind of came out with towards the start of this year. So it's kind of interesting to see that in the US, DoD is actually focusing on increasing trust within AI.
H
Host22:20
It's good to hear you mention defense because I think that's been one of the more interesting areas, particularly what the Department of Defense has been doing with AI position papers. It's as you said, the joint center that they put together. I think both of the last few defense secretaries have had a pretty aggressive focus on how AI has been important on the battlefield and its implications for the country and strategic competitions with China. So I think that's certainly an area to watch, especially because we know so much of the technology pipeline in the US particularly tends to come from defense. Regulation in AI coming out of Europe, but also just more broadly, there are a lot of different areas and sectors obviously where AI can play vastly different impacts. In advertising and commerce, it's playing perhaps a more direct and hands-on role in terms of how it's impacting customers and real people. But when we look at areas like healthcare or even energy or other more advisory scenarios, it's working more with trained professionals, people who may have that ability to control AI as you mentioned the European Union wants to see as part of its regulation. What's your perspective, you know, as an ML ops company and explainability company, on the diversity of use cases?
A
Amit Paka24:11
The diversity of explainability is broad, and we get to hear a whole lot of that from prospective customers across a variety of verticals. The urgent need for explainability comes from verticals that are facing the problem today on the ground and cannot run their business. And so these typically tend to be businesses that want to or are using AI, but they need to make sure that they're meeting some guidelines, some regulations. For example, in banking in the US, there's a stipulation of SR 11-7 that the Fed laid out at the end of the financial crisis, which mandates that you need to have explainability for models after deployment so that you don't take down the financial system again. It was meant for quantitative models, but it is also applicable to AI models. So the need for explainability on the ground, the ones who have the budget and are buying, are where they want to use AI and are hitting into regulations that mandate transparency today. That's one part. And now there are also companies who want to be good corporate citizens and are also seeing the potential pitfalls of not being transparent. For example, there have been recent issues where facial recognition software was not recognizing people well if they have darker skin tones, and because of that problem, a lot of companies exited the facial recognition market due to the groundswell of protest by people. So the second need is coming from these group of companies who want to be good corporate citizens or want to ensure that they don't have any PR issues with the AI that they're deploying. And so the diversity there is pretty broad depending on how widely used their AI is and what kind of potential brand impacts they might face.
H
Host27:10
Models and not just for transparency, trying to deal with existing regulations where for your current products you need to make sure that there is acceptance amongst regulators, acceptance amongst customers. But there's also a sort of bent towards customers that want to use it to develop new products and how it informs new product development. So regular product development and new product development. That makes sense, and I think it really drives with how I have seen this play out in healthcare. And as our listeners know, I work at a medical device company, and in the healthcare industry we've had regulations around software as a medical device, around the FDA's good machine learning practices which have recently been... the FDA has had to bridge the gap between how machine learning impacts today and how machine learning impacts tomorrow, and what does product development tomorrow versus today look like. And as a product manager, I think it probably must be very interesting to you how companies are responding to this in terms of just their management structure and the kind of talent that they're retaining.
A
Amit Paka28:34
Yeah, it's very interesting now as teams go beyond just looking at explainability, beyond the actual regulated use cases like lending, healthcare, recruitment. There are other use cases which are high value, like in oil and natural gas, where a single decision can be worth 10 to 20 million dollars. So explainability there is more like a need to ensure that you're making decisions with the right ROI. But what we're also seeing is that teams are now coupling explainability with monitoring. So it's not enough to know how the models are behaving, but also ensuring that you know how they're behaving when they're running, operational aspects of it, not just pre-deployment but also post-deployment. Because given the nature of machine learning models which are really probabilistic, they can decay over time. The data that they see might be different than the data they were trained on.
H
Host30:10
That leads into the next question that I wanted to ask about MPM or what you've dubbed as Model Performance Monitoring, a bit of a play off of Application Performance Monitoring, I think. And the first thing that jumps out at me when I read the blog post that Krishna wrote, which we can link below for anybody that's interested, is how is this different or how is this the same as what some have been going around talking about when it comes to the evaluation store? And I think it's like Josh Tobin, who's the Full Stack Deep Learning guy, he has been talking about this need for monitoring. And when I read this blog post, that felt like that's what you were calling for too. So how do these two differ and is it the same?
A
Amit Paka31:24
Yeah, it's a good question. I think different companies in this domain are talking to their customers and they're kind of converging on the same problem that they're seeing. So I think different companies are actually coming up with similar solutions, so they are mostly similar. The premise of MPM, really for folks who are hearing about this for the first time, is that within operationalization, you have to focus on understanding and improving model performance, and you do that through visibility across all the parts of the model life cycle. So that is MPM: how do you improve understanding and model performance through visibility across the life cycle? And what does it mean in different parts of the life cycle? At train time, when you have training sets that you're using for your model, you can assess them for bias, you can assess them for any data quality issues. You can also get feedback from your current model champion running on certain slices of the data. This could mean slices of low performance, or if you're a lending model, slices where you're making loans to men versus women or loans in one state to another. It could be slices that are off, like business interest, or it could be slices that have performance implications. At deployment time, this could mean that you're recording traffic that your model is about to serve and you're about to do a champion-challenger or maybe one champion and multiple challenger tests. At monitoring, of course, and this is where a lot of focus seems to be, but it's one part of MPM. You're assessing this traffic that you're recording for drift within life and trying to analyze that maybe in the context of the slices that you just validated the model against at train time. You can use the same slices and then that gives you a new world view that feeds back into the training site. So this iterative process where constantly the model is converging and diverging on its actual performance because data is changing. The management of this across multiple versions of your models is MPM. And you could also view this in the context of if you want to store everything in one location, it could be an evaluation store. We believe it's like a meta layer that allows you to store all the information about your model or slices of the model so you can be better equipped to ensure that your model is operating at peak performance.
H
Host35:26
So this notion of Model Performance Management and the associated idea of control theory sounds a lot like it's very inspired by other engineering disciplines and how they may approach certain dynamic or probabilistic systems. Could you talk a little bit about how other engineering fields have influenced your approach to thinking about machine learning, or whether that's an accurate assumption on my end, and maybe even what control theory is and how that relates to model performance management?
A
Amit Paka35:54
Control theory really means that you not just measure it but you can influence it, so you can actually keep it at a desired state. You've got sensors that not just tell you what it's doing but you can control the speed based on the input. So if your desired speed is 25 miles per hour, the sensor tells you 27, now you can control the sensor and say come down to 25. In the case of models, that means that, for example, you're not really going for highest accuracy in your model, but you're trying to ensure that the model has a certain acceptable level of bias across protected attributes. And so when you first have visibility that that is being breached, you feed whatever that new data representation is back into your training set so you can create a new representation of that model that meets that new threshold, the new reality that you're seeing within production. So it is an extension of control theory given that you first have to have visibility into how a thing behaves. A whole lot of teams don't even know how their models are behaving when they deploy, so that's number one, not having any operational insights. And once you have operational insights, how do you ensure that it's maintaining your desired metrics? And that is where the control theory comes in.
H
Host38:10
Yeah, that's a great way to think about it. And I think it's something that a lot of engineers, other machine learning engineers in our community talk about, which is how do you actually turn an abstraction into something that's usable? Or the idea for an abstraction into something that's usable. For example, the idea of a pipeline, that's a commonly used design pattern or abstraction that we talk a lot about in the community. Google Cloud has written about it and Pachyderm has some really nice features that talk about how to do a build pipeline, and that's how that idea of what a model development pipeline should look like is translated into a particular software or cloud provider's tooling. In your case, could you talk a little bit more about how this abstraction around control theory and Model Performance Management how that translates into the actual tooling that you're building at Fiddler?
A
Amit Paka39:12
Today, teams typically have some data, they train a model, either it could be handcrafted or it could be using AutoML, and it's likely tuned for highest accuracy. And then you just take this model. If it's through AutoML, you don't really know what kind of relationships it dropped. You then go and deploy it, and then you don't monitor it because you assume that the reality that your model faces meets the hypothetical training set. So this is what teams do today. In the MPM world, what you would do is when you get that data, you would assess it for data quality. You would ask: should you have more samples of one versus the other based on the fairness metric that you're trying to assess? How many instances? What is the distribution of the data? Does it have a whole lot of missing values where you want it to grab relationships without the missing values? And so on. Then you go on to the post-model training phase where when your model is trained, how do you learn what relationships the model is grabbing today? That's typically not what a data scientist does. You might do some ad hoc testing on a test data set, but you don't really look at what are the feature importances, what are the important features of this model, what are the lowest performing regions of this model, how is this model behaving in the context of certain protected classes or certain slices that might be relevant to data, for example loans made to certain states. That is not typically done today, and that's one of the aspects of MPM. So teams would be logging this model in the context of the data and running the analysis to ensure they know how the model is behaving. At the monitoring part, you would then be looking at not just in the context of the data that the model is facing in live, you would look at it in the context of slices that you just validated the model against. If for example you have set a threshold for bias, or if loans made in California are really important, you would look at it from the lens of that. So you would need tooling not just to help you monitor broadly but also at a very specific level if your bias thresholds, if you have certain slices that you want to monitor, certain metrics. It might be that you could have trained the model for accuracy, you want to ensure that you're meeting the accuracy. At times you might not even have ground truth. If you have data integrity issues happening, then you use that representation of data that your model is facing in live and feed that back into your training set so you can train a model that accurately represents the reality versus the hypothetical. None of these things are done by a typical data scientist today, and that's what they would need to do within an MPM workflow.
H
Host43:40
Got it, got it. All an incredible set of very helpful activities, things that we would like to do that it's hard to do correctly, and I think that's kind of your point right now. It's sufficiently hard today, and as a result people don't have the investment and the time that they need on these things to do those things.
A
Amit Paka44:11
It's not just an ML engineer or a software engineer focused on machine learning, but it's more actually the model development experts themselves. So it's a mix of data scientists and machine learning engineers. In fact, what we're also seeing in some teams is the traditional DevOps folks who monitored web apps are being repurposed to monitor machine learning apps. That's also happening. But our typical users are data scientists who want to know more about this model that they've trained, ensuring that it's not biased, ensuring that they know if there are regions of lower performance. And machine learning engineers as well. This is where I think the job functions are still being formalized across data science and this world of ML ops. Especially within ML ops, we've not seen a whole lot of purpose-built ML ops teams that handle these deployments. They're being repurposed from DevOps. We have seen ML engineers also kind of do the deployment but maybe not stay on top of the operational model.
H
Host46:10
That reminds me of what you're saying, and even this question that you had, Vishnu. Last week when I was talking to Chris Berg, he was saying one of the quotes that I took away from that conversation was when he said, 'You, as someone who is working on a data product, you own the result, you don't own your piece.' And I found that to be fascinating because you have to make sure that whoever you're working with, whatever they're doing, is also making the result be better. And it's not that 'Okay, I built my model and I'm good and now it's your problem' or 'I did this and now whatever, you figure it out.' There's no more throwing it over the wall. How there's these different roles and there's all of this that hasn't really decided who is going to be the owner of this, but there are people that are interested and there are people that want to ensure that it's happening.
A
Amit Paka47:26
Yeah, we've got secondary users beyond just data scientists and the actual MLE. These are folks who are decision makers running the business who want to know how the ML is working. Then there are folks within ops who are taking calls from customers on why they're seeing the output they're seeing. And then there are customers themselves who want to know more and build trust with these models that impact them. So while there are primary users, there's a whole host of secondary users, and we've built the product to be able to be consumed easily by this range of users who might not be as technical as the primary users.
H
Host48:35
I think that's so crucial. I know that a lot of times on the community we have these intense technical discussions: which AWS tool to use, is GCP doing this, should Vertex be a product that I should adopt? But the really hard question is what are we doing here and who's doing what based on the talent that we have? And I think that's really the key to the success of ML because we can talk until the end of time about the relative merits of different tools and different development processes, but if you don't really have a clear target or if you don't really have a good idea of what success looks like for AI in the context of your business, what success looks like for machine learning and how different people fit into that, it doesn't really matter what tools you use. I mean it really doesn't. And I think today, the company that I work at, that's the clarity that we're pushing towards, and it really has not been driven by the fact that I have adopted one experiment management tool versus another. I think one of the questions I have because it's consultancies, ops companies themselves, and also some consumer type companies: what would you say is one industry where the movement around explainability you were surprised by how robust it was? I guess I'm just looking for your favorite surprise.
A
Amit Paka50:30
Well, I think banking is still one where you would expect, 'Hey, there's regulation and there's a little bit of an unknown, so that really causes teams to take a step back and maybe become a lot more conservative about the new technology.' However, I was surprised that they all wanted to push the boundaries of what AI models can do. We've seen as we work with banks, we've seen complex models, not just your typical ML models like XGBoost or logistic regression, we've seen deep learning models with complex data types. So the surprise was how much they want to adopt AI and how much they want to push that boundary and ensure that they're still abiding by the regulations and not just staying on the sidelines and saying 'We'll wait.' So banking, which has historically been very conservative with high regulations, in the domain of AI they are forming teams, for example, they are forming AI governance teams where the teams are formulated only to validate machine learning models. So they try to crack machine learning models using tools like ours and validate them. So they're rapidly forming these new teams and doing a lot of investments. Very interesting.
H
Host52:44
Yeah, it's actually very interesting to hear you say that because I certainly had that preconceived notion. You know, so many pop culture descriptions of banking and finance kind of say the traders want to do something and then compliance or risk management says no. So that's interesting. What about insurance? I would have thought insurance would have taken up a lot of AI/ML, but we are not seeing as much adoption as you would have expected. It's lightly regulated, not as much as banking, so you would have seen a lot more adoption through premium pricing and whatnot. And from what we've learned, it's due to a lack of data. Not sure if this is applicable to all the insurance companies, but that's a surprise in the opposite way where I would have expected a lot more adoption here but it's not, and in banking it's like there's a lot of adoption where I thought they'd be more conservative. So banking is ahead and insurance is behind.
Regionally and geographically, you tend to find a lot of, at least in health insurance, you have regional providers in different markets. Even in auto and life and different markets, you'll have places that have their own sort of geographic focus, whereas I think in banking and finance, with what we've seen through the 80s and 90s, it has been this movement towards big national and multinational companies that maybe are pioneering this investment more so. That's interesting. I always like to ask that because these different industries impact one another in terms of how they go about things. In terms of another question I kind of had in a slightly different direction: for the ML engineer or the data scientist that's listening to this podcast, what advice would you give to them in terms of how to get started and how to get their company, maybe their engineering group, maybe the executive buy-in? How would you advise them to kind of lead that change in the same way that DevOps has changed models? What advice would you give to that profession?
A
Amit Paka55:28
Different functions have different incentives that move them towards this goal. For machine learning teams that deploy models and use cases, given the investment that you need to make in building machine learning teams and tooling, the use cases that they deploy to first are the highest value. So you've made these investments in machine learning, and you built this model. Do you want to make sure that your business metrics are still working well and that you stay on top of it in real time? You know that if your model has operational issues, you can jump on top of it and address it. Or if you're in a lightly regulated or highly regulated domain, you have visibility into how your model works so you actually protect it against not just fines but also PR risks. So that resonates with a business buyer or a user on the ground. Within DevOps, within the actual management part of these models, you can call it old-school DevOps or new school ML ops. They have seen the pain of managing just five models. So how do you manage this massive scale of AI models that is coming your way? If you have five models, save your time. You don't want to spend your time babysitting these models, and that resonates to that audience that's managing these deployed models, whether it's an MLE in newer teams or old-school DevOps. And for data scientists, the thing that resonates is helping business teams understand how their models that they're working on are behaving. Whether you need to explain it to a regulator or a compliance officer, but beyond that, if you're not in a regulated domain, knowing how your model works can help you unlock potential performance benefits, can help you debug the model. Models are incredibly hard to debug. You're just training a model for high accuracy but you don't know how the accuracy is actually distributed. So the thing that resonates with a data scientist is helping others understand how the model is doing, helping them understand it so they can debug it or improve the performance. Those are the three kind of groups and the three incentives that resonate. So for a data scientist who's looking to get started, you've got operational insights. Do you have tools to help you ensure that you stay on top of when these models go bad or go awry? And then go and talk to the actual business owner and ensure that they have the right tools to ensure that the performance of these models in high ROI cases are still being maintained, and they have the right tooling because typically the actual budgets might come from somewhere else but the user might be somebody else. So bringing this all back full circle to where we started the conversation.