Back
Eric Colson
Chief Executive Officer & Director, ARTISAN PARTNERS ASSET MGMT

Why 90% of Data Science Fails — And How to Fix It — With Eric Colson

🎥 Jan 30, 2025 📺 Delphina ⏱ 69m 👁 11954 views
Eric Colson—former Chief Algorithms Officer at Stitch Fix and VP of Data Science and Machine Learning at Netflix—explains why ...
Watch on YouTube

About Eric Colson

Eric Colson, CEO of Artisan Partners and a former data science executive at Netflix and Stitch Fix, has discussed the challenges companies face in leveraging data science effectively. In a February 2025 podcast, Colson argued that many firms treat data scientists as a support function, limiting their impact by only executing ideas from business teams. He advocated for giving data scientists autonomy and accountability for measurable outcomes, and for using trial-and-error experimentation with cheap failures to find winshol. Colson also emphasized the importance of decoupling algorithms from applications and enabling data scientists to frame problems rather than simply optimize within inherited constraints. In earlier appearances, Colson addressed the asset management industry, stating that "a lot of true active management got diluted" as firms prioritized growth over differentiation. He described Artisan's model as centered on investments, people, and trust, and noted the firm's introduction of "investment degrees of freedom" to allow teams to deviate from benchmarks. Colson also discussed value investing, saying Artisan seeks stocks that are out of favor and positions itself differently from the herd, and highlighted the firm's expansion into global and alternative strategies, including a China post-venture strategy.

Source: AI-verified profile updated from Eric Colson's recent appearances. Browse all interviews →

Transcript (31 segments)
E
Eric Colson0:00
The main challenge is that there's a lot of companies that treat their data scientists as a support function. Their role is to help the various business teams like product management and marketers, and the challenge with that is the ideas are coming from the business teams to the data scientist. That could have some value, but it leaves a lot on the table because there's a lot of things that we can only get the other way from data scientists.
I
Interviewer0:25
That's Eric Colson, data science and machine learning advisor, former Chief Algorithms Officer at Stitch Fix and former VP of Data Science and Machine Learning at Netflix. He's seen firsthand how companies fail to use data science properly and what it takes to get it right. Let's get into it with Eric. Hey there Eric, welcome to the show.
E
Eric Colson0:50
Hello Hugo, great to be here. It's so much fun to be doing another podcast with you. It's been maybe five or seven years since we last spoke.
I
Interviewer0:58
I don't know how long it's been. I guess probably about that. It was definitely not in the last three years.
E
Eric Colson1:04
Well, it definitely was in the before times, before COVID. And you were at the time you were Chief Algorithms Officer at Stitch Fix.
I
Interviewer1:09
Okay, that makes sense. And we talked a lot about what you did at Stitch Fix and what you've done at Netflix before, in both roles that really elevated the role of data and the data function at organizations that were already data-centric. So much has changed in the space since we last did a podcast, so much has stayed the same as well. I'm just wondering to start off, how you see the role of data science evolving in organizations today and what are you most excited about?
E
Eric Colson1:37
Well, let's see. I do think data is becoming more and more central to companies. Companies finally understand this. It's not really the network, it's the data. It's the data that we can learn so much from. So I think we'll continue to see that it becomes more and more central to companies, and companies continue to organize around the data functions. In fact, we'll see even more representation at the C-level: more Chief Data Officers or Chief Algorithms Officers, that kind of thing. And what excites me about it is actually not really as relevant to the businesses as it is to society. The learning that comes out of studying complex systems can be really valuable. Businesses are like the lab rats of the market. At the macro level, we have all these startups that are trying different ideas, and some of them manage to find product market fit for things you would have never even imagined. If we didn't have so many startups trying these things, we wouldn't have found the successes that we had. So that's great learning. At the micro level, this is what's even more near and dear to my heart: internal to any one company, you have some companies that have millions or billions of customer interactions that are so good to learn from. These are things that we probably would have never dreamed of. Some of the learnings that we can pull out of this data, I think some of this learning is applicable outside the world of business to society at large. I think this could have big implications for how we see the world and how we even do policy decisions and so forth. It could be some really valuable stuff that can be applied to society at large, and it's exciting that this field called data science is really the medium for this type of research.
I
Interviewer3:27
I love that framing. I think in a lot of ways we forget that we're still in the early days of data science, a lot because the space has moved so quickly. But I'm interested in we've seen the data science function leveraged in a lot of different ways, and something you've written about—I'll link to your wonderful essay called 'Beyond Skills: Unlocking the Full Potential of Data Scientists'—you've written about the untapped potential of data scientists and in particular their ideas. So why do you think so many organizations fail to leverage data scientist ideas effectively and what are the most significant missed opportunities?
E
Eric Colson4:06
Well, I think the main challenge is that there's a lot of companies that treat their data scientists as a support function. Their role is to help the various business teams like product management, marketers, merchants, finance, and so forth to help them in their efforts, which of course they should be working with those teams, but it should be more of a collaboration. The way they work together is important. What happens in a lot of companies if they're treated as a support team: they're valued for their skills—things like Python, SQL, statistics, and so forth—as well as just being a resource, a body that they can dump work on. People in finance or marketing have some data-related tasks they need done, and they ask the data scientists to do it for them. The challenge with that is the ideas are coming from the business teams to the data scientist. That could have some value, but it leaves a lot on the table because there's a lot of things that we can only get the other way, from data scientist to the business. So you need to at least establish a bidirectional flow of ideas, and that's what doesn't seem to happen when you treat them as a support staff, just dinging them for their skills. Even well-intentioned companies—a lot of companies say, 'Oh we don't do that, our data scientists we want their ideas'—and they may say that, but sometimes they behave in ways that suggest otherwise. One of the symptoms of this is when the data science team is handed down requirements. There's some new initiative, it's clear it's going to be a data science initiative, but it came from one of those business teams and the problem's already been framed and speced out, and they just need the data scientist to execute it. That will limit the way the data scientists contribute. That's one bad symptom: if you're handing off requirements, that's probably a warning sign. The other one is simply overwhelming data scientists with tasks. This happens in most companies. I talked to a lot of people, and this is a very common symptom: the demands for data, ad hoc requests, dashboards, data pulls—all that kind of stuff is vast. Everybody in the company needs some information, and if you have to ask a data scientist for any bit of it, they can easily get overwhelmed. Unfortunately, it creates a vicious cycle because anytime you ask a data-related question, it tends to evoke more questions than it answers, so there's a lot of follow-up work and the next thing you know, you've consumed 100% of your data science resources. Those are some of the symptoms that happen, just not allowing their ideas to come to the table. In my experience, this is the missed opportunity because I think the value of data scientists does not reside in their ability to answer questions or complete ad hoc requests, but rather in their ideas that they can bring forward. By ideas, I mean capabilities, new things that can move the company in better or new directions. These usually manifest as algorithms—a new recommendation algorithm is a good example, or even in inventory management systems, there's often a place where you can plug in new algorithms to figure out how to buy more merchandise or buy it more effectively. These are all examples of business ideas, but they're very unlikely to come from the business teams because they rely on certain qualities that only the data scientists have. Specifically, the ideas from data scientists are uniquely valuable because they have two things: a different cognitive repertoire and different sets of information. We should go into each of those. Cognitive repertoire is a term I got from the complexity researcher Scott Page. He describes it as a set of tools or ways of thinking that an individual can draw upon to frame a problem. It's not the skills they have, but rather a way of thinking about a problem. Every function has their own cognitive repertoire: marketers have theirs, merchants have theirs, finance people, etc. Data scientists certainly have their own set, and they tend to be really relevant to business. These cognitive repertoires include knowledge of various machine learning techniques—boosting, regression, deep neural networks, all these types of techniques for training models. I don't count it as a skill; the ability to execute on it is one thing, but it's the ability to recognize when it's applicable to a business problem—that's a way of thinking. For example, 'Hey, we should frame this problem as a Markov chain or something like that.' That is cognitive repertoire. These also include knowledge of classic papers or other framings that can be reused—the secretary problem, news vendor model, traveling salesman problem. These were written 50 years ago but are wonderfully relevant to a whole myriad of business problems. Data scientists have had those in their education or training, and many other functions have not. They can bring that framing to an idea, and the framing is really important. The way you frame a problem can determine its robustness, its scalability, and so forth. Just as a trivial example, I remember the early days of Stitch Fix. The founder had maybe a handful of employees and had hired some contractors to build some things for her. One was a styling algorithm, a recommendation engine, and she would put air quotes around the word 'algorithm' because she knew it wasn't the greatest thing. It was akin to what they used to call in the 80s and 90s 'expert systems'—a lot of if-then statements. 'If this dress is blue and the customer said she likes blue, then that's a decent match, so let's add 130 points to the score.' The score was this arbitrary unit that had no meaning other than higher was better, but there was no way to interpret 130 or 1500. It was just some number generated by all these if-then statements that came out of conventional wisdom or were just made up using intuition. When we were able to put a data scientist on this, we didn't say, 'Hey, can you come up with some new rules or new amounts to assign to those conditions?' Instead, we said, 'Can you reframe this problem? What would you do?' The data scientist came up with a fairly basic but effective framing: they framed it as logistic regression. They said, 'Okay, we want to see which of these pieces of merchandise are the best selections for this customer. Let's run them each through a logistic regression algorithm, and let's use a score but assign it some meaning. We'll scale it so it's bounded between zero and one, representing the probability of purchase.' So a number like 0.42 would mean a 42% chance the customer is going to buy this thing. It gave it some meaning, it also adhered to a logic curve which fairly well represents consumer decision making, where it's a little more elastic in the middle and less so at the ends. The fact that you can never get to 100% certainty—you can't say with 100% certainty she's going to buy this, and likewise you can't say with 0% certainty she won't buy this—the logic curve enforces good principles. And instead of making up numbers to assign to rules, a logistic regression algorithm can learn its parameters from data. This is going to be a much better solution: easier to interpret, the score has meaning, it's bounded, and it's going to be more accurate because its parameters are learned from data. It's going to be easier to extend as you add more features. The same basic algorithm will work and get retrained. So much easier to scale and extend. While it's very basic, it still requires at least undergraduate statistics, which may not be something a product manager, merchant, or marketer has in their repertoire. That's why when you come to data scientists, don't tell them exactly what to do—give them the problem and let them frame it. You'll leverage their cognitive repertoires. There are so many other ideas—logistic regression is fairly simple, but also using embeddings to figure out customer preferences, or even using a genetic algorithm to design clothes. These are ideas born out of the cognitive repertoire of the data scientist and brought forth to the business.
I
Interviewer13:27
I love that example in particular as well, because logistic regression is an algorithm which we all have people who work in data science have strong intuition for now. It's also something we can explain and show a figure of to non-technical stakeholders. It may be a relatively good first model, and it's something we may get significant lift on if we do more deep learning techniques, or we may not. But it is something as a baseline model that we can build on to develop more sophisticated models if more lift is important.
E
Eric Colson13:55
Absolutely. I think that is a great way to start: start with something that's very transparent like logistic regression, where it's very interpretable, you can actually poke around and get the first derivative of each of those parameters, and it makes sense. I think what you should do is you can add the complexity later—you can move to neural networks or something that may improve your accuracy but might diminish your ability to understand it. You should make that trade-off consciously. Stitch Fix has moved on from logistic regression. There may be increased costs like compute cost, but also headcount cost for maintenance, latency concerns. There are all types of trade-offs we're talking about here. Data scientists once again have an intuition for these things, and you should be conscious about those trade-offs. If it's just a tiny bit more accurate, it may not be worth all the complexity or opacity that it brings. You may say, 'Well, it's good but not good enough to justify that additional burden.' We'd rather be transparent, have algorithms that are transparent. So I do advocate to start with the simplest thing and then increase complexity as it's justified.
I
Interviewer15:04
Absolutely, but right. So this is a good example of a cognitive repertoire. Those things the data scientists have—the knowledge of those things—you want to bring them out. You don't want to just impose your will on them. So that's one reason. The other reason you want to get their ideas from data scientists is because they have different information than you may have. All the teams in a company—marketing, merchandising, product managers—they all have different sets of information, or let's call it deeper knowledge of different areas. For example, product managers may better know how customers are using the service, or marketers may be more in tune with the addressable market and how people perceive the brand. But data scientists have information too. They are pretty tuned in because they are deeply ingrained in data all day, all night, playing with things. In so doing, they see a lot of patterns, distributions, and relationships that nobody else has seen. A lot of these are completely unintuitive, a lot of these are stumbled upon by accident, and those that are particularly unintuitive can be very valuable. Often these new insights can be brought forth from data scientists and can lead to really startling new novel ideas. The one example I gave in the article was about a data scientist tinkering with some data, noticing that the customer segments the company used were not meaningful, at least to explaining customer behavior. These were segments or personas that the marketing team created and customers opted into, so they self-identified with them. Groups like 'edgy,' 'casual,' 'bohemian chic.' In the signup flow, there were pictures of these personas, and customers would opt into one and say, 'Yeah, this one's me.' So a customer says, 'Yep, I'm preppy, that's me.' Out of curiosity, the data scientist was digging through the data to see what these different personas buy. To her surprise, the lists were the same. All groups bought all products at about the same rate. The groups were completely indiscriminate of behaviors. That was really curious. So she dug further. The inspiration was, 'Oh, that's interesting, I got to look into this.' Nobody was asking her to do it. What she did was very clever: instead of using these marketing-defined labels, she decided to let the data do the talking. She created features out of customer behaviors rather than what they say—things like what they click on, what they view, what they like and dislike. She used unsupervised methods to find the directions in the data that explain the most variation. I think it was principal components plus matrix factorization, if I remember correctly. The outcome of this little curiosity dig was she created this multi-dimensional space. The number of dimensions was vast, but she was able to collapse it down to relatively few, like 12 or 14. These dimensions you can place the customers within that space. Groups of adjacent customers became a cluster that you could actually put a label to after the fact. You didn't create the label first; you found the cluster of customers and then assigned a label after the fact. She had to come up with some clever words like 'femininity' or 'rustic' something. So she had to come up with labels after the fact, but the beauty is these groups were now very meaningful. They had different behaviors, they bought different products. This was a major breakthrough because we had believed those marketing-defined segments were real, and we were managing inventory towards them, and our marketing messages were geared towards them. To find out they're all buying the same stuff—that's not good. With these new segments that were more data-driven, we were able to make major improvements to the recommendation, and that was the first obvious application, but also to the way we manage inventory, and marketing messaging—all that got changed to adopt this new way of defining style that was data-driven, not what the customer said, but what they did. That started from a curiosity project that no one was asking for. The observation drove the idea, not the other way around. We don't say, 'I have an idea, now let me go find the data to support that idea.' Instead, the data presents the opportunity, and then we start saying, 'What ideas can we do to leverage that bit of data?' That's what I mean when I say data scientists have access to more or different information, different from the rest of the company. That's why you want to leverage those ideas. These are things that couldn't possibly have been asked for by a product manager, marketer, merchant, or even a data science manager, because they weren't so much conceived of as they were revealed by the data. It really comes down to those two things: the cognitive repertoire and their observations in the data that allow data scientists to bring really valuable ideas to the table.
I also just want to say I love the idea of the cognitive repertoire allowing data scientists to reveal patterns in the data and then that being able to impact business decisions. I do like this example a lot as well, because not only does it tell us that perhaps some teams even with domain expertise find things challenging without having data scientist cognitive skills or parts of the data scientist cognitive repertoire, it also highlights how self-reporting bias can be something we very much need to consider. But on top of that, it really does this idea of clustering these things to allow the patterns to reveal themselves and impact business is so useful. So I am interested in: we started this conversation talking about how a lot of data functions are treated like service centers and assigned tickets on Jira boards or whatever it is. This seems like a challenge of incentives among other things, and challenges between short-term incentives which are like 'we need dashboards now' and medium and long-term incentives. So I'm wondering just culturally how you think about managing these incentives and creating an organization where data scientists can be leveraged for their ideas.
E
Eric Colson21:51
Yeah, it's a great point and it's not easy. This is where the soft skills come out. How do you set it up so that motivations are aligned? Data scientists want to do this type of work, and the company wants someone to do that work. How do you just let it come together? There are a few tactics you can do. I remember at Netflix we had a value—I don't know if it still is—but it was 'context not control.' I've adopted that even after leaving Netflix. I brought it over to Stitch Fix and changed it a little bit: 'Give them context, not tasks.' My counterpart, the original CTO at Stitch Fix, he actually had a better phrase. He kept saying, 'Tell me the problem you're trying to solve.' I'm sure he wasn't the original author of that phrase, but he said it very effectively. He would say it over and over again. I think everybody in the company could recite it. So anytime somebody came to him or one of his engineers with a task, 'Hey, I need you to change this to this,' they said, 'Hold on, tell me the problem you're trying to solve. Give me the context, and maybe I can come up with something better.' I followed suit. I did the same thing with my algorithms team. I would encourage people that if you are going to come to them with some new opportunity, let them do the framing. Tell them the general thing, the context, maybe even some constraints, and then let them come up with a solution, because that leverages that cognitive repertoire, lets that come to the table. Oftentimes you don't even have to ask explicitly—'Hey data scientists, what ideas do you have?' All you have to do is invite them into some meetings anywhere where context is shared. I think context will collide with their cognitive repertoire or their extra information that they have and spawn new ideas. In the article, I give that narrative about a data scientist attending an operations meeting and somebody from merchandising says, 'Gosh, we really need to figure this out. We need to know how to buy enough inventory but not too much.' That triggers the data scientist: 'Oh my God, this is the news vendor model. I know this one. I could solve this.' They almost can't help themselves but start writing down a solution immediately. There are millions of examples of this type of thing. If you receive the context and you have that kind of cognitive repertoire or extra information, you can often solve these problems or come up with a new way of solving or a new idea for something. So just expose them to context is another way to do it. I also mentioned in the article about getting rid of Jira. In my opinion, Jira is just not the way to engage with data scientists for several reasons. As you enter in whatever it is you're entering into Jira, it tends to strip it of all context. It just becomes a pithy little command almost, and there's not any context conveyed. Also, they're just frankly too easy to submit. If something is important enough, take the time to set up a meeting with a data scientist and meet with them face to face where context and ideas can be shared. That'll be much more effective. The simplest thing to do is just make them accountable for impact. Say, 'Hey, you're in charge of making some revenue or improving retention.' That'll almost immediately flush out some ideas from them. It'll also help them reprioritize, and it'll help them if they don't have the context, they'll go out and seek it. It has an amazing impact. Just saddle them with a little bit of accountability.
I
Interviewer25:17
I love the idea of everyone being responsible for impact and value. I do wonder—firstly, I'm going to break the cardinal sin and ask two questions at once, I think, but they're so coupled. I feel like for this to be the case, we really require buy-in from leadership on the data function to let the data function be free, in a lot of respects, because it isn't cheap among other things. But on the other hand, do we need to decouple the building of algorithms and analytics from infrastructure and deployment in order to—because in my mind, if you build a recommendation system model or system, your ability to deliver impact is actually intricately tied into how it's deployed and even the front end for it. So how do you think about this matrix of considerations?
E
Eric Colson26:07
I think what you're asking about is how do you decouple the near-term things—like implementing this specific algorithm and maybe you'll get a win—versus building the infrastructure to not just run that algorithm but any future algorithm. That is a tricky balancing act. That's part of the judgment of the data scientist. Engineers do the same thing on their side. Whether they're building a transaction processing system, they're not asked to build the back end; they're usually asked for front-end things, and they're using their judgment. They're asking for something very specific, but let's build something more general because there'll be more stuff coming. Same thing happens on the data science side. You need to really invest in a platform that anticipates future needs. Right now we're just going to be running algorithms that are linear combinations of things, so we can build that now. But also lean in that probably won't always be the case. We're going to get into ensemble methods and other things that are more complex, so we'll also need to build that stuff too. It's a bit of a juggling act of how much resources to put on near-term things that could have a very visible win versus the behind-the-scenes infrastructure that is less visible to the business people but absolutely necessary. That's where you rely on really good judgment from the leaders of your data science team to make those decisions.
I
Interviewer27:37
Fantastic. Something we're speaking around, I think, is experimentation and trial and error, because you've cited two successful examples, but data scientists try a lot of things that don't work, and we want to do it very rapidly as well. I think I want to set the scene and lead the witness slightly by giving an example where we hear things like 80% or 90% of ML models don't make it to production, and that's bad. My response always is: maybe that's great. If you're testing enough good models and 90% don't make it into production, but the 10% that do are fantastic and do their jobs, and not only that, the wins of them outweigh the losses by the ones that didn't work. So I'm interested in your thoughts on your experience of the role of trial and error and experimentation in data science and in business at large.
E
Eric Colson28:31
Well, I've heard this similar statistic that lots of data science-related projects are failing. I haven't dug in enough to know whether it was a successful trial—meaning you tried something and the idea didn't work out—or the execution itself failed. I'm not sure which one it is. But I will say that I believe for some companies, trial and error is the way to go. What do I mean by trial and error? It's hard to define, but let's go with the opposite. The opposite I'll call overly rigorous planning and execution. That's where the company may take a lot of ideas and instead of trying a lot, they take just one or two, put all their eggs in a few baskets, and do endless research to plan them, making sure they convince themselves that these things have a 100% chance of success. They plan them out to the nth degree, done their market research, customer surveys say these are for sure going to work, the only thing that could possibly go wrong is execution error. The problem is that a lot of companies haven't done a lot of experimentation. One of the first things you learn when you embark on experimentation is that you realize you're wrong a lot. A lot of your brilliant ideas do not pan out. Almost all business ideas are intended to improve things—revenue, retention, customer experience—but when you properly measure them with a randomized controlled trial, 80 to 90% of them fail. They either do nothing at all—no statistically significant result—or they actually fail; they did the opposite. You intended to improve revenue, you actually hurt revenue. That is a startling thing to learn, it's very sobering. I believe most companies do not know this. It's only the ones that do a lot of experimentation that have learned this. Once you learn that, you realize you have to do something different. You can't bet on just one or two ideas. The opposite of overly rigorous planning and execution is trial and error. That's where instead of trying just a few ideas, you try more. Even if you can't change your success rate, at the same success rate, if you just try more ideas due to sheer volume alone, you're going to find more wins. You'll get more failures too, but if you can figure out how to make the failures as cheap as possible, then it could be very lucrative. I think it's a general philosophy that could apply to a lot of companies and a lot of functions—marketing, merchandising, all the functions could benefit from doing this under certain conditions. That said, I think data science has it better in most cases. They have three properties that make them super amenable to trial and error. I distilled it down to at least three, sometimes I describe it as four or five. The first reason why data science ideas—and I usually mean algorithm ideas—are more amenable to trial and error is because they're cheaper to explore and try. Let's separate exploring idea versus trying idea. Exploring really refers to doing enough research to flush it out, to know if it's viable, to see if it's a promising idea. I think these data science ideas can be explored really cheaply. In fact, it's happening all the time from your data scientists. They're a curious group, as I described earlier. Nobody was asking to look into finding new customer embeddings; she did that out of curiosity. That happens more often than people know. The reason is that the data is right there at their fingertips, so exploring ideas is really easy. Example: there's a new data set that came to the company, maybe social media posts are now included in a data warehouse, and all your algorithm developers have access to that. You don't even have to ask them; they're going to hit that really quickly. They're going to just simply open a new tab in their Jupyter notebook and they're off to the races sifting through that data. It's amazing because they can explore it really quickly. Within a few hours, they can get a sense of the distribution of values for the data. They could even, if they have the right infrastructure, try it as new features in an existing algorithm. This is against historical data, not a real A/B test, but they could add the new features and try them out just to see if there's signal there. If it's promising, then maybe they can take it to the next step to be a real trial. I think this exploration of ideas is constantly happening with your data scientists even if you're not asking them to do it. They don't have to ask permission because the data is right there at their fingertips. By contrast, I like to tell this story—it's an apples to oranges comparison, but—I remember a colleague. She was in marketing, exploring the idea of a new loyalty program. Apples to oranges: a loyalty program versus a new feature in an algorithm is very different, but the reason they were somewhat apples to apples was because the expected impact was about the same. For her to explore this loyalty program idea, she didn't have the data at her fingertips. She has to get time to dedicate to this, she has to ask permission, she has to go outside the company sometimes to look around and do some research on what other companies have done, she may have to engage with consultants, she at least has to chat with engineers to see how hard this would be to build. It's a lot of information that she doesn't have at her fingertips, and she has to gather. It's much more expensive to gather. The key is that she's going to ask permission to do that. She has this idea, she's going to go look into it. 'Can I get some time?' Where the data scientists are not asking; they're just going because the data is right there at their fingertips. That's the difference in low-cost exploration. Now that's just to gather enough information to say, 'Should I keep going?' In the case of the data scientist exploring new features for an algorithm, they may have flushed it out enough that they're confident enough to try it in production. This is where infrastructure assumptions come into play. If you have a good infrastructure, that code can be improved a little bit just enough to drop it onto a data platform, and it can take care of it, abstracting the data scientist from all the complexities of distributed processing, automatic failover, containerization. I know Metaflow does a lot of this stuff. Metaflow came after my time at Netflix; we had to build our own type of solution at Stitch Fix. This could really enable data scientists to actually try something in production for remarkably cheap, nearly free. They may still not have even asked permission. They can allocate themselves a small amount of sample to try a few customers to try out their new version of the algorithm. The main difference is no capital outlay is asked for. There's no funding needed from finance to try out this new algorithm idea. By contrast, if we go back to that loyalty program idea, it's a great idea, but you're going to need funding because you're going to have engineers build this thing, and you have to get designers, and depending on the nature of the loyalty program, there might be prizes and awards that you need to be able to purchase, and also an expectation to keep it running from the demand side. So there's a big difference in the cost profile to explore and try ideas. I think the algorithm ideas can be done relatively cheaply with some infrastructure assumption versus a lot of the other domains that need capital outlay. That's a big difference. So that's the first reason that makes them really amenable to trial and error. The second is what I call evidence. Data science ideas like algorithms typically come with some evidence as to their merit. During the exploration phase, for example, you can use some historical data, add your new features to an existing algorithm, and you can run it and get some feedback in the form of how it improves accuracy or AUC. This is good feedback for the data scientist to have. If there is no signal, it's not doing anything, they might just put it away and move on to their next idea. But if there is signal—big increases—that gives them more confidence to go forward to the next stage, the trial phase: 'All right, well, I better fix up my coding a little bit because I'm going to actually try this in production.' That little bit of evidence is really helpful. It tells you when to stop, it can also compel you to keep going. It is hard for other functions. Again, with that loyalty program example, she may explore the idea gathering all that information we talked about, but at the end of the day, all she has is a bunch of assumptions. She doesn't have the actual empirical data to support this decision. So that's a big difference in the amount of evidence you get. But that's just from exploring. Data scientists can try this out in production. They can allocate themselves some sample and try it on a real A/B test, and that's trial feedback. That is really solid evidence. The exploration stuff is better than nothing, but there's no guarantee that just because you have AUC through the roof, it's going to manifest in production. But it's also fairly cheap to try out. Again, if you have the infrastructure set up, you can easily allocate yourself an A/B test and try it out on real customers. By contrast, for other functions—marketing, merchandising, product—oftentimes it's just not the case. A lot of these things simply can't be A/B tested. Imagine a brand campaign, you can't A/B test that. Opening a new physical store, you can't A/B test it. New partnerships can't be A/B tested. So you can't really get that evidence that you need. In the case of a loyalty program, actually probably you could most companies won't, but you could feasibly try it on a couple hundred thousand customers and only expose it to them. They get the perks of points on every purchase, and you can run it for several months enough to get the feedback you need. But it's going to be costly. You have to get the engineers, the designers, customer service on board even though it's only going to be an experiment. It's still quite an investment to roll that out to get that evidence. By contrast, I think the algorithm ideas are generally fairly cheap to get the evidence you need, whether during the exploration phase or the trial phase. It will produce some pretty solid evidence that can give you confidence. The last reason, the third reason, is that data science ideas are more amenable to trial and error: optionality. What I mean by this is when you try an idea and it doesn't work, you don't have an obligation to keep it going. You can generally pull it down pretty easy. Some ideas are hard to back out of. That loyalty program: suppose you did try that as an experiment and you rolled it out to several hundred thousand customers and let it bake for several months, six months, and you get your results back and find it's just not doing what you thought. It's not moving the needle or even hurting retention. You might have to make the very difficult decision to dismantle it, to pull it down, and it's not so easy. You have to send notifications to the customers that have been exposed to it: 'Hey, you've been using this new feature, but we're going to be getting rid of it.' Some may have earned a lot of points, they'll be disappointed. You have to figure out some way to compensate them. Also when you send notices like that, the press usually picks up on them. You can imagine the article: 'E-commerce company pulls back experimental loyalty program.' That could be embarrassing. Even internally, that merits a broad communication. You have to send it to employees: 'Hey, this thing we were trying, it's coming down, it didn't work.' In my mind, there should be no shame in doing that. It was a well-reasoned hypothesis, a great idea, and it didn't have the outcome we wanted, so we're pulling it down. That's awesome, you tried it. But I do remember a certain case—it wasn't this loyalty program, but something similar—that email went out saying after a year of trying, we decided to abandon this idea. I had that reaction: 'Oh, that was good they tried something.' But a different exact reply-all said, 'We need a postmortem on this. We need to figure out why this didn't work.' I thought to myself, 'Oh, that's kind of going to quell future innovation. Nobody's going to want to try things after that.' By contrast, algorithms are usually very easy to pull down. They're usually baked into the system behind a feature flag, meaning we're going to try it out for this many customers for this long, then it'll automatically shut off. They're behind the scenes; customers don't even know that we're swapping out and trying a new algorithm. No notifications needed. If it doesn't work out, you just revert it back to what they were using before. No messages go out, so no press to deal with. Even internally, you don't need to send a broad communication. Maybe locally with an algorithms team, you might want to share that knowledge: 'We tried this, didn't work.' But no need for broad communication. So that's optionality. Algorithms are really easy to pull back; it's kind of baked into the way we deploy them anyway, versus other things that are highly visible to customers. Those are some of the properties: low cost exploration and trial, the evidence, and the optionality. These make algorithm ideas very amenable to trial and error. I came up with that upon reflection after sitting for years and years in these executive meetings with my peers—the CFO, COO, CMO—where we all had similar pressure to deliver business value. Quarterly or half yearly, we would all present to each other our big ideas that we're going to be trying out. I remember thinking, 'Boy, their ideas all had these big capital outlays. They had to ask for funding from finance. They really had no evidence, just strong conviction or opinion to lean on. And almost all of them were going to be very public, so if they failed, they'd have to do a very public apology either internally or externally or both to pull them down.' I felt like I always had a little bit easier. It made me really appreciate my peers. They deal with far more uncertainty and far more risk-taking than I do. I actually had a lot of gratitude. They're doing the heavy lifting in the company. I'm not taking that kind of risk. I mean, I remember even the facilities manager would have to take more risk than I do. She had to sign like a five-year lease. That's a big capital outlay. There's no evidence that she can use to say where we're going to be in five years. She didn't have any data that's going to really help her with that, and it's going to be very hard to undo if she's wrong. I appreciated her because she deals with more uncertainty. Here I am running my biggest expense of employees, the AWS bill. At the time, I didn't even do reserved instances. I just used the spot instances that were more expensive because I didn't want to commit to one to three years of reserved instances, even though the cost was cheaper. It just really gave me a fond appreciation for the business teams and how much they bring to the table and their risk-taking and their comfort with uncertainty.
I
Interviewer44:33
I love that and I do love that you also mentioned the ability to roll out new features or new algorithms or challenger algorithms to small parts of the user base, even internally first, and that type of thing to test things. I always said with some of the best products out there, I used to joke like Google Search for example: there's no actual search product that we all use. Each of us is using a slightly different version depending on a lot of different things because of the constant rapid iteration that software being in the software game affords us as well. So something we've danced around is the asymmetry of wins and losses in experimentation and how data enables rapid iteration. I'm wondering how you recommend companies scale experimentation while mitigating the risk of failure as much as possible.
E
Eric Colson45:12
Yeah, it's a fascinating thing to look into. I generally do advocate companies should try more ideas rather than fewer. They should scale their experimentation. This isn't obvious when we discuss the outcome. When I share with people that most of our ideas fail—you gave some statistics. If you look at Ronny Kohavi's book, I think it's called 'Trustworthy Experimentation,' he has some great stats. Netflix like 90% of their ideas fail. Google I think it was like 96% of their ideas fail. Airbnb like 80-85% of their ideas fail. So the success rate for ideas is very low. While those might be indicative of online experimentation—each of those companies were probably doing things online like marketing messages and buttons—I actually believe it to be a more general phenomenon. At Netflix and Stitch Fix, we were able to experiment more broadly in the company in different areas like operations, finance, merchandising, content, and we would try these large-scale experiments on things like inventory decisions or even entire product lines. We found the success rate to be similar. It led me to do my own kind of research, calling around different colleagues in entirely different domains like pharmaceuticals and even government policy. I have a friend who does large-scale A/B testing on government policy, where they could take an entire region and try an experiment on. Sadly in those domains too, most ideas fail. So there's at least anecdotal evidence that this is a more general phenomenon. When I explain this to people, they ask, 'Well, if most of your ideas are failing, aren't you just eroding business value? Shouldn't you stop trying?' The answer is no. Even with a very low success rate, trying new ideas and decisions can still be very lucrative. This is unintuitive to people. The key to understanding is that experimentation along with optionality creates this asymmetry between wins and losses. Simply put, the failures are mitigated and the wins are amplified. I need to explain this. We have to go a little bit into A/B testing 101 for companies who don't currently do any A/B testing. The simple message is: before you roll out your idea to all the customers, you should try it first as an experiment. Allocate yourself a random sample of customers that will get your new idea or new decision, and a random sample of customers that will not—they'll experience the absence of your idea. Let that bake for a few months, and then you can compare their behaviors. That will get you to causality: did this intervention really make a difference? When you do that, you will find that more often than not, you're wrong. Your great idea didn't do what you thought it was going to. It either did nothing at all or actually hurt the very metric you were trying to improve. This is sobering, but the good news is you didn't roll it out to all customers. You tried it on just a small sample, and that is part of what creates this asymmetry. The key is to try it on as small a sample as possible. There are power calculations you can do to inform you on how small the sample you can get away with, such that you can still get a good read on a result. It varies dramatically depending on the effect. At Stitch Fix, it was roughly about 50,000 customers that we typically used. Netflix, at least when I was there, was like 300,000 customers. But I've read papers on the big search engines Google and Bing that use tens of millions in their samples. Of course, what matters is the size of the effect you're trying to detect as well as the variance of that metric. You can let the power analysis tell you what size sample you need to detect an effect of at least X. You should make it no bigger than what you absolutely need, as small as possible, because it's likely that your idea is not going to do anything. In the event it actually hurts things, you only hurt a few. The exposure was quite small. And of course, as soon as you know, you can terminate the experiment, which further mitigates any downside. But on the other hand, if you get one that wins, that sample of 50,000 customers all of a sudden is higher retention, higher revenue, higher customer satisfaction. You are not limited to that 50,000. You can now roll it out to all five million or however many other customers you have. This is going to greatly amplify that outcome. Not only that, it's not just a one-time thing. For all future periods, you rolled it out for a quarter and got your result; now you can roll it out to all customers for all future periods. So your upside is nearly unbounded. It can be greatly amplified, at the same time you're greatly mitigating the downside. This creates this asymmetry. It means that even with just a few successes, they can greatly outweigh the cost of all those failures. VCs know this in and out. This is the whole VC community in a nutshell. The VC will invest in a few dozen companies just to find one or two winners. It's been the same for business for a long time. Movie studios and record labels have done exactly that: mostly losses, and then you get Elvis or whatever. Bill Gurley, one of the most well-respected VCs in Silicon Valley—he was our VC at Stitch Fix—used to say, 'If I invest a dollar in a company and it fails, I lose a dollar. But if I invest a dollar in a company and it wins, I win like a hundred dollars.' His downside is limited to what he put in, but the upside is unbounded. It could be many orders of magnitude bigger than the downside. The same thing I think applies to our ideas. Most of them are going to fail, but the few that succeed can greatly outweigh the cost of the failures. It does depend on some assumptions, by the way. What is the distribution of outcomes? Most companies won't know this until they've done enough experimentation to build up a database of these things that they can analyze. But you can even run simulations. Even if you assume a Gaussian distribution, it may have a negative mean—meaning your average idea is going to be negative—and it's Gaussian meaning the chances of a really big advantage aren't very high. Even running with the assumption of a Gaussian, you can find that with 80-90% of your ideas failing, it's still worth it to keep trying because of that asymmetry. There's good news: there's a paper a few years old now called 'A/B Testing with Fat Tails.' They studied Bing, which does thousands of experiments, and they got a good sense of the distribution of outcomes. It suggested it was more fat-tailed. So this makes it even a rosier picture. With fat tails, there is some likelihood that you might find some really big win, some crazy win that pays for hundreds or thousands of losses. That resonates with me. It matches anecdotally my experience. You try many, they fail. You get some moderate wins. Then every now and then, you get this outside win that really lights up, gets you kudos for years to come. So all this is to say, it pays to keep trying even in the face of a high probability of loss, because that exposes you to some chance that you hit the big one. That's really the justification for scaling up experimentation.
I
Interviewer53:22
Amazing. I love that you referenced the 'A/B Testing with Fat Tails' paper, which we'll link to in the show notes. This is the second time we've discussed that paper on this podcast. Ravi Garg, who's at Stanford but does a lot of advising on online experimentation for Airbnb, Bumble, Uber—this is a paper he is very adamant about the importance of, particularly when running large online experiments at massive tech companies. Anyone interested in this type of stuff, please do check out that paper. Eric, we're going to have to wrap up soon, sadly. There's so much more fantastic stuff to talk about. But I would love, particularly as data teams aren't cheap, and I want all leadership teams to have as much buying as possible, that's why I love your ideas of rethinking data science teams as revenue generators and impact generators as well. I'm just wondering specifically what structural and organizational and cultural changes are required to enable this shift, and what impacts can it have?
E
Eric Colson54:23
Yeah, it's one I think it's more available to companies than they know. If they were to make some tweaks to how they're organized, as well as maybe some technical changes, I think they could enable their data scientists to be revenue generators or whatever their objective function is—improving retention, profit, whatever. It can be so clear to me that it's the algorithms that are driving a lot of this impact. They're very easy to A/B test; we can get the causal impact that they're making, and it can be tremendous. Not only that, it can be done independently of other functions. If that's the case, why shouldn't you saddle them with a little bit of accountability? 'We got to put your money where your mouth is. You can try your ideas, but you got to return something at the end of the year.' It's hard to know which ideas will hit, but you can try a portfolio of ideas. Likely you'll get some hits, and you should make the data scientists accountable for that. What I like to do is decouple the algorithms from the applications that house them. For an e-commerce company, a recommender system: engineers and product managers own the website and the app, even the page that contains the 'suggestions for you' page. They build that page, but that little space in the middle where the recommendations go, that is the result from an API to the algorithms team. The engineers on the page call the API, it returns product IDs, and they have to do final assembly—get all the content, images, product descriptions to render the page. But that decision of what to put there, that was algorithms team, not engineering. Their responsibility is just to render the page. The products and their ordering should come from the algorithm team. Likewise, inventory management systems are also owned by engineering. They built it, it's a transaction processing machine mostly to manage the state of inventory. Something gets shipped, you mark that item as out to customer; something gets returned, you mark it as in stock. That's the primary goal. But there are places where there are decision points to be made, such as when to buy more of a product and how much to buy. That could be coming from outside the system from an algorithm, with more information and time to process data, and the results can be inserted back into the transaction processing system. This is how we set things up at Stitch Fix. We had many such systems: recommendation engines, a matching algorithm to match a customer to a stylist—again built by engineering but with a call to an algorithm API. Even things like a visitor hits the web page, what landing page should we show them? Algorithmically determined. The engineers are building the scaffolding, but they're going to make a call to the algorithms team, and we'll tell them, 'Oh, show them this page.' If we show them the wrong page, that's on us. If the page was ineffective at converting the customer, that's our fault. So it's a separation of duties. At Stitch Fix and Netflix, I was never part of engineering. It was separate. At Stitch Fix, we had a CTO that ran engineering, and we also had me, a CAO. We were peers, both reported to the CEO. Engineering owned a lot of the applications, but they left these spaces for the algorithms team to insert their logic. That was good decoupling. It actually even leads to good coding practices. It does take some trust. Creating those spaces in those applications takes cross-functional work. You have to work with engineers to remove whatever logic was there and instead call your API. This is where the trust comes in. Engineers are like, 'This is an API written by the algorithm team. What about our SLAs?' We'll meet them. Sometimes it took a little time to meet their SLAs, even though these were not very stringent—500 milliseconds or even a full second in some cases, not low latency things. But we had to get good at that to earn the trust of the engineers. Once you have it, the magical thing that happens is you're no longer reliant on engineering for trial and error. Engineers have a very different workflow. That's why I recommend them being in separate departments. They like to work with a lot of upfront design, then build their code very robustly from the beginning, and then hopefully move on to another thing. Algorithms people like to learn as we go. We need to iterate. Often that first implementation is just a start, and our best ideas are soon to come after that. Since engineers are ready to roll off to something else, they don't want to be burdened with a lot of changes. By this way, we're able to decouple. We took some time together cross-functionally to build the space, but now we're decoupled. We on the algorithm side can iterate as much as we want, try all our experiments, and we don't even have to burden engineering with anything. They're just calling the same API. They don't even know that we're running experiments behind the scenes. That was a huge unlock. At a lot of companies, engineers are the hot commodity. You need them to make all your changes. Everybody's competing for engineering resources—marketing, merchandising, operations—they all want the engineers' time. By asking them to create the space for us to plug in our output, we are freed up. We don't need to burden them with our changes. We can now work at our own pace and try dozens of different versions of algorithms without any coordination needed. When you have that, it really sets up the case: 'Okay, algorithm team, you're autonomous now. That means you should be on the hook. You can't just be trying anything you want and be satisfied with no results at the end of the year. You have to come with something.' Saddle them with some accountability. Challenge them: 'We want to see at the end of a year or two years that much improvement from you—revenue or retention, whatever the metric is.' But now you've granted them the space to do this. They can work far more effectively in their own way with trial and error. I think it really lends itself to better roles. Nothing is more satisfying than a win—actually creating some business impact that was not just a narrative but causally detected through an A/B test. That's something you go home and tell your spouse about: 'Wow, I had a great impact at work today.' By creating these spaces for algorithms to insert their stuff, you set that up. Not only that, you're no longer reliant on somebody else for success. You don't have to wait to get the engineer's time to try that crazy idea. You can just try it on your own. It could also give data science the justification to say no to all the constant ad hoc requests. It gives them an opportunity to say, 'Actually, I would like to work on that ad hoc request, it sounds really interesting, but I gotta hit my revenue quote. I have to focus on that algorithm I was working on.' It makes the trade-offs much more clear. They can even make those decisions we talked about earlier on when to work on infrastructure or refactoring code. I'm a big fan of refactoring code, but it's not something any business person would ever ask for. It's not visible to them. Often the algorithm developer will know better when it's time: 'Okay, now that I know what I know, I should rewrite this.' They can prioritize appropriately. When is it time to lean in on delivering that revenue goal versus I need to worry about next year's revenue goal? If I don't fix up this infrastructure, we're not going to get there. I love that it puts that in their hands. From Dan Pink's book 'Drive,' he says autonomy, mastery, purpose. I like to modify it to autonomy, mastery, impact. This gives data scientists the autonomy they need, it allows them mastery to get really good at doing these things, and then impact—they can measure their impact. Nothing is stopping them. I also think it makes for happier engineering roles. They don't like to be burdened with these handoffs—'Data science, can you make this change?' We're just as guilty at stripping out the context as anyone else, just handing it off. They don't like that constant iteration. Certainly for a wild idea on a hunch, you're not going to be thrilled if you're an engineer: 'You want me to make this change on a hunch? You don't have any evidence? No, I'm not doing it.' But if it's your own time, all you're wasting is your own time, then you'll try your wild ideas as well. I think that's a healthy thing to do. There may be huge upside there, you don't know. You got to keep trying ideas. So decouple the algorithms from the applications that house them. That grants the autonomy, and then that makes it appropriate to saddle your data scientists with a lot of revenue commitments or improving retention, whatever you're optimizing for.
I
Interviewer1:03:57
And also this does avoid engineering and platform and infrastructure people feeling like a service center as well, and gives them more autonomy around how they build what they build. I haven't thought about this a lot, but I have thought about how to reduce or change the incentives for marketers or people on commercial teams treating data scientists like a service function. I do wonder whether introducing some sort of cost for them as well if they make a request that a data scientist does—there should be a cost for them because there's a potential benefit as well. I don't know what that would be.
E
Eric Colson1:04:27
Yeah, I don't know either. It's tricky because some amount of ad hoc requests is healthy. It forces exploration of the data in ways you might not have done otherwise. But it can be absolutely incessant. You can just bury them. I know a lot of companies try, 'Well, let's get more data scientists then.' That's not going to help because it's a vicious cycle. Asking more questions gets more questions. You could add 10x time, but you could still bury them. It's kind of like when you increase the number of lanes in a highway, it doesn't improve. We've seen what Los Angeles turned into when they tried that—they're still doing it.
I
Interviewer1:05:04
So to wrap up, I'd love—there's so many practical takeaways in here for people working in data, for data science team leads, and chief algorithm officers alike. I'm wondering for senior data leaders listening, what are the top lessons or changes you hope they'll take away from our discussion on how to better leverage their functions?
E
Eric Colson1:05:26
Well, if I could say one thing, it would be to not do merely what's asked, but rather bring some ideas to the table. The worst thing you do is just do what people are asking of you. That's not going to bring out the best of your data scientists. It doesn't leverage their cognitive repertoires or their extra information. You've got to bring those ideas to the table. The other thing is to really embrace trial and error. You have to have the right conditions. I wouldn't recommend it in medicine or manufacturing where the cost of iteration is exorbitant, but in a lot of domains—e-commerce, streaming media, all that kind of stuff—it can be really effective. To embrace trial and error means a few things: you've got to lower the cost of trying, so get that infrastructure in there, develop the muscle to do experimentation and abandon the failures. This way you can leverage that asymmetry—the wins are huge, the losses are not so bad. The other thing is to create those spaces so that algorithms can play more autonomously. Work with your engineering team to create all those different places where you can plug in the results of algorithms. The last thing I'll say is you've got to be a good partner. Other teams—marketing, merchandising, product—face far more uncertainty than you do and take far more risks. Appreciate what they bring to the table as well. I have a great fondness for all my former peers at Stitch Fix and Netflix who did those things. I'm just glad they did them and I didn't have to do them. Those are some of the big rocks that really move the company forward. I was very appreciative of them. It is amazing what can happen when all these different disparate teams work together, combining their repertoires to create something better than any one of them could have done on their own.
I
Interviewer1:07:15
Fantastic. So don't do merely what you're asked, bring your ideas to the table, orient the team to make real business impact, trial and error, and be a good partner. I think those are wonderful takeaways for everyone. Eric, as always, I love speaking with you and appreciate all the wisdom from all the work you've done and continued to do in advisory roles. I'm also—we didn't discuss this today—but I'm super excited that you're finding more time to write and think and contribute to the space in having these types of conversations in ways that you perhaps weren't able to when you were leading massive functions as well. So that's super cool.
E
Eric Colson1:07:52
Yeah, thank you. It is a fun thing. It's hard to let it go. I'm no longer working at a company, but I still tinker, and I love to talk to anybody who will listen to me about this kind of stuff. I find myself picking up the phone to talk to different people in very different companies, and it's amazing to hear their stories. Yet there's a certain commonality across all companies that leverage data scientists—they all seem to have the same problems. So it's great to have some space to be able to think about how we solve these things, how we get data scientists out of those ruts they're in at these companies, and how we spread that knowledge around so that we can create better roles.
I
Interviewer1:08:34
Without a doubt. And to your point, coming back to the start of our conversation, we've always thought that data would be a valuable resource for organizations, but as it turns out, it's becoming more and more the basis of a lot of defensible modes. Even with the rise of LLMs, which may make you think everyone has access to the same models, lo and behold, if you can fine-tune them or use prompt engineering with your own data, that's where your moat becomes far more defensible. So it is a very exciting time for this space and the conversations around the principles of data—AI stuff aside, just the general principles of how to leverage data to improve product customer experiences—all of these types of things. Yes, absolutely. Cool. Thanks once again, Eric.
E
Eric Colson1:09:14
Of course. Thanks so much.
I
Interviewer1:09:16
Thanks for listening to High Signal, brought to you by Delina. If you enjoyed this episode, don't forget to sign up for our newsletter, follow us on YouTube, and share the podcast with your friends and colleagues. Like and subscribe on YouTube and give us five stars and a review on iTunes and Spotify. This will help us bring you more of the conversations you love. All the links are in the show notes. We'll catch you next time.