Back
Boris Mouzykantskii
Chief Architect, Criteo (founder & former CEO of IPONWEB), IPONWEB (part of Criteo S.A.)

Iponweb CEO Boris Mouzykantskii - ATS London 2012:

🎥 Sep 26, 2012 📺 ExchangeWireTV ⏱ 28m 👁 3054 views
... to part and pass some of this money over to the publisher uh so it's in their self-interest to buy the cheapest possible impression ...
Watch on YouTube
Transcript (12 segments)
B
Boris Mouzykantskii0:06
You might not have heard much about us. We're not really used to advertising; we're very busy building advertising technology. So I can't resist the temptation to put a couple of slides about our company first and then switch to something which is probably more interesting. So that's what I will try to talk about today.
Here are a few things about us. We cut our teeth building early RTB media platforms with some of the people in the audience. Exciting times, we learned a lot, made a lot of mistakes, and fixed those mistakes a few times. These days we are very busy building custom media trading solutions, trying to take the learning from the early days and apply it at huge scale. Every time our clients ask us to build something in this space, there are several really complicated, difficult engineering problems which we are facing. So by providing solutions again and again to these problems, we sort of became better in machine learning and optimization and that big data analytics stuff. We operate in a few countries across several markets. Here is roughly the platforms which we were building.
Now I put a timeline on the type of clients who work with us. One interesting thing we see is that in early days, the need for technology was coming mostly from intermediaries, from companies like exchanges who sit way away from advertisers and away from publishers. As time goes by, there are DSPs and SSPs appearing which are, we can argue, one step closer to either advertisers or to the publishers, so they want custom technology. Then later on we see big agencies, we see e-commerce players, so that's basically an advertiser as far to the end of the value chain as it gets, brand advertisers and publishers. So the interesting trend which I want to focus on and pick up on is how the need for custom technology moves closer and closer to the actual end customers. We think we know what drives this. It's really not us, it's what we see in the market. Our guess is that it's all down to conflicts which are embedded in the ecosystem, and it's all down to how those conflicts are refereed or mediated or balanced by the party who controls technology. Now if you think about an end-to-end stack providing a technology solution all the way from publishers to the advertisers, then there are so many conflicts embedded in it that I can't really talk about it properly. So much so that big players like Google will tend to actually separate these stacks in almost independent companies like Invite Media for the demand side and AdX for supply side and almost manage them separately. So that's one way to avoid it. Now if you move to just DSPs and SSPs and consider them separately, then on the supply side there are many very interesting, very weird interactions between what different parties want. And actually the technology solutions which come about to manage these conflicts are sort of controversial; we're probably a bit afraid to speak about them. So instead we'll focus solely on the demand side.
So what's the problem? What are the biggest conflicts which we see on the demand side? Here is the dilemma in very simple terms which face you as a demand-side RTB player. Let's say you're a sort of decent-sized player. You listen to say 10 billion impression bid requests a day, and you run some campaigns, maybe a couple of campaigns, maybe a few thousands of campaigns. Your good typical campaign wants to buy a million impressions a day. That means that for every single impression which you need to satisfy for this particular campaign, you face a choice of 10,000 bid requests. The question, and this is the question which technology solves for you, is how do you choose? Now every time there is a choice, there is actually a conflict. How does it play on the demand side? Well, there are three parties, or at least three parties which we can identify: there is advertiser, there is agency, and there is somebody doing the actual work, DSP or trading desk. Now let's put a very simple business model just for illustration purposes. Don't get me wrong, it's not the best model, I'm not really analyzing it or trying to defend it, it's an illustration. So advertiser pays a pound, not a dollar, a pound CPM to the agency. Agency takes 10 pence out of this pound and pays the trading desk to do this service, and they basically run after the advertiser budget which we discussed is a million impressions a day. Now what do these parties want? Again, don't get me wrong, I'm talking about a simplified case to illustrate my point. Let's start with agency. Well, agency in this model gets paid for a thousand impressions no matter what sort of impressions they are. One of their major costs is the cost of buying the media, so they need to pass some of this money over to the publisher. So it's in their self-interest to buy the cheapest possible impression. I wanted to add a few minutes of apologies about not being aggressive towards agencies, but fortunately at the previous panel the news has already been spilled; it does happen in real life, so I don't need to be particularly apologetic. But I still spend like 10 seconds: it's actually a very simplified model. If you shift the business model, then those conflicts shift, but they're still there, just more difficult to analyze.
Okay, so now what about the trading desk? Well, trading desk really doesn't care whether the impressions which they supply are expensive, come from expensive premium publisher or cheap, because they're not really paying for impressions. What they do care is that they want to do the least possible work to select. They face this selection issue: they have 10,000 bid requests, they need to pick up the one request where the right creative will be shown, they need to compute bids for those 10,000 requests, and they want to do it very efficiently because ultimately that's their main cost. In this model at least, their revenue is set in stone, it's 10 pence no matter what. What about advertiser? Well, advertiser in the game has this great panel: advertiser wants quality. They want this selection to be made properly in their interest; they want to pick up the best. Now my real point here is that the technology provider or technology controller, whoever controls the technology, really referees across these competing agendas. Not that they particularly want to referee; it's actually a very difficult, tedious job and everybody gets upset at the very end once they see what exactly you've done. But they don't have a choice. There are 10,000 opportunities, they know they must pick one, and if the system works it makes those choices many hundreds of thousands per second, and those choices essentially direct how those interests are balanced. So this guy really becomes a referee. So let's go one step deeper and let's try to see how the technology controller would shape the technology in order to satisfy either this end of the interest or this party or that party.
Okay, so let's start with the trading desk. So what if they ask us to build technology which is just in their short-term pure self-interest? How would it look like? Well, here is the strategy for them. The bidding strategy is actually very simple, I can explain it on a single slide. A bid request comes in, and all you do is toss a coin. It's a weighted coin; it produces only one success for every hundred tosses. When it does produce success, you actually do some work: you check targeting. If the targeting passes, you bid and you always bid the same one pound CPM, you don't hold your bid. What you do alter every few minutes is you go and check the reporting. You collect reporting from all your billing service and say, 'Okay, I'm tracking to budget.' If we seem to be over budget, then all you do is reduce the success rate. So now this coin toss gives you success once in every 200 tosses or once in every 300 tosses. You keep adjusting that to make sure you're not over budget. If you happen to be under budget, the same thing adjusted the other way around: you basically increase the success rate. With this very simple loop, you basically deliver exactly what's required. You do a lot of work just tossing the coin and nothing else. So that's a very efficient strategy; you can run it at scale, you can run it at 500,000 QPS on not so many computers, and it just delivers to the interests of a trading desk here.
Okay, let's look at the other example: how the bidding strategy would look if you want to buy the cheapest possible impression, if you want to make agency as happy as you can within the framework of my illustrative simple model. Okay, well that's media cost minimization technology. Again, relatively simple, I can put it on a single slide. Bid request comes in, you check targeting, no coin tossing this time. If the targeting passes, you bid something, and you bid X. This X can be say 40 cents, and then again every few minutes you check how you track into budget. Remember you have this one million impressions to satisfy in a day. If you're over budget, you just start bidding less: you bid 35 cents CPM, then 30 pounds CPM, and so on. If you're under budget, you increase, you just move it up. Again, relatively simple control loop, and you know, relatively straightforward strategy, no need for machine learning, big data, or anything like that, just works. Let's compare the two in terms of their technical details and technical cost, and also in terms of their results for the business. What's the technical outcome for the sampling strategy? To get a single impression, you do 10,000 coin tosses which are incredibly cheap, you do 100 targeting checks (here I assume that about 5% of your targeting checks actually pass, so five times you need to evaluate the bid). What about this strategy? Well, this strategy doesn't do coin tosses; it does 10,000 targeting checks. Targeting checks are way more expensive. So net effect is that the cost of running this strategy, the cost for the technology provider, is actually higher, and not just 10% higher, 50% higher, it's 100 times higher compared to the sampling technology. Now what's the business outcome of these two strategies? For the sampling strategy, because you bid all you've done, you sampled, you basically get average impression quality, average media cost, and average agency margin. Down here, you basically did a lot of work, you tried really hard, so now your impression quality is really low, the media cost is really low, and the agency margin is really high. Now the third party, remember in this triangle of competing priorities was the advertiser. So what would advertiser think looking at these business outcomes? And again, I don't need to invent this; we had the previous panel, they seem to be disappointed somehow.
Okay, so what do they want? How do you build technology just for the advertiser? Believe it or not, that's when it gets really complicated because the advertiser wants quality. Realistically for them, quality means that those ad impressions need to produce engagements, and those engagements could be clicks, could be post-click conversions, could be leads, purchases, but something very tangible for their end business. That's what they want. How do you do it? How do you execute this choice between 10,000 possible bid requests to select the single impression to maximize the engagement? Well, for that you need predictive analytics. You need to look at every opportunity and you need to predict what's the likelihood of this opportunity resulting in click, or what's the likelihood of this opportunity resulting in purchase. In order to do this predictive analytic, the only way to do it is really to look at your past historical data and do machine learning or some other type of data mining on your past data, and find out those probabilities. Then you basically out of 10,000 you pick up the one with the highest probability and you run with it. Now the balance or the key key technology question here is how much data do you use? Because there is a lot of data. Some of it is available for free: your friendly SSP will send you quite a bit of it, and some of it you need to pay for, and some of it is your first-party data because you have it as part of your business. Your job is to collect it, clean it up, and act on it. So how much do you use? If we take this limiting case of advertiser who really wants the best possible quality, totally insensitive to all the costs, well they would want to use all the data. And all the data realistically today means several terabytes of data which needs to be mined every day. Is it possible to do it? Well, sort of, and the IBM guys here will tell you more about it. It is possible, it's very expensive. Is it worth doing it? Well, again, I let them answer this question. From our perspective, it depends. Very often mining huge amounts of data costs you so much that whichever incremental lift you have almost doesn't pay for it. So let me give you a very, very rundown example of this choice.
Or just analyzing this very simple piece of data which is available pretty much on every single impression bid request: user agent strings. Those are the cryptic things which browsers send back to us, and they look like that. Actually this one is truncated; it probably runs up to here if I put all of it. You have a choice what to do with this. You have this data on every single bid request, you don't pay for it, but you do pay to process it. The simplest thing you can do is say, 'Okay, fine, let me look at it. It's actually you can tell it's Microsoft Explorer, so I know that the browser is Microsoft Explorer. Is it useful? Does it help me to predict whether this user clicks on this particular campaign? Well, it turns out it does. People who use Microsoft Explorer are different from people who use Chrome, so they would like different things, so that's quite predictive. And it's actually a very small part of data to use because realistically there will probably be half a dozen major browsers and that's it. So this field you will just use half a dozen different values. Okay, you can go deeper and you can actually discover it's Microsoft Explorer version 8. Does it help? Well it does, but you know, incremental value is lower because it's still Microsoft Explorer user, 8 or 7 or 9, not a huge difference. Okay, it's on Windows, it's on Windows XP. So all these pieces help, but each subsequent one helps less and less. Now if you look at it backwards and ask, 'What's the cost of machine learning on this data or this data?' Well, the difference is that there are only six different values here give or take, and there are probably a few hundred different values here, and just for the record there are about five million different strings together which you can see on every single day in the entire ecosystem. So generally what happens: incremental value goes down while the costs skyrocket. The interesting anomaly which I just can't resist the temptation to talk about is about this piece. It turns out somewhat surprisingly that if you're actually willing to do the data mining on full user agent string, all 200 characters of it, then suddenly it gives a lot of value, it's very predictive, actually sometimes more predictive than the browser itself. How does it come? Where it comes from? Well, that's all about traffic anomalies, another very, very important choice to make.
So far we covered two choices. One choice is about building strategy: how you bid, and depending on how you bid you benefit different parties who use your technology. The other choice was how much data to use: little data, more data, whatever. So that's another choice. What to do with traffic anomalies? Where they come from? Well, conceptually there are certain user agent strings today in the open ecosystem which generate 40% click rate. There are certain IP addresses out there today on major SSPs which generate 70% click rate. Now we call them traffic anomalies; our engineers tend to call them fraud, and that's a huge problem. Now why it's a huge problem? Conceptually, the impressions out there which are delivered from those IP addresses or delivered from those final user agent strings are a tiny fraction of impressions. But remember, we're trying to build something for advertiser, we're trying to find the most clicks, most engagement for his system. Any machine learning system will basically go and shift your media buy to those places which click very well. So in this particular case, it will turn to go to the small publishers who for whatever reason need to check whether their creative goes to correct landing pages five times a second. So unless you take this data and carefully move it away, your entire media buying, at least as far as the clicks are concerned, will deliver a very strange traffic which most probably is not even originating from humans. Now that's a threat to the ecosystem. We see it as an emergent threat. We see it's very important for a couple of reasons. One reason is that it actually costs a lot to find it. Remember, in order to identify it with some reliability, I need to do data mining on those full user agent strings, five million of them, I need to look at all IP addresses, I need to look at a lot of data which I would otherwise ignore. Another point why we think it's an important threat is that other forms of advertising in certain stages were swamped by people who are using this advertising in a sort of inappropriate way. Like we all remember email marketing used to be a great thing until spammers started faking return addresses, email addresses, and basically now a big part of the ecosystem is referred to as spamming. Something similar happened to search engines. We don't really see why the successful business model around retargeting should be immune to a similar kind of threat. There is nothing technically in the technology which stops somebody from taking the targeting signal. We don't understand why, I mean we don't see much of it, but we don't see why it shouldn't rise at a certain point. So anyway, threat or not threat, again we have our conflict parties. What do they want to do with this particular fraud issue? Well, start with advertiser. Advertiser absolutely wants to have nothing of it. I mean they definitely paid you good money for humans, they don't want to see any fraudulent clicks, they want it removed, they want it removed from the learning data, and that's their interest. Now agency, okay fine, I mean agencies sort of with the advertiser here, they don't like this, but on the other hand at the end of the day they send reports to the advertisers, those reports have CTRs in it, and the CTR is far better if you include a bit of this strange traffic, or at least if you're not that diligent in removing it. Finally the trading desk, that's really the party which takes most of the blow. They must work very, very hard to identify it, to remove it, to fix reporting. It's a lot of extra work for which they don't seem to be entirely appreciated by their immediate paymaster.
So just to go back to my example, I covered a few extreme cases of how you can tailor technology to satisfy different agendas. In real life, any successful technology controller needs to go and try to find the right balance, which is in itself a very difficult problem. And just for the record, we're not controlling the technology; it's the job for our clients. It's very tedious and not enviable job at all, and they keep tweaking the solutions to make sure that the balance is right. And we're actually very grateful that they take this load off us, because ultimately we're too far from the market to actually make it work, while they seem to be interested in taking this challenge and rising up to this challenge. I'll give you a slide just to try to indicate how it works in real life, how what they actually do to try to find this balance, and then I guess we'll be close to finish. So how do you find the balance? For starters, you can't find balance once and for all for your entire technology solution; you go and do it campaign by campaign by campaign. Again, just to simplify things quite a bit, let's assume that you have a bidding strategy with just two different parameters. One parameter is how much you pay, how much you bid: half a pound, a pound, a pound and a half. And here is the CTR floor, so you basically say, 'Look, if it seems to me that the click rate will be less than 0.05%, I'm not even bidding there, I just let it pass.' That way you control quality. So with these two parameters, there will be somewhere on this line a point where you just deliver to budget. If you bid less and currently we don't care about quality at all, we set the floor to zero, if you bid less you will be under budget, you wouldn't get this million impressions a day. If you bid more, you will be over budget. So there is a point somewhere here. Now you actually can bid more and simultaneously raise the floor, and because raising the floor restricts the amount of bid requests which are available for your bidding, that means that you compensate raising the price by raising the floor. There is an entire curve which we call budget curve where you can be to satisfy budget. Now on top of that, you draw something like margin curve. You say, 'Look, if you are here for this particular pair of floor and bid price, I am breaking even, so my publisher, sorry, advertiser pays me one pound and I'm spending one pound of media, so that's a break-even curve.' So net net over here, you have this line where you can deliver on budget and be profitable, as in spend less on media compared to what advertiser pays you. Now where do you want to be on this line? Well, if it's pure arbitrage, we just take as much margin as possible, so kill the floor, you need to be here. Now for this campaign, well this is a campaign for an advertiser with whom I have this trust relationship, who pays me for my services, so I really want, given the constraints, to deliver as high quality as possible, so I want to be here. And that's just the start of it. Because what if the budget is higher? What if they ask you to spend, to buy not a million impressions but 10 million impressions? In order to satisfy it with other conditions unchanged, you need to bid more, so this entire budget curve basically shifts to the right, and now you can't possibly deliver the budget and be profitable. So well, what do you want to do? You can say, 'Okay fine, I just deliver to my zero margin point and I will probably deliver only five million out of this 10 million.' That's one option, one strategy. Or I can say, 'Well, you know, I have a great relationship with advertiser, I'll go negative on my margin on this particular campaign and I'll be here.' Now that's a choice for technology control, and they need to execute it, they need to define it. And it actually gets only more complicated because on top of all that, sometimes advertisers set some quality targets, say more than 0.1% click rate, so there is another curve here which is about quality. Another complication is that so far we just changed bid and we changed the floor. You can actually add other variables to control your strategy, you can have three, four, five, so then all these curves become surfaces in multi-dimensional spaces. You just intersect them, and it's the client who needs to run through the analysis and to understand where on this complicated surfaces you want to end up. Now I appreciate it might be not easy to understand, and it's not because it doesn't exist, it's because I just don't have time and I'm not doing a very good job at explaining it. My point here is that it is sort of complicated work which our clients do basically every single moment every single day.
So here is the summary. There are conflicts in what you want to do. The technology controller, those conflicts are sort of resolved or refereed by a particular choice which whoever controls technology must make on behalf of their clients. And they make those choices in bidding strategies, in finding which data to use, in finding how you treat fraud, and in many, many other areas of technology which I didn't even discuss. Now it's not the providers of technology companies like Criteo that make those choices; it's the people who control technology. In our case, it's our clients who sit there and make those choices. Now those choices are very important. Being transparent about them, owning them, is something which affects the results of your business, affects your long-term marketing position, and so on. Critically, I mean, thank you for the time, and thanks Kieran for inviting me.