David Stephenson0:10
Take it home with me, but thanks for everyone who's come, especially thanks to the woman in the front who helped me find the coffee bar earlier this morning. That was tremendously helpful after my morning flight from Amsterdam. It's good to be here. My name is David Stephenson, and what the talk is about is how we've built a real-time streaming recommendation engine for an e-commerce site with Neo4j and incorporating some streaming properties. We did it in two weeks with four people. I'm going to talk about how we did that, and most importantly, what I want to do in this talk is not just talk about the fact that Neo4j helps with recommendation engines—we all know that. Peter talked about that earlier this afternoon, and it's a common use case for Neo4j. So I'm going to talk a little bit about that, but a little bit more about being data-driven and as analytically fueled as we could, maximizing the resources we had available as quickly as possible and in an agile way. A big theme of the talk is really how Neo4j fit into that bigger framework of quickly bringing analytic value to a company.
To set the framework for that, I should introduce myself. My name is David Stephenson. I'm from the United States, have been living in the Netherlands for about 10 years. I did my PhD in mathematics and computer science in the US, and I've been working with a variety of different companies—a lot of finance and insurance companies consulting, more recently at eBay where I had a global role in analytics. I left eBay a couple of years ago and I've been doing independent consulting. The reason I'm here is to talk about the data initiative for us. Let me introduce the company to you. The company is Belvilla. It's well known in the Netherlands and on the continent, less so in the UK, but for those of you from other countries, it's a very familiar name. It's a company which has been around for over 30 years. It's a very traditional company which, of course, entered the internet age and had a website online and such, and wasn't doing a whole lot with data and analytics until about two years ago when the majority share of Belvilla was acquired by Axel Springer, the large German publisher. Part of the initiative for Axel Springer was to grow through acquisitions but also organic growth through better use of data and analytics. They made a strategic decision to say, 'We haven't been doing much with data and analytics, so let's bring someone in.' My role was to come in, help them identify ways to utilize technology, and help build up a team. That's important to keep in mind when we talk about the solution, why we chose the path we did, and in terms of the timeframe.
Belvilla is one of the largest online portals for vacation rentals—holiday homes, vacation parks, that type of thing. You can think of it as sort of halfway between Airbnb and Hotels.com. We've got about half a million properties which are spread out in almost 40 different countries, so it has a pretty big footprint. In terms of guests who use Belvilla every year, about 1.3 million people from 140 different countries booked their holidays through Belvilla. It's pretty extensive, and you know that your product is very diverse—everything is a little bit different, with the exception of some holiday parks where there are multiple versions. That added a challenge to the recommender engine. Taking a step back, you come into the company and they say, 'Hey, we want to use data. You've got an open slate. We have a budget for it. What should we be doing with data and analytics?' So essentially what you start to do is generate focus areas. What's important to your business? What are your pain points? What are your strategic ambitions? Where should we focus on with data analytics? You develop focus areas and then you start to say, 'Explain to me in business terms what you want.' They'll say something like, 'We want more effective marketing,' or 'segmentation,' 'marketing ROI,' 'conversion rate,' whatever. Those are the sort of analytic terms you won't necessarily use immediately with the business. Once you've done that, you're looking at your data and then you're looking at technology and your staffing, and that's where things get much more challenging because then you have to look at the organization. You have to say, 'What technology is in place? What is the appetite for change? What staffing is available? Should I use externals? If so, who should I use?' That's where the challenge really was in the situation, and that's a little bit more of what I want to focus on. But in the context of the Neo4j recommender engine, it was, 'Hey, we've got a clean slate for our company. Lots of options. What's the best solution?'
I should have done this in Neo4j, but this is actually what I used—except it wasn't actually in Cyrillic. Those of you in the front row, I had to randomize things for privacy purposes, so bear with me. But what this is meant to show is that as you're looking at applications, you're looking at milestones which span a long time period and have interdependencies. Part of this Cyrillic chart you're seeing has to do with customer journey, and part of the customer journey milestones have to do with recommender systems. We laid this out in the first month I was there. Someone's going to take a screenshot of this and go home and do an algorithm and translate it and give me trouble, but that's fine. Now, if you look at the Belvilla website, I want to motivate the recommender engine problem because it was high on the priority list. Why is that? We're looking at vacation rentals. It's not Airbnb where you're going away for a weekend and you need a room ten times a year. These are typically one, two, three-week bookings which are major holiday events. Most of your customers are coming maybe once a year to book. You've got them once a year. Likely you don't have any cookies; they're not logging in when they're browsing, even the repeat visitors. What's more, there's such an abundance of comparable sites that if they come to your site and don't find something quickly enough, they can go to a competitor. What happens? They come in via a landing page essentially. We know three things about them: they're not logged in, they haven't visited in probably nine months. The only thing we know about them when they land is we know where they want to go in terms of a very large region like southern France, we know when they want to go (arrival date and duration), and we know the occupancy they're asking for. We know three things about this anonymous person, and now we've got to win that person's trust before they go to a competitor. They're only going to come once a year. So what happens? They hit this orange button 'Zoeken,' which is Dutch for 'search.' For this anonymous person, I've got to get their interest and eventually have them book. But among all these 1,191 properties that Elasticsearch returns to me, I've got to find the one which is most interesting to this client who's not logged in and only visits once a year. You can start to understand the importance of the recommendation and you can start to get an idea of what information I have to work with. The only information I'm really going to have to work with now is the next few clicks that my user puts in. I know three things so far. Now I'm giving them a filter list on the left—these are more query filters they're going to put in. I'm going to start finding out more about them and I'm going to see what they do on the website. I'm starting to think, 'Hey, I need a recommender engine.' The data available to me, the primary data, is going to be the browsing data, and I need to get it right away and start working with it right away because if I don't—if I wait for them to log in, if I wait for an overnight batch job—they may not be back and they'll be at a competitor.
We started looking at how well we were doing so far. We started getting hit-level data from several weeks back and analyzing what search results we had returned and how the conversion rate had been on them. For something like this with infrequent, high-ticket purchases, essentially what you've got is very few conversion events. They're not purchasing groceries where they're coming back all the time; they're making one purchase a year. We started looking at those and said, 'Okay, here are some pictures. Let's see how well they did.' How many of these pictures that we returned in search results—we choose them as the six out of a thousand—how many of them actually got clicked to book? Just click for more information. Look at this property in the middle. This one stood out to us. It looks like kind of a nice property from southern France, but it's in the Alps. We looked at it, and in the time period that we looked at, this one had been returned nearly 2,000 times in search queries. It had been selected about 2,000 times over the previous few weeks. We started seeing all the times we've returned it and how many times it actually got clicked. As it happened, do you know how many times it got clicked? Zero times. So we really took that to mean something needs to change in how we're doing the recommender engine. Probably this is actually translating into a very large amount of lost revenue. So we put it under the customer journey stream; it was one of the things that went in. From a business perspective, the business goal was for guests to quickly find properties that appeal to them. If you want to make it really flowery, you say 'guests fall in love with our properties,' whatever. Then we translate that into—me as an analyst, I say, 'Okay, that's a recommender engine. What to do it?' I know the data that's available: we've got the features, the booking history, the clickstream. I probably should go into more detail about that, but I don't have a whole lot of time.
There was me, who was doing everything essentially. I was building up the program, hiring people, doing strategy, doing roadmapping, and when I had some time, I was actually doing Python coding. We were hiring, so we had to figure out how can we get a recommender engine. We started thinking of the technologies and the staffing. Over the course of several months, I talked to a lot of different vendors. I talked to software providers and I talked to service providers and considered different options for that. What we finally landed on was the Neo4j system. We knew Neo4j had a big use case for recommender engines. For me, I've got a pretty extensive background in terms of network flows and modeling, and I knew that a lot of business problems could also be modeled that way. There's his famous beer graph—you guys know Rick's beer graph? How many of you have seen Rick's beer graph? That of course is enough to convince anyone. So we decided we wanted to go with Neo4j. Then the question was, 'Service provider—who's going to do this?' Having spent a long time myself learning a lot of languages and such, I realized that if you're not really sharp with a language, it's going to take you about ten times as long. So we did quite a bit of interviewing people about that. A lot of people were eager to get involved with a Neo4j project. For a lot of service providers, it's a really sexy project and they want to build up experience. Most of them actually don't have experience when you probe them—experience in terms of what they'd built with recommenders and with some of the plugins and such. That's kind of a key part of it. The fact that we were able to get that in two weeks was because of the choice of service providers. I'm sure that if we had chosen some of the other people I talked to, they would have done their best, but I'm also certain we wouldn't have gotten things done nearly as quickly. So that's how we ended up going with the Neo4j solution and going with GraphAware. Then what we did is we said, 'Okay, we want to keep this agile. We want to think about how it's going to work and do something basic first, and then we'll scale it up later.' That's the 'nail it and scale it' philosophy for this.
In terms of how we're thinking of setting up the graph, essentially we have dozens of feature things like 'does it have a dishwasher or not?' But obviously, 'does it have a dishwasher or not' is much less important than something like 'is it in Italy or is it in Norway?' 'It doesn't take pets' is much more important than 'does it have a fireplace?' That's what we wanted to try and understand and put the proper weights on. All those features from the properties we could feed into the nodes to get the content-based recommendations. But what was also really interesting for us—because it was so key for us to get the clickstream data in, because we're depending on it—we would put the clickstream data in as nodes and arcs in order to build a collaborative filtering part. So that's sort of how the model was built. Two weeks, but you know, it's a bit experimental. I'm like, 'Okay, fine. I know how these things work. I've been around a long time, and everything is always a bit of a risk, so we'll just take it like that.' The end goal, of course, was that we wanted to be able to have these properties appearing offline, sort of as a little iframe or something in the search result page. That was sort of the initial goal, as opposed to interacting with Elasticsearch or even streaming real-time, which was our eventual goal. This is what we're starting with. We had Google Analytics already in place, and we had just gone with Google Analytics Premium. Google Analytics Premium, if you don't know what that does, will give you the hit-level data and you can have that in place already when you're looking at streaming data. When you're looking at this sort of fast analytics, typically what happens if you want client-side data is you put on some sort of JavaScript or a pixel or something that is going to send information immediately to something like Kafka or RabbitMQ—some kind of message queue—immediately to your server. Your server will then process it in several milliseconds, and then you'll do what you want with it. A lot of companies will sell you that JavaScript code or pixel code; you can implement it, they'll put it in HDFS for you or whatever. That's kind of a common solution. Whether or not you end up doing anything with that data is another question. I think a lot of companies are building up data lakes with such client-side data but not necessarily using it. What I'm talking about is we had this Google Analytics, and the data is going straight at the hit level to the Google server. Maybe we can bifurcate that call and essentially send it not only to Google but send it straight to Neo4j. That was a bit of an epiphany, and it turns out it's not hard to do. It's not hard to do that at all. If you have Google Analytics, you just overwrite what's called the 'send hit task' function, and it'll send your hit-level data. That's essentially everything you've already tagged in this case in Google Analytics. All the work has gone into it, all the business logic is in it, and then it will at the same time send it straight to whatever endpoint you have. In this case, we said let's go with RabbitMQ. It could have just as easily been Kafka or something. What hadn't been scoped as a real-time streaming project very quickly—within that two weeks—we turned it into this streaming real-time project. So then what we ended up with is we had the Neo4j cluster which we were feeding from the data warehouse: all the property features and the purchases, customer attributes and what they purchased. We were feeding it from there. We could also feed it from BigQuery in terms of session-level data, but that was with a delay. That would come back to the web server in the form of recommendations. But what we were able to do then with this bifurcation of the GA script was to send that through to Rabbit and ingest that directly into Neo4j. These things of course you always have to tune and do a lot of further analysis, but really with two weeks of effort—essentially it was these two guys, me putting in three-quarter time, and a web guy who was maybe half a resource from the web side—we were up and running with this, which I felt was pretty amazing. What we realized too is that the recommender was additionally planned to be in an iframe, then we could also very quickly incorporate it into the Elasticsearch. There's a lot of that—that's a whole huge issue. If you saw the Airbnb presentation, they also talk about the plugin for Elasticsearch. So we started doing that, which is even more value because now you're influencing your search results. Whatever—customer service would be left saying, 'Okay, we're not sure how to rebook you, what property to put you in.' They could then hook into the recommender system using the REST API and say, 'Here's your customer ID. Based on the recommender engine, here are the ten properties that seem most appealing to you.' When you go to marketing direct mails, obviously that's also very useful. We found that having the recommender engine had a lot of additional possibilities to it, which is quite nice.
Just to give you an idea of how much faster this moved than I expected, this was sort of an update slide that I presented after—I'm not sure how far within eight months or something—where we put a little note that we were already in the alpha stage of that by the end of that first two weeks. That was really cool to be able to present that and say, 'Look, we thought it would take this long to get there, but it's actually only taken us much less time to get the initial results.' Not a full-blown optimized version, but showing a proof of concept and ready to be improved and developed further. That was actually really cool to be able to do that. So that's my story. Feel free to contact me if you want, or catch me afterwards in one of these—what are they called? The little Swedish things with the coffee breaks. I'll be around.