Raj Neervannan8:31
All right, cool. So my topic here is how to build a better search engine. As I mentioned, in 2006 through 2008 when I met my co-founder, the idea was how do we find information that is easier to do our work as an analyst. As I mentioned, he didn't have a great tool. During our work during an MBA, you must have been asked to do some research work, go find information for a particular essay. What you do is you go to Google and look for some information, and that's what business people do at work. We felt that it was not working so well. I'll go into that. So that was the topic. I'm going to explain why we did this and how we did that. The paradigm here is, as Martin mentioned, you want to discuss the problems, then what kind of innovation we did, and then what kind of solutions we propose and how we experimented. There is a slight change to this: the real company is that you don't just have an innovation and then a solution, you go through a lot of iterations in the process. Way too many iterations you go through, but you always get market feedback from real users. So with that as a backdrop, the main problem we were addressing was when we're trying to look for information, Google, when you put in a keyword, gives you top 10 results. It tells you 'I found these keywords found in this document' but they didn't really understand how often you mentioned something about a word. It doesn't always come up with other variants of the word. You may mean one thing, and the word you type may have multiple meanings - these are called synonyms. It doesn't really understand that. So that's one problem: it has a strict keyword limitation. Things have gotten better since the AI days, but this is 2008, 2009, so that wasn't so. Then what you get is top 10 links, you wonder what else is out there, and you don't have the patience to go through each one. That's the second problem. Third, you have to open each document and see what's inside, then come back again, so that slows you down. Fourth, what else has happened since I tried last time? That doesn't really flow into keeping you on a time scale. Something came up in the last two hours, something came up in the last day, it doesn't tell you any of that. Finally, as I mentioned, even if you open up a document, you have to go and find one item at a time. These problems don't seem obvious on the surface initially when you are using Google for fun, for simple projects. But when you are doing like simple things, Google kicks you out into a new zone. Most of the products that people used had used first search only did this. Nobody ever bothered to ask the question 'Could this be made better?' And we decided to ask because we felt there was a pain in the process. The first step is not just identify the problems you have, but you want to amplify the magnitude of the problem. How big is the problem really? How much time does it take? You want to magnify and quantify that problem in terms of the irritation you feel or the time it takes to accomplish something. That's step one. Then you want to come up with ways to solve and try it out. So going back to why - you have to go back before you propose something to innovate or solve, you've got to understand why certain things work. You can't just assume that people haven't thought of this; obviously people have ended up stuck in a certain place because of some reason. So you have to give credit to people who have been working on it. To that end, you understand how search engines work. Search engines work by crawling all the web pages, storing all of them, organizing them by some categories, and then ranking them by certain order. When you put in a query, it tries to match your query to the pages that seem to come closest to what you have, and then based on some popularity contest - it used to be called the PageRank algorithm by Sergey Brin who founded Google in the late 90s - based on how many people link back to it, they ranked based on that and some other factors, and then give you the top 10. It worked really well for internet consumer businesses, but if you go down into it, it doesn't work for businesses because the search index process is the culprit. The main problem is when it crawls and finds the links, it just visits each page and grabs the content and extracts the key important pieces of data here and there - sort of like you go to a mall, you don't remember everything about it, you just remember one or two things. It loses all the other context and all the other information you could possibly be interested in. So it loses vital information about all the semantics, all the information inside the page - just key things that it thinks matter to the world. Then it indexes on them. Index is nothing but a glossary of a book. If you go to the back of the book, you'll find that for the 100 pages you have in the first 100 pages, the last five pages are a glossary where it says for these keywords you can find these keywords in this page. You go back and look. That's exactly how the search index works. You just have to go back and find. Then they store that information, and when you ask a question, they just look it up and go back. As you can clearly tell, this is the one that scales because the internet scaled in trillions and trillions of documents. There's no way for them to index every single line and word in each page, so they have to use this glossary of a page to get to the keyword and the page you're interested in. That's it. That's the way a search engine worked. And that is great for casual, quick one-stop shop, looking for any information at a surface level. Then you are thrown into a link for you to dig deeper after that. This is fine for normal purposes, but if you're trying to look for business information - I'm going to go into the business side. Obviously, you are in school, you probably haven't heard of these terms, but business people do. When I say business people, they work in companies that are looking to understand about other companies, and they go look for - or even if you're in the stock market, if you are trading anything, you want to know what is the size of Apple the company. Well, how big is the market? You find the stock market size: 'Oh, it's a trillion dollar company.' That's not the same as the market. There could be multiple products inside Apple, each one has its own market. The iPhone market is different from the Apple Watch market. So each one has a market or a sub-market. This is called total addressable market or TAM. If you just put in 'TAM' in Google and it says 'TAM Apple', it just says things about Apple - 'you want some apple? TAM is 20th anniversary Macintosh' - it doesn't. That's not what TAM means. That's not what business people mean. So clearly, many of the search engines don't work for business people. So this has been the state for Google for a long time. It's been getting better, but this is a recent search, even now it doesn't work that well. So you think, how many of you have used ChatGPT? I can see you all, but I'm assuming most of you have heard of it. ChatGPT is this new wonderful innovation from OpenAI company which allows you to ask questions like this and it comes back and gives you a long essay answer. It will help you, it gives you an answer not just in terms of links but summarizes the answer for you in a way that is like I'm talking to you. There's a very human sense, a human-centric way. But you think that it would know more? Unfortunately, it doesn't. When I say 'TAM Apple', I'm not sure what you mean by TAM, and then I explained it: 'Okay, it's total addressable market for Apple.' Then it goes on to explain what total addressable market means. So I said 'Okay, what is the total addressable market?' Then it kind of repeats. As I mentioned earlier, even on the most sensible latest interface, it doesn't seem to quite address that. So I'm now going to jump straight into our solution. Then I'll go back to why because I want to give you the solutions quickly.
I'm now going to jump straight into our solution. Then I'll go back to why because I want to give you the solutions quickly. Then go if you put the same thing in AlphaSense, my product, 'Apple and TAM', it will find all variations of TAM, it understands what TAM means and actually give you the addressable market and actual size and give you the specific information coming from experts. It curates all of that to you and says 'I have found tons of documents, but you only need to read one particular line in the specific book.' It's kind of like asking a question to a librarian: 'Where can I find this information?' And then she tells you 'Go to this building, go to the shelf, you may find out.' That's not enough. You want to know the specific book, the specific page, and specific line. Then you may want to email the other information that may be in another shelf, another book, another line. So the real search answer should be like that, or a summarized answer would be just 'The answer is this.' Like an expert would. That's what search engines should be doing, but that's not what it's doing. So that's where the source of our idea was: we want to solve the problem in such a way that it is very specific and very clear. Why is this this way? What's the problem with search engines? Why didn't they do it? Normal search doesn't understand keyword values like 'R&D' means research and development. Some of you may know, and you would expect the machines to understand that 'R&D' actually means new capability, demonstrated potential, commercialization - a lot of different ways to describe the same thing. So all those variants are sort of distributed. So you kind of have an explosion of number of terms you should search for, not just one. You should actually search for more variants of the term. Search engines have a hard time ranking even for one term because they have to come up with the top 10 out of thousands of links. How can they find top 10 out of thousands of words at the same time? They clearly wouldn't. So it's not meant for them to solve. So if you go back into the counter argument for the problem, not only do they need to know the ranking, more importantly they need to know what are the words that look like the similar word but not quite. There are other words like 'internal research', 'research operations' that are not quite 'R&D' or variants of the term but not quite similar to the main term. So you ought to know the exact meaning. So you would think at this point, 'Hey, that's not very difficult. I'm just going to go to a dictionary and find out.' You can go to Wikipedia or something, you'll find out all the various meanings, and maybe you can get to the point. But it really requires an expert to understand what it actually means. The dictionary wouldn't understand the real use of the word. You'll use other terminologies. It's quite a complex problem. If you look at the word 'outlook', outlook may be described differently in a dictionary, whereas in the real world, 'outlook' means 'market share' in the business term. So the way we use terminologies is different from how a typical dictionary applies. That's the main problem. It is really a hard problem to solve. What I'm going to do now is get into how we now that we sort of understood the pain point of the problem and the core problems, and then how we try to solve them. What we decided to do was: what was clear was that the business terms were not understood by search engines. There was a clear takeaway. So we decided, 'Okay, we're going to but search engines are really good at finding a keyword once you give it to them.' So we said, 'Okay, we're going to then tag it. If there is a word like 'outlook', I'm just going to tag the word 'outlook' with a lot of different variants of the term in the book itself. It's almost like going back to a book and you find a word, you go to the page and you scribble on the side, 'Oh, this word means the following things.' Then you go re-index it and go back to the glossary and say, 'Now finally, where that is?' So now not only is the original book being tagged, but also the explanations that you have given tagged on top of it is also being re-indexed, so you can now find the actual meaning. In other words, if you don't know, I'm going to tell you what it is, then you're going to re-index and share me the results. So that's kind of one solution to this problem. The second thing was what you want to find is not exactly the keyword on the thing that you want to find; it could be around it somewhere nearby. That's called proximity. Proximity means close by. So not only do you want to find the word, you want to find related words. Not only do you want to find the related words, but also related words that are close by - not right next to it, but next to it - because this allows you to find the semantics a little better. That's the second part of the innovation. Third, you want to show more of this. People had a problem seeing all information in one shot, so we want to see more information but not overwhelm them. So we want to show them not just 10, we want to show more. Now I'll get to how to not overwhelm them, but first we wanted to make sure that at least we got to the key pieces of information and show as much as possible. These are the three main innovations we wanted to do. There are more, but I'll come to that. These three main ones: the way we tag the related words, the way we make the search engine sort of go back and re-index itself and find the new related terms, and also find the proximity close to it. These two dramatically increased scope. Our goal was to first increase the scope for all the missing ones. It was missing keywords before, it was missing the words around it before, it was missing other documents before. So we want to fix all of that. Now in terms of solving how to then present them better, we then decided to build a much better user interface, and I'll come to that in a second. So as I continue further, the first example of how this problem multiplies: when you look at 'revenue outlook for Apple' type example, if you go to Google and say 'Revenue outlook for Apple', Google comes back and gives you some examples. The real clear problem with Google is that it doesn't use the keyword hits, it doesn't use the actual location in the document. We have to go through each - I went through this before, I'm reiterating here - and you have to download each one, and you cannot generate a summary. There are a lot of problems with getting to an answer as opposed to finding some link. Putting in a keyword and finding some link is not an answer. That's kind of like, to my prior example, going to a restaurant, if you ask someone 'What's there to eat?' and the person just says 'Here is a long 40-page menu.' That's like 'Well, that's too many for me to look at.' Or then they ask you 'Well, I don't know, whatever you want to eat.' That's not an answer either. You want to be able to give them specifically 'What's a special today? What would be great over here?' There's a lot more information you can share and make this process a lot quicker. That's the problem with all search engines: they don't get you to an answer, they just give you even more information. So we want to avoid that. Good that we expanded the keywords, good that we expanded more documents, now we made the problem even worse in some ways, right? So we have to fix that. How did we fix that? Now comes the interface. This is the problem: now we find more important things we were missing before. Now how do we solve going back to what actually mattered? We said, 'You know what, we're going to find relevant pages and first you're going to make it easy for people to scan inside the document without necessarily opening each document.' So we're going to provide them a list of documents on the left side as opposed to what you found in Google which is over like this. As opposed to that, we're going to provide this on the left side. And Google didn't really do anything on the content inside the document; they just gave one line here, that's it. Instead, we said we're going to provide more insights inside each document, exactly what you should be reading in response to your query. If you type 'Apple' and you know what Apple means because it's a ticker symbol that you can find in the business, in any stock market, and you can find 'Revenue Outlook'. We have, if you look at it as a yellow line underneath it, it means that we understand what you mean. That's the yellow line should say 'AlphaSense understands what you mean. I got you covered.' And so it expands that to this type of a 'Revenue Outlook'. It expands this understanding: that's what it means. 'I get what you're saying.' That's what the yellow line means: 'I get what you're saying. Now I'm going to give you exactly what you want inside this document. I'm going to give you the document as well. I'm going to show you exactly where to look and save you time by clicking through. When you click on it, it takes you right inside the page.' By doing these things, we were able to not only show more information, pertinent information, rank them even better, and we are able to cut the time so people that used to take hours and hours now take minutes or seconds.
Now, this is great, but how do we possibly understand the semantics for all possible words? After all, there are thousands, millions of words out there, each one has multiple ways of saying that. How do you possibly teach a machine so many different variants? Well, it has been a difficult problem, clearly that's why it wasn't solved. But over time, in the last seven, eight years, it's been solved better with AI, artificial intelligence. We are able to make the machines not just respond to what you're trying to say with actual code, we are also able to make the machine understand the relationship between words and able to learn by itself. When you learn more math, you will appreciate the algorithms behind how the machine learns. So I would really urge you to pay more attention to calculus next time when you're paying attention to your equations and algebra next time. It really is there a reason why you learn algebra 1, algebra 2, calculus 1, calculus 2. It's going to get really difficult to understand how you can make a machine learn because the machine is just going to say 'You wanted an answer, why I got an answer? It's different from why I want to really get to that particular answer. How do I learn?' You're going to say 'Well, it's kind of like a blind man leading a blind man. You're going to say make a left, make a right, make a left, make a right.' That algorithm is all based on math. It's going to compute what is called the difference between the two words and then it's going to iterate through the whole process. I'm not going to go into the math level here, that would be another topic. But I want to just focus on the innovation portions at the high level, at the interface level, at the product level. Maybe another topic, another day we can look into how the math works. But the idea is that you want to use machines to understand words and their meanings and their related meanings, and that helps us solve the problem of coaching the machine manually one word at a time. That's really our problem. Now we solved it, but we can scale it. Scaling means making it work from - okay, we can teach the machine a thousand words, two thousand words. How can you teach the machine a million words? That's where the AI comes in. Second thing is ranking these words. Ranking is supposed to assign scores where we rank based on how it understood what they find inside, what kind of information is there, how relevant is it based on what's how much information was used to describe my problem, how much information was used, how much importance was given. If you read a newspaper, there's a headline, the thick bold mark, and then there's paragraph headings and stuff. They are using words in the paragraph for a reason because they know that it's important. So if you want to know what's important in the document, you will simply one approach would be to simply start reading headlines and paragraph headlines. It's the simplest way. Same thing here: if a word is important for a particular document, you look at the headlines and look at the paragraphs, first line of each paragraph. Things like that. You start translating what humans understand to be important back to the machine. That's how we started ranking each page on the importance of what it means. Then we looked at the specific hits like a search engine would do, then we rank them to its proximity to the important words lines. So there are a lot of thought processes that go on into figuring out what's important. What's important to us humans should be translated to what's important to computers if it really has to understand what you're trying to say. If you want to use what Google does, which is rank based on popularity - it used to be at least that way, after it obviously understands your content way much better - but then it was ranking based on just popularity, and it was not quite getting there. So those are the kind of - I'll leave you with the top approaches on how to rank, how to index, and how to collect this information, how to present this. So those sort of turned out to be the solution for our innovation big ten for the problems we're trying to address. And how does it actually map into a solution for end users? It makes them save a lot of time. As I said before, they can work faster and smarter. They don't have to control-F. If you use Windows, you use control-F to find something inside a document - many of you are probably using Mac even - for document it's control-F, a keyword combination you use. It's really annoying to have to go back one document at a time. But in our business where time is money, people are annoyed by this and they don't want to lose time. So once we solved the problem, we were able to surface meaningful results quicker, and they can spend more time for higher value tasks. Same thing with - I'm sorry - so it's not just time savings. What happens is that when you have more time, you're able to pay attention to more important things. That's the kind of important side benefit. Second is the missing information is a huge problem. What you don't know, you don't care for, and that can be a huge blind spot. It's like driving and not looking at your rearview mirror at all. You probably don't want to do that. Whenever you learn driving, you would look at all sides and behind constantly. So it's important to not miss critical information that's around you. Same thing with searches as well. You can't just do one search and accept it; you have to be critical about what you're missing. That's the second eye-opener. Third is when you're able to do things better, guess what, you are able to focus on higher order activities which is idea generation. You're not toiling over basic stuff, so you can actually think of what else you can do. That's idea generation. Third is when you're able to do this, you are able to pay attention to each other and make the next person smarter, and by sharing these ideas and collaborating, your productivity goes up as a team, not just as an individual. So these things have orders of magnitude productivity benefits. When you have productivity benefits starting with time savings, then you have typically a winning product because it's going to be sticky. Sticky means it's going to make you want to come back, like Facebook or Snapchat or whatever tools you guys use. It's sticky because you want to come back to listen to what your friends are saying. So it's kind of like that: it makes it better. So when you want to come back and look for it, so that was a story. And then even though in 2008 there was AI, wasn't quite...
As we evolved and built the product and proved the value, AI kicked in around 2014 time frame and was beginning to show its value. We pivoted to doing way more AI-centric search as opposed to statistical machine learning type approaches that were prevalent before then. And that is the foundation of most successful search companies out there, including Amazon, Google, Netflix, etc. So what does AI mean to us? As I mentioned before, finding words and related words would be very important. But AI gets better through interactions, especially through real-world interactions. So we want to apply to those problems that cannot be defined by if-then rules. Like trying to find out what's a cat and what's a dog: you can't just tell a cat is the one that has got furry hair and pointy ears; there are way too many descriptions you have to give, so it becomes very difficult. Image recognition is one of the problems that AI can solve, like thousands of problems. Pretty much everything that human does is a problem that AI can potentially solve. So that is the main sort of breakthrough.
The semantic problem, which means meaning, and making machines understand the meaning of a certain word is similar to making machines understand the meaning of a cat. So it's very similar to that. For us, AI is the way that we make machines understand the intent of a search query, meaning what it is you really mean when you put a word out there. When you put Revenue Outlook, as a business person, what do you mean by Revenue? What do you mean by Outlook? The machine has to understand, and it does understand in a way that doesn't require me to teach it one word at a time. That's a problem that AI can solve. Then you want to have other ways of using traditional filtering and whatnot, but the machines really help you understand the semantics. That's really the biggest problem that we have. So we use machine learning as I mentioned before. There is a specific field of machine learning and machine understanding called Natural Language Processing, because we are dealing with natural language. Of course we also understand more about images, charts, graphs, even videos. Business information is everywhere, so we want to make sure we get that right. The biggest reason we want to use this is to apply the same semantic search, making machines understand the meaning of words, to images or videos over time. Searching by intent is not just a keyword meaning searching; it also applies to the literal picture you see. You want to understand the context of a particular image, be able to say this person is dancing on the beach. So you want to understand the intent. This idea of semantic intent translates beyond words, but it all starts with natural language for us.
So that's one thing you have to understand, that these things are constantly evolving. Second is the typical entity recognition. When someone talks about Apple, we don't know whether they are asking about an apple to eat or Apple the company. Or when someone tells you don't keep slacking, in a business context, are you slacking? It could be an insult, it could be a bad thing. Slacking could now mean you're being productive communicating with your colleague, whereas 10 years back if you said you're slacking, that meant you're goofing off. These days if you're slacking it's actually a good thing. Being able to know the difference between slack the verb and Slack the company is important. Words that have become verbs have changed meaning from negative to positive. It's really important to keep the machines up to speed on the terms we use in the real world. This is one example of why it helps to understand what companies are coming up, how to understand products and services as we understand the real world. If the machine doesn't understand those words, it wouldn't be able to help you. That goes to relevance. I'm not going to go into all the details here, but essentially all this information feeds into the model. I don't know if I'm going too fast; I want to be conscious of time. It's 9:40, so I can go a little quicker to get through the Q&A in the next six to seven minutes.