Back
Jason Mars
Cofounder, Clinc

Why conversational AI is taking over our world | Jason Mars | TEDxUofM

🎥 Apr 09, 2020 📺 TEDxTalks ⏱ 16m
Dr. Jason Mars is co-founder of Clinc and professor of Computer Science at the University of Michigan. He has showcased his talents in the natural processing field by becoming one of the leading entrepreneurs in conversational AI. His mission is to help solve complex, real-world issues through AI that learns on the go, empowering enterprises to use revolutionary AI to improve experience for their customers. Jason not only has an impact locally but is recognized globally as a top innovator. He continues to impact the technology industry and academia by pushing the boundaries of AI as we know i...
Watch on YouTube

About Jason Mars

Jason Mars, a computer science professor at the University of Michigan and co-founder of Clinc, has discussed his work in conversational AI and his evolving motivations. In a 2021 interview, Mars stated that he is no longer motivated by money or success, but by "creating something beautiful that helps people." He described Clinc as a "rocket ship" that raised funding from $250,000 to $53 million and reached tens of millions of users, beating incumbents like Google. Mars also noted that he has been building a new AI computational paradigm called Jaseki, with a programming language named Jack, and three products built on it that are live in invite-only access. In a 2020 TEDx talk, Mars argued that conversational AI should aspire to "human and a rule‑level understanding" rather than competing with existing products from Google, Amazon, and Apple. He described a deep learning approach trained entirely on data, comparing it to how a four‑year‑old learns language without knowing grammatical rules. Mars highlighted a system launched to six million users that saw increasing engagement over time, and identified the next challenge as creating systems that "learn with you as you interact with them." He also mentioned that his academic project Lucida was originally called Sirius, but was renamed after Apple raised concerns about the name being too close to Siri.

Source: AI-verified profile updated from Jason Mars's recent appearances. Browse all interviews →

Transcript (5 segments)
J
Jason Mars0:12
Happening in the future, but to really think about the future, you have to think about what innovation means, what technology means for us as a society, as a humanity. We have to reflect on what progress looks like. When you think about it, as a species, we started innovating from the moment we were us, right? We started to develop tools, and with those tools, we developed fire. If you think about the tenants of innovation, it's things that help us in the convenience of our lives, to help us do more, to help us go farther distances by exerting less energy than what we ran on, anatomically capable of doing. Electricity then came in; we can have light so we can read, we can work and be more productive for longer hours. When you look at this timeline, it's hundreds of thousands of years ago we cultivated fire, the wheel similar kind of time frame, and then electricity is in the hundreds, not very long ago. If you think about the computer, which allowed us to capture knowledge and then access that knowledge in a little device that carries energy with itself in our pocket, innovation happens faster and faster, enhancing our ability to do more in the world and do more with ourselves. So we envision, what does the next key major innovation look like? This is a scene from a movie called Her. We often romanticize the future and then realize it; we have this tendency to realize the future as we imagine. This is a future that was imagined very recently, and in the same pace of picking up, right now you might think that they're kind of silly, but it is our imagination and we will get there.
One important question is why now? Why is conversational AI a thing now? I would argue that when we think of intelligence, when we think of AI, we think of the fundamental sociological and anthropological meaning of intelligence. Before we even had fire, when you think of intelligence, you think of your neighbor who has some special intelligence on how to build that tool, that fork or that sphere. You go to them and you interact; what naturally evolves is conversation. 'Hey, you build the best spheres, how do you do that?' Today, well, why is it happening now? I'll make an argument that manifests itself in recent Turing Award winners in just the last few years. If you think about three key pillars that are making it possible today, we have the three Turing Awards of reason that have won the Nobel Prize for computer science. The first Turing Award winner I'm going to talk about is Michael Stonebraker. He won the Turing Award for databases in 2014. We have this phenomenal capability of capturing data in the world and in our lives and storing it in computers. He pioneered what compute would look like in our generation. That compute has gotten smaller and tinier; it's in our pocket and does phenomenal things. People are playing sophisticated 3D games on their cell phone for seven hours. The kinds of innovations that came from John and David allowed us to do more with compute, data, and compute. Lastly, just recently in 2018, it seems like yesterday but it's already 2020, Geoffrey Hinton is one of the key winners of the innovations as it relates to models that are used to capture data and innovations in all those. Without any of the other pillars, we would not be able to do interesting AI. The models themselves require data to gain knowledge because that's all they learn from is data. These kinds of machine learning, deep learning, recurrent neural networks, deep neural networks, without the compute to train those models to be able to do the processing to embed that model with the knowledge it needs from the data, we wouldn't be able to have it. So we needed all three pillars to come together at the same time, represented in the three recent Turing Award winners.
Now my work has been around what I believe to be one of the transformational technologies: conversational AI. It started with a company that I started in 2015 with some bright colleagues and grad students, and we've grown into great success. We'll talk a little bit about that success. Next, I'm moving on to a new endeavor that you'll all learn about in the coming months, maybe a year, and it's quite exciting about the next generation of technology. But let's understand why we're having that revolution today and what it looks like. On this slide, you see on the left, that's actually the old way things were done. Natural language processing, the foundation of conversational AI from a pedagogical standpoint, has traditionally been based on computational linguistics and linguists in general, based on ways that we represent language. The noun and the verb from the synonyms that I know about, I can infer that. The problem with this technique is you have to enumerate those rules, you have to enumerate the grammars, the templates as to what you can say to a system. That's why you have so many silly systems today. You say something outside of that and it breaks because no one anticipated that rule. The beauty of a deep learning approach that's trained entirely on data, an approach that doesn't know nouns and adjectives, just like your four-year-old at home doesn't know what a noun is but can speak complex sentences to you and understand you. Neural networks have a capacity to learn messy things and to infer from noisy information. So when you take a sentence like this, if you have a system for which it must have a rule to match to that, if you have to have the synonyms and dictionaries that have gobbledygook in it, if you were required to have a particular grammatical structure, this might break a system. But a system trained with deep learning and only data, by learning from the experiences of listening to how folks are talking in the world, might be able to do better, build systems that don't know what nouns are, and you may achieve the kind of AI that we use, the kind of intelligence that we use in our biological brains. In a sense like this, you might ask someone what's gobbledygook in the sentence. Anyone in this room would be able to say instantly that gobbledygook must be a place we spend money at, it must be a merchant. If you say how many verbs came before that, you've seen it before, and you may not have seen this particular combination of things before, but no matter what comes after it, you could infer that's probably a place you could spend money, just given the patterns. That's what we embed in these models. This uses a recurrent neural network that is an LSTM, and you can actually have it read the sentence backwards and forwards using a bidirectional LSTM to use future contexts in previous contexts of the patterns, the sequence of things that recognize that resembles something you've seen before to infer what that is. Train your systems with that, and actually you can do that for anything in the sentence. Gobbledygook is a merchant. On smacky wacky, that's something you could spend money on. The flippity-flop is overdrawn. Given my knowledge, that might be a nickname for an account.
I'm sixty-six dollars and seventy-seven cents, and you have five thousand five hundred twenty-five dollars and sixty-four cents outstanding on your statement balance, which needs to be paid before Tuesday, October 8, 2019. Cool, let me make a payment for the minimum amount, please. Yeah sure, let's do it for that Delta flight where I went to California to visit my brother. That was pretty pricey. Please choose between the three available payment plans. Let's do it for the 12-month one, give me a little time. For your one thousand two hundred thirty-five point five six transaction at Delta, would you like to confirm this payment plan? Actually, I changed my mind. Can we do it for that hospital bill where I went to get stitches? That was extremely expensive. I should stop playing hockey. Also, can you give me about 18 months to pay that bad boy off? Yes, please.
So that is the level of human expectation that we would want conversational AI to understand. We can realize that there are other open problems. There's multiple distinct problems that need to be solved in tandem to create an experience like that. Take a deep dive and have a look at some of the work that's coming next. But what does this mean in our lives? We expect this is what happened with Alexa and Google Assistant, Siri, and all that. Everyone's excited when it launches, and then nobody talks to it after a while. The engagement goes down. Everyone bought it and now no one talks about it. Well, what we saw when we deployed a system in your life, improving your capabilities live now, four million out of 6.5 million people are interacting with it regularly. This is a native Turkish system in Turkey. It turns out that when you just train with data, it doesn't matter what language it's trained in; it's agnostic to language. It learns like a four-year-old. What was even more impressive is that the average user of those four million uses it more and more over time. It's creating value in their lives; they're returning to it. That's the goal. What I would argue, and I think what is required for real innovation, is you want to aspire to human-level understanding. Challenge the work you do to be work where you can interact with it freely without having an instruction manual as to what you can say to it. That's how I've driven all of the work that I've done in this space. That's why we were able to do something that's much closer to that ultimate goal. I will continue this work. I would say the next really important challenge is how can these systems learn with you as you interact with it. How do you take some of that neural modeling-inspired insights and then create a system that, as it interacts with you, gets better and better?