Jeffery Hawkins0:01
Thank you and thanks for inviting me once again. I don't know what I can do to morat so he doesn't invite me again but I keep coming back. Many of you in the audience probably know a little bit about Numenta but some of you may not know about Numenta. So briefly, Numenta is a small team of about 15 scientists and engineers here in Northern California, and we are focused mostly on theory, information theoretic principles of how the cortex works. We also believe that those theories will inform machine intelligence. My talk today is a bit aspirational. I'm going to do less on the details of what we do in our modeling and talk more about the big picture, and make some suggestions about how we ought to be thinking about our work as a collective. As we all know, it's funny: 10 years ago when I started Numenta, the term AI was a really negative term—no one wanted to use that term for anything—and now it's become very much in vogue again, as well as machine learning and machine intelligence. And the progress it's been making recently is really great. But there is some fear that we may be getting off track. So what I want to talk about today is what should we be shooting for? Not for a year from now or five years from now, but let's say 30 years from now or 100 years from now. What should be the legacy we should be working towards? It doesn't mean it'll take that long; it just that we ought to be thinking about what we ultimately want to get to. It may only take five years, but we ought to be thinking about the long term. There are some reasons to be worried about this because the field of machine learning and deep learning is a bit hyped these days. I became aware of this; it really struck me once there was a recent headline that said something along the lines of Microsoft investing heavily in machine intelligence. The sub-headline was they agreed to acquire Swype, the predictive keyboard company. And that was a big investment in machine intelligence. I don't think that's what machine intelligence is. So what I wanted to do today is talk about what intelligence is. I want to take it from a perspective of biology and see what we can constrain our view about intelligence, and then ask ourselves: what should we be doing when we're trying to build intelligent machines? That's the title of my talk: 'What Is Intelligence That a Machine Might Have?' The talk is in three sections. The bulk of it is this first section, which is basically going through the biological components of intelligence. We all agree—in fact, the only thing we all agree on—is that intelligent is a human nervous system. So outside of that, there's disagreement. We're going to focus on that and say what does that tell us what intelligence is? Then we can talk about what are the functional components of intelligence that we can pick up from that. Take yourself outside of the biology: what are the functional components? And then finally I'll talk a bit about the diversity of intelligent machines—what might they look like in the future?
Let's talk about the first one. I hope you can see this is a cartoon drawing of a nervous system. I'm going to use a lot of cartoon drawings here—trust me, I know the complexity of the nervous system, but it's not worth putting it up in pictures here. This is showing what a reptile brain might look like. The brain and the nervous system evolved hierarchically. We started with a spinal cord, which has sensory inputs and reflex behaviors. Then on top of that, eventually evolved the brain stem, which is mostly autonomic behaviors like blood pressure and gustatory functions. Eventually we added what we might call midbrain structures in human—that would be like cerebellum, basal ganglia—and these have basic emotions and behaviors, and there can be learned behaviors. Then at the top of the reptile's brain there's something called the pallium, which is very similar to the hippocampus in mammals. It's a very fast memory of where the animal's been so it can recognize where it's been before. A reptile is pretty creative—think about an alligator, for example—it has a lot of behaviors, abilities to rear its children, territorial fighting, eating, sex. But what came along is in mammals we added one more component. Basically all this was preserved, and in mammals we added the neocortex. The neocortex is actually sandwiched between the hippocampus logically, and it represents about 75% of the volume of a human brain—not 75% of the cells, but 75% of the volume. It's pretty expensive. In humans, it's so big that we're the only species who regularly die in childbirth. These days, that's because we have such a big head it doesn't fit through the birth canal. We also take our offspring—they can't even walk for about a year, can't really do anything on their own for about four or five years, and it takes about 18 years for them to become fully mature. These are pretty expensive things, so there ought to be a good reason for having a neocortex. We should ask what it is we're getting over the reptile. It better be pretty good. If I wanted to try to summarize this in one word or one small phrase, I would say the following: the neocortex learns a spatial model—excuse me—a sensory-motor model of the world. What do I mean by sensory-motor model of the world? It basically learned how the world is—the structure of the world—and it's a sensory-motor model because mostly what it learns is how the world behaves when you act upon it. Remember, the world to the brain is just a bunch of patterns coming in on the sensory organs. What you perceive is the model in your brain; it's not the world—the brain doesn't actually deal with that. The brain has to construct that. So it says: when I act, what are the patterns that come in? And when I act again, what are the patterns that come in? Through that interaction we build this model of the world which tells us how everything works—from computers to doors to food and cars and all the things we do every day. We have a pretty sophisticated model of the world, and that's what really makes us tick.
Let's jump into that and look at it in more detail. If you look at the neocortex, everyone knows this—most people do—this is of course the classic Felleman and Van Essen diagram from 1991. The cortex is a sheet of cells divided into regions, and those regions are connected together in a hierarchy. This is the macaque monkey hierarchy: on the left you see the somatosensory hierarchy, on the right the visual hierarchy. The little rectangles—if you're not familiar with this drawing—are the cortical regions, and all the lines are the interconnections, massive interconnections between those regions. Each one of those lines represents millions of nerve fibers going both ways, forming a hierarchical representation of the cortical regions. Information in this diagram comes in at the bottom and flows up the hierarchy and back down. You see that there are different hierarchies from the different modalities that are connected at the top. The first thing we can say is this looks awfully complicated—how are we ever going to figure this out? Well, it's a hierarchy of regions, but as you heard, the regions are remarkably preserved. They're almost identical everywhere. They're not identical; there are differences—some are noted, some are very subtle—but they're remarkably conserved. The regions have a lot of detail, and everywhere you look that detail exists. The basic conclusion is all the regions are doing something very similar. There's so much evidence for this I don't want to debate it, although there are people who would like to debate it. We're not going to debate that here today. The second thing we can argue is that the hierarchy itself varies significantly across species. So if I look at the visual hierarchy of a monkey vs. a human vs. a cat and a dog, they're all quite different. There's nothing particular about this specific graph. It turns out you can take these cortical regions and hook them in different ways and they generally work pretty well. We know from sensory substitution you can put the wrong type of information into one end of the hierarchy and it still works. So we don't have to worry about the complexity of that Felleman and Van Essen diagram. What we need to worry about is: what is each cortical region doing? How do they work in a hierarchy? Then later you can figure out what hierarchy you want to build.
If we look at what each cortical region looks like, the first thing you see is it's structured as layers. Everywhere you see layers of cells—typically six, depending on who's counting, but that's the typical number people use. They're very clearly there; no doubt about this. The second part of the organization is that there are these mini-columns. The individual excitatory neurons are arranged in these mini-columns. They are physically true—you can't always see them, but they're part of the development of the brain. There's a lot of debate about whether they're functional. A mini-column is really skinny, only about 50 microns wide, with about 100 to 120 cells in it. It's a very skinny column of cells as an organizing principle in the neocortex. There are a few things that some people forget about these cortical regions. The first is the input. Everyone knows sensory input comes into the primary visual cortex or the primary auditory cortex, and then it gets passed to the next one and so on. But you need to remember that that's not all—that's only half of what a region gets. The region also gets a copy of motor commands that are being executed by the rest of the brain, either subcortical or other parts—it knows what behavior is being produced by somebody else at this point in time. This is essential because if you didn't have this and the input changes on your retina, you'd think the world was moving. But you don't think the world is moving when you move your eyes; it seems stable. The only way that can happen is that every region of cortex—and this is true for every region in the brain that we know of—is getting both sensory input and copies of motor commands. This is essential for how it understands the world, because most of the changes on the sensory organs are because of your movement—not all, but most. Therefore, it's inferring both sensory data and motor behavior. That's a common principle throughout the cortex.
The second thing is that every region we know of in the cortex has layer five cells that project out of the cortex and generate motor behavior. Everyone—even V1—has cells that project down to the superior colliculus and impact eye movement. So this is a universal property. That same output gets split in two and goes up to the next region in the hierarchy. So the next region is getting a copy of the motor behavior that this region is generating. This is a general principle throughout the neocortex. Every region of the cortex also has cells in layer six which project back to the thalamus and are believed to be involved in attention. There's a lot more to this, but I want you to make some observations. First, every region is recognizing sensory sequences. Every region is getting the stream of data coming in, and just like my speech right now or music or a bird flying across, it's time-based sequences. One type of sequence is pure sensory sequence like my speech right now—you have to infer it by hearing it through time; it's an inference over time, not a spatial inference. The second thing is that every region has to recognize sensory-motor sequences. When you feel something, if I put my hand down and touch something, I know that because I know how I move my hand. The inference is both a combination of my motor behavior and what I'm sensing. Every region generates motor sequences, so it builds a model of sensory data, sensory-motor data, and generates output. You might argue that every region is doing what the entire cortex is doing if the cortex is trying to build a model of the world through sensory-motor interaction. That's what every region is doing, which makes sense—it's very copacetic. It's not like some property that comes out of the hierarchy; this is what every part of the neural tissue is doing. When you hook them in hierarchies, you get nice other properties. So this is basically the core item of understanding how the neocortex works: understanding how one of these regions works, and it's the same pretty much everywhere you go.
Now we hypothesize—and I think it's pretty strong—you can deduce that to do these functions you need to have memories of sequences. You need to have memory of how things change over time. My speech right now is layer five cells in one part of my neocortex firing in a complex pattern; every word is a complex pattern, my phrases are complex patterns. I can repeat them—I can say these sentences over and over again, I can give this talk twice. I have these sequences memorized, and I can put them together in complex ways. But it's all playing back sequences, and that has to be stored in cortical tissue. Similarly, when I recognize speech or music or anything in the visual world, I'm recognizing patterns I've learned before. These are all sequences of pattern streaming data. So it's all about sequence memory here. This is the primary memory of the cortex: how things change over time. We have a theory about this which is that every layer of cells is actually implementing sequence memory. The reason they differ is because we're doing different things with them. Our current best guess—it's pretty clear that layer five cells are playing back motor sequences. Our guess is about layer four: they're learning sensory-motor sequences or inferring sensory-motor sequences, not generating them—recognizing them. Layer three is recognizing pure sensory sequences. I don't want to get into the details, but you can deduce that these properties must exist in the cortex for you to understand the world.
How does that work? We have a theory about exactly how this works. I'll give you the basics. First, we have to talk about neurons. The brain is made of neurons. This is a pyramidal neuron—it's 80% of the excitatory cells in the brain. I just learned recently that spiny stellate cells are actually also pyramidal neurons that lost their apical dendrite. So we can say this is the excitatory cell of the neocortex. These cells have thousands of synapses—it varies, maybe up to 30,000 synapses on a pyramidal cell in the hippocampus. These synapses are expensive. They're there for a purpose. 10% of the synapses are near the cell body, proximal, and they're able to make the cell fire in the classic way we think about in neural networks. 90% of these synapses are so far away from the soma that if you activate one of them, it either has almost indetectable effect or is indetectable at the soma. For many years people said, 'What the hell are they good for? What are 90% of these synapses doing if they don't really have any effect at the cell body?' We had some of this debate yesterday when Christof Koch was here. We now know that these synapses are all along the dendrites. There are no excitatory synapses on the soma; they're only on the dendrites. The spines are about a micron part. We now know that there are active properties to dendrites. This work is from Larkum, Major, and Schiller. The basic idea is if you have a number of synapses that are co-located on a dendritic segment within about 40 microns of each other and you activate multiple ones at the same time, they sum nonlinearly. Typically we talk about an NMDA spike, which is a much larger depolarization and longer, and it has a significant effect on the soma. An NMDA spike is very measurable at the soma—not sufficient to make the cell spike, but sufficient to make a significant depolarization. This diagram shows the difference between going from seven synapses in blue to eight synapses activated in red, showing that nonlinear property. As I said earlier, usually 15 to 20 synapses are required to create an NMDA spike. The basic theory is that there is a coincidence detector—you're detecting a pattern out in some larger neural tissue by detecting 15 to 20 synapses active at the same time, and you recognize that and put the cell into a depolarized state. If you follow that and do the math—built on sparse representations that Rod was just talking about—each pyramidal neuron can recognize hundreds of unique and independent patterns on its dendrites. It's not recognizing one thing; it's like hundreds of them. The basic theory is that most of those detection patterns are patterns that preceded the cell becoming active, and therefore they're predictions. The cell is saying, 'I have seen this pattern in the past; I typically become active afterwards. I'm going to depolarize and be in a more primed state to fire if I do get input.' These are patterns on the basal dendrites that depolarize the cell, and this is a form of prediction. The advantage is when the cell does become active, it will spike a little bit sooner than other cells with similar input. In a sense, cells that are predicted will become active and inhibit others, giving a sparser activation when you have a correct prediction. These are all observations known in the brain.
If you put a bunch of these neurons in a layer of cells, in a columnar representation, you end up developing a very powerful sequence memory. That's what we've been testing for many years. If you want the details, you can come to our poster. I just want to point out that we model these neurons. This is a picture of our software model. We have to model the individual dendrites, different integration zones—the apical, basal, and proximal—these are important. We have to model individual synapses. But there are a lot of things we don't model. Our neuron model is not a spiking model because we haven't found a need for that from an information theoretic point of view. We use a representational scheme for how to represent information in sequences. We use something like this picture: ABCD and XBCY because they have subsequences that are the same. If I show you XBC and I've learned that sequence, I predict D; if I see XBC, I have to predict Y. The point is that sequences are very complex in real data. These little panels represent how in a layer of cells with columns you would represent inputs that are predicted and inputs that are not predicted. It's really cool theory. I encourage you to learn about it—come by the poster later. It explains how a layer of cells can learn very complex sequences of any duration, merge and separate them, making predictions constantly, even multiple predictions at the same time. We also model the apical dendrite synapses. I want to talk about this: in this model, learning is not by modification of synaptic weights; it's by synaptogenesis, which is a much more powerful form of learning. We know this is going on all the time in the cortex. We heard yesterday it's a fact that individual synapses are largely stochastic—you cannot rely on them for any amount of fidelity. Even one digit is more than you can get out of a synapse. If your neural model requires one digit of precision, it's not going to work in a real neuron. But if you learn sets of neurons at once, you can get something like if you can learn 15 or 20 form new synapses, then you have something reliable. That's what we think is going on. The way we learn is we model growth instead of Hebbian-type learning—we model the growth of synapses, not synaptic weight change. We know this is going on all the time. In this picture on the left, you can see an axon and a dendrite, and we have something we call synapse permanence, which represents the growth of the synapse. At zero, there's no growth between these two. Then when I train, I increase the permanence. At some point you get to a threshold where you've actually formed the synapse and it's now active. You can continue training; the strength of the synapse does not change—it's a binary synapse—but the permanence increases. Why? First, it kind of models what's going on in biology. The importance is that you can start training on patterns before you know they're real—they could be noise. You have to see it several times before it becomes an actionable thing. Then you can continue to train, making it much harder to forget. Something you've been exposed to over and over will last longer than something only seen a few times. It allows continuous learning in the face of noise. The system is constantly trying to learn but only acts when it's seen a pattern multiple times. This is a very powerful form of learning.
We've tested this and built these things; we've applied them to commercial products. I just want to point out a couple things. The top point shows in the blue line one of these HTM sequence memories learning a predictive model of some complex data stream that's partly noise and partly structured data. You can see time on the horizontal axis; we start feeding in a stream of data—no batch here—and it starts getting up to perfect accuracy. At some point we change the data; it falls back in accuracy and learns again, constantly adjusting and learning even as the data changes. The bottom shows that the systems are very robust. This is cell death. If I train a system and then kill a bunch of neurons—up to about 40%—it's hardly noticeable. After that you see a sharp drop-off in performance, but even without all those neurons, it relearns how to do the same problem. It's extremely robust in every way: neurons, synapses, dendrites. This is an important property if you're ever going to build these things in hardware. I now want to switch topics to the functional components of intelligence. I just gave you the biological substrate: a system with a cortex and a hierarchy, regions doing sequence memory. What are the functional components? This is Hawkins' list—my list because it's subjective. I made it up, so I call it Hawkins' list of functional components. I think it's a pretty good one. First, if you're going to build an intelligent system, it will have to have networks of neurons that learn and recall sequences. That is the fundamental premise. All inference—auditory, visual, somatosensory—is inference of sequences. All motor behavior is playing back sequences. This is not something to be added to a system; it's the core fundamental principle. Key principles include continuous learning—it's not a batch system; it has to be continuous learning to be intelligent. I didn't go into this, but it has to make multiple simultaneous predictions and be very robust. Those are requirements for an intelligent system. I propose one model for doing this: the HTM model which uses active dendrites, synaptogenesis, and no spikes. There might be other ways; we've looked at LSTM. I believe this is how the brain does it, but it doesn't have to be how we implement an intelligent system. Second, we have to have regions that use sequence memory for sensory inference, sensory-motor inference, and motor generation. I think this is a fundamental requirement for intelligent machines or intelligent people. We don't do most of this today. Then you have to have a hierarchy of regions. The hierarchy is required, but there are many parameters: number of regions, size of regions, connectivity graph. You could have a system with one region or 200 regions, teeny little regions or big regions. That's a design parameter, not essential to the overall function. Finally, if we're modeling the neocortex, it has to have an embodiment—it has to exist in something. It has to exist with one or more sensors, one or more built-in behaviors (behaviors that exist outside the cortex, like old behaviors of the body). The cortex controls the rest of the body; it has to have emotions and motivations—they can be very simple or complex. It also has to have something equivalent to the hippocampus. I shouldn't say it has to have an embodiment, but these individual parameters are things you can adjust. It's not one-size-fits-all. Those are my functional requirements for building intelligent machines in the future.
Now let's talk about the diversity of intelligent machines. If you accept this list, let's look at the things we can play with—what parameters can we play with and what could that lead to? First, you realize you could build a system that works on cortical principles that is very small or very big. Yesterday Bruno talked about tiny brains—insect brains—those are not intelligent brains, but you can build small, tiny intelligent brains. We have rats and mice, but we could go even smaller. At Numenta, we build very small pieces of cortex with 65,000 neurons and several hundred million synapses that learn the spatial-temporal data patterns from sensors on machines, buildings, etc. They work on cortical principles but are very small. Would I call it super intelligent? No. Is it working on intelligent principles? Yes. So you can go from that variety. On the other hand, we could build things with very complex sensors unlike anything a human has, with very big hierarchies. I'll leave you with two of my personal aspirational goals. I doubt I'll see them in my lifetime, but I'd like to see them. One is a machine that is like a super mathematician or super physicist. You could build a hierarchy designed where part of it is used for mathematical behaviors, mathematical functions and operations, so the system operates in the space of mathematics. You might have one section dealing with topology and another with another aspect of mathematics. This system could be a super brain on mathematics—huge, fast, working non-stop 24 hours a day with the right motivations. This is possible using these principles. Another thing I'd love to see: one day we won't be able to live on this planet. I hope it's a long time from now, but it might not be. Elon Musk talks about sending people to Mars. NASA is serious about that, but they think they'll send people there and first build factories and mines. I don't think I want that job. So we need to make super engineer-scientist robots that go and do this stuff. That sounds crazy, but I'm not joking. We need to be able to send machines that can solve engineering problems, construction problems. If we're ever going to get off this planet, explore the rest of the solar system, we have to have machines that can do that because we are not going to survive in those environments. Why can't we make a machine that is really good at using tools, really a good engineer that can sit and solve problems and build things? We can send them out in advance of us. It's crazy, but we ought to be aspiring to something great. We shouldn't be settling for what we can do today. I've argued that these are the principles we need: embodiment, sensory-motor inference, building complex models. There's a huge amount to do, but it's not impossible, and we're making great progress. That's the end of my talk. Thank you very much.