If people started to pull money from OpenAI and they couldn't raise money, that would be a cascading effect in terms of the entire ecosystem.
Exactly. That's exactly what I think is going to happen. Our guest today is Gary Marcus. He is a real critic of large language models and what they're capable of.
The world has gone all-in on neural networks and invested massively in it on an idea that to me never made sense. Large language models are not going to get us to the holy grail of so-called artificial general intelligence. In the investment community, people are starting to say, "Hey, I see all this circular financing. Results on return on investment are not that great. Maybe this doesn't make sense."
Hey, this is Steve Eisman and welcome to another edition of the Real Eisman. And today we have a different kind of guest. Today, our guest today is Gary Marcus, who some of you may have heard of because I've discussed him on my podcast many, many times. He is a real critic of large language models and what they're capable of, which is really the foundation of the entire AI story. And so, Gary is going to discuss his theory, what LLM models are all about. Gary, welcome to the Real Eisman.
Thanks for having me. And also, thanks for that incredible shout out you gave me on CNBC a month or two ago.
Oh, you're very welcome. Well, it was well deserved. Before we get started, why don't you give my audience, most of them don't know you, so why don't you just give people your background so that they understand you're actually entitled to have an opinion on this subject.
So, I've been studying intelligence kind of all my life. I started working on AI when I was 10 years old, as soon as I learned to code. And I spent a lot of my career studying natural intelligence, humans, how kids learn language and things like that. But my dissertation, which I did at MIT, was on two things. It was on how kids learn language, but it was also on neural networks, which are a particular approach to artificial intelligence or to modeling the human mind that is, I would say, loosely inspired by the brain. It's a very nice piece of marketing. It makes you think that it's really based on the brain. It's not really, it has a loose affiliation, but these things called neural networks were popular. I looked at them in the '90s and they weren't really such good models of how human minds work. But I got involved in trying to understand what they actually did. And then when deep learning came back in fashion in 2012, I was like, hold on, I've seen this stuff before. It's a lot like what they did my dissertation on. And I had written a book in 2001 called The Algebraic Mind where I actually anticipated the problem of hallucinations and some of the problems of reasoning that we're seeing.
Which we're going to talk about.
And we'll get there. And so a lot of the issues to me were familiar when this new thing came back in fashion and I knew immediately that it was going to have problems in certain things. So I wrote a piece in the New Yorker in 2012 called "Is Deep Learning a Revolution in Artificial Intelligence?" And I said, "Look, this stuff is really interesting. I admire Jeff Hinton for having stuck to his ground for a long time."
Jeff Hinton. He just won the Nobel Prize last year and he's one of the major players in deep learning.
As is his students, who's someone who recently came around to my side. So, he was a major player in this field. He kind of kept the torch going when neural networks were not very popular to his credit. But, you know, not everything goes to his credit. We don't have to go into all of it, but he kept the torch going. He resuscitated this, and relevant to your listeners, what really resuscitated was his student Ilya Sutskever and maybe a couple other people figured out how to take these systems that they had been building for a long time. I mean really they go back to the 1940s and Hinton did some important work in the mid-80s. They showed you could do the same stuff but with GPUs, graphics processing units that Nvidia was building. And at the time, Nvidia was just building them for video games mostly.
And so these things existed for video games. And the student had the clever idea what they were really good at was parallel computation, which is to say computing something in many little pieces at the same time instead of one at a time. So the classic paradigm is step by step. You know, the central processing unit goes through some software program line by line. That isn't entirely true anymore, but it's kind of the simplification that people might learn if they took computer science 101. And what GPUs did is they let you break up a problem into a bunch of little pieces that you do at the same time. They were built to do this for computer graphics. You'd look at a screen or you would want to render a screen. You'd want to say, "What's the next frame in this video game?" Instead of doing it one thing at a time, which would take forever, you would do the whole scene at once. So, you'd have one little subprocessor doing this pixel and another one doing that pixel and so on. And I mean, they work great for that. I've occasionally played video games and what graphics processing units could do is incredible. So Sutskever and I'm blanking on the other guy who was an author on the paper showed was a good way of running these so-called neural networks and we can talk about what those really are and what that means but they found a way to run them on graphics processing units and that meant that they could go much faster and they could use much more data. Everything that people had done in that field at that point, I don't know, 60 years old or whatever it was, was kind of like a toy model and they showed you could do it for real if you did it on a GPU. You could do it at a much bigger scale. And so everything that we look at was kind of born in that moment in 2012. And as soon as it happened, two things happened. The New York Times wrote an article about how amazing this stuff was. And the next day, I wrote a piece in the New Yorker. I was writing a lot for the New Yorker then for their blog. I wrote a piece in the New Yorker saying, "It's great, but it does have these problems. It's always going to be good at some things, but not good at other things. It's always going to be good at pattern recognition and statistical analysis, and that's great, but we also do a lot of abstraction. So, we know how a family tree works, so we can reason about things in the world. And it's never going to be really good at that. It's not really suited to that." And this was obvious to me from the work I had done on these systems in their earlier incarnations and also from the work I had done on human cognition, understanding how human minds work. If you know Daniel Kahneman's famous book Thinking, Fast and Slow,
You know, he argues for what he calls system one and system two cognition. The system one is fast, it's automatic, it's statistical, it's reflexive. And system two is slower. It's more deliberative. It's more about reasoning. Neural networks are basically like system one. And that's fine. That's part of what we do as humans. But part of what we do is the system two stuff. Especially on a good day, we do the slower stuff. We're more deliberative. We're more reasoned about it. And these systems have never been good at that and they still aren't good at that. And what I said in 2012 is, you know, they do the one but not the other. And at one level of abstraction, what has happened in the 14 years since is that the world has gone all-in on neural networks thinking.
And when you say neural networks, this is what are called large language models.
So large language models are a form of neural networks. Sorry, I didn't spell that out. And in fact they didn't exist in 2012. So a few things happened but basically since 2012 people developed the large language models. The origins of that is really the transformer paper that came out in 2017 and invested massively in it, trillions of dollars at this point. One or two trillion dollars would be my back-of-the-envelope estimate on an idea that to me never made sense. You know, their idea was we were going to get to everything we need for intelligence, artificial general intelligence, but they weren't careful about the system two stuff. At first they just treated it like a giant black box. And still many people do. It's like we'll throw all this data in there and we'll get this system that will do what we need for intelligence without any sophistication from a kind of scientific perspective about what that would really look like. I think these people were very naive and I tried to point this out and this goes to the lone wolf and stuff. For a long time people were just dismissive. They're like we have this thing called.
They were more than dismissive of you. They were disdainful. I mean really disdainful.
We could probably come up with some other stuff. I was disillusioned with them. We could go on that for a little while. And they were openly hostile to me like OpenAI has an emoji for Gary Marcus. I read this in here now.
Well, it's flattering in a way.
It's flattering and next level at the same time, right? I tried to take it with a sense of humor as you can see. But, you know, that is a measure. And, you know, Sam Altman called me a troll on Twitter or whatever. You know, they really didn't want to hear what I had to say. The core of what I had to say was in 2022, I wrote a paper called "Deep Learning Is Hitting a Wall." And what I said in that paper is you can keep scaling this stuff which was the idea that was already starting to get popular. Maybe I should pause for a second. The idea of scaling was that we could just pour more data and more GPUs.
And make the model bigger, bigger, bigger, bigger.
Make the model bigger and bigger and it would get to be amazing. And they had some data in support of this but it was also naive. I call it the trillion-pound baby fallacy. Right? So your baby weighs, you know, 8 lb at birth, a month later it's 16 lb. That doesn't mean it's going to be 32 and 64 and then it's going to be a trillion pounds when it goes to college, right? They made this very naive kind of inference, which I'm sure you see all the time. In the business world, a lot of people, a lot of smart people with a lot of money made this kind of bet. They said, "We see these laws and we project that we will get to intelligence by putting in such and such amount of data." And it's.
Kind of just pause just for a second. Let's just backtrack for a second. So, large language models, like what do they do? And what do these people think they're supposed to do? But I really want to kind of break that down. There's a really great way to ask the question.
So, what they do fundamentally is they predict next things in a sequence. So, think about autocorrect on your iPhone if you have an iPhone, right? Similar.
Which sometimes drives me absolutely insane, but go ahead.
It doesn't always work, but the idea is you're typing out a sentence and it predicts what might come next. So you say, "Meet me at the" and restaurants a pretty good guess, right? So you make a statistical analysis of what people say. And you know, you do okay, not perfectly, and sometimes it makes mistakes and it's annoying, but that's what we call autocomplete. And I call LLMs autocomplete on steroids. They're a special way of doing that prediction process. That's fundamentally what they do. And there's some interesting things about how they do it. One of which is they break everything into little bits and then they reconstruct things which means they actually lose connections between information which means they sometimes hallucinate. They make stuff up.
Well, let's we'll come back to the hallucination.
So come back to hallucinations. That's one of the characteristic errors and that's one I pointed out in 2001 before LLMs were even invented. I said if you keep following this approach to its kind of logical destination, you're going to have this problem. And we turned out to have that problem. So LLMs break stuff up into little bits and they make predictions about what might come next. That works surprisingly well if what you train them on, what you expose them to is the entire internet because almost any question, but almost is a critical word there. Almost any question that you might think of somebody's asked before and somebody's answered before.
And at some level these are like glorified memorization machines.
There was an article in The Atlantic about this just the other day and there's lots of evidence for this right along. So, for example, if you type in part of Harry Potter, it will just finish the, you know, the paragraph or something like that because it's basically memorized things. And if you basically memorize the entire internet, that's kind of special. Like, because when you go and ask a question like, you know, where did the Dodgers play, you know, before they were in Los Angeles, there's lots of sentences. They'll tell you they were born in Brooklyn. And so you have a pretty good chance of coming up with the answer. However, just doing that doesn't give you abstract concepts and ideas. And sometimes you have problems where these little bits get broken up and then reassembled incorrectly.
So like can we talk about hallucinations for a second? That is what hallucinations.
Define halluc, like give us an example of a hallucination and why does that happen?
Hallucinations are when it makes something up, presents it with perfect confidence and it's just not true.
My favorite example involves a guy you might know, Harry Shearer. I don't know if you ever saw Spinal Tap.
So, he plays the bass in Spinal Tap. Okay. And he's, as it happens, he's a friend of mine. So, he plays bass in Spinal Tap. And also, and it matters to the story that he's reasonably well-known. So, he played bass in Spinal Tap. He was in a bunch of those movies with Christopher Guest. He was in The Truman Show. And he does the voices for Mr. Burns and The Simpsons and a few other characters.
So this is what makes this story interesting. So I'm going to rewind for one second, which is to say my old favorite example of hallucinations involve me. Somebody sent me a biography of me that said I owned a pet chicken named Henrietta, which I don't. So that's a pretty good example of hallucination. It's just made up. Turns out there was an author named or an illustrator named Gary Oswald or something who wrote a book about Henrietta goes to school or and it's just like munching together all these little bits of information. Why does it hallucinate?
So, this has to do with the breaking up of little bits of information. So, let me walk you through the Harry Shearer example. So, I kept using the pet chicken Henrietta thing. So, one day he writes to me and says, "No, Henrietta," but then he gives an example where the hallucination is about him. He's much more famous than I am, or at least I used to be. I was starting to get a little known. And it says that he's a British voiceover actor and comedian only he's not British. If you went to Wikipedia for two seconds, you would see he was born in Los Angeles. And because he's famous, you could also go to Rotten Tomatoes or to IMDb and he's done lots of interviews and talked about where he's growing up. He was a child star on the Jack Benny Show in Los Angeles. Like, it's not hard to find the right information. We imagine falsely that large language models are intelligent beings like us, but really all they're doing is reconstructing statistically probable relationships between bits of information.
And so they can be wrong.
And those can be wrong. Sometimes those reconstructions are wrong and in this case it was wrong. In some sense, what it does is it builds like a cluster of things that predict other things statistically. And it turns out there's a lot of British voiceover actors and comedians. There's, you know, Ricky Gervais and John Cleese and etc., etc., right? And so it just blurs all of that stuff together. And the blurring process on the whole works reasonably well, but you can never trust any particular thing is going to be right. And so these hallucinations happen all the time. There was a guy who was tracking legal cases where lawyers published or not published but submitted briefs where they were made up references, citations to cases that didn't exist. The first time I looked he had found like 300 cases. The next time it was like 3 months later, he found 600 cases of lawyers not only did this but got caught, got busted by judges doing this. They were using these tools like ChatGPT to do their work for them. But it makes mistakes and those mistakes, this is the insidious part, slip by. People don't notice. Another example is CNET was one of the first places to use AI to write its articles. And in the first batch of 75, like half of them had errors. The editors didn't notice them because everything is grammatical and well-formed and there aren't typos and people tune out. I call it the "looks good to me" effect. LLMs give you the "looks good to me" effect which has led to another new term that I wish I had invented called work slop. So work slop, which a couple professors I guess invented last year, is a term for people write these reports, they submit them, you know, to their employer, they look superficially good, but they're wrong. They have mistakes in them because LLMs don't really understand.
And what you're saying is LLMs don't think.
They just slam things together that statistically normally makes sense to slam them together.
Exactly. Or glom is another verb I like there. They glom things together. And you know a lot of them are right statistically speaking and some of them are wrong and the systems don't know the difference. They can't tell you the difference. They don't ever say like well it seems to me that you know everything like Wikipedia says that Harry Shearer was born in Los Angeles but I have the vibe as an LLM that it's London and maybe you should check that. I mean they never give you any of that. They just present everything as if they were an encyclopedia no matter whether it's true or not which is one of the reasons why these systems are insidious. In my Substack that's going to come out before this airs that I've already written in my next Substack.
Thank you for subscribing. I'm going to talk about a new article I saw this morning that blew my mind that talked about basically how LLMs are undermining the institutions of society, including things like democracy and civic order and so forth in no small part because of this tendency towards errors.
Right? So, they're undermining the quality of almost all the institutions in life by making fake information and basically undermining the information ecosphere. I mean, just like think about democracy. Like democracy only works if the voters actually understand what's going on, right? The whole point is the voters reflect on the world and then they bring their perspective from, you know, their religion, their childhood, their neighborhood or whatever. But if everybody's getting garbage all the time, that doesn't work anymore. Well, I was reading on the day that Maduro got taken by the United States. You wrote here we were all watching taken up like the whole country is sitting there watching like CNN or whomever and we're literally watching this guy get taken out of Venezuela and ChatGPT is telling you that it hasn't happened.
That's right. That it is and.
That it's fake news that we're watching it on the TV and they're telling you it's fake news.
Yeah. I mean there are many problems that's a different problem but it's not unrelated.
That's what is that problem?
That problem is these things all have a cutoff date so they get trained at a particular moment and there's kind of a core model that only knows what's happened up to then. They put different band-aids on it like they make it do web search and whatever but none of the band-aids are very well integrated and you know some of the systems do that better than others has a problem with novelty.
That is the deepest problem actually going back to my work in 1998 I realized very early on that was the deepest problem with these systems is if you're basically a glorified memorization machine and I bring you something that's far enough away from what you've seen before you're in trouble. Amazing example of this. I don't know the underlying details, but Tesla does a lot of this kind of memorization kind of stuff, but its AI systems are not that sophisticated. And one day someone used the summon feature. You remember how Elon said that you're going to be able to call your car from Los Angeles to New York. You can't really do that, but you can call it across a parking lot apparently. So, somebody did this at an airplane trade show. You can find the video on YouTube. They called their car. They wanted to be like, "Cool, I'm at an airplane trade show and I'm going to have my Tesla come over to me." And it ran directly into a three-and-a-half-million-dollar jet, right? The system didn't have in its training data what to do with a jet because who trains a car to avoid a jet?
And so it didn't have a general understanding of the world, like don't drive into things that are expensive or big or physical objects in your path. Didn't understand any of that. Just like looking for things that match bicycles and pedestrians and so forth. They didn't have a category for jets and so we ran into it. Let me press a question. So,
Hi, Steve Eisman here. You listen to my show because I try to give you the facts about the market and my opinion. But it's not always easy to filter through all the information and hype out there, especially when it comes to figuring out how to invest. And if you've had some success in life, you are probably getting bombarded with investment advice. Suddenly, everyone has a great new idea or product for you, a new fund, a tax scheme, a once-in-a-lifetime private deal. How do you know what's real? By joining a community of people who face the same concerns as you. That is why I'm recommending you take a look at Long Angle. Long Angle is not a wealth manager. They aren't trying to sell you a product. It is a private vetted community of high-net-worth investors, entrepreneurs, executives, and professionals. It's a community filled with people like you. Members use it to compare notes on everything from investment ideas to due diligence on private deals to navigating complex tax codes to the personal stuff like how to raise grounded kids when you have money. I talk a lot about old-school due diligence and that is what this is all about. It's 6,500 vetted members sharing unbiased intelligence. No salesman allowed. If you're looking for a place to ask questions and get answers from people who aren't trying to earn a commission off you, this is it. Go to longangle.com/eisman. Membership is free, but you have to qualify. Again, that's longangle.com/eisman to apply.
These models have scaled. You know, it was ChatGPT 1, ChatGPT 2, 3, 4, 5, and now Gemini is supposed to be better. When people say that the new Gemini is better than, let's say, ChatGPT 5, what does that mean in practicality?
That's actually a really interesting question because there isn't a simple answer to it. What they're really saying is like for the things I do, I happen to get better results. Most of the people who are saying that are saying that intuitively without even quantitative data. And the thing about these systems is we are pressing them into a very broad service unlike almost anything else. Like if you have a car, you can test it. How does it go in the snow? How does it go in the rain, etc. There's a kind of regime that you can test and it's pretty well known. If you had a calculator, we don't really need to do this because we know almost by proof by construction. We know that the calculator will be correct. When the Intel chip made an error in floating point numbers, I think this is about 20 years ago. It was a huge scandal that it wasn't 100% correct, but it was like 99.9999% correct. But there's only so many things you can ask a calculator to do until you can basically prove that it is correct. We can't do that with LLMs, with large language models, because they can be asked to do anything and different people ask them to do different things and so everybody has their own opinion about which model is better. Like almost certainly on almost everything with some qualification GPT-5 is going to be better than GPT-4 and is going to be better than 3. And it's a question of how much better but it's a hard question to answer. It depends very much on what you're testing it on. So when people say, you know, Gemini is better than GPT-5.2 or whatever, they really mean for the things I do and people do very different things. So some people use these systems, for example, to help them with coding. Other people use them to help them with brainstorming. Some people may do financial analyses and so forth and so on. There's sort of endless and it makes it very hard to make a definitive statement. Instead, you know, people kind of try it out and they have the intuition. I actually haven't played that much with Gemini. I played with a lot of these models and kept seeing the same patterns over and over again. And to me, that's boring. Like I come to all of this from a scientific perspective. I want to see something that works differently. And when I see something that works differently, I'm going to spend a lot of time testing it.