Back
Andy Chen
Vice President & Treasurer, General Dynamics Corp

Andrew Chen: "Is Everything I was Taught About Cross-Sectional Asset Pricing Wrong?!" | RR 316

🎥 Aug 01, 2024 📺 TheRationalReminderPodcast ⏱ 59m
Meet with PWL Capital: https://calendly.com/d/cpws-jyp-znp Are you curious about the hidden factors driving your investment decisions? Today’s guest is Andrew Chen, a Principal Economist at the Federal Reserve Board who focuses on monetary policy and financial stability. Published in leading journals, his research informs key policy decisions and helps shape the Federal Reserve’s strategy for managing economic challenges effectively. In this episode, Andrew delves into the intricacies of meta-research and asset pricing, focusing on cross-sectional asset pricing predictors, replication, and ou...
Watch on YouTube

About Andy Chen

Andrew Chen, a principal economist at the Federal Reserve Board, appeared on the Rational Reminder podcast in August 2024 to discuss his research on cross-sectional asset pricing. Chen stated that his views are his own and not necessarily those of the Federal Reserve. He described a predictor as a variable that forecasts asset returns, which can be turned into a factor or trading strategy. Chen outlined three explanations for why cross-sectional predictors exist: they compensate for risk, they result from mispricing, or the statistics are wrong. He noted that academic finance faced a trade-off between open collaboration and close competition, leading his team to create an open-source dataset. Chen reported that of roughly 300 variables collected, about 200 predicted returns in original papers, and his replication judgment replicated all but three, a failure rate of roughly 1–2%. Chen discussed the decay of anomaly returns, stating that if a paper's sample ends in 1989, anomaly returns decline by about 50% from 1990 onward. He said transaction costs eat up about 25–30% of trading-strategy returns in original samples. Chen observed a kink around the mid-2000s in factor returns, which he hypothesized was due to the internet making information more accessible. He suggested that data-mining accounting ratios in 1980 would have uncovered anomalies decades before publication. Chen advised trusting numbers more than text in academic papers, verifying that documented data supports claimed mechanisms.

Source: AI-verified profile updated from Andy Chen's recent appearances. Browse all interviews →

Transcript (75 segments)
B
Benjamin Felix0:10
Weekly reality check on sensible investing and financial decision-making from two Canadians. We are hosted by me, Benjamin Felix, portfolio manager at PWL Capital, and Mark McGrath, associate portfolio manager at PWL Capital. Welcome to episode 316. So Cameron is obviously not here today, so Ben, good job on the mod intro there. Two Canadians, not three. This is a big thing.
M
Mark McGrath0:33
Can we just real quick for a second pause here? This is your first episode with a guest and the first time you and I are together doing an episode without Cameron. I mentioned that to my wife. She's away right now, but I text her this morning like, 'You're not going to believe this, I'm replacing Cameron today.' I still pinch myself over it. It's wild to me after being such a fan and a listener for all these years.
B
Benjamin Felix1:11
I've blown multiple times through this conversation. Yeah, so Andrew's research looks at meta research, so looking at meta studies, large scale replications, large scale statistical inference. Meta science being the broad term of all of that. But he's doing this on cross-sectional asset pricing literature. Factor investing is based on that, and so he's looked at this from a whole bunch of different angles. He's looked at replication, out of sample performance, whether factor premiums survive transaction costs. I'm not going to give away the answer on that, but it kind of turns it on its head. He kind of says that for a very long time, cross-sectional asset pricing research was kind of not going in a great direction, and it's improved more recently. We talked about that near the end of the conversation. Some pretty serious stuff. Any other thoughts before I talk a little bit more about who Andrew is?
M
Mark McGrath2:37
No, just like you said, incredible conversation and really very insightful. To your point, it has me thinking about the topic more and more. What I do with the information that I learned in this episode remains to be seen, but as you said, I won't spoil it for the listeners.
B
Benjamin Felix3:10
Andrew Chen is a research fellow at the National Institutes of Health, which is interesting. His primary research interest, as I mentioned a minute ago, is in financial economics, specifically understanding whether finance research is going in the right direction using tools from meta science. Most of his papers are around that. His research has been published in Management Science, the Journal of Financial Economics, Critical Finance Review, the Journal of Finance, the Journal of Financial and Quantitative Analysis, the Review of Asset Pricing Studies, and the Review of Financial Studies. For anyone unfamiliar, those are the top journals in financial economics. His focus on cross-sectional asset pricing is really going to challenge the beliefs of everyone listening. Okay, it's a great introduction. Let's go to the episode. Andrew Chen, welcome to the Rational Reminder podcast.
A
Andrew Chen4:24
It's great to be here. I just need to put the disclaimer up front that my views are not necessarily those of the Federal Reserve Board or the Federal Reserve System.
B
Benjamin Felix4:36
Perfect, understood. Andrew, what is an asset pricing factor and how is it different from a predictor?
A
Andrew Chen4:43
Okay, so a predictor is like how it sounds. It's something that predicts asset returns. A very popular one is when you take the book value of a company and divide by its market value. That predicts returns. But there are a lot of other things people mean when they say 'factor,' so I prefer the word 'predictor' and try to use it as much as possible. But the language has gotten all squished. People say predictor, factor, and anomaly and just squish it all together. I try to stay with 'predictor,' but it's okay if we deviate a bit.
B
Benjamin Felix5:27
Okay, makes sense. What are the plausible explanations for why cross-sectional predictors exist?
A
Andrew Chen5:33
There are basically three explanations. The first is that it's due to risk, that the strategy is giving you high returns because it's exposing you to something that you should fear and it's rewarding you for that exposure. Another one is of course mispricing, just that the price isn't right, it's going to move up in the future or it's going to move down. The third is just statistics.
B
Benjamin Felix6:11
Interesting. So we've heard about the predictor zoo or the factor zoo with over 400 predictors. How many predictors are there really?
A
Andrew Chen6:21
If you're talking about predictors that are clearly documented in the academic literature, I would say there's roughly 200. I say roughly because it's not always very clear if one predictor is different from another. Also, it's possible to say if you really have the whole literature captured. People have this idea of 450 anomalies out there from a paper which I think has some issues. They're pretty creative with their language, as you can see. I try to be strict about my language, but I also make mistakes too.
B
Benjamin Felix7:10
So you've done a lot of work on the open source asset pricing project. Can you tell us about that?
A
Andrew Chen7:12
Sure. So academic finance and science broadly always has this trade-off between open collaboration, standing on the shoulders of giants, and close competition. If you don't have any competition, you might not get those cutting edge insights. When we started on this project, academic finance was very much close competition, the business school culture to compete. But then we got to the state where we're publishing these papers with machine learning algorithms that are really complicated, fitted to dozens and dozens of predictors, and it seemed like it was impossible to openly collaborate at that point. So that's what motivated the open source asset pricing project. There were also some high quality papers that were misleading people into thinking there's some kind of replication crisis, and we wanted to put out a paper to really counter this narrative.
B
Benjamin Felix8:20
Can you talk a little bit more about the open source asset pricing project?
A
Andrew Chen8:25
Sure. What we did there is we took about 150 papers on cross-sectional asset pricing and we found any kind of result about predictability in the paper. To be clear, it's about predictability. What is a factor? I don't know, everyone disagrees, but you can see it's a predictor, right? There's a regression that shows predictability. We tried to replicate every single one of those predictors. Not only that, we posted all the code online. Finance researchers don't have the best code, but I think it was a good step up from what people often provide. If you go to some Stanford professor's website, you might not like what you see. I hope ours is a bit better.
B
Benjamin Felix9:24
Great. And how did you choose which predictors to study?
A
Andrew Chen9:30
As I alluded to before, there are some papers out there pushing this replication crisis narrative. One is by Hou and Zong, published in the Review of Financial Studies. It has a ton of citations. That's the paper where they have 450 anomalies, but there's really only 150 predictors underlying there. Their main result was that they failed to replicate the anomalies, so we wanted to clarify that that's not really what you see. There's another meta study by McLean and Pontiff that shows how predictability decays out of sample. If you read some numbers in the paper, it tells you that you've got to divide by two in the future if you're looking forward. So we wanted to replicate everything in there too, and we got that pretty much done. Then there is another meta study that didn't actually replicate anything. That was more challenging. They had a lot of stuff that was very difficult to deal with, and we still tried to replicate everything that's in there. Ultimately, we got this list of 300 variables that may or may not predict returns, but they were discussed in these previous meta studies. We wanted to be comprehensive and cover all of that.
B
Benjamin Felix10:50
So how many of those were you actually able to reproduce?
A
Andrew Chen11:11
Well, first of all, some of these papers never really look to see whether there's anything to replicate in the first place. That's something to document, and I think we document it clearly as part of the open source project. We have these spreadsheets where we actually write down the page that you should look at for predictability, and if there's no predictability, we just say there's nothing in there. Of those 300 variables, about 200 actually were shown to predict returns, and we replicated them. Replication is a bit subjective. It's not clear if you get the number seven and the original number was eight, is that good or not? But our judgment was that we replicated everything but three. There were three failures that were pretty clear.
B
Benjamin Felix12:11
You mentioned the other paper that showed maybe replication was not so great. Why were your results different from that?
A
Andrew Chen12:18
If I'm frank about it, it's because they're very loose with their language. It is pretty surprising because this is a really influential paper, still highly cited all the time in this top journal. But really, they never define what an anomaly is, and they never define what a failed replication is. They define what a failed replication is, but it's not a good definition because they don't define an anomaly. Let me describe an example. One of the papers in our data set, both of us share this paper, is by Campbell, Hilscher, and Szilagyi. It's about bankruptcy. The author wants to know, do you get a reward for bearing bankruptcy risk? Do stocks that have high bankruptcy probability have high returns in the future? He finds that the answer is no. They actually tend to have low returns. The relationship is the opposite of what you would expect in an efficient market. You're not going to be compensated for that. But he doesn't find that this relationship is statistically significant. It depends on the test. Some of the tests show significance, some of them don't. But the relationship goes the wrong way. That's a very interesting result, definitely worth knowing about. But when you replicate it, you need to be careful. If you find an insignificant result, that's actually a successful replication if you get the direction right. In the Hou and Song paper, they don't care about that. They call it a failed replication. To make it more interesting, they make three or four versions of this bankruptcy variable, and then they get four failed replications, although they should all be called successful replications if you did this carefully.
B
Benjamin Felix14:34
Okay, very interesting. How do anomaly returns change post-publication in your data?
A
Andrew Chen14:41
I want to be a bit more precise about that. Instead of post-publication returns, let's talk about out of sample returns. It takes a long time to publish a paper, so usually if someone is working on a paper, they have an in-sample period that ends before they start writing. The out of sample period is after that. On average, the returns decline by about 50%. So if you had a 10% per year return in sample before 1989, from 1990 to present you'll have a return of about 5%. But there's a big nuance there. In the first few years out of sample, from about 1990 to 1993, the decay will only be about 25%.
B
Benjamin Felix15:37
So you find you replicate a lot of these things, you find the premiums decline but don't go away. What would you say are the implications of this research for the so-called replication crisis that people have talked about in cross-sectional asset pricing?
A
Andrew Chen15:52
The main implication is that the idea that there's a replication crisis in cross-sectional asset pricing is just false. I think it's become well accepted that for the equities literature at least, that's the case. But then you could also think about replication crisis more broadly. I'm not a fan of how this is being used, but people also interpret this as external validity being wrapped up in replication crisis, and that makes sense. But when you're talking about stock markets, you have to be a bit more careful. If a predictor is based on mispricing and the markets are reasonably efficient, which we hope they are, you would expect the predictability to decay out of sample. What you're really afraid of is that these predictors are just statistical figments. If that was the case, you would expect the returns to vanish immediately out of sample. It shouldn't take five years. Right after 1990, you should see the returns go to zero. But that's not what we see. We see them gradually decay to 50%. That tells you that's also a relatively minor problem. I should credit McLean and Pontiff for that result. In academics, we care a lot about that. They are the ones who really discovered it. We replicate it, which I think is pretty cool. They're a meta study that replicates many papers, and we're a replication of their replications, and you could still see this result. It's perhaps one of the most solid results in the literature.
B
Benjamin Felix17:54
What is a false discovery in statistics?
A
Andrew Chen18:10
A discovery is something that's statistically significant using some tests. But even though something is statistically significant, it could still represent no effect at all. They call that null in statistics. In traditionally in asset pricing, that would mean the true alpha or the true expected return is just zero. A false discovery is when you have a statistically significant result, the statistics look good, but the actual underlying effect is zero. There's actually no alpha there. It's just a false positive.
B
Benjamin Felix18:45
How large is the false discovery rate in cross-sectional asset pricing research?
A
Andrew Chen18:51
I think you could reliably say, and this is just a statistical concept so there are caveats, but it's about 10% to 20% of the predictors that are statistically significant. This is based on our meta studies. This is another case where we might be getting a bit controversial because we're finding stuff at odds with other people's results. The way that things are at odds is also a bit striking. The previous paper on this, which is still probably the most well-known paper by Harvey, Liu, and Zhu, also in the Review of Financial Studies, extremely influential, basically they conflate insignificant with false. There's a difference between insignificant and false. Just because you have a negative result doesn't mean it's false. It could just be that you don't have enough data. If you equate false and insignificant, then you're ignoring this uncertainty.
B
Benjamin Felix20:24
So that's something that they just didn't think about, and you did and put it into your research. I can't get into the heads of what those authors were thinking. I think the rhetoric actually just got too strong. Another thing in the paper which is probably more striking than equating false and insignificant is they equate many with most. You actually have to dig very deep into the paper to understand any of this. It took a long time for even me to understand as someone who specializes in this. In the abstract, they loudly proclaim 'most claimed research findings in financial economics are likely false.' But you have to get to page 26 of the paper, get through all these statistics to really understand what's going on there. If you get there, you'll see they mix up both false and insignificant and also mix up many with most. It's pretty incredible. I think it's because of the incentives. I'm vulnerable to this too. Perhaps I've also overstated things before, but there's a big incentive to make big claims in academia.
I want to keep going on that. Can you talk about publication bias and how it relates to false discovery rates?
A
Andrew Chen22:10
Publication bias is like I was alluding to. No one wants to read something that's not interesting, and no one wants to write something uninteresting. People who write for a living want to write things that are interesting. So we all get together and look for only the interesting stuff, and sometimes we push it too far. Another way to think about it is it's an inevitable fact that if we're only looking for interesting stuff, some of that stuff will be not interesting. If we're biasing our reading or our writing towards stuff that's interesting, some of it will actually not be there. That's publication bias. How that relates to false discovery rates is that you could think about it as a multiple testing problem. Traditionally, if you go to your high school level statistics, you'll say, 'Oh yeah, you have a hypothesis, you construct this test. If the test statistic exceeds some value, then it's significant.' But that procedure assumes that you have a single test. What happens if you have many, many tests and you only think about the very extreme test in the tail? That's when you need to use these false discovery rate methods. It's interesting that they're relatively recent in statistics. They didn't really get developed until the late 20th century.
B
Benjamin Felix24:10
So false discovery rates and publication bias are really statistical concepts, at least the way I've been using them. The statistics are only as good as the ingredients you put in. They only reflect the economics that you put into it. The economics that we've been talking about, perhaps surprisingly, don't account for transaction costs at all. I don't want to go too long on a tangent about this, but there's perhaps a good reason for that. No one agrees on how to measure transaction costs. If someone puts in some number, someone's going to want to take it out and put in their own number. So in a way, it makes sense for the academics to just focus on gross returns. But if you try to measure transaction costs, you find that they eat up about 25% to 30% of the strategies' returns in the original samples.
Wow, that's a lot. When you think also about the out of sample performance declining, does that eat up all the out of sample performance?
A
Andrew Chen25:25
It's a good question. If you know the literature, you know that people have these equal weighted trading strategies. They weight every stock equally no matter how liquid it is. My intuition was that 25% to 30% is about right of what you would have in sample. But I've seen some recent work that says that in the recent period, from about 2005 to the present, even the best of these anomalies or predictors earns nothing in this very baseline transaction cost case.
B
Benjamin Felix26:34
Wow. So on the transaction costs though, are those worst case transaction costs or are you using cost mitigation techniques that a real asset manager might use?
A
Andrew Chen26:43
No, they're not worst case. We are using some transaction cost mitigation techniques. We look to see if we can modify the strategy to eliminate these tiny stocks, or we can try to tilt towards the large stocks. But there are definitely caveats on this. People argue about how to measure transaction costs. Our cost is the effective bid-ask spread, which you can think of as assuming aggressive market orders for a small trader who doesn't have price impact. We're ignoring short sale costs. There's a recent paper that says that just short sale costs will eliminate these anomalies altogether. But at the same time, it can get messier than this. The real way to think about these predictors is in combination. There's an economics to this that transaction costs are actually lower if you use more predictors at the same time because if you combine them, you can reduce trading. There are some recent studies that find that even net of transaction costs, if you can combine predictors, you can still get some returns.
B
Benjamin Felix28:22
Wow, that's really interesting. I hadn't thought about that, that the combination of predictors can actually reduce trading costs. What are the problems with estimating factor expected returns using historical data?
A
Andrew Chen28:36
Of course everyone knows that past performance doesn't necessarily reflect future performance. But one thing that is kind of coming out of the anomalies literature is that the markets get more efficient over time. If you're using historical data, you need to think about the era that your data reflects. For example, a lot of this data, like I'll keep using this 1990 sample end date because Fama and French 1993 or 1992, there are these very influential papers at that time using data from 1963 to 1990 or 1989. You think, how would you trade on book to market in that time? For a lot of people, you'd have to get these accounting statements in the mail and write everything down and make a ton of phone calls. It's a completely different world now. Now it could take a microsecond to do that once you get the data. So that's the big problem. Think about the era that your data reflects and how you need to take a haircut off of that because of the new era.
B
Benjamin Felix29:56
So what is the kink? When does it happen?
A
Andrew Chen30:11
First, before I speculate or try to give my view on what it is, the fact is that from a lot of studies, you see these anomaly returns or predictor returns or factor returns, they all decay. There's a kink around the mid-2000s. It's pretty cool that you see this across many studies. The hypothesis that I prefer is that was when the internet took off. Thinking back when I was an MBA student, even then it was still kind of growing. It could be somewhat difficult to get an annual report and digest it. You'd have to go through a little bit of hurdles. Now it's just all there. Now AI will read these things. There's a competing hypothesis though, which is that it's publication bias. If you look at our open source data set, one of the things we provide is all this documentation for every predictor, what table to look for for predictability, sample dates. You could plot the distribution of publication years, and there's a big peak in the mid-2000s. So that's the competing hypothesis.
B
Benjamin Felix31:30
Interesting. So the empirical fact is that after 2005, there's kind of a kink where factor performance decreases. It could be because markets got more efficient with easier access to information, or the publication thing could also be related to market efficiency. But in either case, ultimately it's market efficiency. But was it just academic, or was it a real effect?
A
Andrew Chen32:12
I think people disagree on that. If you take the transaction costs we talked about and the stale data idea, how much of a combined effect does that have on factor expected returns? The two effects are transaction costs and the post-2005 kink. After that, it's pretty much gone. The returns are pretty much gone for individual anomalies using our baseline effective bid-ask costs, even the best one. There's an important thing to do here. If you look from 2005 to now, of course there's going to be some things that got lucky. Once you adjust for the luck, you shrink the returns. You could call it shrinkage. Basically, the returns are just very close to zero. I say effectively zero because we're ignoring short sale costs and price impact. I want to be kind of qualitative with this because everyone disagrees on how to measure transaction costs. But of course you'll get some things that are lucky. If you plot the distribution of returns and compare that to a simulation where there's actually no real predictability, they look exactly the same. Or not exactly, they look very similar. The simulation of nothing there looks very similar to the actual data.
B
Benjamin Felix34:10
So what are the most robust factors or predictors after accounting for these things?
A
Andrew Chen34:14
Once you account for trading costs and account for the modern era of technology, nothing really does much. There's some stuff in the tails, but you want to shrink that towards zero. It's actually funny. The paper was published a few years ago, and one of the things that sticks out in the details doesn't work anymore just a few years after that because there's luck. If you talk about factor combinations, our paper doesn't really study that in detail. We have a bit about that in the paper, but there are other papers on this. If you take a look at them, they actually disagree on what is the most important stuff. There's a paper by Gu, Kelly, and Peterson that finds that the most important stuff is something different. I think it really depends. In my view, it's not really the best way to think about this. The big picture from this literature is that none of these predictors are actually super special. They all kind of decay out of sample. They're really all just kind of these mispricings that get corrected. What you really want to think about is the method, the meta method for collecting these predictors and for combining them and for eliminating them once they get stale.
B
Benjamin Felix35:47
So you just told us that factor premiums are zero after costs. What are the implications of this for investors?
A
Andrew Chen36:11
I think the main implication is that you can't just read this old paper and implement it. You could try to avoid the transaction costs we measure. You have to be really good at placing your orders, really optimize. But I don't think you can just do that. There could be a reason why these predictors, if it's not mispricing and it's due to some kind of fundamental thing that's not going to be corrected, and you don't mind bearing that risk, then that would work out. But it doesn't seem like that's the case. It seems like there's nothing really special in this zoo. I wouldn't say this is the consensus from the literature, but it's kind of what you'd expect in an efficient market. If they're mispricing, the premiums would decrease up to the point of the transaction costs. The complication is that this is kind of like the Adaptive Markets Hypothesis. If you've heard of this concept from Andrew Lo, I think it's an influential concept but it doesn't get enough credit. The classical idea of efficiency is that the market is efficient all the time. In my view, that's how these predictors became factors. In 1980, Dennis Statman finds that book to market predicts returns, and the efficient market view is that shouldn't exist. 13 years later, Fama and French repackaged that. They said, 'Oh, this is a risk factor.' They wrote down a model and it became a factor. But it doesn't seem to be super special. I've been talking about the anomalies as a whole and how they decay, and book to market is in the middle of the pack.
B
Benjamin Felix38:22
So what about peer reviewed factors with strong theoretical underpinnings? How do those perform relative to a naively data mined factor or predictor?
A
Andrew Chen38:32
This touches on a recent paper by myself, Alejandra Lopez-Lira, and Tom Zimmerman. The short answer is that even the stuff with the strongest theoretical underpinning, stuff with these fancy equilibrium models, they don't do any different or any better than data mined predictors. We look at the theoretical underpinnings. Are they a mispricing idea or risk idea? We also look at the modeling. Do they just wave their hands and say, 'Oh, this should predict returns'? Or did they write down a toy model? Or do they have these quantitative equilibrium models that are calibrated to capture key moments of the US economy? These are the holy grail that people in my world want to create. My dissertation was actually one of these really fancy models of the value premium. It doesn't matter. If anything, the fancier the theoretical underpinning, the worse they seem out of sample. In one sense, they actually tend to underperform mispricing based predictors. When we're talking about risk based on what the words in the paper say, that's kind of an innovation there. Usually people use factor models to measure risk, but we're just like, let's see what the pure review process says. All these sentences in the papers, they're not just the view of the author. They have to get through referees and editors. It's such a thing to get through that process. If I could describe to you the sweat and toil and the brain power that goes into these things, I thought there would be something there.
B
Benjamin Felix40:52
So what does that tell us about the value of the theoretical stories?
A
Andrew Chen41:10
It tells us that you could just use accounting ratios instead. I keep going back to this example of 1963 to 1989 data. You find book to market predicts returns using that sample. What happens if instead you use that same sample, you're kind of living in 1989, and you just take one accounting variable and divide by another one, and then keep repeating that until you search through thousands of accounting variables? You construct this trading strategy that has a t-stat bigger than two. Is that statistically significant? You'd get almost the same returns as from trading off stuff that's published in the literature. One of the responses we get to that is the hypothesis that the publication process really matters. When people publish these ideas, then they get traded away. But the publication dates are all around 2000, so it's all mixed up with this technology story.
B
Benjamin Felix42:17
I think I remember reading a tweet from you about data mining the accounting ratios. I follow you on Twitter and I remember reading that and being quite concerned. It's very fascinating.
A
Andrew Chen42:32
Yeah, it's really weird. It's something that makes me slightly emotional. So much work goes into these things. I think partly what's happening here is that people love these theories so much, and that biases them towards wanting to make these theories work. You think that they'd have some incremental impact after how hard it is to run these things. I used to wake up in the middle of the night to make sure that my software for calculating the equilibrium was converging, because it would take days for the computer to solve it if you don't have a fancy computer. In the end, I wonder if it has not been the right track.
B
Benjamin Felix43:44
Wow. So you talked about some of these biases in the publication process. What do you think this tells us about the academic peer review process in finance?
A
Andrew Chen43:52
I think it tells us that the peer review process is good at verifying numbers but not so good at verifying text. It's kind of like what happens if you extrapolate. If you ask an LLM something that's really in its data set, it'll get it right. If you need a gin and tonic recipe, it's going to be in there. But if you're going to make up some kind of weird cocktail, you have no idea what it'll hallucinate because that's outside of its training data. What we could say about the cross-sectional asset pricing literature as it was done from the 1980s to 2015 or so is that you could trust the numbers. That's what the replication open source paper shows. If you read a number, you could trust that that's probably the right number. But the text, whether the text actually adds any information, it's unclear. In my view, I don't think it really adds anything. As I mentioned, there's this competing hypothesis.
B
Benjamin Felix45:08
So you're saying the numbers are right but the stories might not be?
A
Andrew Chen45:10
Yes, just about this. The numbers being right. The open source asset pricing project was really about equities. It really should be a cross-sectional stock return predictability data set. I'm not sure if you've heard about this, but there's been a recent kind of scandal or crisis in corporate bonds, and now it's kind of emerging in option pricing. Option return predictability seems to have some issues. I think the option one is less widespread. I'm not an expert in this. But both of these literatures have messier data, something very different than equities. Equities data is so standardized and so clean that it's probably hard to pass off a crap number. But corporate bond data is kind of a mess. The nature of corporate bond trading is so much messier. So you don't necessarily want to extrapolate what I'm saying to other fields and to the future. I think hopefully people will read our stuff and improve the literature.
B
Benjamin Felix47:09
What role does machine learning play in the future of cross-sectional asset pricing research?
A
Andrew Chen47:13
There are two ways to think about this question. One way is to think about machine learning just as purely statistical methods for approaching research or for approaching investing. A derogatory way to describe this is data mining. What is machine learning? You just plow through a ton of data looking for patterns. That was previously really frowned upon, it was even taboo to talk about data mining. But I think our research shows that data mining was undervalued in the 80s, 90s, and 2000s. One of the striking results from the paper with Alejandro and Tom is that if you could look for the stuff with the strongest predictability in 1980, you'll find the investment anomaly wasn't published until at least 2004, so 24 years in advance of that. You also find external financing anomalies, a cash flow anomaly, inventory investment anomaly, and also the earnings surprise anomaly. Those were clearly undervalued then. If you're talking about machine learning research going forward, I think the logic of what we see with markets getting more efficient over time means that machine learning that works now will probably not work later. You'll always have to adapt.
B
Benjamin Felix49:10
So you're saying that instead of reading what kind of machine learning academics are doing in the journals, if you just did the machine learning yourself, you would do just as well?
A
Andrew Chen49:27
No, no. I'm saying that if markets are efficient, machine learning won't necessarily give you an edge in markets. But that's not the same as the academic literature. Oh, I see. To be clear, I think you'll get an edge, but it'll be fleeting. You have to act quickly, which in retrospect seems kind of obvious. People do talk about these behavioral biases sometimes as being some kind of fundamental part of our personalities. Sometimes people come up with stories about how as a whole we will interact in this way and they'll create this kind of predictability that is just mispricing. But the empirical evidence doesn't support any of that being very permanent so far.
B
Benjamin Felix51:10
So given everything we've just discussed, what are the key takeaways for investors and researchers?
A
Andrew Chen51:13
You're kind of asking me to extrapolate, but if I'm going to extrapolate, I think one of the lessons is that it's probably better to trust the numbers than the text. If you see some people claim some big thing, you probably want to verify that the numbers actually support that. That's been a theme in my research. The other wrinkle is that the numbers for more complex data sets might not be as reliable. You might want to check them and be skeptical. My research doesn't touch directly on that, but that's kind of a boundary. The last thing is, as we keep talking about, you have to act quickly. Don't expect it to persist even if someone says, 'Oh, this is a new risk factor.' They write down a fancy equilibrium model to say there's a reason why this should exist and it should continue to exist. I would be skeptical of that. Hopefully the profession will get to writing ideas that are more grounded and more real, reflecting the real world. We'll have to see.
B
Benjamin Felix53:09
So is data mining a problem or a solution?
A
Andrew Chen53:11
I have to keep trying to be cautious here. Data mining has a negative connotation. Data mining is something that you could do as an individual or you could do as a profession. For the most part, people are not sitting around just searching through Compustat for statistical significance and writing a paper about that. It's so tabooed. But as a collective, all these professors trying to publish papers, it could just turn out to look that way. This guy looks in this corner of the data, this guy looks in that corner, and as a whole as a community we're doing that. There's nothing negative about that in my opinion. What's negative is that you would hope that the theories would add value beyond data mining. Unfortunately, for cross-sectional stock return predictability as it was practiced from 1980 to 2015, that doesn't seem to be the case. So a long-winded answer, but basically that's plausible.
B
Benjamin Felix54:28
That's a good answer. So based on all your research on this, how would you describe the current state of cross-sectional asset pricing? Is it still useful? Is it productive?
A
Andrew Chen54:39
Cross-sectional asset pricing has evolved a lot. One reason why I've been critical of the Harvey, Liu, and Zhu paper for confusing false and insignificant and conflating many and most, but that paper really did change the literature. I think the paper with Alejandro and Tom more clearly shows that there's something really wrong here. We should add value beyond data mining. Data mining can work too, just to be clear. We should add value beyond it. After Harvey, Liu, and Zhu was published in 2016, I think the literature has tried to move beyond this kind of 'here's one predictor, let's make up a story around it and publish it.' It's changed a lot. I don't know where it's going. It's promising to see so much machine learning being done in the literature, but I would be skeptical of the claims. A lot of times I feel like random forests sometimes are even worse than OLS, but it won't be portrayed that way because then there'll be less to talk about. I guess I'm cautiously optimistic. The literature is changing, but I'm still wary of text. I hope it's rational. We all have behavioral biases, but I've become more skeptical of the text.
M
Mark McGrath56:44
No kidding. Crazy to think about it. It shatters so much of what the idea of factor investing kind of represents.
A
Andrew Chen57:12
It shatters a lot of stories people have in industry. It's true that if you really have a brilliant idea, you're not going to run around sharing it. Although getting a publication in the Journal of Finance does guarantee you a pretty nice life, so there are incentives to do that too. I still feel uncomfortable agreeing with you on this. But I think it's kind of shocking how all that effort was, all these stories, all this economics. It's not just stories. Stories makes it sound pejorative. I like this term that John Cochrane uses, 'paradigms.' But we make too many of them. In the big picture, we've been talking this whole time about the Adaptive Markets Hypothesis parable. There's this tendency for markets to be efficient. If there's predictability, people will buy the underpriced stocks, sell the overpriced ones, and trade it away. This is a parable, and I think it's an overwhelmingly powerful parable that's still underappreciated. We make a lot of parables, and maybe a lot of these underlying the factors are not so helpful.
B
Benjamin Felix58:49
Incredible. All right, our final question for you, Andrew. How do you define success in your own life?
A
Andrew Chen59:11
I define success as being happy about what I've done.
B
Benjamin Felix59:14
That's a great answer. I do the same thing. All right, Andrew, this has been a fantastic conversation. We really appreciate you coming on the podcast.
A
Andrew Chen59:23
Thanks. It was a pleasure to be here. I think you guys have a great program and I love how you get research out to the public. Happy to contribute.
B
Benjamin Felix59:32
Awesome. Thanks, Andrew.