Key takeaways
- Fama and French's 1992 paper reported a 1.53% per month gap between its highest and lowest book-to-market portfolios, measured over July 1963 to December 1990.
- Their HML factor compounded at 4.84% a year across that same window, and at 1.50% a year over the 34 years since the paper appeared.
- HML fell 57.8% from its December 2006 peak to its September 2020 trough, and by June 2026 it was still 30.6% below that peak.
- McLean and Pontiff found 97 published predictors returned 26% less out of sample and 58% less after publication. HML's own post-publication decline was 68%.
- Developed markets outside the US compounded HML at 4.70% a year from July 1990 to June 2026, with a t-statistic of 3.65 on the monthly mean.
Out of sample the premium is positive, and about a third of its in-sample size
You want to know whether value investing still pays now that everyone knows about it. The data to answer that comes from Fama and French themselves, who publish the factor series free.
HML — the return on a portfolio of high book-to-market stocks minus low book-to-market ones, where book-to-market is accounting net worth divided by stock market value — compounded at 4.84% a year over the July 1963 to December 1990 window their 1992 paper studied. Over the 34 years since the paper appeared, it has compounded at 1.50% a year. That's a decline of 68%.
It's also still a positive number, and that is where the argument starts. Across the whole series, July 1926 to June 2026, HML compounded at 3.59% a year, and the t-statistic on its monthly mean is 3.46. Over the post-publication stretch alone the same t-statistic is 1.09. A t-statistic near 1 means the average is small next to how much it moves around month to month. You can't tell that number apart from zero. You also can't tell it apart from the old number. Both of those things are true at once, and most of the disagreement about value investing lives in that gap.
What the 1992 paper actually claimed
Fama and French sorted US stocks into twelve portfolios by book-to-market equity and looked at what each earned. Their finding, in their words: "Average returns rise from 0.30% for the lowest BE/ME portfolio to 1.83% for the highest, a difference of 1.53% per month." The sample ran from July 1963 to December 1990, on NYSE, AMEX and NASDAQ stocks.
They were careful about what it meant. The paper offers a risk reading — "It is possible that the risk captured by BE/ME is the relative distress factor of Chan and Chen (1991)" — and treats it as a hypothesis, not a result. High book-to-market firms are the ones the market has marked down. If they're fragile, a higher expected return is payment for holding fragile things.
The rival reading arrived almost immediately. Lakonishok, Shleifer and Vishny circulated their contrarian-investment paper in 1993 and argued the opposite: value works because investors over-extrapolate past growth, not because value stocks are riskier. Over a sample the paper describes as running "from the end of April, 1963, to the end of April, 1990", they measured a size-adjusted gap of 7.8% a year between value and glamour portfolios, and dismissed the risk story on the numbers. Their value portfolios carried slightly higher market betas, which they judged could "explain the difference of returns of perhaps up to 1 percent per year, and surely not 8 percent that we find."
The premium by decade, computed from the series itself
The chart above compounds the monthly HML factor within each calendar decade. The source is the Fama/French three-factor file from Kenneth French's data library, built from the June 2026 CRSP database, so nothing here is an estimate of the series — it is the series.
Reading left to right: 0.99% a year in the 1930s, 9.65% in the 1940s, 3.30% in the 1950s, 3.28% in the 1960s, 7.67% in the 1970s, 5.61% in the 1980s, -0.35% in the 1990s, 7.38% in the 2000s, -2.63% in the 2010s, and 1.33% a year so far in the 2020s.
Two decades are negative, and both of them run mostly after 1992. That's the fact anyone arguing either side has to work with, and it's thinner evidence than it looks — ten decades is ten observations, and the decade boundaries are an accident of the calendar.
The drawdown that started the obituaries
The decade averages hide the shape of it. Run HML as a cumulative index and it peaked in December 2006, then lost 57.8% over the following 13 years and 9 months, bottoming in September 2020. That is a drawdown deep enough to end most institutional mandates twice over. From January 2007 to December 2020 the factor compounded at -5.67% a year.
Arnott, Harvey, Kalesnik and Linnainmaa measured the same episode in the Financial Analysts Journal and put it at a "drawdown of 55% as of mid-2020", which they noted was "the largest drawdown observed since June 1963". The small difference from 57.8% is the endpoint: they stopped at June 2020, the trough came three months later.
Since then it has recovered, without recovering. From January 2021 to June 2026 HML compounded at 8.57% a year, a cumulative 57.2%. And as of June 2026 the index was still 30.6% below its December 2006 level. Almost two decades on, and the money is not back. If you have ever wondered why a modest-looking annual number is hard to hold, that's the answer — the arithmetic of a long drawdown versus a volatility number is not the same experience at all.
The case that publication competed the premium away
The strongest version of this argument isn't about value specifically. McLean and Pontiff studied 97 variables that academic papers had shown to predict returns, and tracked what happened to each after its paper came out. Their result: "Portfolio returns are 26% lower out-of-sample and 58% lower post-publication." They read the 26% as an upper bound on data-mining bias, which leaves "a 32% (58% - 26%) lower return from publication-informed trading".
That 58% is the average across 97 predictors. HML's own decline, from 4.68% a year before June 1992 to 1.50% after, is 68% — worse than average. McLean and Pontiff also found decay was steepest for the predictors with the biggest in-sample returns, which is what you'd expect if capital chases the loudest results first. One caveat belongs with that number: "Our sample ends in 2013." The 58% is measured on returns through 2013, so it takes in the opening years of value's drawdown and none of the rest of it.
Lakonishok, Shleifer and Vishny had already named the mechanism, in 1993, without knowing they were describing their own future. Asking why the premium had lasted, they wrote: "how can the 7-8% per year in extra returns on value stocks have persisted for so long? One possible explanation is that investors simply did not know about them."
Fama and French put the same logic more sharply in 2020: "If investors do not judge that value stocks are, on some multifactor dimension, riskier than growth stocks, discovery of the value premium should lead to its demise." That is the test. Risk premiums survive publication because you still have to bear the risk. Mispricings don't, because once enough people can see them, buying pressure removes them.
The counter-case: none of this is statistically settled
Here is the objection, at full strength, from the people best placed to make it.
Fama and French ran the out-of-sample test themselves, in a Chicago Booth working paper circulated in January 2020. They split July 1963 to June 2019 into two 28-year halves, breaking at June 1991. The average premium over the market for their broad value portfolio fell from 0.42% a month (t = 3.25) to 0.11% (t = 0.60). Big value fell from 0.36% (t = 2.91) to 0.05% (t = 0.24). Small value fell from 0.58% (t = 3.19) to 0.33% (t = 1.52). Large declines, on the face of it.
Then they tested whether the two halves are statistically different. The p-value is 87.1%. The six differences they measured, in their words, "are far from unusual if expected premiums do not change from the first to the second half of the sample." Their conclusion is that monthly value premiums are so volatile that 28 years of out-of-sample data cannot separate "the premium halved" from "the premium is unchanged and we got a bad draw". They are the authors of the original paper, and they do not claim their own finding has decayed.
Arnott and co-authors go further. Their decomposition attributes the 2007-2020 drawdown to the value portfolio getting cheaper relative to growth, not to value companies doing worse: "changes in the valuation spread between the growth and value portfolios explain the entire drawdown, with room to spare." The relative valuation of the value factor "falls from the top quartile of the historical distribution at the start of 2007 to the bottom percentile as of June 2020". A style that gets cheaper produces bad past returns by construction. It says nothing about the future either way.
Israel, Laursen and Richardson at AQR add the sharpest logical point against the crowding story. If everyone had piled into value, cheap stocks would be less cheap. Instead "value spreads have widened in recent years, making crowding an unlikely explanation for the recent drawdown". And they note the awkwardness of using a drawdown as evidence of arbitrage at all: "Risk-based explanations, however, explicitly allow for negative return realization, so it is difficult to reconcile large drawdowns with awareness/crowding concerns."
That last sentence deserves a moment. A risk premium that never loses money isn't a risk premium. So a 57.8% loss is exactly what the risk story predicts should occasionally happen, and exactly what the arbitrage story predicts as well. The same observation supports both explanations, which means it discriminates between neither.
Refereeing it: the two stories differ on what happens next, not on what happened
Strip the rhetoric and the two camps agree on the record. They disagree about the generating process.
The mispricing account has the better fit to the timing. Publication in 1992, then two negative decades out of the four that followed, and a decline steeper than the 58% average McLean and Pontiff measured across 97 predictors. It also has a mechanism nobody disputes: index funds, factor ETFs and quant managers made a book-to-market screen executable by anyone, which is not true of the 1970s.
The risk account has the better fit to the statistics. Fama and French's 87.1% p-value is not a rhetorical dodge — it's what happens when you divide a small mean by a large standard deviation. Value's monthly volatility is high enough that decades of data buy you very little precision. Anyone claiming the premium is dead is claiming to have resolved a question the data does not resolve.
The point on which the mispricing story looks weakest is the AQR observation about spreads. Arbitrage that removes a mispricing compresses the valuation gap. This drawdown widened it.
Outside the US, the out-of-sample record is much stronger
One test rarely put in front of retail readers: the developed-markets-ex-US series. It begins in July 1990, so almost all of it postdates the 1992 paper, and none of it is the sample Fama and French mined.
From July 1990 to June 2026, ex-US HML compounded at 4.70% a year, with a t-statistic of 3.65. That's a stronger out-of-sample result than the US full-sample number, let alone the US post-publication one. Through the worst of the US drawdown, January 2007 to December 2020, ex-US HML lost only 1.73% a year. Since January 2021 it has compounded at 15.39% a year, t-statistic 3.52.
Both camps can use this. If publication killed the premium, why is it alive in Tokyo and Frankfurt? But equally: US equities are the deepest, most arbitraged market on earth, so the arbitrage story predicts the decay should show up there first and hardest. That is what the data shows. This is also a live reason the split between domestic and foreign equity in a portfolio is not a rounding error — the share of your equities held outside your home market determined which of these two records you actually lived through.
What this data cannot tell you
HML is one definition of value: book equity over market equity, with the portfolios re-formed once a year. Arnott and co-authors argue this definition is now broken, because book value "fails to capture increasingly important intangible assets". A software firm that expenses its research looks expensive on book value and may not be. Every number in this piece inherits that measurement choice.
HML is also a paper portfolio. It is long and short in equal size, rebalanced on a schedule, with no trading costs, no borrow fees and no taxes. A real value fund pays all four. The premium you can capture is smaller than the premium in the series, and the gap is largest in exactly the small illiquid names where the raw premium is biggest.
The t-statistics assume monthly returns are independent draws from one stable distribution. Value returns cluster — the 2007-2020 stretch was one long correlated episode, not 168 coin flips — so the effective sample is smaller than the count of months suggests. That cuts against strong claims in both directions.
And the decade table is a presentation device. Ten decades is ten numbers. Sliding the boundaries by three years changes which decades read as negative. The limitations of a backtest are not fixed by lengthening it; a longer sample of one market's history is still one history.
What would change the conclusion
A US decade at the old rate. If HML in the 2030s compounds near the 4.84% of the 1963-1990 window, the decay story loses its best evidence, and the two negative decades look like the 1930s did — a bad stretch inside a positive series.
Valuation spreads compressing while returns stay poor. That's the combination neither camp has an answer for. Cheap stocks becoming expensive relative to growth, with no premium earned along the way, is what genuine arbitrage looks like. It has not happened yet.
The ex-US series rolling over. The developed-markets result outside America is currently doing most of the work holding up the risk explanation. If it converged on the US post-1992 path, the case that publication matters would get much harder to argue against.
The number worth watching isn't the annual return. It's the valuation spread between the cheap and expensive halves of the market, because that is the only variable on which the risk story and the arbitrage story make opposite predictions. Everything else — the decade table, the drawdown, the t-statistics — both sides can already explain. Sizing a value tilt is then a question about how much tracking error you can hold through 14 bad years, which is closer to a satellite-sizing decision than a forecast.
