Key takeaways
- From July 1963 to June 2026, the lowest-beta tenth of US stocks compounded at 10.79% a year with volatility of 12.01%. The highest-beta tenth managed 9.53% at 28.40%. The market returned 10.87% at 15.42%.
- Return per unit of risk fell from 0.55 in the lowest-beta decile to 0.31 in the highest, against 0.47 for the market. That gap, not a return gap, is the anomaly.
- Buying the low-beta decile and shorting the high-beta decile pound for pound returned -2.18% a year, with a t-statistic of -0.70. Levering the low-beta decile to match market risk returned 12.36% a year instead of the market's 10.87%. The leverage is the strategy.
- Novy-Marx and Velikov rebuilt the betting-against-beta factor with value weights and its Sharpe ratio fell from 1.08 to 0.49, with a five-factor alpha of 24 basis points a month and a t-statistic of 1.63.
- Arnott and co-authors put the low-beta factor's return from January 1967 to September 2015 at 1.64% a year, and the part coming from rising relative valuation at 1.64%. Net of that, 0.00%.
The record: a line that should slope upward, and doesn't
Higher risk, higher return. It's the first thing anybody learns about investing, and the cross-section of US stocks has never really agreed.
Kenneth French's data library sorts US stocks every June into ten portfolios by estimated market beta — how much a stock has tended to move when the market moves. Run the value-weighted returns from July 1963 to June 2026 and the line is flat, not sloped. The lowest-beta tenth compounded at 10.79% a year. The highest-beta tenth managed 9.53%. The market did 10.87%.
Risk was not flat at all. Annualised volatility ran 12.01% in the lowest-beta decile against 28.40% in the highest, and 15.42% for the market. Average beta at portfolio formation ran from 0.23 to 2.66. So the two ends of the sort differ enormously in risk and barely at all in return.
The chart above shows what that does to the Sharpe ratio — excess return divided by volatility, or return per unit of risk. It fell from 0.55 in the lowest-beta decile to 0.31 in the highest, with the market at 0.47. The slide isn't smooth. The second decile scores 0.47 and the fourth 0.52. But the two ends are not ambiguous.
Compounding is the part a person feels. One dollar in the lowest-beta decile in July 1963 became $636 by June 2026. The same dollar in the highest-beta decile became $310. In the market it became $668. More than twice the volatility, half the money.
The CAPM alpha — the return left over once you've paid for market exposure — was 2.37% a year for the safest decile, with a t-statistic of 2.41. For the riskiest it was -2.71%, at -1.52. A t-statistic measures how big an average is next to how much it jumps around; past about three, luck stops being a comfortable story.
Sort on volatility instead of beta and the gap gets far wider
Beta and volatility aren't the same measurement. Beta captures only the part of a stock's movement that tracks the market. Volatility captures all of it. French publishes a second set of portfolios sorted on total return variance, and the result there is much more brutal.
Over the same window the least volatile tenth of US stocks compounded at 10.20% a year at 11.66% volatility. The most volatile tenth compounded at 1.52% at 31.67%. Sharpe ratios of 0.52 and 0.07. One dollar in the first became $454. One dollar in the second became $2.59.
That most volatile decile carried a CAPM alpha of -9.85% a year with a t-statistic of -4.19, and a worst peak-to-trough fall of 93.5%. Baker, Bradley and Wurgler found the same shape on an earlier window in their 2011 Financial Analysts Journal paper. A dollar in the lowest-volatility quintile in January 1968 grew to $59.55 by December 2008. In the highest-volatility quintile it was worth 58 cents.
So the anomaly isn't symmetric, and that matters for every explanation that follows. The safe end is unremarkable: it roughly matched the market for less risk. The wild end is where the damage sits.
The obvious version of the trade earns nothing. Leverage is the whole strategy.
If safe stocks beat risky ones, the trade looks easy. Long the low-beta decile, short the high-beta decile, dollar for dollar, and collect the difference.
Do that with French's deciles and you'd have earned -2.18% a year from 1963 to 2026, with a t-statistic of -0.70. Nothing at all. The reason is arithmetic rather than mystery: that portfolio is net short beta, so it's short the equity risk premium, and equities rose over the period.
Now run it the other way. Take the lowest-beta decile and lever it 1.28 times, borrowing at the Treasury bill rate, so its volatility matches the market's. It compounded at 12.36% a year against the market's 10.87%, at 15.41% volatility against 15.42%, with a worst fall of 47.3% against 50.3%. Same risk, more money, slightly smaller hole.
Nobody borrows at the Treasury bill rate, so that number is an upper bound. Charge a 1% annual spread on the borrowed portion and 12.36% becomes 12.04%. Charge 2% and it becomes 11.73%. The edge survives that on paper, before trading costs, borrow fees and tax.
Rebuild it as a proper factor — long the low-beta decile scaled up to a beta of one, short the high-beta decile scaled down to a beta of one — and it returned 5.68% a year with a t-statistic of 2.36. Real, and yet its Sharpe ratio is only 0.30, because levering a portfolio levers its volatility too.
The leverage-constraint explanation, at its strongest
This is the case Andrea Frazzini and Lasse Pedersen made in Betting Against Beta. If you want more return than the market and can't borrow, you can't lever a safe portfolio, so investors buy risky assets instead. Enough investors doing that bids up high-beta stocks and leaves low-beta stocks cheap. The security market line goes flat.
Their evidence is large. Their US factor, long levered low-beta stocks and short de-levered high-beta stocks, realised a Sharpe ratio of 0.78 between 1926 and March 2012. Across their beta-sorted portfolios, monthly excess returns barely moved — 0.91% at the safe end and 0.97% at the risky end — while volatility ran from 15.7% to 41.7%, and Sharpe ratios fell from 0.70 to 0.28. The factor delivered a positive Sharpe ratio in 18 of 19 MSCI developed countries.
Baker, Bradley and Wurgler supply the mechanism that makes leverage constraints bite hardest. Most professional money is judged against a fixed index. Citing a 2009 study by Sensoy, they report that 94.6% of US mutual fund assets are benchmarked to some popular US index. A manager maximising information ratio against that benchmark, without leverage, treats a low-beta stock as tracking error. Their arithmetic: such a manager wouldn't overweight an undervalued stock with a beta of 0.75 until its alpha exceeded 2.5% a year. Below that, underweighting it is the better trade.
The constraint is real and partly statutory. The Investment Company Act of 1940 caps mutual fund leverage at 33%, and few funds use any. In December 2008 the average equity beta of US mutual funds outside balanced funds was 1.10. The people best placed to arbitrage this away are paid not to.
Asness, Frazzini and Pedersen then closed off the obvious escape route — that this is just a bet on stodgy industries. Building the factor inside each industry, they found all 49 US industry portfolios produced a positive Sharpe ratio, and 26 had significant alphas.
The measurement objection: the famous factor is close to an equal-weighted bet on tiny stocks
Robert Novy-Marx and Mihail Velikov aimed their critique at the construction rather than the idea. Frazzini and Pedersen weight stocks by the rank of their beta, not by market value. Novy-Marx and Velikov show that this produces something almost identical to equal weighting: the two versions are 99.6% correlated month to month.
Equal weighting a US stock portfolio means loading it with very small companies. For every dollar invested, the factor holds on average $1.05 of stocks in the bottom 1% of market capitalisation. Rebuild it with value weights, which is what an investor can actually own at scale, and the Sharpe ratio drops from 1.08 to 0.49. It still earns 56 basis points a month, with a t-statistic of 3.48 — but its alpha against the Fama-French five-factor model is 24 basis points, with a t-statistic of 1.63. Not significant.
Costs make it worse. Turnover runs 214.7% a year, of which 121.1 points sits in the smallest NYSE size decile, where trading is dearest. That comes to about 60 basis points a month. Net, the original factor still earns 48 basis points a month with a t-statistic of 3.30, so it isn't an illusion. But its cost-adjusted five-factor alpha is 16 basis points, with a t-statistic of 1.20.
The reading that leaves is uncomfortable rather than fatal. Something is there. It is much smaller than the headline, and most of what remains is explained by profitability and investment exposures you could buy directly.
Where the alpha actually sits: the short leg, and unprofitable small growth stocks
Novy-Marx's separate paper on defensive equity puts a finer point on it, using the volatility sort rather than the beta sort. The long-short strategy generated a three-factor alpha of 68 basis points a month. Of that, 57 came from the aggressive stocks on the short side. Only eleven came from the defensive stocks people actually buy.
Split it by style and the concentration is stark. One dollar in a defensive strategy built inside small-cap growth stocks in 1968 became $429 by the end of 2013. The same dollar inside small-cap value became $1.22, inside large-cap growth $1.34, and inside large-cap value $0.18. His conclusion is that the effect is a way of excluding unprofitable small growth companies, which could be done directly and more cheaply. He also finds the beta-sorted version fails to produce significant abnormal returns once size and value are controlled for.
This is the same pattern the value premium and momentum literatures show, computed from the same French library by the same method: a headline number from an equal-weighted paper portfolio, and a smaller number once weights, costs and other factors are accounted for.
The crowding case: net of rising valuations, the long-run edge was zero
The other critique is that whatever was there has been bought. Rob Arnott, Noah Beck, Vitali Kalesnik and John West decomposed six factors into the part that came from fundamentals and the part that came from getting more expensive.
Low beta was their most extreme case. From January 1967 to September 2015 its return was 1.64% a year, and the return from changing relative valuation was 1.64% a year. Net of valuation change: 0.00%. Their more conservative regression-adjusted figure leaves 1.11%. Over the ten years to September 2015 the factor returned 2.67%, in the decade the products launched. Their words: "Large asset flows into low beta products are now driving valuation levels far above their historical norms."
They also report the honest weakness in their own result. The correlation between low beta's valuation and its subsequent five-year return is -0.07, with a t-statistic of -0.67. Of the six factors they test, low beta is the only one where the valuation link isn't statistically significant, which they attribute to its turnover and to the valuation series ending near a peak.
The best answer to the crowding case, and what the data since 2011 shows
David Blitz, Pim van Vliet and Guido Baltussen make the strongest rebuttal, and they make it by asking who is on the other side. Their finding: "there is little evidence that the low-risk effect is being arbitraged away because many investors are either neutrally positioned or even on the other side of the low-risk trade."
The detail matters. Active mutual funds have kept a persistent negative exposure to low-volatility stocks even as the style grew popular. Across all US-listed ETFs, low-risk and high-risk exposures roughly cancel, leaving investors in aggregate close to neutral. Hedge funds, who face no leverage constraint, have been positioned against the anomaly rather than for it. They put low-volatility ETF assets at $24 billion at the end of 2016, which is small against the US market. They're candid about one crowding cost: transparent indexes that trade on two fixed days a year, which they estimate costs about 16 basis points a year in front-running.
The French data can referee part of this. Split the beta sort at January 2011, roughly when the products arrived. Before that, the lowest-beta decile ran a Sharpe ratio of 0.45 against 0.34 for the market. From 2011 to June 2026 it ran 0.91 against 0.87. The advantage over the market shrank by about two-thirds, even though the level looks better.
The volatility sort says the opposite. The least volatile quintile ran 0.44 against the market's 0.34 before 2011, and 1.04 against 0.87 after. That advantage widened. Two sorts of the same idea, two verdicts, one sample. Anyone claiming the crowding question is settled is choosing a sort.
What holding it cost: 1999, 2020, and the years the gap ran the wrong way
Whatever the explanation, the experience of owning this is the part most write-ups skip. Low beta is a benchmark-relative torture device in a melt-up.
In 1999 the lowest-beta decile returned -3.7% while the market returned 25.2% and the highest-beta decile returned 63.6%. In 2020 it returned 2.4% against the market's 24.1% and the high-beta decile's 89.8%. Those are the two worst years of relative pain in the whole record, and both were bull markets.
The compensation arrives late and in bad weather. In 2000 the lowest-beta decile returned 25.5% while the market lost 11.6%. In 2008 it lost 27.0% against the market's 36.7% and the high-beta decile's 49.9%. In 2022 it was roughly flat at -0.7% while the market lost 19.9%. Its worst drawdown across the whole period was 38.6%, against 50.3% for the market and 77.6% for high beta — which is the difference between volatility and the loss you actually live through.
What this evidence can't tell you
Start with what the French portfolios are. They're paper portfolios, rebuilt on a rule, paying no spread, no commission, no borrow fee and no tax. The levered version above also assumes you finance at Treasury bill rates, which no individual does, and the two spread scenarios are illustrations rather than quotes from a broker.
The record is US. Frazzini and Pedersen's international results are the main evidence that this travels, and they end in March 2012. The beta and variance files begin in July 1963, so nothing here says anything about the decades before.
Almost every author here has a commercial stake, and the papers say so themselves. Frazzini and Pedersen were both at AQR Capital Management. Baker, Bradley and Wurgler wrote theirs while at or consulting for Acadian Asset Management, where Bradley was director of managed volatility strategies. Blitz, van Vliet and Baltussen all work at Robeco, where van Vliet heads conservative equities. Novy-Marx consults for Dimensional Fund Advisors. Research Affiliates scores its own Fundamental Index in the same table. That doesn't make any of their numbers wrong — the ones checked here reproduce — but it's why this piece leans on figures computed from the raw files where it can.
Finally, a backtest is not a forecast, one sample is not the future, and a Sharpe ratio computed on returns that cluster in time overstates its own confidence. The 2011 split above is a single break in a single sample, chosen because that's when the products launched. It isn't a test.
What would change the conclusion
Three things would, and each is watchable.
The first is leverage getting cheap and easy for ordinary investors. If the constraint is the cause, removing it removes the premium. Margin costs and the spread of leveraged products are the thing to watch, not fund flows.
The second is the two sorts converging. Right now the beta sort's edge over the market has narrowed since 2011 and the volatility sort's has widened. If both narrowed together for another decade, the crowding critics would have the argument, and Arnott's 0.00% net of valuation change would look like the honest number rather than the provocative one.
The third is a defensive portfolio that loses badly in a falling market. The entire case rests on low-beta stocks doing what they did in 2000, 2008 and 2022. A crash in which they fall as hard as the market would leave a strategy with the tracking error of 1999 and none of the payoff. That, rather than a bad year in a bull market, is what would settle it.