Key takeaways
- Fifteen candidate sleeves were added at 10% each to the same 60/40 over 433 months. The best changed its Sharpe ratio by 0.043, the worst by 0.066.
- Gold bullion was the only sleeve that both raised the Sharpe ratio and made the drawdown shallower, from 28.4% to 24.1%.
- Listed real estate was the most costly addition, taking 0.066 off the Sharpe ratio and 7.42 percentage points off the drawdown.
- Correlation with the base ran from 0.04 for Treasury bills to 0.95 for US growth stocks, and it ranked the sleeves only loosely.
- Ranking the sleeves separately on each half of the window gave a correlation of 0.3 between the two orderings.
One more asset class moved a 60/40 by less than a tenth of a Sharpe ratio
You hold something close to a 60/40, and someone tells you it needs gold. Or emerging markets. Or an infrastructure sleeve, a commodity sleeve, a value tilt. The question underneath all of them is the same one: how much does one more asset class actually change the portfolio you already own?
Here's the measured answer, on one consistent test. Fifteen candidate sleeves, each added at 10% to the same base portfolio, one at a time, across the 433 months from July 1990 to July 2026. The best of them raised the Sharpe ratio by 0.043. The worst cut it by 0.066. That's a spread of 0.109 between the most and least helpful thing you could bolt on, and twelve of the fifteen finished inside 0.03 of where they started.
For context, the base portfolio's Sharpe ratio was 0.688. So the entire fifteen-way argument about what to add is worth roughly a sixth of the risk-adjusted return the plain 60/40 already produced. That doesn't make the choice meaningless. It does mean the choice is smaller than the marketing around it. Measured this way, the diversification benefit of one extra asset class is real but narrow.
What was tested, and what it was tested against
The base is 60% US equities and 40% ten-year Treasuries, rebalanced every month. The equity leg is the value-weighted return of all NYSE, AMEX and NASDAQ stocks from the Kenneth French data library, which is the same series academic work uses for the US market. The bond leg is the monthly total bond return column in Robert Shiller's dataset, which runs alongside his ten-year constant-maturity yield. Over 1960 to 2026 that return series correlates minus 0.98 with the monthly change in that yield, which is what a ten-year bond position should do.
Over the window the 60/40 portfolio returned 8.98% a year with 9.36% annualised volatility, a Sharpe ratio of 0.688 measured over the one-month Treasury bill rate, and a worst month-end drawdown of 28.4%, running from the October 2007 peak to February 2009. The equity leg on its own returned 11.08% at 15.17% volatility and fell 50.3%. The Treasury leg returned 5.01% at 6.19% volatility and fell 25.2%.
Each sleeve then goes in at 10%, funded proportionally from both legs, so the test portfolio is 54% equities, 36% Treasuries and 10% sleeve. Nothing else changes. The sleeves are eleven US equity groups from the French library (size, book-to-market, profitability and prior-return sorts, plus six of the 49 industry portfolios), developed ex-US and emerging market equities from the same library, the LBMA Gold Price PM in dollars an ounce, and one-month Treasury bills as a control. Gold is quoted in dollars a troy ounce and administered by ICE Benchmark Administration; the World Bank's monthly gold series matches the mean of those daily fixings to the cent, at 362.53 dollars in July 1990.
Every figure below is computed from those monthly series. None of it is a live product, so nothing here carries a fund fee, a bid-offer spread or a tax bill, and all three of those subtract from the small numbers you're about to read.
The correlation matrix ranks the candidates, and then stops explaining them
Asset class correlation is the number every pitch leads with, so start there. Below is each sleeve's monthly correlation with the base, with the equity leg, and with the Treasury leg.
| Sleeve | With the 60/40 | With the equity leg | With the 10-year Treasury leg |
|---|---|---|---|
| Treasury bills (cash) | 0.04 | 0.01 | 0.13 |
| Gold bullion | 0.06 | 0.02 | 0.17 |
| Gold mining equities | 0.21 | 0.16 | 0.19 |
| Utilities | 0.47 | 0.46 | 0.10 |
| Oil and gas | 0.48 | 0.54 | -0.14 |
| Packaged food | 0.53 | 0.53 | 0.06 |
| Pharmaceuticals | 0.62 | 0.62 | 0.08 |
| Listed real estate | 0.67 | 0.70 | -0.02 |
| Emerging market equities | 0.68 | 0.71 | -0.05 |
| Developed ex-US equities | 0.74 | 0.77 | -0.01 |
| US small caps (smallest 30%) | 0.77 | 0.82 | -0.10 |
| US momentum (top decile) | 0.77 | 0.80 | -0.03 |
| US value (highest 30% BE/ME) | 0.81 | 0.86 | -0.09 |
| US growth (lowest 30% BE/ME) | 0.95 | 0.98 | -0.00 |
| High-profitability US equities | 0.95 | 0.97 | 0.01 |
The range is wide. Treasury bills sit at 0.04 and gold bullion at 0.06, both effectively unrelated to the base month to month. US growth stocks and high-profitability US equities sit at 0.95, which is another way of saying they are the base.
Now the part the pitch leaves out. Gold mining equities correlate 0.21 with the base, lower than every sleeve except gold and cash, and adding them still cost 0.029 of Sharpe ratio. Their own volatility was 37.59% a year and they returned 3.25%. A low correlation multiplied by a large enough volatility and a poor enough return is not a diversifier. It's a drag with a good story.
The same logic runs the other way. High-profitability US equities correlate 0.95 with the base, which by the usual rule should make them useless, and they still added 0.011 of Sharpe ratio because they returned 12.81% a year at 14.46% volatility. If you want the single sentence: correlation tells you how a sleeve moves, not whether it's worth owning. Correlation instability makes that worse, because the number you estimate from one decade is not the number you get in the next.
The full result: fifteen sleeves, ranked by what they did to the Sharpe ratio
Here is every sleeve, its own annualised return and volatility over the window, and what a 10% weight did to the base portfolio's marginal Sharpe ratio and maximum drawdown. A positive drawdown figure means the drawdown got shallower.
| Sleeve | Return a year | Volatility | Change in Sharpe | Change in drawdown (pp) |
|---|---|---|---|---|
| Gold bullion | 6.99% | 15.67% | +0.043 | +4.25 |
| Utilities | 9.74% | 14.02% | +0.028 | -0.88 |
| Pharmaceuticals | 11.79% | 15.81% | +0.027 | +0.04 |
| US momentum (top decile) | 15.39% | 21.93% | +0.018 | -2.42 |
| High-profitability US equities | 12.81% | 14.46% | +0.011 | -1.30 |
| Packaged food | 8.12% | 13.93% | +0.006 | +0.33 |
| Oil and gas | 9.94% | 22.67% | +0.005 | -1.10 |
| US value (highest 30% BE/ME) | 12.64% | 18.29% | +0.004 | -3.72 |
| Treasury bills (cash) | 2.66% | 0.63% | +0.000 | +2.70 |
| US growth (lowest 30% BE/ME) | 11.69% | 15.49% | -0.005 | -1.83 |
| US small caps (smallest 30%) | 11.08% | 21.17% | -0.017 | -3.27 |
| Gold mining equities | 3.25% | 37.59% | -0.029 | +1.28 |
| Emerging market equities | 8.23% | 20.49% | -0.030 | -3.94 |
| Developed ex-US equities | 6.33% | 16.32% | -0.040 | -3.26 |
| Listed real estate | 5.97% | 26.05% | -0.066 | -7.42 |
The chart above plots that fourth column. Read it and the story is how little happens. Read the fifth and it's less comforting. Ten of the fifteen sleeves made the worst drawdown deeper, and several of the ones that helped the Sharpe ratio were among them. US value added 0.004 of Sharpe ratio and 3.72 percentage points of extra drawdown. Momentum added 0.018 and 2.42 points of extra drawdown. If volatility versus drawdown matters to you, those two sleeves were charging you in the currency you care about.
Gold bullion was the only sleeve that raised the Sharpe ratio and cut the drawdown
Gold added the most on the risk-adjusted measure, 0.043, and it's the only candidate that did so while also making the fall shallower. The base's worst drawdown of 28.4% became 24.1%, an improvement of 4.25 percentage points. Through the crisis window from November 2007 to February 2009, the base lost 28.4% and the version with gold in it lost 24.1%.
The mechanism isn't that gold went up. Gold returned 6.99% a year over the window against the base's 8.98%, so the sleeve cost return: the blended portfolio returned 8.93%. What gold did was arrive with a 0.06 correlation and stay there, including in 2022, when the base fell 17.6%, US equities fell 19.9%, ten-year Treasuries fell 15.0%, and gold returned 0.4%. That year is the one the classic 60/40 has no answer for, because both legs fell together.
Sizing matters more than the pitch usually admits. At 5% gold the Sharpe ratio was 0.712, at 10% it was 0.731, at 20% it was 0.75, and at 30% it fell back to 0.737. The drawdown kept improving the whole way, from 26.3% at a 5% weight to 17.8% at 30%, while the annualised return slid from 8.96% to 8.73%. There's no single right weight in that table, only a trade you can price. Our separate piece on how much gold in a portfolio works through the same trade on a different sample.
Listed real estate was the most expensive thing you could add
Listed real estate is sold as a distinct asset class with an inflation link and a low correlation. Over this window it correlated 0.67 with the base, which is higher than gold, cash, gold miners, utilities and oil. It returned 5.97% a year, barely above the Treasury leg's 5.01%, with 26.05% volatility, and its own worst drawdown was 83.05%.
Adding it at 10% cost 0.066 of Sharpe ratio and pushed the base's worst drawdown from 28.4% to 35.8%. In the crisis window the version with real estate lost 35.8% against the base's 28.4%. It was the worst sleeve of the fifteen on the full window, the worst in the first half, and the worst in the second. Nothing else in the test was that consistent.
The fair objection is that listed real estate is not property. It's equity in property companies, priced daily by an equity market, and it behaves like it. That's the point of testing it as a sleeve rather than accepting the label. Two things with the same name did not do the same thing.
The diversification benefit does not hold still
Split the window in half and the ordering falls apart. Over July 1990 to June 2008, oil and gas was the best sleeve of the fifteen, worth 0.07 of Sharpe ratio on a base that scored 0.6. Over July 2008 to July 2026 the same sleeve was thirteenth, at minus 0.052, on a base that scored 0.768. Gold went the other way, from 0.026 in the first half to 0.056 in the second, as its annualised return went from 5.54% to 8.44%.
Rank the fifteen sleeves in each half and the two orderings correlate 0.3. That is close to no relationship. The one exception is listed real estate, last in both halves at minus 0.046 and minus 0.084. Everything else moved around.
This is the strongest argument against the whole exercise, and it deserves stating plainly. A measured diversification benefit is a description of one sample. It is not a property of the asset class. If you had run this study in mid-2008 and acted on the top of the table, you'd have added an oil and gas sleeve just before the eighteen years in which it ranked thirteenth.
Adding all fifteen at once bought nothing
If each sleeve helps a little, a basket of all of them should help more. It didn't. Splitting a 10% allocation equally across all fifteen sleeves produced a Sharpe ratio of 0.692 against the base's 0.688, a return of 9.13%, and a worst drawdown of 29.8%, which is deeper than the 28.4% you started with.
That result is not a paradox. Thirteen of the fifteen sleeves are equities, and the base is already 60% equities. Blending thirteen equity sleeves together reproduces something close to the equity market, so the basket mostly moved money from Treasuries into more of what the portfolio already had. The diversification benefit of a sleeve comes from a risk you don't already own, and thirteen of these owned the one the base was already full of. Diversification counts distinct risks, not distinct product names.
The equity leg owns 93.8% of the risk, which caps what a 10% sleeve can do
The arithmetic behind the flat table is worth seeing. In the base portfolio the equity leg contributes 93.8% of the monthly variance, because 60% of the money sits in something with 15.17% volatility while 40% sits in something with 6.19% volatility, and the two correlate minus 0.03.
A 10% sleeve is therefore competing against a risk budget that is already almost entirely equity risk. To move the Sharpe ratio, it has to either bring a return the base lacks or shrink that equity share. Cash does the second: a 10% cash sleeve left the Sharpe ratio unchanged at 0.688 while cutting the drawdown by 2.7 percentage points, which is exactly what happens when you scale a portfolio down without changing its shape. Gold managed to do a bit of both. Nothing else in the test managed either at any scale worth noticing.
The same reasoning explains why risk contribution is a more useful lens than weight when you're looking at a sleeve, and why a 1% to 5% bitcoin allocation is a different question again: a very high volatility sleeve moves the risk budget at weights where a normal one doesn't.
What this test cannot tell you
It's one sample, in one currency, on one base. The window is 36 years of US-dollar returns against a US-centric 60/40, and it contains a fall in the ten-year Treasury yield from 8.47% in July 1990 to 0.62% in July 2020 that flattered the Treasury leg for most of the way. A sterling investor holding global equities and gilts would get different numbers, and possibly a different ordering.
The results also carry no costs. Every sleeve here is a free index; in practice you'd pay a fund fee, a spread and, outside a tax wrapper, tax on the rebalancing. Against changes in the Sharpe ratio of 0.03, those frictions are not a rounding error.
The drawdown figures are month-end, which understates intramonth falls. And the sample predates or barely touches several sleeves people are now offered, including crypto, private credit and listed infrastructure, none of which have 433 months of independent monthly history to test on the same footing. What happened in diversification in a crisis is a useful companion here, because averages across a whole window hide what a sleeve does in the months you need it. If you want the same arithmetic on your own holdings rather than on an index, LedgerTouch computes the correlation matrix and drawdown contribution from your actual positions.
One robustness check does survive all of that. Rerun the test on the thirteen sleeves with data back to January 1973, and across 643 months the base returned 9.52% at 10.19% volatility, a Sharpe ratio of 0.521 and a 29.4% drawdown. Gold was still the best sleeve, at 0.04 of Sharpe ratio and 5.28 percentage points of shallower drawdown, on an 8.01% annualised return. Listed real estate was still the worst, at minus 0.051. Over 53 years the size of the effect barely changed.
What would change the conclusion
Three things would, and only one of them is about the sleeves.
The first is the base. This result depends on a 60/40 whose two legs correlated minus 0.03 over the window. If equities and bonds correlate positively for a decade, as they did in 2022, the base gets riskier and a genuinely uncorrelated sleeve becomes worth more than 0.043 of Sharpe ratio. The gold result strengthened in the second half of the sample for exactly that reason.
The second is the weight. Everything here is a 10% sleeve. At 30% gold the Sharpe ratio was still above the base and the drawdown was 10.6 percentage points shallower, so a reader who cares about drawdown rather than Sharpe ratio reads this table differently, and reads it correctly.
The third is cost and time. A 0.03 change in Sharpe ratio is inside the noise of a 36-year sample and inside the cost of implementing it. If a sleeve's case rests on a number that small, the honest description is that the evidence cannot separate it from the base. That's true of twelve of the fifteen tested here, and it stayed true when the window was extended to 53 years.