Correlation Is a Fair-Weather Friend

10 min read

Take two markets whose true correlation is exactly 0.50 and never changes. Split the months into a calm half and a volatile half, sorted by the size of one market's moves. The calm half measures 0.21. The volatile half measures 0.62. Nothing about the underlying relationship moved. Only the sample did.

That arithmetic opens François Longin and Bruno Solnik's 2001 study of extreme correlation in international equity markets, and it's the cleanest available warning about what a correlation matrix really is. It isn't a property of the assets. It's an estimate, drawn from one stretch of history, carrying an error bar that nobody prints next to it.

Two separate problems follow from that. The first is statistical: correlations estimated from a finite sample are noisy, and the noise compounds as you add assets. The second is economic: the true correlation isn't fixed either. It shifts with the macro regime, and the shift can run all the way to a change of sign.

A matrix is mostly parameters you never look at

A five-asset portfolio has ten distinct pairwise correlations. A twenty-asset portfolio has 190. A fifty-asset portfolio has 1,225. The count grows with the square of the number of holdings, while the data available to estimate each one stays exactly as long as your price history.

Nobody inspects 190 numbers. What people inspect is a portfolio volatility figure, or a risk contribution chart, and every one of those outputs is a weighted blend of parameters that were never checked individually. A handful of them are wrong by a lot. You don't know which.

The direction of the error isn't random once the numbers reach an optimiser. Olivier Ledoit and Michael Wolf put it bluntly in their 2003 paper on covariance shrinkage: extreme coefficients take extreme values not because that's the truth, but because they carry an extreme amount of error, and the optimiser then places its biggest bets on exactly those. Richard Michaud named the phenomenon error maximisation. A calm estimation sample makes it worse, because calm samples produce low measured correlations, and low correlations are what an optimiser rewards.

The estimation problem grows faster than the portfolio

Ledoit and Wolf ran a controlled experiment on this. They simulated a skilled active manager, calibrated so an unconstrained information ratio of about 1.5 was theoretically available, then fed the optimiser a covariance matrix built from the last 60 monthly returns. Out-of-sample results run from February 1983 to December 2002.

With a 30-stock benchmark, the sample covariance matrix delivered a realised information ratio of 0.97. At 50 stocks it fell to 0.79. At 100 stocks, 0.59. At 225, 0.37. At 500 stocks it was 0.20, and the standard deviation of excess return had climbed from 2.26% to 8.53%. Same manager, same skill, same 60 months of data. The only thing that changed was how many correlations the risk model had to guess.

Their shrinkage estimator improved every one of those scenarios, but it didn't repair the pattern: the shrinkage information ratio also fell, from 1.24 at 30 stocks to 0.30 at 500. Worth flagging that this is a simulation with manufactured return forecasts, not a live track record. The degradation with asset count is the finding; the absolute ratios are an artefact of how the experiment was calibrated.

The blunter version of the same result comes from Victor DeMiguel, Lorenzo Garlappi and Raman Uppal. Testing fourteen optimisation models across seven datasets, they found none beat a naive equal-weight rule consistently on Sharpe ratio, certainty-equivalent return or turnover. Their published abstract puts a number on why: to reliably beat equal weights, a sample-based mean-variance strategy would need an estimation window of roughly 3,000 months for 25 assets, and about 6,000 months for 50. That's 250 years and 500 years. Nobody has that, and even if the data existed the parameters wouldn't have held still across it.

This is one reason a simple sleeve structure often survives contact with reality better than a finely tuned one, a point that also shows up in how far a 60/40 portfolio drifts across a single year.

Equity and bond correlation has already changed sign once

The estimation problem would be tolerable if the true parameters sat still. They don't.

John Campbell, Carolin Pflueger and Luis Viceira document the clearest case. Using quarterly log excess returns on five-year US Treasuries against US equities, the empirical bond-stock correlation was +0.21 over 1979Q3 to 2001Q1 and −0.64 over 2001Q2 to 2011Q4. The regression beta of bond returns on stock returns went from +0.11 to −0.19. A formal break test on daily returns puts the break at 6 December 2000, significant at the 95% level.

Their explanation is macroeconomic rather than financial. Nominal bond returns fall when inflation rises. Equity returns rise with the output gap. So the sign of the bond-stock correlation should track the sign of the inflation-output gap correlation, inverted. That's what the data show: the inflation-output gap correlation was −0.28 before the break and +0.65 after. Before 2001 the US economy was in a stagflationary pattern, where inflation rose when output was weak. After it, inflation rose in expansions.

The mechanism is symmetric, which means it runs backwards too. Marco Lombardi and Vladyslav Sushko, writing in the December 2023 BIS Quarterly Review, date the sign switch back to positive at mid-2021 for US equities and government bonds, measured as monthly realised correlations of daily returns. Their regressions show the coefficient on inflation surprises turning positive and statistically significant at that point, while growth-news coefficients became insignificant. The last comparably prolonged positive-correlation stretch was the 1980s and early 1990s.

So an investor estimating a stock-bond correlation from the fifteen years to 2020 would have measured a reliable negative number, from a sample that happened to cover a single, unusually stable inflation regime. The estimate was accurate. It was also about to stop describing anything. That distinction matters more than the size of any confidence interval, and it's a different question from which assets actually held up in past crises.

A published forecast, and what happened to it

It's worth watching a specific prediction get tested, because it shows how the regime problem behaves in practice.

In September 2021, a Vanguard team of Boyu Wu, Beatrice Yeo, Kevin DiCiurcio and Qian Wang published a machine-learning study of the stock-bond correlation. Their conclusion was that a return to the pre-2000s positive-correlation regime was unlikely. The reasoning was carefully quantified. Breaking the regime, on their five-factor model, needed ten-year trailing inflation of around 3% sustained over five years, which in turn required annual core inflation of at least 5.7% across that period. The pre-2000 positive regimes had run with ten-year trailing inflation near 5.3%. Their baseline had the 24-month rolling correlation at about −0.27 five years out.

The BIS dates the actual sign switch to mid-2021, the same quarter the paper appeared. US core inflation did exceed 5% within the following year.

That's not a criticism of the modelling, which was more transparent than most. It's the point of the article. A correlation regime model conditioned on the past thirty years assigned low probability to a macro path that then occurred within twelve months. The Vanguard paper's own scenario table is the useful part: under a 0.25 correlation regime, a 60% global equity / 40% Treasury portfolio showed median volatility of 10.30% against 9.60% in the baseline, a median Sharpe ratio of 0.27 against 0.29, and an expected maximum drawdown of −13.10% against −11.70%. The 95th-percentile drawdown widened from −25.50% to −27.90%.

Read that as the size of the prize. A full correlation regime change moved modelled 60/40 volatility by about 70 basis points and the tail drawdown by roughly two and a half points. Real, but not the end of diversification, and much smaller than the gap between measured volatility and the drawdown a portfolio actually delivers.

The counter-argument: crisis correlation is partly a measurement artefact

Here's the strongest objection to everything above, and it deserves a fair hearing.

The claim that correlations spike in crises rests on comparing a correlation measured in a turbulent window against one measured in a calm window. Longin and Solnik's opening example already showed why that comparison is broken. Conditioning on large returns raises the measured correlation even when the true one is fixed.

Kristin Forbes and Roberto Rigobon built the correction and applied it. Their test compares cross-market correlations in stable and turmoil periods, then adjusts for the fact that the turmoil correlation is conditioned on higher volatility. Applied to the Hong Kong crash of October 1997, the unadjusted numbers look damning. Average cross-market correlation was 0.20 in the stable period and 0.53 during turmoil. Hong Kong against the Netherlands jumped from 0.35 over the full sample to 0.74 in turmoil. Hong Kong against Belgium went from 0.14 to 0.71. Fifteen countries showed a statistically significant increase, which the standard reading calls contagion.

After the adjustment, the turmoil average falls to 0.32 while the stable-period average rises slightly to 0.22. The Netherlands pair moves from 0.35 to 0.40 rather than to 0.74. Contagion survives in exactly one country, Italy. Forbes and Rigobon reached the same verdict for the 1994 Mexican peso collapse and the 1987 US crash: high co-movement in a crisis was a continuation of existing linkages, not a break. Their title is the summary. No contagion, only interdependence.

If that's right, a good part of the folklore about diversification failing when you need it is an artefact of measuring correlation the wrong way. Markets that look loosely linked in calm periods were always more tightly linked than the calm-sample estimate suggested. The correlation didn't rise. The estimate was simply too low to begin with, which is a different diagnosis with different implications for how much of a portfolio's equity sits outside the home market.

Why the artefact argument doesn't close the case

Two things stop the measurement-artefact story from settling the question.

First, Longin and Solnik didn't stop at the warning. They built a test designed to be immune to it, using extreme value theory to model the tails of the joint distribution directly. Under multivariate normality with constant correlation, the correlation of returns beyond a threshold should fall toward zero as the threshold rises. Their data are monthly MSCI index returns for the US, UK, France, Germany and Japan from January 1959 to December 1996, 456 observations.

Positive tails behaved as normality predicts. Negative tails did not. Using optimal thresholds, the average correlation between negative return exceedances was 0.505 against 0.124 for positive exceedances. For the US-UK pair the figures were 0.578 and 0.226, a t-statistic of 2.066. Normality was rejected for high negative thresholds in every country pair at the 5% level. Their conclusion is precise, and it isn't the folklore version: correlation isn't related to volatility as such, but to market direction. It rises in bear markets and not in bull markets.

Second, the Forbes-Rigobon adjustment has a published rebuttal. Giancarlo Corsetti, Marcello Pericoli and Massimo Sbracia argue that the no-contagion result depends on an implicit and unrealistic restriction on the variance of country-specific shocks in the crisis country. Their generalised test, applied to the same Hong Kong episode, finds evidence of contagion in 5 countries out of 17 for plausible variance values. I read the Bank of Italy working paper version rather than the 2005 journal article, and parts of that PDF extract poorly, so I'm reporting the count from the published abstract rather than from a table I read directly.

So the honest position sits between the two camps. Naive crisis-versus-calm comparisons overstate how much correlation rises, sometimes by a lot. But the asymmetry doesn't vanish when you measure it properly. It just gets smaller and more specific: a downside-tail phenomenon, not a volatility phenomenon.

What would change the conclusion

Three things would move me.

The first is a long stretch of high, unstable inflation with a persistently negative stock-bond correlation. Campbell, Pflueger and Viceira's mechanism predicts that shouldn't happen, and their model was fitted only on macroeconomic moments, which makes the asset-pricing fit a genuine out-of-sample check. If the sign held negative through a real inflation shock, the macro story would be in trouble.

The second is out-of-sample evidence that an optimiser fed a shrunk or factor-based covariance matrix reliably beats equal weighting on live money at realistic asset counts. DeMiguel, Garlappi and Uppal tested through 2009 data, and estimation methods have improved since. Their result is about a specific class of models, not a theorem.

The third is a replication of Longin and Solnik on post-1996 data. Their sample ends in December 1996, so it covers neither 2008 nor 2020 nor 2022. Five developed markets over 456 monthly observations is not a large tail sample. If the bear-market asymmetry weakened after correcting for conditioning bias in modern data, the counter-argument would be stronger than I've credited it.

What none of this supports is treating a single correlation number as a fact about two assets. The estimate depends on the window, the frequency, the tail you condition on, and the macro regime the window happened to cover. A calm sample gives you a confident, precise number, and confidence is the part that doesn't travel.

Sources

  1. Longin and Solnik, 'Extreme Correlation of International Equity Markets', Journal of Finance 56(2), April 2001 (the 0.50/0.21/0.62 conditioning example; monthly MSCI data for five markets, January 1959 to December 1996, 456 observations; negative-tail exceedance correlation 0.505 versus 0.124 positive; US-UK 0.578 versus 0.226, t = 2.066; normality rejected for high negative thresholds) (solnik.people.ust.hk)
  2. Forbes and Rigobon, 'No Contagion, Only Interdependence: Measuring Stock Market Co-movements', NBER Working Paper 7267, July 1999 (heteroskedasticity bias in crisis correlations; Hong Kong 1997 average unadjusted 0.20 calm and 0.53 turmoil versus adjusted 0.22 and 0.32; Netherlands 0.35 to 0.74 unadjusted and 0.35 to 0.40 adjusted; Belgium 0.14 to 0.71; contagion in 15 countries unadjusted and 1 adjusted) (nber.org)
  3. Campbell, Pflueger and Viceira, 'Macroeconomic Drivers of Bond and Equity Risks', August 2018 draft of the Journal of Political Economy 128(8), 2020 article (Table 5 empirical bond-stock correlation +0.21 in 1979Q3-2001Q1 and -0.64 in 2001Q2-2011Q4; beta +0.11 and -0.19; inflation-output gap correlation -0.28 and +0.65; QLR break date 6 December 2000) (economics.nd.edu)
  4. Ledoit and Wolf, 'Honey, I Shrunk the Sample Covariance Matrix', November 2003 working paper, published in the Journal of Portfolio Management 30(4), 2004 (Table 2 information ratios of 0.97, 0.79, 0.59, 0.37 and 0.20 for the sample covariance matrix at N = 30 to 500 with T = 60 monthly returns; excess-return standard deviation 2.26 to 8.53; error maximisation; out-of-sample period February 1983 to December 2002) (econ.uzh.ch)
  5. Wu, Yeo, DiCiurcio and Wang, 'The stock/bond correlation: Increasing amid inflation, but not a regime change', Vanguard Research, September 2021 (3% ten-year trailing inflation threshold requiring 5.7% minimum annual core inflation; pre-2000 regimes near 5.3%; baseline correlation forecast of -0.27; 60/40 median volatility 10.30% versus 9.60%, Sharpe 0.27 versus 0.29, expected maximum drawdown -13.10% versus -11.70%, 95th-percentile drawdown -27.90% versus -25.50%) (nl.vanguard)
  6. Lombardi and Sushko, 'The correlation of equity and bond returns', box in the BIS Quarterly Review, 4 December 2023 (US equity-government bond correlation switching sign in mid-2021 on monthly realised correlations of daily returns; inflation-surprise coefficient turning positive and significant; last prolonged positive regime in the 1980s and early 1990s) (bis.org)
  7. DeMiguel, Garlappi and Uppal, 'Optimal Versus Naive Diversification: How Inefficient is the 1/N Portfolio Strategy?', Review of Financial Studies 22(5), 2009, pages 1915-1953 (published abstract: 14 models across 7 datasets, none consistently beating 1/N; estimation window of about 3,000 months for 25 assets and 6,000 months for 50 assets) (academic.oup.com)
  8. Corsetti, Pericoli and Sbracia, 'Correlation Analysis of Financial Contagion: What One Should Know before Running a Test', Banca d'Italia Temi di discussione 408, June 2001, published as 'Some contagion, some interdependence' in the Journal of International Money and Finance 24(8), 2005 (rebuttal of the Forbes-Rigobon restriction on country-specific shock variance; contagion in 5 of 17 countries per the published abstract) (bancaditalia.it)

Research Disclosure

This content is for informational purposes only and does not constitute financial advice. Always do your own research or consult a qualified financial advisor before making investment decisions.

Published . Data can revise after publication, so validate critical figures at source before making allocation changes.