Key takeaways
- A true correlation of 0.50, held constant by construction, measures 0.21 in calm months and 0.62 in volatile ones purely from sample conditioning.
- US bond-stock correlation ran plus 0.21 from 1979 to 2001, then minus 0.64 to 2011, tracking an inflation-output gap flip from minus 0.28 to plus 0.65.
- With 60 months of data, a sample covariance matrix delivered an information ratio of 0.97 at 30 stocks and 0.20 at 500 stocks.
- Adjusting Hong Kong's 1997 crisis correlations for volatility conditioning cut the turmoil average from 0.53 to 0.32 and contagion cases from 15 to 1.
- Corrected with extreme value theory, negative-tail correlation across five markets averaged 0.505 against 0.124 in positive tails.
It's 2022, and you open your portfolio expecting one half to be cushioning the other. That was the deal you thought you'd signed. Equities fall, bonds rise, the mix bends instead of breaking. Instead both halves are down together, and the number that promised otherwise has quietly stopped being true.
Your diversification rests on that number — the correlation between the things you own — and you've almost certainly never checked it. It turns up looking like a fact about the assets, the way a fund's fee is a fact. Watch what happens to it.
Say you're looking at two markets whose true correlation is exactly 0.50 and never changes. Split the months into a calm half and a volatile half, sorted by the size of one market's moves. The calm half measures 0.21. The volatile half measures 0.62. Nothing about the underlying relationship moved. Only the sample did.
That arithmetic opens François Longin and Bruno Solnik's 2001 study of extreme correlation in international equity markets, and it's the cleanest available warning about what a correlation matrix really is. It isn't a property of the assets. It's an estimate, drawn from one stretch of history, carrying an error bar that nobody prints next to it.
Two separate problems follow. The first is statistical: correlations estimated from a finite sample are noisy, and the noise compounds as you add assets. The second is economic: the true correlation isn't fixed either. It shifts with the macro regime, and the shift can run all the way to a change of sign.
A matrix is mostly parameters you never look at
Count the parameters you're implicitly trusting. A five-asset portfolio has ten distinct pairwise correlations. A twenty-asset portfolio has 190. A fifty-asset portfolio has 1,225. The count grows with the square of the number of holdings, while the data available to estimate each one stays exactly as long as your price history.
Nobody inspects 190 numbers. Would you? What you inspect is a portfolio volatility figure, or a risk contribution chart, and every one of those outputs is a weighted blend of parameters nobody ever checked individually. A handful of them are wrong by a lot. You don't know which.
The direction of the error isn't random once the numbers reach an optimiser. Olivier Ledoit and Michael Wolf put it bluntly in their 2003 paper on covariance shrinkage: extreme coefficients take extreme values not because that's the truth, but because they carry an extreme amount of error, and the optimiser then places its biggest bets on exactly those. Richard Michaud named the phenomenon error maximisation. A calm estimation sample makes it worse, because calm samples produce low measured correlations, and low correlations are what an optimiser rewards.
The estimation problem behind correlation instability grows fastest
Ledoit and Wolf ran a controlled experiment on exactly this. Picture a genuinely skilled active manager — calibrated so an unconstrained information ratio of about 1.5 was theoretically available — handed a covariance matrix built from the last 60 monthly returns. Out-of-sample results run from February 1983 to December 2002.
Now watch what happens as you widen the universe that manager is picking from.
| Stocks in the benchmark | Realised information ratio, sample covariance matrix |
|---|---|
| 30 | 0.97 |
| 50 | 0.79 |
| 100 | 0.59 |
| 225 | 0.37 |
| 500 | 0.20 |
Across that range the standard deviation of excess return climbed from 2.26% to 8.53%. Same manager, same skill, same 60 months of data. The only thing that changed was how many correlations the risk model had to guess.
Their shrinkage estimator improved every one of those scenarios, but it didn't repair the pattern: the shrinkage information ratio also fell, from 1.24 at 30 stocks to 0.30 at 500. And this is a simulation with manufactured return forecasts, not a live track record. The decay with asset count is the finding; the absolute ratios are an artefact of how the experiment was calibrated.
The blunter version of the same result comes from Victor DeMiguel, Lorenzo Garlappi and Raman Uppal. Testing fourteen optimisation models across seven datasets, they found none beat a naive equal-weight rule consistently on Sharpe ratio, certainty-equivalent return or turnover. Their published abstract puts a number on why. To reliably beat equal weights, a sample-based mean-variance strategy would need an estimation window of roughly 3,000 months for 25 assets, and about 6,000 months for 50. That's 250 years and 500 years. Nobody has that, and even if you did, the parameters wouldn't have held still across it.
Which is one reason a simple structure often survives contact with reality better than a finely tuned one — the same lesson that shows up in how far a 60/40 portfolio drifts across a single year.
Equity and bond correlation has already changed sign once
All of that would be tolerable if the true parameters sat still. They don't, and the clearest evidence sits inside the pair you probably own most of.
John Campbell, Carolin Pflueger and Luis Viceira document it. Using quarterly log excess returns on five-year US Treasuries against US equities, the empirical bond-stock correlation was +0.21 over 1979Q3 to 2001Q1 and −0.64 over 2001Q2 to 2011Q4. The regression beta of bond returns on stock returns went from +0.11 to −0.19. A formal break test on daily returns puts the break at 6 December 2000, significant at the 95% level.
Their explanation is macroeconomic rather than financial. Nominal bond returns fall when inflation rises. Equity returns rise with the output gap. So the sign of the bond-stock correlation should track the sign of the inflation-output gap correlation, inverted. That's what the data show: the inflation-output gap correlation was −0.28 before the break and +0.65 after. Before 2001 the US economy was in a stagflationary pattern, where inflation rose when output was weak. After it, inflation rose in expansions.
The mechanism is symmetric, which means it runs backwards too. Marco Lombardi and Vladyslav Sushko, writing in the December 2023 BIS Quarterly Review, date the sign switch back to positive at mid-2021 for US equities and government bonds, measured as monthly realised correlations of daily returns. Their regressions show the coefficient on inflation surprises turning positive and statistically significant at that point, while growth-news coefficients became insignificant. The last comparably prolonged positive-correlation stretch was the 1980s and early 1990s. Sorting the annual record by the direction of the ten-year yield rather than by inflation leaves the stock bond correlation almost unchanged, at +0.05 in rising-rate years against -0.18 in flat ones, which points the explanation back at the inflation regime.
So which half of your portfolio is the hedge now?
Imagine you'd estimated your own stock-bond correlation from the fifteen years to 2020. You'd have measured a reliable negative number, drawn from a sample that happened to cover a single, unusually stable inflation regime. Your estimate was accurate. It was also about to stop describing anything. That distinction matters more than the size of any confidence interval, and it's a different question again from which assets actually held up in past crises.
A published forecast, and what correlation instability did to it
It's worth watching one specific prediction get tested, because it shows how the regime problem behaves when careful people take it seriously.
In September 2021, a Vanguard team of Boyu Wu, Beatrice Yeo, Kevin DiCiurcio and Qian Wang published a machine-learning study of the stock-bond correlation. Their conclusion was that a return to the pre-2000s positive-correlation regime was unlikely. The reasoning was carefully quantified. Breaking the regime, on their five-factor model, needed ten-year trailing inflation of around 3% sustained over five years, which in turn required annual core inflation of at least 5.7% across that period. The pre-2000 positive regimes had run with ten-year trailing inflation near 5.3%. Their baseline had the 24-month rolling correlation at about −0.27 five years out.
The BIS dates the actual sign switch to mid-2021, the same quarter the paper appeared. US core inflation did exceed 5% within the following year.
That isn't a criticism of the modelling, which was more transparent than most. It's the point of the article. A correlation regime model conditioned on the past thirty years assigned low probability to a macro path that then arrived within twelve months.
So what does a regime change of that size actually cost you? The Vanguard paper's own scenario table is the useful part. Under a 0.25 correlation regime, a 60% global equity / 40% Treasury portfolio showed median volatility of 10.30% against 9.60% in the baseline, a median Sharpe ratio of 0.27 against 0.29, and an expected maximum drawdown of −13.10% against −11.70%. The 95th-percentile drawdown widened from −25.50% to −27.90%. The same instability shows up when candidate sleeves are ranked on a fixed base: split the sample in half and the marginal Sharpe ratio ordering of fifteen sleeves correlates just 0.3 between the two halves.
Put that in money. Say you're running a hundred thousand pounds in that 60/40, a round number chosen purely to illustrate. The baseline expected drawdown of 11.70% is about eleven thousand seven hundred pounds. Shift to the 0.25 regime and 13.10% is roughly thirteen thousand one hundred. In the bad tail the gap widens a little more: twenty-five and a half thousand pounds becomes nearly twenty-eight. That's the price of the regime change everyone worried about. Real money, and nothing like the wipeout the phrase "diversification stopped working" tends to conjure.
Read that as the size of the prize. A full correlation regime change moved modelled 60/40 volatility by about 70 basis points and the tail drawdown by roughly two and a half points. Real, and thoroughly unpleasant in the year it lands, but not the end of your diversification — and much smaller than the gap between measured volatility and the drawdown a portfolio actually delivers.
The counter-argument: crisis correlation is partly a measurement artefact
Here's the strongest objection to everything above, and it deserves a fair hearing. What if the crisis spike is mostly an illusion?
The claim that correlations jump in a crisis rests on comparing a correlation measured in a turbulent window against one measured in a calm window. Longin and Solnik's opening example already showed why that comparison is broken. Conditioning on large returns raises the measured correlation even when the true one hasn't moved.
Kristin Forbes and Roberto Rigobon built the correction and applied it. Their test compares cross-market correlations in stable and turmoil periods, then adjusts for the fact that the turmoil correlation is conditioned on higher volatility. Applied to the Hong Kong crash of October 1997, the unadjusted numbers look damning. Average cross-market correlation was 0.20 in the stable period and 0.53 during turmoil. Hong Kong against the Netherlands jumped from 0.35 over the full sample to 0.74 in turmoil. Hong Kong against Belgium went from 0.14 to 0.71. Fifteen countries showed a statistically significant increase, which the standard reading calls contagion.
Then the adjustment goes in. The turmoil average falls to 0.32 while the stable-period average rises slightly to 0.22. The Netherlands pair moves from 0.35 to 0.40 rather than to 0.74. Contagion survives in exactly one country, Italy. Forbes and Rigobon reached the same verdict for the 1994 Mexican peso collapse and the 1987 US crash: high co-movement in a crisis was a continuation of existing linkages, not a break. Their title is the summary. No contagion, only interdependence.
If that's right, a good part of the folklore about diversification failing when you need it is an artefact of measuring correlation the wrong way. The markets you thought were loosely linked in calm periods were always more tightly linked than the calm-sample estimate suggested. The correlation didn't rise. Your estimate was simply too low to begin with — a different diagnosis, with different implications for how much of a portfolio's equity sits outside the home market.
Why the artefact argument doesn't close the correlation instability case
Two things stop the measurement-artefact story from settling the question.
First, Longin and Solnik didn't stop at the warning. They built a test designed to be immune to it, using extreme value theory to model the tails of the joint distribution directly. Under multivariate normality with constant correlation, the correlation of returns beyond a threshold should fall toward zero as the threshold rises. Their data are monthly MSCI index returns for the US, UK, France, Germany and Japan from January 1959 to December 1996, 456 observations.
Positive tails behaved as normality predicts. Negative tails did not. Using optimal thresholds, the average correlation between negative return exceedances was 0.505 against 0.124 for positive exceedances. For the US-UK pair the figures were 0.578 and 0.226, a t-statistic of 2.066. Normality was rejected for high negative thresholds in every country pair at the 5% level. Their conclusion is precise, and it isn't the folklore version: correlation isn't related to volatility as such, but to market direction. It rises in bear markets and not in bull markets. That's the version that costs you money, because a bear market is when you're deciding whether to sell. What that decision costs is arithmetic rather than judgement, and moving to cash in a crash settles the bill on the day you buy back rather than years later.
Second, the Forbes-Rigobon adjustment has a published rebuttal. Giancarlo Corsetti, Marcello Pericoli and Massimo Sbracia argue that the no-contagion result depends on an implicit and unrealistic restriction on the variance of country-specific shocks in the crisis country. Their generalised test, applied to the same Hong Kong episode, finds evidence of contagion in 5 countries out of 17 for plausible variance values. I read the Bank of Italy working paper version rather than the 2005 journal article, and parts of that PDF extract poorly, so I'm reporting the count from the published abstract rather than from a table I read myself.
So the honest position sits between the two camps. Naive crisis-versus-calm comparisons overstate how much correlation rises, sometimes by a lot. But the asymmetry doesn't vanish when you measure it properly. It gets smaller and more specific: a downside-tail phenomenon, not a volatility phenomenon.
What would change the conclusion
Three things would move me.
The first is a long stretch of high, unstable inflation with a persistently negative stock-bond correlation. Campbell, Pflueger and Viceira's mechanism predicts that shouldn't happen, and their model was fitted only on macroeconomic moments, which makes the asset-pricing fit a genuine out-of-sample check. If the sign held negative through a real inflation shock, the macro story would be in trouble.
The second is out-of-sample evidence that an optimiser fed a shrunk or factor-based covariance matrix reliably beats equal weighting on live money at realistic asset counts. DeMiguel, Garlappi and Uppal tested through 2009 data, and estimation methods have improved since. Their result is about a specific class of models, not a theorem. A rerun of that test on 10 asset menus through July 2026 leaves the headline standing, and the rule that beats naive diversification turns out to be the one that estimates no expected returns at all.
The third is a replication of Longin and Solnik on post-1996 data. Their sample ends in December 1996, so it covers neither 2008 nor 2020 nor 2022. Five developed markets over 456 monthly observations is not a large tail sample. If the bear-market asymmetry weakened after correcting for conditioning bias in modern data, the counter-argument would be stronger than I've credited it.
What none of this supports is treating a single correlation number as a fact about two assets. So when a factsheet hands you one, what has it actually told you? The estimate depends on the window, the frequency, the tail it conditions on, and the macro regime that window happened to cover. A calm sample gives you a confident, precise number, and confidence is the part that doesn't travel.