Key takeaways
- Portfolio tracking error equals satellite weight times satellite tracking error, so a 2.5% budget bought 31.4% in NASDAQ-100, 12.2% in gold or 3.5% in bitcoin.
- A 10% bitcoin sleeve supplied 28.1% of portfolio variance over the ten years to July 2026, and half of it at a 17.2% weight.
- Gold at a 10% weight supplied only 2.2% of variance yet still left the portfolio 4.5% behind a pure core at the worst point.
- Over ten years to June 2025, Morningstar found 21% of active strategies survived and beat their average passive peer, and 15% of US stock funds.
- Cremers and Petajisto found tracking error alone doesn't predict fund returns, while the top Active Share funds beat benchmarks by 1.49% to 1.59% after fees.
Two people you could easily be. Same money, same plain S&P 500 core, and both want a satellite beside it — something with a bit more in it than the index. One picks gold. The other picks bitcoin. Both settle on 10%, because 10% of the money sounds like 10% of the risk.
It isn't. Over the ten years to July 2026, that 10% bitcoin sleeve held beside a 90% core supplied 28.1% of the combined portfolio's variance. The gold sleeve at the same 10% weight supplied 2.2%. Same weight, same core, a thirteen-fold gap in what the sleeve did to the portfolio it sat in.
So gold was the safe choice?
That's the trap. The gold sleeve looked close to free on a variance test and still left its owner 4.5% behind a pure core at the worst point. The number that saw that coming wasn't the variance share. It was tracking error, and it's the one almost nobody checks before deciding how big a sleeve should be.
One identity does most of the core-satellite portfolio sizing
Amenc, Malaise and Martellini set out the arithmetic in the Journal of Portfolio Management in 2004. Write the portfolio as P = wS + (1 − w)C, where w is the satellite weight, S the satellite and C the core. Measure both against a benchmark B. Then P − B = w(S − B) + (1 − w)(C − B).
If the core replicates the benchmark, the second term vanishes and P − B = w(S − B). Take the standard deviation of both sides and TE(P) = w × TE(S). Your portfolio's tracking error is the satellite's tracking error multiplied by its weight. Nothing else enters.
That linearity is what makes the thing usable at a kitchen table. Fix how far behind a plain core you're willing to fall, and the largest sensible satellite weight falls straight out of it: w = budget divided by TE(S). The authors' own worked example uses a satellite running 5% tracking error with an information ratio of 0.5. At a relative risk-aversion coefficient of 0.2, their formula returns an optimal satellite weight of 25% and a portfolio tracking error of 1.25%.
Their Exhibit 1 makes a cost point alongside it, and it's worth knowing before you pay anyone. Say you've settled on a 5% tracking-error budget. You can hire one manager to run the whole lot at 5% tracking error for 40 basis points. Or you hold 75% in an ETF core at 16 basis points and put 25% in a satellite running 20% tracking error at 40 basis points. Both land on 5% at portfolio level. The second costs 22 basis points.
What three satellites actually did
To put numbers on the identity I took month-end levels for four series over the ten years to July 2026. The S&P 500 stands in as the core. The NASDAQ-100, London gold and bitcoin stand in as candidate satellites. If you want to check any of this yourself, here's exactly what it's built from: index levels come from FRED, bitcoin from FRED's Coinbase series, and gold from the LBMA afternoon auction in dollars. That yields 119 monthly returns.
One caveat before the figures, because it flatters every satellite here. All four series are price-only. FRED's index levels exclude dividends, and neither gold nor bitcoin pays any income at all. The S&P 500 pays a dividend, so the core's total return was higher than the 13.3% shown here, and every satellite's excess return is flattered by however much that yield came to. Tracking error itself is barely affected, because a smooth dividend stream adds almost nothing to the volatility of a return difference.
The core ran at 15.4% annualised volatility and returned 13.3% a year on price alone. Here's how the three candidates measured up beside it.
| Satellite | Return a year | Volatility | Correlation with core | Tracking error |
|---|---|---|---|---|
| NASDAQ-100 | 19.7% | 19.1% | 0.92 | 8.0% |
| Gold | 12.0% | 15.1% | 0.10 | 20.5% |
| Bitcoin | 60.6% | 74.2% | 0.32 | 70.8% |
Now put gold and the core side by side, the way you would on a fact sheet. Their volatilities are nearly identical, 15.1% against 15.4%. Yet gold's tracking error against that core is 20.5%, two and a half times the NASDAQ-100's. How does the calmer-looking asset end up the more disruptive one?
Low correlation is the answer, and it's the whole answer. An asset that moves independently of your core departs from it constantly, by construction. That's the same property behind the 45-year record underneath a 5-10% strategic gold weight, seen from the relative side rather than the absolute one.
Risk contribution and tracking error disagree
Risk contribution splits portfolio variance among the holdings. A sleeve's share is its weight times its covariance with the whole portfolio, divided by portfolio variance. The shares sum to one, so the number reads as a percentage of total risk.
Follow the bitcoin sleeve up the scale. At a 5% weight it supplied 11.9% of portfolio variance, 2.4 times its weight. At 10% it supplied 28.1%. At 17.2% it supplied exactly half. At 20% it supplied 57.0% — a fifth of the money accounting for the majority of the risk. That's the concentration effect satellite sizing rules are written to prevent, and it's real.
The NASDAQ-100 barely levered its weight at all: 5.7% of variance at a 5% weight, 11.4% at 10%. It doesn't reach half the variance until a 44.7% weight. Part of that is the 0.92 correlation, which is another way of saying a growth satellite next to a US core is largely the same portfolio twice. It's the relative-risk version of what happens when five funds end up 39% invested in the same ten companies.
Gold went the other way entirely. At a 5% weight it supplied 0.8% of variance, a sixth of its weight. At 10% it supplied 2.2%, and it pulled portfolio volatility down from 15.4% to 14.1%.
So on a variance test, gold at 10% looks close to free and bitcoin at 10% looks alarming. Tracking error tells the story backwards. At a 10% weight the gold sleeve produced 2.05% portfolio tracking error against the pure core; the NASDAQ-100 sleeve produced 0.80%. Gold is two and a half times as disruptive relative to the core while contributing a fifth as much absolute risk.
Which reading is right? Both, because they answer different questions. Variance share tells you what makes your portfolio's value move about. Tracking error tells you how far it can drift from the thing you'd otherwise have held — for most private investors, the plain core they nearly bought instead. It's the same split that separates a portfolio's standard deviation from the drawdown its owner actually feels.
How far behind the core you'd have fallen
Tracking error is a volatility, not a loss. The loss it implies is the relative drawdown: the worst your core-plus-satellite portfolio ever fell behind a pure core. I computed that on the same monthly series, rebalancing back to the target weight every month.
At a 10% weight, the NASDAQ-100 sleeve trailed the pure core by at most 2.1% and finished the decade 5.9 percentage points ahead of it. Gold trailed by at most 4.5% and finished 0.7 points ahead, which is a lot of deviation with very little to show for it. Bitcoin trailed by at most 10.5% and finished 72.6 points ahead.
Put money on it. Imagine you'd started with a hundred thousand pounds — a round number, purely as an illustration — and handed a tenth of it to bitcoin, rebalancing monthly and never flinching. A decade later you'd be roughly seventy-two thousand six hundred pounds ahead of the friend who just held the core. And at the worst moment along the way you'd have been about ten and a half thousand pounds behind that same friend, looking at a decision you made years earlier and wondering what you were thinking. How long would you sit with that?
Those relative drawdowns ran between about 1.4 and 2.6 times each portfolio's tracking error, and the multiple held steady within each satellite as the weight changed. They scale almost exactly with weight, as the identity says they must. Doubling bitcoin from 10% to 20% moved the worst relative shortfall from 10.5% to 20.3%, and the worst rolling twelve-month gap from 9.1% to 17.9%. The relationship holds in the data, not only on paper.
That's the practical content of a tracking-error budget. It isn't an abstract risk figure. It's a rough forecast of how far behind a simpler portfolio your plan can fall before you have to decide whether to keep going.
The case against the core-satellite portfolio framing
Here's the strongest objection, and it has real force. Nobody buys a satellite for its tracking error. You buy it for what it might return, and the identity that makes tracking error linear in weight makes expected excess return linear in weight too.
Amenc and his co-authors prove precisely that. When the core replicates the benchmark, the information ratio of the whole portfolio equals the information ratio of the satellite, and it doesn't depend on w for any positive weight. Sizing the satellite makes the package neither more nor less efficient. It only scales the bet. The same arithmetic holds when the satellite is small caps, where sizing a small cap tilt leaves the information ratio identical at 5%, 10% and 20%.
So the critic's case runs like this. If you've genuinely found a satellite with a positive expected information ratio, a tracking-error cap is a self-imposed limit on how much of a good thing you'll accept, dressed up as prudence. The realised figures back the complaint up. Over this decade the NASDAQ-100 sleeve delivered an information ratio of 0.80 against the core, and bitcoin delivered 0.67. A 10% bitcoin satellite beat the pure core by 72.6 percentage points across ten years, and the worst it ever trailed by was 10.5%. Most people would take that trade twice.
What the counter-argument has to survive
Three things, and the first is that a realised information ratio isn't an expected one. Bitcoin's 0.67 and the NASDAQ-100's 0.80 come from the decade in which those two assets were among the best-performing things anyone could own. The bitcoin figures start from a close of $573 at the end of August 2016, and the coin fell 83.8% peak-to-trough inside the window. Start dates do most of the work in figures like these.
The second is the odds that the satellite you pick turns out to be a good one. Morningstar's US Active/Passive Barometer for midyear 2025 covers 9,204 funds and roughly $24 trillion, about 68% of the US fund market. Over the ten years to June 2025, 21% of active strategies both survived and beat their average passive peer. For active US stock funds the figure was 15%. Median ten-year excess returns for surviving US large-cap active funds were negative in all three style categories, and the distribution skewed negative, so the penalty for a bad pick outweighed the payoff from a good one. Cost sorted the field: 27% of the cheapest quintile beat their passive average against 15% of the priciest. That barometer runs to June 2025 and says nothing about the year since.
The third is that tracking error has never predicted returns. Cremers and Petajisto, studying US equity funds from 1980 to 2003, found the highest Active Share funds beat their benchmarks by 2.00% to 2.71% a year before fees and 1.49% to 1.59% after. The lowest Active Share funds underperformed by 1.41% to 1.76% after fees. On tracking error specifically they wrote that it "by itself is not related to fund returns", and that higher tracking error, if anything, predicts slightly poorer performance. A long manager track record fares little better as a predictor once the length of the record is separated from the skill it is meant to demonstrate.
That paper's headline claim didn't survive intact, and the caveat belongs here rather than in a footnote. Frazzini, Friedman and Pomorski re-ran the same sample for the Financial Analysts Journal in 2016 and found Active Share correlates with benchmark returns rather than predicting fund returns. Within a single benchmark it was as likely to correlate negatively with performance as positively. Their conclusion was that Active Share doesn't work as a manager-selection tool. The rebuttal targets Active Share rather than the tracking-error result I'm leaning on, but the episode is a fair warning about how durable any of these cross-sectional findings are.
Then there's the holding-period problem, and it's plain arithmetic rather than a finding. The t-statistic on a satellite's outperformance is about its information ratio times the square root of the years observed. An information ratio of 0.5 needs 16 years before the excess return sits two standard errors from zero. Even the NASDAQ-100's realised 0.80 over ten years only reaches a t of 2.53. You have to hold the sleeve through all of that to collect, which is where relative drawdown starts to matter more than expected return.
What would change the conclusion
Four things, roughly in order of how likely they are.
The first is your reference point. All of this assumes the core is what you'd judge the portfolio against. Suppose instead your yardstick is a spending target — a number you need by a date, not an index. Then tracking error is close to meaningless and variance share becomes the binding constraint. Bitcoin at 17.2% supplying half the risk is the figure that matters, and gold at 10% supplying 2.2% really is nearly free.
The second is instability in the inputs. Bitcoin's 70.8% tracking error is a ten-year average, and it isn't stable. Its annualised volatility ran at 86.3% over the first five years of the window and 56.1% over the second five. If a satellite's tracking error halves, the weight a fixed budget permits doubles. Sizing rules built on a single full-period estimate inherit every bit of that estimation error.
The third is correlation. Gold's 0.10 correlation with the core did the whole job of keeping its variance contribution near zero. Correlations move, and they tend to move at the worst moments. A gold sleeve that started behaving like an equity sleeve would keep its 20.5% tracking error and lose its diversification credit. Measured over 655 months, the gold equity correlation did rise in the worst 5% of equity months, from 0.01 to 0.31, though gold's average return in those months stayed positive.
The fourth is how you run the sleeve. These figures assume monthly rebalancing back to target. Leave a satellite alone and it compounds its weight upward. An untouched 10% bitcoin sleeve would have reached 88.3% of the portfolio at its peak and ended the decade at 78.0%, dragging its variance share and its tracking error up with it. Which would you rather live with — trimming a winner every single month, or waking up one morning inside what is really a bitcoin fund? That trade-off is the subject of annual rebalancing measured against 5% threshold bands. Fees and capital gains tax on the trimming aren't in any of these numbers either.
What survives all four caveats is the shape of the thing. A satellite's contribution to relative risk is its weight times its tracking error, and that product sets how far behind the plan your portfolio can fall. On this decade's numbers, a 2.5% tracking-error budget bought a 31.4% NASDAQ-100 sleeve, a 12.2% gold sleeve, or a 3.5% bitcoin sleeve. Those are wildly different weights for identical relative risk, and the weight number on its own would never have told you that. Run the same logic across every sleeve rather than one satellite and you arrive at the lopsided weights of a risk parity ETF, where the calmest holding takes the largest cheque.