Key takeaways
- Over 69 test years from July 1957 to June 2026, a rule that scaled value exposure with the value spread returned 3.24% a year, against 3.99% for holding the value factor at a constant weight.
- 12 versions of that rule were tested. Three raised the raw return, by as much as 1.58 percentage points a year. None of the 12 improved return per unit of risk.
- The value spread explained 2.9% of the variation in the next year's value premium. After a wide spread the premium was positive 57.9% of the time, against 60.0% after a narrow one.
- Research Affiliates, who argue for contrarian timing, report their own factor version earning 3.3% a year against 2.4% for equal weighting, with the Sharpe ratio falling from 0.52 to 0.39.
- In June 2026 the cheap third of the US market traded at 9.90 times the book-to-market ratio of the expensive third, against a median of 4.89 since 1927 and 9.89 in June 2000.
The short answer: the signal is real, and it's too weak to trade
Can you do better by owning a factor when it's cheap and dropping it when it's expensive? That question got its own public argument in 2016, between Research Affiliates on one side and AQR on the other, and it has never really closed.
Run it on Kenneth French's own data and the answer comes out no, at least for the cleanest version of the trade. A rule that scaled exposure to the value factor up and down with the value spread returned 3.24% a year from July 1957 to June 2026. Holding the same factor at a constant weight returned 3.99%. The timed version also carried a standard deviation of 20.55 points against 13.56, so it paid less and it hurt more.
12 variants of that rule were tested, changing the length of the history used to calibrate it and the shape of the response. Three of them beat the constant weight on raw return. Not one of the 12 improved return per unit of risk. The widest gap in either direction carried a t-statistic of 0.86, which is indistinguishable from noise on 59 to 79 observations.
That doesn't make the signal empty. The correlation between the spread and the following year's premium came out at 0.17, positive and in the direction the theory predicts. It's just far too small to build a position on, and the rest of this piece is about why factor timing at that correlation behaves the way it does inside a portfolio.
What the value spread actually measures
The value spread is the gap in valuation between the cheap end of the market and the expensive end. Here it's measured with book-to-market, the ratio of a company's accounting net worth to its share price. A high book-to-market means you're paying little for each pound of book equity, which is what "cheap" means in this literature.
French's library publishes the aggregate ratio for portfolios sorted on book-to-market every June, using the 30th and 70th NYSE percentiles as the breakpoints. His documentation is explicit about the timing: "BE/ME for June of year t is the book equity for the last fiscal year end in t-1 divided by ME for December of t-1." Everything in the June reading is public by June, which is what makes an honest test possible.
The factor being timed here is French's HML, which his documentation defines as "the average return on the two value portfolios minus the average return on the two growth portfolios." It's a long-short return: what you'd have earned owning the cheap end and shorting the expensive end, with the market's own return netted out. Every return in this piece is that number, not the return on a value fund. His monthly series for it begins in July 1926, which is as far back as this kind of test can go.
In June 2026 the cheap third of the market carried an aggregate book-to-market of 0.8392. The expensive third carried 0.0848. Divide one by the other and the cheap third carried 9.90 times the book-to-market of the expensive third. The median reading since 1927 is 4.89. Only five of the past 100 Junes have been wider: 1932, 1935, 1939, 2020 and 2025.
So the spread is at an extreme, and that's the whole reason anyone asks this question now. June 2000 read 9.89, within a rounding error of today. Anyone who bought the argument in June 2000 was handsomely paid. Anyone who bought it in June 1999, when the spread was already elevated, lost 20.25% on the value factor over the following twelve months before the payoff arrived.
The test: 69 years, one rule, and nothing known in advance
The design is deliberately plain, because an elaborate rule is easier to fit to the past. At the end of each June, take the log of the value spread and convert it to a z-score using every year of history available up to that point and no further. A z-score just says how many standard deviations above or below its own past average a reading sits.
Then set next year's weight in the value factor to 1 plus that z-score, floored at 0 and capped at 2. A spread at its historical average leaves you at a normal weight. A spread two standard deviations wide doubles you up. A spread two standard deviations narrow takes you to nothing. Hold from July to the following June, then repeat. The first test year is 1957, which leaves 30 years of history to calibrate the very first reading.
Across those 69 years the constant-weight factor returned 3.99% a year with a standard deviation of 13.56, a ratio of 0.29. The timed version returned 3.24% with a standard deviation of 20.55, a ratio of 0.16. The shortfall of 0.75 percentage points a year carries a t-statistic of 0.51, so the underperformance isn't reliable either. Both numbers are consistent with a rule that does nothing except add risk.
Splitting the sample makes the instability obvious. From 1957 to 1990 the timed version lost 3.01 percentage points a year to the constant weight. From 1991 to 2025 it gained 1.45. Same rule, same signal, opposite sign, depending on which half you look at.
12 versions of the rule, and none improved return per unit of risk
One specification proves nothing, so the same test was run 12 ways: three calibration windows of 20, 30 and 40 years, and four response shapes. Those are the capped rule above, the same rule with the cap removed so the weight can run past 2 and go negative, a binary rule that holds double weight when the spread is above its own average and nothing when it's below, and a tercile rule that goes to double, normal or zero.
The results run from 1.35 percentage points a year worse than the constant weight to 1.58 percentage points better. Every version that made money did it by taking more risk, and the ratio of return to standard deviation fell in all 12. The largest absolute t-statistic in the set was 0.86. A researcher who ran only the version that came out at plus 1.58 could write a persuasive paper. That's the mechanism behind a lot of factor timing research.
Two consecutive years show what the rule feels like from inside. In June 1999 the spread was already a standard deviation wide, so the rule went to double weight and the value factor fell 20.25% over the next twelve months. Doubled up, that's a loss of 40.47%. In June 2000 the rule stayed at double weight, the factor returned 61.01%, and the timed position made 122.02%. The signal was right, a year early. Twelve months early is long enough to lose a mandate.
The worst single year in the whole test was the other way around. In June 2019 the spread sat 1.14 standard deviations wide, the rule doubled up, and the value factor lost 28.80% in the year to June 2020. The timed position lost 57.60%. The spread had been telling the truth about value being cheap for years by then. It just hadn't said when.
Factor premiums by decade, and the spread that opened each one
The table below is the whole argument in one place: what each factor paid over each decade, and how wide the value spread was in the June that opened it. All returns are annualised from the monthly Fama-French factor series, and the 2020s row runs to December 2025. Those same decade premia, priced against the measured factor loadings of five multifactor products, show how little of each decade a long-only holder could collect.
| Decade | Value spread at the start | Market minus cash | Small minus big | Value minus growth | Momentum |
|---|---|---|---|---|---|
| 1930s | 6.07 | -0.82% | 8.76% | 0.98% | -6.95% |
| 1940s | 9.24 | 9.05% | 4.53% | 9.55% | 6.40% |
| 1950s | 5.17 | 16.19% | -0.87% | 3.32% | 10.86% |
| 1960s | 5.74 | 4.22% | 4.52% | 3.23% | 11.06% |
| 1970s | 5.00 | -0.26% | 2.82% | 7.65% | 9.33% |
| 1980s | 3.57 | 7.34% | -0.39% | 5.59% | 8.61% |
| 1990s | 4.71 | 12.50% | -2.05% | -0.35% | 13.73% |
| 2000s | 9.89 | -3.07% | 4.20% | 7.38% | -2.00% |
| 2010s | 4.10 | 13.03% | -0.49% | -2.62% | 2.62% |
| 2020s | 11.03 | 11.83% | -3.10% | -0.27% | 1.06% |
The chart plots that second column on its own, one reading for the June that opened each decade plus June 2026. Read the first two columns of the table together and you get the case for timing and the case against it in the same 10 rows. The 2000s opened at 9.89 and the value premium paid 7.38% a year. The 1940s opened at 9.24 and it paid 9.55%. Both are exactly what a contrarian would predict. The size premium and the momentum premium in the next two columns move to a rhythm of their own, which is the first hint that one signal does not govern them all.
Then look at the 2020s. That decade opened at 11.03, the widest reading of any decade start in the table, and the value premium has run at -0.27% a year through December 2025. The 1980s opened at 3.57, the narrowest start in the table, and value paid 5.59%. Ten observations, two clear hits, two clear misses. That's the sample size the argument is actually working with.
The strongest case for factor timing, made by the people who make it
Rob Arnott, Noah Beck and Vitali Kalesnik at Research Affiliates published the best-known argument for this in September 2016, under the title "Timing 'Smart Beta' Strategies? Of Course! Buy Low, Sell High!" Their finding, over January 1977 to August 2016, is that a contrarian investor who bought the three cheapest of eight factors earned 3.3% a year, against 1.2% for one who chased the best recent performers and 2.4% for one who held all eight equally.
That's a large gap and it's in their favour. What matters is the second number they print next to it. The contrarian's Sharpe ratio came out at 0.39. The equal-weighted investor's was 0.52. The paper says so plainly: "although value-add is higher (3.3% versus 2.4%) compared to the equally weighted portfolio, the Sharpe ratio is lower (0.39 versus 0.52) due to lower factor diversification and higher risk." Their own summary concedes the point in advance: "it's not easy to garner a materially higher Sharpe ratio. Many would view this as an acceptable outcome; after all, we can't spend a Sharpe ratio."
That is the same result this test produced, from different data and a different rule. More return, worse risk-adjusted return. Whether that trade is worth making is a judgement about your own tolerance, not a finding in the data.
The paper also carries its own most serious limitation in a footnote to Figure 2, which reports a five-year correlation of -0.31 between valuation and subsequent return for its equally weighted smart beta strategy, with a t-statistic of -8.01. Of that January 1967 to August 2016 window, the note reads: "has just under 10 non-overlapping 5-year periods and just under 5 non-overlapping 10-year periods." A t-statistic of 8 sounds decisive. 10 independent observations do not support one.
Why timing value with the value spread is mostly just more value
Cliff Asness made the structural objection in June 2016, in a paper called "My Factor Philippic." He argued that the long-horizon regressions overstate the case, and then went further: "these strategies add little to portfolios that are already invested in the value factor. It turns out that this 'newly' discovered timing tool is, yet again, mostly just a version of regular old value investing."
The arithmetic in this test says the same thing. Regress the timed strategy's annual returns on the constant-weight factor's and the slope comes out at 1.25, with a correlation of 0.82. The intercept is -1.75 percentage points a year. Nearly everything the timing rule did was take the same bet with more size, and what little was left over was negative.
That's the trap. Value being cheap is what a wide value spread means, definitionally. So a rule that adds to value when the spread is wide isn't a new source of return sitting alongside value. It's the value trade with a variable position size, and the variation in size is where the extra volatility comes from.
AQR's follow-up in the Journal of Portfolio Management in 2017, by Asness, Swati Chandra, Antti Ilmanen and Ronen Israel, landed in the same place. Its summary reads: "At first glance, valuation-based timing of styles appears promising." Then the qualifier: "Yet when the authors implement value timing in a multi-style framework that already includes the value style, they find somewhat disappointing results." The reason they give is the one the regression above produced. Value timing "adds further value exposure" rather than a separate one, and their conclusion is that contrarian value timing "is, generally, a weak addition for long-term investors holding well-diversified factors including value."
What this test cannot tell you
It's one country, one factor and one signal. The Fama-French data here is US only, and the sample of genuinely independent wide-spread episodes is small. 19 of the 69 test years opened with a spread above its own historical average, and those episodes cluster, so the effective count is lower still. Backtests aren't forecasts, and a 69-year sample is not a law.
The hit rates make the weakness concrete. After a wide spread the value premium averaged 5.63% over the next twelve months, against 3.37% after a narrow one. That gap looks useful until you check how often the bet was simply right. Value was positive in 57.9% of the years that followed a wide reading and in 60.0% of the years that followed a narrow one. The signal shifts the average outcome without improving the odds, which means it works through a handful of enormous years like 2000 and not through consistency.
There are no trading costs in any of these numbers. The French factors are academic portfolios, rebalanced on paper, with no spreads, no borrow cost on the short leg and no tax. Every version of the timed rule trades more than the constant weight, so real costs would push the comparison further against timing, not towards it. The gap between paper factors and what investors actually received shows up in factor ETF returns, where all four MSCI factor indices beat the market by less once live than in backtest.
And the test says nothing about the other factors. Momentum, profitability and low volatility all have their own valuation measures and their own literature, and the size premium in the table above swung from 8.76% a year in the 1930s to -3.10% in the 2020s without any reference to the value spread. Research Affiliates counted, via Harvey, Liu and Heqing, some 314 published factors by the end of 2012, which is itself a reason to treat any single factor timing result with suspicion.
What would change the conclusion
Three things would. The first is a genuinely extreme reading. Asness's own position is not that valuations never matter, but that you'd need "far more extreme pricing than we do today" before acting on them, and June 2016 was not that. June 2026 reads 9.90, against 6.19 in the June he was writing in and 9.89 in June 2000. If the spread pushed past its 1935 high of 13.71 and the premium still failed to appear, the signal itself would be in question rather than the trade built on it.
The second is out-of-sample evidence since the argument was published. The Research Affiliates paper came out in September 2016. Over the 10 years from July 2016 to June 2026 the timed rule beat the constant weight by 1.41 percentage points a year, with a correlation of 0.55 between the signal and the following year's return. That is the strongest showing anywhere in the test. It also rests on 10 observations and a t-statistic of 0.27, which is why it changes nothing yet. Another decade like it would.
The third is a cost estimate. Every result here is gross. If somebody ran the same 12 rules through a live implementation with real spreads and real borrow, and the timed versions still landed inside a percentage point of the constant weight, that would settle more than another backtest can.
Until one of those lands, the honest reading of 69 years is that the value spread carries real information about the next decade and almost none about the next year, and that a rule sized to act on it ends up holding the same position it always held, just less comfortably.