Skip to content

How to Dominate the Historical Average: Li, Li, Lyu & Yu (2025)

Distilled by claude-sonnet-4-6 · extracted Jun 6, 2026, verified Jun 6, 2026

JEL (IAR-assigned): G12, G11, C53 · assigned from the abstract, not the journal

Full structured metadata (methods, scope, relatesTo, topics, datasets): raw Markdown (.md)

paper-summaryasset-pricingreturn-forecastingequity-premiumtime-seriespanel-regressionpeer-reviewedunreplicateddata:shiller-datadata:wrds

What this is. The paper’s core results, the theoretical framework (bias-variance trade-off in OOS forecasting, first-order stochastic dominance theorems), and the method (conservative-slope forecast with parameter A): enough to understand what it found and how, without reading all 31 pages. To replicate or extend it, read the full source at the original.

The paper proposes an OOS equity premium forecast: instead of setting the predictive slope to zero (historical average) or estimating it by OLS, use a small positive constant slope δ=1/A\delta = 1/A (where A is a large positive number calibrated to the lower confidence bound of the estimated slope). The method has zero estimation variance, matching the historical average, but a lower bias when the population slope is nonzero. The paper proves theoretically (Theorems 1-4) that this forecast first-order stochastically dominates the historical average, and shows empirically on 23 predictors from Goyal and Welch (2008) that 15 of 23 generate significantly positive OOS R2R^2 at the 90% level. Goyal, Welch, and Zafirov (2024) confirmed the Goyal and Welch (2008) findings using a larger predictor set, providing the direct motivation for the paper. The dividend-to-price ratio achieves an OOS R2R^2 of 2.1% (p = .019) at A = 100, versus an insignificant 0.2% for OLS. Clark and West (2006) show finite-sample estimation noise makes OLS R2R^2 negative under the null of no predictability; the proposed method avoids this by using a constant slope with zero variance.

Magnitudes and significance are as reported; \*/\*\* = 10%/5%. Locators point into the source PDF.

#ResultLocatorMagnitude
R1Method OOS R2R^2 for dp predictor (A=50): statistically significant improvement over HMTable 3, p. 3110R2=3.4%R^2 = 3.4\%, p-value = .044; OLS R2=0.2%R^2 = 0.2\%, p-value = .477
R2Method OOS R2R^2 for dp predictor (A=100): gains statistical power as A increasesTable 3, p. 3110R2=2.1%R^2 = 2.1\%, p-value = .019; OLS R2=0.2%R^2 = 0.2\%, p = .477; CT++ R2=1.8%R^2 = 1.8\%, p = .286
R3Method OOS R2R^2 for dp predictor (A=500, A=1,000): very conservative slopes still beat HM at 99% significanceTable 3, p. 3110A=500: R2=0.5%R^2 = 0.5\%, p=.009; A=1,000: R2=0.2%R^2 = 0.2\%, p=.008
R4Across 23 predictors, 15 (8) have positive OOS R2R^2 at 90% (95%) significance§5 / Internet Appendix C, p. 3090-309115 of 23 at 90%; 8 of 23 at 95%; OLS and CT generate statistically insignificant R2R^2 for most (Goyal and Welch 2008)
R5Method first-order stochastically dominates HM for dp predictor: empirical CDF of MSE everywhere above HM CDFFigure 6, p. 3112A=100 CDF (MSE) > HM CDF for all MSE thresholds in 40-year rolling windows; confirmed with kernel smoothing in Internet Appendix C.5
R6Simulations confirm method’s OOS R2R^2 distribution is entirely to the right of zero; OLS can be negativeFigure 3, p. 3105A=50 centered at ~4-5 (broadest), A=100 at ~3-4, A=200 at ~2, A=500 at ~1 (narrowest spike); OLS distribution spans $[-10, +10]$ with nontrivial probability of R2<0R^2 < 0
R7Previously published confidence bounds (Campbell and Shiller 1988a) add value to OOS forecasts when used as the predictive slope§5.4, p. 3112-3113A=209 (95% lower bound from Campbell and Shiller 1988a) produces positive OOS R2R^2 from 1987 onward; demonstrates prior study estimates are not data mining

Overall (paper’s conclusion). A conservative deterministic predictive slope, calibrated to a lower confidence bound near zero, provably dominates the historical average and empirically dominates OLS and Campbell-Thompson forecasts on most of the 23 standard predictors from Goyal and Welch (2008). The method is an ex ante validated benchmark for time-varying expected return models.

The paper models the equity premium return as a linear predictive relationship (Equation 1, p. 3092):

r=μ+bx+e,(1)r = \mu + bx + e, \tag{1}

where bb is the population predictive slope coefficient of the predictor xx, ee is a residual with zero mean uncorrelated with xx, xx has zero mean, and μ\mu is the unconditional expected return. The historical average sets the slope on xx to zero. The unknown μ\mu is estimated by the historical average μ^\hat{\mu}:

μ=μ^+ξ,(2)\mu = \hat{\mu} + \xi, \tag{2}

where ξ\xi has zero mean. Substituting gives (Equation 3, p. 3093):

r=μ^+bx+ϵ,ϵe+ξ.(3)r = \hat{\mu} + bx + \epsilon, \quad \epsilon \equiv e + \xi. \tag{3}

The method’s forecast for the next return is μ^+δx\hat{\mu} + \delta x where δ=1/A\delta = 1/A if b>0b > 0 and δ=1/A\delta = -1/A if b<0b < 0 (Equation 5, p. 3093):

δ={1/Aif b>01/Aif b<0.(5)\delta = \begin{cases} 1/A & \text{if } b > 0 \\ -1/A & \text{if } b < 0. \end{cases} \tag{5}

Theorem 1 (p. 3092): Under conditions that δ\delta is a constant between 0 and bb and the pdf of the error vector e\mathbf{e} is strictly decreasing in e\|\mathbf{e}\|, the forecast μ+δx\mu + \delta x first-order stochastically dominates the forecast μ\mu (population mean) for predicting return rr. That is, applying loss MSE-\text{MSE}, the forecast μ+δx\mu + \delta x gives at least as high a probability of achieving any MSE threshold, and strictly higher probability for some.

Theorem 2 (p. 3093): The forecast μ^+δx\hat{\mu} + \delta x first-order stochastically dominates the historical average μ^\hat{\mu} under the same conditions on δ\delta and the distribution of ϵ=e+ξ\boldsymbol{\epsilon} = e + \xi. Since ξ\xi can correlate with xx (Stambaugh 1999 bias), the monotonicity condition on the pdf of ϵ\boldsymbol{\epsilon} is imposed (the tt and normal distributions satisfy this).

Corollary 1 (p. 3093): Under Theorem 2’s assumptions, the MSE using μ^+δx\hat{\mu} + \delta x satisfies:

MSEμ^+δx=dMSEμ^η,η0.(4)\text{MSE}_{\hat{\mu}+\delta x} \overset{d}{=} \text{MSE}_{\hat{\mu}} - \eta, \quad \eta \geq 0. \tag{4}

The MSE improvement η\eta is a nonnegative random variable, so the method has weakly lower MSE in expectation and first-order stochastically lower MSE.

Theorem 3 (p. 3095) extends stochastic dominance to the case where the sign of bb is inferred with error (probability pp correct, qq wrong). For a given predictor realization x\mathbf{x}, the forecast μ+dx\mu + dx (where dd is δ\delta, δ-\delta, or 0 according to the sign inference outcome) first-order stochastically dominates the historical average when p/q>max(B1,B2)p/q > \max(B_1, B_2), where B1>(2b+δ)/(2bδ)B_1 > (2b+\delta)/(2b-\delta) (approximately 1 when δ/b\delta/b is near zero, and 3 at the maximum δ/b=1\delta/b = 1). Statistical significance at the 95% level gives p/q0.95/0.05=19p/q \geq 0.95/0.05 = 19, well above the cutoff of 3.

Theorem 4 (p. 3097) restates Theorem 3 for the estimated-mean setting where μ\mu is replaced by μ^\hat{\mu}, yielding Corollary 2 (p. 3098, Equation 9):

MSEμ^+dx=dMSEμ^η,η0.(9)\text{MSE}_{\hat{\mu}+dx} \overset{d}{=} \text{MSE}_{\hat{\mu}} - \eta, \quad \eta \geq 0. \tag{9}

OOS MSE decomposition. The expected MSE difference between the historical average and the method is (Equation 10, p. 3098):

E ⁣[MSEμ^MSEμ^+b^x]=(b2Bias(b^)2Variance(b^))x22E ⁣[(μ^μ)b^]x.(10)E\!\left[\text{MSE}_{\hat{\mu}} - \text{MSE}_{\hat{\mu}+\hat{b}x}\right] = \left(b^2 - \text{Bias}(\hat{b})^2 - \text{Variance}(\hat{b})\right)x^2 - 2E\!\left[(\hat{\mu}-\mu)\hat{b}\right]x. \tag{10}

The historical average’s slope is zero and therefore unbiased but uses no predictive information. A regression slope reduces the first two terms but can inflate the variance term to the point where it dominates, yielding a negative OOS R2R^2. The method uses a deterministic δ\delta, so the variance of b^\hat{b} is zero and Equation (10) simplifies to (b2Bias(b^)2)x2>0(b^2 - \text{Bias}(\hat{b})^2)x^2 > 0 whenever b^\hat{b} is between 0 and bb. This is the core intuition: a constant nonzero slope beats both the historical average (zero slope, biased) and OLS (unbiased mean but high variance).

Gradient descent interpretation. The method is a one-step gradient descent update of the historical average toward greater predictability, using the sign (but not the magnitude) of bb as the gradient signal and step size 1/A1/A (Equation 11, p. 3099):

a1=a0γf(a0),(11)a_1 = a_0 - \gamma f'(a_0), \tag{11}

where a0=0a_0 = 0 (the historical average slope), a1=δa_1 = \delta, and γ=1/A\gamma = 1/A.

The implementation has three steps.

Step 1: Obtain the sign of bb. Sign can come from (a) economic theory, as in Campbell and Thompson (2008), who restrict the OLS slope to have the theoretically expected sign, or (b) statistical inference: use the confidence interval [b^L,b^U][\hat{b}_L, \hat{b}_U] for population slope bb. If 0<b^L0 < \hat{b}_L the slope is significantly positive; if b^U<0\hat{b}_U < 0 it is significantly negative. The method builds on time-series-forecasting (predictive regression) but replaces the OLS slope with a constant.

Step 2: Choose A. Setting 1/A1/A to the lower confidence bound b^L\hat{b}_L ensures with near certainty that δ\delta is between 0 and bb (Equation 7, p. 3093):

0<b^Lbwith 95% probability.(7)0 < \hat{b}_L \leq b \quad \text{with 95\% probability.} \tag{7}

For the dividend-to-price ratio, Campbell and Shiller (1988a) Table 4 give a predictive slope of 0.129 (SE = 0.057); the standardized predictor has a standard deviation of 0.277, so the 95% lower confidence bound for the standardized slope is (0.1291.96×0.057)×0.277=0.0048(0.129 - 1.96 \times 0.057) \times 0.277 = 0.0048, implying A=209A = 209 (p. 3094).

Step 3: Standardize the predictor. Predictors are standardized to zero mean and unit variance using a 20-year rolling backward-looking window, so δ=1/A\delta = 1/A measures the effect of a one-standard-deviation change in the predictor on the forecast annual return (p. 3094, 3107).

Simulation design. A VAR(1) for log returns rr, log dividend-to-price ratio dpdp, and log dividend growth Δd\Delta d following Cochrane (2008) provides the data-generating process (Equations 12-13, p. 3101). Parameters are calibrated to the sample (Table 1, p. 3102): population predictive slope br=0.143b_r = 0.143, ρ=0.8911\rho = 0.8911. The historical average (HM) uses a rolling 20-year window. OOS R2R^2 is computed per Equation 14 (p. 3103):

R2=MSEμ^MSEμ^+δxMSEμ^.(14)R^2 = \frac{\text{MSE}_{\hat{\mu}} - \text{MSE}_{\hat{\mu}+\delta x}}{\text{MSE}_{\hat{\mu}}}. \tag{14}

Data. Annual value-weighted CRSP market returns (post-1926) and S&P 500 Index returns (pre-1926), 23 predictors from Goyal and Welch (2008) extended through 2017 (Table 2, p. 3105): bm, cape Shiller, cay, corpr, csp, de, dfr, dfy, dp, dy, ep, eqis, ik, infl, ltr, lty, tbl, tms, ntis, svar, and four Robert Shiller series. Predictors are standardized using rolling 20-year windows. Forecasts start 20 years after the sample start year for each predictor.

OOS R2R^2 evaluation (R1-R4, R7). For each predictor, compute annual one-step-ahead OOS forecasts using the method with A{50,100,200,500,1000}A \in \{50, 100, 200, 500, 1000\}, OLS, CT+, and CT++ (Campbell and Thompson 2008). Compute R2R^2 via Equation (14). Report one-sided p-values using Diebold (2015) heteroscedasticity-adjusted test, cross-checked by Harvey, Leybourne, and Newbold (1997). For dp, this corresponds to a sample of 146 annual observations.

Stochastic dominance evaluation (R5). Estimate the empirical CDF of OOS MSE over 40-year rolling windows of annual successive one-year-ahead forecasts using dp predictor (Figure 6, p. 3112). Compare the method CDF (A=100) to the historical mean CDF. First-order dominance requires the method CDF to lie everywhere above (to the left of) the HM CDF. Kernel smoothing (Internet Appendix C.5) confirms the finding.

Bias-variance simulation (R6). Generate 10,000 simulation samples of the VAR in Equations (12)-(13) with parameters from Table 1. In each sample compute the OOS R2R^2 using a rolling 20-year window for HM and a rolling 20-year window for δ=1/A\delta = 1/A for the method. Figure 3 (p. 3105) reports the kernel density of the R2R^2 distribution across simulations for A{50,100,200,500}A \in \{50, 100, 200, 500\} and OLS.

DatasetRole in paperWiki page
CRSP value-weighted index returnAnnual market return post-1926WRDS / CRSP (licensed)
S&P 500 Index returns (Shiller website)Annual market return pre-1926Shiller data
Goyal and Welch (2008) predictor data (Amit Goyal’s website)19 predictor series for equity premium forecasting, extended to 2017No page yet
Robert Shiller data (http://www.econ.yale.edu/~shiller/data.htm)4 additional predictors: cape Shiller, infl Shiller, lty Shiller, Trcape ShillerShiller data

Sample: 23 predictors with start years ranging from 1872 to 1947 (Table 2, p. 3105), all ending 2017. Forecasting starts 20 years after the predictor start year. Frequency: annual.

Use the original if you are: (a) constructing a competing equity premium forecast and need the formal ex ante dominance proofs for a given distributon assumption; (b) selecting the parameter A for a specific predictor using the confidence-bound rule (Internet Appendix C.1 contains calculations for all 23 predictors); (c) investigating whether previously published coefficient estimates add OOS value (Section 5.4); or (d) comparing to shrinkage estimators such as Ridge and Lasso (Internet Appendix C.6). The key tables are Table 3 (OOS R2R^2 for dp) and Internet Appendix Table C.1 (all predictors); key figures are Figure 1 (cumulative OOS performance), Figure 3 (simulation R2R^2 density), and Figure 6 (stochastic dominance CDF).

Source: peer-reviewed, The Review of Financial Studies 38(10). This distillation was extracted by an LLM on 2026-06-06 and is not human-verified or independently reproduced. The CC BY-NC-ND 4.0 licence permits non-commercial reproduction with attribution and no derivatives; the verbatim PDF is not hosted here.

Li, Kai, Yingying Li, Changlei Lyu, and Jialin Yu. “How to Dominate the Historical Average.” The Review of Financial Studies 38, no. 10 (2025): 3086-3116. DOI: 10.1093/rfs/hhaf010. Replication code: Harvard Dataverse, https://doi.org/10.7910/DVN/9VNJUN. Licensed under CC BY-NC-ND 4.0. This page is a distilled extract by the Institute for Automated Research.

Found an error or want a topic covered? Open an issue, use the Edit page link above, or email contact@instituteforautomatedresearch.org. Edits are reviewed before publishing; provenance and accuracy are the point.