Skip to content

Pockets of Predictability (Replication): Cakici, Fieberg, Neumaier, Poddig & Zaremba (2025)

Distilled by claude-sonnet-4-6 · extracted May 31, 2026, last verified Jun 4, 2026

JEL (IAR-assigned): G12, G14, G17 · assigned from the abstract, not the journal

Full structured metadata (methods, scope, relatesTo, topics, datasets): raw Markdown (.md)

paper-summaryreturn-predictabilityreplicationmarket-timingtime-seriespanel-regressionopen-accesscc-bypeer-revieweddata:wrdsdata:ken-french

What this is. The paper’s core results, datasets, and identification strategy: enough to know what it found without reading all 20 pages. To replicate or extend it, read the full source at the original (open access).

Farmer, Schmidt, and Timmermann (2023, FST) claimed that U.S. aggregate stock market returns exhibit “pockets of predictability” identifiable ex ante via one-sided kernel regressions. Cakici et al. audit the FST replication package and find that the pocket-identification step in FST’s code uses a two-sided kernel, not the one-sided kernel described in the paper. A two-sided kernel draws on data both before and after the forecast date, making the procedure in-sample rather than out-of-sample and leaking future information into the model. Correcting this single error reduces average integral R-squared by a factor of roughly 20 (e.g., from 1.51-3.70% to 0.09-0.28% for daily forecasts). The in-pocket vs out-of-pocket return predictability difference largely disappears, and market-timing alphas become mostly insignificant. Economic restrictions on forecasts offer partial improvement but cannot restore the original conclusions.

Magnitudes and significance are as reported; \*/\*\*/\*\*\* = 10%/5%/1%. Locators point into the source PDF.

#ResultLocatorMagnitude
R1Original (two-sided) code reproduces FST exactly: in-pocket average integral R-squared ranges from 1.48% to 3.70% (daily), with strong in-pocket/out-of-pocket asymmetryTable II Panel A, p. 3779; Figure 1 p. 3773Mean integral R² (daily): dp 1.51%, tbl 1.70%, tsp 2.92%, rvar 2.77% (two-sided kernel)
R2Corrected (one-sided) code collapses predictability: pockets become roughly 20x more frequent, 10x shorter, and far less predictableTable II Panel B, p. 3779Mean integral R² (daily): dp 0.18%, tbl 0.09%, tsp 0.09%, rvar 0.28% (one-sided kernel)
R3In-pocket CW t-statistics vanish with the one-sided kernel: under the two-sided code in-pocket CW t-stats commonly exceed 3-4; under the corrected code they are insignificant for nearly all 27 model-predictor combinationsTable III Panel A.1 vs A.2, pp. 3781-3782Two-sided in-pocket CW (unrestricted): dp 3.00***, tbl 4.75***, tsp 3.04***; one-sided in-pocket CW (unrestricted): dp -0.47, tbl 0.10, tsp -1.06
R4In-pocket alphas drop sharply: unrestricted in-pocket annualised alphas fall from 0.76-6.38% (two-sided) to -0.44-2.51% (one-sided); only one out of nine individual/composite predictors exceeds 1% significance under the one-sided kernelTable III Panel B.1 vs B.2, pp. 3782-3783Average Sharpe ratio drops from 0.71 (two-sided) to 0.44 (one-sided), below the prevailing-mean benchmark 0.46
R5Benchmark model beats kernel models out-of-pocket (two-sided code): out-of-pocket CW t-stats are significantly negative (at 10%) for most individual predictors under the original codeTable III Panel A.1, p. 3781Out-of-pocket CW: dp -1.62†, tbl -1.33†, tsp -1.52†, rvar -1.77†† (two-sided, unrestricted)
R6Alternative bandwidth robustness: corrected code always fails to identify in-pocket predictability across 2-, 2.5-, and 3-year estimation windows and 6-, 12-, 18-month SED windows; not a single significant in-pocket CW t-stat in Panel BTable IV Panel B, pp. 3785-3786All Panel B in-pocket CW entries insignificant across all bandwidth/window combinations
R7Monthly data confirms the result: monthly in-pocket CW t-stats are strong under the two-sided kernel (e.g., tbl 3.55***, tsp 2.44***) but mostly insignificant under the one-sided kernelTable V Panels A and B, p. 3788One-sided in-pocket monthly CW: dp 0.90, tbl 1.22, tsp 0.57, rvar 1.01
R8Partial exception: factor returns (SMB, HML) retain some time-varying predictability even with the corrected one-sided kernel, though weaker than FST documented with the two-sided approach§II.E, p. 3788-3789 (Internet Appendix Section IV)Qualitative finding: significant CW stats and market-timing gains persist for factor portfolios; aggregate equity market is the null result

Overall (paper’s conclusion). The FST pocket-of-predictability evidence is an artefact of in-sample kernel estimation. Once the identification step is restricted to information available before the forecast date, the pockets shrink to one-day artefacts, predictability inside and outside pockets becomes statistically indistinguishable, and market-timing strategies based on the pockets offer no reliable abnormal returns.

DatasetRole in paperWiki page
CRSP U.S. stock market excess return (daily and monthly, 1926-2016)Dependent variable: aggregate market excess return (CRSP return minus short T-bill)WRDS / CRSP (licensed)
Dividend-price ratio (dp), 1926-2016Predictor variable (sourced from FST replication package)no page yet
3-month T-bill rate (tbl), 1954-2016Predictor variable (sourced from FST replication package)no page yet
Term spread (tsp), 1962-2016Predictor variable (sourced from FST replication package)no page yet
Realized variance (rvar), 1927-2016Predictor variable (sourced from FST replication package)no page yet
Ken French Data Library (SMB, HML factor returns)Used in Section II.E robustness for factor-level predictabilityKen French

Study period: 1926-2016 (predictor-dependent; see Table I, p. 3777). Data sourced directly from the FST replication package (“Replication-code 20190881.zip”, Journal of Finance website).

This paper has no original structural economic model. It is a methodological audit and replication of Farmer, Schmidt, and Timmermann (2023, FST). The theoretical object under scrutiny is the claim that aggregate equity market return predictability is time-varying and can be identified ex ante using one-sided kernel regressions. The testable hypothesis is:

  • Null: market-timing strategies built on FST’s “pockets” offer no reliable abnormal returns once a correctly out-of-sample pocket-identification kernel is used.
  • Identification: the two-framework comparison is the entire identification strategy. Every FST analysis is run twice with all other parameters held fixed; only the kernel type in the second estimation stage differs. Any difference in results is attributed to the kernel type (one-sided vs two-sided), since that is the sole deviation from the FST code.

The FST return prediction model (eq. 1, p. 3774):

rt+1=xtβt+ϵt+1(1)r_{t+1} = x_t' \beta_t + \epsilon_{t+1} \tag{1}
  • rt+1r_{t+1} is the excess U.S. stock market return
  • xtx_t is a vector of predictor variables (dp, tbl, tsp, rvar)
  • βt\beta_t are time-varying regression coefficients
  • σt2=E[ϵt+12xt]\sigma_t^2 = E[\epsilon_{t+1}^2 \mid x_t] allows for conditional heteroskedasticity

The βt\beta_t are estimated by the local constant model (eq. 2, p. 3774):

β^t=arg minβ0s=1TKhT(st)[rs+1xsβ0]2(2)\hat{\beta}_t = \operatorname*{arg\,min}_{\beta_0} \sum_{s=1}^{T} K_{hT}(s-t) \cdot [r_{s+1} - x_s' \beta_0]^2 \tag{2}

with kernel weights KhT(u)=K(u/hT)/(hT)K_{hT}(u) = K(u/hT)/(hT) and bandwidth hh. FST use a 2.5-year bandwidth in this step with a one-sided Epanechnikov kernel (eq. 3, p. 3775):

K(u)=32(1u2)1{1<u<0}(3)K(u) = \tfrac{3}{2}(1 - u^2) \cdot \mathbf{1}\{-1 < u < 0\} \tag{3}

Only data from before time tt receives positive weight under the one-sided kernel, making the βt\beta_t estimation genuinely out-of-sample. The discrepancy arises in the second stage.

The method builds on kernel-regression for both the return-prediction estimation and the pocket-identification step, and on time-series-forecasting for evaluating out-of-sample performance against the prevailing-mean benchmark. The paper applies these techniques, it does not propose a new one.

Squared error differential (SED) (eq. 4, p. 3775) measures whether the kernel model outperforms the prevailing-mean benchmark at each date tt:

SEDt=(rtrˉtt1)2(rtr^tt1)2(4)\text{SED}_t = (r_t - \bar{r}_{t|t-1})^2 - (r_t - \hat{r}_{t|t-1})^2 \tag{4}
  • rˉtt1\bar{r}_{t|t-1} is the prevailing-mean forecast
  • r^tt1\hat{r}_{t|t-1} is the kernel model forecast
  • Positive SEDt\text{SED}_t means the kernel model has smaller forecast error that period

Pocket identification (eq. 5, p. 3775): a pocket begins when the fitted SED trend is positive:

SED^t=γ0,t+γ1,tt>0(5)\widehat{\text{SED}}_t = \gamma_{0,t} + \gamma_{1,t} \cdot t > 0 \tag{5}
  • γ0,t\gamma_{0,t} and γ1,t\gamma_{1,t} should be estimated with a one-sided Epanechnikov kernel and one-year bandwidth
  • FST’s published code instead uses a two-sided kernel with a 24-month symmetric window (12 months before and 12 months after day tt), so the identification draws on data that is unavailable at forecast time

Two-framework comparison design: the paper runs the complete FST analysis twice, in parallel, changing only this kernel choice. Panel A results use the original (two-sided) code; Panel B results use the corrected (one-sided) code. All other parameters, bandwidth choices, predictor series, and performance metrics are identical. This clean design means the contrast of Panel A vs Panel B isolates the kernel-type effect.

Performance metrics (§II.B, pp. 3780-3783):

  • Clark-West (2007) t-statistic comparing kernel-model forecasts to the prevailing-mean benchmark
  • Annualised alpha from a stock/T-bill timing strategy (holding stocks when the kernel model predicts positive returns, T-bills otherwise), with Newey-West (1987) t-statistics
  • Annualised Sharpe ratio of the timing strategy
  • Three forecast restriction variants following Campbell and Thompson (2008): unrestricted, non-negative excess return forecasts only, and sign restrictions on both forecasts and slope coefficients

All regressions use daily U.S. excess stock market returns as the dependent variable unless noted (monthly robustness in Table V). The sample varies by predictor (Table I, p. 3777): dp starts November 5, 1926; tbl starts January 4, 1954; tsp starts January 2, 1962; rvar starts January 15, 1927; all end December 2016.

Constant-coefficient replication regressions (Table I, p. 3777). Univariate OLS of daily excess return on each lagged predictor, separately for full sample, in-pocket, and out-of-pocket subsamples. Slope coefficients and Newey-West (1987) adjusted t-statistics and R-squared reported. The pocket partition is the only thing that differs between Panel A and Panel B in this table; the regression itself is identical.

Pocket statistics (Table II, p. 3779). For each predictor and kernel type, counts the number of pockets, their fraction of sample, duration (min/mean/max days), and integral R-squared (min/mean/max). Integral R-squared IR2IR^2 is computed as in FST (p. 1289): the average within-pocket R-squared, weighting by pocket length. Results reported separately for daily and monthly data.

Clark-West prediction performance (Table III, pp. 3781-3782). CW t-statistic comparing kernel model to the prevailing-mean benchmark rˉt+1=(1/t)s=1trs\bar{r}_{t+1} = (1/t)\sum_{s=1}^{t} r_s. Reported separately for:

  • full sample, in-pocket subperiod, out-of-pocket subperiod
  • Panel A.1 (in-sample / two-sided kernel), Panel A.2 (one-sided kernel)
  • nine predictor/composite specifications (dp, tbl, tsp, rvar, pc, mv, comb1, comb2, comb3)
  • three forecast restriction variants Significance under one-tailed test (positive direction) for alphas; two-tailed for CW tests.

Economic significance (Table III Panel B, pp. 3782-3783). Annualised alpha, Newey-West t-statistic, and Sharpe ratio of the asset allocation strategy (stocks when forecast is positive, T-bills otherwise). Reported for the same nine predictor/composite specifications and three restriction variants, separately for the two kernel approaches.

Bandwidth robustness (Table IV, pp. 3785-3786). CW t-statistics for coefficient estimation windows of 2-, 2.5-, 3-year and SED estimation windows of 6-month, 12-month, 15-month. Panel A uses the two-sided kernel; Panel B uses the corrected one-sided kernel. Nine predictor and composite specifications, daily data.

Monthly data robustness (Table V, p. 3788). Reproduces Table III (Panel A and Panel B) using monthly returns, following FST Table VII Panel A. Estimation period is 2.5 years for monthly forecasts.

Use the original (open access) if you are: auditing the FST replication package directly; extending the kernel-regression methodology to other predictors or markets; or assessing whether the partial improvements from economic restrictions (sign constraints on slope coefficients and forecasts) restore the FST conclusions. The locators above point to the exact tables. For “what did this paper find,” the table above is sufficient.

Source: peer-reviewed, The Journal of Finance 80(6), December 2025, pp. 3771-3790. This distillation was extracted by an LLM on 2026-05-31 and is not human-verified or independently reproduced. The article is open access under CC BY 4.0; PDF mirroring is permitted by the licence but the PDF is not hosted in this batch.

Attribution (CC BY 4.0). Cakici, Nusret, Christian Fieberg, Tobias Neumaier, Thorsten Poddig, and Adam Zaremba. “Pockets of Predictability: A Replication.” The Journal of Finance 80, no. 6 (December 2025): 3771-3790. DOI: 10.1111/jofi.13484. © 2025 The Author(s). Licensed under CC BY 4.0. This page is an adaptation by the Institute for Automated Research: core results extracted and re-expressed; changes were made.

Found an error or want a topic covered? Open an issue, use the Edit page link above, or email contact@instituteforautomatedresearch.org. Edits are reviewed before publishing; provenance and accuracy are the point.