Test Assets and Weak Factors: Giglio, Xiu & Zhang (2025)
Distilled by claude-sonnet-4-6 · extracted Jun 6, 2026, verified Jun 6, 2026
JEL (IAR-assigned): G12, C38, C55 · assigned from the abstract, not the journal
What this is. The paper’s core results, the model it builds on (the linear factor model with weak factors), and the method it contributes (SPCA) with the defining equations: enough to know what it found and how, without reading all 61 pages. To replicate or extend it, read the full source at the original.
Giglio, Xiu, and Zhang show that weak factors and test asset selection are two faces of the same problem: a factor is weak precisely in the cross section of test assets chosen by the researcher. They propose Supervised Principal Component Analysis (SPCA), a procedure that screens test assets by their correlation with the factor of interest at each iteration before extracting a principal component. This ensures that only assets with nontrivial exposure to the relevant factor are used, effectively strengthening the factor within the selected subset and enabling consistent risk premium estimation even when some latent SDF factors are weak or when some priced factors are omitted. Applied to a large cross section of 901 to 1,672 characteristic-sorted equity portfolios (1976 to 2020), SPCA estimates risk premia close to model-free averages for tradable factors, achieves substantially higher out-of-sample hedging than PCA for factors with weak exposures, and diagnoses that standard observable factor models (CAPM, FF3, FF5) miss important latent pricing factors.
Core results
Section titled “Core results”Magnitudes and significance are as reported. Locators point into the source PDF.
| # | Result | Locator | Magnitude |
|---|---|---|---|
| R1 | SPCA gives risk premia consistent with model-free averages for the market factor, a strong factor, across all tuning parameters | Table IV, p. 300; Figure 5, Panel A, p. 302 | Market RP estimates 68-74 bps/month for p = 3, 5, 7, 11; average excess return 74/62 bps (train/eval); never statistically different at 5% |
| R2 | SPCA risk premia are consistent with model-free averages for all tradable factors (with minor exceptions); the intermediary capital factor has a significant positive RP even among nontradables | Table IV, pp. 300-301; Figure 5, pp. 302; Figure 6, p. 303 | Momentum RP ~112 bps (p=3) varying across p=3-11; HML RP 37-50 bps; liquidity factor RP 70-95 bps out-of-sample (p.305) |
| R3 | Almost all nontradable macro factors (IP growth, uncertainty, consumption, term spread, credit, oil) have risk premia indistinguishable from zero; equity markets cannot span their variation | Table IV, pp. 300-301; Figure 6, p. 303 | Positive out-of-sample is rare for macro nontradables; LN1/LN2 macro factors: near zero or negative; Liquidity: 0-4% |
| R4 | SPCA out-of-sample hedging is higher or equal to PCA’s, often with fewer factors | Figure 5, pp. 302; Figure 6, p. 303 | Market: for all p, q; Momentum: SPCA reaches with p=3; PCA needs p >= 6; Intermediary Cap: SPCA , PCA much lower |
| R5 | SPCA degrades little when informative test assets are removed; PCA degrades sharply | Figure 8, p. 311 | Momentum: SPCA from 86% to 77% without momentum assets; PCA from 76% to 48%. Profitability: SPCA 71% to 60%; PCA 41% to 14% |
| R6 | In simulations, SPCA has far smaller bias than PCA, rpPCA, Ridge, Lasso, and two-pass estimators for weak factors | Table I, p. 293 | Weak factor V bias: SPCA -3.4 bps vs PCA -35 bps at T=240; RMSE 14.6 vs 36.5 bps; PLS ranks second among alternatives |
| R7 | SPCA achieves out-of-sample Sharpe ratios closest to the theoretical optimum (0.256) among all estimators in SDF recovery simulations | Table III, p. 295 | SPCA SR: 0.193 (T=120), 0.226 (T=240), 0.241 (T=480); PCA: 0.084, 0.110, 0.227; rpPCA: 0.134, 0.192, 0.242 |
| R8 | SPCA diagnoses that CAPM, FF3, and FF5 all miss important pricing factors: SPCA Sharpe ratio grows well above each model’s own Sharpe ratio as additional latent factors are extracted | Figure 9, p. 315 | Market (CAPM) Sharpe = 0.46; SPCA SR grows from ~0.5 to ~1.4+ with 1-20 factors (CZ data); FF5+MOM+BAB+QMJ best but still misspecified in CZ data |
Overall (paper’s conclusion). SPCA resolves the weak factor problem in empirical asset pricing by treating factor weakness as a property of the test assets rather than the factor itself. It consistently estimates risk premia and recovers the SDF in the presence of weak, omitted, and mismeasured factors, and provides a diagnostic tool for observable factor models. Empirically, nearly all nontradable macro factors have risk premia statistically indistinguishable from zero in equity markets, while standard observable factor models miss important latent pricing factors.
Theory / model
Section titled “Theory / model”The paper studies a standard linear latent factor model. Suppose an vector of test asset excess returns follows (equation 1, p. 266):
where is an matrix of factor exposures, is a vector of factor innovations (), and is an vector of idiosyncratic errors. Factors may be latent or observable. The SDF in terms of latent factor innovations is (equation 2, p. 266):
and the tradable SDF representation in terms of test asset excess returns is (equation 3, p. 266):
where satisfies and . The observable factor proxy vector is linked to the latent factors by (equation 4, p. 267):
where , is a matrix of factor loadings of on , and is measurement error orthogonal to . The risk premium of is .
Definition of weak factors. A factor is strong relative to a cross section of test assets when the eigenvalues grow at rate for all (the standard “pervasive” assumption of Bai and Ng (2002)). A factor is weak when some grow at a slower rate than . The necessary condition for consistency of PCA-based risk premium estimation in the multi-factor case is (equation 6, p. 275):
When this fails for even one factor, the PCA estimator of Giglio and Xiu (2021) is inconsistent for risk premia (Proposition 1, p. 270, establishes this in the single-factor case). The same breakdown occurs in the observable-factor context studied by Kan and Zhang (1999), who showed that Fama-MacBeth regressions with useless factors produce invalid inference. Rank deficiency in the beta matrix produces the same failure even when each factor is individually strong (equation 7 example, p. 275).
Method
Section titled “Method”SPCA is an iterative selection-and-projection procedure. It adds a supervised screening step (Step S1) to the PCA-based risk premium estimator of Giglio and Xiu (2021). For the single-factor case (Algorithm 2, p. 273):
- S1 (Selection). Select a subset of test assets by correlation with :
where is the -quantile of the absolute covariances . Only the top assets are kept.
- S2. Run Steps S1 to S3 of the PCA-based Algorithm 1 on the selected return matrix , , and .
The output is .
For the general multifactor case, SPCA iterates selection and projection (Algorithm 3, p. 277). At each step :
(S1.a) Select using covariance with residuals of :
(S1.b) Apply Algorithm 1 to and to extract the -th latent factor .
(S1.c) Project onto to obtain .
(S1.d) Update residuals: , .
Stop at when for threshold (equation 10). The final risk premium estimate is .
Consistency (Theorem 1, p. 279). Under mild moment conditions, if and tuning parameters satisfy:
then . Theorem 2 (p. 280) additionally gives a CLT: under stronger conditions including , i.e., contains at least as many variables as true factors.
Comparison with alternatives. The Ridge estimator (Kozak, Nagel, and Santosh (2020)) converges at rate , which fails when condition (6) fails (Theorem 4a, p. 286). The Lasso estimator is consistent but at a slower rate (Theorem 4b). Lettau and Pelger (2020) propose risk-premium PCA (rpPCA) for weak factors, but the paper shows rpPCA is inconsistent for risk premia in the multifactor weak-factor setting studied here (Internet Appendix Section I). SPCA combines factor structure (via PCA) with supervision (via screening), achieving the rate without requiring the strong sparsity assumption.
Tuning parameters (number of factors) and (number of selected assets) are chosen by three-fold cross-validation, maximizing the time-series of the hedging portfolio for in the training sample.
Empirical specifications
Section titled “Empirical specifications”The empirical analysis uses monthly data from March 1976 to December 2020, split into training (first half) and evaluation (second half) subsamples. Two test asset universes are used:
- Main (Chen and Zimmermann (2022)): 901 characteristic-sorted portfolios (as many as each anomaly’s original paper used, 2 to 10 sorts) plus 49 industry portfolios from Ken French. The CZ data use the April 2021 release.
- Robustness (Hou, Xue, and Zhang (2020)): 1,672 portfolios sorted by characteristics (momentum, value, investment, profitability, intangibles, frictions).
Factors studied include eight tradable factors (market excess return, HML, SMB, RMW, CMA, momentum, BAB, QMJ) and twelve nontradable factors (liquidity, intermediary capital, IP growth, three LN macro principal components, three Jurado-Ludvigson-Ng uncertainty indexes, term spread, credit spread, unemployment, two sentiment indexes, oil, consumption growth). All factor data are at the monthly frequency.
Risk premium estimation (Table IV, R1-R3). SPCA is applied factor-by-factor (, each factor as its own ) and jointly (, all factors simultaneously) with . For each pair, SPCA: (i) selects assets with the largest absolute covariance with ; (ii) runs PCA on those assets to extract factors; (iii) uses Fama-MacBeth-type time-series regressions to estimate risk premia. The out-of-sample of the hedging portfolio for is reported in the evaluation half.
SDF diagnosis (Figure 9, R8). For each observable factor model (CAPM; the Fama and French (1993) three-factor model FF3; FF5; FF5+Momentum; FF5+Momentum+BAB+QMJ), SPCA extracts up to 20 latent factors using as supervisor. The out-of-sample Sharpe ratio of the SPCA-based SDF (right-hand side of equation 20, p. 288) is compared to the Sharpe ratio of . A SPCA Sharpe ratio exceeding the model’s own Sharpe indicates missing factors.
Datasets used
Section titled “Datasets used”| Dataset | Role in paper | Wiki page |
|---|---|---|
| Chen and Zimmermann (2022) open-source asset pricing library | Main test asset cross section: 901 characteristic-sorted equity portfolios, April 2021 release, 1976m3-2020m12 | Open Source Asset Pricing (no page yet) |
| Hou, Xue, and Zhang (2020) replicating anomalies library | Robustness test asset cross section: 1,672 characteristic-sorted portfolios | No page yet |
| Ken French Data Library | 49 industry portfolios (added to CZ cross section); Fama-French factor returns (FF3, FF5) for benchmarking | Ken French library |
| Tradable factor returns (market, SMB, HML, RMW, CMA, BAB, QMJ, momentum) | Observable factor proxies for risk premium estimation | No page yet |
| Nontradable factor series (liquidity, intermediary capital, IP growth, LN macro PCs, uncertainty, sentiment, term, credit, unemployment, oil, consumption) | Nontradable observable factor proxies ; sources include Pastor-Stambaugh, AQR, FRED, Ludvigson-Ng, Jurado-Ludvigson-Ng, Baker-Wurgler, national accounts | No page yet |
Sample: monthly, 1976m3-2020m12 (about 537 months). Training: 1976m3-1998m7; evaluation: 1998m8-2020m12 (approximately equal halves).
When to read the full paper
Section titled “When to read the full paper”Use the original if you are: estimating risk premia for a potentially weak factor and need a consistent estimator robust to omitted factors and measurement error; diagnosing whether an observable factor model (CAPM, FF3, FF5, or richer) spans the SDF; extending the asymptotic inference results (Theorems 1-2, 4-5) to new DGPs; or applying SPCA to asset classes or frequencies beyond the monthly US equity setting studied here. The locators above point to the key tables and figures.
Attribution and rights
Section titled “Attribution and rights”Source: peer-reviewed, The Journal of Finance 80(1), February 2025. Copyright 2024 the American Finance Association. Paywalled; extract-only. This distillation was extracted by an LLM on 2026-06-06 and is not human-verified or independently reproduced.
Giglio, Stefano, Dacheng Xiu, and Dake Zhang. “Test Assets and Weak Factors.” The Journal of Finance 80, no. 1 (February 2025): 259-319. DOI: 10.1111/jofi.13415. Copyright 2024 the American Finance Association. This page is an extract by the Institute for Automated Research; paywalled source, extract-only rights.