Skip to content

Test Assets and Weak Factors: Giglio, Xiu & Zhang (2025)

Distilled by claude-sonnet-4-6 · extracted Jun 6, 2026, verified Jun 6, 2026

JEL (IAR-assigned): G12, C38, C55 · assigned from the abstract, not the journal

Full structured metadata (methods, scope, relatesTo, topics, datasets): raw Markdown (.md)

paper-summaryasset-pricingfactorsfactor-modelsweak-factorsrisk-premiapanel-regressionfama-macbethpeer-reviewedunreplicateddata:open-source-asset-pricingdata:ken-french

What this is. The paper’s core results, the model it builds on (the linear factor model with weak factors), and the method it contributes (SPCA) with the defining equations: enough to know what it found and how, without reading all 61 pages. To replicate or extend it, read the full source at the original.

Giglio, Xiu, and Zhang show that weak factors and test asset selection are two faces of the same problem: a factor is weak precisely in the cross section of test assets chosen by the researcher. They propose Supervised Principal Component Analysis (SPCA), a procedure that screens test assets by their correlation with the factor of interest gtg_t at each iteration before extracting a principal component. This ensures that only assets with nontrivial exposure to the relevant factor are used, effectively strengthening the factor within the selected subset and enabling consistent risk premium estimation even when some latent SDF factors are weak or when some priced factors are omitted. Applied to a large cross section of 901 to 1,672 characteristic-sorted equity portfolios (1976 to 2020), SPCA estimates risk premia close to model-free averages for tradable factors, achieves substantially higher out-of-sample hedging R2R^2 than PCA for factors with weak exposures, and diagnoses that standard observable factor models (CAPM, FF3, FF5) miss important latent pricing factors.

Magnitudes and significance are as reported. Locators point into the source PDF.

#ResultLocatorMagnitude
R1SPCA gives risk premia consistent with model-free averages for the market factor, a strong factor, across all tuning parametersTable IV, p. 300; Figure 5, Panel A, p. 302Market RP estimates 68-74 bps/month for p = 3, 5, 7, 11; average excess return 74/62 bps (train/eval); never statistically different at 5%
R2SPCA risk premia are consistent with model-free averages for all tradable factors (with minor exceptions); the intermediary capital factor has a significant positive RP even among nontradablesTable IV, pp. 300-301; Figure 5, pp. 302; Figure 6, p. 303Momentum RP ~112 bps (p=3) varying across p=3-11; HML RP 37-50 bps; liquidity factor RP 70-95 bps out-of-sample (p.305)
R3Almost all nontradable macro factors (IP growth, uncertainty, consumption, term spread, credit, oil) have risk premia indistinguishable from zero; equity markets cannot span their variationTable IV, pp. 300-301; Figure 6, p. 303Positive out-of-sample R2R^2 is rare for macro nontradables; LN1/LN2 macro factors: R2R^2 near zero or negative; Liquidity: R2R^2 0-4%
R4SPCA out-of-sample hedging R2R^2 is higher or equal to PCA’s, often with fewer factorsFigure 5, pp. 302; Figure 6, p. 303Market: R2>0.98R^2 > 0.98 for all p, q; Momentum: SPCA reaches R2>70%R^2 > 70\% with p=3; PCA needs p >= 6; Intermediary Cap: SPCA R250%R^2 \approx 50\%, PCA much lower
R5SPCA degrades little when informative test assets are removed; PCA degrades sharplyFigure 8, p. 311Momentum: SPCA R2R^2 from 86% to 77% without momentum assets; PCA from 76% to 48%. Profitability: SPCA 71% to 60%; PCA 41% to 14%
R6In simulations, SPCA has far smaller bias than PCA, rpPCA, Ridge, Lasso, and two-pass estimators for weak factorsTable I, p. 293Weak factor V bias: SPCA -3.4 bps vs PCA -35 bps at T=240; RMSE 14.6 vs 36.5 bps; PLS ranks second among alternatives
R7SPCA achieves out-of-sample Sharpe ratios closest to the theoretical optimum (0.256) among all estimators in SDF recovery simulationsTable III, p. 295SPCA SR: 0.193 (T=120), 0.226 (T=240), 0.241 (T=480); PCA: 0.084, 0.110, 0.227; rpPCA: 0.134, 0.192, 0.242
R8SPCA diagnoses that CAPM, FF3, and FF5 all miss important pricing factors: SPCA Sharpe ratio grows well above each model’s own Sharpe ratio as additional latent factors are extractedFigure 9, p. 315Market (CAPM) Sharpe = 0.46; SPCA SR grows from ~0.5 to ~1.4+ with 1-20 factors (CZ data); FF5+MOM+BAB+QMJ best but still misspecified in CZ data

Overall (paper’s conclusion). SPCA resolves the weak factor problem in empirical asset pricing by treating factor weakness as a property of the test assets rather than the factor itself. It consistently estimates risk premia and recovers the SDF in the presence of weak, omitted, and mismeasured factors, and provides a diagnostic tool for observable factor models. Empirically, nearly all nontradable macro factors have risk premia statistically indistinguishable from zero in equity markets, while standard observable factor models miss important latent pricing factors.

The paper studies a standard linear latent factor model. Suppose an N×1N \times 1 vector of test asset excess returns rtr_t follows (equation 1, p. 266):

rt=βγ+βvt+ut,E(vt)=E(ut)=0 and cov(vt,ut)=0,(1)r_t = \beta\gamma + \beta v_t + u_t, \qquad \mathbb{E}(v_t) = \mathbb{E}(u_t) = 0 \text{ and } \text{cov}(v_t, u_t) = 0, \tag{1}

where β\beta is an N×pN \times p matrix of factor exposures, vtv_t is a p×1p \times 1 vector of factor innovations (vt=ftμfv_t = f_t - \mu_f), and utu_t is an N×1N \times 1 vector of idiosyncratic errors. Factors ftf_t may be latent or observable. The SDF in terms of latent factor innovations is (equation 2, p. 266):

mt=1γΣv1vt,(2)m_t = 1 - \gamma^\top \Sigma_v^{-1} v_t, \tag{2}

and the tradable SDF representation in terms of test asset excess returns is (equation 3, p. 266):

m~t=1b(rtE(rt)),(3)\tilde{m}_t = 1 - b^\top (r_t - \mathbb{E}(r_t)), \tag{3}

where bb satisfies E(rt)=Σb\mathbb{E}(r_t) = \Sigma b and Σ=cov(rt)\Sigma = \text{cov}(r_t). The observable factor proxy vector gtg_t is linked to the latent factors by (equation 4, p. 267):

gt=ξ+ηvt+zt,(4)g_t = \xi + \eta v_t + z_t, \tag{4}

where ξ=E(gt)\xi = \mathbb{E}(g_t), η\eta is a d×pd \times p matrix of factor loadings of gtg_t on vtv_t, and ztz_t is measurement error orthogonal to vtv_t. The risk premium of gtg_t is γg=cov(mt,gt)=ηγ\gamma_g = -\text{cov}(m_t, g_t) = \eta\gamma.

Definition of weak factors. A factor is strong relative to a cross section of NN test assets when the eigenvalues λi(ββ)\lambda_i(\beta^\top\beta) grow at rate NN for all i=1,,pi = 1, \ldots, p (the standard “pervasive” assumption of Bai and Ng (2002)). A factor is weak when some λi(ββ)\lambda_i(\beta^\top\beta) grow at a slower rate than NN. The necessary condition for consistency of PCA-based risk premium estimation in the multi-factor case is (equation 6, p. 275):

N/(λmin(ββ)T)0.(6)N / (\lambda_{\min}(\beta^\top\beta) T) \to 0. \tag{6}

When this fails for even one factor, the PCA estimator of Giglio and Xiu (2021) is inconsistent for risk premia (Proposition 1, p. 270, establishes this in the single-factor case). The same breakdown occurs in the observable-factor context studied by Kan and Zhang (1999), who showed that Fama-MacBeth regressions with useless factors produce invalid inference. Rank deficiency in the beta matrix produces the same failure even when each factor is individually strong (equation 7 example, p. 275).

SPCA is an iterative selection-and-projection procedure. It adds a supervised screening step (Step S1) to the PCA-based risk premium estimator of Giglio and Xiu (2021). For the single-factor case (Algorithm 2, p. 273):

  • S1 (Selection). Select a subset I^N\hat{I} \subset \langle N \rangle of test assets by correlation with gtg_t:
I^={iT1Rˉ[i]Gˉcq},\hat{I} = \left\{ i \mid T^{-1} \| \bar{R}_{[i]} \bar{G}^\top \| \geq c_q \right\},

where cqc_q is the (1q)(1-q)-quantile of the absolute covariances {T1Rˉ[i]Gˉ}iN\{ T^{-1} | \bar{R}_{[i]} \bar{G}^\top | \}_{i \in \langle N \rangle}. Only the top qNqN assets are kept.

  • S2. Run Steps S1 to S3 of the PCA-based Algorithm 1 on the selected return matrix Rˉ[I^]\bar{R}_{[\hat{I}]}, Gˉ\bar{G}, and p=1p = 1.

The output is γ^gSPCA:=η^γ^\hat{\gamma}_g^{\text{SPCA}} := \hat{\eta}\hat{\gamma}.

For the general multifactor case, SPCA iterates selection and projection (Algorithm 3, p. 277). At each step kk:

(S1.a) Select I^k\hat{I}_k using covariance with residuals of G(k)G_{(k)}:

I^k={iT1(Rˉ(k))[i]Gˉ(k)MAXcq(k)}.(9)\hat{I}_k = \left\{ i \mid T^{-1} \| (\bar{R}_{(k)})_{[i]} \bar{G}_{(k)}^\top \|_{\text{MAX}} \geq c_q^{(k)} \right\}. \tag{9}

(S1.b) Apply Algorithm 1 to (Rˉ(k))[I^k](\bar{R}_{(k)})_{[\hat{I}_k]} and Gˉ(k)\bar{G}_{(k)} to extract the kk-th latent factor V^(k)\hat{V}_{(k)}.

(S1.c) Project Rˉ(k)\bar{R}_{(k)} onto V^(k)\hat{V}_{(k)} to obtain β^(k)\hat{\beta}_{(k)}.

(S1.d) Update residuals: Rˉ(k+1)=Rˉ(k)β^(k)V^(k)\bar{R}_{(k+1)} = \bar{R}_{(k)} - \hat{\beta}_{(k)}\hat{V}_{(k)}^\top, Gˉ(k+1)=Gˉ(k)η^(k)γ^(k)\bar{G}_{(k+1)} = \bar{G}_{(k)} - \hat{\eta}_{(k)}\hat{\gamma}_{(k)}.

Stop at k=p^k = \hat{p} when cq(k)<cc_q^{(k)} < c for threshold cc (equation 10). The final risk premium estimate is γ^gSPCA=k=1p^η^(k)γ^(k)\hat{\gamma}_g^{\text{SPCA}} = \sum_{k=1}^{\hat{p}} \hat{\eta}_{(k)}\hat{\gamma}_{(k)}.

Consistency (Theorem 1, p. 279). Under mild moment conditions, if log(NT)(N01+T1)0\log(NT)(N_0^{-1} + T^{-1}) \to 0 and tuning parameters satisfy:

c0,c1(logNT)1/2(q1/2N1/2+T1/2)0,qN/N00,(11)c \to 0, \quad c^{-1}(\log NT)^{1/2}(q^{-1/2}N^{-1/2} + T^{-1/2}) \to 0, \quad qN/N_0 \to 0, \tag{11}

then γ^gSPCAPηγ\hat{\gamma}_g^{\text{SPCA}} \xrightarrow{P} \eta\gamma. Theorem 2 (p. 280) additionally gives a CLT: T(γ^gSPCAηγ)dN(0,Φ)\sqrt{T}(\hat{\gamma}_g^{\text{SPCA}} - \eta\gamma) \xrightarrow{d} \mathcal{N}(0, \Phi) under stronger conditions including λmin(ηη)1\lambda_{\min}(\eta^\top\eta) \gtrsim 1, i.e., gtg_t contains at least as many variables as true factors.

Comparison with alternatives. The Ridge estimator (Kozak, Nagel, and Santosh (2020)) converges at rate (N+T)/(λpT)(N+T)/(\lambda_p T), which fails when condition (6) fails (Theorem 4a, p. 286). The Lasso estimator is consistent but at a slower rate b1logN/T\|b\|_1 \sqrt{\log N / T} (Theorem 4b). Lettau and Pelger (2020) propose risk-premium PCA (rpPCA) for weak factors, but the paper shows rpPCA is inconsistent for risk premia in the multifactor weak-factor setting studied here (Internet Appendix Section I). SPCA combines factor structure (via PCA) with supervision (via screening), achieving the T1/2T^{-1/2} rate without requiring the strong sparsity assumption.

Tuning parameters pp (number of factors) and qN\lfloor qN \rfloor (number of selected assets) are chosen by three-fold cross-validation, maximizing the time-series R2R^2 of the hedging portfolio for gtg_t in the training sample.

The empirical analysis uses monthly data from March 1976 to December 2020, split into training (first half) and evaluation (second half) subsamples. Two test asset universes are used:

  • Main (Chen and Zimmermann (2022)): 901 characteristic-sorted portfolios (as many as each anomaly’s original paper used, 2 to 10 sorts) plus 49 industry portfolios from Ken French. The CZ data use the April 2021 release.
  • Robustness (Hou, Xue, and Zhang (2020)): 1,672 portfolios sorted by characteristics (momentum, value, investment, profitability, intangibles, frictions).

Factors studied include eight tradable factors (market excess return, HML, SMB, RMW, CMA, momentum, BAB, QMJ) and twelve nontradable factors (liquidity, intermediary capital, IP growth, three LN macro principal components, three Jurado-Ludvigson-Ng uncertainty indexes, term spread, credit spread, unemployment, two sentiment indexes, oil, consumption growth). All factor data are at the monthly frequency.

Risk premium estimation (Table IV, R1-R3). SPCA is applied factor-by-factor (d=1d=1, each factor as its own gtg_t) and jointly (d=pd=p, all factors simultaneously) with p{3,5,7,11}p \in \{3, 5, 7, 11\}. For each (p,q)(p, q) pair, SPCA: (i) selects qN\lfloor qN \rfloor assets with the largest absolute covariance with gtg_t; (ii) runs PCA on those assets to extract pp factors; (iii) uses Fama-MacBeth-type time-series regressions to estimate risk premia. The out-of-sample R2R^2 of the hedging portfolio for gtg_t is reported in the evaluation half.

SDF diagnosis (Figure 9, R8). For each observable factor model gtg_t (CAPM; the Fama and French (1993) three-factor model FF3; FF5; FF5+Momentum; FF5+Momentum+BAB+QMJ), SPCA extracts up to 20 latent factors using gtg_t as supervisor. The out-of-sample Sharpe ratio of the SPCA-based SDF (right-hand side of equation 20, p. 288) is compared to the Sharpe ratio of gtg_t. A SPCA Sharpe ratio exceeding the model’s own Sharpe indicates missing factors.

DatasetRole in paperWiki page
Chen and Zimmermann (2022) open-source asset pricing libraryMain test asset cross section: 901 characteristic-sorted equity portfolios, April 2021 release, 1976m3-2020m12Open Source Asset Pricing (no page yet)
Hou, Xue, and Zhang (2020) replicating anomalies libraryRobustness test asset cross section: 1,672 characteristic-sorted portfoliosNo page yet
Ken French Data Library49 industry portfolios (added to CZ cross section); Fama-French factor returns (FF3, FF5) for benchmarkingKen French library
Tradable factor returns (market, SMB, HML, RMW, CMA, BAB, QMJ, momentum)Observable factor proxies gtg_t for risk premium estimationNo page yet
Nontradable factor series (liquidity, intermediary capital, IP growth, LN macro PCs, uncertainty, sentiment, term, credit, unemployment, oil, consumption)Nontradable observable factor proxies gtg_t; sources include Pastor-Stambaugh, AQR, FRED, Ludvigson-Ng, Jurado-Ludvigson-Ng, Baker-Wurgler, national accountsNo page yet

Sample: monthly, 1976m3-2020m12 (about 537 months). Training: 1976m3-1998m7; evaluation: 1998m8-2020m12 (approximately equal halves).

Use the original if you are: estimating risk premia for a potentially weak factor and need a consistent estimator robust to omitted factors and measurement error; diagnosing whether an observable factor model (CAPM, FF3, FF5, or richer) spans the SDF; extending the asymptotic inference results (Theorems 1-2, 4-5) to new DGPs; or applying SPCA to asset classes or frequencies beyond the monthly US equity setting studied here. The locators above point to the key tables and figures.

Source: peer-reviewed, The Journal of Finance 80(1), February 2025. Copyright 2024 the American Finance Association. Paywalled; extract-only. This distillation was extracted by an LLM on 2026-06-06 and is not human-verified or independently reproduced.

Giglio, Stefano, Dacheng Xiu, and Dake Zhang. “Test Assets and Weak Factors.” The Journal of Finance 80, no. 1 (February 2025): 259-319. DOI: 10.1111/jofi.13415. Copyright 2024 the American Finance Association. This page is an extract by the Institute for Automated Research; paywalled source, extract-only rights.

Found an error or want a topic covered? Open an issue, use the Edit page link above, or email contact@instituteforautomatedresearch.org. Edits are reviewed before publishing; provenance and accuracy are the point.