Skip to content

War Discourse and the Cross Section: Hirshleifer, Mai & Pukthuanthong (2025)

Distilled by claude-sonnet-4-6 · extracted May 31, 2026, last verified Jun 4, 2026

JEL (IAR-assigned): G12, G14, G41 · assigned from the abstract, not the journal

Full structured metadata (methods, scope, relatesTo, topics, datasets): raw Markdown (.md)

paper-summaryasset-pricinganomaliestext-as-datafactorsdisaster-riskreturn-predictabilityfama-macbethportfolio-sortpeer-reviewedunreplicateddata:nyt-newsdata:wrdsdata:ken-frenchdata:open-source-asset-pricing

What this is. The paper’s core results, datasets, and theory: enough to know what it found without reading all 49 pages. To replicate or extend it, consult the original (paywalled) or request data and code from the authors.

Using 7,000,000 New York Times articles spanning 1871 to 2019, the paper builds a monthly war-discourse index (War) via a semisupervised topic model (sLDA, one seed word: war), then defines the war factor WarFac as the AR(1) innovation in War. Loadings on WarFac significantly and negatively predict expected returns across six sets of test assets covering up to 4,964 portfolios (138 HXZ long-short anomalies, 1,372 HXZ single-sorted, 904 CZ single-sorted, 360 ML-based nonlinear, and 128 and 2,190 own-constructed portfolios), with a monthly return premium ranging from about -0.66% to -3.32% per month. WarFac is incremental to the Fama-French six-factor model and to news-based uncertainty indexes (NVIX, GPR). A mimicking portfolio (WMP) earns an annualised Sharpe ratio of 1.73 and passes the Pukthuanthong et al. (2019) factor-identification protocol and the Giglio-Xiu three-pass test.

Magnitudes and significance are as reported; \*/\*\*/\*\*\* = 10%/5%/1%. Locators point into the source PDF.

#ResultLocatorMagnitude
R1WarFac commands a significant negative return premium on 138 HXZ long-short anomaly portfoliosTable I Panel A, Table II Panel A, p. 3607–3611λ = -1.33%/mo (t = -2.87***) standalone; remains -0.47%** with all FF6+M4+DHS+Q5 factors; R² = 48% as single factor vs 51–77% for multifactor benchmarks (FF6 59%, M4 65%, DHS 51%, Q5 77%)
R2WarFac return premium is negative and significant for all six sets of test assets; consistently ranks in the top three among 11 nontraded factorsTable I (all panels), p. 3607–3610λ ranges from -0.66%** (HXZ single-sorted, t = -2.25) to -3.32%*** (ML portfolios, t = -3.42); no other nontraded factor achieves this across all six sets
R3WarFac explains 62% of cross-sectional variance in ML-based nonlinear portfolio returns as a single factor, outperforming FF6 (41%), M4 (40%), DHS (35%), Q5 (58%)Table II Panel D, p. 3617–3618Single-factor R² = 62%; adding WarFac to FF6 raises R² by 34%; common pricing error falls from 3.3% to near zero
R4WarFac return premium is incremental to traded factors (WMP, MKT, SMB, HML, RMW, CMA, MOM and mispricing factors); CMA and WarFac are the only factors significant across all six test-asset setsTable III (all panels), p. 3619–3622WarFac: -1.33%*** to -3.32%*** depending on test assets; CMA also consistently significant; WMP -2.19%** to -3.32%***
R5WarFac is incremental to NVIX and GPR uncertainty indexes; NVIX and GPR do not command significant return premia across all test assetsTable IV (all panels), p. 3623–3624With all three factors, WarFac: -1.04%** (HXZ long-short), -2.76%* (ML portfolios); NVIX_War2Fac: insignificant for HXZ and ML; GPRFac: insignificant across all panels
R6WarFac prices industry portfolios with a negative premium, incremental to the CrisisFac (crisis event counts) of Berkman et al. (2011)Table V, p. 362630-industry portfolios: λ(WarFac) = -0.24%* (t = -1.89) alone; -0.32%** (t = -2.39) with CrisisFac and CWarFac jointly; 49-industry: -0.28%** (t = -2.16) in joint specification
R7WMP (the traded mimicking portfolio for WarFac) has a Sharpe ratio of 1.73, the highest among all factors in the sample, and generates significant alphas against all benchmark factor models§V.A and Internet Appendix Table IA.III, p. 3627–3628Monthly average return = -3.32%, monthly SD = 6.64%; annualised Sharpe = 1.73; monthly alpha vs all factors ≈ 3.10%*** (t-stat in Internet Appendix); WMP passes three-pass test and factor-identification protocol
R8WarFac captures a distinct tail risk: return premium survives after controlling for CAPM beta, bear beta, downside beta, VIX beta, volatility beta, jump beta, coskewness, skewness beta, tail beta, and idiosyncratic volatility§VII.A and Internet Appendix Table IA.VII, p. 3632WarFac premium remains significant with all tail-risk mimicking portfolios included; it is the only nontraded factor with a significant beta return premium on HXZ single-sorted portfolios in this horse race

Overall (paper’s conclusion, p. 3634). Loadings on the war-discourse factor strongly predict the cross section of stock returns with a negative premium, consistent with rational rare-disaster hedging (good hedges earn low premia) or with behavioral overweighting of war prospects (war-sensitive stocks are overpriced). The war premium is incremental to all standard factor models and to other news-based uncertainty measures, and is driven by factual war news rather than opinion articles.

DatasetRole in paperWiki page
New York Times full text, Jan 1871–Oct 2019 (~7M articles)Source corpus for sLDA topic modelling; constructs the War indexno page yet; proprietary/licensed archive
CRSP monthly stock returns and characteristicsReturns for all six test-asset sets; underlying portfolio construction dataWRDS / CRSP / Compustat (licensed)
CompustatFirm fundamentals for anomaly characteristic constructionWRDS / CRSP / Compustat (licensed)
Hou, Xue & Zhang (2020) HXZ anomaly portfolios (138 long-short, 1,372 single-sorted)Primary test-asset sets; Jul 1972–Dec 2016no page yet; available from HXZ replication files
Chen & Zimmermann (2022) single-sorted portfolios (904)Third test-asset set; available from open-source-asset-pricing projectOpen Source Asset Pricing
Bryzgalova, Huang & Julliard (2023) ML-based nonlinear portfolios (360)Fourth test-asset set; tree-based nonlinear portfoliosno page yet
Ken French Data Library (size, B/M, momentum portfolios)Basis assets for WMP time-series mimicking; Fama-French factor benchmarksKen French Data Library
Berkman, Jacobsen & Lee (2011) crisis event countsCrisisFac and CWarFac benchmarks for §IV.B industry tests; data from sites.duke.edu/icbdatano page yet
NVIX (Manela & Moreira 2017)Benchmark news-based uncertainty index; horse-race test in §IV.Ano page yet
GPR index (Caldara & Iacoviello 2022)Benchmark geopolitical risk index; horse-race test in §IV.Ano page yet

Sample period for asset pricing tests: Jul 1972–Dec 2016 (532 months). War index: Jan 1871–Oct 2019.

This paper has no original structural model. It tests two competing, observationally equivalent theoretical frameworks:

  1. Rational rare-disaster risk (Barro 2006, 2009; Gourio 2008; Gabaix 2012, p. 3601): investors demand a risk premium for bearing war-related disaster risk. Assets that pay off when war risk is high are good hedges and therefore command lower expected returns. A negative cross-sectional return premium on WarFac betas is the central prediction.

  2. Behavioral overweighting (Daniel, Hirshleifer & Subrahmanyam 2001; Tversky & Kahneman 1992 cumulative prospect theory, p. 3601): investors overweight the probability of rare salient disasters such as war, overvaluing stocks that do well under high war risk, so those stocks subsequently earn lower returns. The same negative premium arises for behavioral reasons.

The paper tests both frameworks by constructing a text-based proxy for investor attention to war risk rather than relying on realized war events, which have small sample sizes. Identification rests on the rolling-forward estimation of the sLDA model and the AR(1) residual (WarFac), which ensures only past information is used at each point in time, avoiding look-ahead bias (p. 3591).

The method has two components: (1) constructing the War index via semisupervised topic modelling (sLDA), and (2) building WarFac as the innovation in War via a rolling AR(1). It builds on slda-topic-model and ar1-innovation. The War index used here is the same one that the companion aggregate-return study of Hirshleifer, Mai, and Pukthuanthong (2025, Review of Financial Studies) uses; this paper extends it to cross-sectional pricing (p. 3590).

sLDA topic model (pp. 3596-3598). Each month tt, the model is estimated on all New York Times articles in the preceding 120 months (the rolling window [t119,t][t-119, t]). Using Gibbs sampling, the model infers, for each document dd, the document-topic distribution θd\theta_d (a vector of topic probabilities) and, for each topic kk, the topic-word distribution ϕk\phi_k (a vector of word probabilities). The seed word for the War topic is war (a single word, for parsimony and to avoid researcher discretion in seed-word selection). The global monthly weight of topic kk in month tt is the length-weighted average across all articles dd in month tt:

Wart=1dlenddlendθd,k=War\text{War}_t = \frac{1}{\sum_d \text{len}_d} \sum_d \text{len}_d \cdot \theta_{d,k=\text{War}}
  • lend\text{len}_d is article length in n-gram count.

The rolling window allows topic-word distributions ϕk\phi_k to shift with language over time, which is essential for a corpus spanning 1871 to 2019 (p. 3597).

AR(1) innovation (p. 3603, equations 3 and 4). Following Berkman, Jacobsen & Lee (2011), Liu & Matthies (2022), and Giglio & Xiu (2021), WarFac is defined as the residual from a rolling AR(1) fit to War, estimated at each month tt using data from 1926 to tt to avoid look-ahead bias:

Wart=ρ0+ρWart1+ut(3)\text{War}_t = \rho_0 + \rho \cdot \text{War}_{t-1} + u_t \tag{3} WarFact=ut(4)\text{WarFac}_t = u_t \tag{4}

The AR(1) coefficients (ρ0,ρ)(\rho_0, \rho) are re-estimated each month on the growing window of available data. Results are robust to using an ARMA(1,1) residual or a rolling-regression residual (pp. 3591, 3603 fn. 13).

War-mimicking portfolio (WMP). The traded version of WarFac is constructed using the cross-sectional approach of Lehmann & Modest (1988): the slope from the monthly second-pass cross-sectional regression of asset returns on WarFac betas is the monthly WMP return (p. 3627). As a robustness check, the time-series approach projects WarFac onto the space of excess returns of 360 tree-based portfolios plus basis assets:

WarFact=α+βRte+ϵt(8)\text{WarFac}_t = \alpha + \beta' R^e_t + \epsilon_t \tag{8} WMPt=β^Rte(9)\text{WMP}_t = \hat{\beta}' R^e_t \tag{9}
  • ReR^e is the vector of excess returns on basis assets.
  • β^\hat{\beta} is estimated by OLS on the full sample.

All asset pricing tests use monthly data, July 1972 to December 2016 (T = 532 months for WarFac; T = 522 for tests including NVIX_War and GPR).

First pass: factor loadings (eq. 1, p. 3602). For each test asset i=1,,Ni = 1, \ldots, N, excess returns are regressed on a vector of factors FtF_t in a multivariate time-series regression:

Rite=αi+βiFFt+ϵit,i=1,,N(1)R^e_{it} = \alpha_i + \beta_{iF}' F_t + \epsilon_{it}, \quad i = 1, \ldots, N \tag{1}

The paper reports avg(t)\text{avg}(|t|) (average absolute beta t-statistic) and the number of assets with t1.65|t| \geq 1.65 (the 5% one-sided threshold). This first pass is run for each of the six test-asset sets separately.

Second pass: cross-sectional return premium (eq. 2, p. 3602). Time-series average excess returns are regressed cross-sectionally on the estimated factor loadings:

Rˉie=λ0+βiFλF+ei(2)\bar{R}^e_{i} = \lambda_0 + \beta_{iF}' \lambda_F + e_i \tag{2}
  • λF\lambda_F is the vector of return premium slopes.
  • Standard errors are Shanken (1992) corrected.
  • The paper reports λ\lambda and its tt-statistic, cross-sectional R2=1σe2/σμ2R^2 = 1 - \sigma^2_e / \sigma^2_{\mu}, and mean absolute pricing error MAPE=eˉ\text{MAPE} = |\bar{e}|.
  • Under rational pricing, λ0=0\lambda_0 = 0.

Industry portfolios: rolling Fama-MacBeth with betas (eqs. 5-6, p. 3625). For the industry pricing tests (§IV.B), betas are estimated over a rolling 60-month window for excess returns on factor XX (WarFac, CrisisFac, or CWarFac) plus market, size, and value controls:

Rite=αi+βitXt+βitMKTMKTt+βitSMBSMBt+βitHMLHMLt+ϵit,window: t-59 to t(5)R^e_{it} = \alpha_i + \beta_{it} X_t + \beta^{\text{MKT}}_{it} \text{MKT}_t + \beta^{\text{SMB}}_{it} \text{SMB}_t + \beta^{\text{HML}}_{it} \text{HML}_t + \epsilon_{it}, \quad \text{window: } t\text{-}59 \text{ to } t \tag{5}

Cross-sectional betas are ranked into quintiles each month tt and rescaled to [0,1][0, 1]. The monthly return premium is estimated by:

Rite=λ0t+λtβi,t1+λtMKTβi,t1MKT+λtSMBβi,t1SMB+λtHMLβi,t1HML+eit(6)R^e_{it} = \lambda_{0t} + \lambda_t \beta_{i,t-1} + \lambda^{\text{MKT}}_t \beta^{\text{MKT}}_{i,t-1} + \lambda^{\text{SMB}}_t \beta^{\text{SMB}}_{i,t-1} + \lambda^{\text{HML}}_t \beta^{\text{HML}}_{i,t-1} + e_{it} \tag{6}
  • Time-series averages of λt\lambda_t are reported; statistical significance uses Newey-West (1987) standard errors.
  • Sample period for industry tests: July 1926 to December 2018 (T = 1,110 months for Panels A and B of Table V, p. 3626).

WMP spanning test (eq. 7, p. 3628).

WMPt=α+βFt+ϵt(7)\text{WMP}_t = \alpha + \beta' F_t + \epsilon_t \tag{7}
  • FtF_t is the vector of benchmark traded factors.
  • α\alpha measures whether WMP expands the mean-variance frontier.
  • Monthly alpha of WMP against all factors combined is approximately 3.10%, significant at the 1% level (Internet Appendix Table IA.III, p. 3628).

Use the original at doi.org/10.1111/jofi.13482 (paywalled) if you are: replicating (data and code available from authors on request per fn. 16); extending the sLDA war-discourse methodology to other corpora or time periods; doing a literature review where the full robustness battery (seed-word variants, ARMA(1,1), sLDA vs LDA comparisons, tail-risk horse races) matters; or auditing a specific coefficient. The locators above point to the exact table.

Source: peer-reviewed, The Journal of Finance 80(6), December 2025, pp. 3589–3637. © 2025 the American Finance Association. This distillation was extracted by an LLM on 2026-05-31 and is not human-verified or independently reproduced. The underlying article is paywalled; no verbatim PDF is hosted here.

Hirshleifer, David, Dat Mai, and Kuntara Pukthuanthong. “War Discourse and the Cross Section of Expected Stock Returns.” The Journal of Finance 80, no. 6 (December 2025): 3589–3637. DOI: 10.1111/jofi.13482. © 2025 the American Finance Association. Extract-only; redistribution of the original article is subject to Wiley/AFA terms.

Found an error or want a topic covered? Open an issue, use the Edit page link above, or email contact@instituteforautomatedresearch.org. Edits are reviewed before publishing; provenance and accuracy are the point.