Mutual Fund Stars: Hounyo & Lin (2026)
Distilled by claude-sonnet-4-6 · extracted Jun 25, 2026, verified Jun 25, 2026
JEL (IAR-assigned): C33, G11, G12, G14 · assigned from the abstract, not the journal
What this is. The paper’s core results, the statistical model of fund performance, and the cross-sectional dependent wild bootstrap (CSDWB) method with its defining equations: enough to understand what was found and why existing bootstrap tests fail, without reading all 18 pages. To replicate or extend, read the full source at DOI 10.1016/j.jempfin.2025.101673.
The Fama-French (2010) bootstrap test for mutual fund skill is severely undersized, primarily because it jointly resamples factors and residuals, producing duplicate observation pairs in the bootstrap sample. This distortion far exceeds the undersampling effect identified by Harvey and Liu (2022) and persists even at large sample sizes. Hounyo and Lin propose the cross-sectional dependent wild bootstrap (CSDWB), which multiplies cross-sectional residuals by random Rademacher weights at each time point, preserving unbalanced panel structure and cross-sectional dependence while eliminating duplicate observations. Simulations confirm CSDWB achieves near-optimal test size under a range of cross-sectional dependence scenarios. Applied to 3,630 U.S. equity mutual funds from January 1984 to May 2019, CSDWB finds that a measurable fraction of funds outperform the market on the four-factor Carhart (1997) model, with outperformance concentrated before 2003 and absent afterwards.
Core results
Section titled “Core results”Magnitudes are as reported; \* = significant at the 5% level (likelihood exceeds 0.95).
Locators point to tables, figures, and sections of the source PDF.
| # | Result | Locator | Magnitude |
|---|---|---|---|
| R1 | FF joint resampling generates far more extreme bootstrap t-statistics than Fixed (no-duplicate) resampling, even at T=60, confirming duplicate observations (not undersampling) as the dominant distortion | Fig. 1, p. 4 | Joint & Complete bootstrap right-tail density is visibly larger than Fixed & Complete at both T=13 and T=60; the gap persists as T increases, showing duplicate-observation inflation does not vanish with sample size |
| R2 | CSDWB_I and CSDWB_II have near-optimal test size (~10%) at all percentiles across three cross-sectional dependence scenarios; FF is oversized at the max percentile and HL methods are unreliable for small T | Figs. 7-9, pp. 11-12 | CSDWB_I and CSDWB_II size ≈10% at max, 99.5, 99, 95, 90th percentiles under no, weak, and strong cross-sectional dependence; HL_II size diverges from 10% at small T |
| R3 | CSDWB detects genuine mutual fund outperformers across the performance distribution; the FF method misses them entirely | Table 2, p. 15 | CSDWB_I: max 0.915, 2nd 0.989*, 5th 0.992*, 99.9 0.988*, 99.5 0.972*, 99th 0.920; FF: max 0.154, 2nd 0.330, 5th 0.568; KTWW: max 0.424, 2nd 0.806, 5th 0.982* |
| R4 | Outperformance is concentrated entirely in the pre-2003 period; no evidence of outperformance after 2003 | §4.1, p. 13-14; Table E.3 | Pre-2003 subperiod: strong evidence of outperformance using CSDWB; post-2003: no evidence; pattern consistent across bootstrap methods |
| R5 | Best-performing funds portfolio (238 funds, selected via MCS-based Algorithm 1 using 1984-1993 data) earns cumulative excess returns of 8.91% above the market over 4 years and 15.12% over 9 years | §4.2, p. 15; Fig. 11, p. 16 | Superior performance significant at 10% level for first 4 years; advantage fades over longer horizons; most notable in 1998, when the portfolio briefly underperforms the market |
| R6 | Time-varying (nonparametric) alpha for top-percentile funds is stable pre-2005 near the upper bootstrap confidence bound; post-2005 alpha becomes volatile and approaches the lower bound | §4.3, p. 16; Fig. 12, p. 17 | 99.5th percentile fund alpha ≈ 0.05-0.25% per month pre-2005; post-2005 widening confidence interval and alpha near lower bound suggest reduced and uncertain skill |
Overall (paper’s conclusion). Duplicate observations in the Fama-French (2010) bootstrap are the dominant source of its severe undersizing, not the undersampling problem previously emphasized by Harvey and Liu (2022). The proposed CSDWB methods correct this flaw and confirm that a small but measurable fraction of U.S. equity mutual funds outperform the market, with outperformance concentrated in the pre-2003 period. Outperforming funds tend to have lower expense ratios, higher total net assets, and higher turnover, consistent with active management in a competitive environment before low-cost ETFs and product proliferation compressed alpha opportunities post-2003.
Theory / model
Section titled “Theory / model”The paper has no structural economic model. The statistical framework is a four-factor panel regression for fund excess returns with the null hypothesis of zero true alpha.
Huang et al. (2023) analyzed mutual fund performance using a bootstrap approach but used the FF method to generate simulated data, embedding the same duplicate-observations distortion into the DGP. The current paper shows that a valid bootstrap method applied to FF-generated data will detect excess extreme t-statistics, suggesting the presence of outperformers even under the null, and proposes CSDWB as the correct alternative.
The benchmark regression (equation 2.1, p. 2) for fund at time is:
where is fund ‘s excess return over the risk-free rate, is fund ‘s true abnormal return (alpha), is the factor loading on factor , is the return on factor at time , and is the residual. The factors include the market excess return (), the size factor () and value-growth factor () from Fama and French (1993), and the momentum factor () from Carhart (1997).
The null hypothesis of no skill is:
A positive indicates fund outperforms the market after adjusting for factor exposures. The paper tests this null by comparing the distribution of estimated statistics (ranked order statistics) against a bootstrap-approximated distribution under , focusing on extreme-right-tail percentiles (99th, 99.5th, max) following Kosowski et al. (2006) and Fama and French (2010).
The key insight is that the Fama-French (2010) implementation distorts the bootstrap null distribution by jointly resampling factor returns and fund residuals, so the same pair may appear multiple times in a bootstrap sample, inflating extreme t-statistics and reducing the likelihood of rejection.
Method
Section titled “Method”The paper proposes two variants of a cross-sectional dependent wild bootstrap (CSDWB).
Both variants build on panel-regression (the Carhart regression framework, equation 2.1)
and the wild-bootstrap primitive (Wu 1986, Liu 1988, Mammen 1993).
Bootstrap pseudo-returns under the null (equation 2.3, p. 3): factor returns and residuals are combined to produce bootstrap pseudo-returns with zero true alpha:
where is the estimated factor loading, and and are bootstrap factor returns and residuals. The CSDWB methods differ from Kosowski et al. (2006) and Fama and French (2010) in how these quantities are generated.
CSDWB_I (the preferred variant, p. 8): Let denote fund ‘s observed period and if , 0 otherwise. Let be a random weight with mean 0 and variance 1, drawn independently at each period (Rademacher weights: with probability 0.5 each, following Cameron et al. 2008 and Davidson and Flachaire 2008). The CSDWB_I bootstrap residuals and factor returns are:
The same scalar weight is applied to the entire cross-section of residuals at each period , preserving cross-sectional dependence. Missing observations remain intact. Because each observation uses a different realization of , no exact duplicate pairs are generated, eliminating the FF duplicate-observations distortion.
CSDWB_II differs from CSDWB_I solely in that factor returns are not perturbed: . This makes CSDWB_II closely related to both the KTWW and CSDB methods and is useful for elucidating differences across approaches.
Fund selection algorithm (Algorithm 1, p. 9): A Model Confidence Set (Hansen et al. 2011) sequential procedure identifies the best funds. At each iteration, CSDWB is used to derive a 90% confidence level for the performance of the current best fund relative to all other funds in the active set . The worst performer is eliminated if the null of equal performance is rejected. The procedure iterates until no further rejection occurs, yielding a final set of best-performing funds.
Empirical specifications
Section titled “Empirical specifications”Data sample: CRSP Survivor-Bias-Free U.S. Mutual Fund Database, 3,630 open-end U.S. domestic equity funds, January 1984 to May 2019, 425 months. Factor returns from the Kenneth French Data Library (four-factor model: RMRF, SMB, HML, MOM). The final dataset is an unbalanced panel with missing observations reflecting fund entry and exit.
Bootstrap implementation: For each fund , OLS estimates of model (2.1) yield and residuals . Bootstrap pseudo-returns are generated by subtracting and applying CSDWB weights. Each bootstrap iteration produces a ranking of the fund t-statistics; the procedure is repeated times to build the empirical null distribution. The test focuses on extreme right-tail percentiles (max, 2nd, 5th, 99.9th, 99.5th, 99th, 98th, 97th, 95th, 90th) following Kosowski et al. (2006).
Performance likelihood (Table 2, p. 15): For each percentile, the reported likelihood is
the proportion of 10,000 bootstrap replications in which the simulated t-statistics at that
percentile fall below the observed t-statistic. A likelihood exceeding 0.95 (marked with \*)
provides strong evidence of genuine outperformance at that percentile.
Simulation design (Section 3, pp. 9-12): Three parametric scenarios (no, weak, strong cross-sectional dependence; ) with funds over periods, 5% of funds having positive alpha, . An enhanced empirical simulation uses CSDWB_I to generate data under the null (Section 3.2.2), avoiding the FF duplicate-observations distortion in the simulated DGP.
Subperiod analysis (Section 4.1, §4.2): Pre/post-2003 split to test persistence. Fund selection period: January 1984 to December 1993 (10 years of training data, 238 funds selected via Algorithm 1). Out-of-sample evaluation: January 1984 to December 2002.
Time-varying alpha (Section 4.3, equation 4.1, p. 16): A local least-squares estimator estimates the time-varying alpha in:
using a rolling-window approach following Cai (2007) and Cai et al. (2018). Bootstrap confidence intervals are constructed with CSDWB_I fixing .
Datasets used
Section titled “Datasets used”| Dataset | Role in paper | Wiki page |
|---|---|---|
| CRSP Survivor-Bias-Free U.S. Mutual Fund Database | Monthly excess returns for 3,630 open-end U.S. domestic equity mutual funds, 1984-2019 | WRDS / CRSP (licensed) |
| Kenneth French Data Library | Four-factor model returns (RMRF, SMB, HML, MOM) used as benchmark factors in regression (2.1) | Ken French library |
Sample: January 1984 to May 2019, N = 3,630 funds, T = 425 months (unbalanced panel). Data access provided via the University at Albany Center for Institutional Investment Management (CIIM).
When to read the full paper
Section titled “When to read the full paper”Use the original if you are: implementing a bootstrap procedure for mutual fund or hedge fund performance in an unbalanced panel (the CSDWB algorithm is fully described in Section 2.3 and Internet Appendices); diagnosing why the Fama-French (2010) method undersizes (Sections 2.1-2.2 with the dominant-effect analysis); comparing CSDWB against HL, KTWW, CSDB, and FF across simulation designs (Section 3); or applying the MCS-based fund-selection Algorithm 1 to build a best-funds portfolio.
Attribution and rights
Section titled “Attribution and rights”Source: peer-reviewed, Journal of Empirical Finance 85 (2026) 101673. Paywalled; no open-access or CC licence. This distillation was extracted by an LLM on 2026-06-25 and is not human-verified or independently reproduced. Quotation of results is for scholarly reference only (extract-only).
Hounyo, Ulrich, and Jiahao Lin. “Can mutual fund ‘stars’ really pick stocks? New evidence from a wild bootstrap analysis.” Journal of Empirical Finance 85 (2026): 101673. DOI: 10.1016/j.jempfin.2025.101673. © 2025 Elsevier B.V. All rights reserved.