Equilibrium Data Mining and Data Abundance: Dugast & Foucault (2025)
Distilled by claude-sonnet-4-6 · extracted Jun 6, 2026, verified Jun 6, 2026
JEL (IAR-assigned): G14, G23, D83 · assigned from the abstract, not the journal
What this is. The paper’s core propositions, the model it builds on (a noisy rational-expectations equilibrium with two types of asset managers), and the mechanism it establishes (the price-informativeness effect dominates the hidden-gold-nugget effect for a large enough data frontier): enough to know what it found and how, without reading all 48 pages. To replicate or extend it, read the full source at the original.
The paper builds a rational-expectations equilibrium model to study how the big data revolution affects the market for active asset management. Two types of managers coexist: “experts” (discretionary funds) with a fixed signal precision, and “data miners” (quant funds) who discover predictors through a sequential search process. The key finding is that the two dimensions of the big data revolution, lower information-processing costs () and a larger data frontier (, reflecting more available data sets), have asymmetric and sometimes opposite effects. Reducing search costs always raises data miners’ search intensity and capital allocated to quants. But a larger data frontier can reduce quant search intensity and the allocation to quants once it is large enough, because greater price informativeness erodes the value of any given signal. Despite this, a larger data frontier always raises price informativeness. Asset managers’ average gross performance is hump-shaped in both and , implying the big data revolution should eventually erode active managers’ performance.
Core results
Section titled “Core results”| # | Result | Locator | Magnitude |
|---|---|---|---|
| R1 | An asset manager’s optimal position is proportional to the gap between her signal and the asset price; the equilibrium price is a sufficient statistic for the aggregate demand | Proposition 1, eq. 11-13, p. 223 | Trading aggressiveness ; equilibrium price where |
| R2 | Price informativeness always increases with average signal quality and therefore with data miners’ search intensity | Lemma 1, eq. 14, p. 224 | ; strictly increasing in |
| R3 | A decrease in data miners’ search costs always increases data miners’ equilibrium search intensity , raising capital allocated to quants (), average signal quality, and price informativeness | Proposition 4, p. 227 | ; converges to as ; capital to data miners increases with |
| R4 | A larger data frontier () reduces data miners’ search intensity and capital allocated to quants () once exceeds a threshold , because the price informativeness effect dominates the hidden gold-nugget effect | Proposition 5, p. 227 | Threshold exists for all ; for : and |
| R5 | Despite reducing quant search intensity when is large, a push back of the data frontier always raises average signal quality and therefore price informativeness ; price informativeness is bounded above as | Proposition 5, p. 227 | always; bounded above as (Assumption 1 + proof) |
| R6 | Asset managers’ average gross excess return is hump-shaped in search costs and in the data frontier ; the big data revolution is predicted to first raise, then reduce, average active management performance | Corollary 1, Figure 3, pp. 232-233 | ; peaks at ; hump-shaped in and |
| R7 | In the fee extension (Nash bargaining, ), data miners charge no rents; experts’ fees are set by their scarcity and decline when data miners’ search intensity rises, whether from lower or a data-frontier increase that raises | Corollary 6, Section VI, eq. 39-41, pp. 237-238 | ; ; experts’ fees decline with a fall in and decline (for low-skill experts) or may rise (high-skill experts) with |
Overall (paper’s conclusion). The two dimensions of the big data revolution, lower data-processing costs and data abundance, have the same effect on the allocation of capital to quants when (both increase ) but opposite effects when (lower raises ; larger reduces ). The model predicts that the rise of quant funds driven by both forces should eventually reverse as data abundance grows, and that average active management performance is eventually eroded by greater price informativeness. Distinguishing these two dimensions of the big data revolution is therefore essential for empirical analysis.
Theory / model
Section titled “Theory / model”The model has four periods (Figure 1, p. 218). Period 0: investors (mass one) allocate savings to either an expert or a data miner, as in Garleanu and Pedersen (2018). Period 1: data miners conduct sequential search for a predictor. Period 2: trading occurs. Period 3: the risky asset payoff is realized.
Signals. All asset managers receive a noisy signal before trading (eq. 1, p. 218):
where is signal precision (“quality”). Experts’ skill is fixed and drawn from cumulative distribution (density ) on . Data miners discover their predictor through search: each round costs and yields a precision draw on (eq. 2, p. 219):
The data frontier captures the maximum attainable signal precision from available data sets; higher reflects data abundance. A data miner with stopping threshold stops when her draw exceeds , so the likelihood of stopping in a given round is:
A higher means more demanding search (fewer stops per round on average), so is called the “search intensity.”
Capital allocation. Let denote the fraction of investor capital allocated to data miners. Investors observe experts’ skills and anticipate data miners’ search strategy. They optimally allocate to experts with skill (the marginal expert) until each expert is at capacity. The rest goes to data miners. In a stable interior equilibrium (Proposition 3, eq. 8, p. 222):
where is data miners’ equilibrium search intensity, and the marginal expert skill equals .
Trading. The market for the risky asset is as in Vives (1995): noise traders with aggregate demand trade alongside asset managers. Price informativeness is measured by the inverse of residual payoff variance, following Grossman and Stiglitz (1980) and Verrecchia (1982). Risk-neutral dealers post a price equal to their expectation of the payoff conditional on aggregate demand (eq. 4, p. 221):
Asset managers have constant absolute risk aversion . Asset manager returns to her client (eq. 5, p. 221):
where is the number of search rounds conducted. Investor utility from trading with an expert of skill is (eq. 6, p. 221):
and investor utility from trading with a data miner of search intensity is (eq. 7, p. 221):
Method
Section titled “Method”The paper characterizes the equilibrium in three steps.
Step 1: Trading equilibrium (Proposition 1, p. 223). Taking (and hence ) as given, the paper solves for the trading equilibrium. In equilibrium, each asset manager’s demand is proportional to her signal minus the price (eq. 11):
The equilibrium price (eq. 12) is:
where is the average signal quality across all asset managers and is defined in eq. 13. Price informativeness is (eq. 14, p. 224):
Step 2: Equilibrium data mining (Proposition 2, p. 225). The trading value of a signal of quality is (Lemma 2, eq. 16, p. 224):
A data miner’s continuation value after finding and then rejecting a predictor of quality is (eq. 17-18, pp. 224-225):
In a symmetric equilibrium, solves , which reduces to (eq. 21, p. 225):
where
Proposition 2 establishes that this equation has a unique solution whenever .
Step 3: Full equilibrium (Proposition 3, p. 226). Combining Steps 1 and 2 with the capital-allocation condition (eq. 8) yields the full equilibrium characterization. The model is solved analytically under the parameterization and for numerical illustrations (Figure 2, p. 230).
The decomposition of the data-frontier effect (eq. A26, p. 250 Appendix; related text at eq. 25, pp. 228-229) separates the hidden gold-nugget effect (a larger raises the value of the best possible predictor) from the price informativeness effect (more data raises , reducing the value of any given signal). The second term captures the price informativeness effect:
When is large enough, the positive informativeness effect dominates, so and decreases with (Proposition 5).
Empirical specifications
Section titled “Empirical specifications”This is a pure theory paper. It derives no estimating equations and runs no regressions. The empirical implications are summarized in Table I (p. 243), which characterizes the directional effects of lower search costs () and data abundance () on:
- Allocation of capital to data miners (): increases with lower ; hump-shaped in
- Price informativeness (): always increases with both shocks
- Average signal quality (): always increases with both shocks
- Data miners’ relative performance (): increases with (Corollary 4); ambiguous with lower
- Within-group performance dispersion (): decreases with lower ; increases with above threshold (Corollaries 2-3)
- Average performance (): hump-shaped in both and (Corollary 1)
Section VII (pp. 242-244) suggests three types of empirical tests: (i) cross-sectional variation in quant and discretionary fund holdings to capture differential exposure to alternative data (shocks to ), exploiting the finding of Abis (2022) that quant funds grew from 6.1% to 18.6% of U.S. equity AUM between 2000 and 2017; (ii) regulatory changes that reduce information processing costs (e.g., the SEC’s XBLR mandate lowering for IT-intensive funds, consistent with evidence from Zhao (2021) that this mandate reduced the performance gap between quant and discretionary funds); and (iii) the introduction of cloud computing (e.g., Amazon Web Services in 2006) as shocks to . Measurement of signal quality follows Proposition 1: the theoretical coefficient from regressing a fund’s holdings on is (eq. 45, p. 244), which is strictly positive for informed managers and increases with skill. The paper also relates to evidence from Pastor, Stambaugh, and Taylor (2015) that the size of the active management industry is negatively related to funds’ performance. Han and Sangiorgi (2018) and Banerjee and Breon-Drish (2021) model information acquisition as search but analyze different questions; the key distinction here is the simultaneous variation in both the intensive and extensive margins. Stambaugh (2020) documents a related result that improved manager skills reduce average performance, which the model endogenizes through the price informativeness channel.
Datasets used
Section titled “Datasets used”This paper uses no empirical datasets. All results are derived analytically from the theoretical model. Numerical illustrations use the parametric family and , with calibrated values for , , , , and (Figures 2-6, pp. 230-239).
| Dataset | Role in paper | Wiki page |
|---|---|---|
| No empirical data used | Theory paper only | n/a |
When to read the full paper
Section titled “When to read the full paper”Use the original if you are: building a model of quant fund behavior in equilibrium; studying the welfare effects of the big data revolution on price efficiency and active management; looking for testable predictions on the cross-sectional variation in fund performance as a function of data availability; or extending the model to allow learning about signal quality (fn. 10, p. 220) or non-extreme decreasing returns to scale (Section II.E, Internet Appendix). The appendix (pp. 245-252) contains the full proofs of all propositions.
Attribution and rights
Section titled “Attribution and rights”Source: peer-reviewed, The Journal of Finance 80(1). This distillation was extracted by an LLM on 2026-06-06 and is not human-verified or independently reproduced. The CC BY-NC 4.0 licence permits sharing with attribution for non-commercial purposes; the verbatim PDF is not hosted in this batch.
Citation. Dugast, Jérôme, and Thierry Foucault. “Equilibrium Data Mining and Data Abundance.” The Journal of Finance 80, no. 1 (February 2025): 211-258. DOI: 10.1111/jofi.13397. CC BY-NC 4.0. This page is an extract by the Institute for Automated Research: core results re-expressed for research reference; changes were made.