Skip to content

Equilibrium Data Mining and Data Abundance: Dugast & Foucault (2025)

Distilled by claude-sonnet-4-6 · extracted Jun 6, 2026, verified Jun 6, 2026

JEL (IAR-assigned): G14, G23, D83 · assigned from the abstract, not the journal

Full structured metadata (methods, scope, relatesTo, topics, datasets): raw Markdown (.md)

paper-summaryasset-pricinginformation-economicsmarket-microstructurequant-fundsinstitutional-investorspeer-reviewedunreplicated

What this is. The paper’s core propositions, the model it builds on (a noisy rational-expectations equilibrium with two types of asset managers), and the mechanism it establishes (the price-informativeness effect dominates the hidden-gold-nugget effect for a large enough data frontier): enough to know what it found and how, without reading all 48 pages. To replicate or extend it, read the full source at the original.

The paper builds a rational-expectations equilibrium model to study how the big data revolution affects the market for active asset management. Two types of managers coexist: “experts” (discretionary funds) with a fixed signal precision, and “data miners” (quant funds) who discover predictors through a sequential search process. The key finding is that the two dimensions of the big data revolution, lower information-processing costs (cc) and a larger data frontier (τdmmax\tau_{dm}^{\text{max}}, reflecting more available data sets), have asymmetric and sometimes opposite effects. Reducing search costs always raises data miners’ search intensity and capital allocated to quants. But a larger data frontier can reduce quant search intensity and the allocation to quants once it is large enough, because greater price informativeness erodes the value of any given signal. Despite this, a larger data frontier always raises price informativeness. Asset managers’ average gross performance is hump-shaped in both cc and τdmmax\tau_{dm}^{\text{max}}, implying the big data revolution should eventually erode active managers’ performance.

#ResultLocatorMagnitude
R1An asset manager’s optimal position is proportional to the gap between her signal and the asset price; the equilibrium price is a sufficient statistic for the aggregate demandProposition 1, eq. 11-13, p. 223Trading aggressiveness β(τ)=τρσω2\beta(\tau) = \frac{\tau}{\rho \sigma_\omega^2}; equilibrium price p=λ(τ)ξp^* = \lambda(\tau^*) \xi where λ(τ)τˉ2τˉ2+ρ2σω4ση2\lambda(\tau^*) \equiv \frac{\bar{\tau}^2}{\bar{\tau}^2 + \rho^2 \sigma_\omega^4 \sigma_\eta^2}
R2Price informativeness always increases with average signal quality and therefore with data miners’ search intensity τ\tau^*Lemma 1, eq. 14, p. 224I(τ;τdmmax)=Var[ωp]1=σω2+τˉ2/(ρ2σω4ση2)\mathcal{I}(\tau^*; \tau_{dm}^{\text{max}}) = \text{Var}[\omega \mid p^*]^{-1} = \sigma_\omega^{-2} + \bar{\tau}^2 / (\rho^2 \sigma_\omega^4 \sigma_\eta^2); strictly increasing in τ\tau^*
R3A decrease in data miners’ search costs cc always increases data miners’ equilibrium search intensity τ\tau^*, raising capital allocated to quants (μ\mu^*), average signal quality, and price informativenessProposition 4, p. 227τ/c<0\partial \tau^* / \partial c < 0; τ\tau^* converges to τdmmax\tau_{dm}^{\text{max}} as c0c \to 0; capital to data miners μ=Γ(τ)\mu^* = \Gamma(\tau^*) increases with τ\tau^*
R4A larger data frontier (τdmmax\tau_{dm}^{\text{max}}) reduces data miners’ search intensity τ\tau^* and capital allocated to quants (μ\mu^*) once τdmmax\tau_{dm}^{\text{max}} exceeds a threshold τtr(c)\tau^{tr}(c), because the price informativeness effect dominates the hidden gold-nugget effectProposition 5, p. 227Threshold τtr(c)\tau^{tr}(c) exists for all c>0c > 0; for τdmmax>τtr(c)\tau_{dm}^{\text{max}} > \tau^{tr}(c): τ/τdmmax<0\partial \tau^* / \partial \tau_{dm}^{\text{max}} < 0 and μ/τdmmax<0\partial \mu^* / \partial \tau_{dm}^{\text{max}} < 0
R5Despite reducing quant search intensity when τdmmax\tau_{dm}^{\text{max}} is large, a push back of the data frontier always raises average signal quality τˉ\bar{\tau} and therefore price informativeness I\mathcal{I}; price informativeness is bounded above as τdmmax\tau_{dm}^{\text{max}} \to \inftyProposition 5, p. 227I/τdmmax>0\partial \mathcal{I} / \partial \tau_{dm}^{\text{max}} > 0 always; I\mathcal{I} bounded above as τdmmax\tau_{dm}^{\text{max}} \to \infty (Assumption 1 + proof)
R6Asset managers’ average gross excess return is hump-shaped in search costs cc and in the data frontier τdmmax\tau_{dm}^{\text{max}}; the big data revolution is predicted to first raise, then reduce, average active management performanceCorollary 1, Figure 3, pp. 232-233E[Rˉe(τ)]=1W0ρ(1τˉ+τˉρ2σω2ση2)1\mathbb{E}[\bar{R}^e(\tau)] = \frac{1}{W_0 \rho} \left( \frac{1}{\bar{\tau}} + \frac{\bar{\tau}}{\rho^2 \sigma_\omega^2 \sigma_\eta^2} \right)^{-1}; peaks at τˉ=ρσωση\bar{\tau} = \rho \sigma_\omega \sigma_\eta; hump-shaped in cc and τdmmax\tau_{dm}^{\text{max}}
R7In the fee extension (Nash bargaining, κ>0\kappa > 0), data miners charge no rents; experts’ fees are set by their scarcity and decline when data miners’ search intensity rises, whether from lower cc or a data-frontier increase that raises τ\tau^*Corollary 6, Section VI, eq. 39-41, pp. 237-238fdm=0f_{dm}^* = 0; fex(τ)=κ(w(τ)w(τ))f_{ex}^*(\tau) = \kappa(w(\tau) - w(\tau^*)); experts’ fees decline with a fall in cc and decline (for low-skill experts) or may rise (high-skill experts) with τdmmax\tau_{dm}^{\text{max}}

Overall (paper’s conclusion). The two dimensions of the big data revolution, lower data-processing costs and data abundance, have the same effect on the allocation of capital to quants when τdmmaxτtr(c)\tau_{dm}^{\text{max}} \leq \tau^{tr}(c) (both increase μ\mu^*) but opposite effects when τdmmax>τtr(c)\tau_{dm}^{\text{max}} > \tau^{tr}(c) (lower cc raises μ\mu^*; larger τdmmax\tau_{dm}^{\text{max}} reduces μ\mu^*). The model predicts that the rise of quant funds driven by both forces should eventually reverse as data abundance grows, and that average active management performance is eventually eroded by greater price informativeness. Distinguishing these two dimensions of the big data revolution is therefore essential for empirical analysis.

The model has four periods (Figure 1, p. 218). Period 0: investors (mass one) allocate savings W0W_0 to either an expert or a data miner, as in Garleanu and Pedersen (2018). Period 1: data miners conduct sequential search for a predictor. Period 2: trading occurs. Period 3: the risky asset payoff ωN(0,σω2)\omega \sim \mathcal{N}(0, \sigma_\omega^2) is realized.

Signals. All asset managers receive a noisy signal before trading (eq. 1, p. 218):

sτi=ω+τi1/2εi,εiN(0,σω2)(1)s_{\tau_i} = \omega + \tau_i^{-1/2} \varepsilon_i, \qquad \varepsilon_i \sim \mathcal{N}(0, \sigma_\omega^2) \tag{1}

where τi\tau_i is signal precision (“quality”). Experts’ skill τ\tau is fixed and drawn from cumulative distribution Γ()\Gamma(\cdot) (density γ()\gamma(\cdot)) on [0,τexmax][0, \tau_{ex}^{\text{max}}]. Data miners discover their predictor through search: each round costs cc and yields a precision draw τΦ()\tau \sim \Phi(\cdot) on [0,τdmmax][0, \tau_{dm}^{\text{max}}] (eq. 2, p. 219):

Φ(τ)=Pr(τ~τ)=Ψ(τ)Ψ(τdmmax),τ[0,τdmmax](2)\Phi(\tau) = \Pr(\tilde{\tau} \leq \tau) = \frac{\Psi(\tau)}{\Psi(\tau_{dm}^{\text{max}})}, \qquad \tau \in [0, \tau_{dm}^{\text{max}}] \tag{2}

The data frontier τdmmax\tau_{dm}^{\text{max}} captures the maximum attainable signal precision from available data sets; higher τdmmax\tau_{dm}^{\text{max}} reflects data abundance. A data miner with stopping threshold τi\tau_i^* stops when her draw exceeds τi\tau_i^*, so the likelihood of stopping in a given round is:

Λ(τi;τdmmax)Pr(τ[τi,τdmmax])=1Φ(τi)(3)\Lambda(\tau_i^*; \tau_{dm}^{\text{max}}) \equiv \Pr(\tau \in [\tau_i^*, \tau_{dm}^{\text{max}}]) = 1 - \Phi(\tau_i^*) \tag{3}

A higher τi\tau_i^* means more demanding search (fewer stops per round on average), so τ\tau^* is called the “search intensity.”

Capital allocation. Let μ\mu denote the fraction of investor capital allocated to data miners. Investors observe experts’ skills and anticipate data miners’ search strategy. They optimally allocate to experts with skill ττ\tau \geq \underline{\tau} (the marginal expert) until each expert is at capacity. The rest goes to data miners. In a stable interior equilibrium (Proposition 3, eq. 8, p. 222):

μ=Γ(τ)(8)\mu^* = \Gamma(\tau^*) \tag{8}

where τ\tau^* is data miners’ equilibrium search intensity, and the marginal expert skill equals τ\tau^*.

Trading. The market for the risky asset is as in Vives (1995): noise traders with aggregate demand ηN(0,ση2)\eta \sim \mathcal{N}(0, \sigma_\eta^2) trade alongside asset managers. Price informativeness is measured by the inverse of residual payoff variance, following Grossman and Stiglitz (1980) and Verrecchia (1982). Risk-neutral dealers post a price equal to their expectation of the payoff conditional on aggregate demand (eq. 4, p. 221):

p=E[ωD(p)](4)p^* = \mathbb{E}[\omega \mid D(p^*)] \tag{4}

Asset managers have constant absolute risk aversion ρ\rho. Asset manager ii returns to her client (eq. 5, p. 221):

Wi,j=W0+xi(sτi,p)(ωp)(nic)1{j=dm}(5)W_{i,j} = W_0 + x_i(s_{\tau_i}, p)(\omega - p) - (n_i c)\mathbb{1}_{\{j=dm\}} \tag{5}

where nin_i is the number of search rounds conducted. Investor utility from trading with an expert of skill τ\tau is (eq. 6, p. 221):

H(τ)=E[exp(ρ(W0+xi(sτi,p)(ωp)))](6)H(\tau) = \mathbb{E}\left[ -\exp\left( -\rho(W_0 + x_i(s_{\tau_i}, p)(\omega - p)) \right) \right] \tag{6}

and investor utility from trading with a data miner of search intensity τi\tau_i^* is (eq. 7, p. 221):

V(τi)=E[exp(ρ(W0+xi(sτi,p)(ωp)))]Expected utility from trading×E[exp(ρ(nic))]Expected utility cost of exploration(7)V(\tau_i^*) = \underbrace{\mathbb{E}\left[ -\exp\left(-\rho(W_0 + x_i(s_{\tau_i}, p)(\omega - p))\right) \right]}_{\text{Expected utility from trading}} \times \underbrace{\mathbb{E}\left[\exp(\rho(n_i c))\right]}_{\text{Expected utility cost of exploration}} \tag{7}

The paper characterizes the equilibrium in three steps.

Step 1: Trading equilibrium (Proposition 1, p. 223). Taking τ\tau^* (and hence μ\mu^*) as given, the paper solves for the trading equilibrium. In equilibrium, each asset manager’s demand is proportional to her signal minus the price (eq. 11):

x(sτ,p)=β(τ)(sτp),β(τ)=τρσω2(11)x^*(s_\tau, p) = \beta(\tau)(s_\tau - p), \qquad \beta(\tau) = \frac{\tau}{\rho \sigma_\omega^2} \tag{11}

The equilibrium price (eq. 12) is:

p=E[ωD(p)]=λ(τ)ξ,ξω+ρσω2τˉ(τ;τdmmax)1η(12)p^* = \mathbb{E}[\omega \mid D(p^*)] = \lambda(\tau^*)\xi, \qquad \xi \equiv \omega + \rho \sigma_\omega^2 \bar{\tau}(\tau^*; \tau_{dm}^{\text{max}})^{-1} \eta \tag{12}

where τˉ(τ;τdmmax)\bar{\tau}(\tau^*; \tau_{dm}^{\text{max}}) is the average signal quality across all asset managers and λ(τ)\lambda(\tau^*) is defined in eq. 13. Price informativeness is (eq. 14, p. 224):

I(τ;τdmmax)Var[ωp]1=1σω2+τˉ(τ;τdmmax)2ρ2σω4ση2(14)\mathcal{I}(\tau^*; \tau_{dm}^{\text{max}}) \equiv \text{Var}[\omega \mid p^*]^{-1} = \frac{1}{\sigma_\omega^2} + \frac{\bar{\tau}(\tau^*; \tau_{dm}^{\text{max}})^2}{\rho^2 \sigma_\omega^4 \sigma_\eta^2} \tag{14}

Step 2: Equilibrium data mining (Proposition 2, p. 225). The trading value of a signal of quality τ\tau is (Lemma 2, eq. 16, p. 224):

g(τ,τ)=(1+τσω2I(τ;τdmmax))12(16)g(\tau, \tau^*) = -\left(1 + \frac{\tau}{\sigma_\omega^2 \mathcal{I}(\tau^*; \tau_{dm}^{\text{max}})}\right)^{-\frac{1}{2}} \tag{16}

A data miner’s continuation value after finding and then rejecting a predictor of quality τ^i\hat{\tau}_i is (eq. 17-18, pp. 224-225):

J(τ^i,τ)=exp(ρc)Λ(τ^i;τdmmax)1exp(ρc)(1Λ(τ^i;τdmmax))×Eϕ[g(τ,τ)τ^iττdmmax](18)J(\hat{\tau}_i, \tau^*) = \frac{\exp(\rho c) \Lambda(\hat{\tau}_i; \tau_{dm}^{\text{max}})}{1 - \exp(\rho c)(1 - \Lambda(\hat{\tau}_i; \tau_{dm}^{\text{max}}))} \times \mathbb{E}_\phi\left[g(\tau, \tau^*) \mid \hat{\tau}_i \leq \tau \leq \tau_{dm}^{\text{max}}\right] \tag{18}

In a symmetric equilibrium, τ\tau^* solves g(τ,τ)=J(τ,τ)g(\tau^*, \tau^*) = J(\tau^*, \tau^*), which reduces to (eq. 21, p. 225):

F(τ)=exp(ρc),(21)F(\tau^*) = \exp(-\rho c), \tag{21}

where

F(τ)ττdmmaxr(τ,τ)ϕ(τ)dτ+(1Λ(τ;τdmmax)),r(τ,τ)(τ+σω2I(τ;τdmmax)τ+σω2I(τ;τdmmax))12(22-23)F(\tau^*) \equiv \int_{\tau^*}^{\tau_{dm}^{\text{max}}} r(\tau, \tau^*)\phi(\tau)d\tau + (1 - \Lambda(\tau^*; \tau_{dm}^{\text{max}})), \qquad r(\tau, \tau^*) \equiv \left(\frac{\tau^* + \sigma_\omega^2 \mathcal{I}(\tau^*; \tau_{dm}^{\text{max}})}{\tau + \sigma_\omega^2 \mathcal{I}(\tau^*; \tau_{dm}^{\text{max}})}\right)^{\frac{1}{2}} \tag{22-23}

Proposition 2 establishes that this equation has a unique solution τ(0,τdmmax)\tau^* \in (0, \tau_{dm}^{\text{max}}) whenever F(0)<exp(ρc)F(0) < \exp(-\rho c).

Step 3: Full equilibrium (Proposition 3, p. 226). Combining Steps 1 and 2 with the capital-allocation condition μ=Γ(τ)\mu^* = \Gamma(\tau^*) (eq. 8) yields the full equilibrium characterization. The model is solved analytically under the parameterization Φ(τ)=1(1+τ)3/21(1+τdmmax)3/2\Phi(\tau) = \frac{1-(1+\tau)^{-3/2}}{1-(1+\tau_{dm}^{\text{max}})^{-3/2}} and Γ(τ)=1(1+τ)3/2\Gamma(\tau) = 1-(1+\tau)^{-3/2} for numerical illustrations (Figure 2, p. 230).

The decomposition of the data-frontier effect (eq. A26, p. 250 Appendix; related text at eq. 25, pp. 228-229) separates the hidden gold-nugget effect (a larger τdmmax\tau_{dm}^{\text{max}} raises the value of the best possible predictor) from the price informativeness effect (more data raises I\mathcal{I}, reducing the value of any given signal). The second term captures the price informativeness effect:

Fτdmmax=ϕ(τdmmax)(r(τdmmax,τ)Eϕ[min{1,r(τ,τ)}])<0:Gold Nugget Effect+(ττdmmaxr(τ,τ)Iϕ(τ)dτ)Iτdmmax>0:Informativeness Effect(26)\frac{\partial F}{\partial \tau_{dm}^{\text{max}}} = \underbrace{\phi(\tau_{dm}^{\text{max}})(r(\tau_{dm}^{\text{max}}, \tau^*) - \mathbb{E}_\phi[\min\{1, r(\tau, \tau^*)\}])}_{<0: \text{Gold Nugget Effect}} + \underbrace{\left(\int_{\tau^*}^{\tau_{dm}^{\text{max}}} \frac{\partial r(\tau, \tau^*)}{\partial \mathcal{I}} \phi(\tau)d\tau\right) \frac{\partial \mathcal{I}}{\partial \tau_{dm}^{\text{max}}}}_{>0: \text{Informativeness Effect}} \tag{26}

When τdmmax\tau_{dm}^{\text{max}} is large enough, the positive informativeness effect dominates, so F/τdmmax>0\partial F / \partial \tau_{dm}^{\text{max}} > 0 and τ\tau^* decreases with τdmmax\tau_{dm}^{\text{max}} (Proposition 5).

This is a pure theory paper. It derives no estimating equations and runs no regressions. The empirical implications are summarized in Table I (p. 243), which characterizes the directional effects of lower search costs (cc \searrow) and data abundance (τdmmax\tau_{dm}^{\text{max}} \nearrow) on:

  • Allocation of capital to data miners (μ\mu^*): increases with lower cc; hump-shaped in τdmmax\tau_{dm}^{\text{max}}
  • Price informativeness (I\mathcal{I}): always increases with both shocks
  • Average signal quality (τˉ\bar{\tau}): always increases with both shocks
  • Data miners’ relative performance (RPRP): increases with τdmmax\tau_{dm}^{\text{max}} (Corollary 4); ambiguous with lower cc
  • Within-group performance dispersion (ΔRα\Delta R_\alpha): decreases with lower cc; increases with τdmmax\tau_{dm}^{\text{max}} above threshold (Corollaries 2-3)
  • Average performance (E[Rˉe]\mathbb{E}[\bar{R}^e]): hump-shaped in both cc and τdmmax\tau_{dm}^{\text{max}} (Corollary 1)

Section VII (pp. 242-244) suggests three types of empirical tests: (i) cross-sectional variation in quant and discretionary fund holdings to capture differential exposure to alternative data (shocks to τdmmax\tau_{dm}^{\text{max}}), exploiting the finding of Abis (2022) that quant funds grew from 6.1% to 18.6% of U.S. equity AUM between 2000 and 2017; (ii) regulatory changes that reduce information processing costs (e.g., the SEC’s XBLR mandate lowering cc for IT-intensive funds, consistent with evidence from Zhao (2021) that this mandate reduced the performance gap between quant and discretionary funds); and (iii) the introduction of cloud computing (e.g., Amazon Web Services in 2006) as shocks to cc. Measurement of signal quality τ\tau follows Proposition 1: the theoretical coefficient from regressing a fund’s holdings x(sτ,p)x^*(s_\tau, p^*) on ωp\omega - p^* is β(τ)=τ/(ρσω2)\beta(\tau) = \tau / (\rho \sigma_\omega^2) (eq. 45, p. 244), which is strictly positive for informed managers and increases with skill. The paper also relates to evidence from Pastor, Stambaugh, and Taylor (2015) that the size of the active management industry is negatively related to funds’ performance. Han and Sangiorgi (2018) and Banerjee and Breon-Drish (2021) model information acquisition as search but analyze different questions; the key distinction here is the simultaneous variation in both the intensive and extensive margins. Stambaugh (2020) documents a related result that improved manager skills reduce average performance, which the model endogenizes through the price informativeness channel.

This paper uses no empirical datasets. All results are derived analytically from the theoretical model. Numerical illustrations use the parametric family Φ(τ)=1(1+τ)3/21(1+τdmmax)3/2\Phi(\tau) = \frac{1-(1+\tau)^{-3/2}}{1-(1+\tau_{dm}^{\text{max}})^{-3/2}} and Γ(τ)=1(1+τ)3/2\Gamma(\tau) = 1-(1+\tau)^{-3/2}, with calibrated values for ρ\rho, σω\sigma_\omega, ση\sigma_\eta, W0W_0, and cc (Figures 2-6, pp. 230-239).

DatasetRole in paperWiki page
No empirical data usedTheory paper onlyn/a

Use the original if you are: building a model of quant fund behavior in equilibrium; studying the welfare effects of the big data revolution on price efficiency and active management; looking for testable predictions on the cross-sectional variation in fund performance as a function of data availability; or extending the model to allow learning about signal quality (fn. 10, p. 220) or non-extreme decreasing returns to scale (Section II.E, Internet Appendix). The appendix (pp. 245-252) contains the full proofs of all propositions.

Source: peer-reviewed, The Journal of Finance 80(1). This distillation was extracted by an LLM on 2026-06-06 and is not human-verified or independently reproduced. The CC BY-NC 4.0 licence permits sharing with attribution for non-commercial purposes; the verbatim PDF is not hosted in this batch.

Citation. Dugast, Jérôme, and Thierry Foucault. “Equilibrium Data Mining and Data Abundance.” The Journal of Finance 80, no. 1 (February 2025): 211-258. DOI: 10.1111/jofi.13397. CC BY-NC 4.0. This page is an extract by the Institute for Automated Research: core results re-expressed for research reference; changes were made.

Found an error or want a topic covered? Open an issue, use the Edit page link above, or email contact@instituteforautomatedresearch.org. Edits are reviewed before publishing; provenance and accuracy are the point.