Skip to content

Confidence, Self-Selection, and Bias in the Aggregate: Enke, Graeber & Oprea (2023)

Distilled by claude-sonnet-4-6 · extracted Jun 25, 2026, verified Jun 25, 2026

JEL (IAR-assigned): C91, D44, D91 · assigned from the abstract, not the journal

Full structured metadata (methods, scope, relatesTo, topics, datasets): raw Markdown (.md)

paper-summarybehavioral-economicscognitive-biasesself-selectionoverconfidenceexperimentalpeer-reviewedunreplicated

What this is. The paper’s core results, the theoretical framework linking confidence to institutional filtering, and the three experimental institutions (parimutuel betting market, discriminatory auction, committee voting) with their defining equations: enough to know what it found and how, without reading all 34 pages. To replicate or extend, read the full source at the original.

Enke, Graeber, and Oprea run a large preregistered online experiment on Prolific (2,153 subjects, June 2021) exposing participants to 15 canonical cognitive biases from behavioral economics and three simple social institutions (betting markets, auctions, committees) that allow voluntary self-selection. They find that institutions filter biases on average, but with large cross-task variation: exponential growth bias (EGB) is reduced by roughly 17 percentage points, while base-rate neglect and correlation neglect are barely affected, and the winner’s curse is even amplified. Almost all of this cross-task heterogeneity (r = 0.76 to 0.93) is explained by a single sufficient statistic: the within-task Pearson correlation between subjects’ stated confidence and their decision optimality. When better performers are also more confident, they self-select more intensively and the institution de-biases effectively. When confidence and performance are uncorrelated or negatively correlated, the institution cannot filter, regardless of average overconfidence levels.

Magnitudes as reported; Locators point into the source PDF.

#ResultLocatorMagnitude
R1Positive self-selection in all three institutions on average across tasks: optimal decision makers bet, bid, and vote more intensively than suboptimal onesFigure 2, p. 1952; p. 1953Betting: 64.8 avg. bet (optimal) vs 47.4 (suboptimal), 37% more; Auction: 56.4 vs 43.6, 29% more; Committee: 75 vs 57.9 votes, 29% more
R2Large cross-task variation in institutional filtering: EGB and iterated reasoning (IR) strongly improved, some tasks near-zero or negativeFigure 3, p. 1954EGB: ~17 pp improvement across institutions; IR: ~8 pp; RM, AC, EQ near-zero or negative (approx. -4 to 0 pp); pairwise correlations across institutions 0.85-0.91
R3Confidence-performance correlation varies widely across tasks: from negative (RM = -0.13, TM significantly negative) to moderately positive (GF = 0.39); 6 of 15 tasks negativeFigure 4, p. 1956Pearson r ranges -0.13 (RM) to 0.39 (GF); N = 334 in Confidence treatment; no task exceeds r = 0.5
R4Confidence-performance correlation strongly predicts institutional improvement across the 15 tasksFigure 5, p. 1957r = 0.76 (between-subjects), r = 0.93 (within-subjects); robust to leave-two-out: between-subjects range 0.61-0.83 (mean 0.76)
R5Predictive power is consistent across all three institutionsp. 1958r^auction = 0.69, r^betting = 0.73, r^committee = 0.77 (between); r^auction,within = 0.90, r^betting,within = 0.90, r^committee,within = 0.91
R6Confidence-performance correlation predicts institutional efficiency (fraction of theoretically possible improvement realized) even more stronglyp. 1959r = 0.87 (between-subjects), r = 0.94 (within-subjects)
R7Average overconfidence (d = c - p) shows a weak, statistically insignificant negative relationship with institutional improvementp. 1961r = -0.34 (between-subjects), r = -0.32 (within-subjects); neither significantly different from 0 at conventional levels

Overall (paper’s conclusion). Average overconfidence, the traditional focus of most confidence research, is largely irrelevant for predicting whether markets and organizations de-bias economic aggregates. The relevant object is the confidence-performance correlation, which determines whether the biased individuals who self-select out of institutions are actually the ones making worse decisions. This implies a simple methodological blueprint: researchers studying cognitive biases can estimate the likely institutional impact by appending an unincentivized confidence question and reporting the resulting correlation with performance.

The paper lays out a simple analytical framework (Section II, pp. 1947-1950) to derive testable predictions about institutional filtering. There is no fully specified equilibrium model; the framework is a linear approximation designed to connect observable confidence and performance to institutional outcomes.

Setup. Each of NN agents forms a judgment on a cognitive task. Agent ii‘s solution is optimal (Xi=1X_i = 1) with probability pip_i and incorrect (Xi=0X_i = 0) with probability 1pi1 - p_i. Pre-institutional aggregate performance equals the raw optimality rate

Θpre=1Ni=1NXi,θpreE[Θpre]=1Ni=1Npi.\Theta^{\text{pre}} = \frac{1}{N}\sum_{i=1}^N X_i, \qquad \theta^{\text{pre}} \equiv E[\Theta^{\text{pre}}] = \frac{1}{N}\sum_{i=1}^N p_i.

Each agent then makes an institutional decision ki[0,1]k_i \in [0, 1], representing bet intensity, bid size, or vote share, depending on the institution. Institutional filtering G=θpostθpre\mathbb{G} = \theta^{\text{post}} - \theta^{\text{pre}} is positive when the institution makes aggregate outcomes appear as if participants were more rational than they actually are.

Institutional performance metrics (pp. 1948). For betting markets and committees, the performance metric is a weighted average of agents’ optimality, weighted by their participation intensity:

θbet,compost=ikipiiki.(5)\theta^{\text{post}}_{\text{bet,com}} = \frac{\sum_i k_i p_i}{\sum_i k_i}. \tag{5}

The expected institutional gain is

Gbet,com=θpostθpre=ipi(kikˉ)Nkˉ.(6)\mathbb{G}_{\text{bet,com}} = \theta^{\text{post}} - \theta^{\text{pre}} = \frac{\sum_i p_i (k_i - \bar{k})}{N\bar{k}}. \tag{6}

This expression is positive if and only if better-performing agents participate more intensively than the average. For auctions, only the top-5 bidders win and the metric is their average optimality rate, so

Gauc=1WjΩpj1Nipi,(7)\mathbb{G}_{\text{auc}} = \frac{1}{|W|}\sum_{j \in \Omega} p_j - \frac{1}{N}\sum_i p_i, \tag{7}

where Ω\Omega is the set of winners (highest five bids).

Confidence and performance (p. 1949). Institutional self-selection is assumed to depend on agents’ stated confidence cic_i about the ex ante optimality of their decision. The paper models the within-task confidence-performance relationship as approximately linear:

ci=α+βpi.(8)c_i = \alpha + \beta \cdot p_i. \tag{8}

The slope β\beta is the confidence-performance correlation (the key object). Average overconfidence is d=cˉpˉ=α+(β1)pˉd = \bar{c} - \bar{p} = \alpha + (\beta - 1)\bar{p}, which combines both the intercept α\alpha and the slope β\beta. Institutional self-selection is taken to be proportional to confidence:

ki=ωci[0,1].(9)k_i = \omega \cdot c_i \in [0,1]. \tag{9}

Here ω>0\omega > 0 captures the degree to which self-selection actually depends on confidence as opposed to other factors.

Predictions. Substituting equations (8) and (9) into (6) and (7) yields two preregistered predictions (pp. 1950):

  • Prediction 1: If β>0\beta > 0, then G>0\mathbb{G} > 0 (institutions filter biases). Institutional improvement G\mathbb{G} increases in the confidence-performance correlation β\beta.
  • Prediction 2: The effect of average overconfidence dd on G\mathbb{G} is ambiguous. In auctions, there is no relationship (only the ordering of bids matters, not their level). In betting and committees with β>0\beta > 0, the effect of dd is weakly negative.

Fehr and Tyran (2005) provide foundational evidence that individual irrationality sometimes survives in aggregate market outcomes, and sometimes does not; this framework clarifies that the confidence-performance correlation is the sufficient statistic for predicting which case applies. The paper complements List (2003), who shows that market experience reduces anomalies through learning; here the channel is purely self-selection, which operates even in the absence of feedback or repeated play.

The experiment (Section I, pp. 1938-1946) implements three maximally simple static variants of canonical economic institutions to isolate the self-selection mechanism. All three are implemented on Prolific with identical slider-based interfaces (0-100) so the self-selection decision is comparable across institutions. The three institutions and their performance metrics are:

Betting market (parimutuel), p. 1942. Ten subjects are grouped into a parimutuel betting market. Each subject ii bets bi[0,100]b_i \in [0, 100] ECUs on the proposition that her own part-1 response was optimal. The market price on the optimal-decision security is

θBetting=i=110xibii=110bi[0,1].(1)\theta^{\text{Betting}} = \frac{\sum_{i=1}^{10} x_i b_i}{\sum_{i=1}^{10} b_i} \in [0,1]. \tag{1}

If subject ii‘s part-1 decision was optimal, her payoff is

πiBetting=biθBetting+(100bi).(2)\pi_i^{\text{Betting}} = \frac{b_i}{\theta^{\text{Betting}}} + (100 - b_i). \tag{2}

The market price in equation (1) is simply a reweighting of individual part-1 decisions xix_i by how much each subject bets; if no self-selection occurs (everyone bets equally), the price equals the raw optimality rate.

Discriminatory auction (5 winners), p. 1943. Subjects submit sealed bids bi[0,100]b_i \in [0, 100] ECUs. The five highest bidders win and receive a bonus of 100 ECUs if their part-1 decision was optimal. The institutional performance metric is the optimality rate among winners:

θAuction=iΩxi5,(3)\theta^{\text{Auction}} = \frac{\sum_{i \in \Omega} x_i}{5}, \tag{3}

where Ω\Omega is the set of five highest bidders. Under standard assumptions this auction implements an efficient allocation to the highest-value bidders (Krishna 2009).

Utilitarian committee voting, p. 1943. Each subject receives 100 votes and submits vi[0,100]v_i \in [0, 100] votes in favor of her own part-1 answer. The fraction of votes on the optimal answer is

θCommittee=i=110Xivii=110vi[0,1].(4)\theta^{\text{Committee}} = \frac{\sum_{i=1}^{10} X_i v_i}{\sum_{i=1}^{10} v_i} \in [0,1]. \tag{4}

All subjects earn 100×θCommittee100 \times \theta^{\text{Committee}} regardless of their own vote.

Confidence elicitation and treatments. After each part-1 cognitive task, subjects in the Confidence treatment (N = 334) and the Within treatments (N = 314) are asked an unincentivized slider question: “How certain are you that your decision in Part 1 was optimal?” (0-100%). In between-subjects treatments (Betting, Auction, Committee), confidence is never elicited; the confidence-performance correlation is measured in a separate Confidence treatment. Table 2 (p. 1945) summarizes the full experimental design: Betting (N = 387), Auction (N = 323), Committee (N = 337), Confidence (N = 334) between-subjects; and Betting Within (N = 105), Auction Within (N = 105), Committee Within (N = 104) within-subjects.

The 15 cognitive tasks (Table 1, p. 1940) cover information processing and statistical reasoning (base rate neglect, correlation neglect, balls-and-urns belief updating, gambler’s fallacy, sample size neglect, regression to mean), logic (Wason task, cognitive reflection test), strategic reasoning (backward induction, equilibrium reasoning), constrained optimization (knapsack), and financial reasoning (thinking at the margin, portfolio choice, exponential growth bias, acquiring a company).

The main empirical analysis works at the task level (15 observations) rather than at the individual level. There is no single regression equation; the core result is a cross-task bivariate correlation.

Step 1: Measure the confidence-performance correlation per task. For each of the 15 tasks kk in the Confidence treatment (between-subjects) or the Within treatments (within-subjects), compute the Pearson correlation between the binary optimality indicator xix_i (part 1) and stated confidence cic_i (part 2):

β^k=Corr(xi,ci)in task k.\hat{\beta}_k = \text{Corr}(x_i, c_i) \quad \text{in task } k.

This is computed separately for each task across the 334 (between) or 314 (within) subjects who see the confidence elicitation.

Step 2: Measure institutional improvement per task. For each task kk and institution j{Betting, Auction, Committee}j \in \{\text{Betting, Auction, Committee}\}, simulate 10,000 random 10-subject cohorts by drawing with replacement from the pool of part-1 and part-2 decisions. Compute θk,jpost\theta^{\text{post}}_{k,j} for each cohort using equations (1)-(4), and compare to the cohort’s raw optimality rate θkpre\theta^{\text{pre}}_k. The institutional improvement is the mean of θpostθpre\theta^{\text{post}} - \theta^{\text{pre}} over the 10,000 cohorts. Standard errors are computed conservatively as the standard deviation of cohort-level improvements divided by N/10\sqrt{N/10}, where NN is the treatment sample size (e.g., 387/10=38.7387/10 = 38.7 cohorts in Betting; Figure 3 notes, p. 1954).

Step 3: Cross-task regression. The main result (Figure 5, p. 1957) is the Pearson correlation between β^k\hat{\beta}_k (step 1) and the average institutional improvement in task kk (step 2) across the 15 tasks. The between-subjects correlation pools improvements across Betting, Auction, and Committee. The within-subjects correlation uses the same subjects for both confidence and institutional decisions.

Average overconfidence check. As an ancillary test, the paper replaces β^k\hat{\beta}_k with task-level average overconfidence d^k=cˉkpˉk\hat{d}_k = \bar{c}_k - \bar{p}_k and repeats step 3. This directly tests Prediction 2: the resulting correlation is weakly negative (r = -0.34 between, r = -0.32 within) but not statistically distinguishable from zero (p. 1961).

Expert survey. A separate sample of 38 behavioral economists (CESifo/VIBES panel, November 2021) predicted institutional improvements and confidence differences for 7 of the 15 tasks in the Auction treatment. Experts’ median forecasts are compared to actual outcomes in Figure 6 (p. 1962). The analysis uses a paired comparison of forecast vs. actual for each of the 7 tasks; no regression is reported. Camerer and Lovallo (1999) document overconfidence in entry decisions; the expert results here parallel that finding in the prediction domain. Moore and Healy (2008) provide the taxonomy of overconfidence types that the paper uses to frame what experts miss. Kendall and Oprea (2018) study the market selection hypothesis in a related laboratory design.

DatasetRole in paperWiki page
Primary experimental data (Prolific online experiment, June 2021)15 cognitive tasks x 3 institutions x 2,153 subjects; ~70,000 individual decisionsNo page yet
Expert survey (Social Science Prediction Platform, November 2021)38 behavioral economists predicting institutional filtering and confidence differences for 7 tasksNo page yet

Sample: 1,381 subjects in between-subjects treatments (Betting, Auction, Committee, Confidence); 314 subjects in within-subjects treatments (Betting Within, Auction Within, Committee Within); June 2021 on Prolific. Replication data available at ICPSR (https://doi.org/10.3886/E185741V1).

Use the original if you are: studying which specific cognitive biases survive market aggregation (the paper gives results for all 15 tasks by institution); designing experiments to study self-selection through institutions (the exact institution implementations are in online appendices); investigating what determines the confidence-performance correlation itself and why it varies across tasks (Section IV.C); or replicating the expert-survey methodology (the SSPP survey instruments are reproduced in online appendices). Locators above point to the exact figures and pages.

Source: peer-reviewed, American Economic Review 113(7), July 2023. Published under AEA copyright with 12-month delayed open access. No CC license detected. This distillation was extracted by an LLM on 2026-06-25 and is not human-verified or independently reproduced. Redistribution is extract-only.

Enke, Benjamin, Thomas Graeber, and Ryan Oprea. “Confidence, Self-Selection, and Bias in the Aggregate.” American Economic Review 113, no. 7 (July 2023): 1933-1966. DOI: 10.1257/aer.20220915. Replication data: https://doi.org/10.3886/E185741V1 (AEA/ICPSR). This page is an extraction by the Institute for Automated Research: core results and equations summarized; not reproduced.

Found an error or want a topic covered? Open an issue, use the Edit page link above, or email contact@instituteforautomatedresearch.org. Edits are reviewed before publishing; provenance and accuracy are the point.