Skip to content

Second-Best Fairness: Cappelen, Cappelen & Tungodden (2023)

Distilled by claude-sonnet-4-6 · extracted Jun 24, 2026, verified Jun 25, 2026

JEL (IAR-assigned): D63, D72, D78, H23, I38 · assigned from the abstract, not the journal

Full structured metadata (methods, scope, relatesTo, topics, datasets): raw Markdown (.md)

paper-summarybehavioral-economicsfairnessredistributionsocial-insurancepolitical-economypanel-regressionpeer-reviewedunreplicated

What this is. The core results, theoretical model, and empirical strategy from this paper: enough to understand what was found and how, without reading all 28 pages. To replicate or extend, read the original at the DOI.

The paper examines how people trade off false positives (paying an undeserving individual) against false negatives (not paying a deserving individual) in second-best fairness decisions. Across three large-scale experiments in the United States and Norway (26,500 spectators total), the large majority of spectators are false negative averse: they prefer risking a false positive over a false negative. When the probability of a false claim is 50 percent, 72.4 percent of spectators still choose to pay, consistent with placing higher weight on avoiding a false negative. However, about 20 percent are strongly false positive averse. Americans are more false positive averse and less false negative averse than Norwegians, and right-wing spectators exhibit the same pattern within both countries. These second-best fairness preferences strongly predict policy attitudes on unemployment benefits and income redistribution, above and beyond stated fairness views and altruism.

Magnitudes are as reported; \*\*\* = 1%. Locators point into the source PDF.

#ResultLocatorMagnitude
R1Paying falls monotonically with false-claim probability; 72.4% still pay at Pr(f)=0.5, implying a FN-FP gap of 44.8 ppTable 3 col 1, p. 2473; Figure 1, p. 2472Coeff at 50% false-claim prob: -17.6 pp (SE=0.016, p<0.001); baseline (Pr(f)=0) = 90.0%; at 100%: -79.7 pp; FN-FP difference = 44.8 pp (p<0.001)
R2Type-share lower bounds: FP averse 20.3%, FN averse 65.2% (pooled compensation experiment)Table 4 upper panel, p. 2474FP averse LB 20.3% (SE=0.015); symmetric UB 14.5% (SE=0.037); FN averse LB 65.2% (SE=0.027)
R3Majority have highly asymmetric preferences: 20.3% strongly FP averse, 43.5% strongly FN averseFigure 2 upper-left panel, p. 2475; text p. 2474Strongly FP averse (beta <= 0.25): 20.3%; strongly FN averse (beta >= 0.75): 43.5% of pooled sample
R4US spectators are more FP averse and less FN averse than Norwegians across all experimentsTable 6 right panel, p. 2481Strongly FP averse US vs Norway: +11.8 pp (SE=0.017, p<0.001); strongly FN averse: -10.0 pp (SE=0.021, p<0.001)
R5Right-wing spectators are less FN averse and more FP averse than non-right-wing spectatorsTable 6 left panel, p. 2481FN averse -9.5 pp (SE=0.010, p<0.001); strongly FP averse +6.2 pp (SE=0.019, p<0.001); strongly FN averse -12.1 pp (SE=0.022, p<0.001)
R6Second-best fairness preferences predict policy attitudes independently of stated fairness viewsTable 7, p. 2482Paying predicts support for generous unemployment benefits: coeff 0.562 (SE=0.025, p<0.001); income inequality: 0.354 (SE=0.025, p<0.001); survives controls for fairness views, efficiency costs, altruism, religiosity

Overall (paper’s conclusion). The majority of spectators in both countries are false negative averse in all three experiments. A significant minority is strongly false positive averse. Country and political differences are of similar magnitude: the US-Norway gap mirrors the right-wing/non-right-wing gap within each country. These second-best fairness preferences are strongly predictive of real-world policy attitudes on redistribution and social insurance, suggesting they are a fundamental ingredient in the political economy of welfare institutions.

The paper proposes a simple expected utility framework to characterize second-best fairness preferences (Section I, pp. 2461-2462). Consider an environment where a spectator must choose a payment yy for an individual whose claim may or may not be false. Let m(f)m(f) be the fair payment if the claim is false and m(c)m(c) the fair payment if it is correct, with m(c)>m(f)m(c) > m(f). The probability the claim is false is Pr(f)\Pr(f) and the probability it is correct is 1Pr(f)1 - \Pr(f).

In line with Cappelen et al. (2013a), the spectator dislikes any payment that deviates from what is fair. The expected utility of paying yy is (equation 1, p. 2461):

EU(y)=Pr(f)u ⁣(ym(f))+[1Pr(f)]u ⁣(ym(c)),(1)EU(y) = \Pr(f)\, u\!\left(y - m(f)\right) + \bigl[1 - \Pr(f)\bigr]\, u\!\left(y - m(c)\right), \tag{1}

where u()u(\cdot) is weakly decreasing in the deviation from the fair payment and u(0)=0u(0) = 0. The spectator faces a binary choice between paying, y=m(c)y = m(c), and not paying, y=m(f)y = m(f). These yield (equations 2 and 3, p. 2461):

EU ⁣(y=m(c))=Pr(f)u ⁣(m(c)m(f)),(2)EU\!\left(y = m(c)\right) = \Pr(f)\, u\!\left(m(c) - m(f)\right), \tag{2} EU ⁣(y=m(f))=[1Pr(f)]u ⁣(m(f)m(c)).(3)EU\!\left(y = m(f)\right) = \bigl[1 - \Pr(f)\bigr]\, u\!\left(m(f) - m(c)\right). \tag{3}

Since m(c)m(f)=m(f)m(c)|m(c) - m(f)| = |m(f) - m(c)|, define β\beta as the relative weight the spectator places on avoiding a false negative versus a false positive (p. 2462):

β=u ⁣(m(f)m(c))u ⁣(m(c)m(f))+u ⁣(m(f)m(c)).\beta = \frac{u\!\left(m(f) - m(c)\right)}{u\!\left(m(c) - m(f)\right) + u\!\left(m(f) - m(c)\right)}.

Three types follow: False Positive Averse (β<1/2\beta < 1/2), Symmetric (β=1/2\beta = 1/2), False Negative Averse (β>1/2\beta > 1/2).

Observation 1 (p. 2462): The spectator is indifferent between paying and not paying when Pr(f)=β\Pr(f) = \beta; strictly prefers not to pay when Pr(f)>β\Pr(f) > \beta; and strictly prefers to pay when Pr(f)<β\Pr(f) < \beta. The individual switching threshold directly reveals β\beta. This also predicts that the choice between the two options should be independent of the size of m(c)m(f)m(c) - m(f) (the payoff size), which Table 5 confirms (the high-stakes treatment changes behavior by -4.3 pp, not significant after multiple-testing correction).

Upper and lower bounds on type shares are derived from sp(Pr(f))sp(\Pr(f)), the share of spectators paying at each treatment value (Observation 4, p. 2469):

SU=2×max ⁣{0,min ⁣(sp(0.25)sp(0.5),  sp(0.5)sp(0.75))},S_U = 2 \times \max\!\left\{0,\, \min\!\bigl(sp(0.25) - sp(0.5),\; sp(0.5) - sp(0.75)\bigr)\right\}, FPU=1sp(0.5),FNU=sp(0.5),FPL=FPU0.5SU,FNL=FNU0.5SU.FP_U = 1 - sp(0.5), \quad FN_U = sp(0.5), \quad FP_L = FP_U - 0.5\, S_U, \quad FN_L = FN_U - 0.5\, S_U.

The estimation strategy is between-subject OLS on the binary payment indicator, applied to the pooled sample and separately for each country (Section III, pp. 2468-2471). The approach builds on randomized-survey-experiment (spectators randomly assigned to treatment arms) and panel-regression (linear probability model with controls and population weights).

The main specification for treatment effects (equation 4, p. 2468):

ei=α+α1P(0.25)i+α2P(0.5)i+α3P(0.75)i+α4P(1)i+γXi+εi,(4)e_i = \alpha + \alpha_1 P(0.25)_i + \alpha_2 P(0.5)_i + \alpha_3 P(0.75)_i + \alpha_4 P(1)_i + \gamma \mathbf{X}_i + \varepsilon_i, \tag{4}

where eie_i is an indicator equal to one if spectator ii pays, P(0.25)iP(0.25)_i through P(1)iP(1)_i are treatment indicators for each false-claim probability level (baseline: Pr(f)=0\Pr(f) = 0), and Xi\mathbf{X}_i is a vector of controls (income, education, gender, age, political ideology). Estimates are population-weighted. Multiple testing corrected via Holm-Bonferroni and Romano-Wolf procedures.

For additional robustness treatments at Pr(f)=0.5\Pr(f) = 0.5 (equation 5, p. 2470):

ei=α+α1Mi+γXi+εi,(5)e_i = \alpha + \alpha_1 M_i + \gamma \mathbf{X}_i + \varepsilon_i, \tag{5}

where MiM_i indicates the specific additional treatment (doubled stakes, nationality framing, or endowment introduction). For the cost treatments (equation 6, p. 2470):

ei=α+α1C(0.1)i+α1C(0.3)i+γXi+εi,(6)e_i = \alpha + \alpha_1 C(0.1)_i + \alpha_1 C(0.3)_i + \gamma \mathbf{X}_i + \varepsilon_i, \tag{6}

where C(0.1)iC(0.1)_i and C(0.3)iC(0.3)_i indicate treatments where the spectator bears a personal cost of US$0.1 or US$0.3 for paying.

For policy attitudes (equation 7, p. 2471):

poli=α+α1payi+γXi+εi,(7)\text{pol}_i = \alpha + \alpha_1\, \text{pay}_i + \gamma \mathbf{X}_i + \varepsilon_i, \tag{7}

where poli\text{pol}_i is stated support for unemployment benefit generosity or income equalization on a seven-point scale, and payi\text{pay}_i is an indicator for having paid in the experiment.

Three experiments share the same between-subject design. In each, spectators are randomly assigned to one of five treatments where Pr(f){0,0.25,0.5,0.75,1}\Pr(f) \in \{0, 0.25, 0.5, 0.75, 1\}, then decide whether to pay a worker whose claim may be false.

Compensation experiment (Section II.A, pp. 2463-2465). Workers are recruited on an international online labor market platform. It is randomly determined whether they are offered work. Those not offered work are entitled to a compensation of US$4 (m(c)=4m(c) = 4, m(f)=0m(f) = 0). Spectators are told the false-claim probability for the matched worker and decide whether to pay. The main sample is 5,395 spectators (2,695 US, 2,700 Norway). Additional treatments at Pr(f)=0.5\Pr(f) = 0.5 test: (i) high stakes (US$8 compensation, Panel A Table 5); (ii) nationality framing, where stakes are reported in local currency and workers are implied to be compatriots (Panel A); (iii) three endowment-plus-cost arms (Panel B). All produce null effects, consistent with the theoretical prediction that stake size and in-group salience should not affect the FP-FN trade-off (RESULT 3).

Earnings experiment (Section II.B, pp. 2465-2466). Same structure but the claim is for earnings from completing a 15-minute task rather than compensation for not being offered work. Spectators in 5,391 observations (main study, excluding pilot). Treatment effects are tested for differences from the Compensation experiment via interaction terms (Figure 3, Panel A). All interaction effects are small and not robust to multiple-testing correction (RESULT 4), confirming the preferences are not specific to the compensation context.

Unemployment experiment (Section II.C, p. 2466). A nonincentivized survey experiment where respondents decide whether to hypothetically pay unemployment benefits to someone with a known false-claim probability. The pattern of results closely matches the incentivized experiments (RESULT 5). Respondents are somewhat more false positive averse in this policy domain, consistent with the policy debate around welfare fraud (Figure 3, Panel B).

Country-level regressions run equation (4) separately for the US and Norway. Pooled cross-country and cross-political-spectrum comparisons use 22,476 observations (all three experiments combined) with interaction terms for Norwegian nationality and for right-wing political affiliation (Table 6). Almås, Cappelen, and Tungodden (2020) provide the comparison framework for cross-country distributive preference differences. Policy attitude regressions (Table 7) pool all experiments and add controls for fairness views and efficiency beliefs; Alesina and Angeletos (2005) is the background reference for the fairness-redistribution link.

DatasetRole in paperWiki page
Norstat population panel (US and Norway)Recruitment of 22,500 spectators, quota-matched on age, gender, and geography to be nationally representativeNo page yet
Amazon Mechanical TurkRecruitment of 6,250 workers (consequential decisions) and 4,000 pilot spectators from earlier Earnings versionsNo page yet

All primary data are author-generated experimental data. Replication data and code are publicly available at the AEA/ICPSR OpenICPSR archive (https://www.openicpsr.org/openicpsr/188201). Sample: summer 2022 (main study); pilot collected 2019. Workers: 2,250 AMT main study + 4,000 pilot.

Read the original if you are: designing second-best social insurance policies where eligibility is uncertain; studying cross-country or political heterogeneity in redistributive preferences; classifying spectator types (FP vs FN averse) in a behavioral experiment; or extending the framework to judicial or disability contexts. The locators above point to the exact tables and figures for each result.

Source: peer-reviewed, American Economic Review 113(9), September 2023. AEA copyright; not yet freely available on AEAweb (3-year embargo expires September 2026). This distillation was extracted by an LLM on 2026-06-24 and is not human-verified or independently reproduced. Redistribution is extract-only; the PDF is not hosted here.

Cappelen, Alexander W., Cornelius Cappelen, and Bertil Tungodden. “Second-Best Fairness: The Trade-Off between False Positives and False Negatives.” American Economic Review 113, no. 9 (September 2023): 2458-2485. DOI: 10.1257/aer.20211015.

Found an error or want a topic covered? Open an issue, use the Edit page link above, or email contact@instituteforautomatedresearch.org. Edits are reviewed before publishing; provenance and accuracy are the point.