Diversifying Society's Leaders: Chetty, Deming & Friedman (2026)
Distilled by claude-sonnet-4-6 · extracted Jun 28, 2026, verified Jun 28, 2026
JEL (IAR-assigned): I23, J24, J62 · assigned from the abstract, not the journal
What this is. A distilled skeleton of Chetty, Deming, and Friedman (2026). Read the original article to replicate or extend. Equations, tables, and figures referenced below are from that source. This page is LLM-extracted and has not been human-verified.
The paper uses a newly linked panel dataset combining federal income tax records, college attendance records, SAT/ACT scores, and internal applications data from Ivy-Plus and flagship public colleges to study two questions: (i) why children from top-income families disproportionately attend Ivy-Plus colleges (Harvard, Yale, Princeton, and the other eight Ivy League colleges, Chicago, Duke, MIT, and Stanford), and (ii) whether attending those colleges causally improves students’ postcollege outcomes. The analysis proceeds in four parts: characterizing the pipeline from application through matriculation, identifying the mechanisms driving the high-income admissions advantage, estimating causal effects using two quasi-experimental designs, and predicting the effects of counterfactual admissions policies on socioeconomic diversity.
The headline findings are that (i) Ivy-Plus attendance does causally improve upper-tail outcomes by substantial magnitudes relative to attending a flagship public college, and (ii) the credentials that give high-income applicants their admissions advantage (legacy status, nonacademic ratings, athletic recruitment) do not predict better postcollege outcomes once college quality is held fixed, while academic credentials do. Contrary to the well-known findings of Dale and Krueger (2002) on mean log earnings, large causal effects of Ivy-Plus attendance on upper-tail income and nonmonetary outcomes emerge once richer data allow college value-added to be measured directly rather than through test-score proxies. The paper reconciles with Dale and Krueger (2002) and Dale and Krueger (2014): both papers agree on effects on mean log earnings; the divergence arises entirely on upper-tail outcomes where Ivy-Plus colleges have disproportionate effects.
Core results
Section titled “Core results”| # | Result | Locator | Magnitude as reported |
|---|---|---|---|
| R1 | Top 0.1% income families are 2.5x more likely to gain admission to Ivy-Plus than middle-class applicants (70th-80th pctile) with the same test scores; 99th-99.9th pctile are 44% more likely; flagship public admissions rates are uncorrelated with parental income conditional on test scores | Figure III Panel B, p. 80 | Relative admission rate: 2.5x for top 0.1%; 1.44x for 99th-99.9th pctile; roughly 1.0x at flagship publics |
| R2 | 68% of the income gap in Ivy-Plus attendance conditional on test scores arises from admissions rather than applications or matriculation; decomposed into legacy preferences (31%), nonacademic credentials (21%), and athletic recruitment (16%), accounting for 114 of 168 extra top-1% students | Table II, pp. 77, 82; Figures V, VI | 52 extra top-1% students from legacy preferences; 35 from nonacademic credentials; 27 from athletic recruitment; 141 of 168 from admissions-related factors combined including athletes |
| R3 | Attending an Ivy-Plus college instead of the average flagship public college causally increases the predicted probability of reaching the top 1% of income at age 33 by 5 pp (+42%) | Table IV col. 1 and 6, p. 114 | TOT = 5.01 pp (SE 1.31), ; from 11.8% to 16.8% |
| R4 | Ivy-Plus attendance nearly doubles the probability of attending an elite graduate school | Table IV Panel B, p. 114 | TOT = 5.64 pp (SE 2.79); from 6.1% to 11.7%, +92% |
| R5 | Ivy-Plus attendance more than triples the probability of working at an elite firm at age 25 | Table IV Panel B, p. 114 | TOT = 16.96 pp (SE 4.01); from 8.5% to 25.5%, +199% |
| R6 | Ivy-Plus attendance nearly quadruples the probability of working at a prestigious firm at age 25 | Table IV Panel B, p. 114 | TOT = 17.51 pp (SE 4.26); from 7.2% to 24.7%, +245% |
| R7 | The credentials underlying the high-income admissions advantage (legacy status, nonacademic ratings, athletic recruitment) have zero or negative association with postcollege success after adjusting for college quality; high academic ratings have a +4.8 pp effect on top-1% probability | Figure XV Panel B, p. 129 | Legacy: negative (negatively associated, p. 129); nonacademic rating: approx. 0 (no significant association); athlete: approx. 0 (no significant association); high academic rating: +4.8 pp on top-1% probability |
| R8 | Eliminating all three high-income admissions advantages (legacy preferences, nonacademic-credentials boost, athletic recruitment income gradient) would increase the share of Ivy-Plus students from the bottom 95% of parental income by 8.8 pp, with no reduction in average student outcomes | Table V Panel A rows 1-4, p. 133 | Top-1% parental income share falls from 15.8% to 9.9%; bottom-60% share rises from 15.7% to 20.0%; average predicted outcomes unchanged or improved |
Overall (paper’s conclusion). Ivy-Plus colleges have large causal effects on students’ chances of achieving upper-tail earnings and nonmonetary leadership outcomes, but they also substantially over-admit students from high-income families relative to what academic credentials alone would predict. The three factors driving this admissions advantage (legacies, nonacademic credentials, athletes) are uncorrelated with, or negatively predictive of, postcollege success, meaning that admissions policy changes that eliminate these advantages would increase socioeconomic diversity by an amount comparable to race-based affirmative action without reducing student-body quality. Because Ivy-Plus colleges account for a relatively small share of all Americans, changes in admissions policy have small effects on the share of top-1% earners from low-income families but could meaningfully diversify the socioeconomic backgrounds of people in nonmonetary leadership positions (senators, Supreme Court justices, Nobel laureates).
Theory / model
Section titled “Theory / model”The paper presents a formal statistical model (Section IV.A, pp. 92-97) to clarify what each research design identifies.
Admissions ratings. College assigns applicant a composite rating
Z_{ij} = \gamma_{1j} X_{1i} + \gamma_{2j} X_{2i} + \eta_i + \epsilon_{ij}, \tag{3}
where is observable (e.g., SAT/ACT score), is unobservable but correlated with long-term outcomes (e.g., intrinsic ability or motivation), is a common idiosyncratic component uncorrelated with (e.g., a strong guidance counselor letter that helps at all colleges), and is pure noise at college uncorrelated with across all colleges (e.g., whether the student plays an instrument needed for the college’s orchestra in the application year). Colleges admit student to college if , where is a college-specific cutoff. Colleges are assumed to decide independently.
Postcollege outcomes. The student’s outcome (e.g., earnings or one of the leadership proxies in Figure I) follows
Y_i = \sum_{j \in J_i} D_{ij} \phi_j + \beta_1 X_{1i} + \beta_2 X_{2i} + \epsilon_i^Y, \tag{4}
where is an enrollment indicator, is college ‘s causal value added (normalized to zero for the average flagship public, the outside option ), and is an outcome error orthogonal to and . The goal is to estimate , the causal effect of attending an Ivy-Plus college instead of college .
Identification. OLS on admitted students is biased because affects both admission and outcomes. The paper offers two designs to remove this bias, both exploiting data on admissions decisions at multiple colleges.
Research Design 1 (idiosyncratic-admissions IV, Section IV.A.2, p. 93). Among students on the waitlist at college , the paper uses being admitted off the waitlist as a quasi-instrument. The rescaled waitlist estimator is
r_A = \frac{E[Y_i | P_{iA}=1, X_{1i}, \tilde{X}_{2i}] - E[Y_i | P_{iA}=0, X_{1i}, \tilde{X}_{2i}]}{E[D_{iA} | P_{iA}=1, X_{1i}, \tilde{X}_{2i}]}, \tag{5}
where is a proxy for (e.g., whether the student was placed on the waitlist, itself a signal of near-marginal quality). Under the correlated-admissions-criteria assumption (Assumption 1, p. 95) that for colleges with similar holistic admissions processes, the estimator identifies if and only if the test statistic , where measures whether being admitted vs. rejected from college ‘s waitlist predicts admission at college . This multiple-rater test (Figure VII, pp. 102-104) passes empirically: waitlisted students’ admission outcomes at other Ivy-Plus colleges are statistically indistinguishable from each other, regardless of whether they are admitted from the waitlist at the reference college.
Research Design 2 (matriculation design, Section IV.A.3, p. 96). Following Mountjoy and Hickman (2021) and Dale and Krueger (2002), the paper compares outcomes for students admitted to the same portfolio of colleges who choose to attend different colleges:
r_M = E[Y_i | D_{iA}=1, X_{1i}, J_i = \{A, O\}] - E[Y_i | D_{iO}=1, X_{1i}, J_i = \{A, O\}], \tag{6}
under Assumption 2 that conditional on the admissions portfolio and , the unobservable is orthogonal to the matriculation choice (p. 96). Both designs yield consistent estimates (Table IV, p. 114), which strengthens the credibility of both.
Method
Section titled “Method”Surrogate index for early-career outcomes (Section II.C.4, p. 67; Section IV.B, pp. 97-100). Because income ranks at age 33 are not observed for recent cohorts, the paper constructs a surrogate index (Athey et al., forthcoming) using employers and graduate schools at ages 22-25 to predict the probability of reaching the top 1% of income at age 33. This is motivated by the finding that firms’ employment composition at ages 22-25 strongly predicts income at 33 (Figure IX and Online Appendix Figure A.23a, pp. 107-108). The paper verifies that early-career employers and graduate schools capture the income dynamics that produce age-33 outcomes, with a near-zero treatment effect at age 25 that grows steadily to ~5 pp by age 33 (Figure IX Panel A, p. 107).
Treatment-effect heterogeneity by outside options (Section IV.C.5, pp. 110-113). To identify (the causal effect relative to the average flagship public as outside option), the paper exploits variation in the value added of each applicant’s outside option. Applicants are grouped by home state, parental income, and race; the outside-option quality is measured as the average observational value added of colleges attended by rejected non-waitlisted applicants in each group. The paper then estimates:
(Table IV col. 1 and Figure X Panel A, pp. 112-113). The slope of the heterogeneity relationship equals -0.79 (Figure X Panel A), indicating that most variation in outcomes between colleges is driven by genuine causal effects rather than selection, with students facing weaker outside options gaining most from Ivy-Plus attendance.
Multiple-rater admissions test (Section IV.C.1, pp. 100-104; Figure VII, p. 102). The paper develops a new validation test for the quasi-random variation assumption in Research Design 1. The test compares admission rates at lower-ranked Ivy-Plus colleges (ranked by revealed student preferences) for students who are admitted vs. rejected from the waitlist at college . Under the correlated-admissions-criteria assumption, if and only if the residual variation in ‘s admissions decisions among waitlisted students is orthogonal to . The paper implements this test for three specifications (no controls, with controls, dropping legacies/athletes/top-1%) and finds statistically indistinguishable from zero across all three (Figure VII, p. 102), supporting the identification assumption.
Empirical specifications
Section titled “Empirical specifications”Pipeline analysis: counterfactual attendance rate under income-neutral admissions (Section III.A.1, p. 75, Equation 1). To quantify how many extra top-1% students are in the Ivy-Plus class conditional on test scores, the paper computes:
\text{Counterfactual Attendance Rate}_c = \sum_a N_{Top 1\%, a} \times \text{Attendance Rate}_{P70-80, ac}, \tag{1}
where is the number of test takers with score from families in the top 1% and is the fraction attending college among students with score from the 70th-80th percentile. Scaling to a class of 1,650 students, this counterfactual implies 93 students from the top 1% rather than the observed 261, a gap of 168 “extra” top-1% students (10.2% of enrollment).
Pipeline decomposition (Section III.B.4, p. 81, Equation 2). For non-athletes, the paper decomposes the 168-student gap by sequentially equalizing applications, admissions, and matriculation rates across income groups:
\text{Equal Admit CF}_c = \sum_a N_{Top 1\%, a} \times \text{Application Rate}_{Top 1\%, ac} \times \text{Admission Rate}_{P70-80, ac} \times \text{Matriculation Rate}_{Top 1\%, ac}. \tag{2}
Setting application, admission, and matriculation rates to those of the middle class, and averaging across orderings, admissions account for 58% (96 students) of the 168-student gap (Online Appendix Table A.6, p. 82). Including athletes, 114 of 168 extra students (68%) are from admissions-related factors.
Causal-effect regression (Section IV.C.3, pp. 105-106). The paper estimates the treatment-on-the-treated (TOT) effect from Research Design 1 by regressing an outcome indicator on a waitlist-admission indicator plus college-by-cohort fixed effects, clustering standard errors by student (to account for students on multiple waitlists), and dividing the reduced-form coefficient by the first-stage effect (the probability of attending college conditional on admission). In the primary specification (Table IV col. 1, p. 114), the outcome is the predicted top-1% income probability based on age-22-25 employers. Controls include a quintic in SAT/ACT scores, parent income bin dummies, race/ethnicity indicators, gender, home-state indicators, recruited-athlete and legacy status indicators, and college-by-cohort fixed effects. Robustness to dropping legacies, athletes, and top-1% applicants (the third bar set in Figures VIII and XI) confirms the estimates are not driven by the same characteristics that generate the admissions imbalance.
Datasets used
Section titled “Datasets used”| Dataset | Role in paper | Wiki page |
|---|---|---|
| Federal income tax records (IRS, 1996-2021) | Parental income (1040, W-2), children’s individual income (W-2, Form SE, 1040), employer identification (W-2) | no page yet |
| 1098-T college attendance forms (Dept of Education via NSLDS) | College attendance indicator for all US colleges; linked to tax records at individual level | no page yet |
| SAT scores (College Board, 2001-2005 and odd years 2007-2015) | Standardized academic qualification measure; composite score (math + critical reading) | no page yet |
| ACT scores (ACT, 2001-2015) | Standardized test scores converted to SAT equivalents via ACT 2016 concordance tables | no page yet |
| Pell Grant records (NSLDS, 1999-2013) | Low-income student identification; supplementary college attendance signal | no page yet |
| Application and admissions records (Ivy-Plus and flagship public colleges, 1998-2015) | Admission indicators, legacy/athlete/faculty-child flags, admissions-office ratings (academic and nonacademic), application round, GPA; from several Ivy-Plus colleges plus 9 flagship public systems | no page yet |
Sample note. The pipeline analysis sample covers 5,063,263 students on pace to graduate high school in 2011, 2013, or 2015 (Table I col. 1, p. 70). The college-specific analysis sample covers 486,150 Ivy-Plus applicants and 1,877,770 flagship public applicants for whom internal admissions records are available (Table I cols. 3-4, p. 72). All data were linked at the individual level using Social Security numbers and stripped of personally identifiable information before analysis; the IRS component was accessed under IRS contract TIRNO-16-E-00013. The dataset construction and income variable definitions build on the earlier linked administrative dataset of Chetty et al. (2020) on income segregation across US colleges; the current paper extends that work with internal admissions records and later cohorts. Application-rate differences as a driver of access gaps were documented by Hoxby and Avery (2013) using geographic imputations of family income; this paper finds admissions rates are the primary driver at Ivy-Plus colleges in the more recent period studied, after private colleges expanded low-income recruitment programs.
When to read the full paper
Section titled “When to read the full paper”Read Chetty, Deming, and Friedman (2026) if you need:
- The exact statistical model and identification assumptions (Equations 3-6, Online Appendix H with formal proofs), especially the multiple-rater test logic and how the two designs nest into a unified framework (Section IV.A, pp. 92-97).
- Heterogeneity in treatment effects by parental income, race, and outside-option quality (Figure XII Panel D, p. 120; Table IV, p. 114; Online Appendix Table A.10).
- The admissions-ratings analysis showing how legacy preferences and nonacademic credentials mechanically generate the high-income admissions advantage and how it is measured via counterfactual admissions predictions (Figures V and VI, pp. 85, 89; Section III.C, pp. 84-91).
- The counterfactual admissions simulations and their predicted effects on leadership outcomes across specific categories (senators, Supreme Court justices, Nobel laureates) under alternative admissions policies (Table V, pp. 133-134; Section VI, pp. 131-141).
- The quantile treatment effects analysis showing why Ivy-Plus effects are concentrated at the very top of the income distribution rather than distributed proportionally (Figure XIII, p. 124; Section IV.E, pp. 122-125).
- The college-level data on parental income distributions at each stage of the application process, publicly released at www.opportunityinsights.org/data (Online Appendix O, p. 142).
Attribution and rights
Section titled “Attribution and rights”Chetty, Raj, David J. Deming, and John N. Friedman. “Diversifying Society’s Leaders? The Determinants and Causal Effects of Admission to Highly Selective Private Colleges.” The Quarterly Journal of Economics 141(1), 2026, 51-145. DOI: 10.1093/qje/qjaf050.
Copyright (c) The Author(s) 2025. Published by Oxford University Press on behalf of President and Fellows of Harvard College. All rights reserved. Replication code and data available at Harvard Dataverse: https://doi.org/10.7910/DVN/YMVK4K.
This page is a machine-generated distillation (LLM-extracted). It has not been human-verified and reproduces no paywalled content, only brief quotations of results as permitted for commentary and research purposes (extract-only).