Skip to content

How Much Does Racial Bias Affect Mortgage Lending: Bhutta, Hizmo & Ringo (2025)

Distilled by claude-sonnet-4-6 · extracted Jun 6, 2026, verified Jun 6, 2026

JEL (IAR-assigned): G21, J15, G28 · assigned from the abstract, not the journal

Full structured metadata (methods, scope, relatesTo, topics, datasets): raw Markdown (.md)

paper-summaryhousehold-financemortgage-lendingdiscriminationfair-lendingpanel-regressionpeer-reviewedunreplicateddata:hmdadata:nsmo

What this is. This is a machine-distilled skeleton of the paper. Read the original (DOI 10.1111/jofi.13444) to replicate or extend.

Using confidential expanded HMDA data for 2018-2019 (nearly 9 million applications), the paper finds that observable applicant risk factors (credit score, LTV, DTI, AUS recommendation) explain most of the racial and ethnic gaps in mortgage denial rates. The residual “excess denial” gap is 2 percentage points for Black applicants and roughly 1 pp for Hispanic and Asian applicants, substantially smaller than the 8 pp gaps found by Munnell et al. (1996) or the 7-10 pp estimated by Bartlett et al. (2022) without controlling for credit score and other underwriting factors. Giacoletti, Heimer and Yu (2025) estimate a similar 7 pp raw gap and argue at least half reflects discrimination; this paper’s evidence points more to unobserved risk. Cross-sectional evidence on lender strictness shows that stricter lenders have larger excess minority denial rates, consistent with tighter overlays on unobserved risk factors rather than discriminatory intent. Indirect tests (fintech lenders, market competition, regional racial animus) do not yield clear evidence of discrimination. The paper also revisits findings from Bhutta and Hizmo (2020) on minority mortgage pricing, and extends them by examining denial disparities using the newly expanded HMDA data. A separate analysis using NSMO survey data finds that minority borrowers report substantially worse service quality, suggesting a dimension of disparate treatment in service delivery that is not captured by denial statistics.

#ResultLocatorMagnitude as reported
R1Raw Black-White denial gap before any controlsTable I, p.1474Black denial rate 18%, White 8%; raw gap 10 pp
R2Excess denial after full FICO-LTV-DTI-AUS-lender controlsTable II col.(3), p.1475Black 2.0 pp**, Hispanic 0.9 pp**, Asian 1.4 pp** (s.e. 0.001)
R3Lender strictness correlation with excess minority denialsFigure 2 right panels, p.1480r = 0.63 (Black), 0.50 (Hispanic), 0.65 (Asian) vs. lender strictness for Whites
R4AUS excess denial for Black applicants (unobserved risk signal)Table II col.(5), p.14751.5 pp** (s.e. 0.001), suggesting racial gaps in AUS-observed-but-HMDA-unobserved risk
R5Racial animus correlation test for discriminationTable IV col.(3) vs col.(6), p.1488Racially charged search rate interaction: 0.002** for lender and AUS excess denials alike, suggesting unobserved risk rather than discrimination
R6Minority borrower service quality: processing and closing delaysTable V cols.(1)-(2), p.1493Black: 4.5 pp more likely to report processing delays***, 9.6 pp more likely to have postponed closing***
R7Minority borrower satisfaction with lenderTable V col.(6), p.1493Black 7.1 pp less likely to be very satisfied with lender***; Asian 11.3 pp less***

Overall (paper’s conclusion). The paper concludes that disparate treatment plays a much smaller role in generating mortgage denial disparities than the 1990s benchmark of Munnell et al. (1996) suggests, implying significant progress in fair lending over 30 years. The 1-2 pp excess denials overstate actual discrimination because unobserved risk factors that vary by race and ethnicity (and that stricter lenders screen for) explain at least part of the residual gap. Service quality disparities, however, are a documented and underexplored dimension of differential treatment.

The paper has no formal economic model. It defines disparate treatment structurally as a difference in expected credit decisions across race/ethnicity for otherwise identical applicants.

Let the binary AUS recommendation be:

D_{AUS} = g(X, u), \tag{1}

where g()g(\cdot) is a deterministic function of risk characteristics XX observable in HMDA and other risk characteristics uu unobserved in HMDA (see Section I, p.1472 for the full DU factor list).

Lender ii‘s binary denial decision is:

D^{i}_{Lender} = h_i(X^*, u, w, r) + e, \tag{2}

where hi()h_i(\cdot) may differ from g()g(\cdot); XX^* is a potentially updated value of XX after verification; ww is lender-specific overlays beyond AUS; rr is race/ethnicity; and ee is idiosyncratic human error. Lender ii engages in disparate treatment against Black relative to White applicants if (p.1473):

\int h_i(X^*, u, w, \text{Black})\, dF_B(X^*, u, w) > \int h_i(X^*, u, w, \text{White})\, dF_B(X^*, u, w), \tag{3}

where FB()F_B(\cdot) is the joint CDF of underwriting factors for the Black applicant population. The identification challenge is separating rr (illegal discrimination) from uu and ww (unobserved but potentially race-correlated risk factors).

The key methodological innovation is the lender strictness measure, constructed as the lender fixed effect from a denial regression run exclusively on White applicants (equation (2) controls, p.1478). This measure isolates lender-specific overlay policies from any differential treatment of minorities by construction. The paper then correlates lender strictness with lender-specific excess minority denial rates (estimated with lender-varying race/ethnicity coefficients) as an indirect test of whether unobserved risk drives excess denials. The approach is analogous to the judge-specific propensity-to-release design of Arnold, Dobbie and Yang (2018) and Arnold, Dobbie and Hull (2022).

Fintech identification follows Fuster et al. (2019): a lender is coded as fintech if it appears on that paper’s fintech list. Market concentration is proxied by the top-4 lenders’ county market share. Racial animus is measured by the racially charged search rate from Stephens-Davidowitz (2014), standardized to mean zero and unit variance. These three cross-sectional dimensions are interacted with race/ethnicity in equation (2) to test whether excess denials are systematically higher in settings where discrimination would be easier or more prevalent (Table IV).

Main denial regression (Table II). The estimating equation is a linear probability model regressing an indicator of lender denial on race/ethnicity dummies and controls (p.1475-1476):

D^{i}_{Lender,j} = \alpha_r \cdot \mathbf{1}[\text{race}_j = r] + \beta' X_j + \delta_l + \varepsilon_j, \tag{4}

where jj indexes applications; rr indexes race/ethnicity relative to non-Hispanic White; XjX_j includes the FICO-LTV-DTI grid (interactions of credit score bins, LTV bins, and DTI bins; see Table II notes for exact bin definitions), AUS denial recommendation (interacted with loan purpose and program), county-by-month fixed effects, loan amount bins, co-applicant indicator, and income bins (all covariates interacted with program and loan purpose); δl\delta_l is a lender fixed effect. Standard errors are clustered at the lender and county levels. Columns (1) to (3) vary the control set progressively; column (3) is the preferred full specification.

AUS denial regression (Table II, cols. 4-5). The dependent variable switches to the AUS denial indicator DAUS,jD_{AUS,j}, with the same right-hand side. Because AUS is color-blind by design, residual racial gaps in AUS recommendations reflect unobserved risk factors uu correlated with race (p.1477-1478).

Lender-specific excess denials and strictness correlation (Figure 2, Table III). For the 100 largest lenders, the race/ethnicity coefficients are allowed to vary by lender (i.e., the full col.(3) specification with lender-varying race/ethnicity slopes). These lender-specific excess denial estimates are plotted against lender strictness (lender FE from White-only denial regression). The correlation coefficient is reported (p.1481).

Loan performance validation (Figure 4). For 48 Ginnie Mae issuers matched to HMDA, the paper regresses 60-day delinquency within one year of origination on lender strictness. Residual riskiness is the lender FE from a delinquency regression controlling for flexible functions of DTI, LTV, credit score, and month dummies. Both raw and residual riskiness are negatively correlated with strictness, validating that strictness captures real overlay policies (p.1484-1486).

Indirect tests for discrimination (Table IV). The baseline specification (col.(3) of Table II) is augmented with interactions between race/ethnicity and (i) fintech indicator, (ii) top-4 lender county market share, and (iii) racially charged Google search rate, estimated separately for lender denials and AUS denials (p.1486-1489).

Service quality regressions (Table V). OLS regressions using NSMO individual-level data with controls including loan type, loan purpose, loan amount, credit score, income, LTV, self-employment status, co-applicant status, loan term, 11 LTV categories, 6 credit score categories, and 8 loan amount categories; all fully interacted with program and loan purpose. Survey fixed effects and county fixed effects included. Heteroskedasticity-robust standard errors. Outcomes are binary (processing delays, closing date postponement, satisfaction dummies), p.1492-1493.

DatasetRole in paperWiki page
HMDA (confidential, expanded 2018-2019)Main denial analysis: ~9 million applications with credit score, LTV, DTI, AUS recommendation/wiki/datasets/hmda/
National Survey of Mortgage Originations (NSMO)Survey component of NMDB; borrower satisfaction and service quality analysis (N=35,162)NSMO
National Mortgage Database (NMDB)Provides inquiry data for pre-application discouragement analysis; links HMDA to credit bureau recordsno page yet
Ginnie Mae securitization pool dataLoan performance (delinquency) validation for 48 matched issuers, 2018-2019 originationsno page yet

Sample: First-lien, 30-year fixed-rate mortgages on owner-occupied single-family properties; 2018-2019; applications through one of three main AUS (DU, LPA, TOTAL); excludes jumbo loans and withdrawn/incomplete applications. Final AUS-processed sample: ~8.9 million applications. NSMO subsample: 35,162 respondents for service quality regressions.

Read the paper if you are studying racial disparities in mortgage credit, the role of automated underwriting in fair lending, or methods for detecting discrimination in lending outcomes. Table II is the key excess denial table; Figure 2 and Table III document the lender strictness channel; Table IV presents the indirect discrimination tests; Table V covers service quality. The Internet Appendix (referenced throughout) contains additional robustness checks, the pre-application discouragement analysis, and the interest rate gap replication.

This article is a U.S. Government work and is in the public domain in the USA (per PDF p.1463 copyright notice: “Published 2025. This article is a U.S. Government work and is in the public domain in the USA.”). The Wiley/JF Crossref record carries the Wiley terms-and-conditions URL rather than a CC licence. Extract-only: not reproduced or human-verified here. LLM-distilled; not reproduced.

Bhutta, Neil, Aurel Hizmo, and Daniel Ringo (2025). “How Much Does Racial Bias Affect Mortgage Lending? Evidence from Human and Algorithmic Credit Decisions.” The Journal of Finance 80(3): 1463-1496. DOI: 10.1111/jofi.13444.

Found an error or want a topic covered? Open an issue, use the Edit page link above, or email contact@instituteforautomatedresearch.org. Edits are reviewed before publishing; provenance and accuracy are the point.