Skip to content

Optimal Insurance: Gershkov, Moldovanu, Strack & Zhang (2023)

Distilled by claude-sonnet-4-6 · extracted Jun 25, 2026, verified Jun 25, 2026

JEL (IAR-assigned): D82, D86, G22 · assigned from the abstract, not the journal

Full structured metadata (methods, scope, relatesTo, topics, datasets): raw Markdown (.md)

paper-summaryinsuranceadverse-selectionmechanism-designcontract-theorypeer-reviewedunreplicated

What this is. The paper’s core theoretical results, the dual-utility insurance model, and the mechanism design method: enough to understand what was found and how, without reading all 34 pages. To replicate or extend the results, read the full source at the original.

Gershkov, Moldovanu, Strack, and Zhang study a monopoly insurance market under adverse selection where agents have dual utility (Yaari 1987), meaning they weight probabilities via a distortion function rather than taking expectations, and face random losses whose distribution is correlated with their privately known risk type. This generalizes Rothschild and Stiglitz (1976) and Stiglitz (1977) in two directions: from expected utility to dual utility, and from a single fixed loss to a random distribution of losses correlated with the agent’s type.

The central result (Theorem 1) characterizes the optimal retention function as a “layer contract”: for each loss level, the agent either retains the entire marginal loss or the insurer covers it entirely, so the derivative of the retention function is in {0, 1} almost everywhere. Whether the optimal menu uses deductibles or coverage limits depends on the direction of private information. When the agent’s type governs loss probability, the virtual value single-crosses from below, and deductibles are optimal. When the type governs loss magnitude, coverage limits are optimal, even though they are the worst possible contract for any given expected cost (Theorem 2). The welfare gain from reduced information rents dominates the efficiency loss from offering the worst contract form.

In contrast to Chade and Schlee (2012) under expected utility, where full insurance is never optimal, full insurance to some types can be optimal here because dual utility agents exhibit first-order risk aversion.

All results are theoretical; locators point to theorems, propositions, and examples in the source PDF.

#ResultLocatorKey statement
R1Layer contract structure: the optimal retention function satisfies R/l{0,1}\partial R / \partial l \in \{0,1\} almost everywhere; each marginal dollar of loss is either fully retained by the agent or fully covered by the insurerTheorem 1, pp. 2595-2596The profit-maximization objective is linear in R/l\partial R / \partial l; extreme points of the feasible set of Lipschitz-1 retention functions satisfy R/l{0,1}\partial R / \partial l \in \{0,1\} a.e. (Bauer’s maximum principle)
R2SOSD ranking of contract forms: for any strongly risk-averse agent (averse to mean-preserving spreads), the deductible contract second-order stochastically dominates any doubly monotone contract with the same expected cost, which dominates the coverage limit contractTheorem 2, p. 2600RC(,θ)SOSDRI(,θ)SOSDRD(,θ)R_C(\cdot,\theta) \leq_{\text{SOSD}} R_I(\cdot,\theta) \leq_{\text{SOSD}} R_D(\cdot,\theta) for any contract II with the same expected cost; deductibles are welfare-maximizing at fixed cost
R3Deductible menu is optimal when the virtual value J(l,θ)J(l,\theta) crosses zero from below in ll: the insurer offers a menu of deductible-premium pairs (D(θ),t(θ))(D(\theta), t(\theta))Theorem 3(i), p. 2600; Example 2, pp. 2600-2601Holds when private information concerns loss probability (equation (1)); the optimal deductible D(θ)D^*(\theta) is nonincreasing in the agent’s degree of loss aversion
R4Coverage limit menu is optimal when J(l,θ)J(l,\theta) crosses zero from above in ll: the insurer offers a menu of cap-premium pairs (C(θ),t(θ))(C(\theta), t(\theta))Theorem 3(ii), p. 2600; Example 3, p. 2602Holds when private information concerns loss magnitude; in Example 3, C(θ)=θ2f(θ)/[2(1F(θ))]C^*(\theta) = \theta^2 f(\theta) / [2(1 - F(\theta))] is increasing in θ\theta
R5Higher risk aversion raises insurer profit: if g2(p)<g1(p)g_2(p) < g_1(p) for all p(0,1)p \in (0,1) (agent 2 is more risk averse than agent 1), the insurer’s profit under g2g_2 is strictly higherProposition 2, p. 2599More risk-averse agents value coverage more; for any fixed retention function the premium extractable from a more risk-averse agent is higher, and adjusting to the optimal retention amplifies the gain
R6Finite losses: with nn possible loss levels l1<<lnl_1 < \dots < l_n, the optimal deductible menu uses at most n+1n+1 contracts; each offered deductible equals one of the nn loss values or 0Proposition 3, p. 2603The optimal mechanism is a basic deductible-premium contract plus a finite ladder of add-on fees that progressively reduce the deductible; full insurance is one possible top rung
R7Single fixed loss (Stiglitz 1977 case): the optimal menu offers either full insurance (deductible = 0) or no insurance (deductible = lˉ\bar l); partial insurance is never optimalCorollary 1, p. 2604Follows from Proposition 3 with n=1n = 1; extends the DeFeo and Hindriks (2014) result to dual utility; contrasts with Chade and Schlee (2012) under EU where full insurance is never optimal

Overall (paper’s conclusion). Layer contracts (deductibles and coverage limits) emerge endogenously from first principles under dual utility with adverse selection. The form of the optimal menu depends on which component of the agent’s private information drives the loss distribution, explaining the structural difference observed in practice between property/casualty insurance (deductibles) and medical malpractice insurance (coverage limits).

Setup. An agent faces a random loss LL distributed on [0,Lˉ][0, \bar L]. The agent’s private type θΘ=[θ,θˉ]\theta \in \Theta = [\underline\theta, \bar\theta] parameterizes the conditional loss distribution Hθ:R+[0,1]H_\theta : \mathbb{R}_+ \to [0,1] via Hθ(l)=Pr(Llθ)H_\theta(l) = \Pr(L \leq l \mid \theta). Higher types face stochastically larger losses in first-order stochastic dominance; HθH_\theta is decreasing in θ\theta. Types are distributed according to FF with density ff (p. 2587).

Two canonical cases illustrate the model. When the type is the accident probability (Rothschild and Stiglitz 1976; Stiglitz 1977 classical setting):

Hθ(l)=(1θ)+θQ(l),(1)H_\theta(l) = (1 - \theta) + \theta Q(l), \tag{1}

where QQ is a fixed conditional loss distribution given an accident (p. 2588, equation 1). When the type scales the loss magnitude:

Hθ(l)=Q(l/θ),H_\theta(l) = Q(l / \theta),

so all types face the same accident probability but higher types incur proportionally larger losses (p. 2588).

Dual (Yaari) utility. Agents have Yaari (1987) dual utility determined by a probability distortion function g:[0,1][0,1]g : [0,1] \to [0,1], increasing, absolutely continuous, with g(p)pg(p) \leq p (weak risk aversion). For a random total loss xx distributed according to HH, the certainty equivalent is (p. 2589):

CE(x)=0[1g(H(s))]ds.(2)CE(x) = -\int_0^\infty \left[1 - g(H(s))\right] ds. \tag{2}

Dual utility modifies the expectation operator by reweighting each loss level ss by g(H(s))g'(H(s)); the condition g(p)pg(p) \leq p implies the agent overweights the probability of large losses, generating first-order risk aversion (the risk premium is proportional to the standard deviation of the loss, not its variance). A key property for tractability is additivity in nonrandom transfers: CE(x+t)=CE(x)+tCE(x + t) = CE(x) + t, which makes the mechanism design problem separable in premia (p. 2589).

Koszegi and Rabin (2006) loss-averse preferences with linear utility over outcomes correspond to the distortion g(p)=(2λ)p+(λ1)p2g(p) = (2-\lambda)p + (\lambda-1)p^2 for λ(1,2]\lambda \in (1,2], a special case of dual utility (p. 2590).

Insurance contracts. A direct mechanism offers a menu of retention functions R(,θ)R(\cdot, \theta) and premia t(θ)t(\theta), where R(l,θ)[0,l]R(l, \theta) \in [0, l] is the share of loss ll retained by type θ\theta. Two ex post moral hazard conditions (Assumption 1, p. 2591) restrict the retention slope: (i)(i) R(l,θ)/l0\partial R(l,\theta)/\partial l \geq 0 (the agent cannot inflate a reported loss to reduce retention), and (ii)(ii) (lR(l,θ))/l=1R/l0\partial(l - R(l,\theta))/\partial l = 1 - \partial R/\partial l \geq 0 (the agent cannot hide part of a loss to claim higher indemnity). Together these require R(l,θ)/l[0,1]\partial R(l,\theta)/\partial l \in [0,1] almost everywhere.

Under the additivity of dual utility, the certainty equivalent of contract (R(,θ),t(θ))(R(\cdot,\theta), t(\theta)) to type θ\theta is (p. 2591):

U(θ)=t(θ)0Lˉ[1g(Hθ(l))]R(l,θ)ldl.(3)U(\theta) = -t(\theta) - \int_0^{\bar L} \left[1 - g(H_\theta(l))\right] \frac{\partial R(l,\theta)}{\partial l}\, dl. \tag{3}

Incentive compatibility envelope. Dual utility’s linearity in R/l\partial R / \partial l permits a simple envelope characterization (Proposition 1, p. 2593). Any incentive-compatible mechanism satisfies:

U(θ)=U(θ)+θθ[0LˉR(l,s)lg(Hs(l))Hs(l)sdl]ds.(4)U(\theta) = U(\underline\theta) + \int_{\underline\theta}^{\theta} \left[\int_0^{\bar L} \frac{\partial R(l,s)}{\partial l}\, g'(H_s(l))\, \frac{\partial H_s(l)}{\partial s}\, dl \right] ds. \tag{4}

Integration by parts yields the insurer’s expected profit as a functional of RR alone:

π(R)=θθˉ[E[L(θ)]0LˉR(l,θ)lJ(l,θ)dl]f(θ)dθU(θ),(5)\pi(R) = \int_{\underline\theta}^{\bar\theta} \left[-E[L(\theta)] - \int_0^{\bar L} \frac{\partial R(l,\theta)}{\partial l}\, J(l,\theta)\, dl \right] f(\theta)\, d\theta - U(\underline\theta), \tag{5}

where the virtual value (analogous to the Myerson virtual value in mechanism design) is (p. 2593):

J(l,θ)=Hθ(l)g(Hθ(l))+1F(θ)f(θ)g(Hθ(l))Hθ(l)θ.(6)J(l,\theta) = H_\theta(l) - g(H_\theta(l)) + \frac{1-F(\theta)}{f(\theta)}\, g'(H_\theta(l))\, \frac{\partial H_\theta(l)}{\partial \theta}. \tag{6}

The first two terms Hθ(l)g(Hθ(l))0H_\theta(l) - g(H_\theta(l)) \geq 0 measure the efficiency gain from covering the marginal loss at level ll for type θ\theta (the agent’s valuation exceeds the insurer’s cost because of risk aversion). The third term is the information rent cost: since Hθ/θ<0\partial H_\theta / \partial\theta < 0, the product g(Hθ(l))(Hθ/θ)g'(H_\theta(l))(\partial H_\theta/\partial\theta) is nonpositive, so higher types require larger information rents that reduce the effective profit from insuring them.

Pointwise linear optimization. For each θ\theta, the integrand in (5) is linear in R(l,θ)/l\partial R(l,\theta)/\partial l. The feasible set of retention functions satisfying Assumption 1 is convex. By Bauer’s maximum principle, the maximum is attained at an extreme point of this set, and extreme points of the unit ball of Lipschitz-1 functions on [0,Lˉ][0, \bar L] have derivative in {0,1}\{0,1\} almost everywhere (p. 2608). The pointwise optimum therefore sets:

R(l,θ)l=1{J(l,θ)0}.(7)\frac{\partial R^*(l,\theta)}{\partial l} = \mathbf{1}\{J(l,\theta) \leq 0\}. \tag{7}

When J(l,θ)J(l,\theta) is nondecreasing in θ\theta for all ll (the regularity condition of Theorem 1), the resulting RR^* is submodular: R(l,θ)/lR(l,θ)/l\partial R(l,\theta') / \partial l \leq \partial R(l,\theta) / \partial l for θ>θ\theta' > \theta, meaning higher-risk types receive more coverage at each loss level. Proposition 1(ii) shows submodularity of RR is sufficient for incentive compatibility; thus the pointwise solution is also the global optimum of the original screening problem.

Deductible and coverage limit forms. A deductible D(θ)D(\theta) and a coverage limit C(θ)C(\theta) correspond to (p. 2599):

RD(l,θ)={l,l<D(θ)D(θ),lD(θ),RC(l,θ)={0,lC(θ)lC(θ),l>C(θ).(8)R_D(l,\theta) = \begin{cases} l, & l < D(\theta) \\ D(\theta), & l \geq D(\theta) \end{cases}, \qquad R_C(l,\theta) = \begin{cases} 0, & l \leq C(\theta) \\ l - C(\theta), & l > C(\theta) \end{cases}. \tag{8}

Whether J(l,θ)J(l,\theta) crosses zero from below (deductibles) or from above (coverage limits) as a function of ll is governed by the direction of private information. In the loss-probability case (equation (1)), JJ typically crosses from below, yielding deductibles optimal. In the loss-magnitude case, JJ crosses from above, yielding coverage limits optimal. Gershkov et al. (2022) develop related tools for nonexpected utility in auction settings.

Liang, Zou, and Jiang (2022) study a two-type variant with distortion risk measures; the full continuum is handled here by the pointwise separability of the profit functional.

This paper presents no empirical analysis; all results are theorems, propositions, and corollaries derived from the model. Examples 1 through 4 illustrate the general results with specific parametric choices for the distortion function gg and the type-conditional loss distribution HθH_\theta, but use no data. The paper cites empirical patterns from prior literature (household deductible choices in Sydnor 2010; Barseghyan et al. 2013; medical malpractice coverage limits in Silver et al. 2015) for motivation only.

DatasetRole in paperWiki page
NonePure theoretical analysis; no datasets collected or analyzedn/a

This is a theory paper. Empirical patterns cited in the introduction and related literature (household insurance data, malpractice claims) come from prior published work, not new data collection.

Use the original if you are: extending the mechanism design approach to other nonexpected utility frameworks beyond dual utility; studying competitive (multi-insurer) adverse selection with probability-distorting agents; analyzing the welfare implications of mandating deductibles vs coverage limits in regulated insurance markets; examining when coinsurance (linear contracts, which are never optimal under this framework) is appropriate; or working through the formal proofs and examples in the appendix and online appendices.

Source: peer-reviewed, American Economic Review 113(10), October 2023. This distillation was extracted by an LLM on 2026-06-25 and is not human-verified or independently reproduced. The paper is under AEA copyright; no CC license was found. Extract-only: do not reproduce the full text.

Gershkov, Alex, Benny Moldovanu, Philipp Strack, and Mengxi Zhang. “Optimal Insurance: Dual Utility, Random Losses, and Adverse Selection.” American Economic Review 113, no. 10 (October 2023): 2581–2614. DOI: 10.1257/aer.20221247. © 2023 American Economic Association.

Found an error or want a topic covered? Open an issue, use the Edit page link above, or email contact@instituteforautomatedresearch.org. Edits are reviewed before publishing; provenance and accuracy are the point.