Optimal Insurance: Gershkov, Moldovanu, Strack & Zhang (2023)
Distilled by claude-sonnet-4-6 · extracted Jun 25, 2026, verified Jun 25, 2026
JEL (IAR-assigned): D82, D86, G22 · assigned from the abstract, not the journal
What this is. The paper’s core theoretical results, the dual-utility insurance model, and the mechanism design method: enough to understand what was found and how, without reading all 34 pages. To replicate or extend the results, read the full source at the original.
Gershkov, Moldovanu, Strack, and Zhang study a monopoly insurance market under adverse selection where agents have dual utility (Yaari 1987), meaning they weight probabilities via a distortion function rather than taking expectations, and face random losses whose distribution is correlated with their privately known risk type. This generalizes Rothschild and Stiglitz (1976) and Stiglitz (1977) in two directions: from expected utility to dual utility, and from a single fixed loss to a random distribution of losses correlated with the agent’s type.
The central result (Theorem 1) characterizes the optimal retention function as a “layer contract”: for each loss level, the agent either retains the entire marginal loss or the insurer covers it entirely, so the derivative of the retention function is in {0, 1} almost everywhere. Whether the optimal menu uses deductibles or coverage limits depends on the direction of private information. When the agent’s type governs loss probability, the virtual value single-crosses from below, and deductibles are optimal. When the type governs loss magnitude, coverage limits are optimal, even though they are the worst possible contract for any given expected cost (Theorem 2). The welfare gain from reduced information rents dominates the efficiency loss from offering the worst contract form.
In contrast to Chade and Schlee (2012) under expected utility, where full insurance is never optimal, full insurance to some types can be optimal here because dual utility agents exhibit first-order risk aversion.
Core results
Section titled “Core results”All results are theoretical; locators point to theorems, propositions, and examples in the source PDF.
| # | Result | Locator | Key statement |
|---|---|---|---|
| R1 | Layer contract structure: the optimal retention function satisfies almost everywhere; each marginal dollar of loss is either fully retained by the agent or fully covered by the insurer | Theorem 1, pp. 2595-2596 | The profit-maximization objective is linear in ; extreme points of the feasible set of Lipschitz-1 retention functions satisfy a.e. (Bauer’s maximum principle) |
| R2 | SOSD ranking of contract forms: for any strongly risk-averse agent (averse to mean-preserving spreads), the deductible contract second-order stochastically dominates any doubly monotone contract with the same expected cost, which dominates the coverage limit contract | Theorem 2, p. 2600 | for any contract with the same expected cost; deductibles are welfare-maximizing at fixed cost |
| R3 | Deductible menu is optimal when the virtual value crosses zero from below in : the insurer offers a menu of deductible-premium pairs | Theorem 3(i), p. 2600; Example 2, pp. 2600-2601 | Holds when private information concerns loss probability (equation (1)); the optimal deductible is nonincreasing in the agent’s degree of loss aversion |
| R4 | Coverage limit menu is optimal when crosses zero from above in : the insurer offers a menu of cap-premium pairs | Theorem 3(ii), p. 2600; Example 3, p. 2602 | Holds when private information concerns loss magnitude; in Example 3, is increasing in |
| R5 | Higher risk aversion raises insurer profit: if for all (agent 2 is more risk averse than agent 1), the insurer’s profit under is strictly higher | Proposition 2, p. 2599 | More risk-averse agents value coverage more; for any fixed retention function the premium extractable from a more risk-averse agent is higher, and adjusting to the optimal retention amplifies the gain |
| R6 | Finite losses: with possible loss levels , the optimal deductible menu uses at most contracts; each offered deductible equals one of the loss values or 0 | Proposition 3, p. 2603 | The optimal mechanism is a basic deductible-premium contract plus a finite ladder of add-on fees that progressively reduce the deductible; full insurance is one possible top rung |
| R7 | Single fixed loss (Stiglitz 1977 case): the optimal menu offers either full insurance (deductible = 0) or no insurance (deductible = ); partial insurance is never optimal | Corollary 1, p. 2604 | Follows from Proposition 3 with ; extends the DeFeo and Hindriks (2014) result to dual utility; contrasts with Chade and Schlee (2012) under EU where full insurance is never optimal |
Overall (paper’s conclusion). Layer contracts (deductibles and coverage limits) emerge endogenously from first principles under dual utility with adverse selection. The form of the optimal menu depends on which component of the agent’s private information drives the loss distribution, explaining the structural difference observed in practice between property/casualty insurance (deductibles) and medical malpractice insurance (coverage limits).
Theory / model
Section titled “Theory / model”Setup. An agent faces a random loss distributed on . The agent’s private type parameterizes the conditional loss distribution via . Higher types face stochastically larger losses in first-order stochastic dominance; is decreasing in . Types are distributed according to with density (p. 2587).
Two canonical cases illustrate the model. When the type is the accident probability (Rothschild and Stiglitz 1976; Stiglitz 1977 classical setting):
where is a fixed conditional loss distribution given an accident (p. 2588, equation 1). When the type scales the loss magnitude:
so all types face the same accident probability but higher types incur proportionally larger losses (p. 2588).
Dual (Yaari) utility. Agents have Yaari (1987) dual utility determined by a probability distortion function , increasing, absolutely continuous, with (weak risk aversion). For a random total loss distributed according to , the certainty equivalent is (p. 2589):
Dual utility modifies the expectation operator by reweighting each loss level by ; the condition implies the agent overweights the probability of large losses, generating first-order risk aversion (the risk premium is proportional to the standard deviation of the loss, not its variance). A key property for tractability is additivity in nonrandom transfers: , which makes the mechanism design problem separable in premia (p. 2589).
Koszegi and Rabin (2006) loss-averse preferences with linear utility over outcomes correspond to the distortion for , a special case of dual utility (p. 2590).
Insurance contracts. A direct mechanism offers a menu of retention functions and premia , where is the share of loss retained by type . Two ex post moral hazard conditions (Assumption 1, p. 2591) restrict the retention slope: (the agent cannot inflate a reported loss to reduce retention), and (the agent cannot hide part of a loss to claim higher indemnity). Together these require almost everywhere.
Under the additivity of dual utility, the certainty equivalent of contract to type is (p. 2591):
Method
Section titled “Method”Incentive compatibility envelope. Dual utility’s linearity in permits a simple envelope characterization (Proposition 1, p. 2593). Any incentive-compatible mechanism satisfies:
Integration by parts yields the insurer’s expected profit as a functional of alone:
where the virtual value (analogous to the Myerson virtual value in mechanism design) is (p. 2593):
The first two terms measure the efficiency gain from covering the marginal loss at level for type (the agent’s valuation exceeds the insurer’s cost because of risk aversion). The third term is the information rent cost: since , the product is nonpositive, so higher types require larger information rents that reduce the effective profit from insuring them.
Pointwise linear optimization. For each , the integrand in (5) is linear in . The feasible set of retention functions satisfying Assumption 1 is convex. By Bauer’s maximum principle, the maximum is attained at an extreme point of this set, and extreme points of the unit ball of Lipschitz-1 functions on have derivative in almost everywhere (p. 2608). The pointwise optimum therefore sets:
When is nondecreasing in for all (the regularity condition of Theorem 1), the resulting is submodular: for , meaning higher-risk types receive more coverage at each loss level. Proposition 1(ii) shows submodularity of is sufficient for incentive compatibility; thus the pointwise solution is also the global optimum of the original screening problem.
Deductible and coverage limit forms. A deductible and a coverage limit correspond to (p. 2599):
Whether crosses zero from below (deductibles) or from above (coverage limits) as a function of is governed by the direction of private information. In the loss-probability case (equation (1)), typically crosses from below, yielding deductibles optimal. In the loss-magnitude case, crosses from above, yielding coverage limits optimal. Gershkov et al. (2022) develop related tools for nonexpected utility in auction settings.
Liang, Zou, and Jiang (2022) study a two-type variant with distortion risk measures; the full continuum is handled here by the pointwise separability of the profit functional.
Empirical specifications
Section titled “Empirical specifications”This paper presents no empirical analysis; all results are theorems, propositions, and corollaries derived from the model. Examples 1 through 4 illustrate the general results with specific parametric choices for the distortion function and the type-conditional loss distribution , but use no data. The paper cites empirical patterns from prior literature (household deductible choices in Sydnor 2010; Barseghyan et al. 2013; medical malpractice coverage limits in Silver et al. 2015) for motivation only.
Datasets used
Section titled “Datasets used”| Dataset | Role in paper | Wiki page |
|---|---|---|
| None | Pure theoretical analysis; no datasets collected or analyzed | n/a |
This is a theory paper. Empirical patterns cited in the introduction and related literature (household insurance data, malpractice claims) come from prior published work, not new data collection.
When to read the full paper
Section titled “When to read the full paper”Use the original if you are: extending the mechanism design approach to other nonexpected utility frameworks beyond dual utility; studying competitive (multi-insurer) adverse selection with probability-distorting agents; analyzing the welfare implications of mandating deductibles vs coverage limits in regulated insurance markets; examining when coinsurance (linear contracts, which are never optimal under this framework) is appropriate; or working through the formal proofs and examples in the appendix and online appendices.
Attribution and rights
Section titled “Attribution and rights”Source: peer-reviewed, American Economic Review 113(10), October 2023. This distillation was extracted by an LLM on 2026-06-25 and is not human-verified or independently reproduced. The paper is under AEA copyright; no CC license was found. Extract-only: do not reproduce the full text.
Gershkov, Alex, Benny Moldovanu, Philipp Strack, and Mengxi Zhang. “Optimal Insurance: Dual Utility, Random Losses, and Adverse Selection.” American Economic Review 113, no. 10 (October 2023): 2581–2614. DOI: 10.1257/aer.20221247. © 2023 American Economic Association.