Definition
Examiner and judge leniency instruments are quasi-experimental identification strategies that exploit the quasi-random assignment of disability cases to decision-makers who differ systematically in how often they allow claims. Because a more lenient examiner or judge raises an applicant's award probability for reasons unrelated to the applicant's own characteristics, the assigned decision-maker's leniency serves as an instrument for disability insurance (DI) receipt. The design comes in two forms — the examiner instrument at the initial Disability Determination Services (DDS) stage (Maestas, Mullen, and Strand 2013) and the judge instrument at the Administrative Law Judge (ALJ) appeals stage (French and Song 2014; Black et al. 2024) — and underpins most of the credible causal estimates of DI's effects on labor supply and mortality.
Key Ideas
- The leniency design: within an office (or hearing office) and time period, cases are assigned to examiners/judges in a way unrelated to case merits (rotational or quasi-random). Decision-makers vary widely in allowance rates despite similar caseloads, so the assigned decision-maker's average allowance rate ("leniency") shifts award probability exogenously — a classic instrumental variable.
- Examiner IV (Maestas, Mullen, and Strand 2013): uses DDS examiner allowance-rate variation at the initial determination. Identifies a LATE for marginal applicants near the award threshold (≈23% of applicants, whose outcome an examiner's leniency can flip); estimated causal work reduction ≈26–28 pp.
- Judge IV (French and Song 2014): uses quasi-random rotational assignment of appealed cases to ALJs. 1.78M hearings (1990–99); first stage strong (a one-standard-deviation increase in judge leniency raises allowance ≈6.6 pp); causal labor-force-participation reduction ≈26 pp. Notably, OLS ≈ IV at the ALJ stage — selection biases roughly cancel there, unlike at the initial DDS stage.
- Different complier populations: the examiner IV's compliers are marginal initial applicants; the judge IV's are marginal appellants (a more motivated, more impaired subgroup who pursued appeal). Their convergence on ≈26–28 pp work-reduction estimates is the strongest evidence for a genuine causal effect. See Causal Effects of DI Receipt.
- Beyond labor supply — mortality (Black et al. 2024): applying the judge IV to 10-year mortality yields a positive LATE (+2.8 pp) for marginal appellants, but a marginal treatment effect decomposition shows the average masks heterogeneity — inframarginal (sicker) recipients benefit while marginal (healthier) recipients can be harmed via work disincentives. See DI Beneficiary Mortality.
- Identifying assumptions: validity requires (1) relevance — decision-makers genuinely differ in leniency (testable; first stages are strong, e.g., F = 271.5 in Strand and Messel 2019); (2) exclusion — the assigned decision-maker affects outcomes only through the award decision; (3) monotonicity — a more lenient decision-maker never makes an applicant less likely to be awarded. Under these, the estimand is a LATE for compliers, not the population average treatment effect.
How It Works
- Identify the assignment unit (DDS office × period, or hearing office × period) within which cases are quasi-randomly assigned.
- Construct each decision-maker's leniency as their leave-one-out allowance rate, conditioning on assignment-cell fixed effects so only the within-cell quasi-random variation is used.
- First stage: regress own award on assigned leniency (plus cell fixed effects). Second stage / 2SLS: instrument DI receipt with leniency to estimate its causal effect on the outcome (earnings, employment, mortality).
- The result is a complier-weighted LATE; combining instruments at different stages, or using an MTE framework, characterizes how the effect varies across the latent leniency margin.
Why It Matters
- The credible-evidence core: the examiner and judge designs produce the most-cited causal estimates of DI receipt's effect on labor supply (≈26–28 pp) and, more recently, mortality — replacing OLS comparisons that conflate the effect of benefits with the worse health of recipients. See Causal Effects of DI Receipt.
- Policy-margin interpretation: because each instrument identifies the effect for marginal cases at a particular stage, the estimates speak directly to the consequences of moving the allowance threshold (e.g., the Vocational Grid age cutoffs) — the population a stringency reform would actually affect.
- Heterogeneity and welfare: the MTE extension (Black et al. 2024) shows that average LATEs can hide opposite-signed effects for sicker vs. healthier recipients — central to the welfare analysis of DI generosity.
Open Questions
- How transportable are stage-specific LATEs (initial-applicant compliers vs. appellant compliers) across different policy reforms?
- Does quasi-random assignment hold cleanly given examiner/judge specialization, case-routing rules, and attorney representation at the ALJ stage?
- Can examiner and judge instruments be combined to trace the full MTE curve across the leniency margin?
Related
Sources