Definition
Essential heterogeneity is the condition in which individuals have private information about their own returns to a treatment — specifically, that the individual-level treatment effect βi=Yi(1)−Yi(0) is partly known to the agent at the time of the treatment decision and influences selection into treatment. Named and formalized by Heckman and Vytlacil (1999, 2005), essential heterogeneity is "essential" because it cannot be removed by conditioning on observed covariates: the selection process is driven in part by private returns unobservable to the econometrician. When essential heterogeneity is present, different valid instruments systematically yield different instrumental variable (IV) estimates — not because any instrument violates exclusion, but because each instrument moves a different complier population at a different region of the Marginal Treatment Effect (MTE) function.
Key Ideas
- Selection on private gains: Agents with higher returns to treatment are more likely to select into it. The treated population therefore has above-average treatment effects (average treatment effect on the treated [ATT] exceeds the average treatment effect [ATE]). Ordinary least squares (OLS) conflates this selection with the causal return. IV corrects for selection bias but does not recover ATE — it recovers a local average treatment effect (LATE) for the specific complier population moved by the instrument.
- IV estimates as MTE integrals: Every IV estimator identifies a weighted integral ∫MTE(u)⋅wIV(u)du of the MTE function, with instrument-specific weights wIV(u) concentrated at the propensity-score thresholds where the instrument moves individuals into or out of treatment. OLS, IV, matching, and difference-in-differences (DiD) all recover different weighted averages of the same underlying MTE.
- Direction of bias is instrument-specific: Whether IV > OLS or IV < OLS depends on whether the compliers identified by a given instrument have above- or below-average private returns. For schooling instruments like college proximity (which move credit-constrained students), IV may identify low-propensity compliers with high returns (IV > OLS). For the GI Bill instrument, IV identifies compliers with below-average returns (IV < OLS).
- Willis and Rosen (1979) as precursor: Willis and Rosen applied Roy's (1951) comparative advantage model to educational choice fifteen years before the LATE framework was formalized. They showed that college attendance is driven by comparative advantage in two-sector earnings (w1,w0) and that marginal students induced into college by any instrument have lower average returns than infra-marginal attenders — the core essential heterogeneity prediction applied to education. They did not use the language of LATE or MTE, but their factor model implicitly defines what would later be called the selection propensity and the MTE function. See Willis and Rosen 1979 — Education and Self-Selection, Sherwin Rosen, Robert J. Willis.
- Inter-study inconsistency without contradiction: When essential heterogeneity is present, two valid IV studies with different instruments may produce significantly different estimates without either being wrong. Each study correctly identifies the LATE for its own complier population; the divergence reflects genuine MTE heterogeneity, not identification failure.
- Testable implication: A regression of Y on the estimated propensity score P^(Z) will be nonlinear if essential heterogeneity is present (because MTE varies with u). A flat MTE function implies zero essential heterogeneity and predicts a linear relationship. The nonlinearity test has low power in typical samples.
- Ex ante vs. ex post returns (Cunha and Heckman 2007): Approximately half of ex post return variability in schooling was not foreseeable at decision time — agents partly select on noise. This dampens essential heterogeneity relative to a world of perfect private information: agents cannot fully sort on returns they cannot observe.
How It Works
Let UD be the latent propensity index driving treatment selection: individual i selects into treatment when UDi≤P(Zi) (where P(Z) is the propensity score). The MTE is defined as MTE(u)=E[βi∣UDi=u] — the average treatment effect for individuals who are exactly indifferent between treatment and non-treatment at propensity-score level u. Under essential heterogeneity, MTE(u) is decreasing in u: individuals who are "hardest to convince" (high u, need a strong instrument nudge) have the lowest average returns. An instrument that shifts P(Z) across a narrow range [u0,u1] identifies MTE integrated over [u0,u1] — a local average over a specific segment of the selection distribution.
Why It Matters
- External validity across studies: Essential heterogeneity implies that LATE estimates from different instruments are not portable across contexts — even if all instruments are valid. A collection of internally valid studies may not aggregate into a single structural parameter without knowledge of the MTE function.
- Policy extrapolation: Predicting the effect of a new policy requires knowing MTE across the full propensity-score distribution, which requires structural modeling (the P2/P3 hierarchy in Instrumental Variables). Reduced-form IV alone is insufficient for policy design beyond the margin already identified.
- Rationalization of the schooling IV literature: Essential heterogeneity provides a coherent explanation for why different valid schooling instruments yield different returns (Card 1999), and why GI Bill estimates fall below OLS while proximity-to-college estimates exceed OLS — without positing instrument invalidity in either case. See Returns to Schooling.
Open Questions
- How prevalent is essential heterogeneity empirically? The nonlinearity test has low power and has rarely been applied systematically across literatures.
- In the Disability Insurance (DI) context (examiner IV vs. administrative law judge [ALJ]-lottery IV), does the convergence of the Maestas, Mullen, and Strand (MMS, 2013) and French and Song (2014) estimates rule out substantial essential heterogeneity, or does it reflect coincidentally similar MTE values at the two margins?
- Can the ex ante/ex post decomposition of Cunha and Heckman be extended to DI: how much of a worker's disability risk was knowable at career entry, and does this private knowledge shape application behavior?
Related
Sources