Definition
Causal effects are heterogeneous when the treatment effect Y1−Y0 varies across individuals — so that no single average adequately summarizes the distribution. Heterogeneous treatment effects (HTE) is the recognition that most treatments affect different people differently, and the research agenda of characterizing how and for whom effects vary. The central question is whether the average effect estimated by a particular design (ATE, ATT, LATE, MTE) generalizes to the population of interest, or whether it masks opposing effects that partially cancel in the average.
Key Ideas
Treatment Effect Parameters Under Heterogeneity
| Parameter |
Definition |
Identifies Effect For |
| ATE |
E[Y1−Y0] |
Full population |
| ATT |
E[Y1−Y0∣D=1] |
The treated |
| ATC |
E[Y1−Y0∣D=0] |
The untreated |
| LATE |
E[Y1−Y0∣complier] |
Compliers defined by a specific instrument |
| MTE |
E[Y1−Y0∣V=u] |
Individuals at the margin of indifference at propensity threshold u |
| CATE |
E[Y1−Y0∣X=x] |
Subgroup with observed characteristics X=x |
When treatment effects are homogeneous, all five coincide. In practice they rarely do, and the difference matters for policy.
Propensity-Based Heterogeneity
A key approach, developed in Brand and colleagues' work, stratifies the sample by estimated propensity score p^(X) and estimates conditional average treatment effects (CATE) within strata. This reveals how effects vary as a function of how likely individuals are to receive the treatment, going beyond simple subgroup breakdowns.
Negative selection (Brand and Xie 2010): If treatment effects are negatively correlated with the propensity to receive the treatment, those least likely to be treated benefit most from it. Brand and Xie (2010) document this for college wage returns across the National Longitudinal Survey of Youth 1979 (NLSY79) and the Wisconsin Longitudinal Study — 10 consistent negative Level-2 slopes across two datasets, two sexes, and multiple career stages. For NLSY men, the estimated wage return ranges from ~30% in the lowest propensity stratum to ~10% in the highest — a 20 pp gap. The mechanism is counterfactual deprivation: the negative pattern arises not because low-propensity graduates earn more, but because low-propensity non-graduates earn very little without a degree; high-propensity individuals can fall back on their superior resources and abilities even without college. This directly contradicts the common assumption of positive selection (that the most capable, most likely to attend, also benefit most). Under positive selection, ordinary least squares (OLS) overestimates effects on the untreated; under negative selection, OLS underestimates the benefit of expanding treatment to reluctant participants — ATT < ATE < ATC.
Negative selection in fertility effects (Brand and Davis 2011): The same propensity-score hierarchical linear model (HLM) framework extended to fertility outcomes. Women least likely to attend college experience the largest fertility-decreasing effects of college. For attendance, the Level-2 slope = +0.10 (p<0.05); stratum 1 women have 65% fewer children than comparable non-attenders. For college completion, the Level-2 slope = +0.17 (p<0.01), and the effect fully reverses for high-propensity women (+42% more children by age 41) — advantaged women use college as a complement to, rather than substitute for, family formation. Unlike earnings, however, the normative valence of the treatment effect is ambiguous: lower fertility for disadvantaged women may represent escape from early disadvantaged family formation, but also forecloses desired childbearing.
Multi-Outcome Sorting: "Sorting on the Mix" (Brooks, Chapman, and Schroeder 2018)
Prior discussions of essential heterogeneity assumed agents sort on expected gains from a single outcome of interest, with treatment effects on other outcomes (e.g. costs, adverse events) either absent or uncorrelated. In real-world clinical and policy settings, treatment choices reflect a weighted assessment of expected effects across multiple beneficial and detrimental outcomes. Brooks, Chapman, and Schroeder (2018) label this "sorting on the mix" and use simulation (5 scenarios × 1000 runs × 5000 patients) to derive its implications:
- Identification still holds per outcome: Regression → ATT and instrumental variables (IV) → LATE for each outcome separately. The single-outcome estimand results generalize.
- But ATT and LATE are sensitive to the benefit/detriment correlation structure: Even when the full distribution of benefit effects is identical across simulated populations, the true values of ATTB and LATEB differ substantially across scenarios because the correlation between benefit and detriment effects shifts which patients choose treatment and who the marginal patients are.
- Treatment-risk paradox (positive benefit/detriment correlation): Patients with the highest expected benefit from treatment also have the highest expected detriment risk. Rational sorting on the mix leads them to forgo treatment — so high-benefit patients are under-treated relative to naive single-outcome reasoning. Forcing high-risk patients toward higher treatment rates (to match lower-risk patients' rates) would cause intolerable adverse outcomes.
- Diagnosing over/underuse: When LATE estimates of benefit and detriment are combined with outcome valuations, the joint pattern of ATT vs. LATE across outcomes reveals whether treatment rates are optimal in the study population. In scenarios where detriment expectations understate true detriment risk, LATED × value exceeds LATEB × value → overuse.
- External validity requires matching correlation structures: A LATE from one population cannot be safely generalized to another unless the correlations between benefit and detriment effects across patients are similar. This is a substantially stronger external validity condition than is typically acknowledged in observational healthcare studies.
Essential Heterogeneity (Heckman-Vytlacil)
Essential heterogeneity exists when individuals select into treatment partly because they privately know their own above-average returns: Cov[Y1−Y0,D]>0. Under essential heterogeneity:
- OLS overestimates ATE (treated have higher returns, not just higher baseline)
- LATE is instrument-specific — different instruments induce different complier margins and thus different LATEs
- No single LATE generalizes unless the policy to be evaluated creates the same complier margin as the instrument
The Marginal Treatment Effect (MTE) framework (Heckman and Vytlacil 1999, 2005) makes essential heterogeneity precise and recovers the full distribution of effects as a function of the selection index. See Marginal Treatment Effect.
Observable vs. Unobservable Heterogeneity
- Observable heterogeneity: Effects vary as a function of measured pre-treatment covariates X — estimated by subgroup analysis or interaction models.
- Unobservable heterogeneity: Effects vary as a function of unmeasured characteristics, including private information about one's own returns (essential heterogeneity) — requires structural assumptions or instruments to recover.
Propensity-score stratification (Brand and Simon-Thomas 2012) works under selection-on-observables; MTE methods address unobservable selection on gains.
Why It Matters
- External validity: A LATE identifies the effect for compliers of a specific instrument. Whether that generalizes depends on whether the complier population is representative. HTE analysis characterizes the full distribution, making external validity an empirical question rather than an assumption.
- Policy targeting: If effect heterogeneity is systematic — e.g., effects are larger for low–socioeconomic-status (SES) students (negative selection in education) or for young workers with musculoskeletal disorders (Disability Insurance [DI] work disincentive) — policies can be targeted to the subgroups where benefit-cost ratios are highest.
- Interpreting conflicting estimates: When different studies find different effect magnitudes, HTE is the first explanation to check: the studies may have identified effects for different complier or subgroup populations.
- Welfare analysis: A positive ATE can coexist with harm to some recipients (negative individual effects in the distribution); HTE characterizes the distribution of gains and losses, which matters for distributional welfare assessments.
Open Questions
- Machine learning methods: Causal forests, causal trees, and LASSO-based interaction selection can uncover HTE in high-dimensional covariate spaces (Wager and Athey 2018; Brand, Zhou, and Xie 2023). How to do valid inference on discovered heterogeneity without double-dipping remains active research.
- HTE with unobservable selection: Propensity-based stratification requires selection-on-observables. When selection is on unobservables (gains), MTE or other structural approaches are needed. Bridging the two frameworks remains incomplete.
- Publication bias: Subgroup analyses are prone to false positive heterogeneity from multiple testing. Pre-specified heterogeneity analyses and correction methods are needed.
Related
Sources