Heterogeneous Treatment Effects

causal-inferencetreatment-effectsheterogeneous-treatment-effectspropensity-scoreLATEATEMTEnegative-selectionessential-heterogeneitysocial-stratificationmultiple-outcomestreatment-risk-paradoxexternal-validitysimulation

Definition

Causal effects are heterogeneous when the treatment effect Y1Y0Y_1 - Y_0 varies across individuals — so that no single average adequately summarizes the distribution. Heterogeneous treatment effects (HTE) is the recognition that most treatments affect different people differently, and the research agenda of characterizing how and for whom effects vary. The central question is whether the average effect estimated by a particular design (ATE, ATT, LATE, MTE) generalizes to the population of interest, or whether it masks opposing effects that partially cancel in the average.

Key Ideas

Treatment Effect Parameters Under Heterogeneity

Parameter Definition Identifies Effect For
ATE E[Y1Y0]E[Y_1 - Y_0] Full population
ATT E[Y1Y0D=1]E[Y_1 - Y_0 \mid D=1] The treated
ATC E[Y1Y0D=0]E[Y_1 - Y_0 \mid D=0] The untreated
LATE E[Y1Y0complier]E[Y_1 - Y_0 \mid \text{complier}] Compliers defined by a specific instrument
MTE E[Y1Y0V=u]E[Y_1 - Y_0 \mid V=u] Individuals at the margin of indifference at propensity threshold uu
CATE E[Y1Y0X=x]E[Y_1 - Y_0 \mid X=x] Subgroup with observed characteristics X=xX=x

When treatment effects are homogeneous, all five coincide. In practice they rarely do, and the difference matters for policy.

Propensity-Based Heterogeneity

A key approach, developed in Brand and colleagues' work, stratifies the sample by estimated propensity score p^(X)\hat{p}(X) and estimates conditional average treatment effects (CATE) within strata. This reveals how effects vary as a function of how likely individuals are to receive the treatment, going beyond simple subgroup breakdowns.

Negative selection (Brand and Xie 2010): If treatment effects are negatively correlated with the propensity to receive the treatment, those least likely to be treated benefit most from it. Brand and Xie (2010) document this for college wage returns across the National Longitudinal Survey of Youth 1979 (NLSY79) and the Wisconsin Longitudinal Study — 10 consistent negative Level-2 slopes across two datasets, two sexes, and multiple career stages. For NLSY men, the estimated wage return ranges from ~30% in the lowest propensity stratum to ~10% in the highest — a 20 pp gap. The mechanism is counterfactual deprivation: the negative pattern arises not because low-propensity graduates earn more, but because low-propensity non-graduates earn very little without a degree; high-propensity individuals can fall back on their superior resources and abilities even without college. This directly contradicts the common assumption of positive selection (that the most capable, most likely to attend, also benefit most). Under positive selection, ordinary least squares (OLS) overestimates effects on the untreated; under negative selection, OLS underestimates the benefit of expanding treatment to reluctant participants — ATT < ATE < ATC.

Negative selection in fertility effects (Brand and Davis 2011): The same propensity-score hierarchical linear model (HLM) framework extended to fertility outcomes. Women least likely to attend college experience the largest fertility-decreasing effects of college. For attendance, the Level-2 slope = +0.10 (p<0.05); stratum 1 women have 65% fewer children than comparable non-attenders. For college completion, the Level-2 slope = +0.17 (p<0.01), and the effect fully reverses for high-propensity women (+42% more children by age 41) — advantaged women use college as a complement to, rather than substitute for, family formation. Unlike earnings, however, the normative valence of the treatment effect is ambiguous: lower fertility for disadvantaged women may represent escape from early disadvantaged family formation, but also forecloses desired childbearing.

Multi-Outcome Sorting: "Sorting on the Mix" (Brooks, Chapman, and Schroeder 2018)

Prior discussions of essential heterogeneity assumed agents sort on expected gains from a single outcome of interest, with treatment effects on other outcomes (e.g. costs, adverse events) either absent or uncorrelated. In real-world clinical and policy settings, treatment choices reflect a weighted assessment of expected effects across multiple beneficial and detrimental outcomes. Brooks, Chapman, and Schroeder (2018) label this "sorting on the mix" and use simulation (5 scenarios × 1000 runs × 5000 patients) to derive its implications:

Essential Heterogeneity (Heckman-Vytlacil)

Essential heterogeneity exists when individuals select into treatment partly because they privately know their own above-average returns: Cov[Y1Y0,D]>0\text{Cov}[Y_1 - Y_0, D] > 0. Under essential heterogeneity:

The Marginal Treatment Effect (MTE) framework (Heckman and Vytlacil 1999, 2005) makes essential heterogeneity precise and recovers the full distribution of effects as a function of the selection index. See Marginal Treatment Effect.

Observable vs. Unobservable Heterogeneity

Propensity-score stratification (Brand and Simon-Thomas 2012) works under selection-on-observables; MTE methods address unobservable selection on gains.

Why It Matters

Open Questions

Related

Sources