Summary
Heckman surveys the econometric approach to causality for a statistical audience, distinguishing it from the Neyman-Rubin (NR) statistical treatment-effects framework. The paper argues that the two traditions share the potential-outcomes starting point but diverge on whether to model selection explicitly (econometric approach: yes, via a Roy-model selection equation) or to finesse selection through randomization or instrumental variables (statistical approach). The central organizing concept is the Marginal Treatment Effect (MTE), which Heckman shows unifies ordinary least squares (OLS), instrumental variables (IV), matching, and difference-in-differences (DiD) as different weighted averages of individual treatment effects. The paper introduces a P1/P2/P3 policy hierarchy to clarify what each estimation approach can and cannot answer.
Key Claims
- Three-task distinction: Any causal analysis must separately address (a) defining the counterfactual — what would have happened to the treated under the alternative?; (b) identifying the effect from hypothetical ideal data; and (c) identifying the effect from real, available data. Holland (1986) conflates (a) and (b); the local average treatment effect (LATE) conflates definition and identification.
- P1/P2/P3 hierarchy: P1 = evaluating a historical program for the treated population (internal validity); P2 = forecasting effects in new environments not yet experienced (external validity); P3 = evaluating new policies never experienced anywhere. The statistical treatment-effects framework answers P1 well, is inadequate for P2, and cannot address P3 — which requires a structural model.
- Neyman-Rubin postulates (R-1 to R-4): The NR model (i) postulates counterfactuals without modelling the selection mechanism (R-1); (ii) considers only ex post objective outcomes, not subjective evaluations (R-2); (iii) assumes the Stable Unit Treatment Value Assumption (SUTVA) (R-3); and (iv) is recursive/triangular — no simultaneity (R-4). The econometric model relaxes R-1 by explicitly modelling selection, R-2 by allowing subjective valuations, and R-4 via simultaneous equations.
- Fixing vs. conditioning (Haavelmo 1943): The structural equation E(Y∣do(X=x),U=u)=xβ+u is the causal estimand; E(Y∣X=x)=xβ+E(U∣X=x) is not causal when X is endogenous (E[U∣X]=0). This distinction between "conditioning on" (passive observation) and "fixing/intervening on" (active manipulation of the data-generating process, DGP) is the formal foundation of Pearl's "do" operator and of structural identification.
- Marginal Treatment Effect (MTE): Introduced by Björklund and Moffitt (1987) and developed by Heckman and Vytlacil (1999, 2005), the MTE is the average treatment effect for individuals at the margin of indifference given a particular value of the propensity score. OLS, IV, matching, and DiD all identify different weighted averages of MTE — with different instrument-determined weights that depend on the propensity score distribution. This unifying framework reveals that LATE is a weighted average of MTE at specific propensity-score thresholds.
- Marschak's Maxim: Marschak (1953) argued that economists should use the minimum model needed to answer the policy question. Statistical treatment effects implement this maxim for P1 questions (evaluating a past program): no structural model is needed because the policy variable itself is the treatment. But P2 and P3 questions require identifying structural parameters that remain invariant across environments — which the treatment-effects framework cannot provide without additional assumptions.
- Simultaneous causality: The Neyman-Rubin framework is inherently recursive (triangular): treatment is exogenous or instrumented, outcomes are downstream. Haavelmo (1943, 1944) defined causal effects in simultaneous systems via exclusion restrictions on the structural coefficient matrix. Simultaneous causality — where two variables jointly determine each other — requires the full simultaneous equations apparatus; the LATE framework has no analog for this case.
Concepts Introduced or Extended
Entities Mentioned
Quotes
"The problem of causal inference is distinct from the problem of statistical inference. Causal inference requires a model; statistical inference does not."
"The LATE parameter identifies the average treatment effect for persons who are induced to switch treatment status by an instrument. It is a well-defined causal parameter. But it does not answer the question asked by program evaluators, which is the effect of the program on the people in the program."
My Take
The paper's main intellectual contribution is clarifying what economists and statisticians mean — and what they don't mean — when they say "treatment effect," and showing that the two traditions have largely been talking past each other. The three-task distinction is genuinely clarifying: Holland's "no causation without manipulation" conflates (a) and (b) because it rules out defining counterfactuals for attributes that cannot be manipulated, but that is a definitional choice, not a statistical theorem. The MTE synthesis is a real advance: it turns the Imbens-Angrist/Heckman debate from a dispute about which estimator is "right" into a question about which weighted average of MTE a researcher's design identifies — and whether those weights correspond to a policy-relevant parameter. The P1/P2/P3 hierarchy is useful but underspecified: Heckman is right that structural models are needed for P3, but he does not fully confront the difficulty of validating the invariance assumptions structural models require. The critique of LATE's policy relevance is well-taken: the complier population is instrument-specific, unobservable, and often not the population of policy interest.