Rubin Causal Model

methodscausal-inferencepotential-outcomesstatisticseconometricsstructural-modelscounterfactuals

Definition

The Rubin Causal Model (RCM) is a framework for causal inference built on potential outcomes: for every unit ii and every possible treatment value dd, there exists a potential outcome Yi(d)Y_i(d) representing what would be observed if unit ii received treatment dd. The causal effect of treatment on unit ii is the comparison Yi(1)Yi(0)Y_i(1) - Y_i(0). Because only one potential outcome is ever realized — the one corresponding to the treatment actually received — individual causal effects are unobservable. This "fundamental problem of causal inference" makes statistical assumptions unavoidable for identifying average causal effects.

Key Ideas

How It Works

  1. Define potential outcomes Yi(0),Yi(1)Y_i(0), Y_i(1) for each unit before observing any data.
  2. Choose a treatment assignment mechanism (random, observational, or instrument-based).
  3. State assumptions sufficient to identify the desired estimand (ATE, ATT, LATE) — typically some form of ignorability, exclusion, or monotonicity.
  4. Estimate the estimand from observed (Yi,Di,Zi,Xi)(Y_i, D_i, Z_i, X_i) using the assumptions.
  5. Perform sensitivity analysis to violations of key assumptions.

Why It Matters

The RCM provides a unified language for causal inference across statistics, economics, epidemiology, and social science. Before the RCM, causal claims from observational data were made informally or through structural equation models whose assumptions were opaque ("the disturbances are uncorrelated with the instrument"). The RCM makes every identifying assumption a statement about observable or potentially observable quantities — e.g., "holding DD fixed, the instrument ZZ has no direct effect on YY" (exclusion restriction) — which is substantively evaluable even if not directly testable.

Heckman's Critique: R-1 to R-4 (Heckman 2008)

Heckman characterizes the Neyman-Rubin (NR) model by four postulates that distinguish it from the econometric structural approach:

Three-Task Distinction (Heckman 2008)

Any causal analysis implicitly involves three separate tasks that must not be conflated:

  1. Define the counterfactual: What would have happened to the treated units under the alternative? Holland's (1986) "no causation without manipulation" conflates this with identification — ruling out counterfactuals for attributes that cannot be manipulated. But defining counterfactuals and identifying them from data are logically separate steps.
  2. Identify from ideal data: Given complete data on potential outcomes (e.g., a hypothetical randomized experiment), can the parameter of interest be computed? This is a mathematical question about estimands.
  3. Identify from real data: Given the actual data available (observational, with selection), can the parameter be estimated? LATE conflates (2) and (3): it is simultaneously defined as the estimand and identified by the Wald ratio, so it is often unclear whether a researcher is making a definitional choice or an identification claim.

Open Questions

Related

Sources