Natural Experiments

methodscausal-inferencenatural-experimentsquasi-experimentsinternal-validityexternal-validitypolicy-evaluationidentificationvalidity-threatsarea-fixed-effectsextrapolation

Definition

A natural experiment is a research design that exploits exogenous, quasi-random variation in treatment generated by policy changes, government rules, institutional events, or other forces outside the researcher's control. Unlike a true randomized controlled trial, the researcher does not assign treatment; an identifiable external event does so in a way that is plausibly independent of potential outcomes. The goal is to find variation in a key explanatory variable whose determinants are transparent — clearly understood and scrutinizable for exogeneity. "Quasi-experiment" (the term preferred in psychology) is more precise: such studies are not quite experiments. The term "natural experiment" is the dominant usage in economics but misleadingly implies spontaneity and experimental design.

Key Ideas

Area Fixed-Effects Model (Moffitt 2005)

The area fixed-effects model formalizes ecological IV for panel or repeated cross-section data. Write the structural system in first differences:

ΔYi=α+βiΔTi+γΔXi+εi\Delta Y_i = \alpha + \beta_i \Delta T_i + \gamma \Delta X_i + \varepsilon_i ΔTi=δ+θΔXi+ϕΔZi+υi\Delta T_i = \delta + \theta \Delta X_i + \phi \Delta Z_i + \upsilon_i

where Δ\Delta denotes change from Time 1 to Time 2, and ΔZi\Delta Z_i is the change in the area-level environmental variable (a new law, a policy expansion, a price change). Area fixed effects that would be correlated with TT in levels cancel out in differencing. ΔZ\Delta Z then instruments for ΔT\Delta T — individual changes in TT are still presumed endogenous, so differencing alone is not sufficient.

This is not the individual fixed-effects (FE) model. The individual FE model estimates only ΔY=α+βiΔT+γΔX+ε\Delta Y = \alpha + \beta_i \Delta T + \gamma \Delta X + \varepsilon (without an instrument for ΔT\Delta T), assuming that differencing eliminates all bias. Moffitt (2005) argues this assumption "leaves unspecified why individual changes in TT occur" — those changes may be endogenous. The area FE model ties changes in TT to a specific, measurable ΔZ\Delta Z.

The model can be run on repeated cross-sections (not just panels): aggregate up to area means, or use individual data with area and time dummies and the area ×\times time interaction of ZiZ_i.

Objections: (1) Residential sorting — location may not be exogenous. (2) Time-varying unobservables correlated with ΔZ\Delta Z and ΔY\Delta Y simultaneously. (3) Lagged adjustment — the model assumes immediate and permanent response, but norms and behavior may adjust slowly. (4) Mechanism distance — the further ZZ is from individual TT, the more spurious mechanisms can generate a spurious ZYZ \to Y association.

Population-Segment Fixed-Effects (Moffitt 2005)

A related design in which nationwide policies affect demographic groups differently (e.g., low-income single mothers vs. married women). Any difference in outcome trends between the affected and unaffected groups is attributed to differential policy exposure. Mathematically equivalent to the area FE model with demographic groups replacing geographic areas.

Moffitt argues this design has weaker support: the assumption that different demographic groups would evolve in parallel absent the policy is "much more suspect" than the comparable geographic assumption. Groups defined by marital status, income, or family structure differ on many time-varying observables and unobservables. See Difference-in-Differences.

Two Types of Extrapolation Failure (Moffitt 2005)

Any instrumental variable, however valid, faces two distinct limits on generalizability:

  1. Mechanism specificity: Each ZZ represents one specific cause of variation in TT. Whether postponing childbearing via abortion policy access has the same effect as postponing it via a labor market shock is an empirical question the framework cannot answer from within. "The effect of TT" is an ill-posed question without specifying the mechanism inducing TT to change. The structural model assumes βi\beta_i is independent of the particular ZiZ_i that moved TT — but this is an assertion, not a testable property of the data.

  2. Range restriction: ZZ induces variation in TT only across a particular range (e.g., from 30%30\% to 40%40\% smokers). Extrapolation to the rest of the population requires additional assumptions. Instruments inducing larger first-stage variation are preferred for extrapolation but are often harder to defend for internal validity. This is the sharpest formulation of why LATE estimates from examiner IV (Maestas et al. 2013) — which move allowance rates by 10\approx 10 percentage points among marginal applicants — cannot be directly applied to the full applicant pool or to policy interventions of different magnitudes.

Campbell's Threats to Validity

Adapted from Donald Campbell's experimental design literature for economics (Meyer 1995).

Internal Validity (9 threats)

Internal validity: whether the estimated difference was caused by the treatment within the study context.

  1. Omitted variables — events concurrent with the treatment providing alternative explanations for the outcome change
  2. Trends in outcomes — time trends independent of treatment (inflation, secular aging, wage growth)
  3. Misspecified variances — understated standard errors from omitted group error terms; outcomes for units within a group are correlated (clustering)
  4. Mismeasurement — changes in survey methods, question wording, or definitions over the study period
  5. Political economy — endogenous policy changes driven by past or expected future outcomes
  6. Simultaneity — joint determination of treatment and outcome
  7. Selection — nonrandom assignment correlated with potential outcomes; Ashenfelter's dip (earnings dip precedes training program entry)
  8. Attrition — differential loss of respondents from treatment and comparison groups
  9. Omitted interactions — differential trends or omitted variables that affect treatment and comparison groups differently; this is precisely what the parallel trends assumption rules out in Difference-in-Differences

External Validity (3 threats)

External validity: whether effects generalize to other individuals, settings, and time periods.

  1. Interaction of selection and treatment — the treatment group is unrepresentative; the estimated effect applies only to that group at that margin
  2. Interaction of setting and treatment — effects vary across geographic or institutional contexts
  3. Interaction of history and treatment — effects vary across time periods; temporary vs. permanent policy changes may have different effects as institutions adapt

Research Design Hierarchy

  1. One-group before/after (simple differences): Weak. Requires all variation in outcomes attributable to the treatment. Useful as preliminary analysis or when a strong single-group design is available.
  2. Before/after with untreated comparison group (Difference-in-Differences): The dominant design. Differences out common time trends and time-invariant group differences. Key assumption: no omitted interaction between group membership and time period.
  3. Triple-differences (higher-order interactions): Treatment defined as a three-way interaction (state × demographic group × post-period). Removes state×time and state×group trends. Examples: Gruber (1994) maternity mandates; Yelowitz (1994) Medicaid expansions.

Open Questions

Related

Sources