Summary
Meyer surveys the strengths and weaknesses of natural experiment designs in economics, describing how researchers exploit policy changes, government lotteries, and other exogenous events to identify causal effects when randomization is unavailable. The paper advocates for more elaborate designs — multiple comparison groups, multiple pre- and post-intervention time periods, triple-differences — as ways to probe the comparability of treatment and control groups and increase confidence in causal claims. The core lesson: if variation cannot be experimentally controlled, its source must be transparently understood and scrutinized.
Key Claims
- Primary lesson: Understanding the source of variation is the central methodological contribution of the natural experiment literature. Variation whose determinants are unknown or unmodeled cannot support valid causal inference, regardless of what it is called.
- Nomenclature: "Quasi-experiment" (from psychology) is more precise than "natural experiment" (dominant in economics). The latter inappropriately implies spontaneity and experimental design; the former emphasizes that such studies are not quite experiments. Conventional studies without a modeled source of variation correspond to psychology's "correlational designs."
- Research design hierarchy: (1) one-group before/after (simple differences, weak); (2) before/after with untreated comparison group (difference-in-differences [DiD], the dominant design); (3) treatments defined by higher-order interactions (triple-differences, removes additional group×time trends).
- DiD mechanics: β^=(Yˉ11−Yˉ01)−(Yˉ10−Yˉ00); the key identifying assumption is no omitted interaction between treatment-group membership and the post-treatment period (parallel trends).
- Triple-differences: Gruber (1994) — women of certain ages in certain states after maternity mandate; Yelowitz (1994) — mothers with children of certain ages in certain states after Medicaid expansion. Treatment is the 3-way interaction; all main effects and first-order interactions are included as controls.
- Probing comparability: Multiple comparison groups test overidentifying restrictions (Rosenbaum's "control by systematic variation" recast as an economist's overidentification test). Multiple pre/post periods detect seasonality, pre-existing differential trends, and misspecified variances.
- Political economy threat: Not every law change is a good natural experiment. Changes driven by past or anticipated future outcomes create spurious correlations; remedy is understanding political determinants and applying Granger/Sims exogeneity tests (Cook and Tauchen 1982).
- Instrumental variables (IV) for imprecise assignment: When a natural event shifts treatment probabilities rather than deterministically assigning treatment, the event serves as an instrument. Vietnam draft lottery → veteran status (Angrist 1990); quarter of birth → schooling (Angrist and Krueger 1991); policy dummy → continuous benefit level (two-stage least squares [2SLS] where d∗ is the first-stage instrument for benefit amount b).
- External validity limitation: Natural experiments identify effects for specific groups over specific ranges of the explanatory variable — a local average treatment effect (LATE)-like parameter for compliers. Conventional studies assume homogeneous effects but rarely identify which source of variation is influential; the difference is transparency, not necessarily better external validity.
Concepts Introduced or Extended
Entities Mentioned
Quotes
"The natural-experiment approach emphasizes the general issue of understanding the sources of variation used to estimate the key parameters. In my view, this is the main lesson of these studies. If one cannot experimentally control the variation one is using, one should understand its source."
"The term quasi-experiments emphasizes that such studies are not quite experiments. The term natural experiments, which is more commonly used in economics, somewhat inappropriately suggests that these studies are experiments and moreover that they are spontaneous."
"Of course, calling a source of variation a natural experiment does not make that variation exogenous."
My Take
Meyer's framework remains the clearest synthetic treatment of quasi-experimental design in economics. The parallel with Campbell's threats to validity grounds econometrics in the broader experimental design tradition and gives practitioners a structured vocabulary for critiquing identification. The paper predates the Angrist, Imbens, and Rubin (AIR 1996) LATE framework, so it lacks a formal treatment of what parameter natural experiments identify — the "local average treatment effect" logic is implicit in Section 9 but not formalized. The advocacy for multiple comparison groups (as overidentification tests) and multiple pre-periods (as pre-trend diagnostics) anticipated what are now standard practices in applied DiD work. The warning about political economy endogeneity of policy changes, illustrated with Granger tests, remains underemphasized in the applied literature.