Causal Inference: Natural & Quasi-Experiments
The first group had randomization doing the identifying work, and the second replaced it with an assumption nothing can test. This group takes a third route: find a place where the world already ran something close to an experiment, and borrow the variation it produced.
The instruments are circumstantial rather than designed — growing up near a college, being born just before a cutoff date, living on the other side of a state line when a law changed. Nobody assigned them for research purposes, which is exactly why they are credible and exactly why they are fragile. Each design converts a claim about the world into a claim about identification, and each has a characteristic failure: an instrument that barely moves the treatment, a threshold people can manipulate, a parallel trend that was never parallel, a comparison unit that was not comparable.
So the group is built in pairs. Every design is followed by the diagnostic literature that grew up to police it — weak-instrument inference after IV, validity and falsification tests after regression discontinuity, honest confidence intervals after difference-in-differences. The methods are the interesting half; the ways they fail are the useful half.
What each design reports, and what the literature that polices it returns
Every value is committed notebook output. Unlike the previous group there is no answer key here — no randomized benchmark to grade against — so the check on each design is the diagnostic literature that grew up around it, which is why the group is built in pairs.
A is the case for the group, and it is a strong one. Each of these five designs takes a setting where no experiment was run and extracts a credible causal estimate anyway, from variation the world supplied for its own reasons. Card’s instrument turns campus geography into an experiment on schooling; Lee’s threshold turns a coin-flip election into an experiment on incumbency; Proposition 99 turns one state’s policy into a comparative case study. These are genuine achievements, and the headline column is where the field’s confidence in them comes from.
The right-hand column is why the group is not five examples but ten. Every design here has a second literature attached to it, and in each pair that literature was written because practitioners were getting the first column right and the conclusion wrong. The pattern is consistent enough to be worth stating plainly: the diagnostic almost never overturns the estimate, and it very often overturns the confidence attached to it.
B is that pattern isolated. Each row holds its point estimate fixed and changes only the inference procedure, to whichever one the setting actually licenses. Card’s F of 13.3 clears the textbook threshold of 10 and falls an order of magnitude short of the 104.7 that conventional t-inference actually requires, so the tF interval covers zero. Proposition 99 has one treated unit, which rules out the jackknife standard error and leaves a placebo estimator four times larger with an interval spanning zero. The beer-tax coefficient keeps its sign under every specification while its p-value travels from 0.024 to 0.144. In none of these cases is the estimate wrong. In all of them the published-style summary would be.
There is a structural reason this keeps happening. Identification borrowed from the world is narrow — it rests on the compliers, or on the neighbourhood of a cutoff, or on a single treated unit — and narrow identification means a small effective sample no matter how many rows the dataset has. Card has 3,010 men and the instrument moves 12.2% of them. Lee has 6,558 races and the CCT bandwidth keeps those within 13.6 percentage points of a tie. Proposition 99 has 31 years of data and one treated state, which caps the attainable permutation p at 0.026 before any data is examined. The precision the full sample size suggests was never available.
C is the distinction worth carrying out of the group, because the three failures call for three different responses. A biased estimator is a mistake and has a fix — use Callaway–Sant’Anna instead of twoway fixed effects, system GMM instead of difference GMM. An unpinned magnitude is not a mistake: the design is sound and the honest report is an interval, which mainly requires not writing the point estimate as though the third decimal meant something. Unavailable inference is neither — it is the data declining to answer a question that was asked of it, and the correct response is to say so.
Which makes the ordering of the pairs the argument of the group. The design comes first because it works; the diagnostic comes second because working is not the same as being safe to summarise. A borrowed natural experiment is credible precisely to the extent that nobody arranged it, and fragile for the same reason: nobody arranged it to have enough power, either.
How the ten examples relate
Five designs, each immediately followed by the literature that polices it. The pairs are independent of one another and can be read in any order, but each pair should be read in its own order.
Borrow an instrument
Something shifts treatment without touching the outcome any other way.
1 · Instrumental variablesCard’s proximity instrument — and the estimate speaks for the 12.2% it moved 2 · Weak-instrument-robust inferencethe folklore threshold is 10; valid inference needs 104.7Borrow a threshold
A rule assigns treatment at a cutoff nobody near it can control.
3 · Regression discontinuitythe incumbency advantage as a jump — and a bandwidth curve spanning 41% of it 4 · Validity & falsificationwhich checks certify the design and which certify the estimateBorrow repetition
The same units observed again, so each can be its own control.
5 · Panel data & fixed effectsdifference away what is fixed within a unit, measured or not 6 · Dynamic panelsa lagged outcome breaks both workhorses, in opposite directionsBorrow a policy date
A law changed somewhere and not somewhere else.
7 · Difference-in-differencesthe canonical 2×2, then the staggered trap that biases the obvious generalisation 8 · Honest DiDprice the untestable assumption instead of eyeballing a pre-trendBorrow a comparison region
One treated unit, and a control that has to be built rather than found.
9 · Synthetic controla weighted blend of donor states, with permutation for inference 10 · Synthetic DiDboth sets of weights — and the standard error that does not applyWhere it goes next
What none of these designs delivers.
The rest of the arcgraphical models, heterogeneous effects and causal survival analysisOne thread runs the length of the group. Each design earns its causal reading by giving something up, and the thing given up is always the same thing: the number of independent comparisons the data actually contains. An instrument identifies the effect on whoever it moved. A threshold identifies it for whoever was near. A policy date identifies it for whoever was treated, against whoever was not yet. A single treated region identifies it against a donor pool small enough to enumerate. That is the trade the group is named for — and reading the design pages without their diagnostic partners is precisely how the trade goes unnoticed.
Instrumental Variables — the Return to Schooling
Card’s proximity-to-college instrument on 3,010 men, where the confounder — ability — is precisely the thing no covariate captures. Two-stage least squares — 2SLS — puts the return to schooling at 13.2% against OLS’s 7.5%, the opposite direction from what ability bias predicts. LATE — the local average treatment effect — resolves it, and the resolution is smaller than it first looks: the compliers are 12.2% of the sample — roughly one man in eight, whose college decision turned on campus distance — against 42% who went regardless and 46% who never went. The estimate speaks for that eighth alone. A weak-instrument simulation then shows the characteristic failure: as the first-stage F falls below 10 the estimate collapses back toward the confounded OLS value. A weak instrument is worse than none, because it returns a number with a standard error that looks like an answer.
View example →Weak-Instrument-Robust Inference — AR Sets and tF
Diagnosing a weak instrument is not surviving it. With thirty weak instruments — the Angrist–Krueger structure — 2SLS is dragged from a true 1.00 to 1.84, most of the way to the confounded OLS value, and the nominal 95% Wald interval covers the truth 4% of the time. Anderson–Rubin, which inverts an exact test rather than estimating the parameter, holds its coverage at any instrument strength and reports honestly by going unbounded when the data cannot pin the effect down. Then Lee, McCrary, Moon & Weidner recalibrate the folklore: valid conventional inference needs a first-stage F above 104.7, not 10 — more than an order of magnitude too lenient. Applied back to Card, whose F is 13.3, the tidy significance dissolves: AR just excludes zero and the conservative tF interval includes it. Then the study that started the literature. On Angrist & Krueger’s own 329,509 census records, the celebrated 180-instrument specification has a first-stage F of 2.58; replacing quarter of birth with random draws and rerunning it leaves 197 of 200 placebo runs significant at 5%, with standard errors calibrated to their own sampling variation the entire time.
View example →Regression Discontinuity — the Incumbency Advantage
Lee’s 6,558 House races, where winning by 0.1% rather than losing by 0.1%
turns on weather and turnout noise, so the jump at the threshold is causal. At the CCT bandwidth — the window width
that minimises mean squared error — the incumbency advantage is 6.4 points of next-election vote share,
with the from-scratch fit matching rdrobust exactly. Then the bandwidth curve, which
most RD write-ups omit: across h from 0.04 to 0.40 the estimate runs
0.058 to 0.084 — a spread of 41% of the estimate itself,
and not monotone. The density test finds no manipulation (p = 0.151), which is
consistent with a valid design without proving one. And fuzzy RD turns out not to be
like instrumental variables but to be them: outcome jump over treatment jump
recovers a known effect of 3.0 that the sharp estimate puts at 1.77.
RD Validity & Falsification — What Each Check Certifies
A number is not a result. The standard suite — covariate continuity, density, placebo cutoffs, donut and bandwidth — run on Lee’s elections and, for contrast, on a deliberately manipulated design where sorters pile in just above the threshold. The tests have teeth: the covariate jump is −0.150 (p = 0.256) in the valid design and +1.099 (p < 0.001) in the manipulated one, and of five cutoffs only the true one yields an effect. Then the checks stop agreeing, which is the point. The effect is positive and significant at every donut radius and bandwidth, but its magnitude runs 0.046 to 0.093 and spans 41% of its own size. The first three checks license the design; the last two describe the estimate. Here the design is credible and the magnitude is uncertain, and treating a passed falsification test as certifying the point estimate is how a paper claims more robustness than it has.
View example →Panel Data & Fixed Effects — Differencing Confounders Away
Repeated observations let you difference away whatever is fixed within a unit, measured or not.
On Stock & Watson’s 48-state fatality panel, pooling says higher beer taxes go with
more deaths (+0.365); the within estimator flips it to
−0.656, matching PanelOLS exactly. Then two results that
belong beside the headline rather than behind it. First differences give +0.029
— the opposite sign and indistinguishable from zero — and since FD and FE are both
consistent under the same assumption, that disagreement is diagnostic, not a second opinion.
And adding year effects, the more defensible specification, leaves the coefficient almost
unchanged while p moves from 0.024 to 0.095, and to 0.144 with
controls: the standard error grows because year effects absorb the variation the tax was
competing to explain. A negative sign, then, whose magnitude is not pinned down.
Dynamic Panels — Nickell Bias, GMM and the Bond Bracket
How much of a firm’s employment carries over from one year to the next? Put that lagged outcome on the right-hand side and both panel workhorses break, in opposite directions: on Arellano & Bond’s own 140-firm panel, pooled OLS puts the carry-over at 0.932 — a shock taking decades to fade — and fixed effects at 0.514, where half of it washes out each year. They bracket a truth neither can reach. Difference GMM, which instruments the lag with earlier levels, then overshoots below fixed effects at 0.336 — an estimator built to correct a downward bias landing past the estimate it was correcting, which is the signature of lagged levels going weak on a persistent series. System GMM restores 0.635. Then the diagnostics split. The serial-correlation test is clean and the instrument count is a healthy 42 against 140 firms, but the over-identification test, which asks whether the surplus instruments agree with one another, rejects at p = 0.003. It is a test that fails and cannot be trusted to have failed for the right reason. The restrictions are unconfirmed, not confirmed.
View example →Difference-in-Differences — Card–Krueger and the Staggered Trap
The canonical 2×2 on 410 fast-food stores: New Jersey’s minimum wage rose, wages followed, and employment did not fall relative to Pennsylvania — +2.75, reproduced exactly by a regression interaction. Then timing varies, and the same regression stops estimating what it appears to. On a simulation with a known average effect on the treated of 3.580, twoway fixed effects returns 2.677. Goodman-Bacon locates the loss: comparisons against never-treated units average 3.407 and against later-treated 2.756, but those using an already-treated control average 0.843 and carry 27% of the weight — a control group whose own effect is still growing makes the treated look flat. Every ingredient is a legitimate DiD; the blend is biased. Callaway–Sant’Anna, using only clean controls, recovers 3.553.
View example →Honest DiD — How Big a Violation Would Overturn It
Parallel trends is untestable, and the eyeball test everyone uses cannot do the job. Two simulated event studies have pre-trend statistics of 0.30 and 0.39 — indistinguishable, and both near the 0.14 of sampling noise — while their underlying bias differs tenfold. Because the truth is known, the failure is exact: the coefficient estimates τ + g rather than τ, so the naive estimate is overstated by precisely the per-period drift — 10% in one case, 100% in the other, where the reported effect is more than double the truth. Rambachan & Roth price the assumption instead of asserting it: the breakdown value M̄* separates them at 2.5× against 1.25× the pre-period worst. An estimate that breaks below 1 needs the trend to politely halt at the treatment date.
View example →Synthetic Control — Proposition 99
One treated unit and no obvious comparison, so build the control instead of finding it. A non-negative, sum-to-one blend of six states — Utah at 0.394, then Montana, Nevada, Connecticut, New Hampshire, Colorado — tracks pre-1988 California to a root-mean-square fit error of 1.66 packs across eighteen years, and the post-1988 gap widens to −26.6 packs by 2000. With a single treated unit there is no standard error, so inference is a placebo permutation over all 38 donors — and California ranks 3rd of 39, p = 0.077, which does not clear 5%. Suggestive rather than decisive, and structurally so: with 38 placebos the smallest attainable p is 0.026. The effect size is robust across implementations; the inference depends on how the donors were matched.
View example →Synthetic Difference-in-Differences — Unit and Time Weights
The synthesis of the pair before it: unit weights like synthetic control but with a level-shifting intercept, so it matches California’s trend rather than its level, plus time weights favouring the late-1980s years most predictive of what came after. The three estimators separate cleanly — DiD −27.35, synthetic control −19.51, SDID −15.60 — and the from-scratch figure reproduces the published −15.6. Then the inference correction. The jackknife SE of 2.37 usually quoted requires several treated units; Prop 99 has one. The placebo estimator the authors prescribe gives 9.49, four times larger, and its interval includes zero (p = 0.051). The point estimate is solid; its distinguishability from zero is not established by this design.
View example →