← All examples

Causal Inference: Natural & Quasi-Experiments

The first group had randomization doing the identifying work, and the second replaced it with an assumption nothing can test. This group takes a third route: find a place where the world already ran something close to an experiment, and borrow the variation it produced.

The instruments are circumstantial rather than designed — growing up near a college, being born just before a cutoff date, living on the other side of a state line when a law changed. Nobody assigned them for research purposes, which is exactly why they are credible and exactly why they are fragile. Each design converts a claim about the world into a claim about identification, and each has a characteristic failure: an instrument that barely moves the treatment, a threshold people can manipulate, a parallel trend that was never parallel, a comparison unit that was not comparable.

So the group is built in pairs. Every design is followed by the diagnostic literature that grew up to police it — weak-instrument inference after IV, validity and falsification tests after regression discontinuity, honest confidence intervals after difference-in-differences. The methods are the interesting half; the ways they fail are the useful half.

What each design reports, and what the literature that polices it returns

Every value is committed notebook output. Unlike the previous group there is no answer key here — no randomized benchmark to grade against — so the check on each design is the diagnostic literature that grew up around it, which is why the group is built in pairs.

A · Five designs, each followed by the literature that grew up to police it Left: what the design reports. Right: what its paired diagnostic returns on the same data. Red = the headline claim does not survive; amber = it survives in part. design headline result paired diagnostic and what it returns Instrumental variables Card’s proximity instrument, 3,010 men 2SLS 13.2% vs OLS 7.5% first-stage F = 13.3, clears the F > 10 rule Weak-instrument-robust inference valid inference needs F > 104.7; the tF interval includes zero Regression discontinuity Lee’s 6,558 House races incumbency advantage 0.0637 robust CI [0.0348, 0.0839] at the CCT bandwidth Validity & falsification design passes every check; estimate runs 0.046 to 0.093 across them Panel fixed effects 48 states × 7 years of fatalities beer tax −0.656 sign flips from the pooled +0.365 Dynamic panels lagged outcome breaks it; Sargan rejects the restrictions at p = 0.003 Difference-in-differences Card & Krueger, 410 stores +2.75 on the canonical 2×2 staggered twoway FE returns 2.677 against a true 3.580 Honest DiD pre-trends of 0.30 and 0.39 hide bias differing tenfold Synthetic control Proposition 99, 1 treated unit, 38 donors −26.6 packs by 2000 California ranks 3rd of 39, p = 0.077 Synthetic DiD SDID reproduces −15.6; the applicable SE is 9.49, not 2.37 B · The group’s recurring result: the significance is the part that does not survive Each row keeps its point estimate and changes only the inference procedure to the one the setting actually licenses. result as conventionally reported under the applicable procedure Card’s return to schooling F = 13.3 passes the folklore threshold t > 1.96, significant tF interval includes zero Beer taxes and traffic deaths adding year effects, then controls p = 0.024 p = 0.095, then 0.144 Proposition 99, synthetic control 38 placebos, so p = 0.026 is the floor large and widening gap p = 0.077, does not clear 5% Proposition 99, synthetic DiD one treated unit, so the jackknife does not apply SE 2.37, CI [−20.25, −10.96] SE 9.49, CI [−34.21, +3.01] Arellano & Bond persistence AR(2) clean, 42 instruments for 140 firms system GMM 0.635, in bracket Sargan rejects at p = 0.003 C · The failures are not all the same failure, and the distinction changes what to do All ten examples sort into three kinds. Only the first is a broken estimator; the other two are honest estimators reporting less than the write-up claims. The estimator is biased by construction It answers a different question than the one asked, and more data does not help. · pooled OLS on a panel: +0.365 against the within estimate’s −0.656 · twoway fixed effects on staggered adoption: 2.677 against a known 3.580, with forbidden comparisons at 0.843 carrying 27% of the weight · difference GMM on a persistent series: 0.336, below the fixed-effects estimate it was built to correct The design holds; the magnitude does not The causal reading is earned. The number attached to it is a range, and the write-up usually reports a point. · RD across bandwidths: 0.058 to 0.084, a spread of 41% of the estimate itself, and not monotone · RD across the full falsification suite: 0.046 to 0.093, significant throughout · beer taxes: negative under every specification, with p from 0.024 to 0.144 The estimate holds; the inference is unavailable The point estimate reproduces and is credible. The data cannot say how far it is from zero. · Card with F = 13.3: AR just excludes zero, tF includes it · Proposition 99 with 38 donors: p = 0.077, and 0.026 is the smallest value attainable · synthetic DiD with one treated unit: placebo SE four times the jackknife, interval spanning zero

A is the case for the group, and it is a strong one. Each of these five designs takes a setting where no experiment was run and extracts a credible causal estimate anyway, from variation the world supplied for its own reasons. Card’s instrument turns campus geography into an experiment on schooling; Lee’s threshold turns a coin-flip election into an experiment on incumbency; Proposition 99 turns one state’s policy into a comparative case study. These are genuine achievements, and the headline column is where the field’s confidence in them comes from.

The right-hand column is why the group is not five examples but ten. Every design here has a second literature attached to it, and in each pair that literature was written because practitioners were getting the first column right and the conclusion wrong. The pattern is consistent enough to be worth stating plainly: the diagnostic almost never overturns the estimate, and it very often overturns the confidence attached to it.

B is that pattern isolated. Each row holds its point estimate fixed and changes only the inference procedure, to whichever one the setting actually licenses. Card’s F of 13.3 clears the textbook threshold of 10 and falls an order of magnitude short of the 104.7 that conventional t-inference actually requires, so the tF interval covers zero. Proposition 99 has one treated unit, which rules out the jackknife standard error and leaves a placebo estimator four times larger with an interval spanning zero. The beer-tax coefficient keeps its sign under every specification while its p-value travels from 0.024 to 0.144. In none of these cases is the estimate wrong. In all of them the published-style summary would be.

There is a structural reason this keeps happening. Identification borrowed from the world is narrow — it rests on the compliers, or on the neighbourhood of a cutoff, or on a single treated unit — and narrow identification means a small effective sample no matter how many rows the dataset has. Card has 3,010 men and the instrument moves 12.2% of them. Lee has 6,558 races and the CCT bandwidth keeps those within 13.6 percentage points of a tie. Proposition 99 has 31 years of data and one treated state, which caps the attainable permutation p at 0.026 before any data is examined. The precision the full sample size suggests was never available.

C is the distinction worth carrying out of the group, because the three failures call for three different responses. A biased estimator is a mistake and has a fix — use Callaway–Sant’Anna instead of twoway fixed effects, system GMM instead of difference GMM. An unpinned magnitude is not a mistake: the design is sound and the honest report is an interval, which mainly requires not writing the point estimate as though the third decimal meant something. Unavailable inference is neither — it is the data declining to answer a question that was asked of it, and the correct response is to say so.

Which makes the ordering of the pairs the argument of the group. The design comes first because it works; the diagnostic comes second because working is not the same as being safe to summarise. A borrowed natural experiment is credible precisely to the extent that nobody arranged it, and fragile for the same reason: nobody arranged it to have enough power, either.

How the ten examples relate

Five designs, each immediately followed by the literature that polices it. The pairs are independent of one another and can be read in any order, but each pair should be read in its own order.

Where it goes next

What none of these designs delivers.

The rest of the arcgraphical models, heterogeneous effects and causal survival analysis

One thread runs the length of the group. Each design earns its causal reading by giving something up, and the thing given up is always the same thing: the number of independent comparisons the data actually contains. An instrument identifies the effect on whoever it moved. A threshold identifies it for whoever was near. A policy date identifies it for whoever was treated, against whoever was not yet. A single treated region identifies it against a donor pool small enough to enumerate. That is the trade the group is named for — and reading the design pages without their diagnostic partners is precisely how the trade goes unnoticed.

Instrumental Variables — the Return to Schooling

Card’s proximity-to-college instrument on 3,010 men, where the confounder — ability — is precisely the thing no covariate captures. Two-stage least squares — 2SLS — puts the return to schooling at 13.2% against OLS’s 7.5%, the opposite direction from what ability bias predicts. LATE — the local average treatment effect — resolves it, and the resolution is smaller than it first looks: the compliers are 12.2% of the sample — roughly one man in eight, whose college decision turned on campus distance — against 42% who went regardless and 46% who never went. The estimate speaks for that eighth alone. A weak-instrument simulation then shows the characteristic failure: as the first-stage F falls below 10 the estimate collapses back toward the confounded OLS value. A weak instrument is worse than none, because it returns a number with a standard error that looks like an answer.

View example →

Weak-Instrument-Robust Inference — AR Sets and tF

Diagnosing a weak instrument is not surviving it. With thirty weak instruments — the Angrist–Krueger structure — 2SLS is dragged from a true 1.00 to 1.84, most of the way to the confounded OLS value, and the nominal 95% Wald interval covers the truth 4% of the time. Anderson–Rubin, which inverts an exact test rather than estimating the parameter, holds its coverage at any instrument strength and reports honestly by going unbounded when the data cannot pin the effect down. Then Lee, McCrary, Moon & Weidner recalibrate the folklore: valid conventional inference needs a first-stage F above 104.7, not 10 — more than an order of magnitude too lenient. Applied back to Card, whose F is 13.3, the tidy significance dissolves: AR just excludes zero and the conservative tF interval includes it. Then the study that started the literature. On Angrist & Krueger’s own 329,509 census records, the celebrated 180-instrument specification has a first-stage F of 2.58; replacing quarter of birth with random draws and rerunning it leaves 197 of 200 placebo runs significant at 5%, with standard errors calibrated to their own sampling variation the entire time.

View example →

Regression Discontinuity — the Incumbency Advantage

Lee’s 6,558 House races, where winning by 0.1% rather than losing by 0.1% turns on weather and turnout noise, so the jump at the threshold is causal. At the CCT bandwidth — the window width that minimises mean squared error — the incumbency advantage is 6.4 points of next-election vote share, with the from-scratch fit matching rdrobust exactly. Then the bandwidth curve, which most RD write-ups omit: across h from 0.04 to 0.40 the estimate runs 0.058 to 0.084 — a spread of 41% of the estimate itself, and not monotone. The density test finds no manipulation (p = 0.151), which is consistent with a valid design without proving one. And fuzzy RD turns out not to be like instrumental variables but to be them: outcome jump over treatment jump recovers a known effect of 3.0 that the sharp estimate puts at 1.77.

View example →

RD Validity & Falsification — What Each Check Certifies

A number is not a result. The standard suite — covariate continuity, density, placebo cutoffs, donut and bandwidth — run on Lee’s elections and, for contrast, on a deliberately manipulated design where sorters pile in just above the threshold. The tests have teeth: the covariate jump is −0.150 (p = 0.256) in the valid design and +1.099 (p < 0.001) in the manipulated one, and of five cutoffs only the true one yields an effect. Then the checks stop agreeing, which is the point. The effect is positive and significant at every donut radius and bandwidth, but its magnitude runs 0.046 to 0.093 and spans 41% of its own size. The first three checks license the design; the last two describe the estimate. Here the design is credible and the magnitude is uncertain, and treating a passed falsification test as certifying the point estimate is how a paper claims more robustness than it has.

View example →

Panel Data & Fixed Effects — Differencing Confounders Away

Repeated observations let you difference away whatever is fixed within a unit, measured or not. On Stock & Watson’s 48-state fatality panel, pooling says higher beer taxes go with more deaths (+0.365); the within estimator flips it to −0.656, matching PanelOLS exactly. Then two results that belong beside the headline rather than behind it. First differences give +0.029 — the opposite sign and indistinguishable from zero — and since FD and FE are both consistent under the same assumption, that disagreement is diagnostic, not a second opinion. And adding year effects, the more defensible specification, leaves the coefficient almost unchanged while p moves from 0.024 to 0.095, and to 0.144 with controls: the standard error grows because year effects absorb the variation the tax was competing to explain. A negative sign, then, whose magnitude is not pinned down.

View example →

Dynamic Panels — Nickell Bias, GMM and the Bond Bracket

How much of a firm’s employment carries over from one year to the next? Put that lagged outcome on the right-hand side and both panel workhorses break, in opposite directions: on Arellano & Bond’s own 140-firm panel, pooled OLS puts the carry-over at 0.932 — a shock taking decades to fade — and fixed effects at 0.514, where half of it washes out each year. They bracket a truth neither can reach. Difference GMM, which instruments the lag with earlier levels, then overshoots below fixed effects at 0.336 — an estimator built to correct a downward bias landing past the estimate it was correcting, which is the signature of lagged levels going weak on a persistent series. System GMM restores 0.635. Then the diagnostics split. The serial-correlation test is clean and the instrument count is a healthy 42 against 140 firms, but the over-identification test, which asks whether the surplus instruments agree with one another, rejects at p = 0.003. It is a test that fails and cannot be trusted to have failed for the right reason. The restrictions are unconfirmed, not confirmed.

View example →

Difference-in-Differences — Card–Krueger and the Staggered Trap

The canonical 2×2 on 410 fast-food stores: New Jersey’s minimum wage rose, wages followed, and employment did not fall relative to Pennsylvania — +2.75, reproduced exactly by a regression interaction. Then timing varies, and the same regression stops estimating what it appears to. On a simulation with a known average effect on the treated of 3.580, twoway fixed effects returns 2.677. Goodman-Bacon locates the loss: comparisons against never-treated units average 3.407 and against later-treated 2.756, but those using an already-treated control average 0.843 and carry 27% of the weight — a control group whose own effect is still growing makes the treated look flat. Every ingredient is a legitimate DiD; the blend is biased. Callaway–Sant’Anna, using only clean controls, recovers 3.553.

View example →

Honest DiD — How Big a Violation Would Overturn It

Parallel trends is untestable, and the eyeball test everyone uses cannot do the job. Two simulated event studies have pre-trend statistics of 0.30 and 0.39 — indistinguishable, and both near the 0.14 of sampling noise — while their underlying bias differs tenfold. Because the truth is known, the failure is exact: the coefficient estimates τ + g rather than τ, so the naive estimate is overstated by precisely the per-period drift — 10% in one case, 100% in the other, where the reported effect is more than double the truth. Rambachan & Roth price the assumption instead of asserting it: the breakdown value * separates them at 2.5× against 1.25× the pre-period worst. An estimate that breaks below 1 needs the trend to politely halt at the treatment date.

View example →

Synthetic Control — Proposition 99

One treated unit and no obvious comparison, so build the control instead of finding it. A non-negative, sum-to-one blend of six states — Utah at 0.394, then Montana, Nevada, Connecticut, New Hampshire, Colorado — tracks pre-1988 California to a root-mean-square fit error of 1.66 packs across eighteen years, and the post-1988 gap widens to −26.6 packs by 2000. With a single treated unit there is no standard error, so inference is a placebo permutation over all 38 donors — and California ranks 3rd of 39, p = 0.077, which does not clear 5%. Suggestive rather than decisive, and structurally so: with 38 placebos the smallest attainable p is 0.026. The effect size is robust across implementations; the inference depends on how the donors were matched.

View example →

Synthetic Difference-in-Differences — Unit and Time Weights

The synthesis of the pair before it: unit weights like synthetic control but with a level-shifting intercept, so it matches California’s trend rather than its level, plus time weights favouring the late-1980s years most predictive of what came after. The three estimators separate cleanly — DiD −27.35, synthetic control −19.51, SDID −15.60 — and the from-scratch figure reproduces the published −15.6. Then the inference correction. The jackknife SE of 2.37 usually quoted requires several treated units; Prop 99 has one. The placebo estimator the authors prescribe gives 9.49, four times larger, and its interval includes zero (p = 0.051). The point estimate is solid; its distinguishability from zero is not established by this design.

View example →