Causal Inference
A theory of causation must answer the one question association never can: what would have happened otherwise?
Every empirical question that matters to policy is causal — does the training programme raise employment, does the minimum wage cost jobs, does the treatment extend life — and every one of them runs into the same wall. For any unit we observe the outcome under the treatment it actually received, never the outcome under the treatment it did not. This is the fundamental problem of causal inference: the counterfactual is missing by construction, so a causal effect is never observed, only estimated under assumptions. Correlation is what the data hand you for free; causation is what you must earn.
This arc is organised around a single pedagogical claim — move from the setting where causation is cheap to the settings where it is expensive — and around the field’s great synthesising texts.
The Books, and Why the Order Is the Argument
Imbens & Rubin, Causal Inference for Statistics, Social, and Biomedical Sciences, supplies the potential-outcomes spine and, crucially, its ordering. Randomized experiments come first, establishing the counterfactual benchmark in the one setting where the assumptions are credible by design, before the relaxation to observational data and the identification problems that follow. Most causal-inference collections open with instrumental variables or difference-in-differences and never mention that randomization is where the entire framework is defined. This one starts there deliberately: a designed experiment is not a warm-up, it is the standard against which every observational method is secretly measured.
Cunningham, Causal Inference: The Mixtape, supplies the applied toolkit and, just as important, the datasets. Each method is paired with the paper that made it canonical, which is this collection’s format throughout: build the estimator from scratch, validate it against the standard package, and add an R companion where R is the field’s lingua franca.
Morgan & Winship, Counterfactuals and Causal Inference, enters near the end, where Pearl’s structural-graph language earns its place alongside potential outcomes as a second grammar for the same problem. Hernán & Robins, Causal Inference: What If, anchors the final application — the g-methods and marginal structural models built for treatment effects on time-to-event data under censoring and time-varying confounding.
One divergence is deliberate. The Mixtape, following Pearl, introduces directed acyclic graphs early, as the primary language, before potential outcomes. This arc keeps the deep DAG treatment for late, as an explicit bridge — two languages for one problem — because the potential-outcomes-first path makes the assumptions behind each design impossible to hide. The choice is pedagogical, not a substitution of one framework for the other.
The Through-Line
The discipline econometrics teaches and prediction often forgets: a number is only as causal as the design that produced it. Machine learning asks what is associated with what; this arc asks what would happen if we intervened, and is relentless about the assumptions separating the two. Each subsection replicates a landmark study on its real data, builds the estimator from scratch and checks it against the reference package, and states plainly what the design does and does not identify.
Where the Material Is
Ten subsections, built as eight groups. The three on the left are the argument’s spine, ordered by how much has to be assumed; the two on the right are second passes over the same experiments and the same data, asking what priors and hierarchy contribute once identification is settled or once it cannot be.
Identification by design
A researcher assigned the treatment. Subsection 1.
Randomized Experimentsthe counterfactual benchmark, where the assumptions are credible by construction Bayesian Analysis of Randomized Experimentsa second pass — what priors add when identification is not in doubtIdentification by assumption
Nobody assigned anything; the covariates have to carry it. Subsection 2.
Selection on Observablesunconfoundedness, and the diagnostics that stand in for a test of it Bayesian Selection on Observablesa second pass — priors for the quantities no data identifiesIdentification from a graph
Which variables belong in the model, and where does the graph come from? Subsection 8.
Causal Structurethe back-door criterion, the collider trap, and how much of a graph data alone can supplyEstimation, once identification is granted
Conditional effects, flexible nuisance, decisions, and dynamics. Subsections 9–10.
Frontier — ML & Dynamicsτ(x) rather than an average, orthogonalization, policy rules, censoring and treatment–confounder feedbackIdentification borrowed
The world ran something close to an experiment. Subsections 3–7.
Natural & Quasi-Experimentsan instrument, a threshold, a policy date, a comparison region — each paired with the literature that polices itThe three trunk groups run in that order for a reason, and it is the arc’s whole argument: each step outward buys applicability by giving up something that was free the step before. A randomized experiment identifies the effect and needs almost nothing else. Selection on observables applies anywhere covariates were recorded and pays for it with an assumption no diagnostic can check. Borrowed identification applies where no covariate set would have been enough and pays for it in precision — the effect is real, and the design often cannot say how far it is from zero.
The Arc
The list below is the order of the argument as well as the order of construction, and the entries carrying links are the ones with material behind them. Subsections three to seven were built as a single group: each of those designs borrows its identification from something outside the study, and each is paired with the diagnostic literature that grew up to police it, so the design and its critic belong on the same shelf rather than five shelves apart.
-
Randomized Experiments & Randomization Inference
Fisher’s design-of-experiments tradition and the exact test via permutation, set against Neyman’s repeated-sampling framework. The Rubin causal model — potential outcomes, SUTVA — is introduced here, grounded where randomization is genuinely credible, before any identification problem appears. The subsection carries the design questions that follow from it: covariate adjustment for precision, noncompliance and clustering, and the online A/B-testing machinery — sequential testing, ratio metrics, interference, multiple comparisons — which inherits the whole framework and then has to survive contact with traffic.
A second pass: Bayesian Analysis of Randomized Experiments
The Bayesian group is not one of the ten. It is a second pass over the first — the same real experiments re-analysed with priors and hierarchical structure, kept separate because it answers a different question. Not what is the effect, which randomization already settled, but what do priors and hierarchy contribute once identification is not in doubt. Randomized experiments are the cleanest place to ask that, precisely because nothing else is in dispute.
-
Potential Outcomes & Matching
Identification by assumption rather than by design. Unconfoundedness, propensity-score estimation and overlap, matching (nearest-neighbour, Mahalanobis, and genetic in the R companion), covariate-balance diagnostics, IPW and doubly-robust AIPW; sensitivity analysis for the assumption that cannot be tested; and the modern balancing estimators — entropy balancing, CBPS, TMLE — that target balance directly rather than through a model of who gets treated. All three examples run on LaLonde, where a randomized trial supplies the answer key, so the estimators can be graded rather than merely compared. Cross-link → Variable Selection: choosing which confounders to condition on is a variable-selection problem.
What the group turns up
Adjustment works — the naive comparison gets the sign wrong and propensity matching lands on the benchmark almost exactly. But every diagnostic in the group measures something other than what you want to know. Aggregate balance picks the wrong estimator: the better-balanced scheme misses by 60%. Exact balance is bought with 77% of the effective sample, which the balance table does not show. And the estimate that landed on the truth is erased by a confounder shifting within-pair odds by a factor of 1.21.
A second pass: Bayesian Selection on Observables
The same second-pass logic, one setting further out. Where the randomized group asks what priors add once identification is settled, this one asks what they add when identification rests on an assumption nothing can test. The answer is different in kind: here the priors stand in for quantities no amount of data identifies — the unmeasured confounder, and the response surface where treated and control units do not overlap — so nothing ever swamps them.
-
Instrumental Variables
Two-stage least squares from scratch against off-the-shelf, and the LATE framework (Imbens & Angrist 1994) on Card’s proximity-to-college instrument, where the confounder — ability — is precisely what no covariate captures. The local nature of the estimate is made concrete rather than asserted: the instrument moves 12.2% of the sample, and the estimate speaks for that eighth alone. Then → Weak-Instrument-Robust Inference: diagnosing a weak instrument is not surviving it. Anderson–Rubin sets, and the recalibration of the folklore F > 10 to F > 104.7 — which dissolves the tidy significance on Card’s own data. Angrist & Krueger’s study enters here, first as the many-weak-instruments structure and then on its own 329,509 census records, where randomly generated quarters of birth run through the published specification return a significant estimate 197 times out of 200.
-
Regression Discontinuity
Sharp and fuzzy designs, local-linear estimation with a triangular kernel, and bandwidth selection (Calonico–Cattaneo–Titiunik) on Lee’s 6,558 House races. The fuzzy variant turns out not to be like instrumental variables but to be them, applied at a boundary, inheriting the whole LATE apparatus — so the effect is local twice over. Then → Validity & Falsification: the standard suite — covariate continuity, density, placebo cutoffs, donut and bandwidth — separated by what each one actually certifies. The first three license the design; the last two describe the estimate, and here they disagree.
-
Panel Data & Fixed Effects
The within and first-difference estimators, fixed against random effects with a Hausman test, and twoway fixed effects — the workhorse of applied micro, and the estimator that difference-in-differences turns out to be a special case of. On Stock & Watson’s fatality panel, pooling says higher beer taxes go with more deaths; differencing the states away flips the sign. Then → Dynamic Panels: one lagged outcome breaks both workhorses in opposite directions — Nickell bias, the Bond bracket, and Arellano–Bond and Blundell–Bond GMM built from scratch to climb back inside it. Cross-link → Multivariate Time Series / BVAR: the same panel structure, a different question.
-
Difference-in-Differences
The canonical 2×2 on Card & Krueger’s minimum-wage study, then the staggered-adoption revolution that has reshaped applied econometrics since roughly 2018. The Goodman-Bacon decomposition locates where twoway fixed effects loses a known effect: comparisons using an already-treated control carry 27% of the weight and average 0.843 against a truth of 3.580. Callaway & Sant’Anna, using only clean controls, recovers it. Then → Honest DiD: parallel trends is untestable and the eyeball test cannot do the job — two pre-trends that look identical hide bias differing tenfold. Rambachan & Roth price the assumption instead of asserting it.
-
Synthetic Control
Abadie-style comparative case studies on California’s Proposition 99: build the control rather than find it, as a non-negative sum-to-one blend of donor states, with inference by placebo permutation because a single treated unit admits no standard error. A small, high-signal subsection, disproportionately visible in applied policy work. Then → Synthetic Difference-in-Differences: unit and time weights, reproducing the published figure — and the standard error usually quoted with it, which is not valid with one treated unit.
What the middle group turns up
Across all five, the diagnostic almost never overturns the estimate and very often overturns the confidence attached to it. Identification borrowed from the world is narrow — it rests on the compliers, or the neighbourhood of a cutoff, or one treated unit — so the effective sample is small however many rows the dataset has. Card has 3,010 men and the instrument moves 12.2% of them; Proposition 99 has 31 years and one treated state, which caps the attainable permutation p at 0.026 before any data is examined. The group page sorts the failures into the three kinds that call for three different responses.
-
DAGs, Mediation & the Structural Causal Model
Pearl’s do-operator, identification by graph, the front-door criterion, and causal mediation — framed as the explicit bridge: potential outcomes against structural graphs, two languages for the same problem. Morgan & Winship earns its distinct place here. Each rule is verified against a known effect and then broken on purpose, which is where the collider result lands: adding one control to a correct regression moves a true effect of 2 to 0.50, so the instinct that more controls are safer is not imperfect but inverted. Then → Causal Discovery: if identification is read off a graph, where does the graph come from? PC, GES and LiNGAM, the Markov-equivalence limit, and the Sachs single-cell benchmark that interventions had to finish. And → DAGs and Discovery on Survey Data: both methods on NHANES, where a bad control costs 71% of the estimate and the algorithm orients every edge touching age and sex backwards — until one sentence of background knowledge is supplied.
What the group turns up
The sharpest result is about what an experiment does not buy. Randomizing the treatment identifies the total effect and nothing else — add a confounder of the mediator and the outcome, and the total holds to within 0.09 of the truth while the estimated direct effect falls from 0.50 to 0.06. The output then reads “almost entirely mediated” and 37% of the effect never goes through the mediator. Every mechanism claim layered on a randomized experiment is observational, and is printed in the same table with the same standard errors. The group page collects the four rules against the price of each one’s condition.
-
Heterogeneous Effects & Double/Debiased Machine Learning
From the average effect to the conditional effect τ(x): meta-learners (S, T, X, R), causal forests, and Neyman-orthogonal Double ML with cross-fitting. This is where the causal arc meets the machine-learning arc — honest trees, regularized nuisance models, and the same out-of-sample discipline, cross-fitting being purged cross-validation’s cousin. Built as four examples → causal forests, meta-learners, double machine learning and policy learning, all graded against a known τ(x) — which is how the undercovering intervals and the inverting benchmark rankings turned up.
-
Causal Survival Analysis
Treatment effects on time-to-event outcomes under censoring: inverse-probability weighting and the g-formula for survival, marginal structural models for time-varying confounding (Robins), and the modern critique of the naive causal hazard ratio (Hernán, “The hazards of hazard ratios”) in favour of restricted-mean-survival-time and survival-probability contrasts. Cross-link → the Survival catalogue: Weibull PH, Cox-via-Poisson, frailty and interval-censored models — the same likelihoods, now asked a causal question. Then → Marginal Structural Models: when treatment repeats and the confounder responds to it, adjusting for that confounder is worse than not adjusting, and only reweighting recovers the effect.
What the group turns up
Flexibility moves the failure rather than removing it. A forest that ranks people correctly and mis-states their effects; confidence intervals that undercover in both reference implementations; a benchmark ranking that inverts between datasets and between implementations; and a coverage check on 40 replications that could not have detected its own failure, since a clean sweep of forty has probability 0.13 even at exactly nominal coverage. The group page collects them.