Causal Inference: Causal Structure
Every group so far chose an estimator and argued about what it identifies. This one steps back to the language that decides which variables belong in the model at all — Pearl’s structural causal models and the directed acyclic graphs that encode them.
The payoff is that identification becomes something you read off a picture rather than argue about in prose. The back-door criterion names the confounders to adjust for, and turns out to be unconfoundedness restated. The front door identifies an effect through a mediator even when an unobserved confounder makes the back door impossible. And the collider rule delivers the single most consequential result here: adjusting for the wrong variable creates bias rather than removing it, which is the exact opposite of the instinct that more controls are safer.
The two pages divide the subject cleanly. The first assumes the graph is known and reads identification off it. The second asks where the graph comes from, and finds that observational data recovers it only up to an equivalence class — which is why this group closes the arc’s argument rather than replacing it. Graphs tell you what would be identified if your assumptions held; they do not supply the assumptions.
What a graph identifies, and what each rule costs when its condition fails
Every value is committed notebook output. The first two examples run on simulations with known effects, which is the only honest way to test whether an identification rule recovers a known answer, and on the Sachs benchmark. The third repeats both methods on national health-survey data, where the reader can grade the result without deferring to anyone.
A is the group’s substance. Four rules, each recovering a known effect when its condition holds — that is the left-hand column, and it is a genuine achievement, because these are questions no estimator can answer on its own. The back-door criterion does not merely resemble the unconfoundedness assumption behind matching; it is that assumption, written as a property of a picture instead of a sentence about treatment assignment.
The right-hand column is the part usually left implicit. Two of the four rules rest on conditions about unobserved variables, so neither can be checked against the data it is applied to, and both were tested here by violating them deliberately. The front door degrades to 134% above the truth, sliding back toward the confounded answer it was brought in to escape. Mediation degrades further: the estimated direct effect goes from 0.50 to 0.06. In both cases the estimator returns a number in every row, with a standard error, in the same table.
B is the result worth carrying out of the group. Adding one control to a correctly specified regression moves a true effect of 2 down to 0.50. The algebra gives exactly one half for this design, so it is structural rather than an unlucky draw, and nothing in the output looks wrong — the coefficient is precise and the model fits better. The instinct that more controls are safer is not merely imperfect here, it is inverted, and this is the one bias in the whole arc that gets worse the more careful an analyst is by conventional standards.
It also generalises further than a lesson about a third variable. Selection into a sample is conditioning on a collider. Surviving long enough to be surveyed, being admitted to hospital, staying in a panel, appearing in an administrative file — each is a common effect of the things being studied, and each induces the same spurious association silently, in data that looks complete and well-measured.
C is the honest limit, and it is why the second page exists. If identification is read off a graph, everything depends on where the graph came from — so the natural question is whether it can be learned. Partly. On the field’s own benchmark, PC achieves perfect precision, which sounds like a triumph until you notice it proposed 7 adjacencies out of 55 possible pairs. It is being reluctant rather than accurate, and what little it commits to is safe.
The sweep is what settles the interpretation. Loosening the test twentyfold adds three false edges and one true one, so the ten missing edges are not near misses waiting for a friendlier threshold — they are absent from the observational distribution, and no tuning recovers them. Two independent implementations in two languages, using two different algorithm families, return the same seven edges. The boundary is in the data.
Which gives the group its ending, supplied by molecular biology rather than econometrics. Sachs et al. reconstructed the full network by intervening — perturbing each protein in turn — not by analysing the observational data harder. The canonical demonstration that causal structure can be learned from data was completed by running experiments, which is this arc’s through-line arriving from an unexpected direction.
How the two examples relate
The first two are inverses of one another, and the order matters: the rules have to be worth having before the question of where the graph comes from is interesting. The third runs both on data where the reader, rather than a consensus network, is the ground truth.
Given the graph
Which variables belong in the model, and which must be left alone?
1 · DAGs, mediation & the SCMback door, collider, front door and the mediation split — each with the price of its own conditionLearning the graph
Can the structure itself be recovered from data?
2 · Causal discoveryPC, GES and LiNGAM, the Markov-equivalence limit, and the Sachs benchmark that interventions had to finishBoth, where you are the expert
Same two methods, on variables that need no specialist to referee.
3 · DAGs and discovery on survey dataNHANES — a bad control that costs 71%, and eleven of eleven orientations backwards until one sentence of knowledge is suppliedWhat it connects to
The graph language restates the rest of the arc.
Selection on Observablesthe back-door criterion is its identifying assumption, drawn Natural & Quasi-Experimentsinstruments and discontinuities are graphically distinct identification routesThe thread joining them is a warning against reading either page as a promise. A DAG does not supply your assumptions; it makes them explicit and derives their consequences. That is worth a great deal — the collider result is invisible without it, and no amount of care with estimation or standard errors would surface it — but the arrows still come from you. Discovery narrows which sets of arrows are consistent with the data, and stops well short of choosing among them. Both pages end in the same place the rest of the arc does: identification comes from design and from knowledge, and the formalism’s contribution is to stop you from fooling yourself about which of the two you are relying on.
DAGs, Mediation & the Structural Causal Model
Four rules read off a picture, each verified against a known effect and then broken on purpose. The back-door criterion takes a naive 3.48 to 2.00 against a truth of 2 — and is unconfoundedness, restated graphically. The collider rule runs the other way: adding one control to a correct regression moves the estimate to 0.50, exactly a quarter of the truth by algebra rather than luck, which makes selection into a sample a silent bias in data that looks complete. The front door identifies an effect through a mediator when the back door is impossible — then goes 134% wrong when the unobserved confounder also touches the mediator, sliding back toward the naive answer. And the mediation split is the sharpest: a mediator–outcome confounder leaves the randomized total effect intact while driving the estimated direct effect from 0.50 to 0.06, at which point the output reads “almost entirely mediated” and 37% of the effect never goes through the mediator. A randomized experiment buys the total effect and nothing else.
View example →Causal Discovery — PC, GES and LiNGAM
If identification is read off a graph, where does the graph come from? Observational data recovers it only up to a Markov equivalence class: PC and GES, two different algorithm families, return the same CPDAG on simulated data with the same edge left undirected, because both orientations imply identical independencies. LiNGAM escapes that with non-Gaussian noise and recovers every direction — and the 0.15 coefficient cutoff that looks like a tuning knob is not one, since the fitted matrix prunes to exact zeros with nothing between 0.63 and 0.91. Then the real benchmark. On Sachs et al.’s 853 single cells, PC achieves perfect precision by proposing 7 adjacencies out of 55 possible pairs — reluctance rather than accuracy. Loosening the test twentyfold adds three false edges and one true one, so the missing cascade is not a tuning problem. Sachs needed interventions to finish the network.
View example →DAGs and Discovery on Survey Data
The same two methods on NHANES, where every variable is one you already have opinions about — and where you can therefore grade the answer yourself. Estimating whether exercise lowers blood pressure on 4,048 adults, the back-door set is genuinely needed: age alone moves the raw −4.573 mmHg to −1.600. Then the near-universal next step ruins it. Adjusting for BMI costs 71% of what remains, and the arithmetic says exactly where it went — the effect travelling through BMI is −0.601 mmHg and the estimate lost −0.599, the same number to 0.002. Nothing in the output separates the right specification from the wrong one. Then discovery, scored against a ground truth needing no expertise at all: nothing causes your age. PC orients eleven of eleven edges touching age and sex backwards. Supplying that one sentence as background knowledge fixes all eleven, repairs two more by propagation that nobody touched, and resolves the last ambiguous edge.
View example →