Causal Inference: Bayesian Analysis of Randomized Experiments
A Bayesian re-analysis is only worth reading if it produces something the original could not. That is the standard this group is built to, and it is a demanding one — done mechanically, re-fitting a settled experiment in PyMC yields the same estimate with more arithmetic and a longer runtime.
So the examples are ordered by a single spine: how much the prior can actually move the answer, which falls as information accumulates. It starts at fifteen matched pairs, where the prior decides the verdict outright and the frequentist result turns out not to survive a second analysis. It ends at three hundred thousand voters, where the prior is provably irrelevant — included deliberately, because a group about what Bayesian methods add should contain the case where they add nothing.
In between sit the examples where the machinery does structural rather than sensitivity work: partial pooling across a handful of groups, hierarchical modelling of clustered assignment, and compliance types treated as the latent classes they are. Identification is not the subject here. Randomization already settled that in the previous group, and every model on these pages inherits it unexamined — which is precisely what makes them a clean place to see what priors and hierarchy contribute on their own.
What the prior can move, what each example adds, and where the verdict changed
Every value is committed notebook output. The five experiments are real and settled; identification was done by randomization in the previous group and is inherited here unexamined, which is what makes these a clean place to see what priors and hierarchy contribute on their own.
A is the group’s organising claim and its correction in the same picture. The examples are ordered by how far a prior can move the answer, and the intuition for that ordering is sample size: small experiments are prior-sensitive, large ones are not. The first and last rows behave exactly as expected — fifteen matched pairs where a mild sceptic moves the estimate 17.8%, and 229,444 voters where the same sceptic moves it 0.1%.
The middle three are where the intuition fails. Sorted by sample size the bars go 17.8, 8.1, 9.2, 11.4 — rising across a range where n grows from 192 to 23,682. The vitamin A complier effect rests on a hundred and twenty times the Electric Company’s sample and is the more prior-sensitive of the two. The reason is exact rather than anecdotal: a prior removes 1/(k²t²+1) of the estimate, and n does not appear in that expression. Only t does. Sample size buys precision, precision is what resists a prior, and the two come apart whenever an estimator spends its sample on something else — which a complier effect, estimated from the fraction of people an assignment actually moved, does by construction.
B is the entry condition. Re-fitting a settled experiment in PyMC and recovering the same number with a longer runtime is not worth a page, so each example has to return an object the original analysis does not have: a sensitivity curve, a pooling decision converted into a parameter, seventy-nine school effects with an intraclass correlation attached, a never-taker mortality rate, and — in the last case — a proof of its own irrelevance. That last one is included deliberately. A group about what Bayesian methods add is more credible for containing the case where they add nothing.
C is what makes the group more than a methods demonstration. In three of the five the second analysis does not widen an interval around the same conclusion — it reaches a different one. Darwin’s result, in print since 1876 and the origin of the paired t-test, clears zero by 0.004 inches and does not survive a prior nobody would call informative. Project STAR’s tidiest model produces the narrowest interval because its constant-effect assumption is false, which inverts the usual reading of a width comparison and vindicates the cluster-robust standard error the published analysis settled for. And the weak-instrument warning attached to Bloom’s estimator turns out to name the wrong failure: it stays calibrated at 95% down to a compliance rate of 0.05, and becomes useless anyway by widening to [7.98, 165.50] around an effect of size 3.
Read together, those three are one lesson rather than three. Every one of them is a case where the original analysis reported something true and the summary drawn from it was wrong. The t-interval really does exclude zero; the varying-intercept interval really is the narrowest; Bloom’s coverage really is 95%. What the second analysis supplies is not a correction to the arithmetic but the missing question — how much does this depend on what I assumed, what is this interval narrow because of, and calibrated at what width.
How the five examples relate
One spine, running from where the prior decides everything to where it provably decides nothing, with the examples that do structural rather than sensitivity work sitting in between.
Where the prior decides
Fifteen pairs, and the answer depends on what you brought.
1 · A small experimentDarwin’s maize — direction survives, magnitude never settles, and the published verdict does not holdWhere the structure does the work
Not sensitivity — objects the original analysis could not produce.
2 · Partial pooling across subgroupsthe choice between four estimates and one becomes a parameter 3 · Hierarchical models for clustered assignmentseventy-nine school effects, and a refutation of the constant effect 4 · Principal stratificationcompliance type as a latent class, so the control arm is modelled as the mixture it isWhere the prior is irrelevant
Included because the group would be less honest without it.
5 · The large-sample limit305,866 voters, the exact shrinkage law, and the sample-size intuition it correctsWhere it goes next
The same machinery where randomization is not available.
Bayesian Selection on Observablespriors standing in for quantities the data cannot identify at allThe distinction the group turns on is worth carrying forward. Here the prior competes with the data and loses as information accumulates — that is what panel A measures, and it is why the last example can prove its own prior irrelevant. In the observational group that follows, the priors stand in for quantities no amount of data identifies, so nothing ever swamps them. Same machinery, opposite situation, and the reason the second group needs a different justification rather than a longer version of this one’s.
Bayesian Analysis of a Small Experiment — Darwin’s Maize
The published t-interval on Darwin’s 15 pairs clears zero by 0.004 inches. Re-analysed, the verdict does not hold: with a prior standard deviation four times the observed effect — nobody’s idea of informative — the credible interval is [−0.050, 5.256] and contains zero. A sensitivity sweep then separates two claims the p-value fuses together: the direction holds above 0.94 for any prior that allows an effect the size of the one observed, falling to 0.745 only under active scepticism, while the magnitude is never settled at all, running from 0.31 to 2.62 inches across the sweep. A Student-t likelihood with estimated degrees of freedom then handles the two outlying pairs Fisher used a permutation test to avoid — and moves the estimate up to 2.855 with a tighter interval. Three analyses, one experiment, and they disagree about whether the effect is established.
View example →Partial Pooling Across Subgroups — The Electric Company by Grade
Four subgroup effects that look five-fold apart — grade 1 at 8.79 against grade 4 at 1.70 — where the biggest estimate is also the least precise. The frequentist choice is between treating that 8.79 as real and asserting the grades are identical; a hierarchical model makes the choice a parameter. Partial pooling moves grade 1 to 6.14, and the mechanism is not the usual summary: shrinkage is imprecision times distance from the mean, so grade 2 barely moves while the most precise estimate is pulled nearly a point — upward, since shrinkage goes toward the common mean, not toward zero. The between-grade spread turns out barely identified (95% [0.13, 7.44]), and the grade 1 vs grade 4 gap that looked decisive at +7.09 is +3.49 with Pr = 0.888.
View example →Hierarchical Models for Clustered Assignment — Project STAR
Seventy-nine schools, where the published analysis fixed the clustering with a cluster-robust standard error and could say nothing further. Modelling it instead gives the intraclass correlation as a posterior (0.214, 95% [0.160, 0.279]) and seventy-nine school-level effects — and then contradicts the constant-effect assumption outright: the spread of the effect across schools is 11.9 points against an average of 6.6, with school effects from −16 to +37 and 18 of 79 excluding zero where chance gives four. The width comparison then reverses the expected moral. The varying-intercept model produces the tightest interval of four — and is tighter precisely because it assumes the effect is constant, which is false. Cluster-robust was right to be wide: assuming nothing about how schools differ, it absorbed heterogeneity the simpler model wished away.
View example →Principal Stratification and Weak Instruments — Vitamin A
Compliance type is a latent class, so the control arm is a mixture of compliers and never-takers — and modelling it as one, instead of backing the complier effect out of a ratio, returns never-taker mortality as a parameter: 3.2 times the rate compliers would have died at untreated, which is the selection that biases as-treated comparisons. Then a weak-instrument study overturns the standard warning. Bloom’s ratio stays calibrated at 95% all the way down to a compliance rate of 0.05 and becomes useless anyway, its interval widening from 7.98 to 165.50 around an effect of size 3. The Bayesian alternative over-covers at 100% while being a third as wide — narrower and covering more, which means information came from somewhere other than the data. A prior sweep locates only about a third of it in the prior; the rest is the mixture structure.
View example →The Large-Sample Limit — Social Pressure GOTV
The deliberate control: 305,866 voters, where the prior is shown to be irrelevant rather than assumed to be. A prior centred on no effect whose standard deviation equals the entire observed effect moves the answer by 0.0099 percentage points; it takes a prior asserting the effect is one percent of what was found to change anything. Behind that is an exact law — a prior removes 1/(k2t2+1) of the estimate — in which sample size does not appear. That corrects this group’s own ordering: the vitamin A complier effect rests on 23,682 children at t = 2.79 and is more prior-sensitive than the Electric Company’s 192 at t = 3.37. Then the same voters, split by household size and age, give cells of 14 — where pooling moves the sparsest by 11 points. Large was always a property of the question.
View example →