← All examples

Causal Inference: Bayesian Analysis of Randomized Experiments

A Bayesian re-analysis is only worth reading if it produces something the original could not. That is the standard this group is built to, and it is a demanding one — done mechanically, re-fitting a settled experiment in PyMC yields the same estimate with more arithmetic and a longer runtime.

So the examples are ordered by a single spine: how much the prior can actually move the answer, which falls as information accumulates. It starts at fifteen matched pairs, where the prior decides the verdict outright and the frequentist result turns out not to survive a second analysis. It ends at three hundred thousand voters, where the prior is provably irrelevant — included deliberately, because a group about what Bayesian methods add should contain the case where they add nothing.

In between sit the examples where the machinery does structural rather than sensitivity work: partial pooling across a handful of groups, hierarchical modelling of clustered assignment, and compliance types treated as the latent classes they are. Identification is not the subject here. Randomization already settled that in the previous group, and every model on these pages inherits it unexamined — which is precisely what makes them a clean place to see what priors and hierarchy contribute on their own.

What the prior can move, what each example adds, and where the verdict changed

Every value is committed notebook output. The five experiments are real and settled; identification was done by randomization in the previous group and is inherited here unexamined, which is what makes these a clean place to see what priors and hierarchy contribute on their own.

A · The spine the group is ordered by: how far a sceptical prior can move the answer Bars show the shift at k = 1 — a prior centred on no effect whose standard deviation equals the whole observed effect. Rows are sorted by sample size, smallest first. experiment n t shift at k = 1 at k = 0.1 Darwin’s maize 15 matched pairs 15 2.15 17.8% 95.6% Electric Company, grade 1 one subgroup of a reading trial 192 3.37 8.1% 89.8% Project STAR, cluster-robust 79 schools 3,743 3.15 9.2% 91.0% Vitamin A, complier effect the whole Sommer–Zeger trial 23,682 2.79 11.4% 92.8% Social Pressure GOTV the deliberate control case 229,444 30.21 0.1% 9.9% Sorted by n, the bars do not shrink. Vitamin A rests on 23,682 children and is more prior-sensitive than the Electric Company’s 192. The exact law is that a prior removes 1/(k²t²+1) of the estimate. Sample size does not appear in it — only t, and t is what n buys, not what n is. B · What each example returns that the frequentist analysis structurally cannot The entry condition for the group: a rerun that reproduces a settled estimate with more arithmetic is not worth a page. Darwin’s maize a sensitivity curve that separates two claims a p-value fuses direction holds above 0.94 for any prior allowing an effect the size observed • magnitude never settles, running 0.31 to 2.62 inches Electric Company the pooling decision as a parameter rather than a choice no pooling treats a noisy 8.79 as real, complete pooling asserts the grades are identical • the model estimates where between them to sit Project STAR seventy-nine school effects, and the ICC as a posterior intraclass correlation 0.214, 95% [0.160, 0.279] • school effects from −16 to +37, with 18 of 79 excluding zero where chance gives four Vitamin A never-taker mortality as a parameter, not an assumption 3.2× the rate compliers would have died at untreated — the selection that biases every as-treated comparison Social Pressure GOTV a proof that the prior is irrelevant, rather than a hope a prior as wide as the whole effect moves the answer 0.0099 percentage points • it takes asserting 1% of the finding to change anything C · And the three places the re-analysis changed the verdict, not the presentation Not wider intervals around the same conclusion — a different conclusion. Darwin’s published result does not survive a second analysis The t-interval clears zero by 0.004 inches. With a prior standard deviation four times the observed effect — nobody’s idea of informative — the credible interval is [−0.050, 5.256] and contains zero. The tightest interval is tightest because its assumption is false The varying-intercept model gives the narrowest of four intervals, and it is narrow precisely because it assumes a constant effect. STAR’s effect spread across schools is 11.9 points against an average of 6.6. Cluster-robust was right to be wide. A calibrated estimator can be useless, and the warning is the wrong one Bloom’s ratio holds 95% coverage all the way down to a compliance rate of 0.05 — it does not break. Its interval widens to [7.98, 165.50] around an effect of size 3. The weak-instrument failure here is width, not bias.

A is the group’s organising claim and its correction in the same picture. The examples are ordered by how far a prior can move the answer, and the intuition for that ordering is sample size: small experiments are prior-sensitive, large ones are not. The first and last rows behave exactly as expected — fifteen matched pairs where a mild sceptic moves the estimate 17.8%, and 229,444 voters where the same sceptic moves it 0.1%.

The middle three are where the intuition fails. Sorted by sample size the bars go 17.8, 8.1, 9.2, 11.4 — rising across a range where n grows from 192 to 23,682. The vitamin A complier effect rests on a hundred and twenty times the Electric Company’s sample and is the more prior-sensitive of the two. The reason is exact rather than anecdotal: a prior removes 1/(k²t²+1) of the estimate, and n does not appear in that expression. Only t does. Sample size buys precision, precision is what resists a prior, and the two come apart whenever an estimator spends its sample on something else — which a complier effect, estimated from the fraction of people an assignment actually moved, does by construction.

B is the entry condition. Re-fitting a settled experiment in PyMC and recovering the same number with a longer runtime is not worth a page, so each example has to return an object the original analysis does not have: a sensitivity curve, a pooling decision converted into a parameter, seventy-nine school effects with an intraclass correlation attached, a never-taker mortality rate, and — in the last case — a proof of its own irrelevance. That last one is included deliberately. A group about what Bayesian methods add is more credible for containing the case where they add nothing.

C is what makes the group more than a methods demonstration. In three of the five the second analysis does not widen an interval around the same conclusion — it reaches a different one. Darwin’s result, in print since 1876 and the origin of the paired t-test, clears zero by 0.004 inches and does not survive a prior nobody would call informative. Project STAR’s tidiest model produces the narrowest interval because its constant-effect assumption is false, which inverts the usual reading of a width comparison and vindicates the cluster-robust standard error the published analysis settled for. And the weak-instrument warning attached to Bloom’s estimator turns out to name the wrong failure: it stays calibrated at 95% down to a compliance rate of 0.05, and becomes useless anyway by widening to [7.98, 165.50] around an effect of size 3.

Read together, those three are one lesson rather than three. Every one of them is a case where the original analysis reported something true and the summary drawn from it was wrong. The t-interval really does exclude zero; the varying-intercept interval really is the narrowest; Bloom’s coverage really is 95%. What the second analysis supplies is not a correction to the arithmetic but the missing question — how much does this depend on what I assumed, what is this interval narrow because of, and calibrated at what width.

How the five examples relate

One spine, running from where the prior decides everything to where it provably decides nothing, with the examples that do structural rather than sensitivity work sitting in between.

The distinction the group turns on is worth carrying forward. Here the prior competes with the data and loses as information accumulates — that is what panel A measures, and it is why the last example can prove its own prior irrelevant. In the observational group that follows, the priors stand in for quantities no amount of data identifies, so nothing ever swamps them. Same machinery, opposite situation, and the reason the second group needs a different justification rather than a longer version of this one’s.

Bayesian Analysis of a Small Experiment — Darwin’s Maize

The published t-interval on Darwin’s 15 pairs clears zero by 0.004 inches. Re-analysed, the verdict does not hold: with a prior standard deviation four times the observed effect — nobody’s idea of informative — the credible interval is [−0.050, 5.256] and contains zero. A sensitivity sweep then separates two claims the p-value fuses together: the direction holds above 0.94 for any prior that allows an effect the size of the one observed, falling to 0.745 only under active scepticism, while the magnitude is never settled at all, running from 0.31 to 2.62 inches across the sweep. A Student-t likelihood with estimated degrees of freedom then handles the two outlying pairs Fisher used a permutation test to avoid — and moves the estimate up to 2.855 with a tighter interval. Three analyses, one experiment, and they disagree about whether the effect is established.

View example →

Partial Pooling Across Subgroups — The Electric Company by Grade

Four subgroup effects that look five-fold apart — grade 1 at 8.79 against grade 4 at 1.70 — where the biggest estimate is also the least precise. The frequentist choice is between treating that 8.79 as real and asserting the grades are identical; a hierarchical model makes the choice a parameter. Partial pooling moves grade 1 to 6.14, and the mechanism is not the usual summary: shrinkage is imprecision times distance from the mean, so grade 2 barely moves while the most precise estimate is pulled nearly a point — upward, since shrinkage goes toward the common mean, not toward zero. The between-grade spread turns out barely identified (95% [0.13, 7.44]), and the grade 1 vs grade 4 gap that looked decisive at +7.09 is +3.49 with Pr = 0.888.

View example →

Hierarchical Models for Clustered Assignment — Project STAR

Seventy-nine schools, where the published analysis fixed the clustering with a cluster-robust standard error and could say nothing further. Modelling it instead gives the intraclass correlation as a posterior (0.214, 95% [0.160, 0.279]) and seventy-nine school-level effects — and then contradicts the constant-effect assumption outright: the spread of the effect across schools is 11.9 points against an average of 6.6, with school effects from −16 to +37 and 18 of 79 excluding zero where chance gives four. The width comparison then reverses the expected moral. The varying-intercept model produces the tightest interval of four — and is tighter precisely because it assumes the effect is constant, which is false. Cluster-robust was right to be wide: assuming nothing about how schools differ, it absorbed heterogeneity the simpler model wished away.

View example →

Principal Stratification and Weak Instruments — Vitamin A

Compliance type is a latent class, so the control arm is a mixture of compliers and never-takers — and modelling it as one, instead of backing the complier effect out of a ratio, returns never-taker mortality as a parameter: 3.2 times the rate compliers would have died at untreated, which is the selection that biases as-treated comparisons. Then a weak-instrument study overturns the standard warning. Bloom’s ratio stays calibrated at 95% all the way down to a compliance rate of 0.05 and becomes useless anyway, its interval widening from 7.98 to 165.50 around an effect of size 3. The Bayesian alternative over-covers at 100% while being a third as wide — narrower and covering more, which means information came from somewhere other than the data. A prior sweep locates only about a third of it in the prior; the rest is the mixture structure.

View example →

The Large-Sample Limit — Social Pressure GOTV

The deliberate control: 305,866 voters, where the prior is shown to be irrelevant rather than assumed to be. A prior centred on no effect whose standard deviation equals the entire observed effect moves the answer by 0.0099 percentage points; it takes a prior asserting the effect is one percent of what was found to change anything. Behind that is an exact law — a prior removes 1/(k2t2+1) of the estimate — in which sample size does not appear. That corrects this group’s own ordering: the vitamin A complier effect rests on 23,682 children at t = 2.79 and is more prior-sensitive than the Electric Company’s 192 at t = 3.37. Then the same voters, split by household size and age, give cells of 14 — where pooling moves the sparsest by 11 points. Large was always a property of the question.

View example →