← All examples

Causal Inference: Randomized Experiments

Every empirical question that matters is causal, and every one of them meets the same wall: for any unit we observe the outcome under the treatment it actually received, never the outcome under the treatment it did not. The counterfactual is missing by construction, so a causal effect is never observed — only estimated under assumptions.

This group is the one place in the arc where those assumptions are nearly free. When the researcher assigns treatment at random, treated and control groups match in expectation on every pre-treatment characteristic, observed and unobserved, and the plain difference in means is unbiased for the average effect. Identification comes from the design rather than from a model. Everything later in the arc — matching, instruments, discontinuities, difference-in-differences — is an attempt to earn back what randomization gives away here for nothing.

Which makes this group the right place to be precise about what randomization does and does not buy. It does not buy precision, so covariate adjustment still pays. It does not survive units refusing the treatment they were assigned, or treatments that leak between units. And at the scale of online experimentation it meets a new set of failures that have nothing to do with identification and everything to do with how the test is run: peeking, ratio metrics, dashboards of correlated outcomes, and interference between users who share a marketplace.

What the design settles, what it does not, and what each repair costs

Every value is committed notebook output. A draws on the three real-data examples, B and C on the simulations — which is not a shortcut but a requirement, since measuring an estimator’s error rate needs a truth to measure against.

A · Randomization identifies the effect. Everything else in the group is what it does NOT buy. identification confounded assignment vs randomized, true ATE 4.0 9.21 4.00 bias +5.21 removed by design precision Electric Company: unadjusted vs pre-test adjusted SE 2.537 1.169 variance cut 79% — adjustment, not design compliance vitamin A: as-treated vs complier effect, per 1,000 −6.47 −3.23 randomization survives only via the instrument honest SEs Project STAR: naive vs cluster-robust SE 1.042 1.850 naive 95% intervals covered 54% Only the first row is settled by the design. The other three are the price of a real experiment meeting real subjects. B · Four online-experiment failures, each measured where the truth is known by construction 0% 25% 50% 75% 100% intended 5% peeking at 20 looks 24% Type-I error against a nominal 5% Bayesian Pr(B>A)>0.95 rule 58% the same failure, worse ratio metric, naive SE 13% 1 − coverage; nominal 5% miss rate 160-metric dashboard 100% P(at least one false winner) Every one is a false-positive rate that was supposed to be 5%. None of them is an identification problem — the randomization was valid in all four. C · The repairs, and what each one costs cluster-robust SEs 54% → 94% coverage cost: none — the naive SE was simply wrong cluster randomization bias +0.588 → +0.002 cost: SD 0.019 → 0.021, a 30:1 trade mSPRT 24% → 3.6% under any peeking cost: none — and it stops sooner (1,454 vs 1,570) Benjamini–Hochberg FDR 0.040, power 0.613 cost: real — Bonferroni holds FWER but finds 1 in 4 delta method 87% → 95% coverage cost: none — matched by a cluster bootstrap CUPAC variance −75%, SE halved cost: none, IF the covariate is pre-treatment only Five of the six repairs are close to free. That is unusual, and it is a property of experiments: when identification is settled by design, the remaining problems are measurement problems, and measurement problems have clean fixes.

A separates the one thing randomization does from the three it does not. Only the first row is settled by the design — a confounder drives the naive estimate to 9.21 against a true effect of 4.0, and randomizing the same units returns 4.00 averaged over 4,000 re-randomizations. Nothing about the estimator changed; only who was assigned where. The other three rows are the price of a real experiment meeting real subjects: randomization does not make an estimate precise, does not survive subjects declining the treatment they were assigned, and does not by itself produce an honest standard error when whole groups were assigned together.

B is the group’s most transferable finding, and the framing matters. All four bars are false-positive rates that were supposed to be 5%, and in every case the randomization was perfectly valid. Peeking twenty times gives 24%; the Bayesian version of the same impulse gives 58%, worse rather than better, because the inflation belongs to stopping on a wandering statistic and not to a statistical philosophy. Computing a ratio metric at the wrong unit turns a 5% miss rate into 13%. And scoring a dashboard of 160 null metrics produces at least one spurious winner with probability 1.000 — not approaching certainty but at it. None of these is an identification problem, which is exactly why a team confident in its randomization can still be wrong almost all of the time.

C is the reassuring half, and the asymmetry is worth noticing. Five of the six repairs cost essentially nothing: the naive standard error was simply wrong, so replacing it is free; the mSPRT controls error under unlimited peeking and stops sooner than the fixed design; CUPAC halves the standard error using data already collected. Only Benjamini–Hochberg involves a real trade, and even there the alternative is worse than it looks — Bonferroni’s tiny false-discovery rate is the symptom of a method that finds one true effect in four. That so many fixes are close to free is a property of experiments specifically: once identification is settled by design, what remains are measurement problems, and measurement problems have clean solutions. The rest of the causal arc, where identification must be argued for rather than assigned, has no such luxury.

How the nine examples relate

Three foundations on real experiments, then six ways an online experiment fails without any identification problem at all. The split is not cosmetic: the first three could only be done on real data, and the last six could only be done by simulation.

Two threads run the length of the group and are worth following deliberately. The first is estimand discipline: intention-to-treat against complier effect in example 3, direct against total effect in example 7, and testing against deciding in example 8 are all the same instruction — work out which quantity the decision needs before choosing the design, because no analysis recovers an estimand the design threw away. The second is that every claim here is checked against a constructed truth. That is why six of nine are simulations: a real dataset can show a standard error changing, but only a known answer can show which one was right.

Randomized Experiments & Randomization Inference

The Rubin causal model on a simulation where the truth is known: with a confounder driving assignment the naive difference in means reads 9.21 against a true effect of 4.0; randomize the same units and it averages 4.00 over 4,000 re-randomizations. Then Fisher’s exact randomization test, on his own examples. The Lady Tasting Tea gives 1/70 = 0.0143 from the hypergeometric null, matching scipy. Darwin’s maize is settled by enumerating all 32,768 sign patterns rather than sampling them: cross-fertilized plants run 2.617 inches taller, 13 of 15 pairs favour them, 863 patterns are at least as extreme, so p = 0.02634 — against the paired t-test’s 0.0497, the benchmark Fisher used the data to justify. Neyman then supplies what Fisher does not: an interval, [0.229, 5.004], whose t-based version clears zero by four thousandths of an inch.

View example →

Covariate Adjustment — ANCOVA, CUPED and the Freedman–Lin Debate

Randomization makes the estimate unbiased; it says nothing about how sharp it is. Adjustment buys precision, and buys exactly ρ² of it — a law tested at both ends on two real experiments. The Electric Company’s pre-test correlates 0.883 with the outcome, so adjustment cuts the sampling variance 79% (SE 2.537 → 1.169) while the estimate barely moves. Social Pressure (a get-out-the-vote mailing to 305,866 voters) has a baseline correlating 0.163, and buys 3%. The value of a covariate is therefore knowable before you use it. Includes what the shift from +5.66 to +4.73 actually is — chance imbalance being corrected, not bias introduced — and Freedman’s objection to plain ANCOVA with Lin’s centred, interacted fix, which the per-grade slopes show is not a formality.

View example →

Noncompliance and Cluster Designs — Intention-to-Treat, Complier Effects and Clustered Standard Errors

Two ways a real trial departs from the clean case, both on real data. In the Sommer–Zeger vitamin A trial a fifth of the treatment arm never took the vitamin, and the four available estimates split two and two: ITT −2.58 per 1,000 and CACE −3.23 are valid answers to different questions, while as-treated (−6.47) and per-protocol (−5.15) answer neither — non-takers died at 14.06 per 1,000 against takers’ 1.24, so almost all of that gap is who takes vitamins. Bloom’s estimator and two-stage least squares agree exactly, because a noncompliant trial is the textbook instrument. Then Project STAR: the same +5.82 effect, but a cluster-robust interval 1.8× wider — and a known-truth simulation showing why, with naive 95% intervals covering only 54% of the time.

View example →

A/B Testing — Sequential Monitoring and Power

The design is settled; what breaks is the monitoring. On an A/A test where every rejection is a false positive by construction, one look gives the nominal 4.8% and twenty looks give 24% — still climbing at fifty, because a running z-statistic is a random walk and will cross any fixed boundary if you wait. The classical fix is to commit to n in advance, where sample size scales as 1/δ² (halve the effect, quadruple the users: 1,570 → 6,279 per arm). The modern fix removes the tension entirely: the mixture SPRT is a martingale whose Ville bound holds across all sample sizes at once, so it held Type-I at 3.6% while checking every single observation — and still stopped at a median 1,454 per arm, earlier than the fixed horizon. The R side adds the pharma answer: O’Brien–Fleming boundaries costing just 2.8% more sample.

View example →

A/B Testing — Ratio Metrics and the Delta Method

Experiments randomize users; product metrics are computed over sessions. Sessions from one user are correlated, so a session-level standard error is too small — here 0.0041 against an honest 0.0050. The cost is not abstract: over 1,500 simulated experiments the naive 95% interval covered 87.4%, so about one experiment in eight makes a confidence statement that is simply false, and every failure is in the direction of a false positive. The delta method computes the ratio’s variance at the user level and restores 95.0%; a cluster bootstrap that assumes no linearization agrees to four decimals. Closes with the sample-ratio-mismatch guardrail, where a split of 50,600/49,400 — six-tenths of a point, invisible by eye — already gives p = 0.00015.

View example →

A/B Testing — Interference and Switchback Designs

Every method before this leaned on SUTVA — no interference between units; marketplaces and networks break it. Randomizing users within a market holds the treated fraction near 0.5 in both arms, so the spillover term cancels in the difference — the design subtracts it away. With a true direct effect of 1.00 and spillover −0.60, the naive A/B returns 0.988 against a true policy effect of 0.40. The dangerous part is the SD of 0.019: the estimate is not noisy, it is tight around the wrong number, so every internal diagnostic looks excellent. Cluster randomization recovers 0.402 — removing a bias of 0.588 for 0.002 in SD, a trade far more lopsided than the bias-variance framing suggests. Switchbacks do the same for temporal carryover, but only if the window exceeds it: at window = 1 the estimate is 0.998, the naive blind spot reproduced in the time domain.

View example →

A/B Testing — Multiple Comparisons and the False Discovery Rate

Peeking is multiplicity across time; a metric dashboard is the same arithmetic across metrics. With 200 metrics of which 160 are null, the family-wise error rate is 1.000 — not approaching certainty but at it, so every experiment yields at least one spurious winner. Bonferroni and Holm bound that at 5%, and pay for it: power falls to 0.26, about one true mover in four. Benjamini–Hochberg asks the better question for a screen — what fraction of my reported winners are spurious — and holds the FDR at 0.040 while recovering 61% of true movers, 2.4× the power. Bonferroni’s 0.6% false-discovery rate is not a better result; it is the symptom of a method that rejects almost nothing.

View example →

Bayesian A/B Testing and Multi-Armed Bandits

Experimentation as a decision rather than a test. The Beta–Binomial posterior gives Pr(B > A) = 0.911 and an expected loss of shipping B of 0.00047 against 0.01583 for shipping A — so the rule ships B even though the 95% credible interval contains zero. That disagreement is the point: the question is not “am I sure” but “what does being wrong cost”. Then the caveat that gets oversold everywhere — a stop-when-Pr(B>A)>0.95 rule declares a winner 58% of the time on identical arms, worse than the frequentist peeking it was meant to avoid, because the inflation belongs to stopping on a wandering statistic and not to a philosophy. Closes with Thompson sampling cutting regret 3.2× and routing 68% of traffic to the best of four arms unprompted.

View example →

A/B Testing — ML-Based Variance Reduction (CUPAC)

CUPED’s variance reduction is capped at the ρ² of a single covariate. CUPAC keeps the mathematics and removes the ceiling: use a cross-fitted ML prediction of the outcome, built from all pre-experiment features, as the control variate. On an outcome driven nonlinearly by six features, the best single covariate reaches 34% while the booster reaches 75% — halving the standard error from 0.0697 to 0.0352, which by the power arithmetic is worth quadrupling the sample, from data you already had. It holds under one rule: the control variate must use pre-treatment features only, or it absorbs part of the effect itself — the mediator trap in ML costume. Verified across repeated experiments: bias +0.002, coverage 0.94, and only the spread changes.

View example →