Bertrand Duflo and Mullainathan 2003 — How Much Should We Trust Differences-in-Differences Estimates

difference-in-differencesserial-correlationstandard-errorsclusteringinferenceeconometricsmethodspanel-dataMonte-Carlo

Summary

Bertrand, Duflo, and Mullainathan (2003) demonstrate that conventional ordinary least squares (OLS) standard errors in Differences-in-Differences (DiD) regressions are severely biased downward due to serial correlation in outcomes, leading to rejection rates of 444467%67\% on randomly assigned placebo laws instead of the nominal 5%5\%. A survey of 9292 DiD papers (1990–2000) confirms almost none correct for this. The paper evaluates several fixes: parametric autoregressive (AR(kk)) corrections fail; clustering by group (state), pre/post aggregation, and block bootstrap all work, with clustering being the most widely adopted solution. Published in the Quarterly Journal of Economics (QJE) 119(1): 249–275 (2004).

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"Most papers that employ Differences-in-Differences estimation (DD) use many years of data and focus on serially correlated outcomes but ignore that the resulting standard errors are inconsistent."

"For each [placebo] law, we use OLS to compute the DD estimate of its 'effect' as well as the standard error of this estimate. These conventional DD standard errors severely understate the standard deviation of the estimators: we find an 'effect' significant at the 5 percent level for up to 45 percent of the placebo interventions."

"Collapsing the time series information into a 'pre' and 'post' period and explicitly taking into account the effective sample size works well even for small numbers of states."

My Take

This is one of the most influential applied econometrics papers of the 2000s — it changed how practitioners handle inference in DiD settings. The empirical evidence is striking because it is directly actionable: the paper not only identifies a near-universal flaw in the published DiD literature but ranks competing fixes. The recommended solution (clustering at the group level) became standard in virtually every DiD application since, and is now a default in most econometrics software. The paper's survey of 92 published papers showing 87/9287/92 either ignoring or inadequately correcting serial correlation is the most honest audit of a methodology in the literature. The one limitation: the paper focuses on standard error estimation assuming the parallel trends assumption holds; it does not address threats from differential pre-trends or treatment effect heterogeneity across groups — both of which became major topics in the subsequent DiD literature (Callaway and Sant'Anna 2021, de Chaisemartin and D'Haultfoeuille 2020).