Pendergast et al. (1996) A Survey of Methods for Analyzing Clustered Binary Response Data

geeclustered-datacorrelated-binarylongitudinal-datarandom-effectspanel-dataoverdispersionliterature-survey

Summary

A comprehensive survey of statistical methods for analyzing correlated binary responses arising from clustered designs — repeated measures on the same subject, bilateral data (e.g., two eyes), family studies, and cluster-randomized trials. The paper organizes the methods landscape into four families: naive/response-feature approaches (ignore correlation), conditionally specified models (specify joint distribution via pairwise terms), marginal models (generalized estimating equations, GEE), and cluster-specific random-effects models. The central tension is between marginal population-average coefficients (GEE) and cluster-specific subject-specific coefficients (random effects), which diverge in magnitude for nonlinear link functions. GEE, introduced by Liang and Zeger (1986), is highlighted as the most robust and widely applicable marginal approach.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

None of the six authors has a dedicated wiki entity page.

Quotes

"A central distinction in clustered data analysis is between population-averaged and cluster-specific models. The choice depends on the scientific question: population-averaged models answer questions about average effects in the population; cluster-specific models answer questions about effects for a particular individual or cluster."

My Take

A well-organized survey that established the four-family taxonomy now standard in biostatistics textbooks (Diggle-Liang-Zeger 1994; Fitzmaurice-Laird-Ware 2004). The GEE exposition is clear and the marginal vs. cluster-specific distinction is explained more carefully than in most applied references. The goodness-of-fit gap for GEE remains a genuine limitation; the paper is honest about it. The paper predates Bayesian random-effects approaches (Zeger-Karim 1991 used Gibbs sampling but GEE2 for correlation) and does not discuss Markov chain Monte Carlo (MCMC)–based full-likelihood inference for clustered binary data, which has since become practical.