Zeger-Liang-Albert (1988) Models for Longitudinal Data: A Generalized Estimating Equation Approach

generalized-estimating-equationslongitudinal-datageneralized-linear-modelrandom-effectsbiostatisticsquasi-likelihood

Summary

This paper draws the now-standard distinction between two ways to extend generalized linear models to longitudinal / correlated data and shows both can be fit by a generalized estimating equation (GEE) approach. Subject-specific (SS) models put the heterogeneity across subjects explicitly into the regression parameters (as in random-effects / mixed GLMs), so coefficients describe how a given individual's response depends on covariates. Population-averaged (PA, or marginal) models target the aggregate response of the population, so coefficients describe how the population-average response depends on covariates. The authors fit both classes for discrete and continuous outcomes via GEE, and show that when the subject-specific effects are Gaussian there are simple closed-form relationships between the PA and SS parameters. The methods are illustrated on longitudinal data for mothers' smoking and children's respiratory disease. (Biometrics 44(4): 1049–1060.)

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"Two approaches are considered: subject-specific (SS) models in which heterogeneity in regression parameters is explicitly modelled; and population-averaged (PA) models in which the aggregate response for the population is the focus."

"When the subject-specific parameters are assumed to follow a Gaussian distribution, simple relationships between the PA and SS parameters are available."

My Take

This is the paper that made "marginal vs. conditional" (population-averaged vs. subject-specific) a required distinction in longitudinal-data analysis, and its lasting practical message is that the attenuation between the two is real and computable, not a discrepancy to be explained away: a logistic random-effects coefficient and a GEE coefficient estimate different things and differ in magnitude by a factor governed by the random-effect variance. It complements Liang–Zeger's estimating-equation machinery by clarifying what the estimand is under each modeling choice, and it sits directly upstream of the GLMM vs. marginal-model debates and the Bayesian hierarchical treatments (e.g. Zeger–Karim's Gibbs sampler) that followed. The main caveat is that the clean PA–SS relationships rely on the Gaussian-mixing assumption; with non-Gaussian or misspecified random effects the exact correspondence breaks down.