Summary
A case-study comparison of three model families for longitudinal binary outcomes, applied to self-reported regular smoking in the Victorian Adolescent Health Cohort Study (waves 1992–1995). The three models are: a semiparametric generalized estimating equations (GEE)/marginal model, a standard multilevel logistic-normal model (subject-specific), and a discrete latent-class mixture model. Model checking via Gelman's posterior predictive distributions shows the mixture model captures bimodal smoking trajectories that the other two miss.
Key Claims
- Marginal vs. subject-specific effects. For binary outcomes under a logistic link, population-averaged (marginal) effects from GEE and individual-level (subject-specific) effects from the logistic-normal model differ by a shrinkage factor proportional to the random-effect variance. They coincide only under a linear (identity) link. Practitioners must choose which estimand is of interest before selecting a model.
- Three model families. (1) Marginal/GEE: estimates population-averaged odds-ratio effects; does not assume a random-effects distribution; robust to misspecification of within-subject correlation structure. (2) Logistic-normal: estimates subject-specific log-odds trajectories; integrates over a normal random intercept; assumes normal heterogeneity. (3) Discrete mixture: finite mixture of latent trajectory classes with class-specific fixed effects; nonparametric approximation to the random-effects distribution.
- Posterior predictive model checking. Replicate the data many times from the fitted posterior predictive distribution; compare observed summary statistics (e.g. marginal smoking rate by wave, proportion always-smoking) against the distribution of replicated statistics. Extreme observed values relative to replicated distribution indicate model failure.
- GEE and logistic-normal fail the predictive check. Neither can reproduce the observed proportion of persistent non-smokers or the heavy-smoking tail; both yield replicated smoking rates that are too diffuse at the individual level. The discrete mixture with three latent classes passes all checks.
- Identifiability of discrete mixture. With K=3 latent classes the model is identified from the longitudinal structure; class membership probabilities and class-specific trajectories are estimated jointly via expectation-maximization (EM) or Bayesian simulation.
Concepts Introduced or Extended
Entities Mentioned
Quotes
"The appropriate application and interpretation of these models remains somewhat unclear, especially when compared with the computationally more straightforward semiparametric or 'marginal' approach."
My Take
This is a model-selection paper in applied biostatistics, not econometrics. Its relevance to the wiki is methodological: the marginal-vs-subject-specific distinction is directly analogous to random-effects vs. population-averaged estimation in panel probit models (echoing Rossi-Allenby 2003 on individual vs. aggregate effects), and the posterior predictive checking methodology is Gelman-Meng-Stern (1996) applied systematically. The discrete mixture "wins" here because smoking trajectories are genuinely bimodal, not because mixtures are universally superior; the lesson is that predictive checks should drive model selection rather than theoretical elegance alone.