Brown-Greene-Harris (2014) A New Formulation for Latent Class Models

latent-classmixture-modelordered-probitmodel-selectionhealth-economicsconsumer-heterogeneityinformation-criteriaiia

Summary

The paper reforms the standard latent class (finite mixture) model around an observation about how it is actually used: researchers almost always rank and label the estimated classes ex post by their class-specific expected values ("low users" vs. "high users"), yet the model itself is estimated with class probabilities that ignore that ordering. Brown, Greene and Harris propose (i) enforcing the expected-value ordering across classes during estimation through an exponential-increment recursion on the latent indices, and (ii) replacing the standard multinomial-logit (MNL) class-assignment probabilities with a more parsimonious ordered (probit/logit) specification consistent with that ranking. On UK BMI data the MNL formulation supported only two classes while the ordered formulation supported up to five.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"A priori, nothing is known about the behaviours within each class; but ex post, researchers invariably label the classes according to expected values, however defined, within each class."

"We enforce ordering in the EVs across classes and suggest an ordered probabilistic specification for the class assignment probabilities."

My Take

The insight is small but genuinely useful: the field's universal ex post ritual — sort the classes, call them "low/high", interpret the gradient — is smuggled in after estimation, and the paper simply moves it into estimation, where it buys both interpretability (the exp()\exp(\cdot) increments are differential effects) and parsimony (one ordered index instead of a fan of MNL vectors that the information criteria then punish). It is essentially a targeted fix to the latent class model's worst practical failure mode — being unable to fit more than a couple of classes before the covariate-laden MNL assignment equation stops converging. The cost is that it only helps when class membership depends on covariates and the outcome carries a real ordinal/cardinal notion; for unordered typologies (the classic diagnostic-accuracy or brand-choice use) it collapses to a reparameterisation and buys nothing. The Monte Carlo is thin (100 reps, results "available on request"), so I read the empirical BMI 2-vs-5-class contrast as the load-bearing evidence rather than the simulation.