Summary
A didactic discussion-paper treatment of latent class analysis (LCA) as a discrete analogue of factor analysis, aimed at social-science / geography researchers whose data are categorical rather than metric. Ernste and Fischer lay out the standard latent class model for a multi-way contingency table under local independence, its maximum-likelihood estimation and identifiability conditions, and the exploratory vs. confirmatory distinction, then illustrate the method as a tool for typological analysis — deriving a six-class typology of Swiss permanent residents from five categorical demographic indicators.
Key Claims
- LCA = discrete factor analysis. Latent class models reduce many manifest discrete variables to one (or few) latent discrete variable(s); conditional probabilities πit(A∣X) play the role of factor loadings and latent class probabilities πtX describe the class distribution.
- Local independence is the defining assumption. For manifest variables A,B,C and latent class X with T categories, πijktABCX=πit(A∣X)πjt(B∣X)πkt(C∣X)πtX — the latent variable fully accounts for the manifest associations, so the manifest variables are conditionally independent given the class.
- ML estimation via IPF / scoring. The likelihood equations (Goodman 1979, Haberman 1979) are solved by iterative proportional fitting or Fisher scoring; as a log-linear model with a composite link function (Thompson-Baker 1981), LCA is a special case of generalized linear models.
- Identification and multiple solutions. Unlike ordinary log-linear models, LC MLEs are not uniquely determined; a necessary-and-sufficient local-identifiability condition is that the ((I+J+K−2)T−1) matrix of partial derivatives of non-redundant probabilities has full column rank.
- Model selection by chi-square. Fit is judged by the Pearson X2 and the likelihood-ratio L2; L2 is preferred because it is partitionable (once T is fixed, hypotheses about the probabilities can be tested efficiently) — essential for the confirmatory mode.
- Exploratory vs. confirmatory. Exploratory LCA fits unrestricted models (no a-priori constraints); confirmatory LCA imposes size restrictions on conditional and/or latent class probabilities to test structural hypotheses, and allows multi-population comparison (Clogg-Goodman 1984/85).
- Empirical typology. On Swiss resident data (sex, household size, nationality, age, employment/vocational status), the 1-class independence model (L2=1189.5) and 2-class model (L2=616.4) are rejected; a six-class model fits (L2=169.8, 193 d.f., p>0.05), yielding interpretable types (working-age Swiss women; middle-aged part-time women; foreign full-time low-status men; elderly retired Swiss; schoolchildren living with parents; higher-status Swiss family fathers).
Concepts Introduced or Extended
Entities Mentioned
Quotes
"Latent class models may be considered as a discrete analogue of factor analysis."
"Essential to the standard latent class model is the assumption of local independence, i.e. that the latent variable completely accounts for the manifest variables."
My Take
This is a competent expository paper rather than a methodological advance — its value to the wiki is as a clean, self-contained statement of the classical Goodman-Haberman latent class machinery (local independence, IPF/scoring estimation, the full-rank Jacobian identifiability test, L2 partitioning, exploratory-vs-confirmatory) applied to a concrete typology, which usefully complements the diagnostic-accuracy framing in Albert-Dodd (2004) and the ordered-class reformulation of Brown-Greene-Harris (2014) already in Latent Class Model. The "six classes fit at p>0.05" model-selection ritual it performs is exactly the fragility later authors push back on: a chi-square-driven search over T that is sensitive to sample size and offers no guard against near-non-identifiability. Read as a period piece it also documents how latent class analysis entered social science / regional science (via Clogg, McCutcheon, Hagenaars) as the categorical counterpart to factor analysis.