Goodman (1974) introduces a unified maximum likelihood (ML) estimation algorithm for latent class models with polytomous manifest variables and latent classes — an iterative procedure mathematically equivalent to the expectation-maximisation (EM) algorithm, three years before Dempster-Laird-Rubin (1977) formalised the general framework. The paper also provides a practical identifiability diagnostic (a Jacobian rank test evaluated at the maximum likelihood estimate (MLE)), unifies unrestricted and restricted latent structures under a single estimation framework, and proposes an exploratory strategy that uses unidentifiable models as stepping stones toward parsimonious identifiable ones.
Model: For manifest polytomous variables and a -class latent variable , cell probabilities factor as: Within each latent class the manifest variables are mutually independent (local independence). Free parameters: for the basic set; free cell probabilities: .
Estimation algorithm (EM): Starting from initial trial values , iterate:
Necessary condition for identifiability (counting): The number of free cell probabilities must not be less than the number of free structural parameters: When violated, parameters are structurally unidentifiable.
Local identifiability rank test: Form the Jacobian matrix of with respect to the basic parameter set. Parameters are locally identifiable iff this matrix has full column rank. Goodman applies this test at (the MLE), diagnosing whether the estimated model is identified — a practical advance over prior literature, which only checked the population .
Degrees of freedom: Under a locally identifiable -class structure, (likelihood ratio) is asymptotically with .
Restricted latent structures: Equality constraints across latent classes reduce parameter count and can restore identifiability. Three key collapsibility conditions (classes 1 and 2 identical on all , , or manifest variables) cause unidentifiability. For restricted structures the Jacobian has merged columns; local identifiability requires the modified matrix to have full column rank.
Exploratory strategy: Fit the unrestricted -class model. If unidentifiable, the expected cell frequencies are still uniquely determined and can guide the choice of restrictions. Impose restrictions to obtain an equivalent identifiable model yielding the same .
Application 1 — Stouffer-Toby role conflict (, table): 2-class model fits excellently (, ); Stouffer-Toby originally concluded 5 classes were needed. Modal latent class (): intrinsically particularistic except on variable .
Application 2 — Coleman schoolboys (, table): One latent variable: , (very poor fit). Two latent variables (crowd membership + attitude , a restricted 4-class model): , (excellent fit). The dramatic improvement illustrates the value of multi-latent-variable restricted structures.
"For each of the models considered here, a relatively simple method is presented for calculating the maximum likelihood estimate of the frequencies in the -way contingency table expected under the model, and for determining whether the parameters in the estimated model are identifiable." (p. 215)
A foundational paper for latent class analysis that anticipates the EM algorithm three years before Dempster-Laird-Rubin (1977). The iterative procedure is derived directly from the ML equations rather than from an abstract incomplete-data framework, but the mathematics is identical. The identifiability rank test applied at the MLE is a practical innovation: it distinguishes models that are theoretically identified but empirically unidentified (flat ridges in the likelihood). The most influential methodological insight is that fitting unidentifiable models is productive — is uniquely determined regardless of identifiability, and its structure reveals the restrictions needed. The two applications are well chosen: one where parsimony wins (2 classes beat Stouffer-Toby's 5), one where structural complexity is genuinely needed (2 latent dimensions beat 1). The restricted-structure taxonomy in §§4–5 anticipates modern confirmatory latent class analysis and connects directly to the collapsibility concerns in Albert-Dodd (2004): near-identical from structurally distinct models is precisely why diagnostic accuracy estimates diverge.