Albert-Dodd (2004) A Cautionary Note on the Robustness of Latent Class Models for Estimating Diagnostic Error without a Gold Standard

latent-classdiagnostic-accuracymodel-misspecificationsensitivity-specificitybiostatisticssimulationcategorical-datamodel-selectionmixture-model

Summary

Albert and Dodd (2004) study four latent class models for estimating diagnostic accuracy (sensitivity/specificity) when no gold standard exists. They show that even when models fit the observed data equally well by every standard diagnostic, their estimated sensitivities can diverge by 0.2 or more. Misspecified maximum likelihood estimators (MLEs) converge to pseudo-true parameters that are nearly indistinguishable in likelihood yet yield very different sensitivity estimates, and the number of diagnostic tests required to distinguish models in practice is unrealistically large (J10J \geq 10).

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"We have shown that even when models for estimating diagnostic accuracy with no gold standard are not identifiable from the data, they are distinguishable from the true model with more raters."

"We caution that if such a study is not feasible, the best one can do is to fit all reasonable models and present sensitivity analyses."

My Take

A sobering identifiability paper. The core insight — that maximum likelihood can converge to a wrong but nearly indistinguishable pseudo-true parameter — is not new, but the concrete numerical illustration (0.2 divergence in sensitivity, expected log-likelihoods equal to 5 decimal places) makes the problem viscerally clear. The practical upshot is pessimistic: with J=5J = 5 raters you simply cannot know which latent class structure is correct. The three guidelines (use gold standard when possible; sensitivity analysis across models; collect J10J \geq 10 if no gold standard) are sensible but the third is rarely achievable.