White (1982) Maximum Likelihood Estimation of Misspecified Models

model-misspecificationquasi-mlepseudo-true-parametersmaximum-likelihoodeconometricsspecification-testing

Summary

White (1982) establishes the asymptotic theory of maximum likelihood (ML) estimation when the assumed parametric model is misspecified. The key result is that ML converges not to the true parameter but to the pseudo-true parameter θ\theta^* — the point that minimises the Kullback-Leibler (KL) divergence between the true density and the assumed family. The quasi-maximum likelihood estimator (QMLE) is consistent for θ\theta^* and asymptotically normal, but with a sandwich covariance matrix that differs from the standard inverse-Hessian formula valid only under correct specification. The paper provides the foundation for all robust likelihood-based inference in econometrics.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"We show that the QMLE is consistent for a pseudo-true parameter, that is, the parameter that minimizes the Kullback-Leibler information criterion." (p. 1)

"Misspecification of the model in general leads to inconsistency of the usual ML estimator for the parameters of the true data generating process, but the QMLE is consistent for its pseudo-true value." (p. 2)

My Take

The foundational paper for robust likelihood inference. The sandwich covariance result is used everywhere GARCH models are estimated (Bollerslev-Wooldridge 1992 extended it to the GARCH setting) and underpins essentially all "robust standard errors" in time-series econometrics. The KL-divergence framing of pseudo-true parameters is conceptually clean and connects ML estimation to information theory. The main limitation — noted by White himself — is that pseudo-true parameters may not have structural economic meaning; they are artefacts of the misspecification rather than objects of interest. This connects directly to the Albert-Dodd (2004) finding in latent-class models: pseudo-true parameters from different misspecified families can diverge substantially even when the KL loss is nearly identical across models.