White (1982) establishes the asymptotic theory of maximum likelihood (ML) estimation when the assumed parametric model is misspecified. The key result is that ML converges not to the true parameter but to the pseudo-true parameter — the point that minimises the Kullback-Leibler (KL) divergence between the true density and the assumed family. The quasi-maximum likelihood estimator (QMLE) is consistent for and asymptotically normal, but with a sandwich covariance matrix that differs from the standard inverse-Hessian formula valid only under correct specification. The paper provides the foundation for all robust likelihood-based inference in econometrics.
Pseudo-true parameter. Let be the true density and the misspecified family. Under regularity, the quasi-MLE converges almost surely to: where denotes expectation under . This is the best approximation to the truth within the assumed family, measured by KL divergence. Under correct specification (the true parameter), but under misspecification may differ substantially.
Asymptotic normality of QMLE. Under regularity conditions: where (expected Hessian, negative definite) and (outer product of scores). The sandwich matrix is the correct covariance even under misspecification.
Information matrix equality under correct specification. When , the information matrix equality holds: . The sandwich covariance reduces to — the standard inverse-Hessian (or inverse-Fisher-information) formula. Misspecification is thus detectable by testing ; this is the information matrix test.
Misspecified regression example. For a linear regression with non-normal but mean-zero and homoskedastic: the QMLE (= ordinary least squares, OLS) is consistent for (so ), but the standard OLS covariance is wrong unless the true distribution is normal. The sandwich formula reduces to White's (1980) heteroskedasticity-consistent (HC) estimator in the heteroskedastic case.
Misspecified dynamic models. For time series, White extends the results to stationary ergodic processes, providing the foundation for QMLE inference in GARCH, ARCH, and other volatility models estimated under the false assumption of normality.
Practical implication. A QMLE (e.g., Gaussian QMLE for GARCH) is consistent for the conditional mean and variance parameters even when the true innovation distribution is non-normal — provided the conditional mean and variance are correctly specified. Inference uses the sandwich covariance , not the standard inverse Hessian.
"We show that the QMLE is consistent for a pseudo-true parameter, that is, the parameter that minimizes the Kullback-Leibler information criterion." (p. 1)
"Misspecification of the model in general leads to inconsistency of the usual ML estimator for the parameters of the true data generating process, but the QMLE is consistent for its pseudo-true value." (p. 2)
The foundational paper for robust likelihood inference. The sandwich covariance result is used everywhere GARCH models are estimated (Bollerslev-Wooldridge 1992 extended it to the GARCH setting) and underpins essentially all "robust standard errors" in time-series econometrics. The KL-divergence framing of pseudo-true parameters is conceptually clean and connects ML estimation to information theory. The main limitation — noted by White himself — is that pseudo-true parameters may not have structural economic meaning; they are artefacts of the misspecification rather than objects of interest. This connects directly to the Albert-Dodd (2004) finding in latent-class models: pseudo-true parameters from different misspecified families can diverge substantially even when the KL loss is nearly identical across models.