Summary
Kleibergen and Zivot (2003) establish the precise Bayesian analogues of the classical two-stage least squares (2SLS) and limited information maximum likelihood (LIML) estimators by analyzing four diffuse-prior Bayesian approaches to the instrumental variable (IV) regression model. The central result is that the Jeffreys prior on the restricted reduced form (RRF) parameters yields a posterior for structural parameters whose functional form is identical to the exact finite-sample density of the LIML estimator. The Drèze (1976) and Bayesian two-stage (B2S) approaches, by contrast, have more in common with 2SLS than with LIML. The paper uses a reduced-rank-restriction representation of the IV model — expressing the RRF as a rank constraint on the unrestricted reduced form (URF) — to derive all results, and analyzes behavior under weak instruments for each prior.
Key Claims
- Main theorem: The Jeffreys prior (proportional to the square root of the information matrix determinant) applied to the IV model's RRF parameters gives a posterior for the structural parameter β that is identical in functional form to the exact small-sample density of the LIML estimator (proved via Kleibergen 2000a's orthogonal parameterization). This holds for both exactly- and over-identified models.
- Drèze (1976) approach specifies a flat prior on structural form (SF) parameters. Its marginal posterior of β is a mix of the LIML concentrated likelihood and an ordinary least squares (OLS)-centered Student-t kernel. It is not invariant to the ordering of endogenous variables in over-identified models — the same failure as 2SLS — and becomes more sensitive to spurious instruments as the degree of over-identification grows.
- Bayesian two-stage (B2S) approach is explicitly constructed to mimic 2SLS and behaves like it: posterior mode moves toward the asymptotic-bias point of 2SLS as spurious instruments are added; posterior tails are slightly thinner than Drèze.
- Diffuse prior on the URF — treating all reduced-form parameters symmetrically without imposing the rank restriction — has some LIML-like properties but is less well-behaved than the Jeffreys prior.
- Implied priors on the URF: For each diffuse prior on the RRF, the paper derives the implied prior on the URF parameter matrix Π via the Jacobian of the URF→RRF transformation. The Jeffreys prior implies a flat URF prior in the just-identified case, and a prior favoring large reduced-form coefficients in the direction of spurious instruments in the over-identified case — this "pretesting" structure explains LIML's robustness.
- Weak instruments: Under weak instruments (Π≈0), the Jeffreys posterior and LIML density become bimodal — one mode near the OLS limit, one near the IV estimate. The Drèze posterior is even more bimodal. Both the Jeffreys and flat-URF posteriors behave better than the Drèze and B2S posteriors when spurious instruments are added.
- Moments: The Drèze posterior of β has moments up to (but not including) the degree of over-identification d, mirroring the 2SLS finite-sample distribution. The Jeffreys posterior/LIML density has Cauchy-type tails and no finite moments, but is approximately median-unbiased with strong instruments.
- Ordering invariance: LIML and the Jeffreys posterior are invariant to the ordering of endogenous variables (β^LIML−1=β^LIML∗ where β∗=β−1); 2SLS and Drèze are not.
- Fuller's modified LIML and other k-class estimators are noted as alternatives with finite moments but not analyzed in detail.
Concepts Introduced or Extended
- Instrumental Variables — provides Bayesian foundations for classical IV estimators; LIML = Bayesian estimation under Jeffreys prior; Drèze/2SLS = Bayesian estimation under less coherent diffuse priors
Entities Mentioned
Quotes
"The approach based on the Jeffreys prior is the Bayesian counterpart of LIML and the approach using a diffuse prior on the unrestricted reduced form also has some properties in common with LIML."
"We find that the first two Bayesian procedures have more in common with 2SLS than with LIML, the approach based on the Jeffreys prior is the Bayesian counterpart of LIML."
My Take
This is a technical reference paper rather than an applied empirical contribution — its value to the wiki is as the theoretical foundation explaining why LIML is preferred over 2SLS under weak instruments. The main result (Jeffreys prior → LIML) gives LIML a principled Bayesian justification: it is the estimator that emerges from prior ignorance when you parameterize ignorance coherently. The paper also illuminates why the Drèze approach — long presented as the Bayesian version of LIML — actually behaves more like 2SLS: it is sensitive to over-identification in the same way 2SLS is.
Practically, the paper strengthens the applied case for LIML or Fuller's modified LIML over 2SLS in settings with many instruments or weak instruments — exactly the settings common in applied DI and labor economics research in this wiki. The ordering invariance result is also useful: LIML gives the same answer regardless of which endogenous variable is treated as "the" outcome, while 2SLS does not.
Limitation: the analysis is confined to the normal/homoskedastic independent and identically distributed (iid) errors case. Extensions to heteroskedasticity, clustering, and generalized method of moments (GMM) are not covered here (those are in the subsequent Kleibergen 2005 Econometrica paper on the Kleibergen-Paap (KP) rk statistic).