Griffiths-Skeels-Chotikapanich (2002) Sample Size Requirements for Estimation in SUR Models

bayesiansursample-sizemaximum-likelihoodfeasible-glsnoninformative-priorfinite-samplematrix-rank

Summary

Demonstrates that the standard textbook requirement for estimating Seemingly Unrelated Regressions (SUR) models — Tmax(M,kmax+1)T \geq \max(M, k_{\max} + 1) — is both incomplete and potentially seriously misleading. Derives necessary conditions separately for two-stage feasible generalized least squares (FGLS) (Theorem 1: TM+ρηT \geq M + \rho - \eta) and for maximum likelihood (ML)/Bayesian estimation (Theorem 2: TM+ρT \geq M + \rho), where ρ=rank([X1,,XM])\rho = \text{rank}([X_1,\ldots,X_M]) and η\eta captures how many equations have exclusive regressors. The ML/Bayesian requirement is always at least as stringent as two-stage, and strictly more stringent when equations have partially distinct regressors — because ML/Bayes minimizes a degree-2M2M polynomial in β\beta rather than a quadratic. Empirically confirmed via experiments where ML/Gibbs sampler broke down at M7M \geq 7 while two-stage worked up to M=18M=18 (T=19T=19, Kj=3K_j=3 per equation).

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"It is somewhat surprising, therefore, that the sample size requirements for joint estimation of the parameters in this model do not appear to have been correctly stated in the literature."

"For a 10-equation model with 3 distinct explanatory variables and a constant in each equation, two-stage estimation requires T11T \geq 11 … whereas ML and Bayesian estimation need T>40T > 40."

My Take

A short, rigorous correctness paper with immediate practical implications. Theorem 2 is the critical result for Bayesian practitioners: any SUR system with T<M+ρT < M + \rho produces an improper posterior under the noninformative prior, but standard Markov chain Monte Carlo (MCMC) may not detect this (the Gibbs sampler may simply get stuck or produce garbage). The two-stage bound (Theorem 1) is looser but depends on the overlap structure ω\omega in a way that rewards heterogeneous regressor design. The paper motivates careful pre-flight checking of rank conditions before any SUR estimation.