Summary
Geweke (1993) develops exact Bayesian inference for the linear model with i.i.d. Student-t disturbances by exploiting the Student-t = scale-mixture-of-normals equivalence: yi∣X∼t(xi′β,σ2;ν) is identical to yi∣X∼N(xi′β,σ2ωi) with ν/ωi∼χ2(ν) and n latent scale weights ωi. Augmenting with ω renders all full conditionals conjugate and enables a direct four-block Gibbs sampler. Applied to 14 Nelson-Plosser (1982) macroeconomic series, posterior odds strongly favor the Student-t over the normal for 13/14 series; lower degrees of freedom systematically reduce the posterior odds in favor of difference stationarity.
Key Claims
- Student-t = scale mixture of normals: yi∣X∼N(xi′β,σ2ωi) with ν/ωi∼χ2(ν) (Section 2.1, eq. 6) — augmenting with ω makes all full conditionals conjugate.
- Four-block Gibbs (Section 3): (1) ωi∣(β,σ,ν) — Inverse-Gamma: [(σ−2ui2+ν)/ωi]∣(β,σ)∼χ2(ν+1) (eq. 15); (2) σ2∣(β,ω) — Inverse-Gamma: [∑iui2/ωi]/σ2∼χ2(n) (eq. 14); (3) β∣(σ,ω) — generalized-least-squares (GLS)-form Normal β^(ω)=(X′Ω−1X)−1X′Ω−1y (eq. 12); (4) optional ν∣(β,σ,ω) — non-standard kernel (eq. 18) ∝(ν/2)nν/2Γ(ν/2)exp(−ην) with η=(1/2)∑i[logωi+ωi−1]+λ, requiring specialized sampling.
- Posterior odds ratio (eq. 16): POR(ν(1),ν(2)) is an expectation under the ν(2) posterior, computable from standard Gibbs output; limν(1)→∞ gives the odds ratio in favor of normality.
- Unknown ν: flat π(ν)∝const (ν>0) forces normality as the limiting case; exponential prior π(ν)=λexp(−λν) (eq. 17) is the recommended proper prior.
- Convergence: Theorem 4 (Gibbs chain converges in distribution) and Theorem 5 (a.s. ergodic convergence m−1∑g(θ(j))→Ep[g(θ)]) proved for fixed ν and for exponential prior.
- Posterior moments: E[β] exists for ν>2, Var(β) for ν>4 (Theorem 3, improper flat prior).
- Nelson-Plosser (1982) application: 13/14 macroeconomic time series favor Student-t over normal (Table I); posterior mean of ν typically 3–7 (Table II); lower ν systematically reduces posterior odds in favor of difference over trend stationarity (Table V) and shrinks posterior standard deviation (SD) of trend coefficient δ.
Concepts Introduced or Extended
- Mixture of Normals — Student-t = scale-mixture-of-normals representation; ωi latent scale weights as data augmentation device
- Gibbs Sampler — four conjugate blocks via scale-mixture augmentation; convergence Theorems 4 and 5
- Unit Root Inference — tail assumption sensitivity for Nelson-Plosser unit root conclusions
- Stochastic Volatility — ωi latent weights are the static precursor to the dynamic ht in stochastic-volatility (SV) models
Entities Mentioned
Quotes
"The main contribution is to provide a simple and stable computational method for full Bayesian inference in the independent Student-t linear model."
"For all but one of these series posterior odds ratios favour Student-t linear models with degrees of freedom in the range of 3 to 7, over normal linear models."
My Take
The canonical reference for the ωi latent-weight data augmentation device in Bayesian regression with leptokurtic disturbances. The same structure reappears in Jacquier-Polson-Rossi (1994), where ωt becomes time-varying and follows a log-AR(1) (first-order autoregressive) process — the SV model. Geweke-Keane (1999) extends the scale-mixture idea to the probit context. The Nelson-Plosser application is a useful reminder that unit root conclusions are not robust to tail assumptions, a point often neglected in applied work that assumes Gaussian disturbances.