Summary
Ni and Sun (2005) compare Bayesian vector autoregression (VAR) estimators across six prior combinations and two loss functions for coefficients, measuring performance by frequentist (simulation-based) risk. The paper introduces a shrinkage prior on VAR slope coefficients Φ — a two-level hierarchical prior that integrates to πS(ϕ)∝∥ϕ∥−(J−2) — and an asymmetric LINEX (linear-exponential) loss function. Both improve over constant priors and quadratic (posterior-mean) estimators, especially near unit roots. A Student-t error extension with latent scale variables is derived via Gibbs sampling. The main finding is that prior choice for Φ matters more than loss function choice for frequentist risk.
Key Claims
- The shrinkage prior πS(ϕ)∝∥ϕ∥−(J−2), J=p(Lp+1), dominates constant priors (Jeffreys, RATS [Regression Analysis of Time Series], reference) in frequentist risk; implemented as a two-level hierarchy (ϕ∣δ∼NJ(0,δIJ), π(δ)∝1) that integrates to the shrinkage marginal; δ∣ϕ,Σ,Y∼IG(J/2−1,21ϕ′ϕ).
- LINEX loss ϕ^ij=−(1/aij)logE{exp(−aijϕij)∣Y} with aij=−4 for non-intercept coefficients corrects the downward bias of the posterior mean under near-unit-root dynamics.
- Yang-Berger reference prior πR(Σ)∝∣Σ∣−1∏i<j(λi−λj)−1 dominates Jeffreys (πJ∝∣Σ∣−(p+1)/2) and RATS (πA∝∣Σ∣−(L+1)p/2−1) priors for Σ under all three Σ-loss functions.
- Prior choice for Φ dominates loss function choice in determining frequentist risk.
- Best combination overall: shrinkage prior for Φ + Yang-Berger reference prior for Σ.
- Student-t VAR: latent scale variables qt∼Gamma(ν/2,ν/2) give εt∣qt∼Np(0,qt−1Σ); the ν conditional is log-concave → Gilks-Wild (1992) adaptive rejection sampling.
- Simulation (p=5, L=1, T=50, N=1,000, B1=I5 unit root data generating process (DGP)): shrinkage+reference dominates all others; maximum likelihood estimator (MLE) largest eigenvalue avg 0.960 vs shrinkage+reference 0.965; impulse response z1,(5,5)=0.605 (shrinkage reference) vs 0.397 (constant/MLE).
- U.S. macro application (p=6, L=2, T=172, Schwarz selection): posterior risks 0.123/6.470 (shrinkage reference) vs 0.126/13.703 (constant RATS); GDP-to-inflation impulse response function (IRF) qualitatively inverts under shrinkage prior.
Concepts Introduced or Extended
Entities Mentioned
Quotes
"Simulation results show that the shrinkage prior is superior to the constant prior and the reference prior is superior to both the Jeffreys and RATS priors. In general, the choice of priors is more important than the choice of loss functions."
My Take
This paper is a direct companion to Sun-Ni (2005) (Journal of Statistical Planning and Inference), which established reference prior dominance for Σ under quadratic loss with a constant Φ prior. The Journal of Business & Economic Statistics (JBES) paper adds the crucial missing pieces: a proper shrinkage prior for Φ (analogous to ridge regression but Bayesian, implementable via a two-level hierarchy), LINEX loss suited to near-unit-root estimation, and a heavy-tailed error model. The frequentist risk framing is unusual for Bayesian work and makes the dominance results unusually clean. The most striking practical result — that GDP-to-inflation IRFs change sign under the shrinkage prior — should prompt caution about interpreting structural VARs estimated with constant diffuse priors.