Kadiyala-Karlsson (1993) Forecasting with Generalized Bayesian Vector Autoregressions

bayesianvarforecastingminnesota-priornormal-wishartenc-priornormal-diffusemonte-carloprior-comparison

Summary

Kadiyala and Karlsson compare five prior distributions for Bayesian vector autoregression (VAR) forecasting — the Minnesota prior, Normal-Wishart, Diffuse (Jeffreys'), Normal-Diffuse, and the Extended Natural Conjugate (ENC) prior — across three forecasting experiments on Canadian GNP/M2 data and US wheat export data. The central critique of the Minnesota prior is that it forces equation-independence and treats the residual variance-covariance matrix Ψ\Psi as known and diagonal — the wrong tradeoff, since practitioners can form prior beliefs about regression coefficients more easily than about Ψ\Psi. Priors that allow inter-equation dependence consistently match or beat the Minnesota prior; OLS is worst in every experiment.

Key Claims

Model and Notation

The Five Priors

Minnesota prior (Litterman 1986a): equations treated independently; γiN(γ~i,Σ~i)\gamma_i \sim N(\tilde\gamma_i, \tilde\Sigma_i) with Σ~i\tilde\Sigma_i diagonal and ψii\psi_{ii} fixed at 0.81si20.81s_i^2. Prior variances (eq. 2):

Σ~i=diag ⁣{π1k,π2si2ksj2,0.81π3si2}\tilde\Sigma_i = \text{diag}\!\left\{\frac{\pi_1}{k},\quad \frac{\pi_2 s_i^2}{k s_j^2},\quad 0.81\pi_3 s_i^2\right\}

for own-lag kk, cross-lag kk of variable jij\neq i, and exogenous coefficients respectively. Key limitation: treats Ψ\Psi as known and diagonal — the opposite of what practitioners know.

Normal-Wishart prior: γΨN(γ~,ΨΩ~)\gamma|\Psi \sim N(\tilde\gamma, \Psi\otimes\tilde\Omega), ΨIW(Ψ~,α)\Psi \sim IW(\tilde\Psi, \alpha) (inverse-Wishart, IW). Natural conjugate — posterior γΨ,yN(γˉ,ΨΩˉ)\gamma|\Psi,y \sim N(\bar\gamma, \Psi\otimes\bar\Omega), ΨyIW(Ψˉ,T+α)\Psi|y \sim IW(\bar\Psi, T+\alpha); marginal ΓMT(Ω~1,Ψ~,Γˉ,α)\Gamma \sim MT(\tilde\Omega^{-1}, \tilde\Psi, \bar\Gamma, \alpha) (matric-variate tt, MT). Allows non-diagonal Ψ\Psi and yields closed-form posterior moments. Limitation: ΨΩ~\Psi\otimes\tilde\Omega structure forces symmetric treatment of all equations, so own-lag prior variances differ from Minnesota.

Diffuse (Jeffreys') prior: p(γ,Ψ)Ψ(q+1)/2p(\gamma,\Psi) \propto |\Psi|^{-(q+1)/2}. Posterior centered on OLS: E(Γy)=Γ^E(\Gamma|y) = \hat\Gamma; moments analogous to Zellner seemingly unrelated regressions (SUR). Same Ψ(ZZ)1\Psi\otimes(Z'Z)^{-1} restriction as Normal-Wishart.

Normal-Diffuse prior (new): γN(γ~,Σ~)\gamma \sim N(\tilde\gamma,\tilde\Sigma), p(Ψ)Ψ(q+1)/2p(\Psi)\propto|\Psi|^{-(q+1)/2}. Combines a Minnesota-type prior on regression parameters with a diffuse prior on Ψ\Psi; allows non-diagonal Ψ\Psi and inter-equation dependence. Marginal posterior of γ\gamma is product of marginal prior and a matricvariate tt factor — can be bimodal when prior and likelihood strongly conflict.

ENC prior (new, Drèze-Morales 1976): reparameterizes the VAR with block-diagonal Δ\Delta (columns γi\gamma_i on diagonal). Prior: p(Δ)Ψ~+(ΔΔ~)Mˉ(ΔΔ~)α/2p(\Delta)\propto|\tilde\Psi+(\Delta-\tilde\Delta)'\bar{M}(\Delta-\tilde\Delta)|^{-\alpha/2}, ΨΔIW()\Psi|\Delta\sim IW(\cdot). Overcomes the ΨΩ\Psi\otimes\Omega restriction on Var(γ)\text{Var}(\gamma) that constrains the Normal-Wishart. Marginal posterior of Δ\Delta is not matricvariate tt; no closed-form moments — numerical evaluation required.

Forecasting

Empirical Results

Concepts Introduced or Extended

Entities Mentioned

Quotes

"The prior distributions that allow for dependencies between the equations of the VAR give rise to better forecasts." (abstract)

"In no case does the Minnesota prior provide forecasts that are significantly better than the more general prior distributions."

"The Normal-Wishart and Diffuse priors are especially suitable in this context, since analytic expressions for the first and second posterior moments of the parameters are available."

My Take

The paper's central contribution is empirical: the Minnesota prior's forced equation-independence is costly in small samples, and more flexible priors consistently match or beat it. The theoretical critique — treating Ψ\Psi as known diagonal is the wrong tradeoff — is well-argued and anticipates the later Sims-Zha framework. The ENC prior is technically elegant (the Drèze-Morales reparameterization is clean) but computationally demanding; the paper notes Gibbs sampling as an alternative, which Kadiyala and Karlsson (1997) later pursued. The Normal-Wishart prior is the practical choice for applications requiring closed-form posterior moments (impulse response functions, IRFs, and variance decompositions), as the authors note in the conclusion.