George-Sun-Ni (2008) Bayesian Stochastic Search for VAR Model Restrictions

bayesianvarvariable-selectionspike-and-slabgibbs-samplermcmcmodel-selectionregressionstochastic-search

Summary

George, Sun, and Ni extend the stochastic search variable selection (SSVS) framework of George-McCulloch (1993, 1997) from univariate regression to vector autoregression (VAR) models, simultaneously searching restrictions on the VAR coefficient matrix Φ\Phi and the off-diagonal elements of the upper-triangular Cholesky factor Ψ\Psi (where Σ1=ΨΨ\Sigma^{-1} = \Psi\Psi'). The upper-triangular parameterisation is essential for global identification of Σ\Sigma: arbitrary Ψ\Psi satisfies only local rank conditions and may be non-unique (Bekker-Pollock 1986; Amisano-Giannini 1997). A five-component conjugate prior yields a five-step Gibbs sampler whose every conditional is of standard form (Gamma, Normal, or Bernoulli) with no Metropolis-Hastings steps. Rao-Blackwell model-averaged forecasts substantially outperform unrestricted maximum likelihood estimation (MLE) across simulated and empirical settings.

Key Claims

How It Works

Model Setup

VAR: yt=ztC+j=1LytjAj+εty_t' = z_t' C + \sum_{j=1}^L y_{t-j}' A_j + \varepsilon_t', εtNp(0,Σ)\varepsilon_t \sim \mathcal{N}_p(0,\Sigma); stack as Y=XΦ+EY = X\Phi + \mathcal{E} (eq. 4). MLE: Φ^M=(XX)1XY\hat\Phi_M = (X'X)^{-1}X'Y (eq. 8).

Precision decomposition: Σ1=ΨΨ\Sigma^{-1} = \Psi\Psi' where Ψ\Psi is upper triangular with positive diagonal (eq. 9). Row jj of Ψ\Psi: diagonal element ψjj>0\psi_{jj} > 0; off-diagonal row vector ηj=(ψ1j,,ψj1,j)\eta_j = (\psi_{1j},\ldots,\psi_{j-1,j})'; restriction indicator vector ωj\omega_j.

Five-Component Prior (Section 2.2)

  1. Always-included coefficients ϕnon\phi_{non}: Nnm(ϕ~non,Mnon)\mathcal{N}_{n-m}(\tilde\phi_{non}, M_{non}) (eq. 11).
  2. Restricted ϕm\phi_m: mixture (ϕsiγi)(1γi)N(0,τ0i2)+γiN(0,τ1i2)(\phi_{s_i}|\gamma_i) \sim (1-\gamma_i)\mathcal{N}(0,\tau_{0i}^2) + \gamma_i\mathcal{N}(0,\tau_{1i}^2) (eq. 13); γiBernoulli(pi)\gamma_i \sim \text{Bernoulli}(p_i) (eq. 14).
  3. Off-diagonal ψij\psi_{ij}: mixture (ψijωij)(1ωij)N(0,κ0ij2)+ωijN(0,κ1ij2)(\psi_{ij}|\omega_{ij}) \sim (1-\omega_{ij})\mathcal{N}(0,\kappa_{0ij}^2) + \omega_{ij}\mathcal{N}(0,\kappa_{1ij}^2) (eq. 16); ωijBernoulli(qij)\omega_{ij} \sim \text{Bernoulli}(q_{ij}) (eq. 17).
  4. Diagonal ψii2Gamma(ai,bi)\psi_{ii}^2 \sim \text{Gamma}(a_i, b_i), default ai=bi=0.01a_i = b_i = 0.01 (eq. 18).
  5. Semiautomatic tuning: τki=ckσ^ϕsi\tau_{ki} = c_k\hat\sigma_{\phi_{s_i}} from OLS.

Five-Step Gibbs Sampler (Appendix A.3, eqs. 51–56)

  1. (ψii2ϕ,ω;Y)Gamma(ai+T/2,Bi)(\psi_{ii}^2|\phi,\omega;Y) \sim \text{Gamma}(a_i + T/2,\, B_i) where BiB_i involves Schur complements of S(Φ)=(YXΦ)(YXΦ)S(\Phi) = (Y-X\Phi)'(Y-X\Phi) (eq. 51).
  2. (ηjϕ,ω,ψ;Y)Nj1(μj,Δj)(\eta_j|\phi,\omega,\psi;Y) \sim \mathcal{N}_{j-1}(\mu_j, \Delta_j) via normal-normal conjugacy (eq. 53).
  3. ωijBernoulli\omega_{ij} \sim \text{Bernoulli} via ratio of normal densities evaluated at current ψij\psi_{ij} (eq. 30).
  4. (ϕΨ,Y)N(\phi|\Psi,Y) \sim \mathcal{N} from normal-normal conjugate (eq. 56).
  5. (γiϕ,γ(i),η,ψ;Y)Bernoulli(\gamma_i|\phi,\gamma_{(-i)},\eta,\psi;Y) \sim \text{Bernoulli} — does not depend on YY (eq. 37; cf. George-McCulloch (1993)).

Identification Result

Upper-triangular Ψ\Psi is globally identified: unique mapping ΣΨ\Sigma \to \Psi via Cholesky. Arbitrary Ψ\Psi satisfies only local rank conditions (Bekker-Pollock 1986); the decomposition may be non-unique with multiple valid Ψ\Psi matrices.

Simulation Results (Section 4)

Example 1 — 6-variable VAR, L=1L=1, T=50T=50, 100 replications, 50K burn-in + 10K kept:

Example 2 — 4-variable VAR, L=2L=2, 36 Φ\Phi elements + 6 Ψ\Psi elements:

Empirical Application (Section 5)

7-variable PPI-to-CPI VAR, L=12L=12, two periods (1969:1–1980:12 and 1981:1–2001:12):

1969–80: Supply-chain contemporaneous structure confirmed; links FF↔CM, CM↔IM, IM↔FC, IM&FC↔CPI, CPI↔UNEMP, UNEMP↔FFR. PPI lagged coefficients NOT significant predictors of CPI. Key CPI equation: 6.579εCPI=0.060εFF+0.022εCM+1.054εIM+0.533εFC+uCPI6.579\,\varepsilon_{CPI} = -0.060\,\varepsilon_{FF} + 0.022\,\varepsilon_{CM} + 1.054\,\varepsilon_{IM} + 0.533\,\varepsilon_{FC} + u_{CPI} (standard deviation, SD ≈ 0.152%).

Post-1981: FF-CM link breaks; FFR becomes endogenously responsive to IM, CPI, and UNEMP shocks contemporaneously — consistent with Volcker-era inflation targeting.

Concepts Introduced or Extended

Entities Mentioned

Quotes

"The key advantage of SSVS over other model selection approaches is that, rather than selecting a single model, it uses all the information in the data by averaging over a large class of potential models." (paraphrase of paper's motivation)

"The Φ-search Gibbs sampler step 5 shows that γ does not depend on Y — the same property exploited in George and McCulloch (1993)." (eq. 37 and surrounding discussion)

My Take

The paper's main technical contribution is clean: it shows that the upper-triangular Cholesky decomposition of the precision matrix is both necessary for global identification and sufficient to keep all Gibbs conditionals in standard form. The result that restricting Φ helps Ψ selection (and vice versa) is intuitive but the simulation quantification is useful. The empirical application is informative but the two-period comparison is qualitative — no formal test is offered for whether the contemporaneous structure genuinely changed. The semiautomatic tuning rule inherited from George-McCulloch (1997) is convenient but the sensitivity of results to c0c_0 and c1c_1 deserves more attention in high-p settings.