Steady State VAR

varbayesiansteady-stateforecastingcointegrationgibbs-sampler

Definition

The steady state vector autoregression (VAR) (mean-adjusted VAR) is a reparametrization of the standard VAR in which the unconditional mean μt=Ψdt\mu_t = \Psi d_t is an explicit, separately identified parameter, enabling informative Bayesian priors to be placed directly on the steady state of the process. Proposed by Villani (2008) to address the gap that all existing Bayesian VAR priors (Minnesota, Sims-Zha) are effectively non-informative about the long-run level of the system.

Key Ideas

How It Works

Standard Form vs. Steady State Form

The standard VAR parametrization is:

Π(L)xt=Φdt+εt,εtNp(0,Σ)(2.1)\Pi(L)x_t = \Phi d_t + \varepsilon_t, \qquad \varepsilon_t \sim \mathcal{N}_p(0,\Sigma) \tag{2.1}

with unconditional mean μt=Π1(1)Φdt\mu_t = \Pi^{-1}(1)\Phi d_t — a non-linear function of Π1,,Πk\Pi_1,\ldots,\Pi_k and Φ\Phi.

The steady state VAR reparametrizes as:

Π(L)(xtΨdt)=εt(2.2)\Pi(L)(x_t - \Psi d_t) = \varepsilon_t \tag{2.2}

so that the unconditional mean is directly μt=Ψdt\mu_t = \Psi d_t. Common choices for dtd_t: a constant (fixed steady state), a piecewise constant (regime shift), or a linear time trend.

Prior Distribution

Three independent blocks:

p(Σ)Σ(p+1)/2p(\Sigma) \propto |\Sigma|^{-(p+1)/2}

vecΠN(θΠ,ΩΠ)\mathrm{vec}\,\Pi \sim \mathcal{N}(\theta_\Pi,\, \Omega_\Pi)

vecΨNpq(θΨ,ΩΨ)\mathrm{vec}\,\Psi \sim \mathcal{N}_{pq}(\theta_\Psi,\, \Omega_\Psi)

The prior on Π\Pi nests Minnesota/Litterman and Sims-Zha Prior as special cases. The novel component is the prior on Ψ\Psi: θΨ\theta_\Psi encodes the analyst's point belief about the steady state, and ΩΨ\Omega_\Psi its uncertainty. For example, an inflation-targeting central bank sets the inflation equation's row of θΨ\theta_\Psi to 2% with tight ΩΨ\Omega_\Psi.

Three-Block Gibbs Sampler

The joint posterior p(Σ,Π,ΨD)p(\Sigma,\Pi,\Psi \mid D) is intractable (model is non-linear in parameters), but all full conditional posteriors are standard:

Block 1 — ΣΠ,Ψ,D\Sigma \mid \Pi, \Psi, D:

ΣΠ,Ψ,DIW(EE,T),E=[Π(L)(xtΨdt)]t=1T\Sigma \mid \Pi, \Psi, D \sim \mathrm{IW}(E'E,\, T), \qquad E = \bigl[\Pi(L)(x_t - \Psi d_t)\bigr]_{t=1}^T

Block 2 — vecΠΣ,Ψ,D\mathrm{vec}\,\Pi \mid \Sigma, \Psi, D:

vecΠΣ,Ψ,DN(θˉΠ,ΩˉΠ)\mathrm{vec}\,\Pi \mid \Sigma, \Psi, D \sim \mathcal{N}(\bar\theta_\Pi,\, \bar\Omega_\Pi)

ΩˉΠ1=Σ1XΨXΨ+ΩΠ1\bar\Omega_\Pi^{-1} = \Sigma^{-1} \otimes X_\Psi'X_\Psi + \Omega_\Pi^{-1}

where YΨ=[xtΨdt]Y_\Psi = [x_t - \Psi d_t] and XΨ=[xt1Ψdt1,,xtkΨdtk]X_\Psi = [x_{t-1}-\Psi d_{t-1},\ldots,x_{t-k}-\Psi d_{t-k}]. Conditional on Ψ\Psi, the model reduces to a standard VAR for the mean-adjusted series xtΨdtx_t - \Psi d_t.

Block 3 — vecΨΣ,Π,D\mathrm{vec}\,\Psi \mid \Sigma, \Pi, D:

vecΨΣ,Π,DN(θˉΨ,ΩˉΨ)\mathrm{vec}\,\Psi \mid \Sigma, \Pi, D \sim \mathcal{N}(\bar\theta_\Psi,\, \bar\Omega_\Psi)

ΩˉΨ1=U(DDΣ1)U+ΩΨ1\bar\Omega_\Psi^{-1} = U'(D'D \otimes \Sigma^{-1})U + \Omega_\Psi^{-1}

where U=(Ipq,IqΠ1,,IqΠk)U' = (I_{pq},\, I_q \otimes \Pi_1',\ldots, I_q \otimes \Pi_k') encodes the effect of Ψ\Psi on all lags. The Ψ\Psi update is derived by rewriting (2.2) as Y=DΘ+EY = D\Theta + E with vecΘ=UvecΨ\mathrm{vec}\,\Theta' = U\,\mathrm{vec}\,\Psi.

Unit Root Stability

For the constant-mean case (dt=1d_t = 1), the precision of Ψ\Psi's full conditional simplifies to:

ΩˉΨ1=T ⁣(Ipi=1kΠi) ⁣Σ1 ⁣(Ipi=1kΠi)+ΩΨ1\bar\Omega_\Psi^{-1} = T\!\left(I_p - \sum_{i=1}^k \Pi_i\right)'\!\Sigma^{-1}\!\left(I_p - \sum_{i=1}^k \Pi_i\right) + \Omega_\Psi^{-1}

At a unit root, (IpiΠi)(I_p - \sum_i \Pi_i) is rank-deficient: the data carry no information about Ψ\Psi. With a flat prior (ΩΨ10\Omega_\Psi^{-1} \to 0), ΩˉΨ\bar\Omega_\Psi diverges and the Gibbs sampler fails. With an informative prior:

ΩˉΨΩΨas ρ ⁣(IpiΠi)0\bar\Omega_\Psi \to \Omega_\Psi \quad \text{as } \rho\!\left(I_p - \sum_i\Pi_i\right) \to 0

The prior regularizes the singular information matrix. The sampler remains stable even if posterior draws of Π\Pi stray into the non-stationary region.

Simulation evidence (Villani 2005, Figures 1–3). Bivariate mean-adjusted autoregressive (AR)(1) with Π=diag(0.95,0.95)\Pi = \mathrm{diag}(0.95, 0.95), Ψ=(1,4)\Psi^* = (1, 4)', T=100T = 100 (near-unit-root). Three prior regimes:

Prior on Ψ\Psi Behaviour
Flat Gibbs diverges: explosive excursions when Π\Pi-draw enters non-stationary region; chain does not return quickly
Mildly informative: ψ1N(2.5,1.252)\psi_1 \sim \mathcal{N}(2.5, 1.25^2), ψ2N(5,2.52)\psi_2 \sim \mathcal{N}(5, 2.5^2) Drastically better; occasional large jumps but chain recovers
Informative: ψ1N(1,1)\psi_1 \sim \mathcal{N}(1, 1), ψ2N(4,1)\psi_2 \sim \mathcal{N}(4, 1) Excellent mixing throughout; maximal eigenvalue stays below 1

The mildly informative prior already captures the key gain; the informative prior adds little beyond that. The lesson is that even a relatively vague steady-state prior with correct order of magnitude is sufficient to prevent sampler failure.

Steady State Vector Error Correction Model (VECM)

For cointegrated I(1)I(1) systems (Clements and Hendry 1999 parametrization):

Γ(L)(Δxtγ)=α(βxt1μ0μ1t)+εt(3.1)\Gamma(L)(\Delta x_t - \gamma) = \alpha(\beta'x_{t-1} - \mu_0 - \mu_1 t) + \varepsilon_t \tag{3.1}

where:

When μ1=0\mu_1 = 0, the constraint βγ=0\beta'\gamma = 0 is imposed via γ=βλ\gamma = \beta_\perp\lambda, and the prior on γ\gamma is projected:

λβN(βθγ,  βΩγβ)\lambda \mid \beta \sim \mathcal{N}(\beta_\perp'\theta_\gamma,\; \beta_\perp'\Omega_\gamma\beta_\perp)

Posterior sampling uses the decomposition:

p(ϕ,η,Γ,ΣD)=p(ϕ,ηD)p(Γ,Σϕ,η,D)p(\phi, \eta, \Gamma, \Sigma \mid D) = p(\phi, \eta \mid D)\cdot p(\Gamma, \Sigma \mid \phi, \eta, D)

where ϕ\phi collects unrestricted cointegration vector coordinates and η=(μ0,γ)\eta = (\mu_0, \gamma)' (or (μ0,λ)(\mu_0, \lambda)' when μ1=0\mu_1=0). The marginal p(ϕ,ηD)p(\phi,\eta\mid D) is non-standard but low-dimensional and sampled via independence Metropolis-Hastings with a tailored Student-tt proposal. The conditional p(Γ,Σϕ,η,D)p(\Gamma,\Sigma\mid\phi,\eta,D) is Normal-Inverse Wishart and sampled directly.

Normal-Diffuse Prior and Stationarity Constraint (Meseguer 2010)

Meseguer (2010) uses a Normal-diffuse prior rather than the Normal-Inverse Wishart conjugate prior:

p(Σ)Σ(m+1)/2,vec(Φ)N(μΦ,ΣΦ),vec(Ψ)N(μΨ,ΣΨ)p(\Sigma) \propto |\Sigma|^{-(m+1)/2}, \qquad \mathrm{vec}(\Phi) \sim \mathcal{N}(\mu_\Phi, \Sigma_\Phi), \qquad \mathrm{vec}(\Psi) \sim \mathcal{N}(\mu_\Psi, \Sigma_\Psi)

where mm is the number of variables. This is a non-conjugate specification (the Σ\Sigma marginal does not have a closed-form Normal-Inverse Wishart (NIW) posterior), so a three-block Gibbs sampler cycles over (Σ,Φ,Ψ)(\Sigma, \Phi, \Psi). The prior on Φ\Phi encodes the Minnesota hyperparameters (λ1\lambda_1λ4\lambda_4) via the block diagonal ΣΦ\Sigma_\Phi; the prior on Ψ\Psi encodes the demographic steady-state beliefs (e.g., declining mortality trend, replacement-level fertility).

Stationarity rejection constraint. To prevent explosive draws — common in near-unit-root demographic series — a hard stationarity constraint is imposed on each Gibbs draw via an indicator function:

I(Φ)={1if λmax(companion matrix of Φ)<10otherwiseI(\Phi) = \begin{cases} 1 & \text{if } |\lambda_{\max}(\text{companion matrix of } \Phi)| < 1 \\ 0 & \text{otherwise} \end{cases}

Draws with I(Φ)=0I(\Phi) = 0 are rejected and re-drawn. This truncates the prior and posterior to the stationary region without requiring a formal truncated normal distribution — an effective practical solution for large systems where explicit stationarity-constrained priors are computationally intractable.

Demographic application parameters:

Extension to Structural VARs

For identified VARs with Σ=Ip\Sigma = I_p and equation-by-equation restrictions ωi=Giϕi\omega_i = G_i\phi_i, the full conditional of the unrestricted structural coefficients ϕi\phi_i depends on the absolute normal distribution:

fAN(x;μ,ρ)=cxexp ⁣[(xμ)22ρ],xRf_{\mathrm{AN}}(x;\mu,\rho) = c\,|x|\,\exp\!\left[-\frac{(x-\mu)^2}{2\rho}\right], \qquad x \in \mathbb{R}

This arises because Υξ1|\Upsilon| \propto |\xi_1| after the Waggoner-Zha (2003b) change-of-variables. The key extension over Waggoner-Zha is that the centering is now ξ^j=μϕiRivj0\hat\xi_j = \mu_{\phi_i}'R_i^{-\prime}v_j \neq 0 (non-zero prior mean). The distribution is bimodal with modes at μ/2±12μ2+4\mu/2 \pm \frac{1}{2}\sqrt{\mu^2+4} and is approximated by a mixture of two Gaussians, which is essentially exact for ρ=T1\rho = T^{-1} at empirical sample sizes.

Why It Matters

Open Questions

Related