The steady state vector autoregression (VAR) (mean-adjusted VAR) is a reparametrization of the standard VAR in which the unconditional mean μt=Ψdt is an explicit, separately identified parameter, enabling informative Bayesian priors to be placed directly on the steady state of the process. Proposed by Villani (2008) to address the gap that all existing Bayesian VAR priors (Minnesota, Sims-Zha) are effectively non-informative about the long-run level of the system.
Key Ideas
In standard form, the unconditional mean μt=Π−1(1)Φdt is a non-linear function of all lag matrices — not directly accessible to prior elicitation.
Clements and Hendry (1998): a badly estimated process mean is the dominant source of long-horizon forecast failure.
An informative prior on the steady state encodes institutional knowledge naturally available to macroeconomists — inflation targets, dynamic stochastic general equilibrium (DSGE)-model-implied calibrations, historical pre-sample averages.
A second, underappreciated benefit: an informative prior on Ψ stabilizes the Gibbs sampler near unit roots, where the data information about the steady state vanishes.
How It Works
Standard Form vs. Steady State Form
The standard VAR parametrization is:
Π(L)xt=Φdt+εt,εt∼Np(0,Σ)(2.1)
with unconditional mean μt=Π−1(1)Φdt — a non-linear function of Π1,…,Πk and Φ.
The steady state VAR reparametrizes as:
Π(L)(xt−Ψdt)=εt(2.2)
so that the unconditional mean is directly μt=Ψdt. Common choices for dt: a constant (fixed steady state), a piecewise constant (regime shift), or a linear time trend.
Prior Distribution
Three independent blocks:
p(Σ)∝∣Σ∣−(p+1)/2
vecΠ∼N(θΠ,ΩΠ)
vecΨ∼Npq(θΨ,ΩΨ)
The prior on Π nests Minnesota/Litterman and Sims-Zha Prior as special cases. The novel component is the prior on Ψ: θΨ encodes the analyst's point belief about the steady state, and ΩΨ its uncertainty. For example, an inflation-targeting central bank sets the inflation equation's row of θΨ to 2% with tight ΩΨ.
Three-Block Gibbs Sampler
The joint posterior p(Σ,Π,Ψ∣D) is intractable (model is non-linear in parameters), but all full conditional posteriors are standard:
Block 1 — Σ∣Π,Ψ,D:
Σ∣Π,Ψ,D∼IW(E′E,T),E=[Π(L)(xt−Ψdt)]t=1T
Block 2 — vecΠ∣Σ,Ψ,D:
vecΠ∣Σ,Ψ,D∼N(θˉΠ,ΩˉΠ)
ΩˉΠ−1=Σ−1⊗XΨ′XΨ+ΩΠ−1
where YΨ=[xt−Ψdt] and XΨ=[xt−1−Ψdt−1,…,xt−k−Ψdt−k]. Conditional on Ψ, the model reduces to a standard VAR for the mean-adjusted series xt−Ψdt.
Block 3 — vecΨ∣Σ,Π,D:
vecΨ∣Σ,Π,D∼N(θˉΨ,ΩˉΨ)
ΩˉΨ−1=U′(D′D⊗Σ−1)U+ΩΨ−1
where U′=(Ipq,Iq⊗Π1′,…,Iq⊗Πk′) encodes the effect of Ψ on all lags. The Ψ update is derived by rewriting (2.2) as Y=DΘ+E with vecΘ′=UvecΨ.
Unit Root Stability
For the constant-mean case (dt=1), the precision of Ψ's full conditional simplifies to:
ΩˉΨ−1=T(Ip−i=1∑kΠi)′Σ−1(Ip−i=1∑kΠi)+ΩΨ−1
At a unit root, (Ip−∑iΠi) is rank-deficient: the data carry no information about Ψ. With a flat prior (ΩΨ−1→0), ΩˉΨ diverges and the Gibbs sampler fails. With an informative prior:
ΩˉΨ→ΩΨas ρ(Ip−i∑Πi)→0
The prior regularizes the singular information matrix. The sampler remains stable even if posterior draws of Π stray into the non-stationary region.
Simulation evidence (Villani 2005, Figures 1–3). Bivariate mean-adjusted autoregressive (AR)(1) with Π=diag(0.95,0.95), Ψ∗=(1,4)′, T=100 (near-unit-root). Three prior regimes:
Prior on Ψ
Behaviour
Flat
Gibbs diverges: explosive excursions when Π-draw enters non-stationary region; chain does not return quickly
The mildly informative prior already captures the key gain; the informative prior adds little beyond that. The lesson is that even a relatively vague steady-state prior with correct order of magnitude is sufficient to prevent sampler failure.
Steady State Vector Error Correction Model (VECM)
For cointegrated I(1) systems (Clements and Hendry 1999 parametrization):
Γ(L)(Δxt−γ)=α(β′xt−1−μ0−μ1t)+εt(3.1)
where:
β (p×r): cointegration vectors
α (p×r): error-correction adjustment speeds
γ=E(Δxt): mean growth rates — prior centered on, e.g., 2% real GDP growth per annum
μ0: mean of the cointegrating relations — prior centered on, e.g., the purchasing power parity (PPP) equilibrium real exchange rate
μ1=β′γ: drift in cointegrating relations (nonlinear constraint linking β and γ)
When μ1=0, the constraint β′γ=0 is imposed via γ=β⊥λ, and the prior on γ is projected:
λ∣β∼N(β⊥′θγ,β⊥′Ωγβ⊥)
Posterior sampling uses the decomposition:
p(ϕ,η,Γ,Σ∣D)=p(ϕ,η∣D)⋅p(Γ,Σ∣ϕ,η,D)
where ϕ collects unrestricted cointegration vector coordinates and η=(μ0,γ)′ (or (μ0,λ)′ when μ1=0). The marginal p(ϕ,η∣D) is non-standard but low-dimensional and sampled via independence Metropolis-Hastings with a tailored Student-t proposal. The conditional p(Γ,Σ∣ϕ,η,D) is Normal-Inverse Wishart and sampled directly.
Normal-Diffuse Prior and Stationarity Constraint (Meseguer 2010)
Meseguer (2010) uses a Normal-diffuse prior rather than the Normal-Inverse Wishart conjugate prior:
where m is the number of variables. This is a non-conjugate specification (the Σ marginal does not have a closed-form Normal-Inverse Wishart (NIW) posterior), so a three-block Gibbs sampler cycles over (Σ,Φ,Ψ). The prior on Φ encodes the Minnesota hyperparameters (λ1–λ4) via the block diagonal ΣΦ; the prior on Ψ encodes the demographic steady-state beliefs (e.g., declining mortality trend, replacement-level fertility).
Stationarity rejection constraint. To prevent explosive draws — common in near-unit-root demographic series — a hard stationarity constraint is imposed on each Gibbs draw via an indicator function:
I(Φ)={10if ∣λmax(companion matrix of Φ)∣<1otherwise
Draws with I(Φ)=0 are rejected and re-drawn. This truncates the prior and posterior to the stationary region without requiring a formal truncated normal distribution — an effective practical solution for large systems where explicit stationarity-constrained priors are computationally intractable.
Demographic application parameters:
Mortality BVAR: 22 age groups per gender; p=1 lag; series = Δlogma,t; λ1=0.2 (moderate tightness); prior AR mean = 0 (improvement continues at historical rate).
Fertility BVAR: 7 age groups; p=2 lags; series = logfa,t; λ1=0.05 (tight shrinkage toward prior); prior AR mean = 0.8 (strong persistence encoding slow convergence to replacement level).
Population projection: each pair of posterior draws (ma,t(s),fa,t(s)) is fed into the cohort-component identity to generate one full population trajectory; the ensemble of S trajectories provides the forecast distribution. See Cohort Component Method.
Extension to Structural VARs
For identified VARs with Σ=Ip and equation-by-equation restrictions ωi=Giϕi, the full conditional of the unrestricted structural coefficients ϕi depends on the absolute normal distribution:
fAN(x;μ,ρ)=c∣x∣exp[−2ρ(x−μ)2],x∈R
This arises because ∣Υ∣∝∣ξ1∣ after the Waggoner-Zha (2003b) change-of-variables. The key extension over Waggoner-Zha is that the centering is now ξ^j=μϕi′Ri−′vj=0 (non-zero prior mean). The distribution is bimodal with modes at μ/2±21μ2+4 and is approximated by a mixture of two Gaussians, which is essentially exact for ρ=T−1 at empirical sample sizes.
Why It Matters
Closes the gap between what macroeconomists know (steady states, targets, DSGE calibrations) and what standard Bayesian VAR priors can incorporate.
Yields large root mean squared error (RMSE) improvements over Litterman Bayesian VAR (BVAR) at medium-to-long horizons when data are uninformative about the steady state — Swedish inflation and gross domestic product (GDP) growth examples show the posterior mean of Ψ stays near the prior throughout.
Euro area application (Villani 2005): 7-variable system (π,Δw,Δc,Δi,r,e,Δy; 1970Q1–2002Q4); monetary policy regime dummy with break at 1992Q4 (end of Exchange Rate Mechanism (ERM)), 4 lags. Steady-state inflation estimates for post-1992 regime: maximum likelihood (ML) = 6.97% (full sample) / -0.62% (1980-subsample); standard BVAR = 4.56% / 1.39%; mean-adjusted models = ≈2.0% in both cases — consistent with European Central Bank (ECB)'s "below but close to 2%" target and robust to sample-period changes. The ML and standard BVAR estimates are in gross conflict with institutional knowledge that is easily encoded as a steady-state prior.
Riksbank application (Adolfson et al. 2005): the Villani (2005) Gibbs sampler was used for a 7-variable open-economy BVAR (foreign and domestic GDP growth, consumer price index (CPI), interest rate; real exchange rate) with a 2% steady-state inflation prior. This BVAR matched or beat Riksbank official forecasts at 2–8 quarter horizons and was the primary model in a Bayesian forecast combination exercise. See Forecast Combination.
The unit root stability argument provides a theoretical justification for the steady state prior beyond forecasting: it prevents sampler breakdown near the boundary of the stationary region without requiring stationarity to be imposed.
The VECM extension allows economists to specify beliefs about common growth rates and cointegrating equilibria separately, filling a gap in cointegration priors.
Open Questions
How to formally elicit ΩΨ when the analyst has only a point belief (e.g., an official target); sensitivity to this choice is not yet systematically studied.
Extensions to time-varying steady states are not yet developed within this framework — see Markov-Switching VAR for regime-switching approaches.
The Kronecker structure assumed for the prior covariance of Γ in the steady state VECM precludes cross-equation shrinkage of the Litterman type.