A Vector Autoregression (VAR) is a linear dynamic system in which each of endogenous variables is regressed on lags of all variables in the system, plus optional exogenous terms. Introduced by Sims (1980) as a minimally-restricted alternative to large structural econometric models.
where:
The independence of is the defining constraint of the structural VAR: it is what allows each shock to be given an unambiguous economic label (technology shock, policy shock, etc.).
Stacking and :
Pre-multiplying by :
where , , and . The reduced-form covariance matrix is:
The reduced form is identified without further assumptions and can be estimated by OLS. Recovering from is the identification problem.
The VAR(p) process is stable if
Stability implies the process is stationary with time-invariant moments. A root at (unit root) allows I(1) and cointegrated variables; a root inside the unit circle () produces explosive dynamics.
When the VAR includes exogenous variables alongside the endogenous , three exogeneity concepts are relevant (Engle, Hendry, and Richard 1983):
Only super-exogeneity justifies using the conditional VAR model for counterfactual policy experiments; weak exogeneity suffices for estimation.
A VAR with variables, lags, and a constant has free parameters. With and this yields 1,500 parameters per equation; the Minneapolis Fed 47-variable model had ~8,000 total (Litterman 1986a). Unrestricted OLS fits noise rather than signal in such settings. Zellner (1985) sharpens this critique: a 6-variable multivariate ARMA (MVARMA)(r=3,q=4) has 395 parameters against observations (obs/param ); an unrestricted VAR(10) with 6 variables has 387; a VAR(4) has 159. By contrast, the 3-equation Demand-Supply-Entry structural model has 20 parameters — illustrating how theory creates parsimony that data alone cannot.
Standard approaches offer only two options: exclude a variable (exact-zero prior — too rigid) or include it without a prior (flat prior — too diffuse). The Minnesota Prior (Litterman 1986b) resolves this by assigning informative shrinkage priors to every coefficient, encoding the belief that most are near zero while the first own lag is near one (random walk belief). This yields a ridge-type estimator:
Monte Carlo simulation with 3,000 draws shows that the Bayesian posterior mean dominates OLS with any fixed variable set and beats stepwise selection, especially when is small or is low (Litterman 1986b, Section 3).
Five-year real-time comparison (1980:2–1985:1) against DRI, Wharton, and Chase commercial forecasters across real gross national product (GNP), GNP deflator, nominal GNP, and unemployment (1,604 total forecasts): the Bayesian VAR (BVAR) was closest to actual 34.8% of the time vs. 16–27% for each commercial service. The BVAR significantly underperformed on inflation (over-forecast by >2 bootstrap standard errors at all horizons) but outperformed on real GNP at 4–7 quarter horizons.
Longer lags reduce misspecification but increase the parameter count as , rapidly exhausting degrees of freedom. The Bayesian approach (see Sims-Zha Prior) resolves this by shrinking coefficients rather than truncating lags. Frequentist model selection uses information criteria Akaike Information Criterion (AIC), Hannan-Quinn (HQ), and Schwarz Criterion (SC) (see Lag-Order Selection Criteria).
Kadiyala and Karlsson (1993) compare five prior distributions for Bayesian VAR forecasting across three empirical experiments (Canadian GNP/M2 data and US wheat export data):
| Prior | Key property | Posterior |
|---|---|---|
| Minnesota | Equation-independent; fixed diagonal | Analytical (ridge) |
| Normal-Wishart | , ; conjugate | Matrix- closed form |
| Diffuse (Jeffreys') | ; OLS-centered | Matrix- closed form |
| Normal-Diffuse | Minnesota prior + diffuse | Importance sampling |
| ENC | Drèze-Morales reparametrization; unrestricted | Importance sampling |
Main finding: The Minnesota prior never provides forecasts significantly better than the more general alternatives. In small samples (17–26 observations) it is significantly worse than the extended natural conjugate (ENC) and Normal-Diffuse priors. OLS is worst in every experiment. The Normal-Wishart and Diffuse priors are most computationally convenient (closed-form posterior moments for IRFs and variance decompositions).
When a researcher lacks prior information and wishes to use a minimal Bayesian benchmark, two noninformative priors for are commonly considered (Sun-Ni 2005):
See Noninformative Prior for VAR for the propriety conditions and hit-and-run Markov Chain Monte Carlo (MCMC) algorithm.
George, Sun, and Ni extend the stochastic search variable selection (SSVS) framework of George-McCulloch (1993, 1997) to VAR models, simultaneously searching restrictions on the coefficient matrix (of dimension ) and the off-diagonal elements of the upper-triangular Cholesky factor (where ). The total competing submodel space has elements.
Why upper-triangular : The Cholesky decomposition gives a globally unique mapping , ensuring identification. An arbitrary lower-triangular satisfies only local rank conditions and may be non-unique (Bekker-Pollock 1986; Amisano-Giannini 1997).
Mixture prior: Binary indicators for elements and for off-diagonal elements each select between narrow (exclusion) and wide (inclusion) Normal components. Diagonal — always included — guarantees positive definiteness of every sampled .
Five-step Gibbs sampler (all conditionals standard, no Metropolis-Hastings (MH) steps):
Rao-Blackwell forecasts average over all visited models, improving out-of-sample mean squared error (MSE) by 20–67% over MLE in simulation. Restricting improves selection and vice versa.
Empirical application: 7-variable producer price index (PPI)→consumer price index (CPI) VAR (), two periods. In 1969–80, the supply-chain contemporaneous structure is confirmed (IM→FC→CPI, UNEMP→FFR). Post-1981, the FF-CM link breaks and the Fed Funds Rate becomes endogenously responsive to commodity, import, and CPI shocks contemporaneously (Volcker disinflation). See Variable Selection for the full SSVS-VAR prior and Gibbs sampler details.
An alternative parametrization exposes the unconditional mean (steady state) as an explicit parameter:
where is directly identified rather than implicitly defined as . This reparametrization enables informative Bayesian priors on the steady state — for example, centering on a central bank's inflation target or a Dynamic Stochastic General Equilibrium (DSGE) model's long-run calibration. See Steady State VAR.
A VAR may be disciplined by no-arbitrage cross-equation restrictions when the state vector includes bond yields alongside macro variables. In the Ang-Piazzesi (2001) framework, the state follows a companion VAR(1) (eq. 8) and bond yields of maturity are constrained to lie on an affine surface:
where must satisfy the no-arbitrage recursions determined by the stochastic discount factor (see Affine Term Structure Model). These recursions impose tight cross-equation restrictions: all yields must be consistent with a single pricing kernel. Practical consequences:
Any linear (or linearized) DSGE equilibrium can be cast as a state-space system (see State-Space Representation). The Kalman filter applied to this system produces the innovations representation , from which the VAR for the observables follows directly:
The DSGE therefore implies a generically infinite-order VAR (or more precisely a VARMA), not a finite-order VAR. Three conditions govern the quality of a finite-order VAR approximation:
When invertibility holds, the Kalman gain collapses to , the estimation error covariance is , and the identification matrix equals the model's matrix (). In this case VAR impulse responses exactly recover the economic model's impulse responses — no approximation error. When invertibility fails, no choice of can repair the mismatch.
The foundational VAR application in Sims (1980) uses a six-variable quarterly system to demonstrate the methodology.
Variables: money , real output , unemployment , nominal wages , price level , import prices , plus constant and time trend.
Data: U.S. 1949Q1–1975Q4 (); West Germany 1958Q1–1976Q2 ().
Lag length selection. Sims (1980) tests 4 vs. 8 lags using a -corrected likelihood ratio statistic (where is the number of RHS variables per equation) to reduce finite-sample over-rejection:
Results: U.S. ; Germany . Both countries accept at conventional critical values (). The 144 degrees of freedom reflect 144 additional parameters when going from to in a 6-variable system.
Identification. Cholesky triangularization in the ordering : money comes first (does not respond contemporaneously to any other variable within the quarter), import prices last. Orthogonalized innovations are unit-variance by construction.
Variance decompositions (Sims 1980, Tables III–IV). At long horizons, money dominates U.S. nominal variables:
| Horizon | U.S.: % of variance due to | U.S.: % of variance due to |
|---|---|---|
| 1 | 0% | 3% |
| 9 | 37% | 30% |
| 33 | 64% | 60% |
Germany shows a different pattern: price innovations dominate own forecast error at all horizons (86% at ), and real GNP is strongly driven by its own innovations (93% at ), reflecting a more supply-side-driven economy.
Block exogeneity tests. The rational expectations market-clearing hypothesis predicts real-sector variables should not be Granger-caused by money. Testing this as a zero-restriction Wald test on cross-block lag coefficients:
| Hypothesis | U.S. (df) | Germany (df) | Conclusion |
|---|---|---|---|
| Real sector exogenous to | (24), | (32), | Rejected both |
| Real + money jointly exogenous (U.S.) | (36), | — | Not rejected |
Money Granger-causes real variables in both countries, contradicting pure market-clearing. See Granger Causality for the block exogeneity test framework.
When (the number of micro/industry variables) is large, estimating an unrestricted VAR is infeasible. Lastrapes (2005) shows that two over-identifying restrictions make the problem tractable. Partition where () are micro variables and () are common/aggregate factors.
Restrictions:
Estimation: These restrictions allow the system to be separated into two independent sub-systems: with . The residual covariance is diagonal (because is diagonal), so equation-by-equation OLS is fully efficient — no GLS needed. Each equation in the sub-system contains only its own lags plus and its lags.
Identification: Using with , structural identification proceeds in three steps:
The full set of structural impulse responses follows from . The entire identification burden falls on the -dimensional aggregate sub-system, keeping the problem feasible regardless of .
Critique of state-by-state estimation: The Carlino-Defina (1998) approach of estimating a separate small VAR for each state or industry — with its own aggregate variables but without imposing block exogeneity jointly — is misleading because the identification of aggregate shocks varies arbitrarily across the separate systems.
Felix and Nunes (2003) apply Bayesian VAR and Bayesian error correction model (BECM) methods to forecasting six quarterly euro area aggregates — gross domestic product (GDP) (Y), unemployment (U), private consumption deflator (P), nominal wages (W), long-term interest rate (ILT), and nominal exchange rate (S) — with three exogenous variables (external GDP YW, external price index PW, short-term rate IST). Data: 11-country euro area aggregate, 1977:1–1997:4; pseudo out-of-sample evaluation 1989:1–1997:4.
Integration. Augmented Dickey-Fuller (ADF) tests classify Y, U, ILT, S, YW, IST as I(1); P, W, PW as I(2) (unit root in first differences). I(2) variables enter in first differences within levels BVAR specifications.
Twelve models compared (RMSE relative to random walk = 1.000, averaged over 6 variables 12 horizons):
| Model | Avg RMSE |
|---|---|
| BVAR levels | 0.731 |
| BECM(EG)-IP | 0.754 |
| BECM(J)-IP | 0.805 |
| BVAR first-diff | 0.839 |
| BECM(EG)-FP | 0.857 |
| AR | 1.042 |
| VAR levels | 1.188 |
| BECM(J)-FP | 1.278 |
All Bayesian models beat all non-Bayesian models. The Minnesota prior is extended with a real/price block structure (hyperparameters –) and a loading tightness — see Minnesota Prior for full details. The flat-prior danger in BECM(J)-FP is documented in Cointegration (§ Flat-Prior Pathology).
Kim, Miller, and Ozanne (2004) apply BVAR methods to a fiscal-policy forecasting problem: estimating current-year U.S. capital gains realizations and projecting them 10 years ahead for Congressional Budget Office (CBO) revenue scoring. Two BVAR strategies are compared against the CBO historical mean-reversion benchmark (gains/GDP reverts to historical average), which produced a 1-year RMSE of 18.57 pp after tax adjustment.
Two-step method: A BVAR(5) in macro variables (real GDP, labor productivity, output gap, S&P 500, interest-rate spread, compensation per hour, fixed private investment) forecasts explanatory variables; annual tax-adjusted gains growth is then regressed on the forecast changes in those variables. Best 1-year RMSE: 14.80 pp (−20% vs. CBO benchmark).
Integrated quarterly method (best overall): Annual tax-adjusted gains are interpolated to quarterly frequency (three alternatives: linear, economic, in-model) and incorporated directly into a unified BVAR. The Kalman filter delivers current-year estimates that are revised as quarterly macro data arrive throughout the year (E1 = estimation error). At year-end the smoother is applied; the year-ahead projection follows from the BVAR transition equation (E2 = forecast|estimated start). Best variant: linear interpolation BVAR, 1-year RMSE 11.92 pp (−36% vs. CBO).
Why linear interpolation wins: Annual gains residuals are positively serially correlated (Durbin-Watson = 0.78). In the quarterly model, an annual estimation error of propagates as (because the quarterly model revises the current-year path more aggressively than the annual model), while in the annual model it is . The extra revision weight amplifies the benefit of within-year data arrivals.
Prior specification: Follows Robertson-Tallman (1999) / Sims-Zha (1998) style with 6 hyperparameters (–) corresponding to: overall tightness, first-lag diagonal shrinkage, intercept variance, higher-lag decay rate, unit-root dummy observation weight, and cointegration dummy observation weight.
Multi-year horizon: BVAR superiority fades beyond 2–3 years; mean reversion or random-walk with drift performs comparably thereafter. The 2001 out-of-sample evaluation (actual −46%; best model −6.9%) illustrates that capital gains tail events driven by asset-price dynamics lie outside the BVAR's macro information set.
A VAR can circumvent the overlapping-data problem that afflicts long-horizon return regressions. With the VAR state (stock return, bond return, inflation, log dividend-price ratio, change in short rate) estimated at the one-period frequency, multi-period expected values follow directly from the companion matrix :
No time-overlapping observations appear; standard inference applies at any horizon . Engsted-Tanggaard (2000) apply this to test the Fisher Hypothesis for US (1926–1997) and Danish (1922–1996) stocks and bonds at years. Key results: US stocks fail Fisher at (correlation with expected inflation falls from 0.52 at to 0.38); Danish stocks achieve near-perfect Fisher at (corr , SD ratio ). See Engsted-Tanggaard (2000) for full details.
VAR and vector error correction model (VECM) methods extend naturally to stochastic mortality modeling. Under the Lee-Carter structure, the mortality index for each of two related populations follows a non-stationary I(1) process. Modeling jointly with a VAR in differences:
achieves a symmetric model that does not require assuming which population is dominant. The non-divergence condition (equal long-run drift rates) must be imposed as a parameter constraint. Upgrading to VECM adds error-correction terms and enforces non-divergence automatically. See Lee-Carter Model and Cointegration for the mortality VECM setup.