Summary
This paper shows that a Vector Autoregression (VAR) with Bayesian shrinkage is an appropriate and effective tool for large dynamic systems — dozens to over a hundred variables — overturning the long-standing practice of keeping VARs small (typically 3–6 variables) to avoid parameter proliferation. Building on De Mol-Giannone-Reichlin (2008), the central methodological point is that the degree of shrinkage should be set in relation to the model size: as more variables are added, the prior must be tightened correspondingly, calibrated to hold a small benchmark model's in-sample fit constant. With this rule, large Bayesian VARs (implemented via a Minnesota-style Normal-inverted-Wishart prior imposed through dummy observations) forecast key US macro series as well as or better than small VARs and factor models, and produce credible impulse responses suitable for structural analysis.
Key Claims
- Shrinkage tuned to dimension controls over-fitting. The overall prior tightness λ is chosen so that the in-sample fit of the large model matches that of a small benchmark VAR; because a bigger system over-fits more, this forces the degree of shrinkage 1/λ to increase with the number of variables N. This is the key to making large VARs work — shrinkage acts as a regularization that keeps the effective number of parameters in check.
- Adding variables improves forecasts. For the three focus series (Employment, CPI, Federal Funds Rate), moving from a SMALL (3-variable) to a MEDIUM (
20) and LARGE (131-variable) VAR improves out-of-sample forecast accuracy (relative MSFE below one), contradicting the folk theorem that VARs cannot be large. Most of the gain is already realized at the medium (20-variable) size.
- Competitive with factor models. Large shrinkage-VARs match or beat the forecasting performance of factor-based methods — the dynamic factor model / diffusion-index approach of Stock-Watson (2002) and the Factor-Augmented VAR (FAVAR) of Bernanke-Boivin-Eliasz (2005) — while remaining a single, internally consistent model.
- Credible impulse responses for structural analysis. Under a recursive (Cholesky) identification of a monetary-policy shock, the large BVAR yields impulse responses close to those from the FAVAR and consistent with economic priors (no price puzzle), demonstrating that large VARs are usable for structural work, not just forecasting.
- Practical implementation via dummy observations. The Normal-inverted-Wishart prior (retaining the Minnesota principles — random-walk prior mean, harmonic lag decay, scale normalization) is imposed by augmenting the system with artificial ("dummy") observations, which also incorporates a sum-of-coefficients prior (à la Sims-Zha) that discourages an implausible deterministic component / imposes a "no-cointegration"-flexible unit-root belief. This keeps estimation a closed-form conjugate problem with no numerical difficulties even at N=131.
Concepts Introduced or Extended
- Bayesian VAR — large-system BVARs; shrinkage set in relation to cross-sectional dimension; regularization interpretation
- Minnesota Prior — Normal-inverted-Wishart version imposed via dummy observations; overall tightness λ decreasing in N
- Sims-Zha Prior — sum-of-coefficients / dummy-observation prior component
- Factor Model — Factor-Augmented VAR (FAVAR) as the benchmark large-information competitor
- Dynamic Factor Model — Stock-Watson diffusion-index forecasting as the alternative to large VARs
- Forecasting — out-of-sample MSFE evaluation across model sizes
- Impulse Response Function — monetary-policy shock identification in a large system
Entities Mentioned
Quotes
"This paper shows that Vector Autoregression with Bayesian shrinkage is an appropriate tool for large dynamic models."
"When the degree of shrinkage is set in relation to the cross-sectional dimension, the forecasting performance of small monetary VARs can be improved by adding additional macroeconomic variables and sectoral information."
"Large VARs with shrinkage produce credible impulse responses and are suitable for structural analysis."
My Take
This is the paper that made "large BVARs" a standard tool. The single most important idea is deceptively simple: shrink harder as the system grows, with the tightness pinned down by matching a small model's fit. That one rule turns the curse of dimensionality into a controlled bias-variance trade-off and lets a 100+ variable VAR out-forecast a carefully chosen 3-variable one. It reframes the VAR-vs-factor-model debate — rather than compressing many series into a few factors, keep all the series and let the prior do the regularizing — and shows the two approaches are near-equivalent in forecast accuracy while the BVAR retains a transparent structural interpretation. The dummy-observation implementation (Minnesota + sum-of-coefficients) keeps everything conjugate and computationally trivial, which is a large part of why the approach was so widely adopted. The wiki's Giannone-Lenza-Primiceri (2015) "Prior Selection for VARs" is the natural sequel — it turns the ad hoc "match the small model's fit" tuning of λ into a formal hierarchical/marginal-likelihood procedure. Complements the Minnesota and Sims-Zha prior pages and the Kadiyala-Karlsson (1997) Normal-inverted-Wishart machinery it builds on.