Summary
Proposes treating Minnesota prior hyperparameters as unknown parameters estimated from the data by maximising the marginal likelihood, in the spirit of hierarchical/empirical Bayes. The prior tightness λ (or ξ) governs how much the vector autoregression (VAR) is shrunk toward a naïve random-walk benchmark; the optimal λ balances forecast bias (from prior misspecification) against forecast variance (from parameter uncertainty). With a Normal-Wishart conjugate prior, the marginal likelihood is analytic, making the optimisation fast. Applied to U.S. macroeconomic VARs with up to 22 variables, data-driven λ^ substantially outperforms fixed-λ Minnesota and flat priors on out-of-sample log predictive scores.
Key Claims
- Hyperparameter as parameter. The standard Minnesota prior has tightness λ set by researcher judgment. This paper reframes it as a hyperparameter in a hierarchical model and estimates it by λ^=argmaxλp(y∣λ), the marginal (predictive) likelihood of the data integrated over all VAR coefficients and error variances.
- Marginal likelihood is analytic. Under the Normal-Wishart conjugate prior (β∣Σ)∼N(β0,Σ⊗V0(λ)), Σ∼IW(Ψ0,d0) (inverse-Wishart, IW), the marginal likelihood p(y∣λ) has a closed-form expression involving the determinant of the marginal covariance. No Markov chain Monte Carlo (MCMC) is needed to optimize over λ; a standard numerical optimizer suffices.
- Bias-variance trade-off for λ. If λ is too small (dogmatic prior): density forecasts are tightly concentrated around the prior mean; unless the prior is accurate, coverage collapses and log scores are poor. If λ is too large (flat prior): in high-dimensional VARs, estimation uncertainty inflates density forecasts and degrades coverage. The optimal λ^ sits between these extremes and is identified by the marginal likelihood.
- Normal-Wishart vs. original Minnesota. Litterman (1986) fixed Σ at a diagonal matrix estimated equation-by-equation, enabling an analytical posterior but producing a non-conjugate prior. The conjugate Normal-Wishart prior jointly models (β,Σ) and yields cleaner analytical marginal likelihoods at the cost of a slightly different shrinkage structure.
- Large-VAR gains. Applied to U.S. macroeconomic systems with 3, 7, and 22 variables. For small VARs, the gains from data-driven λ^ over sensible fixed values are modest. For large VARs (22 variables), the gains are substantial: the marginal likelihood selects a much tighter prior than the defaults used in the literature, improving density forecast scores markedly.
- Analogy to cross-validation. Maximizing the marginal likelihood is equivalent (asymptotically) to leave-one-out cross-validation for the density forecasts, connecting the approach to predictive model selection without explicit held-out data.
Concepts Introduced or Extended
Entities Mentioned
- (none with existing wiki pages)
Quotes
"This paper studies the optimal choice of the informativeness of these priors, which we treat as additional parameters, in the spirit of hierarchical modeling. This approach is theoretically grounded, easy to implement, and greatly reduces the number and importance of subjective choices in the setting of the prior."
My Take
One of the most practically influential Bayesian VAR (BVAR) papers of the 2010s. The key insight — that the prior tightness can be estimated rather than calibrated — is both theoretically clean and computationally free (analytic marginal likelihood). The Normal-Wishart conjugate is a slight departure from the original Litterman (1986) specification but is now the standard in the large-BVAR literature. A limitation is that the approach treats the prior mean (random walk) as fixed; extensions allowing the prior mean itself to be estimated are developed in Bańbura-Giannone-Reichlin (2010) and Villani (2009). The paper is a direct precursor to the large-BVAR literature (BVAR with 100+ variables, Koop 2013) where prior tightness selection is essential.
Circulated as NBER Working Paper No. 18467 (2012); published as Giannone, D., M. Lenza, and G.E. Primiceri (2015), "Prior Selection for Vector Autoregressions," Review of Economics and Statistics 97(2): 436–451 (the citation of record).