Sims and Zha argue that classical bootstrap confidence intervals for Vector AutoRegression (VAR) impulse responses are conceptually flawed: they confound parameter-location uncertainty with model-fit information, violating the likelihood principle. They advocate flat-prior Bayesian posterior probability bands as the correct reporting standard, and often find that 68% bands are more informative than 95% bootstrap bands. They introduce an eigendecomposition of the posterior covariance of the stacked impulse-response vector to characterise which shapes of departure from the point estimate are most probable. For overidentified Structural Vector AutoRegressions (SVARs) they derive the correct marginal posterior over A0 via Metropolis Markov Chain Monte Carlo (MCMC) and show that the "naive Bayesian" method — drawing from the unrestricted posterior and mapping via Maximum Likelihood Estimate (MLE) formulas — is neither Bayesian nor frequentist and fails when the likelihood has multiple peaks.
Key Claims
Likelihood principle: Classical confidence intervals report a parameter region obtained by inverting test statistics; they conflate where parameters are likely with how well the model fits. Bayesian posterior probability bands avoid this and are the conceptually correct summary of posterior uncertainty.
68% vs. 95%: Bayesian 68% bands often convey more information than 95% bootstrap bands because they concentrate on the high-density region of the posterior.
Eigendecomposition (Section 6): Stack impulse responses into c~; compute posterior mean cˉ and covariance Ω=WΛW′. Display cˉ±W⋅kλk (68%) or cˉ±1.96W⋅kλk (95%) for the leading eigenvectors k=1,2,…. Each band shows the most probable shape of departure from the point estimate in the Impulse Response Function (IRF).
Bootstrap critique: The "other-percentile" bootstrap amplifies estimator bias. Bias-corrected bootstrap (Kilian 1998a) performs best among bootstrap variants but still produces unrealistically narrow bands at long horizons: Bayesian/bootstrap band-length ratio falls to 0.60–0.65 at horizon 32 (Table IV).
Correct overidentified SVAR posterior (Section 8A): Marginal posterior p(A0)∝∣A0∣T−νexp[−21tr(A0S(B^)A0′)]; sampled via random-walk Metropolis with Waggoner-Zha normalisation (minimise distance of A0 from MLE); 480 000 draws per chain, 11 Central Processing Unit (CPU) hours for the 6-variable model.
Naive Bayesian critique: Drawing (B,Σ) from the unrestricted posterior and mapping via the MLE formula for (A0,Γ0) is not a posterior draw — it yields neither a Bayesian nor a frequentist measure and fails badly when the likelihood has multiple peaks (Table X).
Applications: (1) Gross National Product (GNP)–M1 bivariate: first eigenvalue captures >90% of posterior IRF variance — one shape direction dominates. (2) Blanchard-Quah bivariate with long-run restrictions: shape uncertainty is more complex, requiring 3+ eigenvectors; the original B-Q (1989) paper had incorrect bootstrap error bands. (3) Sims (1986) 6-variable overidentified SVAR: money-demand responses are bimodal, naive Bayesian misses the asymmetry, and the price-puzzle magnitude is highly method-dependent.
"The likelihood principle, which is accepted by both Bayesians and many non-Bayesians, says that inference should be based only on the likelihood function… confidence intervals violate the likelihood principle." (p. 1113–1114)
"The naive Bayesian method produces neither Bayesian nor frequentist measures of error bands… it can be badly misleading when the likelihood function has multiple peaks." (p. 1147)
My Take
The likelihood-principle argument is crisp and largely persuasive: there is no good reason to prefer bootstrap Confidence Intervals (CIs) over Bayesian bands in a setting where the latter are computationally feasible. The eigendecomposition is the most novel methodological contribution — it reframes error-band reporting as a principal-component analysis of posterior IRF shape uncertainty, which is particularly valuable when the response path (not just its level at each horizon) is the object of interest. The correct SVAR posterior (Section 8A) is important but often overlooked in applied practice; the naive Bayesian method remains common despite its documented failure with multiple peaks. The main practical limitation is computational cost: 11 CPU hours for a 6-variable model in 1999, though this is now trivial on modern hardware.