Geweke and Amisano compare and evaluate Bayesian predictive distributions of asset returns, illustrated with five models applied to daily S&P 500 returns from 1976 through 2005. Bayesian inference gives exact out-of-sample predictive distributions that fully and coherently incorporate parameter uncertainty. The comparison exercise uses predictive likelihoods and is inherently Bayesian; the evaluation exercise uses the probability integral transform and is inherently frequentist. The two approaches are shown to be complementary — each identifying strengths and weaknesses in models not evident from the other.
"The comparison exercise uses predictive likelihoods and is inherently Bayesian. The evaluation exercise uses the probability integral transform and is inherently frequentist. The illustration shows that the two approaches can be complementary, each identifying strengths and weaknesses in models that are not evident using the other."
A clarifying methodological point dressed as an empirical illustration: comparison and evaluation are not the same, and the Bayesian log-predictive-score answers the first while the frequentist PIT answers the second. The insistence that the winning model still be checked for absolute adequacy is the right discipline — it is exactly the gap the model-comparison machinery leaves open, since a Bayes factor only ever ranks. The Bayesian predictive densities, integrating over parameter uncertainty rather than plugging in estimates, are also the honest object to evaluate. It slots cleanly beside proper scoring rules (the comparison metric) and calibration (the reliability analogue of the PIT), and the five-model S&P 500 horse race is a compact template for density-forecast practice.