Gneiting-Raftery (2007) Strictly Proper Scoring Rules, Prediction, and Estimation

scoring-ruleprobabilistic-forecastingforecast-evaluationproper-scoring-rulecrpsenergy-scorebayes-factorcross-validationquantile-forecastloss-functionliterature-survey

Summary

This JASA review article develops the theory of proper scoring rules — numerical rewards S(P,x)S(P,x) that grade a probabilistic forecast PP against the outcome xx that materializes. A rule is proper if the forecaster maximizes expected score by quoting his or her true belief, and strictly proper if that optimum is unique; propriety is what makes a score a trustworthy tool for both evaluation (ranking competing forecasters/models) and estimation. The paper gives a general characterization of proper scoring rules via convex functions and Bregman divergences, catalogues the standard rules for categorical, density, quantile, and interval forecasts (logarithmic, Brier/quadratic, spherical, CRPS, energy score, interval score), and draws bridges to Bayes factors, cross-validation, and M-estimation. A weather-forecasting case study shows how an improper score can rank a deliberately dishonest forecast above an honest one.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"A scoring rule ... is proper if the forecaster maximizes the expected score for an observation drawn from the distribution FF if he or she issues the probabilistic forecast FF, rather than GFG\ne F. It is strictly proper if the maximum is unique."

"Maximum likelihood estimation forms a special case of optimum score estimation, and optimum score estimation forms a special case of M-estimation."

My Take

This is the reference that turned a scattered meteorology-and-decision-theory folklore into a clean, general theory, and its vocabulary (proper, strictly proper, CRPS, energy score, sharpness-subject-to-calibration) is now standard in Bayesian and econometric forecasting. For this wiki its most useful throughlines are two: (1) the log score = log Bayes factor identity, which makes Bayes-factor model comparison a special case of proper scoring and connects predictive evaluation to marginal-likelihood machinery; and (2) the CRPS/interval score as robust, honesty-enforcing alternatives to the raw MSE/coverage diagnostics used in volatility forecast evaluation and VaR backtesting. The propriety lens also clarifies why the QLIKE/MSE robustness results (Patton 2011) matter: an improper or proxy-sensitive loss can silently reward the wrong model.