Diebold and Mariano propose explicit tests of the null hypothesis of no difference in the accuracy of two competing forecasts. Unlike earlier procedures, the tests accommodate a wide variety of accuracy measures — the loss function need not be quadratic or even symmetric — and forecast errors that are non-Gaussian, non-zero-mean, serially correlated, and contemporaneously correlated. They develop an asymptotic test based on the mean loss differential and its long-run variance, plus exact finite-sample sign and Wilcoxon signed-rank tests, and evaluate their size and power.
"We propose and evaluate explicit tests of the null hypothesis of no difference in the accuracy of two competing forecasts. In contrast to previously developed tests, a wide variety of accuracy measures can be used (in particular, the loss function need not be quadratic, and need not even be symmetric), and forecast errors can be non-Gaussian, non-zero mean, serially correlated, and contemporaneously correlated."
This short paper defined how a generation compares forecasts: reduce two competitors to a single loss-differential series and test its mean, with a long-run-variance correction for the serial dependence that multi-step errors always have. Its lasting strength is loss-function agnosticism — evaluation can finally match the decision problem instead of defaulting to MSE. The now-standard caveats came later and are catalogued in the concept page: the asymptotics get delicate for nested models with estimated parameters (West 1996; Clark-McCracken) and the statistic over-rejects in small windows (Harvey-Leybourne-Newbold correction). Diebold himself later stressed that DM tests a property of forecasts, not of models — the distinction that keeps its application honest.