Christoffersen (1998) Evaluating Interval Forecasts

interval-forecastvalue-at-riskbacktestingconditional-coverageforecast-evaluationlikelihood-ratio-testvolatilityrisk-management

Summary

This paper builds the standard framework for evaluating interval forecasts (prediction intervals) under conditional heteroskedasticity, and is the foundation of modern value-at-risk backtesting. Christoffersen observes that the existing literature tests only for correct unconditional coverage — whether the interval contains the outcome the right fraction of the time on average — which is inadequate when volatility is time-varying, because a forecaster can hit the right average coverage while getting the timing of violations badly wrong (clustered violations in turbulent periods). He formalizes conditional coverage through the "hit" indicator sequence and derives simple likelihood-ratio tests, illustrated on exchange-rate risk-management interval forecasts.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"Most of the literature implicitly assumes homoskedastic errors even when this is clearly violated, and proceeds by merely testing for correct unconditional coverage. Consequently, I set out to build a consistent framework for conditional interval forecast evaluation."

My Take

The elegance is in the reduction: no matter how complicated the forecasting model, its interval forecasts are judged entirely through a sequence of 0/1 hits, and "a good interval forecast" becomes the crisp, testable statement "the hits are i.i.d. Bernoulli(pp)." Splitting that into a coverage part and an independence part is what makes the test diagnostic rather than merely pass/fail — it tells you how the model is wrong. For the wiki this is the evaluation counterpart to the VaR estimation machinery: the HS-VaR pathologies documented there (right average coverage, disastrous conditional coverage after a volatility spike) are exactly what LRindLR_{ind} is built to catch. The main limitation, addressed by later work (Engle-Manganelli's DQ test, Berkowitz density tests), is that the first-order-Markov independence alternative has low power against richer forms of violation dependence.