Andersen and Bollerslev (1998) resolve the apparent paradox that generalized autoregressive conditional heteroscedasticity (GARCH) models with highly significant parameters produce very low values (0.02–0.05) in conventional forecast evaluation regressions. They show analytically that low is a mathematical implication of correctly specified, fat-tailed persistent variance processes — not evidence of misspecification. The core problem is that squared daily returns are an extremely noisy proxy for latent volatility: under GARCH(1,1) fitted to Deutsche Mark–US dollar (DM-$) exchange rates, measurement noise accounts for roughly 93% of regression residuals. Constructing realized volatility — cumulative squared intraday returns — sharply reduces this noise: at the 5-minute frequency (m=288 per day), rises from 0.05 to 0.48, close to the theoretical maximum.
Low is the correct prediction of a well-specified model. The population under GARCH(1,1) is derived analytically. For DM-$ (, , , excess kurtosis ) the formula gives ; the empirical estimate is . The critic's "damning" evidence is exactly what theory predicts.
Squared daily returns are a terrible proxy for latent variance. The forecast mean squared error (MSE) using decomposes into (a) the ideal variance of the Mincer-Zarnowitz regression residual under perfect foresight, plus (b) measurement noise. For DM-$: actual , ideal — measurement noise accounts for 93% of the regression residual. Any evaluation based on measures almost entirely noise.
Realized volatility converges to integrated variance. By quadratic variation theory (Karatzas-Shreve 1988), as . At 5-minute frequency the noise reduction factor equals , collapsing measurement error to near zero.
The improvement is dramatic and exactly as predicted by theory.
| Frequency | DM-$ theory | DM-$ empirical | ¥-$ theory | ¥-$ empirical | |
|---|---|---|---|---|---|
| Daily | 1 | 0.063 | 0.047 | 0.089 | 0.025 |
| 8-hour | 3 | 0.151 | 0.133 | 0.198 | 0.095 |
| Hourly | 24 | 0.383 | 0.331 | 0.419 | 0.237 |
| 5-minute | 288 | 0.483 | 0.479 | 0.488 | 0.392 |
Continuous-time framework links GARCH to latent volatility. Nelson (1990) showed GARCH(1,1) converges in distribution to , identifying structural parameters from . Temporal aggregation uses the weak GARCH theory of Drost-Nijman (1993) and Drost-Werker (1996).
"The low statistics that have plagued the volatility forecast evaluation literature are simply an artifact of the high noise-to-signal ratio in the standard forecast evaluation equations."
"The results are remarkably clear. The sample statistics increase dramatically with the sampling frequency."
"Yes, ARCH and stochastic volatility models do provide good volatility forecasts!"
This paper is the foundational document for the realized volatility literature that dominated financial econometrics throughout the 2000s. The core insight — evaluate GARCH forecasts against high-frequency realized volatility, not squared daily returns — seems obvious in retrospect but was not until this paper made the measurement-error decomposition explicit and quantified it. The analytical derivation of the population under GARCH is clean, and the empirical confirmation across two currencies at five sampling frequencies is compelling. The Nelson (1990) diffusion limit is essential: without it the connection between discrete GARCH parameters and continuous-time integrated variance would be unmotivated. One remaining puzzle: the 5-minute empirical (0.479 DM-$, 0.392 ¥-$) still falls short of the theoretical maximum (0.483, 0.488), hinting at microstructure noise in 5-minute returns — the concern that drove the noise-robust realized variance literature of the 2000s.