Andersen-Bollerslev (1998) Answering the Skeptics

garchrealized-volatilityhigh-frequencyvolatility-forecastingmeasurement-errorcontinuous-timestochastic-volatility

Summary

Andersen and Bollerslev (1998) resolve the apparent paradox that generalized autoregressive conditional heteroscedasticity (GARCH) models with highly significant parameters produce very low R2R^2 values (\approx0.02–0.05) in conventional forecast evaluation regressions. They show analytically that low R2R^2 is a mathematical implication of correctly specified, fat-tailed persistent variance processes — not evidence of misspecification. The core problem is that squared daily returns are an extremely noisy proxy for latent volatility: under GARCH(1,1) fitted to Deutsche Mark–US dollar (DM-$) exchange rates, measurement noise accounts for roughly 93% of regression residuals. Constructing realized volatility — cumulative squared intraday returns — sharply reduces this noise: at the 5-minute frequency (m=288 per day), R2R^2 rises from \approx0.05 to \approx0.48, close to the theoretical maximum.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"The low R2R^2 statistics that have plagued the volatility forecast evaluation literature are simply an artifact of the high noise-to-signal ratio in the standard forecast evaluation equations."

"The results are remarkably clear. The sample R2R^2 statistics increase dramatically with the sampling frequency."

"Yes, ARCH and stochastic volatility models do provide good volatility forecasts!"

My Take

This paper is the foundational document for the realized volatility literature that dominated financial econometrics throughout the 2000s. The core insight — evaluate GARCH forecasts against high-frequency realized volatility, not squared daily returns — seems obvious in retrospect but was not until this paper made the measurement-error decomposition explicit and quantified it. The analytical derivation of the population R2R^2 under GARCH is clean, and the empirical confirmation across two currencies at five sampling frequencies is compelling. The Nelson (1990) diffusion limit is essential: without it the connection between discrete GARCH parameters and continuous-time integrated variance would be unmotivated. One remaining puzzle: the 5-minute empirical R2R^2 (0.479 DM-$, 0.392 ¥-$) still falls short of the theoretical maximum (0.483, 0.488), hinting at microstructure noise in 5-minute returns — the concern that drove the noise-robust realized variance literature of the 2000s.