The statistical problem of evaluating and comparing volatility forecasts when conditional variance is a latent variable that cannot be observed directly. Because ht=logσt2 is never revealed, standard accuracy metrics (mean squared error (MSE), R2) must be adapted or replaced with proxy-robust or proxy-free alternatives.
Key Ideas
Latency of the target. Unlike forecasting inflation or gross domestic product (GDP), there is no ex-post realization of σt2. Any evaluation must work around or through a noisy proxy.
Proxy noise problem. The most natural proxy — the squared daily return rt2 — has enormous measurement error. Under any location-scale model, rt2=σt2εt2 where εt2∼χ2(1) has variance 2, swamping the signal.
R2 puzzle resolved.Andersen-Bollerslev (1998) showed generalized autoregressive conditional heteroskedasticity (GARCH) appeared to have R2≈0.05 using daily rt2 as proxy — not a model failure but pure noise inflation. Replacing the proxy with 5-minute realized volatility recovers R2≈0.48.
Realized volatility.RVt=∑j=1Mrt,j2M→∞∫tt+1σs2ds. Measurement error is O(1/M); at 5-minute sampling (M≈78 for a 6.5-hour day) it is small enough to be nearly negligible.
Predictive density evaluation. Log-score ∑tlogp(rt∣Ft−1) requires no proxy. It is the natural metric for fully specified probabilistic models (GARCH, stochastic volatility (SV), SV-DPM), and corresponds directly to log Bayes factors and marginal likelihoods in Bayesian comparisons. Used in Kim-Shephard-Chib (1998) to decisively favor SV over GARCH.
Robust loss functions.Patton (2011) showed that MSE and QLIKE (h^/h−log(h^/h)−1) preserve the ranking of forecasters under any conditionally unbiased proxy; mean absolute error (MAE) does not. This means MSE/QLIKE comparisons using imperfect proxies are still valid.
Proxy-dependent regression slopes.Bollerslev-Zhou (2006) showed in the Heston model that return–volatility regression slopes change sign depending on which proxy is used. Structural features of continuous-time models determine the proxy-bias direction, not model misspecification.
How It Works
With a realized-volatility proxy:
Compute RVt at a chosen sampling frequency (5-min is standard to balance micro-structure noise vs. estimation variance).
Regress or compare model-forecast σ^t2 against RVt.
Loss functions MSE and QLIKE computed against RVt give the same forecast ranking as if true σt2 were observed (Patton 2011).
Proxy-free (predictive density / likelihood):
For parametric models, compute the one-step predictive density p(rt∣Ft−1,θ^).
Sum log-scores: LS=T1∑tlogp(rt∣Ft−1).
Differences in LS equal log Bayes factor contributions; Diebold-Mariano test applies.
Regress proxy V^t on 1 and forecast σ^t2. Efficient forecast: intercept =0, slope =1.
Under rt2 proxy, slope estimates are biased; use RVt to avoid Bollerslev-Zhou distortion.
Why It Matters
Motivated the move from daily return squares to intraday realized measures as the benchmark for evaluating all volatility models; changed conclusions about GARCH accuracy.
Log-score / predictive likelihood became the dominant model-selection criterion in Bayesian SV literature, replacing frequentist in-sample AIC/BIC.
Patton (2011) robust loss functions provided a theoretical underpinning for industry practice of evaluating forecasts against imperfect but widely available proxies (e.g., VIX2 as a proxy for expected variance).
Ghysels-Santa-Clara-Valkanov (2006) Mixed Data Sampling (MIDAS) approach exploits the same logic: more frequent data reduces proxy noise, improving estimation precision at daily or lower frequencies.
Open Questions
Microstructure noise. At very high sampling frequencies (M≫100), bid-ask bounce and price discretization inflate RVt. Kernel-based (Barndorff-Nielsen et al.) and pre-averaging estimators address this but add complication.
Jump contamination. If the price process has jumps, RVt is not a pure integrated variance (IV) proxy; bipower variation BVt separates continuous and jump components.
Nonparametric tails. Log-score evaluation presupposes a fully specified return density. Models with nonparametric tails (SV-DPM) have heavier tails than Gaussian SV, yielding higher log-scores by construction; whether this reflects genuine forecasting improvement or tail overfitting is debated.
Long horizons.RV compounds cleanly for short horizons; multi-period volatility forecasting compounds non-linearly under stochastic volatility, and proxy construction at h>1 is less settled.