Ghysels-Santa-Clara-Valkanov (2006) Predicting Volatility: Getting the Most Out of Return Data Sampled at Different Frequencies

midasvolatility-forecastinghigh-frequencyrealized-volatilityforecastingrealized-powerdistributed-lagbeta-polynomiallong-memoryempirical-finance

Summary

Introduces the MIDAS (Mixed Data Sampling) regression framework for forecasting volatility using return data at arbitrary sampling frequencies. The central finding is that realized power — the sum of absolute intraday returns — is the single best volatility predictor across all tested specifications and horizons, outperforming the long-memory Andersen-Bollerslev-Diebold-Labys (ABDL) autoregressive fractionally integrated ARFI(5,d) benchmark by 10–40% in mean squared error (MSE). Surprisingly, directly using 5-minute intraday data in a MIDAS regression does not improve on daily aggregates. Beta-polynomial weighting captures the relevant lag structure with only two free parameters.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"Among the five predictors considered, the realized power variation… clearly dominates all other predictors."

"This result is due to the fact that the realized power variation… removes the jump component of the quadratic variation from the predictor."

"Although the use of intra-daily data to compute the predictors yields better estimates of daily integrated variance, it does not yield better volatility forecasts."

My Take

The realized power finding is the sharpest result: it is theoretically motivated (Barndorff-Nielsen–Shephard 2004 bipower theory isolates the continuous component), empirically consistent across DJ + 6 stocks, and robust in- and out-of-sample. The MIDAS framework itself is more general than this application — the 2-parameter Beta polynomial is a practical engineer's tool, not a theoretically derived weighting scheme, so the choice of kmaxk^{\max} and mm involves informal judgment. The "no gain from direct intraday use" finding is surprising but potentially specific to this aggregation: the paper's direct MIDAS uses raw 5-min returns without pre-averaging microstructure noise (Zhang-Mykland-Aït-Sahalia 2005), so the null result may reflect noise rather than genuine information equivalence.