Realized volatility (RV) is a nonparametric, model-free estimate of the latent integrated varianceIVt=∫01σt+s2ds constructed by summing squared intraday returns sampled at frequency m:
RVt(m)=j=1∑mrt,j/m2
By the quadratic variation theorem (Karatzas-Shreve 1988), RVt(m)p∫01σt+s2ds as m→∞ under a continuous semimartingale price process. At 5-minute intervals (m=288 per trading day), the measurement noise is reduced by a factor of m relative to squared daily returns, making RV the standard benchmark for evaluating and comparing volatility models.
Key Ideas
Squared daily returns are nearly pure noise. Under generalized autoregressive conditional heteroskedasticity (GARCH)(1,1) fitted to DM-$ exchange rates, measurement noise accounts for ≈93% of the residual variance in Mincer-Zarnowitz forecast evaluation regressions — the low R2 is a property of the criterion, not a flaw in the model (Andersen-Bollerslev 1998).
RV dramatically improves forecast evaluation. Replacing rt2 with 5-minute RV raises R2 from ≈0.05 to ≈0.48, close to the population maximum predicted by GARCH(1,1).
Returns standardized by RVt are approximately N(0,1). Andersen, Bollerslev, Diebold, and Labys (ABDL, 2001) document this across equity, foreign exchange (FX), and bond markets, implying the heavy-tailed unconditional distribution of returns is driven entirely by time-varying volatility — not by a fixed fat-tailed innovation distribution.
Microstructure noise limits sampling frequency. Bid-ask bounce, price discreteness, and asynchronous trading contaminate very-high-frequency returns. The optimal frequency trades off noise reduction against microstructure contamination, typically yielding 5–15 minute sampling for equity and FX data.
RV is the most precise of several daily volatility proxies. In the measurement hierarchy: RVt (5-minute) ≻ daily high-low range ≻∣rt∣. RV Granger-causes the noisier proxies but not vice versa.
Model realized volatility directly (ABDL 2003): because RV is (nearly) observable, log realized volatilities can be modeled with standard linear Gaussian time series rather than latent-variable filtering. A simple trivariate long-memory (fractionally-integrated) Gaussian VAR for the log realized variances and covariance out-forecasts daily GARCH and RiskMetrics, and — combined with a lognormal-normal mixture — yields well-calibrated density and VaR forecasts.
How It Works
Quadratic Variation Foundation
Under a continuous semimartingale dpt=μtdt+σtdWt (no jumps), the quadratic variation over [t−1,t] equals the integrated variance:
For DM-$ data: model imperfection =0.084, measurement noise =1.137, total =1.221 — noise accounts for 93% of the residual. The population R2 under GARCH(1,1) is determined entirely by the signal-to-noise ratio in the criterion variable, not by the model's forecasting accuracy.
When RV at frequency m replaces rt2, the measurement noise shrinks by the factor m. The R2 improvement is precisely predicted by theory:
Frequency
m
Theory R2 (DM-$)
Empirical R2 (DM-$)
Daily
1
0.063
0.047
8-hour
3
0.151
0.133
Hourly
24
0.383
0.331
5-minute
288
0.483
0.479
The convergence of empirical to theoretical R2 across five frequencies is itself confirmatory evidence of correct GARCH specification.
Microstructure Noise and Optimal Sampling
The observed log-price is p~t,j=pt,j+et,j, where et,j is i.i.d. microstructure noise with variance ξ2 (bid-ask bounce, discreteness, asynchrony). The bias in RV is:
E[RVt(m)]=IVt+2mξ2
which grows without bound as m→∞ — the "signature plot" signature. The optimal frequency m∗ minimizes MSE and corresponds empirically to 5–15 minute intervals. Noise-robust alternatives include:
Bipower variation (Barndorff-Nielsen–Shephard 2004): BVt=2π∑j∣rt,j∣∣rt,j−1∣, consistent for IVt even with finite-activity price jumps. The difference RVt−BVt estimates total jump variation.
Kernel-based RV (Barndorff-Nielsen et al. 2008): applies a smooth weight function to realized autocovariances of intraday returns, removing microstructure bias while remaining consistent for IVt.
Pre-averaging (Jacod et al. 2009): averages returns over a local window before squaring; achieves the optimal m1/4 convergence rate under noise.
Log-Normality of RV (ABDL 2001)
Across equity, FX, and bond markets, logRVt is approximately Gaussian. This motivates the log-normal mixture for the conditional return distribution:
log(σt)∣Ft−1∼N(μt,τ2),rt∣σt∼N(0,σt2)
where σt2≈RVt. This model generates well-calibrated value at risk (VaR) coverage rates, outperforming GARCH with normal innovations by exploiting the observed near-normality of rt/RVt — see Value at Risk.
The autocorrelation function (ACF) of logRVt decays hyperbolically, motivating autoregressive fractionally integrated moving average (ARFIMA) models for RV forecasting (fractional d≈0.4) — see Long Memory and Fractional Integration.
RV Sampling Frequency and Monthly Regression Bias (Bollerslev-Zhou 2006)
Bollerslev and Zhou (BZ, 2006) use a Monte Carlo study (Table 2) to quantify how the choice of intraday sampling frequency affects the three return-volatility regression slopes at monthly frequency. Key findings:
5-minute RV introduces negligible measurement error for monthly volatility feedback, leverage, and implied-vol forecasting regressions. Biases are near zero and comparable to those from the infeasible integrated variance itself.
Daily-squared-return RV introduces large biases. For the volatility feedback slope, the bias toward zero is ~0.08 at T=150 and barely shrinks to 0.078 at T=600 — the bias is not alleviated by larger samples. This is because daily-squared-return noise is a fixed fraction of the monthly integrated variance rather than a measurement-error component that shrinks with T.
Practical implication: monthly return-volatility regressions should use 5-minute (or finer) intraday realized vol; using daily returns severely understates the predictability of volatility by lagged returns.
This result complements Andersen-Bollerslev (AB, 1998): where AB (1998) showed that noise in the criterion variable (squared daily return as volatility proxy) inflates GARCH forecast evaluation residuals, BZ (2006) show that noise in the regressor (daily RV as integrated variance proxy) induces persistent slope biases.
RV as a Volatility Proxy in the MIM (Engle-Gallo 2006)
The Multiple Indicators Model (MIM) treats three daily volatility proxies as a joint vector multiplicative error model (MEM): absolute returns ∣rt∣, daily high-low range hlt, and realized volatility vt. Bayesian information criterion (BIC)-selected cross-equation lags run from vt to the noisier proxies, but not in reverse — confirming the Granger-causality hierarchy that maps to the measurement precision ordering. See MEM and MIM for the full system specification and multi-step forecasting dynamics.
Why It Matters
Resolves the GARCH forecast-evaluation paradox.R2≈0.02–0.05 in squared-return regressions is exactly what a well-specified GARCH(1,1) predicts. The Andersen-Bollerslev (1998) measurement-error analysis closed a decade of incorrect skepticism about GARCH forecasting ability.
Provides a model-free variance proxy. RV requires no parametric model; it is a direct observation of IVt up to sampling error. This makes it usable as a dependent variable in forecasting regressions, as a real-time risk signal, and as a neutral benchmark for model comparison.
Enables distributional regularities. The near-normality of rt/RVt is a sharp empirical regularity that strongly constrains the class of admissible models: the non-Gaussian behavior of raw returns must come from time-varying volatility alone.
Identifies jump risk. The gap RVt−BVt consistently estimates the jump component of quadratic variation, enabling model-free tests for jump activity and jump timing.
Realized Power and MIDAS Volatility Forecasting (Ghysels-Santa-Clara-Valkanov 2006)
Ghysels, Santa-Clara, and Valkanov (2006) compare five daily volatility predictors — realized variance Q~(m), squared returns r2, absolute returns ∣r∣, daily range [hi-lo], and realized power P~(m)=∑j=1m∣rt−(j−1)/m∣ — within the Mixed Data Sampling (MIDAS) distributed-lag regression, benchmarked against the ABDL ARFI(5,d) long-memory model (see MIDAS Regression).
Realized power strictly dominates all other predictors at every forecast horizon H=1,5,10,20 days and for every asset examined (DJ index + DIS, GE, JPM, XOM, MCD, HON). In-sample MSE ratios vs. ABDL range 0.606–0.912 in levels; out-of-sample 0.714–0.897. The predictor ranking is:
P~(m)≻[hi-lo]≻Q~(m)≻∣r∣≻r2
Why realized power outperforms realized variance: Under Barndorff-Nielsen–Shephard (2004) bipower theory, P~(m) converges to the continuous component of quadratic variation (integrated variance without jumps), while Q~(m) includes both continuous and jump parts. Because jumps are transient, their inclusion in the predictor reduces persistence and degrades multi-step forecasts. logP~(m) has higher ACF persistence and less measurement noise than logQ~(m).
No gain from direct intraday use: Using raw 5-min returns as MIDAS regressors (rather than daily realized aggregates) does not improve MSE at any horizon — the relevant information in high-frequency data is fully captured by daily realized power.
RV Inconsistency under Lévy Time-Change (Barndorff-Nielsen-Shephard 2006)
When log-prices follow a time-changed Lévy process Yt=mt+Zτt with Z non-Brownian, RV is an inconsistent estimator of the time-changeτi. The quadratic variation of Y includes a jump contribution, so [Y]d,ip[Y]i=k2τi regardless of sampling frequency. Specifically:
The variance of the RV error [Y]d,i−k2τi contains an irreducible termk4Δξ driven by the 4th cumulant k4 of the Lévy increments, which does not vanish as M→∞
The biasE([Y]d,i−k2τi)=O(M−1) is small; squared bias is O(M−2) and dominated by the O(1) variance in MSE
The ACF inequality: Cor([Y]d,i,[Y]d,i+s)≤Cor(τi,τi+s) when k4>0 — RV underestimates the autocorrelation of the true variance process, so actual volatility predictability exceeds what RV-based studies show
For a symmetric Lévy process, k4=0⟺Z is Brownian motion (by Lévy-Khintchine); thus testing k4=0 is a jump test
The practical implication is that the high predictability of RV found in Andersen-Bollerslev (1998) and subsequent work is a lower bound on the true predictability of τi. See Time-Changed Lévy Process for the full framework.
Open Questions
Multivariate realized covariance raises the Epps effect: non-synchronous trading across assets causes realized covariances to shrink toward zero at high sampling frequency. Refresh-time sampling and Hayashi-Yoshida estimators partially address this, but scalability to large cross-sections remains challenging.
Long memory vs. structural breaks in logRVt. The hyperbolic ACF decay could reflect genuine fractional integration (d≈0.4) or aggregation of short-memory components across volatility regimes. Disentangling these is unresolved in finite samples.
Jumps vs. continuous variation. Finite-sample power of RV−BV jump tests is limited, especially for small jumps. Jump-robust estimators (quantile-based, truncated RV) offer complementary approaches.