GARCH (Generalized Autoregressive Conditional Heteroscedasticity) models the conditional variance of a time series as a function of past squared innovations and past conditional variances. The Baba-Engle-Kraft-Kroner (BEKK) parametrization (Engle and Kroner 1995) extends GARCH to multivariate systems while guaranteeing positive definiteness of the conditional covariance matrix at every time step.
ARCH() — Engle (1982): Let with . The conditional variance follows
GARCH() — Bollerslev (1986): Augments ARCH with autoregressive variance terms:
Covariance stationarity. A GARCH(1,1) process has finite unconditional variance if and only if . The unconditional variance is then .
IGARCH. When the process is integrated in variance (IGARCH): shocks to variance are permanent and the unconditional variance does not exist. IGARCH still admits well-defined multi-step conditional variance forecasts converging to at rate .
Multi-step forecasting (Engle 2001). For a covariance-stationary GARCH(1,1) with , the -step-ahead conditional variance forecast has the closed-form geometric mean-reversion recursion:
where is the long-run (unconditional) variance. For this reduces to the one-step GARCH recursion; as forecasts converge to at geometric rate . When persistence is near unity (e.g., ), a volatility spike takes hundreds of periods to half-revert, producing the empirically observed long memory in squared returns without requiring fractional integration.
Empirical (Engle 2001 — Dow Jones Industrial Average (DJIA) + bond portfolio, 1990–2000): , , ; ; (daily SD ). The 1% daily Value at Risk (VaR) on a $1M portfolio ranged from ≈$8,400 (low-volatility periods) to ≈$52,900 (high-volatility periods) — a six-fold range driven entirely by time-varying .
The basic GARCH(1,1) model treats the long-run variance as a fixed constant. Engle and Lee (1999) allow it to be time-varying by writing the GARCH(2,2) model in a component form:
where is the long-run (trend) variance component and is the short-run (transitory) component. The long-run component follows its own AR(1) with persistence , while governs the speed of mean-reversion to . When , the long-run component moves slowly and the model approximates long memory in volatility (Andersen-Bollerslev 1997), which is routinely found in asset return variances. Gallant, Hsu, and Tauchen (1999) and Alizadeh, Brandt, and Diebold (2002) provide empirical evidence for this component structure.
For an -dimensional innovation vector with conditional covariance , the BEKK(1,1) model (Engle and Kroner 1995) is:
where:
Parameter count. Full BEKK(1,1) has parameters. For : 3 + 8 = 11 parameters.
Diagonal BEKK. Restrict and . This eliminates cross-equation volatility spillovers. For : 3 + 4 = 7 parameters. The diagonal restriction is testable by a likelihood-ratio statistic with degrees of freedom.
General BEKK():
The BEKK(1,1) process is covariance stationary (has finite unconditional covariance matrix) if and only if all eigenvalues of
lie strictly inside the unit circle, i.e., . When the unconditional covariance matrix does not exist — the multivariate analog of IGARCH.
Empirical finding (BDV 1998). Estimating BEKK-GARCH for bivariate long/short interest rate innovations in Belgium, Germany, France, UK, and USA, Bauwens, Deprins, and Vandeuren (1998) find in all five countries, implying multivariate IGARCH behavior: unconditional interest rate volatility does not exist, though conditional forecasts are well-defined for short horizons of one to three months.
When in GARCH(1,1), shocks to variance are permanent (integrated GARCH, Engle and Bollerslev 1986). The special case with gives the EWMA formula:
which is an exponentially weighted moving average of past squared innovations with decay factor . This is the formula underlying the RiskMetrics VaR methodology. Under IGARCH/EWMA the unconditional variance does not exist, but short-horizon conditional variance forecasts converge at rate per step.
The covariance-stationarity condition ensures finite unconditional variance but is not necessary for the process to be well-defined and stationary in distribution. Nelson (1990b) established the exact condition for strict stationarity of GARCH(1,1): the process is strictly stationary (and ergodic) if and only if
where . Because by Jensen's inequality, the strict-stationarity condition is strictly weaker than . In particular, IGARCH () can satisfy the strict-stationarity criterion.
Persistence ambiguity. IGARCH is strictly stationary yet has as for . This means variance shocks never die out in expectation even though the process has a well-defined stationary distribution. Strict stationarity and moment convergence are distinct notions: IGARCH is "stationary" in one sense and "non-stationary" in another. Any empirical finding of IGARCH should therefore distinguish between these senses — a model that is strictly stationary but has diverging conditional variance forecasts is economically very different from a covariance-stationary model.
Practical implication. Near-IGARCH estimates () are common. Nelson's theorem explains why: the strict-stationarity region extends well above , so such processes can be legitimate probability models even when they lack finite unconditional variance. See also the spurious-IGARCH caveat in Open Questions.
Kleibergen and Van Dijk (1993) sharpen the weak vs. quasi-strict distinction using Bayesian inference on US T-bill data.
Weak vs. quasi-strict. The weak condition is strictly more restrictive than the quasi-strict condition : the stationary region under the weak criterion is a proper subset of the quasi-strict stationary region.
Convergence theorem. The conditional variance product converges to 0 a.s. iff , and diverges to otherwise. The mechanism is the law of large numbers on the log scale: a.s. The weak condition incorrectly uses and overestimates non-stationarity.
IGARCH revisited. At with normal errors, : the conditional variance decreases more often than it increases, even though . IGARCH satisfies the quasi-strict condition despite violating weak stationarity.
Empirical application (US 3-month T-bill, Jan 1957–Apr 1989, ). AR(1)-GARCH(1,1)-t model estimated by importance sampling (SISAM, Hop-Van Dijk 1992). Bayesian posterior mean: , , .
| Stationarity criterion | Bayes factor | |
|---|---|---|
| Weak () | 4.95 | |
| Quasi-strict () | 0.434 |
The weak criterion overstates non-stationarity by nearly 8× relative to the quasi-strict criterion.
Fat-tail / unit-root interaction. Posterior odds for a unit root in the AR mean coefficient :
| Model | Conclusion | |
|---|---|---|
| Constant variance, normal | 0.17 | Reject unit root |
| GARCH(1,1), normal | 0.04 | Strongly reject |
| Constant variance, t | 2.30 | Support unit root |
| GARCH(1,1), t | 2.50 | Support unit root |
The AR and tail-thickness parameters are negatively correlated in the posterior: fat-tailed models attribute extreme observations to heavy tails rather than high AR persistence, freeing to concentrate near 1. The t-score function decreases for large , downweighting outliers. Normal likelihoods treat the same outliers as high-volatility evidence, biasing the posterior toward stationarity in mean.
Different ARCH families, when sampled at increasingly fine time intervals, converge to different continuous-time diffusions. This establishes structural links between ARCH models and stochastic-volatility option-pricing models.
GARCH(1,1) limit (Nelson 1990a; Drost-Nijman 1993). As the sampling interval with GARCH parameters rescaled appropriately, the GARCH(1,1) variance process converges to the mean-reverting stochastic differential equation (SDE):
where is the speed of mean-reversion and is a Brownian motion. This is a special case of the Heston (1993) stochastic-volatility model (with zero correlation between and the return Brownian motion). The unconditional distribution of the GARCH(1,1) process converges to an inverted-gamma mixture of normals, producing Student- tails — consistent with the fat tails documented in asset return data.
EGARCH limit. The EGARCH (Exponential GARCH) process converges to a diffusion where follows an Ornstein-Uhlenbeck process, giving a log-normal unconditional distribution for . The resulting returns follow a normal–lognormal mixture. This provides a theoretical connection between EGARCH and the log-normal stochastic volatility model used in continuous-time finance.
ARCH-family selectivity. Different ARCH specifications (power-GARCH, TARCH, etc.) converge to different diffusions. The choice of ARCH model is therefore equivalent to choosing a specific continuous-time volatility specification, which matters for option pricing (Engle-Mustafa 1992) and for the interpretation of the unconditional return distribution.
Option pricing via diffusion limits (Bollerslev-Engle-Nelson 1994). Because the continuous-time limit of a GARCH process has a known form, the Black-Scholes PDE evaluated at the GARCH-implied diffusion provides a structural link between ARCH-based conditional variance estimates and option prices — a connection exploited in ARCH-based option pricing (Engle-Mustafa 1992; Duan 1995).
A natural question is whether a GARCH model at daily frequency implies GARCH at weekly or monthly frequency. The answer depends on whether "GARCH" is interpreted in the strict or weak sense.
Weak GARCH. The weak GARCH class (Drost-Nijman 1993) defines conditional heteroscedasticity only through the second-order moment structure — it does not require the conditional distribution to be Gaussian or even i.i.d. innovations. Weak GARCH is closed under temporal aggregation: if is weak GARCH() at frequency , then the aggregated series is weak GARCH at frequency , with parameters that are explicit functions of the original parameters and .
The aggregation formula implies that the persistence is preserved exactly under aggregation: if daily, the weekly series has the same persistence. What changes is the allocation between the ARCH and GARCH components — aggregation concentrates persistence into the GARCH term and reduces the ARCH term .
Heteroscedasticity fading. As (very low frequency), the ARCH coefficient of the aggregated series and the GARCH coefficient of the original, so the conditional variance becomes nearly constant at the long-run value . If , heteroscedasticity disappears at sufficiently low frequencies — consistent with the empirical observation (Akgiray 1989; BEN 1994) that daily returns show strong GARCH while monthly returns are approximately i.i.d. normal.
Strict GARCH is not closed. If the daily process is strict GARCH (conditional distribution Gaussian), the weekly aggregate is generally not strict GARCH. Aggregation introduces leptokurtosis at the weekly level that cannot be fully captured by a GARCH model with Gaussian innovations. This is why Gaussian GARCH tends to underestimate kurtosis at intermediate frequencies.
Given an underlying continuous-time diffusion , what discrete-time ARCH model minimises the asymptotic mean-squared error of the resulting volatility estimate?
Optimality criterion. Nelson and Foster (1994) frame this as a filter-design problem: among all ARCH processes of the form , find the one for which the asymptotic estimation error variance is minimised. The answer depends on the assumed shape of the conditional distribution.
Result: absolute-value GARCH. When the true conditional distribution of has heavy tails (e.g., Student- or any distribution with excess kurtosis), the optimal filter uses absolute-value residuals rather than squared residuals . This provides a formal theoretical justification for TARCH (Zakoïan 1991) and NARCH/Power-ARCH (Higgins-Bera 1992) over standard GARCH when innovations are fat-tailed.
Intuition. Large squared residuals from heavy-tailed distributions carry excessive noise relative to their information content. Absolute-value residuals weight large observations less aggressively, reducing the estimator's variance. In the extreme (Cauchy innovations), squared residuals are meaningless, and only absolute values carry useful information about scale.
Empirical relevance. BEN (1994) and Hentschel (1995) find estimated power parameters on US equity data — intermediate between absolute-value () and squared () — consistent with Nelson-Foster optimality for moderately fat-tailed distributions.
Just as I(1) series can be co-integrated (a linear combination is stationary in the mean), IGARCH series can be co-persistent in variance: a portfolio has no variance persistence even though each individual series does. The co-persistent vector plays the exact role of the co-integrating vector.
Formal definition. Let measure the influence of time- shocks on future covariance forecasts. The process is persistent in variance if a.s. It is co-persistent if a nonzero exists such that a.s.
Theorem 2. Vector GARCH() is co-persistent iff for all right eigenvectors corresponding to . The co-persistent portfolio eliminates the explosive subspace.
Factor GARCH connection (Lemma 2). In a -factor GARCH model with integrated factors and stationary factors, any portfolio orthogonal to the integrated factors is co-persistent. Co-persistence exists whenever the number of persistent variance factors is less than .
Empirical finding. Daily DM and BP vs. USD (1980–1985, 1,245 obs): both near-IGARCH (); estimated co-persistent vector — the DM/BP bilateral rate. Univariate GARCH for DM/BP: t-stat for equals 6.033 — strongly rejects IGARCH. Interpretation: persistence in USD exchange rate volatility is dollar-specific news; the bilateral European rate shows no variance persistence.
One of the first rigorous GARCH applications. Data: CRSP value-weighted daily index returns, Jan 1963 – Dec 1986 (T = 6,030), split into four 6-year sub-periods.
Key diagnostic finding. After AR(1) pre-filtering (, –), standard independence tests on fail to reject white noise. But and remain significantly autocorrelated at all lags up to 60 — conclusive evidence of nonlinear (non-strict-white-noise) dependence. This is the diagnostic argument for GARCH over linear models.
Estimation. GARCH(1,1) strictly dominates ARCH() (–) by log-likelihood in every period. Representative GARCH(1,1) estimates: –, –, –. Dickey-Fuller rejects in 3 of 4 periods (covariance-stationary, but near-IGARCH). in all periods — past conditional variance dominates new shocks, producing persistent volatility. Standardized residuals pass normality tests; the GARCH model fully absorbs the excess kurtosis of raw returns.
The autocorrelation function (ACF) of closely follows the GARCH(1,1) recursion — 48 of 60 autocorrelations lie within of the implied values.
Volatility forecasting comparison. 24 rolling out-of-sample monthly variance forecasts per period, compared across four methods:
| Method | Description |
|---|---|
| Historical mean | Unconditional sample variance over estimation window |
| EWMA | Exponentially weighted moving average, smoothing constant w estimated from data (w ≈ 0.11–0.24) |
| ARCH | h-step-ahead forecast from fitted ARCH(p) via iterated expectations |
| GARCH | Same formula, , from GARCH(1,1) |
GARCH wins on mean error (ME), root-mean-square error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE) in all four periods. Gains are largest in high-volatility periods (1969-74, 1975-80). Even so, minimum MAPE > 30% across all methods — variance forecasting is inherently difficult.
Frequency dependence. GARCH effects are a daily phenomenon: monthly returns are approximately i.i.d. normal (no significant autocorrelation in {e²} at monthly frequency), likely reflecting a central limit theorem (CLT) for sums of dependent daily returns.
The ARCH-M model (Engle, Lilien, and Robins 1987) augments the conditional mean with a function of conditional variance:
Common choices: (standard deviation), (variance), or . The parameter captures the risk premium: investors demand higher expected returns when conditional risk is higher. Engle et al. (1987) found a significant for excess returns of 6-month US Treasury bills. This model is the first formal link between ARCH volatility and the expected return level.
Fiorentini and Sentana (1998) show that when , the conditional mean itself follows a vector linear process with the same AR polynomial:
Applied to multivariate GARCH, the conditional variance is a stationary ARMA process for past squared innovations and past variances; the Fiorentini-Sentana result pins down the ACF structure. For GARCH the conditional variance displays the ACF of a vector autoregressive moving-average (VARMA) process with the same AR polynomial and . The GARCH(1,1) special case is particularly clean:
where is the innovation to squared returns. This is a vector autoregression (VAR(1)) for the conditional covariance matrix — the one-period persistence is exactly, consistent with the GARCH recursion.
The GARCH-M model augments the conditional mean with conditional variance: . Because is proportional to conditional variance, the autocorrelation structure of is entirely determined by the autocorrelation structure of , which is always nonneg:
This is a testable structural restriction: GARCH-M processes can only generate nonneg autocorrelations in the observable series . Financial returns routinely exhibit near-zero or weakly negative autocorrelations, which explains why GARCH-M models tend to fit financial time series poorly despite their theoretical appeal as risk-return models. The constraint is exact (not approximate) and applies to all parametrisations of the variance function as long as with .
Bollerslev, Engle, and Wooldridge (1988) proposed the vech model for an -vector of returns :
where stacks the lower triangle of a symmetric matrix ( elements). Diagonal restrictions on and reduce the parameter count but eliminate cross-variance dynamics. With (bills, bonds, stocks), BEW show that Capital Asset Pricing Model (CAPM) -coefficients are time-varying — the conditional covariance with the market portfolio fluctuates with .
A practical limitation: the vech model does not automatically guarantee ; extra parameter constraints are required. This motivated the BEKK form (see above).
For large (many assets), the BEKK and vech parameter counts become prohibitive. The factor-ARCH model (Engle 1987; Engle, Ng, and Rothschild 1990) imposes a factor structure:
where () is a loadings matrix, is the number of common volatility factors, is constant (), and with each following an ARCH process. The conditional covariance matrix is:
Common volatility factors drive all cross-asset covariances simultaneously; only ARCH equations (plus and ) need to be estimated. Applied to Diebold-Nerlove (1989) daily exchange rates and to Treasury bills of different maturities (Engle et al. 1990).
Bollerslev (1990) proposed a computationally tractable multivariate GARCH by restricting conditional correlations to be constant through time:
where contains individual GARCH standard deviations and is a time-invariant positive definite correlation matrix. The conditional covariance between series and is — time-varying level driven by each series' own GARCH volatility, but constant relative coherence.
Estimation. The ML estimate of is the sample correlation of standardized residuals: where , positive definite by construction. This concentrates out of the likelihood, reducing matrix inversions to 1. The concentrated log-likelihood reduces to:
SUR connection. With constant correlations the model is an extension of Seemingly Unrelated Regressions (SUR) allowing GARCH heteroskedasticity — a framing that clarifies why has the SUR sample analogue form.
Empirical finding (Bollerslev 1990). Applied to weekly DM, FF, IL, SF, BP returns vs. USD: all conditional correlations significantly higher in the European Monetary System (EMS) period (1979–1985) than in the pre-EMS float (1973–1979). E.g. DM–FF: 0.607 → 0.932. LR test for parameter constancy: 388 ~ — highly significant. Even non-EMS currencies (BP, SF) showed higher correlations, suggesting a common USD factor beyond EMS policy coordination.
Limitation. Constant correlations is a strong assumption relaxed by DCC (Engle 2002) below.
Engle (2002a) introduced the Dynamic Conditional Correlation (DCC)-GARCH model as a computationally tractable multivariate GARCH with time-varying correlations. Write and decompose where contains individual GARCH standard deviations and is the conditional correlation matrix. evolves as:
where are standardized residuals and is the unconditional correlation. DCC extends the CCC model (Bollerslev 1990) where is fixed. DCC is estimated in two stages: (1) fit univariate GARCH models for the variances; (2) estimate from the standardized residuals.
Bauwens, Laurent, and Rombouts (2006) survey the MGARCH literature by organizing all models into three families and documenting their structural properties. See Bauwens-Laurent-Rombouts (2006).
VEC(1,1) (Bollerslev-Engle-Wooldridge 1988). The full vector model is where , , and , are matrices. Parameter count for the full model: — for this is 78, already unidentifiable in practice. Diagonal VEC restricts and to diagonal, giving parameters.
RiskMetrics. Scalar diagonal VEC with pre-set (daily) or (monthly): . Not estimated — calibrated by J.P. Morgan.
BEKK(1,1,K) (Engle-Kroner 1995). . PD by construction. For scalar : parameters. Covariance-stationary iff all eigenvalues of lie strictly inside the unit disk (BLR notation).
Factor GARCH. BEKK with rank-1 matrices , ; each factor is a univariate GARCH(1,1). Common factors drive all cross-asset covariances. For factor: parameters (same as diagonal VEC). All factors share persistence implied by the respective .
FF-GARCH (Vrontos-Dellaportas-Politis 2003). where is lower-triangular with ones on the diagonal and contains individual GARCH variances. Always PD. Generalizes the Cholesky-based Pourahmadi (1999) reparametrization.
O-GARCH (principal component GARCH). Extract eigenvectors from the sample correlation matrix of returns; fit univariate GARCH to each principal component. The conditional covariance is where is the eigenvector matrix and is the diagonal matrix of factor variances. Simple and scalable, but orthogonality of components is a restrictive assumption.
GO-GARCH (van der Weide 2002). Generalizes O-GARCH by allowing the mixing matrix to be where is the eigenvector matrix, is the diagonal eigenvalue matrix, and is an orthogonal matrix parametrized by Givens rotation angles. Two-step estimation: first estimate by PCA, then estimate by ML. Nests O-GARCH as .
CCC (Bollerslev 1990) — recap. with constant . The constant-correlation assumption is testable via the LM test of Tse (2000): under of constant . CCC is not invariant to linear transformations of — multiplying by a non-diagonal matrix changes the model structure.
DCC-T (Tse-Tsui 2002). where is the -period rolling correlation matrix of and . Requires for to be invertible. The baseline and two scalars are free.
DCC-E (Engle 2002) — extended version. Scalar DCC uses two parameters ; extended DCC uses matrices , so , preserving PD via Hadamard products. Scalar DCC has parameters including the entries of .
GDC (Kroner-Ng 1998). where contains GARCH standard deviations, is a dynamic correlation matrix, and adds element-wise cross-product terms. The ADC (Asymmetric Dynamic Conditional Covariance) extends GDC with leverage: the -element is augmented by where (component-wise). GDC nests DCC, CCC, DVEC, BEKK, and Factor GARCH; free parameters.
Copula-GARCH (Patton 2000, 2002; Jondeau-Rockinger 2001). Sklar's (1959) theorem: any multivariate distribution can be written as where is the copula and are marginals. Copula-GARCH fits univariate GARCH models for the margins and a (possibly time-varying) copula for the dependence structure. Time-varying copula parameters evolve as ARMA-type recursions in . Two-step MLE naturally decomposes marginal and dependence likelihoods.
Invariance. A model is invariant if multiplying by a non-singular matrix yields a model of the same type for . General VEC and BEKK are invariant; diagonal VEC and BEKK are NOT (cross-covariance dynamics change). CCC is NOT invariant because is generally not diagonal.
Marginalization. A bivariate VEC(1,1) model implies each univariate margin is at most a weak GARCH(3,3) — not necessarily a strong GARCH. Diagonal VEC(1,1) implies each margin is a strong GARCH(1,1). CCC, DCC, and GDC marginalization properties are unknown.
Temporal aggregation. Weak multivariate GARCH is closed under temporal aggregation (Hafner 2003): if the daily process is weak MGARCH, the weekly aggregate is also weak MGARCH with parameters that are explicit functions of the daily parameters. This generalizes the Drost-Nijman (1993) univariate result.
QML consistency (Jeantheau 1998). The Gaussian quasi-log-likelihood produces a consistent estimator whenever the conditional mean and covariance are correctly specified, regardless of the true innovation distribution. The analytical score (eq. 56 of BLR) avoids numerical differentiation.
QML asymptotic normality. Proved only for BEKK by Comte and Lieberman (2003). The asymptotic distribution theory for CCC, DCC, and GDC remains an open problem.
Variance targeting (Engle-Mezrich 1996). Replace the long-run intercept of (e.g., in BEKK or in DCC) by the sample unconditional covariance . Reduces the free-parameter count substantially without affecting consistency; makes the remaining parameters identifiable from the correlation dynamics alone.
DCC two-step. The full log-likelihood decomposes into where: Step 1 maximizes for each margin separately; step 2 maximizes for correlation parameters given the step-1 estimates. The two-step estimator is consistent but loses efficiency relative to joint maximization; one Newton-Raphson step at the two-step estimates achieves asymptotic efficiency.
Kim and Tsurumi (2000) apply Bollerslev's (1990) CCC-MGARCH to four daily Korean financial returns — KOSPI spot (KS), KOSPI futures (KF), Won/Dollar spot (WS), and 3-month NDF (NF) — over 10 October 1996 – 1 April 1998. Individual GARCH(1,1) estimates yield (IGARCH) for all four series. Full-sample ; pooled log-likelihood = −2100.77.
Structural break date is treated as an unknown parameter. The Bayesian log-posterior is approximated via Laplace: where and are split-sample log-likelihoods and Hessians at the respective MLE. Grid search over 5 July – 10 December 1997 identifies the MGARCH posterior break mode at 20 October 1997 — three weeks before the 8 November official devaluation (pre-break LL = −889.14; post-break LL = −1144.72). Individual-series posteriors find earlier breaks for derivative markets: KF = 11 August 1997; NF = 14 August 1997 — consistent with informed trading anticipating the crisis. The IGARCH finding is consistent with the Hillebrand (2004) / Lamoureux-Lastrapes (1990) result that structural breaks in GARCH parameters can induce spurious near-unit-root variance persistence. See Kim-Tsurumi (2000) and Structural Break Testing.
Given standardized residuals and :
For large N, the DCC model is the most flexible tractable option (N+2 parameters) but still requires N univariate GARCH estimations. An intermediate approach is Flex-GARCH (Ledoit, Santa-Clara, and Wolf 2003): estimate N(N+1)/2 bivariate GARCH(1,1) models with parameter constraints, and "paste" them together into the coefficient matrices A, B, C of a scalar GARCH. Specific transformations of bivariate parameters ensure psd covariance forecasts. Flex-GARCH becomes cumbersome above N ≈ 30; beyond that, the base-asset factor approach is preferred.
The base-asset mapping sets: where () are returns on liquid base assets (e.g., equity indices, FX rates, benchmark rates), is an loading matrix (from regression, pricing model, or expert matching), and captures idiosyncratic risk. A realized covariance matrix for the base assets is feasible because (where is the intraday sampling frequency, e.g., 48 for 30-minute data). Filtered Historical Simulation (see Value at Risk) then resamples factor and idiosyncratic innovations separately.
Standard forecast evaluation regresses ex-post squared daily returns on the GARCH one-day-ahead conditional variance forecast . The resulting R² values of 0.02–0.10 led many researchers to conclude GARCH forecasts are useless. Andersen and Bollerslev (1998) showed this conclusion is wrong.
Analytical population R². Under correctly specified GARCH(1,1) with excess kurtosis , the population R² of the Mincer-Zarnowitz regression () is:
For DM-\alpha=0.068\beta=0.898\kappa=2.8R^2=0.064$, matching the empirical 0.047. The "disappointing" R² is exactly what a correctly specified model predicts — not evidence of misspecification.
Measurement error decomposition. The actual forecast MSE using decomposes as:
For DM-r_t^2$ is almost entirely measuring noise, not forecast quality.
Realized volatility as the solution. By quadratic variation theory (Karatzas-Shreve 1988), as . The noise terms shrink by a factor . At 5-minute frequency (m=288), R² rises from ≈0.05 to ≈0.48 (theory 0.483) for DM-(\alpha+\beta)$.
Continuous-time GARCH limit (Nelson 1990). GARCH(1,1) converges in distribution to the SDE , establishing the structural link between discrete GARCH parameters and continuous-time integrated variance. Temporal aggregation uses the weak GARCH framework of Drost-Nijman (1993) and Drost-Werker (1996). See Andersen-Bollerslev (1998).
For assets with liquid intraday data, the realized variance on day is:
where are returns sampled at the -th interval on day (e.g., 30-minute intervals, ). As (continuous sampling), converges to the integrated variance , the true daily variance of the underlying continuous-time process. At 15–30 min intervals, measurement error from market microstructure noise is small enough to treat as approximately observed.
Key empirical regularities:
Realized covariance generalizes to the multivariate case:
is positive definite as long as (portfolio dimension smaller than the number of intraday observations per day). For large this fails; the base-asset factor approach (above) resolves this. Non-synchronous trading can bias cross-asset covariances; correction methods (analogous to Scholes-Williams 1977 for CAPM betas) add lead and lag cross-products.
The log-normal/normal mixture model connects realized and GARCH-based approaches: if and , then daily returns follow a scale mixture of normals. ABDL show this fits FX returns well in terms of VaR coverage rates. See Value at Risk for the conditional distribution modeling context.
The news impact curve (Engle and Ng 1993) quantifies how the conditional variance at time responds to the single most recent innovation , holding all prior information fixed at its unconditional level:
For standard GARCH(1,1): — symmetric in , opening upward as a parabola. For EGARCH (Nelson 1991): — the NIC is asymmetric: negative shocks () increase variance more than positive shocks of the same magnitude (leverage effect). For power-GARCH (Ding, Granger, and Engle 1993): — estimated on daily S&P 500 (1928–1991), rejecting both and .
Engle and Ng (1993) also provide misspecification tests for GARCH models based on whether the standardized squared residuals correlate with sign and magnitude of past innovations.
Hentschel (1995) nests eight widely-used GARCH specifications under a single variance equation governed by four parameters: a Box-Cox power applied to the conditional standard deviation, a shock exponent , a shift , and a rotation :
The function is a shifted and rotated absolute value. When the left side reduces to (the EGARCH parameterization).
Nesting table — standard models as parameter restrictions on :
| Model | ||||
|---|---|---|---|---|
| EGARCH | 0 | 1 | 0 | free, |
| TGARCH | 1 | 1 | 0 | free |
| AGARCH | 1 | 1 | free | 0 |
| GARCH | 2 | 2 | 0 | 0 |
| NA-GARCH | 2 | 2 | free | 0 |
| GJR-GARCH | 2 | 2 | 0 | free |
| NARCH | 0 | 0 | ||
| A-PARCH | 0 | free |
Two types of asymmetry. The shift moves the news impact curve minimum rightward, dominating the asymmetric response for small shocks. The rotation changes the slopes on either side of the minimum, dominating for large shocks. Prior GARCH models allow only one type at a time, conflating two empirically distinct dimensions of the news impact curve.
Empirical result (CRSP 1926–1990, daily excess returns). All symmetric and standard asymmetric models are rejected by LR statistics exceeding 200. The unconstrained optimum is , — intermediate between the Taylor-Schwert absolute-value GARCH and Bollerslev squared-residual GARCH, closer to the AGARCH family. The shift is significant in every maintained specification; the rotation is not. Asymmetry in U.S. equity volatility is primarily a small-shock phenomenon: the leverage effect conventionally attributed to large negative returns is driven instead by the shift of the news impact curve minimum away from zero.
The standard GARCH(1,1) imposes a single conditional distribution on returns. The MN-GARCH model (Haas, Mittnik, and Paolella 2004a; Bauwens and Rombouts 2005) instead writes the conditional density as a -component normal mixture where each component has its own GARCH(1,1) variance process:
Zero-mean constraint: For the unconditional mean to equal zero, .
Covariance stationarity: The mixture is weakly stationary even when individual components have (explosive), because the stationarity condition involves a weighted combination across all components. This permits the model to capture a heavy-tailed, high-variance regime (explosive component with small weight) alongside a persistent but stationary normal regime (dominant component).
Bayesian estimation (Bauwens-Rombouts 2005): A 4-block Gibbs sampler uses data augmentation via latent state indicators :
Marginal likelihood for selection: A Laplace approximation around the posterior mode gives .
Near-IGARCH as misspecification artifact: Applied to S&P 500 returns (Jan 1994–Jun 2005, ), is strongly preferred (marginal log-likelihood vs. for ). The two-component solution has: Component 1 weight , ; Component 2 weight , (explosive). The standard GARCH(1,1) estimate is a weighted compromise between the two components — the near-IGARCH finding is therefore a misspecification artifact from imposing a single distribution, not evidence of a unit root in variance. See Bauwens-Rombouts (2005).
Standard GARCH assumes a single volatility regime. The Markov Switching GARCH (MS-GARCH) model allows the level of the conditional variance to shift across a discrete set of regimes governed by a hidden Markov chain. Das and Yoo (2004) develop the first workable Bayesian Markov Chain Monte Carlo (MCMC) estimator for this model.
The path-dependence problem. In the model with Markov, the variance recursion implies depends on the entire past history of states . The exact likelihood therefore requires summing over state paths — infeasible for any realistic sample size. This is the fundamental obstacle to MLE for MS-GARCH; it also prevents the multi-move forward–backward state-sampling algorithm (used in simpler Markov switching models) from functioning correctly under GARCH dynamics.
Bayesian solution. Treating the full state vector as a latent parameter and conditioning on it at every MCMC step eliminates path-dependence: given , the variance recursion is deterministic and the likelihood factorizes across as a product of Gaussian densities. This is the core idea that makes the Bayesian approach computationally tractable where MLE is not.
MCMC algorithm (four blocks):
Relation to switching ARCH (Kaufmann–Frühwirth-Schnatter 2002). Earlier Bayesian work on regime-switching volatility was limited to switching ARCH (no lagged variance term ) precisely because multi-move state sampling fails under GARCH path-dependence. Das and Yoo demonstrate that single-move sampling sidesteps this limitation at the cost of slower Markov chain mixing for the state sequence.
Numerical results. MS-GARCH(1,1), : posterior means track true values well; MH acceptance rates 13–19%; the GARCH persistence parameter shows chain autocorrelation ; the regime indicator is imprecisely estimated (posterior mean 0.150 vs. true 0.050). See Das-Yoo (2004).
Let where is the conditional mean (e.g., the Vector Error Correction Model (VECM) component). The quasi log-likelihood is
maximized over by nonlinear numerical optimization. QML is consistent and asymptotically normal even when is non-Gaussian (Weiss 1986 univariate; Bollerslev and Wooldridge 1992; Tuncer 1996 and Bauwens-Vandeuren 1996 for BEKK), provided the conditional mean and variance are correctly specified.
Standard errors. Under non-normality, the usual inverse Hessian SE is invalid. Robust (sandwich) SE use where is the Hessian and is the outer-product-of-gradients matrix.
GARCH-based VaR (see Value at Risk) uses the estimated conditional variance to compute a time-varying quantile. Under conditional normality, the 99% one-day VaR for a portfolio of value is where . Nobel Committee (2003) S&P 500 example: the 99% VaR on a $1M portfolio ranged from $12,400 (September 1995, minimum volatility) to $61,500 (July 2002, maximum volatility) — illustrating the practical importance of time-varying volatility modeling.