A stochastic volatility (SV) model specifies the log-variance of a return series as a latent autoregressive process, rather than as a deterministic function of past observations. Unlike generalized autoregressive conditional heteroskedasticity (GARCH), the conditional variance is not observable given past returns — it is a hidden state that must be inferred alongside the model parameters.
The first-order autoregressive SV model (Taylor 1982) specifies:
with for all . Parameters: (unconditional log-variance level), (persistence), (volatility of volatility).
Unconditional properties. The log-variance is Gaussian autoregressive (AR)(1) with stationary distribution when .
| Feature | GARCH(1,1) | SV (AR(1)) |
|---|---|---|
| observable? | Yes (deterministic given past ) | No (latent state) |
| Likelihood | Closed form | Requires integration over |
| Estimation | Numerical ML (fast) | MCMC or particle filter (slow) |
| Flexibility | Limited tail fit | More flexible via |
| Asymmetry | Via Exponential GARCH (EGARCH)/Glosten-Jagannathan-Runkle GARCH (GJR-GARCH) add-ons | Via leverage effect: |
The continuous-time counterpart specifies:
where . Volatility (not ) follows an arithmetic Ornstein-Uhlenbeck (AR1) process with mean-reversion speed , long-run mean , and vol-of-vol .
Closed-form distribution. The stock price distribution has an exact closed-form (Stein-Stein eq. 9–10) involving sinh/cosh terms and a single numerical integration. Key result from Appendix A–B: , i.e., a mixture of lognormals, where is lognormal with RMS volatility and is the mixing distribution over realized RMS vol. This representation holds for any diffusion SV model, not just OU.
Volatility smile. Stochastic volatility raises all option prices above Black-Scholes. Implied volatility is U-shaped in strike: lowest at-the-money (where Black-Scholes is near-linear in , so only the mean of matters) and rising away-from-the-money (where convexity of Black-Scholes in adds a dispersion-of-mixing-distribution premium).
Fat tails. as (power law, fatter than lognormal):
Matches Bookstaber-McDonald (1987) empirical finding: fat tails at 1–5 day returns, near-lognormal at 250 days — consistent with stationary (mean-reverting) volatility.
Limitation. : no leverage effect. Correlated Wiener processes were analytically intractable; Heston (1993) solved the correlated case via characteristic functions.
Ball and Roma unify the Hull-White (H-W) and Stein-Stein (S-S) approaches through the moment-generating function (MGF) of average variance :
Cox-Ingersoll-Ross (CIR)/Heston square-root model. For , the MGF is available directly from CIR (1985) bond pricing: , where , , , . Simpler than S-S's sinh/cosh formula; square-root process is always non-negative with gamma limiting distribution.
Volatility smile derivation. Implied variance under the (incorrect) constant-volatility assumption is , where: is quadratic in log-moneyness , minimized at at-the-money (ATM) () and rising for out-of-the-money (OTM)/in-the-money (ITM) — formal derivation of the U-shaped smile.
S-S error corrected. S-S compared their SV price to BS at variance , but the correct benchmark is . The excess arises because S-S model (absolute value of OU) rather than a reflected OU process: has long-run mean above by Jensen's inequality. With the correct benchmark, SV exerts a downward bias ATM (BS is concave in variance there) and upward bias away-from-the-money.
Zero-correlation constraint. The average-variance framework — both power series and Fourier — requires . Nonzero correlation breaks the lognormality of conditional on .
Amin and Ng (1993) provide an equilibrium foundation for the expected-Black-Scholes formula, working in a pure discrete-time constant proportional risk aversion (CPRA) framework without continuous-time mathematics.
Systematic variance decomposition (eq. 3): The asset's conditional variance decomposes as , where is the conditional variance of log consumption growth (systematic component) and is idiosyncratic; measures the stock's sensitivity to aggregate consumption shocks. This structure directly anticipates factor GARCH models.
Proposition 1 (Preference-dependent): Under CPRA utility and bivariate conditional lognormality of stock-consumption returns, the call price is a preference-weighted expectation of payoffs. The weights depend on the covariance of the stochastic discount factor with the cumulative asset return.
Proposition 2 (Preference-free): When the variance process is predictable (measurable in the investor's time-0 information set), the call price collapses to: No risk-aversion parameter appears — this is the same expected-Black-Scholes formula derived by Hull-White (1987) and Ball-Roma (1994), now grounded in discrete-time general equilibrium.
Endogenous riskless rate (eq. 11): . The interest rate declines with consumption variance — a flight-to-quality channel — distinguishing Amin-Ng from no-arbitrage SV models that treat the interest rate as exogenous.
Systematic jump formula (eq. 27): When consumption growth contains a compound Poisson jump component, the call price becomes a Merton (1976)-type series , with jump risk priced through the CPRA kernel rather than under risk neutrality.
Bakshi, Cao, and Chen (1997) embed continuous-time SV within a fully general nested model — SVSI-J — and evaluate four special cases on 38,749 S&P 500 call options (June 1988–May 1991) along three dimensions: internal consistency, out-of-sample pricing, and hedging.
General dynamics. The SVSI-J model has three state variables: stock price , instantaneous variance , and short rate :
Jumps arrive as a Poisson process with intensity λ; log jump sizes are . Setting gives SVSI; gives SVJ; both gives SV; all three zero gives BS.
Pricing via characteristic function inversion. The European call price takes the Black-Scholes-Merton (BSM) form , where and are recovered by inverting the characteristic function of . The characteristic function factors into closed-form exponential-affine contributions from , , and the jump process.
Implied parameter estimation. On each trading day, structural parameters are estimated by minimizing the sum of squared pricing errors across all available options. This daily cross-section sum of squared errors (SSE) minimization yields a time series of implied .
Three-yardstick results (Table III, All Options):
| Model | SSE | ||||
|---|---|---|---|---|---|
| BS | — | — | — | — | 69.60 |
| SV | 1.15 | 0.39 | — | 10.63 | |
| SVSI | 0.98 | 0.42 | — | 10.68 | |
| SVJ | 2.03 | 0.38 | 0.59/yr | 6.46 |
Internal consistency: All models are significantly misspecified. Implied (SV) is 2–3 larger than time-series ML estimate (); implied is its ML counterpart. Among SV-class models, SVJ is least misspecified.
Out-of-sample pricing: SVJ SVSI SV BS. SV reduces pricing errors 25–70% over BS; jumps add most value for short-term OTM options.
Hedging (single-instrument): SV SVJ SVSI BS. With /year (one jump per 1.7 years), jumps are essentially never realized at 1–5 day hedging intervals — so jump risk cannot be exploited in the delta-hedge.
BSDV control (delta-plus-vega-neutral BS): A BS hedge augmented with a vega-neutral option position (BSDV) nearly closes the hedging gap for non-ITM calls, confirming that most of the SV models' hedging advantage is mechanical (using a second option instrument) rather than structural.
See Jump-Diffusion Model for the full BCC nested model structure and empirical details.
The foundational MCMC approach (JPR 1994) uses data augmentation: augment with the full latent variance path and sample each one at a time from its full conditional, then draw parameters jointly.
Full conditional for :
where and are the mean and variance of conditional on its AR(1) neighbors. The return piece is inverse-gamma shaped; the state piece is lognormal. An IG blanketing density is formed by moment-matching the lognormal to an IG, and an independence Metropolis step accepts/rejects draws from this IG proposal. Acceptance rates 70–80%.
Parameter block: Given , the log-linear state equation is a standard Gaussian AR(1), so are drawn jointly from a Normal-inverse-gamma (IG) conjugate posterior.
Simulation results: Bayesian root mean squared error (RMSE) is 3.7–4.1 smaller than method of moments (MM) and comparable or better than quasi-maximum likelihood (QML). MM suffers corner solutions (>50% of samples for high-persistence, low vol-of-vol regimes). Bayesian smoother beats the approximate Kalman smoother even when given the true parameters (relative RMSE 1.11 median).
Precursor (Geweke 1993). The latent-variance data augmentation device in JPR (1994) has a static predecessor in Geweke (1993): augmenting a linear regression with independent and identically distributed (i.i.d.) Student-t errors with latent scale weights (where ) renders all full conditionals conjugate and enables a direct Gibbs sampler. JPR's dynamic path is the time-series counterpart of Geweke's static — the device that makes Bayesian SV inference tractable originates in Geweke's cross-sectional regression context.
KSC transform the observation equation by log-squaring:
The noise has mean , variance , and severe left skewness — it is not Gaussian, so a Kalman filter applied directly would be misspecified (QML). KSC approximate it by a fixed mixture of Gaussians:
with weights , means (centred), and variances chosen to match the first four moments plus the mode of (Table 4 of KSC).
MIX Gibbs sampler. The three-block cycle is:
Indicator draw. For each , draw the discrete indicator from its multinomial full conditional:
Path draw. Conditional on and parameters, the system is a linear Gaussian state space: Apply the Carter-Kohn (1994) forward filter / backward simulation smoother to draw the entire path jointly in time.
Parameter draw. Given , the state equation is a Gaussian AR(1). Draw jointly Normal from the conjugate regression of on , and from the conjugate Inverse-Gamma.
Importance reweighting. The MIX sampler targets a pseudo-posterior (under the mixture approximation). Exact inference is recovered by importance reweighting with: The effective sample size ratio exceeds 0.99 in simulations — the 7-component approximation is highly accurate.
Efficiency comparison. The inefficiency factor INEF measures sampling inefficiency relative to i.i.d. draws. Across the three samplers (SM: single-move; MIX: mixture multi-move; INT: integration particle filter):
| Sampler | INEF () | INEF () | Cost per draw |
|---|---|---|---|
| SM (single-move) | 20–200 | 20–150 | Low |
| MIX (mixture) | 1–3 | 1–3 | Low |
| INT (particle filter) | 1–2 | 1–2 | High |
MIX achieves near-i.i.d. efficiency at the same computational cost as SM.
Bayes factors (Chib 1995 identity). Evaluated via at the posterior mode , with the log-likelihood computed by the Kalman filter under the mixture approximation. For daily S&P 500 (1981–1987): ; SV also decisively beats -GARCH.
Mahieu and Schotman (1998) provide the most systematic comparison of SV estimation techniques in an exchange-rate application. The log-squared-return transformation yields a linear state-space model, but the measurement error has a log chi-square(1) distribution — not Gaussian — with , , and severe left skewness. Applying the Kalman filter directly (QML) is therefore misspecified.
Variance decomposition. The total variance of decomposes as:
The log chi-square noise accounts for 60–80% of for weekly foreign exchange (FX) returns, leaving only 20–40% attributable to genuine log-volatility variation. This makes the latent path very difficult to estimate precisely.
Mixture-of-normals approximation. Kim, Shephard, and Chib (1998) approximate the log chi-square distribution by a fixed mixture of normals with pre-specified weights , means , and variances (chosen to match moments). Conditional on discrete mixture indicators :
which is Gaussian, allowing exact Kalman filtering and smoothing. Mahieu-Schotman extend to a flexible mixture: all are estimated from the data jointly with via Simulation EM (SIEM), providing a better fit for fat-tailed exchange rate innovations.
Gibbs sampler for the flexible mixture (SIEM2/Bayesian). The sampler cycles through:
Empirical results (6 FX pairs, weekly 1975–1991):
| Method | Notes | ||
|---|---|---|---|
| QML | Biased downward on both | ||
| SIEM1 (fixed mixture) | Better | ||
| SIEM2 (flexible mixture) | Best ML-type | ||
| Bayesian | Similar to SIEM2 |
QML severely underestimates both and . Standard errors of the smoothed volatility estimate from QML are too small; the Bayesian/SIEM standard errors are 2–3 larger, reflecting the true posterior uncertainty.
Option pricing implication. The option value under SV is , computed by averaging Black-Scholes values over simulated future volatility paths. Because QML underestimates , it also underestimates the distribution of future volatility — producing option price confidence intervals that are too narrow.
JPR (2004) extends the basic single-move MCMC framework to two richer SV specifications, then uses Bayes Factors to compare all four models.
Fat-tailed model (FC). Replace the Gaussian return innovation with via scale mixing:
Augmenting with latent precision weights restores Gaussianity conditional on : . The full conditional for is then log-concave, enabling adaptive rejection sampling (Wild and Gilks 1993) in place of the inverse-gamma M-H step of JPR (1994). The degrees-of-freedom parameter is drawn via a Metropolis step.
Correlated errors (ASV2). Introduce same-period correlation by rewriting:
A fall in price (negative ) directly raises the current-period volatility shock when — the leverage mechanism. Drawing uses a linearization of around its current value to construct a Gaussian M-H proposal with high acceptance rates.
Bayes Factors. All four models — basic SVOL, FC, ASV2, and Full (FC + ASV2) — are compared via importance-sampling BF computed from the MCMC posterior draws. For S&P 500 and CRSP equity indices, BF strongly favors both extensions; the correlated model wins more decisively than fat tails alone. For FX (CAD/USD), BF — no evidence for either extension.
Key empirical findings. For equity: to (strong leverage); –30 for daily data (moderate fat tails), smaller at weekly frequency. Fat-tailed model produces smoother volatility paths: extreme returns are attributed to heavy tails rather than volatility spikes. Fat-tailed SV reduces option pricing RMSE by vs. basic SVOL; leverage creates put-call skew.
The leverage effect — negative correlation between return shocks and future volatility shocks — requires care in discrete-time SV. Two specifications exist:
| ASV1 (Harvey-Shephard 1996) | ASV2 (Jacquier et al. 2004) | |
|---|---|---|
| Correlation | ||
| Timing | contemporaneous return → next-period vol shock | same-period return and vol shock |
| Efficient market hypothesis (EMH) | Martingale difference: ✓ | — predictable returns ✗ |
ASV1 orthogonalised form (nonlinear state space, eq. 2.3): setting ,
is linear in : if , bad news raises future expected log-variance — unambiguous. The analogous ASV2 re-parameterisation is not analytically tractable.
Estimation: Multi-move Kim-Shephard-Chib (1998) algorithms use a log-squared transformation that destroys leverage information; BUGS (single-move MH) is applied directly on the nonlinear state-space form. Marginal likelihood via Chib (1995) identity + Kitagawa (1996) particle filter for log-likelihood evaluation.
Empirical verdict: Bayes factor ASV1/ASV2 (S&P500 1980–1987) — decisive. (ASV1) to ; ASV2 understates by . MCMC outperforms QML: relative efficiency 56–71%.
Omori, Chib, Shephard, and Nakajima (2007) extend the KSC (1998) multi-move MCMC sampler to the ASV model with correlated return and volatility shocks (). The obstacle was that KSC's log-squaring transformation discards sign information and renders the joint density of return and volatility innovations non-separable. The resolution is to approximate the joint density of — where — by a 10-component mixture of bivariate normals, exploiting the return sign to partially recover leverage information.
ASV model. Returns and log-volatility follow: where , . The parameter captures the leverage effect; the difficulty is that standard KSC applies only when .
Bivariate mixture approximation. Within each mixture component , the conditional mean of given involves , which is approximated linearly: The mean-square-optimal closed-form coefficients are: The resulting bivariate mixture approximation is: Crucially, the quality of this approximation is independent of — the bivariate mixture remains accurate across all leverage values, including .
Two-step MCMC. Conditional on mixture indicators , the model is a linear Gaussian state space, enabling the de Jong-Shephard (1995) simulation smoother. The sampler cycles through:
Importance reweighting. Exact inference is recovered by reweighting MCMC draws by , where is the true joint density and is the mixture approximation. For TOPIX (1998–2002, ), the log-weight standard deviation is , confirming negligible approximation error.
Empirical results (TOPIX 1998–2002). Posterior means: , , [95% CI: ]. Inefficiency factors: , , , — all small, indicating near-i.i.d. draws. Log Bayes factor (Chib 1995 identity) for ASV vs. SV : leverage is decisively present. Model ranking: ASV-t ASV superposition SV-t SV.
Jensen and Maheu (2008) attack a different axis of misspecification: not leverage, but the parametric form of the return innovation distribution. Standard SV models assume ; they replace this with a nonparametric Dirichlet Process Mixture (DPM) while keeping the parametric AR(1) log-volatility process intact.
SV-DPM model. Returns where and with Normal-IG base distribution . After stick-breaking, : an infinite Gaussian mixture with volatility-scaled variances. Log-volatility follows the standard AR(1): , , independent of . No leverage effect.
Key finding: misspecification inflates . A Gaussian SV model confounds excess kurtosis and skewness in returns with volatility variation — it attributes non-Gaussian tail behavior to spikes in the latent variance rather than to the innovation distribution. The DPM corrects this.
4-block MCMC. (1) Truncated-Normal/IG for ; (2) Fleming-Kirby (2003) random-length block sampler for — groups consecutive observations into randomly-sized blocks and runs a Kalman smoother within each block, more efficient than single-move for persistent volatility; (3) Chinese Restaurant Process updates for cluster allocations and cluster parameters (West-Müller-Escobar 1994 / MacEachern-Müller 1998); (4) Gamma posterior update for .
Place in literature. Jensen-Maheu (2008) complements Omori-Chib-Shephard-Nakajima (2007): OCSN (2007) handles leverage () with a fixed parametric return distribution; Jensen-Maheu handles distributional misspecification with no leverage. A model combining both extensions — bivariate mixture for leverage + DPM for the innovation shape — remains an open problem. See Dirichlet Process Mixture for the statistical framework.
Chib et al. (2002) extend the basic SV0 model in two directions and develop efficient MCMC algorithms for each.
Model SVt. The observation equation replaces Gaussian errors with Student-t via scale mixing:
The level effect allows variance to scale with the mean level; adds covariates to the log-variance equation; includes lagged returns for the mean equation.
Model SVJt. Adds a Bernoulli jump term where and .
The seven-component Gaussian mixture. After augmenting with , the squared and logged observation where . Kim-Shephard-Chib (1998) approximate the log distribution by a 7-component Gaussian mixture with fixed weights , means , and variances (Table 1). Conditional on discrete indicators , the model is a linear Gaussian state space, enabling the de Jong-Shephard (1995) simulation smoother to draw the entire latent path in one block.
Algorithm 1 (SVt — 4-block Gibbs):
Algorithm 2 (SVJt — 6-block Gibbs): adds blocks for jump indicators , jump sizes , and jump intensity .
Marginal likelihood. The Chib (1995) identity requires evaluating (done via the auxiliary particle filter, Pitt-Shephard 1999, with particles) and the posterior ordinate (done via Chib-Jeliazkov (2001) for the M-H blocks, which uses the proposal and acceptance probability directly).
Inefficiency factor ; its inverse equals Geweke's numerical efficiency. Values of 1–5 obtained with tailored M-H proposals and blocking; much higher without careful implementation.
S&P 500 results (8,849 obs, July 1962 – August 1997):
| Bayes factor () | vs. SV0 |
|---|---|
| SVt | 10.75 — decisive |
| SVJ | 5.18 — decisive |
| SVJt | 9.76 — decisive |
| SVt vs. SVJ | +5.57 in favor of SVt |
| SVJt vs. SVt | −0.99 (essentially tied) |
Key parameter estimates: (strong persistence); for SVt; jump intensity /day in SVJ (1 per days) falls to 0.00195 in SVJt (1 per days). The decrease in jump frequency when moving from SVJ to SVJt shows that SVt absorbs most "jump-like" observations as ordinary tail realizations. SV0 systematically overestimates volatility because large returns are attributed entirely to variance spikes rather than being filtered as tails or jumps.
Prior and simulation robustness: G=5,000 and G=50,000 sweeps yield identical posteriors; rankings robust to 7 alternative prior specifications including priors strongly favoring Gaussian errors.
Chib, Nardari, and Shephard (2006) extend the univariate SV framework to high-dimensional multivariate settings via a factor structure: each of asset returns loads on latent factors, with both asset-specific and factor-specific log-volatilities following independent AR(1) processes.
MSVJt model. The observation equation is:
where are latent factors with diagonal variance ; is a Bernoulli jump indicator; and with (Student-t scale mixing). Each log-volatility follows:
Identification. is lower-triangular with positive diagonal elements, leaving free parameters. At , : 688 total parameters — a scale previously infeasible for multivariate SV estimation.
Reduced blocking scheme. Naive alternation of then produces inefficiency factors exceeding 1,000, because and are nearly collinear. The key innovation: sample marginalized over via a Newton-Raphson-tuned multivariate- Metropolis-Hastings proposal (eq. 7), then draw analytically (). After this step the model separates into conditionally independent univariate SV state spaces, each sampled via the KSC (1998) 7-component mixture + de Jong-Shephard (1995) simulation smoother. Resulting inefficiency factors: 1–30 (vs. >1,000 without the refinement).
Marginal likelihood. An auxiliary particle filter (Pitt-Shephard 1999, particles) estimates at the posterior mode . The simulation-consistent likelihood feeds into the Chib (1995) identity; posterior ordinates for M-H blocks use the Chib-Jeliazkov (2001) framework. Bayes factors select among MSV / MSVt / MSVJ / MSVJt variants.
Application (10 international equity indices, weekly 1973–2003, ). Model selection: Bayes factors favor MSVt with 3 factors (fat tails, no jumps). Forecast comparison: MSV mean absolute deviation (MAD) at 1-week horizon = 3.45 vs. BEKK = 3.63, Dynamic Conditional Correlation (DCC) = 3.72 — MSV best or tied across all multivariate GARCH (MGARCH) alternatives. value at risk (VaR): MSV achieves best conditional coverage across 4 portfolios (World, Europe, Pacific Rim, US) at both 1% and 5% levels.
Jacquier, Johannes, and Polson (2007) show how to compute the frequentist MLE of SV parameters via MCMC, without gradient methods. independent copies of the volatility path are drawn at each Gibbs step. The augmented joint density has a -marginal concentrating at the MLE as .
For the basic SV model (), stacking volatility copies in a single regression gives:
where X is the design matrix of stacked copies and S the residual sum. The volatility copies are drawn via M-H (as in JPR 1994). The convergence theorem guarantees , providing standard errors as . Normality of scaled draws (Jarque-Bera) is the diagnostic.
, suffices for the basic SV model ( min on 2006-era hardware); smoothed volatility estimates are essentially identical to the Bayesian smoother, confirming that the MCMC smoothing efficiency is preserved at the MLE. See MCMC Maximum Likelihood for the general framework.
Whereas the models above treat volatility as a univariate latent process, Brandt and Kang (2004) model both the conditional mean and conditional volatility of stock returns as jointly latent states in a bivariate VAR(1):
The key parameter is — the contemporaneous correlation between mean and volatility innovations. Estimated by simulated maximum likelihood with a VAR importance-sampling correction.
Main finding (Center for Research in Security Prices (CRSP) monthly returns, 1946–1998): () — strongly negative. A positive shock to expected returns is simultaneously accompanied by a drop in volatility. The lag risk-return tradeoff coefficient (lagged volatility predicting mean) is small and insignificant ().
This distinguishes two correlations with opposite signs: the contemporaneous innovation correlation (negative, strong) vs. the unconditional long-run correlation (positive — both mean and volatility peak at recession troughs). See Risk-Return Tradeoff.
Bollerslev and Zhou (2006) show that three apparently contradictory empirical regularities all follow analytically from the Heston (1993) model with just two parameters: (instantaneous leverage) and (negative volatility risk premium).
Model: Heston (1993) objective dynamics and risk-neutral dynamics:
The gap between and is central to all three results.
Proposition 1 — Volatility feedback: Population slope in realized-vol regression is ; can be negative when . Implied-vol slope is always, where due to . S&P500 (Jan 1990–Feb 2002): (realized), (implied).
Proposition 2 — Leverage asymmetry: : implied vol always exhibits stronger asymmetric response to lagged returns than realized vol. Long regressions confirm: realized leverage is statistically insignificant while implied leverage is highly significant ( to ).
Proposition 3 — Implied-vol forecasting bias: Population slope in is . Slower risk-neutral mean-reversion () makes implied vol systematically overshoot future realized vol. Not a market inefficiency — it is structural. Empirical (std), (variance).
Monte Carlo finding: 5-minute realized vol introduces negligible measurement error for monthly regressions. Daily-squared-return-based RV introduces large, persistent biases that do not shrink with sample size.
Standard BVAR estimation fixes the error covariance over the full sample. Uhlig (1997) extends the conjugate BVAR framework to allow a time-varying precision matrix (the inverse covariance), maintaining analytical tractability via a multiplicative random walk driven by multivariate Beta variates.
Law of motion for H_t:
where is drawn from the matrix Beta distribution (Uhlig 1994b: ratio of two independent Wishart variates). The scalar parameter controls persistence — large means H_t is nearly constant, small allows large jumps. In the univariate case, with : a scalar multiplicative random walk for the precision.
Conjugate updating. The prior on is matrix-normal inverse-Wishart (MNIW). The key result: despite time-varying , the posterior at each step remains MNIW, updated by closed-form recursions. At each period:
No MCMC is needed for the per-step update; the recursion is a nonlinear Kalman filter.
Importance sampling for the marginal posterior of B. The full joint posterior is not closed-form. Uhlig evaluates it by: (a) drawing paths from their conditional posteriors, (b) weighting by importance ratios. Figure 2 of the paper shows log-weight vs. log-posterior scatter near-diagonal for 4,000 draws on a 4-variable system — confirming the sampler does not collapse.
Contrast with other approaches:
Empirical finding (Uhlig 1996 draft, 4-variable U.S. macro VAR): "Modest evidence of a genuine change" in volatility. Results are robust to discount factor choices .
Eraker (2001) solves the continuous-time estimation problem by data augmentation: latent data points between each discrete observation set , making the Euler Gaussian transition density arbitrarily accurate. The augmented posterior is sampled by a two-block Gibbs cycle.
CEV one-factor model. . Given augmented data , the drift parameters have a Normal posterior via linear regression; has an IG posterior; (log-concave) uses a second-order Taylor–expanded Metropolis step. Key empirical result on 2,288 weekly US T-bill yields (1954–1997): vs. GMM (Chan et al. 1992) — a large discrepancy attributable to discretization bias in non-Bayesian estimates. CEV fails diagnostically: heavy-tailed residuals (Q-Q) and ARCH-like autocorrelations in squared residuals that the model cannot explain.
CEV+SV two-factor model. (log-volatility) is latent, augmented alongside the missing path. Key results: first-order autocorrelation (AC) (consistent with GARCH/SV literature); (vs. Andersen-Lund 1997a EMM: 0.176 — Eraker attributes the difference to higher short-run volatility variation); entirely negative in SV model (strong mean reversion), whereas CEV finds only posterior mass below 0.
Reparameterization fix for bias (Appendix D). High serial dependence in the volatility chain's running mean biases downward. Fix: at each step subtract and adjust , . This recenters the log-volatility chain without affecting the stationary distribution.
Discretization-convergence trade-off. As Gibbs conditionals collapse toward Dirac measures, slowing convergence. At large, discretization bias dominates. Practical recommendation: start with m=1 and increase until posterior means stabilize (convergence at – for CEV parameters). See Diffusion Process for the Brownian bridge AR-MH proposal and theoretical justification.
Cappuccio, Lubian, and Raggi (2006) extend the standard SV model by replacing Gaussian return shocks with a Skew-GED distribution built via the Azzalini (1985) device applied to the generalized error distribution (GED) base. Two parameters: (skewness; = left-skew) and (tail thickness; is Gaussian). The model nests Normal (), Skew-Normal (), and GED (). No leverage effect: return and volatility shocks are independent.
MCMC: Log-volatilities via Delayed-Rejection MH (Tierney-Mira 1999); and via Adaptive-Rejection Metropolis Sampling. Specification testing via Savage-Dickey density ratios: , computed by kernel smoothing MCMC output at the restriction.
Empirical results (DJ30, S&P500, Nasdaq; daily and weekly from Datastream): – (near-IGARCH); – (heavy tails throughout). Daily: for DJ30/S&P500 (GED sufficient), for Nasdaq. Weekly: decisive left-skew across all indexes ( to ). Gaussianity rejected universally. Frequency dependence: asymmetry is a weekly phenomenon, stronger than at daily frequency — consistent with the broader stylized-fact literature.