A Markov-Switching VAR (MS-VAR) is a Vector Autoregression in which some parameters switch over time according to a discrete Markov chain. Introduced by Sims and Zha (2005a) as a principled alternative to continuous random-walk coefficient drift models, it restricts time variation to a finite number of states while preserving structural shock interpretability.
Standard time-varying parameter (TVP) VARs let all coefficients drift as a random walk:
This is over-parameterized: the dimension of time variation in equals the number of parameters, far exceeding the effective dimension of variation in the data. It also conflates changes in structural coefficients with changes in shock variances, making structural shock identification ambiguous.
Suppose a scalar parameter evolves as:
Discretize the support (two unconditional standard deviations) into states using thresholds . State corresponds to (with , ). The transition probability from state to state is:
where is the midpoint of interval and is the standard normal cumulative distribution function (CDF). This analytical form means the transition matrix is fully determined by .
For states, Zha (2005) uses thresholds (in units of ).
The MS-VAR allows different blocks of parameters to switch independently. For example:
This permits a direct test of whether the post-WWII moderation of US business cycles reflects a change in policy behavior or a lucky reduction in external shock variances.
Albert and Chib (1993) provide the first fully Bayesian estimator for a univariate Markov-switching AR model, and the approach generalizes directly to MS-VARs. The model is:
where , identifies state 1 as the high-mean state, and identifies state 1 as the high-variance state.
Data augmentation key. The full likelihood requires summing over all state paths — Hamilton's filter. Albert and Chib instead treat as missing data. The augmented posterior factors into tractable conditionals:
States — backward pass (single sweep, O(T)): draw from , then for : Each draw is a binary Bernoulli.
Regression block — truncated multivariate Normal (truncation: ), with precision and conditional mean from generalized least squares (GLS).
Error variance — Inverse-Gamma with shape and scale .
Variance scale — Inverse-Gamma truncated to .
AR coefficients — Normal, drawn conditional on the stationary region.
Transition probabilities — independent Beta draws from transition counts (number of transitions in ):
Empirical results (interest rates, 1953:3–1990:4). State 1 precisely identifies 1979:4–1982:3 (Volcker Fed tightening): posterior mean , so state-1 variance is state-0. Bayes Factors strongly favor switching over constant-parameter AR. Gross National Product (GNP) growth prefers AR(0) switching with modest variance change.
See James H. Albert, Siddhartha Chib, and Gibbs Sampler.
Albert and Chib (1993) and McCulloch-Tsay (1994) drew each state individually from its full conditional — requiring extra Gibbs blocks and producing high autocorrelation when is large. Chib (1996) shows the entire state sequence can be drawn jointly in one Gibbs block via a two-pass algorithm.
Forward filter. Maintain an matrix where row stores the (unnormalized) filtered probabilities :
where is the current transition matrix, is element-wise multiplication, and is the row vector of component densities across all states. Initialize row 1 from the chain's stationary distribution. The forward pass costs .
Backward sampler. Draw from the last row of . For :
(row of , element-wise multiplied by column of , then normalized). One draw per period; backward cost .
Full Gibbs cycle. Three blocks: (1) jointly via the above; (2) transition matrix — Dirichlet conjugate with updated counts of transitions in the sampled (eq. 10); (3) component parameters, conjugate conditional on . This is the discrete-state analogue of the Carter-Kohn (1994) simulation smoother for continuous Gaussian state-space models.
See Chib (1996) and Gibbs Sampler.
Before MS-VARs became standard, Engel and Hamilton (1990) applied the simplest Markov-switching model — a two-state switching mean — to quarterly dollar exchange rates. This is a univariate special case with no autoregressive lags:
with Markov transition matrix .
German mark estimates (1973 Q3–1983 Q4): /qtr (depreciation state), /qtr (appreciation state); , — regimes last and 7 quarters on average.
Testing. Under the null (random walk), the parameters , are unidentified — nuisance parameters disappear under the null. This makes classical likelihood-ratio asymptotics invalid. Engel-Hamilton use Wald tests on moment conditions and reject the random walk for Germany, France, and the UK at 5%.
Forecasting. The smoothed regime probabilities (Hamilton filter) correctly identify the dollar appreciation episode of the early 1980s and the preceding depreciation. In-sample MSE improvement over the random walk: 9–14% at 4-quarter horizon. Post-sample (1984–1987) gains are similar.
UIP failure. A key result: interest rate differentials have essentially no predictive content for which regime the exchange rate occupies. During the dollar appreciation episode (state 2), U.S. interest rates were often lower than foreign rates — the wrong sign for uncovered interest parity. See Uncovered Interest Parity.
Engel and Kim (1999) apply a Markov-switching state-space model to 106 years (1890–1995) of monthly U.S./U.K. real exchange rate data. The model decomposes the real exchange rate into two unobserved components with independent Markov chains governing their variances:
Estimated by Gibbs sampling using the Kim (1994) filter/smoother for the latent permanent component .
Key findings: (1) The permanent component's innovation variance dominates the transitory component's — the real exchange rate has a large permanent component, rejecting purchasing power parity (PPP). (2) The transitory variance states align with the international monetary regime: state 1 (gold standard / Bretton Woods) is nearly zero, state 3 (post-1973 float) is high. (3) Augmented Dickey-Fuller (ADF) unit root tests have actual size of 19% at the nominal 5% level under this data-generating process (DGP) — Markov-switching heteroskedasticity severely distorts classical unit root inference. See Purchasing Power Parity and Structural Break Testing.
Kim and Kim (1996) embed Markov-switching heteroscedasticity within a structural unobserved-components (UC) decomposition of stock prices. The model is:
where is a permanent AR(2) component () and is a transient "fad" AR(1) component (). The innovation variances switch independently:
Independent chains. The two Markov chains are independent of each other, so there are combined variance states at each period. An exact filter would require tracking Gaussian mixtures across consecutive periods. Kim's (1993, 1994) collapsing approximation retains only by collapsing the lagged mixture to a product of marginals, maintaining tractable maximum likelihood (ML) estimation.
Empirical results (monthly real S&P 500, 1952:1–1992:12). , (near-random-walk permanent component); (fad is nearly serially uncorrelated). The high-variance fad state has (mean duration months); the low-variance state has (mean duration months). Only two fad episodes are statistically significant: OPEC 1973–74 and the October 1987 crash. Both are negative fad realizations — unwarranted pessimism, not speculative bubbles.
Model comparison. Log-likelihoods: UC-MS > Turner-Startz-Nelson (1989) > GARCH(1,1) (generalized autoregressive conditional heteroskedasticity) . The GARCH near-unit-root persistence () is reinterpreted as a high-persistence low-variance Markov state rather than true I(GARCH) behavior.
See Chang-Jin Kim, Myung-Jig Kim, and Kalman Filter.
Kim and Nelson (1998) operationalize the synthesis proposed by Diebold and Rudebusch (1996): embed Hamilton's (1989) regime-switching mechanism directly inside the Stock-Watson (1989, 1991) linear dynamic factor model. The result is a model that captures both key features of the business cycle identified by Burns and Mitchell (1946) — comovement among indicators and nonlinearity at turning points — within a single framework.
Model structure. Each of monthly coincident indicators depends on current and lagged values of an unobserved common factor (the coincident index) plus an idiosyncratic AR component :
The common factor evolves as an AR process whose mean switches between regimes:
with , , so in recessions () and in booms (). The transition probabilities are:
Gibbs sampling: three latent vectors. The model is cast in state-space form. Three unobserved vectors are treated as missing data and sampled in turn:
Common factor path — Conditional on and parameters , the model is linear Gaussian, so the Carter-Kohn (1994) multimove Gibbs algorithm draws the entire path in one block via the Kalman filter forward pass and backward simulation smoother.
Regime path — Conditional on and , the unobserved index AR(1) in equation (9) is Gaussian, so Hamilton's (1989) basic filter provides ; the full path is then drawn backward via:
Parameters — Conditional on both latent paths, the model factorizes into independent Normal-IG posteriors for each equation's AR coefficients and noise variances, plus Beta conjugate posteriors for and .
The composite coincident index is recovered at each Gibbs iteration from the sampled factor path.
Empirical results. Estimated on U.S. data 1960:1–1995:1. Posterior regime probabilities from the multivariate model match the National Bureau of Economic Research (NBER) business cycle dates precisely; the univariate model's regime probabilities are substantially noisier. The new model-based coincident index correlates 0.9825 with the Department of Commerce (DOC) index in first differences, but shows sharper drops during recessions.
Duration dependence extension. The base model's fixed transition probabilities are replaced by a probit specification for the regime indicator, where the probability of a transition depends on how long the current phase has lasted (duration ). Coefficients (recession duration effect) and (boom duration effect) are given spike-and-slab priors: with probability (no duration dependence) or drawn from a truncated Normal (positive/negative duration dependence). Posterior probabilities of no duration dependence are read directly from the fraction of Gibbs draws in which the indicator . Strong evidence of positive duration dependence in recessions (posterior Pr[no dependence] ≈ 0.01–0.19 across prior configurations); weak and prior-sensitive evidence for booms.
See Chang-Jin Kim, Charles R. Nelson, Kim-Nelson (1998), and Variable Selection.
Kim and Nelson (1999a) provide the first formal econometric model of Friedman's (1964) "plucking" hypothesis: that output cannot exceed a trend ceiling and is occasionally plucked downward by transitory recessions that always rebound. The model decomposes output into a stochastic trend ceiling and a transitory cycle with an asymmetric Markov-switching shock:
where during recessions () and otherwise. The trend evolves as a random walk with stochastic drift: , . Setting and homoskedasticity recovers Clark (1987)'s linear UC; the asymmetry is uniquely placed in the transitory component rather than the trend (unlike Hamilton 1989 / Lam 1990).
Estimation uses the Kim (1994) approximate maximum likelihood estimator (MLE) filter: Kalman updates are run conditional on all pairs, and the resulting Gaussian mixture is collapsed back to 2 components via Harrison-Stevens (1976) approximations at each step. This contrasts with the exact Gibbs-sampling approach in Kim-Nelson (1998).
Empirical results (U.S. real gross domestic product (GDP) and unemployment, 1951:1–1995:3):
See Chang-Jin Kim, Charles R. Nelson, Kim-Nelson (1999a), and Plucking Model.
Kim and Nelson (1999b) ask whether Hamilton's (1989) Markov-switching business-cycle model itself underwent a structural break. They augment the model with a latent absorbing state () that allows a one-time permanent shift in the regime-dependent mean growth rates and/or the innovation variance :
Four models compared via Bayes factors (Chib 1995/1998 reduced-run marginal likelihood):
| Model | Break type | ln |
|---|---|---|
| I | None (Hamilton benchmark) | −267.63 |
| II | Shift parameters only | −247.01 |
| III | Variance only | −253.70 |
| IV | Both | −260.19 |
Key findings: Model II wins under all three prior specifications. The dominant source of the Great Moderation (break at 1984:Q1) is a narrowing boom–recession gap, not a pure decline in shock variance. Pre-break: recession mean , boom mean . Post-break: recession mean , boom mean . Within a linear model, regime-gap narrowing is observationally equivalent to variance decline; the MS framework is necessary to identify both separately.
See Great Moderation, Chang-Jin Kim, Charles R. Nelson, Kim-Nelson (1999b).
MS-VARs provide a middle ground between constant-parameter VARs (too rigid) and random-walk TVP-VARs (too flexible). The discrete-state structure enables: