Structural Break Testing

structural-breaksunit-rootsBai-PerronHansenPerronZivot-Andrewsmultiple-breaksbayesiangibbs-samplermarkov-switchinggreat-moderationtrend-stationary

Definition

Structural break testing refers to procedures for detecting and estimating parameter instability in econometric models — changes in means, trends, slopes, or variances at one or more unknown dates within the sample. The central challenge is that ignoring genuine breaks biases inference toward spurious persistence (unit roots, long memory, integrated GARCH (IGARCH)), while falsely imposing breaks inflates the risk of spurious rejections.

Key Ideas

How It Works

Classical Tests for Coefficient Constancy (Chow 1984 §10)

The following tests all address H0:βt=βH_0: \beta_t = \beta (constant) vs. a changing-coefficient alternative. They predate the Zivot-Andrews / Bai-Perron literature and remain standard diagnostics.

Chow (1960) F-test — requires a known break date TBT_B. Estimate separate regressions on [1,TB][1, T_B] and [TB+1,T][T_B+1, T]; the ratio of the improvement in sum of squared residuals (SSR) to the restricted SSR, scaled by degrees of freedom, is FF-distributed under H0H_0. Critical limitation: the test is not valid when TBT_B is chosen by inspecting the data.

Cumulative Sum (CUSUM) and CUSUM² (Brown, Durbin, Evans 1975) — based on recursive residuals wt=(ytxtβ^t1)/(1+xt(Xt1Xt1)1xt)1/2w_t = (y_t - x_t'\hat{\beta}_{t-1})/(1 + x_t'(X_{t-1}'X_{t-1})^{-1}x_t)^{1/2}. The CUSUM statistic Wr=t=k+1rwt/σ^W_r = \sum_{t=k+1}^r w_t / \hat{\sigma} is compared against straight-line boundaries ±aTk(1+2(rk)/(Tk))\pm a\sqrt{T-k}(1 + 2(r-k)/(T-k)). CUSUM² uses cumulated squared recursive residuals and detects variance change. Both have asymptotically known distributions under H0H_0.

Quandt (1960) maximum likelihood-ratio test (LRT) — maximizes the Chow F-statistic over all interior break dates λ=TB/T[λ0,1λ0]\lambda = T_B/T \in [\lambda_0, 1-\lambda_0]. Precursor to the Zivot-Andrews statistic; critical values are not standard FF and must be tabulated for each (λ0,k)(\lambda_0, k).

Nyblom (1983) LMP test — locally most powerful test of H0:V=0H_0: V = 0 against the random-walk time-varying-parameter (TVP) alternative βt=βt1+ηt\beta_t = \beta_{t-1} + \eta_t. The test statistic is the score evaluated at the null: L=T2t=1TSt(XX)1St/σ^2,St=s=1txs(ysxsβ^)L = T^{-2}\sum_{t=1}^T S_t' (X'X)^{-1} S_t / \hat{\sigma}^2, \qquad S_t = \sum_{s=1}^t x_s(y_s - x_s'\hat{\beta}) Asymptotically distributed as a functional of Brownian bridges; critical values tabulated.

Pagan (1980) score test — also tests H0:V=0H_0: V = 0; similar score-based approach. Detects any form of coefficient variation, not just the random-walk alternative.

Perron (1989) — Single Known Break

Perron (1989) demonstrates that standard DF tests have zero asymptotic power against trend-stationary (TS) alternatives with a single known structural break. Three models are considered, distinguished by the form of the break in the deterministic component; each has a distinct null, alternative, and limiting distribution.

Break date and fraction. Let TBT_B be the break date, λ=TB/T(0,1)\lambda = T_B/T \in (0,1) the break fraction, and define: DUt=1(t>TB)DU_t = \mathbf{1}(t > T_B) (post-break level dummy), DTt=tTBDT^*_t = t - T_B for t>TBt > T_B, else 0 (post-break trend), D(TB)t=1(t=TB+1)D(T_B)_t = \mathbf{1}(t = T_B + 1) (pulse dummy at the break).

Model A — Crash (intercept shift).

Model B — Changing growth (slope shift).

Model C — Both. Both an intercept and slope shift are allowed; the test regression includes DUtDU_t, tt, and DTtDT^*_t.

Theorem 1 (zero power under the alternative). Under the trend-stationary alternative, the probability limit of α^\hat\alpha from the test regression satisfies: for Model A (crash), αˉ[(μ1μ2)2A]/[(μ1μ2)2A+σe2]<1\bar\alpha \to [(\mu_1-\mu_2)^2 A] / [(\mu_1-\mu_2)^2 A + \sigma^2_e] < 1 — power is positive but αˉ↛0\bar\alpha \not\to 0; for Model B (changing growth), αˉ1\bar\alpha \to 1 exactly — the test is asymptotically inconsistent against slope breaks.

Theorem 2 (limiting distribution under the null). Under the unit root null (with σ2=σe2+σu2\sigma^2 = \sigma^2_e + \sigma^2_u, where σe2\sigma^2_e is the regression error variance and σ2\sigma^2 is the long-run variance):

T(α^i1)HiKi,tα^i(σ/σe)Hi(giKi)1/2T(\hat\alpha^i - 1) \Rightarrow \frac{H_i}{K_i}, \qquad t_{\hat\alpha^i} \Rightarrow \frac{(\sigma/\sigma_e) H_i}{(g_i K_i)^{1/2}}

where HiH_i, KiK_i, gig_i are functionals of a standard Brownian motion W(r)W(r) that depend on λ\lambda and differ across Models A, B, C. Because the break variables (DUt,DTt)(DU_t, DT^*_t) are nuisance regressors under the null, the limiting distributions are non-standard and shift with λ\lambda.

Critical values. The 5% critical value for the tt-statistic is approximately: Model A 3.76\approx -3.76; Model B 3.96\approx -3.96; Model C 4.24\approx -4.24 (tabulated for each λ{0.1,,0.9}\lambda \in \{0.1,\ldots,0.9\} in Tables IV–VI). The standard DF value (with constant and trend) is 3.41-3.41; ignoring the break makes rejection too easy.

AO versus IO. The Additive Outlier (AO) model assumes the break in the mean is instantaneous — used for Model A (1929 crash). The Innovational Outlier (IO) model allows the break to propagate gradually through the AR polynomial — used for Model B (1973 oil shock). The two forms require different test regressions and, for the AO model, the two-step procedure of first estimating the break-corrected series and then testing on residuals.

Empirical results. Applied to 13 Nelson-Plosser macroeconomic series: 11 reject the unit root (real gross national product (GNP), nominal GNP, industrial production, employment, wages at 1%; real per capita GNP, GNP deflator, money stock, common stock prices, real wages at 2.5%); non-rejections: consumer prices, velocity, interest rate. Quarterly real GNP (not in NP dataset): t^=3.98\hat{t} = -3.98, α^=0.86\hat\alpha = 0.86, lag k=10k=10.

Exogeneity postulate. The break dates 1929 and 1973 are treated as exogenous, known constants. This delivers sharper critical values than searching over λ\lambda, but raises data-snooping concerns addressed by Zivot-Andrews (1992); see the Zivot-Andrews (1992) subsection below.

Erratum. Perron-Vogelsang (1993) corrected errors in the AO-model asymptotic distributions. The IO-model results and all empirical conclusions of the paper are unaffected.

See Perron (1989).

Zivot-Andrews (1992) — Endogenous Break Date

Treats TBT_B as unknown; searches over λΛ\lambda \in \Lambda (a trimmed interior of (0,1)(0,1)):

infλΛtα(λ)\inf_{\lambda \in \Lambda} t_\alpha(\lambda) This minimizes the t-statistic over all candidate break dates, giving greatest weight to the trend-break hypothesis. Critical values are tabulated for this statistic (stricter than Perron's known-break values). Zivot-Andrews also show the break is frequently identified one period too early, a mis-timing issue.

Bai-Perron (1998) — Multiple Breaks at Unknown Dates

Model:

yt=xtβ+ztδj+uty_t = x_t' \beta + z_t' \delta_j + u_t

for regime j=1,,m+1j = 1,\ldots,m+1, with TjT_j unknown.

Estimation: minimize SSR over all (m+1)-partitions with minimum segment length h; dynamic programming achieves this efficiently.

sup F(k; q): tests stability against k fixed known breaks; generalizes Andrews (1993) sup F.

UDmax F(M, q): maxk=1,,MsupΛkFT(λ1,,λk;q)\max_{k=1,\ldots,M} \sup_{\Lambda_k} F_T(\lambda_1,\ldots,\lambda_k; q) with equal weights ak=1a_k = 1; null is stability, alternative is unknown kMk \leq M breaks.

WDmax F(M, q): maxk=1,,Mc(q,α,1)c(q,α,k)supFT(k;q)\max_{k=1,\ldots,M} \tfrac{c(q,\alpha,1)}{c(q,\alpha,k)} \cdot \sup F_T(k; q); weights equalize marginal p-values across k so that power does not decline as the true number of breaks grows. Critical values tabulated for M=5M=5, α=0.05\alpha=0.05.

Bai (1999) sequential LR test: null l breaks vs. l+1; statistic is the normalized drop in optimal SSR from l to l+1 breaks; limiting distribution has known closed-form density over the search interval, yielding analytic critical values.

Hansen (2000) — Systems with Non-Stationary Regressors

Extends structural change testing to systems where regressors xnix_{ni} may themselves be subject to structural change or non-stationarity. Key insight: standard SupF/ExpF/AveF tests are not invariant to structural change in the regressors — a significant test could reflect instability in xx rather than in β\beta.

Fixed-regressor bootstrap: treats all regressors (including lagged dependents) as if fixed; generates bootstrap null distributions that are valid even under arbitrary structural change in xx and under heteroscedastic errors. Achieves near-correct size in small samples, unlike asymptotic critical values.

Bayesian Random Shift Models — McCulloch-Tsay (1993)

Classical break tests require specifying the number of break points. A fully Bayesian alternative treats each observation's shift indicator as a latent variable, eliminating the model-selection problem.

RLAR (Random Level-shift AR): The mean level evolves as μt=μt1+δtβt\mu_t = \mu_{t-1} + \delta_t\beta_t, where δtiidBernoulli(ε)\delta_t \overset{iid}{\sim}\mathrm{Bernoulli}(\varepsilon) and βtN(0,τ2)\beta_t \sim \mathcal{N}(0, \tau^2) is the jump magnitude. The autoregressive (AR) process yt=μt+xty_t = \mu_t + x_t runs around this stochastically shifting mean.

RVAR (Random Variance-shift AR): The innovation variance shifts multiplicatively — σt2=σ02Gt\sigma_t^2 = \sigma_0^2 \cdot G_t where GtG_t cumulates the multiplicative jumps βj2\beta_j^2 at times when δj=1\delta_j = 1.

Gibbs sampler. All full conditional posteriors are conjugate and closed-form: AR coefficients (N\mathcal{N}), innovation variance (χ2\chi^{-2}), jump sizes (N\mathcal{N} or χ2\chi^{-2}), shift probability ε\varepsilon (Beta\mathrm{Beta}). The binary indicators δt\delta_t are sampled from their posterior odds ratios. Drawing 3 indicators jointly rather than one at a time improves convergence.

Probit extension: Replace constant ε\varepsilon with εt=Φ(wtγ)\varepsilon_t = \Phi(w_t'\gamma) where wtw_t are exogenous predictors (Albert-Chib 1993 data augmentation for γ\gamma). Enables prediction of future shifts: given a known upcoming event, the model returns P(shift at twt)P(\text{shift at }t \mid w_t).

Contrast with Markov-switching VAR (MS-VAR): Markov-switching models (Hamilton 1989) allow persistent states — once a regime change occurs, the system stays in the new state until the next switch. RLAR/RVAR assume each period's shift is independent and identically distributed (i.i.d.) Bernoulli, making them better suited to isolated, infrequent jumps rather than persistent regime shifts.

Gasoline application: Monthly US price changes (1978–1991, n=158n=158). RVAR posterior shift probability ε^0.11\hat\varepsilon \approx 0.11; detected shifts align with OPEC events. RVAR-Probit: P(shiftworld event)0.73P(\text{shift} \mid \text{world event}) \approx 0.73 vs. P(shiftno event)0.12P(\text{shift} \mid \text{no event}) \approx 0.12. See McCulloch-Tsay (1993).

Bayesian Changepoint in Markov-Switching Model — Kim and Nelson (1999b)

When the parameter instability occurs within a regime-switching model — i.e., the hyperparameters of the Markov-switching process themselves shift — classical sup-F tests are non-standard because the transition probability q00q_{00} is a nuisance parameter absent under the null. Kim and Nelson (1999b) resolve this by treating the changepoint as an absorbing Markov state DtD_t:

Pr[Dt=0Dt1=0]=q00,Pr[Dt=1Dt1=1]=1\Pr[D_t=0|D_{t-1}=0] = q_{00}, \qquad \Pr[D_t=1|D_{t-1}=1] = 1

so that once the economy transitions to the post-break regime, it stays there permanently. Conditional on DtD_t, the MS parameters shift: μSt,t=μSt+μSt,postDt\mu^*_{S_t,t} = \mu_{S_t} + \mu_{S_t,\text{post}}D_t and σt2=(1Dt)σ02+Dtσ12\sigma^2_t = (1-D_t)\sigma^2_0 + D_t\sigma^2_1. The Bayesian framework eliminates the nuisance-parameter problem: q00q_{00} appears in the joint posterior and is integrated out via the Gibbs sampler. The changepoint posterior p(τdata)p(\tau|\text{data}) is a byproduct of sampling the DTD_T sequence.

Marginal likelihoods (Chib 1995/1998 reduced-run decomposition) compare four models: no break, break in shift parameters only, break in variance only, break in both. Applied to U.S. real gross domestic product (GDP) 1953:II–1997:I, the break in shift parameters model wins (lnm=247.01\ln m = -247.01), locating the break at 1984:Q1 and identifying a narrowing boom–recession gap as the dominant source of the Great Moderation.

See Great Moderation, Markov-Switching VAR, Kim-Nelson (1999b).

Why It Matters

Open Questions

Related