Definition
Unit root inference concerns the problem of drawing conclusions about the autoregressive parameter ρ in dynamic models when ρ may be at or near 1. In classical statistics, the sampling distribution of the maximum likelihood estimator (MLE) ρ^ changes discontinuously at ρ=1 (Dickey-Fuller 1979), producing non-standard asymptotic distributions that require separate tables. In Bayesian statistics with a flat prior, the posterior is proportional to the likelihood, which retains the same Gaussian shape whether or not ρ=1. The two approaches give qualitatively different — and in some cases opposite — probability statements about the same data.
Key Ideas
- The likelihood function in an autoregressive (AR) model is Gaussian-symmetric around ρ^, exactly as in a static regression, regardless of whether ρ=1.
- The asymmetric Dickey-Fuller distribution is a property of the MLE's sampling distribution, not of the likelihood or the Bayesian posterior.
- Flat-prior Bayesian analysis leads to standard symmetric posteriors; using Dickey-Fuller (DF) p-values as posterior probabilities implicitly imposes a prior that irrationally favors explosive ρ>1 values.
- The likelihood principle (Berger-Wolpert 1984) states that inference should be based on the likelihood function, not on the sampling behavior of estimators; AR models satisfy the principle even at the unit root boundary.
How It Works
The Model and Setup
Consider the AR(1) model
yt=ρyt−1+εt,εt∼N(0,σ2)
with the ordinary least squares (OLS) estimator
ρ^=∑t=1Tyt−12∑t=1Tyt−1yt
Conditional on y0 and with σ2 known, the likelihood is Gaussian in ρ regardless of whether ρ<1, ρ=1, or ρ>1. This is a classical result that follows from the Gaussianity of the disturbances.
The Helicopter Tour (Sims-Uhlig 1991)
Treat ρ as having a flat (uniform) prior over (0.80,1.10) and construct the joint probability density function (pdf) of (ρ,ρ^) by Monte Carlo (31 values of ρ; 10,000 simulations each; T=100). The joint pdf forms a 3D surface. Two cross-sections reveal the classical/Bayesian contrast:
| Cross-section |
Direction |
Shape |
Interpretation |
| Fix ρ=1, vary ρ^ |
Along ρ=1 plane |
Asymmetric, skewed left |
Classical sampling distribution of ρ^ under unit root |
| Fix ρ^=1, vary ρ |
Along ρ^=1 plane |
Symmetric around ρ=1 |
Posterior p(ρ∣ρ^=1) under flat prior |
Both cross-sections are views of the same surface; the asymmetry is a property of how you slice it, not of the underlying likelihood.
DF P-values vs. Posterior Probabilities
For an observed ρ^=0.95 with T=100:
| Hypothesis |
DF one-sided p-value |
Flat-prior posterior probability |
| ρ=0.9 (lower tail) |
0.04 (reject at 5%) |
P(ρ<0.9∣ρ^=0.95)≈0.07 |
| ρ=1.0 (upper tail) |
0.12 (fail to reject) |
P(ρ>1.0∣ρ^=0.95)≈0.07 |
The posterior is symmetric: deviations of equal magnitude above and below ρ^ are equally (un)likely. The DF test procedure treats them very differently.
The Implicit Prior
If a researcher systematically interprets DF p-values as posterior probabilities, the implied prior π(ρ∣ρ^) can be recovered by comparing the pseudo-posterior to the flat-prior posterior. For each observed ρ^:
- The implied prior puts 2–3× more weight on ρ≈1 than on ρ≈0.9.
- The implied prior is sample-dependent: it changes with ρ^, so no single prior rationalizes the procedure.
- The implied prior keeps increasing into ρ>1 without bound — it assigns greater prior probability to explosive dynamics than to moderate stationarity.
The Likelihood Principle Implication
The likelihood principle asserts that all information about ρ from the data is contained in the likelihood function L(ρ∣y). Since L(ρ∣y) is symmetric around ρ^ in AR models, any inference satisfying the likelihood principle must treat ρ=ρ^+δ and ρ=ρ^−δ as equally supported by the data for any δ. The Dickey-Fuller framework violates this because it incorporates information about the behavior of ρ^ in samples other than the one observed.
Why It Matters
- Tail sensitivity (Geweke 1993): Unit root conclusions in the Nelson-Plosser (1982) macroeconomic series are not robust to the assumption of Gaussian disturbances. Replacing normality with a Student-t likelihood (estimated degrees of freedom typically 3–7) systematically reduces posterior odds in favor of difference stationarity; the trend-coefficient posterior standard deviation (SD) shrinks as well. The same data, under a different tail assumption, can support appreciably different conclusions. See Geweke (1993).
- Reporting practice: Conventional t and F statistics describe the likelihood shape correctly in autoregressions; they should be reported even when DF corrected p-values are also provided.
- Model simplification decisions: Whether to difference, allow for cointegration, etc., should be based on the posterior probability of near-unit-root behavior, not on DF p-values that carry implicit explosive priors.
- Bayesian priors for macro models: Flat or near-flat priors on ρ are well-behaved near the unit root boundary; the apparent "non-standard asymptotics" are an artifact of the frequentist framing.
- Connects to Sims (1988): The graphical and quantitative results here extend the analytical argument in Sims (1988) "Bayesian Skepticism on Unit Root Econometrics."
Dickey-Fuller as a Special Case (Sims-Stock-Watson 1990)
Sims, Stock, and Watson (1990) show that the Dickey-Fuller unit-root test is a special case of their general canonical-regressor framework (Section 5). In the univariate AR(p) model, the F-statistic for the unit-root null H0:δ3=1 (where δ3 is the autoregressive coefficient on the lagged level) has the limiting distribution:
F⇒∫01W(t)2dt[∫01W(t)dW(t)]2
where W is a standard Wiener process. This is the nonstandard Dickey-Fuller distribution. It arises because the restriction δ3=1 constrains the coefficient on the I(1) canonical regressor Zt3, not on a mean-zero stationary canonical regressor — exactly the condition identified in Theorem 2 as producing a nonstandard limit. The classical Dickey-Fuller asymptotic theory thus emerges as a consequence of the general principle rather than a sui generis result.
Prior Robustness in Real Data: OECD Unemployment (Summers 2002)
Summers (2002) applies Bayesian unit root inference to unemployment rates in 16 Organisation for Economic Co-operation and Development (OECD) countries, using five priors on ρ: N(1,1000), N(1,0.5), Lubrano's (1995) extended Beta, Beta(10,2), and Berger-Yang (1994). Posterior draws from the base N(1,1000) prior are reweighted to alternative priors via importance sampling (Geweke 1999): {ρji}={ρi⋅fj(ρ)/fN(ρ)} — no additional MCMC required.
The key empirical finding is that unlike the stylized AR(1) example in Sims-Uhlig (1991), OECD unemployment posteriors are nearly identical across all five priors. For 12 of 16 countries, ≥94% of the posterior lies below ρ=0.95 regardless of prior. Only France shows clear unit-root evidence (median ρ≈0.98, 24–32% mass above 1.0); Italy and Spain are borderline.
This illustrates an important distinction: prior sensitivity in unit root inference is high in data-poor or misspecified settings (stylized AR(1) with no structural breaks) but low when the data are informative — as with real macroeconomic series modelled with structural breaks. The theoretical case (unemployment is bounded) and the empirical posterior both point toward stationarity.
Open Questions
- The model analyzed (AR(1), known σ2, condition on y0) is deliberately simplified. The extension to multivariate systems (vector autoregression, VAR), unknown σ2, and multiple lags involves a higher-dimensional joint pdf that is harder to visualize but the qualitative conclusion extends.
- The practical significance of the explosive-prior implication depends on whether applied researchers actually interpret DF p-values as posteriors. If they are used purely for size-correct hypothesis testing, the critique has less force.
- Uhlig (1994) extended the Bayesian perspective to the multivariate case, arguing for agnosticism about exact unit roots in macro models.
Related