Cointegration describes a situation in which individually non-stationary variables share stationary linear combinations — the cointegrating relations. The cointegrating vectors collect the linear weights that render those combinations stationary, and the loading matrix captures how fast each variable corrects toward the long-run equilibria.
The cointegration concept emerged from the spurious regression problem. Granger and Newbold (1974) showed by simulation that regressing two independently generated random walks on each other by ordinary least squares (OLS) produces t-statistics that are systematically too large and that is systematically too high — with strongly autocorrelated residuals (Durbin-Watson ). Phillips (1986) provided the asymptotic theory: the t-statistic diverges to infinity as and does not converge to the true (zero) value. See Spurious Regression for full details.
The cointegration concept (Granger 1981) identifies the exception: if the regression residual is I(0), the regression is not spurious but instead captures a genuine long-run equilibrium. Cointegration is thus the condition under which levels regressions are valid.
Engle and Granger (1987) developed the first systematic estimation method:
Step 1 (OLS for ): Estimate the cointegrating vector by OLS regression of on . Stock (1987) showed the OLS estimator is superconsistent: it converges to at rate rather than the usual . The superconsistency means that pre-estimation error in has no first-order effect on subsequent inference.
Step 2 (ML for , ): Hold fixed, substitute the estimated equilibrium error into the error correction model (ECM), and estimate by maximum likelihood. The second-stage estimators are consistent and asymptotically normal under standard conditions.
This two-step approach "opened the gates for a flood of applications" (Nobel Committee 2003) by giving applied economists a tractable way to incorporate long-run relationships. Johansen (1988, 1991) is the "second generation": rather than separating OLS and ML, he derives the maximum likelihood estimator (MLE) of the cointegrating space directly via reduced-rank regression, simultaneously obtaining β, α, and sequential LR tests for the rank . Johansen builds entirely on ML rather than partly on OLS, which gives efficiency gains and unified hypothesis testing.
Sims, Stock, and Watson (1990) show via a canonical-regressor decomposition that when the Vector Autoregression (VAR) is estimated in levels by OLS, the limiting distribution of all coefficient estimators is identical to the distribution that would obtain if the cointegrating vector were known a priori. The superconsistency of the OLS estimator (convergence at rate rather than ) means that plugging in in place of the true has no first-order effect on any subsequent inference — it contributes nothing to the asymptotic distribution.
The practical implication is that the Engle-Granger two-step procedure is asymptotically redundant: there is no efficiency gain from estimating a cointegrating regression in Step 1 before estimating the error-correction model in Step 2, relative to simply estimating the levels VAR directly. The result rationalizes VAR-in-levels estimation as a complete alternative to the Johansen/Engle-Granger ECM approach for purposes of coefficient inference (though rank testing and structural long-run analysis still require explicit cointegration methods).
Dickey-Jansen-Thornton (1991) survey three competing methods unified by reparameterizing the VAR() in error-correction form:
The rank of equals the number of cointegrating vectors ; the system therefore has common stochastic trends. The three approaches differ in how they estimate rank():
Engle-Granger (1987). Choose a normalization variable, regress it on the others by OLS, and apply an augmented Dickey-Fuller (ADF) test to the residuals. Critical values are non-standard. Main limitation: normalization-sensitive — different choices of dependent variable can yield different cointegrating vector estimates and different test outcomes.
Stock-Watson (1988). Factor-analyze ; the directions of largest variance span the common stochastic trends, and their orthogonal complement contains the cointegrating relations. Avoids normalization sensitivity but requires choosing how many trends to extract.
Johansen (1988, 1991). Performs canonical correlation analysis of and after projecting out lagged-difference regressors. The factorization yields ML estimates of the cointegrating vectors and loading matrix , with sequential trace and max-eigenvalue LR tests for rank. Full ML gives efficiency gains over the OLS-based Engle-Granger approach.
Multiple comparison power loss. Sequentially testing rank 0, then rank 1, then rank 2, … inflates total type I error and shifts critical values far from standard Dickey-Fuller tables. Power falls as grows.
Economic interpretation caveat. The cointegrating vectors estimated from a reduced-form VAR are linear combinations of all structural equations — they cannot be interpreted as individual structural relations. A cointegrating vector that matches a theoretical identity (Fisher equation, quantity-theory identity) does so by coincidence of the reduced form, not because the system identifies structural parameters.
Money demand application (US quarterly, 1953.2–1988.4). Johansen finds 1 cointegrating (CI) vector for {real M1, real income , interest rate}; income elasticity – (unity not rejected); interest elasticity significantly negative. M2 and NM1M2 each yield 1 CI vector. Monetary base combined with yields 2 CI vectors with R3M but only 1 with R10Y — rank finding is sensitive to interest rate specification.
Seasonal cointegration (Hylleberg, Engle, Granger, and Yoo 1990): when series are seasonally integrated ( but ), cointegration is defined at the seasonal frequencies separately. Unit roots at (biannual) and (quarterly) require different test regressions from the standard zero-frequency case.
Multicointegration (Granger and Lee 1990): when and are cointegrated with equilibrium error , the cumulated disequilibrium is I(1). If is cointegrated with or , the variables are multicointegrated — a richer long-run structure relevant to stock-flow models (e.g., sales and production: their difference is change in inventory, and the inventory level may be cointegrated with sales).
Threshold cointegration (Balke and Fomby 1997): error correction is nonlinear — adjustment toward equilibrium only occurs when the disequilibrium exceeds a threshold (transaction costs, menu costs). The ECM becomes where is the threshold. Granger and Swanson (1996) provided the theoretical framework for nonlinear cointegration with transaction costs.
For an -dimensional I(1) vector with cointegrating vectors collected in (so ), the following three representations are equivalent:
1. Error-correction (ECM):
where and are finite-order matrix polynomials and (at least one variable corrects toward the equilibrium ).
2. Moving-average (Wold MA):
where has reduced rank (exactly zero eigenvalues — one for each cointegrating relation). The nonzero rows of span the permanent component; the cointegrating vectors are in the left null space of .
3. Autoregressive:
Standard VAR representation with the ECM constraint imposed (where is the matrix of adjustment speeds and is the cointegrating matrix; Engle-Granger use the notation reversed from Johansen's later convention). Specifically, the VAR has a unit root at of multiplicity , with all remaining roots outside the unit circle.
Key insight: The reduced rank of is what distinguishes a cointegrated system from a plain multivariate random walk. If had full rank , no stationary linear combination would exist. The rank deficiency by exactly is the algebraic fingerprint of cointegrating relations.
All seven tests are applied to the OLS residuals from the cointegrating regression. Critical values are non-standard — obtained by Monte Carlo simulation (Tables I–III of the paper) — because are generated residuals, not observed series. Standard Dickey-Fuller tables are invalid here.
| Test | Statistic | Notes |
|---|---|---|
| CRDW (Cointegrating Regression Durbin-Watson) | Durbin-Watson on | Reject if DW ; simple but low power |
| DF | DF -statistic on | No lags; serially correlated residuals inflate size |
| ADF | ADF -statistic (recommended) | Add lagged ; best power of all seven |
| RVAR | -type test from restricted bivariate VAR | Imposes ; equivalent to DF for |
| UVAR | Likelihood ratio test from unrestricted VAR | Does not impose from step 1 |
| ARVAR | Augmented version of RVAR | Adds lags; like ADF for RVAR |
| DVAR | VAR of differences + EC term | Compares restricted vs unrestricted |
Recommendation: ADF on residuals is preferred. The CRDW is a quick check only. The VAR-based tests have lower power than ADF when is well estimated.
Critical value example (from Table I, , , 5% level): CRDW ; DF ; ADF (2 lags) . These are substantially larger in absolute value than the corresponding unit-root critical values because the residuals are estimated rather than observed.
By the Granger Representation Theorem (Engle and Granger 1987), the cointegration hypothesis implies that the long-run impact matrix in the VAR has rank . Writing the VAR in error correction form:
where , is a vector of deterministic components (constant, trend, dummies), and .
Since , write:
The cointegration rank is determined by Johansen's reduced rank regression. The test statistics are:
where are eigenvalues from the generalized eigenvalue problem involving the concentrated likelihood. Critical values depend on the deterministic specification (constant inside/outside cointegrating space, trend).
The factorization is not unique: for any invertible matrix , . Identifying and separately requires restrictions. Bauwens and Lubrano (1994a) use equation-by-equation linear restrictions on each cointegrating vector:
where () and () are known. A just-identified system imposes exactly restrictions in total (typically normalisations and exclusion restrictions). An over-identified system imposes more, producing testable cross-equation restrictions on .
Let () be the matrix of , () the matrix of , and the matrix of the remaining regressors (, ). Define:
Under a non-informative prior on and a prior on the identified cointegrating vectors, the marginal posterior of is:
where and , and is the number of columns of . This density has no closed form when — integration over requires numerical methods (importance sampling or Gibbs Sampler).
Given , the model in (1) becomes a standard multivariate regression:
The conditional posterior of the stacked parameter block is matrix-Student with expectation equal to the generalized least squares (GLS) estimator:
where . At each Gibbs step, can be drawn jointly from this matrix-Student (or its moments can be accumulated for Rao-Blackwell estimation of ).
The Bayesian ECM (BECM) approach (LeSage 1990) estimates super-consistently via Johansen ML, then includes the disequilibrium terms as regressors in a Bayesian VAR (BVAR) of differences. The -th equation becomes:
The flat-prior problem. A natural prior assigns informative Minnesota-type priors to the lag coefficients (prior variance ) but diffuse priors to (large ). This is inappropriate because converges at the standard rate — not the super-consistent rate of (). Tight priors on combined with flat priors on therefore over-weight the ECM correction relative to short-run dynamics, degrading short- and medium-horizon forecasts.
The fix (IP-BECM). Assign informative priors to with hyperparameters of moderate size, tuned by the same grid-search procedure as the other hyperparameters. Empirically (Italian GDP, consumption, investment, 1970–1995), the IP-BECM model dominates all alternatives in 13/15 forecast horizon-equation combinations. The flat-prior FP-BECM only outperforms a standard Minnesota BVAR at long horizons (12+ steps), where the random-walk prior's ignorance of cointegration becomes the binding constraint.
Felix and Nunes (2003) provide the clearest empirical demonstration of what goes wrong when a diffuse prior is placed on error-correction factor loadings in a BECM estimated with Johansen's multiple cointegrating vectors.
The mechanism. Johansen's trace test on a six-variable quarterly euro area system finds 5–6 cointegrating vectors, of which four are retained. A flat prior () treats all four estimated equilibrium relations as equally informative and allows the model to fit arbitrary linear combinations of them. Because estimates converge only at the slow rate — unlike the super-consistent rate of — a diffuse prior on over-weights the ECM correction relative to the evidence actually contained in the data. With four loosely estimated equilibrium relations, this effectively overfits.
Empirical result (6 variables, 12 forecast horizons, avg root-mean-square error (RMSE) relative to RW = 1.000):
| Model | Avg RMSE ratio |
|---|---|
| BVAR levels | 0.731 |
| BECM(EG)-IP (informative prior on α, single EG vector) | 0.754 |
| BECM(J)-IP (informative prior on α, 4 Johansen vectors) | 0.805 |
| BECM(EG)-FP (flat prior, single EG vector) | 0.857 |
| BECM(J)-FP (flat prior, 4 Johansen vectors) | 1.278 |
The BECM(J)-FP is 28% worse than a random walk. The single Engle-Granger vector is less susceptible because it provides only one (better-estimated) equilibrium correction.
The fix. Assign a finite informative prior variance Ω to α, tuned by the same grid search used for the Minnesota hyperparameters. This shrinks the factor loading estimates toward zero and limits the damage from the four loosely estimated Johansen vectors. BECM(EG)-IP is the second-best model overall and the best for real GDP at forecast horizons beyond 12 quarters. The result corroborates the theoretical argument in Amisano-Serati (1999) that informative priors on α are essential when α is not super-consistently estimated. See Minnesota Prior for the full Felix-Nunes hyperparameterization.
LeSage (1990) provides the largest-scale empirical comparison of ECM and BVAR forecasting to that point: 50 Ohio industries, monthly labor market data (manhours N, nominal wages W, prices P), rolling 25-month evaluation over 1983–1985 with 12-step-ahead horizons. Five models compete: ECM, unrestricted VAR, MVAR (Minnesota prior), BVAR (block-recursive prior), and BECM (Minnesota prior on lags, diffuse prior on EC term).
Cointegrated industries (7 of 50). The ECM dominates at forecasting horizons 3–12 months, producing ~20% less mean absolute percentage error (MAPE) at month 12 than any VAR or BVAR. The improvement is monotonically increasing with horizon — near-term ECM forecasts are similar to or slightly worse than VAR, while long-run forecasts are markedly better. Bayesian shrinkage (MVAR, BVAR) degrades performance for cointegrated series because the random-walk prior is misspecified: it ignores the long-run equilibrium. The BECM performs nearly as well as the ECM at horizons 5–12 months — contradicting Engle and Yoo's (1987) prediction that the misspecified prior would dominate and produce inferior forecasts.
Possibly cointegrated (5 industries, mixed ADF/DF results). The BVAR (block-recursive) wins at all 12 horizons — suggesting the variables are not truly cointegrated despite borderline test statistics. LeSage proposes using forecast comparisons as a practical tie-breaker when cointegration tests give conflicting results.
Non-cointegrated (38 industries). The MVAR (Minnesota prior) produces the best short-horizon forecasts (months 1–7); the BECM produces the best long-horizon forecasts (months 8–12), with ECM second. A regression of ADF t-statistics on forecast errors shows no systematic relationship, ruling out the low-power explanation: the error-correction variable itself is responsible for the long-horizon improvement even without confirmed cointegration.
Mechanism. In the BECM, the EC variable enters with a diffuse prior while the autoregressive lags are shrunk toward a random walk. Tighter shrinkage on the VAR terms increases the relative importance of the EC term, giving the practitioner direct control over the short-run vs. long-run balance. A sufficiently loose prior recovers the ECM as a limiting case.
When the cointegrating vector is known or imposed a priori, Bayesian estimation can proceed with a diffuse prior on the error correction loading while retaining informative (Minnesota-type) shrinkage on the short-run dynamics. Stark (1998) demonstrates this approach for a 7-variable U.S. macroeconomic system.
Model specification. Stark's BVEC(5) system comprises {ΔRGDP, Δ²PGDP, ΔRFF, ΔRPIM, ΔU, ΔRM2, ΔRTB}, estimated 1960Q4–1997Q4. A single EC term (the federal funds rate minus the 10-year Treasury bond rate) enters all equations, imposing the cointegrating vector between the two interest rates as a known I(0) relation. The loading on the EC term receives a diffuse prior in each equation, following LeSage (1990) and Joutz-Maddala-Trost (1995). Estimation uses Theil's mixed estimator, which coincides with the Bayesian posterior mean under normality.
EC significance and forecast relevance. In a rolling regression experiment (mid-1975–1997Q4), the spread enters significantly in the ΔRGDP equation (), ΔRFF (), ΔU (), and ΔRTB (). Removing the EC term raises the two-year unemployment RMSE from 0.805% to 1.113% — a 38% deterioration — with mixed effects on inflation and GDP.
Fisher relation: imposing cointegration can worsen forecasts. When a second EC term is added (imposing real rate stationarity as a Fisher relation), the two-year inflation RMSE rises from 1.52% to 2.36%. This illustrates a general principle: if the economic theory underlying a cointegrating restriction is empirically fragile, imposing it creates more specification error than it eliminates. The Fisher effect is weak enough in the U.S. data that the constraint actively damages inflation forecast accuracy.
In a cointegrated system with I(1) variables and cointegrating rank , there are common stochastic trends — the permanent components shared by all variables. Gonzalo and Granger (1995) provide an explicit identification and estimation method.
Let () be the orthogonal complement of the loading matrix , satisfying . The common trends are defined as:
The identifying restriction is that the equilibrium errors do not Granger-cause at very low frequencies — i.e., the transitory component has no persistent effect on the permanent component. Under this restriction, span the permanent and transitory components in a way that is simultaneously rotation-free and economically interpretable.
The full system decomposition is:
This is the multivariate analogue of the Beveridge-Nelson decomposition for a single series: just as the Beveridge-Nelson decomposition splits into a random-walk (permanent) and a stationary (transitory) part, the Gonzalo-Granger decomposition splits the -dimensional system into random-walk trends and stationary equilibria .
Economic interpretation. In a bivariate system of income and consumption with one cointegrating vector (the consumption-income ratio), there is one common trend (say, permanent income) and one transitory component (cyclical deviations from the long-run ratio). The Gonzalo-Granger decomposition identifies which linear combination of constitutes the common trend without imposing arbitrary normalization.
Granger (1997) caution. The decomposition provides the common stochastic trends even when the series are not strictly I(1) but merely "persistent" — the ECM-based construction is robust to near-unit-root and other persistent processes that fail to reject the I(1) null but are not pure unit root processes.
Despite being the dominant frequentist approach to cointegration, the Johansen (1991) MLE has several practical limitations that applied economists sometimes overlook:
Phillips (1991) establishes the theoretical foundation for why full-system ML dominates OLS for estimating cointegrating vectors. The key is a distinction between two asymptotic frameworks arising from how the cointegrating structure is parameterized.
Triangular ECM representation. Phillips works with the system:
where () is the cointegrating vector appearing linearly in (P1). This triangular form separates the components explicitly and is the key structural feature.
LAMN vs. LGF. The likelihood of the triangular ECM — which eliminates unit roots from the parameterization — is locally asymptotically mixed normal (LAMN): the normalized score converges to a mixture of normals (mixing on the random Fisher information matrix ), enabling chi-squared test statistics and Cramér-Rao efficiency bounds. By contrast, an unrestricted VAR in levels implicitly estimates unit roots alongside all other parameters and delivers a locally Gaussian functional (LGF) likelihood — producing nonstandard, nuisance-parameter-dependent distributions and no optimal inference theory.
Theorem 1 (iid errors): mixed-normal MLE. Under the triangular ECM, the full-system MLE satisfies:
which is conditionally given — a mixed normal (MN) distribution. The MLE is symmetrically distributed, median unbiased, and asymptotically efficient. Wald, LR, and Lagrange multiplier (LM) tests all converge to under this parameterization.
OLS and simultaneous-equations bias. OLS satisfies:
The additional term (cross-equation innovation covariance) induces a simultaneous-equations bias whenever and are correlated. OLS estimators of the cointegrating vector are not median unbiased and are asymptotically inefficient.
Single-equation ECM. The single-equation partial MLE is efficient iff (strict exogeneity of ). In the typical macro setting where all variables are jointly determined, this condition fails and single-equation ML is biased.
Transient dynamics need not be estimated. A key practical implication: only a consistent long-run covariance estimate is required — the transient dynamics (, , etc.) need not be jointly estimated by ML. This justifies the Phillips-Hansen (1990) fully modified OLS (FM-OLS) approach.
General linear processes (Theorem 1'). The mixed-normal limit holds for general linear process errors with long-run covariance — the spectral density at frequency zero replaces the short-run .
Toda and Phillips (1993) showed that standard Wald asymptotics for Granger noncausality tests in levels VARs fail generically when . The critical quantity is the rank of the subblock of the cointegrating matrix corresponding to the candidate causal variable. Partition where () is the potential cause and let denote the last rows of :
The practical implication: Granger causality tests must be conducted in a Johansen ECM framework after testing the cointegrating rank and the rank of . See Granger Causality for the full asymptotic theory and sequential testing procedure.
Warne (1997) extends the Sims-Stock-Watson (1990) result to nonlinear restrictions on VAR coefficients, with particular focus on nonlinear cross-equation (NCE) restrictions implied by rational expectations (RE) models. The key insight is that RE restrictions typically constrain the row space of the long-run impact matrix , and this is precisely why standard χ² asymptotics fail.
Setup. Write the VAR in the canonical regressor decomposition of Sims-Watson (1994): , where collects mean-zero stationary regressors, collects the non-stationary components, and is the deterministic trend. The coefficient blocks converge at rate , but converges at rate — different rates yield different contributions to the limiting Wald distribution.
Proposition 1. Under the null (smooth nonlinear restrictions), the Wald statistic converges to: where has components drawn from Brownian functionals, and is the limiting Jacobian. This distribution is non-χ² unless the restrictions do not constrain the coefficients.
NCE restrictions from RE models. A linear RE model generates restrictions of the form , where are known matrices and is a known vector. Under the null these imply nonlinear cross-equation restrictions on . The Jacobian — the block corresponding to the non-stationary regressors — fails to have full row rank whenever , because the NCE restrictions constrain the row space of and thus act on the cointegrating space. This is the structural reason the Wald statistic has a nonstandard limit.
Lower bound for r. Conveniently, the NCE restrictions provide a lower bound for the cointegration rank: the rows of (the complement induced by the model matrices) must themselves be cointegrating vectors, so . This bound selects the cointegrating space needed for the simulation, making the approach practical.
Expectations hypothesis application. For the term spread between one- and three-month U.S. bond yields (Campbell-Shiller 1991 data, January 1952–February 1987), the cointegration rank lower bound is , (no trend in the spread), giving and explicitly. The Wald statistic (q = 8 restrictions) is decisively rejected at the 99th percentile under all three distributions: χ² (32.00), simulated nonstandard asymptotic (33.86), and empirical bootstrap (37.58). Campbell-Shiller (1991) reached the same conclusion.
Small-sample caution. Monte Carlo evidence indicates the simulated asymptotic critical value at nominal 5% corresponds to an empirical rejection rate of approximately 10% when . Practitioners should use low nominal significance levels (1–2%) when applying this approach.
The limiting null distribution of the Johansen LR trace statistic depends not only on but also on which deterministic terms are assumed present in the data-generating process. Lütkepohl (1999, Table 1) classifies five cases by assumptions on the intercept and trend slope in :
| Case | Assumption | Interpretation |
|---|---|---|
| 1 | No deterministic terms; rarely applicable | |
| 2 | free, | Intercept only, no trend |
| 3 | free, , intercept restricted to cointegrating space | No linear trend in levels; intercept absorbed into |
| 4 | free, , | Linear trend in levels but not in cointegrating relations |
| 5 | , both free | Linear trend allowed in cointegrating relations |
In applied work Case 4 (trend in levels, no trend in cointegrating relations) is most common. For each case the critical values of the trace and max-eigenvalue statistic differ and are tabulated separately.
Saikkonen-Lütkepohl prior trend adjustment. An alternative for Cases 2–4 is to first estimate the mean/trend parameters by a feasible GLS procedure, subtract the fitted deterministic component from , and then apply the rank test to the adjusted series. Saikkonen and Lütkepohl (1997, 1998, 1999) show this "trend-first" procedure attains higher local power than the standard Johansen test under the respective null. In simulations (Lütkepohl-Saikkonen 1999), none of the three standard test variants uniformly dominates, but prior trend adjustment is generally preferred when the trend specification is known.
The cointegrating matrix is only identified up to an invertible rotation — so the object with natural prior meaning is not itself but the cointegrating space , which lives on the Grassmann manifold (the set of all -dimensional subspaces of ). Warne (2006) parametrizes the space by a companion matrix via , where is a fixed normalization matrix.
Flat-prior pathology. A flat prior on does not induce a uniform distribution over ; instead it overweights the region near (i.e., near normalization breakdown), as shown by Strachan and van Dijk (2003). This produces distorted posterior inference on the cointegrating relations.
Villani/Warne prior (uniform on Grassmann manifold). The matrix- prior
where is the orthogonal complement of , is algebraically equivalent to a uniform distribution over and therefore free of normalization artifacts. Warne (2006) combines this with a conjugate prior on the remaining parameters:
The prior is Warne's key addition relative to Villani (2005b): it introduces informative short-run dynamics analogous to a Litterman/Minnesota prior within the cointegrated VECM.
After integrating out analytically, the marginal posterior of is proportional to a matrix- density whose mode solves the generalized eigenvalue problem
where are Johansen-style concentrated moment matrices augmented by prior terms and rescaled by (Warne 2006, Proposition 4). The eigenvectors corresponding to the largest eigenvalues give the posterior modal cointegrating space. This is algebraically identical to the Johansen (1991) MLE but computed on modified moment matrices that incorporate prior information — the frequentist and Bayesian modal estimators coincide in form.
Because the cointegration rank and lag order are discrete unknowns, Warne computes the joint posterior from the marginal likelihood using a Chib (1995) identity applied to the VECM:
evaluated at posterior estimates . The last term is estimated without additional MCMC by averaging the analytic Gaussian conditional over Gibbs draws: .
Key simplification (Corollary 1). Under the informative prior, the marginal likelihood at rank (full rank, unrestricted VAR) is available in closed form — no MCMC required. This allows the entire grid to be evaluated efficiently: Bayesian lag-order selection at full rank uses the closed form; MCMC is only needed for . See Lag-Order Selection Criteria.
Bartlett's paradox. If all parameters — including — are given improper flat priors, the marginal likelihood integrates to infinity and rank posteriors are undefined. Villani's proper prior on is the minimal fix that renders the comparison well-defined.
Cointegration arises naturally in multi-population mortality modeling. Under the Lee-Carter structure (Lee and Carter 1992), the log central death rate for population at age and time is , where is a non-stationary mortality index. The non-divergence hypothesis — that mortality levels in related populations cannot drift apart indefinitely — translates exactly into the requirement that are cointegrated with cointegrating vector :
This motivates the bivariate VECM for the factors:
The adjustment coefficients and (both estimated negative in the UK data: and ) confirm that both populations are pulled back toward equilibrium, but at different speeds.
Key advantages over the RWAR (random walk + autoregressive) baseline (Cairns et al. 2011):
See Lee-Carter Model and Longevity Basis Risk for the broader context.
Standard Johansen (1991) rank tests assume that the VECM innovations are homoscedastic. In practice, financial data exhibit volatility clustering (GARCH effects), which inflates the residual variance estimate and can reduce the apparent spread between eigenvalues in the reduced-rank regression — biasing tests toward non-rejection of the null of no cointegration.
VECM-BEKK model. Bauwens, Deprins, and Vandeuren (BDV, 1998) estimate a bivariate VECM for the long rate and short rate of five countries with BEKK-GARCH innovations (see GARCH and BEKK-GARCH):
where follows a BEKK(1,1) process. Estimation is by quasi-maximum likelihood (QML): the full Gaussian log-likelihood is maximized jointly over all VECM and GARCH parameters.
Present value motivation. Campbell and Shiller (1987) show that if long rates are rational expectations of future short rates, the spread must be stationary, implying . Under the pure expectations hypothesis, . BDV accept for France, UK, and USA but not Belgium and Germany.
Main empirical findings (BDV 1998):
The broader methodological lesson is that joint estimation of long-run (cointegrating) and short-run volatility (GARCH) dynamics is preferable to two-step approaches that ignore the interaction. See GARCH and BEKK-GARCH for the BEKK model details and covariance stationarity condition.
The steady state VECM of Villani (2008) augments (1) with informative priors on the mean growth rate and on the mean of the cointegrating relations :
with the constraint . See Steady State VAR for details.