Structural Identification

varidentificationstructural-shockseconometricsincredible-identificationrational-expectationsnairublanchard-quah

Definition

Structural identification in the vector autoregression (VAR) context is the problem of recovering the contemporaneous coefficient matrix A0A_0 — and hence the structural shocks εt=A0ut\varepsilon_t = A_0 u_t — from the observable reduced-form covariance Σ=(A0A0)1\Sigma = (A_0 A_0')^{-1}. Since Σ\Sigma has n(n+1)/2n(n+1)/2 distinct elements and A0A_0 has n2n^2 free parameters, at least n(n1)/2n(n-1)/2 restrictions must be imposed. Sims (1980) argued that the exclusion restrictions traditionally used to achieve identification in large structural models are "incredible" — normalizations rather than genuine theory — and that the VAR with minimal recursive restrictions is a more honest alternative.

Key Ideas

How It Works

Incredible Identification (Sims 1980)

Sims (1980) identified three compounding sources of identification failure in the Cowles Commission tradition:

1. A priori restrictions as normalizations. The exclusion restrictions that render large structural models identified — which variables appear in which equations — are not derived from optimizing theory. They are imposed because identification requires them, not because theory mandates them. Liu (1960) anticipated this critique: any equation from a correctly-specified model should include all variables, since general equilibrium means all things affect all things. Arbitrary zeros produce a misspecified model whose structural interpretation is not credible.

2. The Dynamics Problem (Hatanaka 1975). Even granting some exclusion restrictions, identification of a dynamic structural model requires an exogenous instrument for every right-hand-side endogenous variable at every lag. Because the true lag length is unknown, the required instrument list is unbounded in principle. Under rational expectations this becomes acute: let the structural model be b(L)yt+b+(L)yt=c(L)xtb^- (L) y_t + b^+(L)^* y_t = c(L) x_t where L=L1L^* = L^{-1} is the forward operator and xtx_t is an exogenous forcing variable. Then:

yt=[b(L)]1[b+(L)yt+c(L)xt]y_t = [b^-(L)]^{-1}\bigl[-b^+(L)^* y_t + c(L)x_t\bigr]

The backward-looking coefficients bb^- are identified from the distributed lag of xtx_t on yty_t under standard rank conditions. But the forward-looking coefficients b+b^+ govern future conditional expectations of yty_t as seen from time tt. Identifying b+b^+ requires cross-equation restrictions from the model's full rational expectations solution — meaning the researcher must know the entire model to identify any one equation. Sims (1980, p. 9) describes this as making forward identification "orders of magnitude more difficult" than backward identification.

3. Forecasting under false restrictions. Misspecified models with false exclusion restrictions can still forecast well because restricted estimators beat unrestricted ones in high-dimensional problems — the variance reduction from restriction can offset the bias introduced by incorrect zeros. However, this statistical argument does not rescue the structural interpretation: a model that forecasts well under false restrictions cannot be used for valid policy analysis, since policy counterfactuals depend on structural parameters that are biased.

Continuous-time hiring example (Sims 1980, Eqs. 1–14). To illustrate the rational expectations dynamics identification problem, Sims works through a firm's optimal hiring model with quadratic adjustment costs. The firm's optimum implies nt=(1ρL)(1ρL1)xt/ρn_t^* = (1-\rho L)(1-\rho L^{-1}) x_t / \rho where xtx_t is the real wage and ρ<1\rho < 1 is a function of the discount rate and adjustment cost. The structural equation is:

nt=a0+a1nt1+j=0bjxtj+j=1bj+Et[xt+j]+utn_t = a_0 + a_1 n_{t-1} + \sum_{j=0}^\infty b_j^- x_{t-j} + \sum_{j=1}^\infty b_j^+ E_t[x_{t+j}] + u_t

Identification of bj+b_j^+ from data on (nt,xt)(n_t, x_t) requires either (i) assuming a specific autoregressive moving average (ARMA) process for xtx_t and imposing the cross-equation restriction that the coefficients on leads of xtx_t in the employment equation equal those implied by the rational expectations solution of the model — a restriction the researcher only knows if the structural model is fully specified — or (ii) having external instruments for Et[xt+j]E_t[x_{t+j}], which are almost never available. This makes the forward-looking coefficients effectively unidentifiable in practice without strong auxiliary assumptions.

Implication. The appropriate response is to minimize restrictions to those with genuine credibility — i.e., the VAR with only a Cholesky ordering assumption, rather than a structural model with hundreds of incredible exclusion restrictions.

The Identification Problem

Start from the structural system Axt=C(L)xt1+DεtAx_t = C(L)x_{t-1} + D\varepsilon_t (using Keating's notation, where AA0A \equiv A_0). The reduced-form residuals satisfy:

et=A1Dεt    Σe=A1DΣεD(A1)e_t = A^{-1}D\varepsilon_t \implies \Sigma_e = A^{-1}D\,\Sigma_\varepsilon\,D'(A^{-1})'

Parameter count (Keating 1992): AA has n2n^2 elements, DD has n2n^2 elements, Σε\Sigma_\varepsilon has n(n+1)/2n(n+1)/2 unique elements — a total of 2n2+n(n+1)/22n^2 + n(n+1)/2 unknowns. Σe\Sigma_e provides only n(n+1)/2n(n+1)/2 equations. Identification requires at least 2n22n^2 restrictions.

Standard normalizations consume 2n2n of these:

Assuming Σε\Sigma_\varepsilon diagonal (independent shocks) provides n(n1)/2n(n-1)/2 additional restrictions. The remaining gap is:

2n22nn(n1)2=3n(n1)22n^2 - 2n - \tfrac{n(n-1)}{2} = \tfrac{3n(n-1)}{2}

So at least 3n(n1)2\tfrac{3n(n-1)}{2} further restrictions on AA and DD must come from economic theory. In the common case D=InD = I_n (no simultaneous shock-to-variable mapping beyond the equation itself), this reduces to n(n1)2\tfrac{n(n-1)}{2} restrictions on AA alone.

Equivalently, in the Zha (2005) notation with A0A_0 absorbing both AA and DD:

Σ=(A0A0)1    A0Σ1A0=In\Sigma = (A_0 A_0')^{-1} \iff A_0' \Sigma^{-1} A_0 = I_n

This system has n(n+1)/2n(n+1)/2 equations in n2n^2 unknowns, under-determined by n(n1)/2n(n-1)/2.

Order condition (necessary): equation jj must have at least n1n-1 restrictions on aja_j (the jjth column of A0A_0).

Rank condition (sufficient): the Jacobian of the restrictions evaluated at the true A0A_0 must have full rank.

The Cooley-LeRoy Critique of Cholesky

The original "atheoretical" VAR practice separated residuals into orthogonal shocks via a Cholesky decomposition of Σe\Sigma_e: find the unique lower-triangular RR such that Σe=RR\Sigma_e = RR', then define vt=R1etv_t = R^{-1}e_t. This was presented as theory-free.

Cooley and LeRoy (1985) showed this claim is false. The Cholesky decomposition of a VAR ordered (x(1),x(2),,x(n))(x^{(1)}, x^{(2)}, \ldots, x^{(n)}) is algebraically equivalent to estimating the recursive system:

et(1)=vt(1)e_t^{(1)} = v_t^{(1)} et(2)=R1et(1)+vt(2)e_t^{(2)} = R_1 e_t^{(1)} + v_t^{(2)} et(3)=R2et(1)+R3et(2)+vt(3)e_t^{(3)} = R_2 e_t^{(1)} + R_3 e_t^{(2)} + v_t^{(3)} \vdots

where each vt(j)v_t^{(j)} is orthogonal to all previous shocks by construction. This is a fully recursive contemporaneous structural model — with n!n! equally valid orderings, each implying a different economic structure. Results sensitive to the ordering have no structural interpretation. This critique directly motivated the development of non-recursive structural VARs.

CEE Recursiveness Assumption (Christiano-Eichenbaum-Evans 1999)

A prominent application of recursive identification to monetary policy. Partition the VAR variables as Zt=(Xt,St,X2t)Z_t = (X_t, S_t, X_{2t}) where:

The recursiveness assumption — Fed observes XtX_t but not X2tX_{2t} before setting StS_t, and policy shocks have no immediate effect on XtX_t — implies A0A_0 is block lower-triangular, giving Cholesky identification.

Identification invariance result: Any A0A_0 in the identified family (block lower-triangular with positive diagonal) generates the same dynamic response of ZtZ_t to the monetary policy shock εts\varepsilon_t^s, regardless of the orthogonal rotation WW applied to the lower-right X2tX_{2t} block. Cholesky is one member of a large equivalence class, all giving the same policy impulse response function (IRF). See Monetary Policy Shocks.

Strategy 1 — Recursive Identification (Cholesky)

Restrict A0A_0 to be lower triangular with positive diagonal. This uniquely identifies A0A_0 as the inverse of the Cholesky factor of Σ\Sigma:

A0=L1,where Σ=LLA_0 = L^{-1}, \quad \text{where } \Sigma = L L'

The ordering of variables in yty_t determines the causal flow: variable jj does not respond contemporaneously to shocks to variables j+1,,nj+1, \ldots, n. Appropriate when theory genuinely predicts a recursive structure; otherwise it is a strong and fragile assumption.

Strategy 2 — Non-Recursive Short-Run Restrictions

A0A_0 is not triangular; instead, specific elements are set to zero based on economic theory. For example, in a monetary policy block:

A0=(00)A_0 = \begin{pmatrix} * & 0 & * \\ * & * & * \\ 0 & * & * \end{pmatrix}

Each column aja_j satisfies linear restrictions:

Qjaj=0,QjRqj×n,rank(Qj)=qjQ_j\, a_j = 0, \quad Q_j \in \mathbb{R}^{q_j \times n}, \quad \text{rank}(Q_j) = q_j

Let UjRn×(nqj)U_j \in \mathbb{R}^{n \times (n - q_j)} be an orthonormal basis for the null space of QjQ_j. Then:

aj=Ujbja_j = U_j\, b_j

where bjRnqjb_j \in \mathbb{R}^{n-q_j} are the free parameters. Analogous restrictions Rjfj=0R_j f_j = 0 apply to the lag coefficients fjf_j, with free parameters gj=Vj1fjg_j = V_j^{-1} f_j.

Strategy 3 — Long-Run Restrictions

Applied to a VAR in first differences (permanent-shock case), the long-run effect matrix Θ(1)=i=0Θi\Theta(1) = \sum_{i=0}^\infty \Theta_i satisfies the key identifying equation (Keating 1992, eq. 17):

[Iβ(1)]1Σe[Iβ(1)]=Θ(1)ΣεΘ(1)[I - \beta(1)]^{-1}\,\Sigma_e\,[I - \beta(1)]^{-\prime} = \Theta(1)\,\Sigma_\varepsilon\,\Theta(1)'

where β(1)==1pβ\beta(1) = \sum_{\ell=1}^p \beta_\ell is the sum of VAR lag coefficient matrices. The left side is fully observable from the estimated reduced-form VAR. Restrictions on Θ(1)\Theta(1) — e.g., that demand shocks have zero long-run effect on output, [Θ(1)]y,demand=0[\Theta(1)]_{y,\text{demand}} = 0 — identify the structural parameters without imposing any contemporaneous restrictions.

In the Zha (2005) notation, the long-run cumulative response is:

Φ=(In=1pB)1A01\Phi_\infty = \left(I_n - \sum_{\ell=1}^p B_\ell\right)^{-1} A_0^{-1}

Long-run restrictions are computationally fragile near unit roots and are not valid when the variables are cointegrated without appropriate adjustment (King-Plosser-Stock-Watson 1991).

Advantage over contemporaneous restrictions: long-run restrictions do not require contemporaneous exclusion restrictions, which Keating (1990) showed are generally invalid under rational expectations — any observable variable can signal future events, making contemporaneous zero restrictions implausible.

Application: Fisher Decomposition (St-Amant 1996)

St-Amant (1996) applies the Blanchard-Quah long-run restriction to a bivariate system xt=(Δit,rt)x_t = (\Delta i_t, r_t)' where iti_t is the nominal interest rate and rt=it,kπtr_t = i_{t,k} - \pi_t is the ex post real rate. The single off-diagonal long-run restriction: ex ante real interest rate shocks have zero permanent effect on the nominal rate level, i.e., A(1)1,real=0A(1)_{1,\text{real}} = 0. This is justified by the long-run Fisher effect — nominal rates and inflation expectations are cointegrated (1,1) and the real rate is stationary. The recovered shocks are: (1) an inflation expectation shock (permanent), and (2) a real rate shock (transitory). Applied to U.S. 1-year and 10-year bond rates, Feb 1957 – Jun 1995. See Fisher Hypothesis.

Application: NAIRU Identification (Zhao undated; Blanchard-Quah 1989)

A canonical bivariate application to the inflation-unemployment system. The VAR is Xt=(Δπt,ut)X_t = (\Delta\pi_t, u_t)' where inflation is first-differenced (I(1)) and unemployment is stationary. The two structural disturbances are:

The NAIRU is the counterfactual path of unemployment when ε20\varepsilon_2 \equiv 0; core inflation is the counterfactual path when ε10\varepsilon_1 \equiv 0. This recovers a time-varying NAIRU without imposing any parametric model of price-setting. See NAIRU and Zhao (undated).

Independence Testing

Once structural shocks ε^t\hat{\varepsilon}_t are recovered, their mutual independence can be tested (Leeper-Zha 2003). This provides a diagnostic for identification validity that is unavailable in classical structural models.

Posterior Simulation for Restricted Structural VARs (SVARs) (Waggoner-Zha 2000)

Under the Sims-Zha (1998) reference prior, the posterior over the contemporaneous matrix AA under linear zero restrictions is non-Gaussian. The posterior lies on a curved, non-elliptic ridge in parameter space, so importance sampling with a Gaussian proposal assigns nearly all weight to a single draw — the sampler collapses.

Waggoner and Zha (2000) derive a Gibbs sampler that resolves this. For each equation ii, write ai=Uibia_i = U_i b_i (free contemporaneous parameters via null-space basis UiU_i) and fi=Vigif_i = V_i g_i (free lag parameters). The sampler alternates:

  1. Draw bibi,g,Yb_i \mid b_{-i}, g, Y: construct an orthonormal basis {wj}\{w_j\} via LU decomposition; draw bi=jβjwjb_i = \sum_j \beta_j w_j where β1UW\beta_1 \sim \text{UW} (Univariate Wishart) and βj2N(0,T1)\beta_{j \geq 2} \sim \mathcal{N}(0, T^{-1}).
  2. Draw gibi,Yg_i \mid b_i, Y: Gaussian conditional — standard restricted ordinary least squares (OLS).

When exclusion restrictions are block-recursive (contemporaneous matrix block-lower-triangular after permutation), draws across equations are exactly independent (Corollary 1), so no burn-in or thinning is needed. See Gibbs Sampler for the full mechanics and independence proof.

Why It Matters

Identification is the central methodological challenge in applied VAR work. The choice of restrictions determines which shocks can be separately labeled and quantified — and hence what policy counterfactuals are possible.

Open Questions

Related