Instrumental Variables

methodscausal-inferenceLATEnatural-experimentsendogeneityidentificationpotential-outcomesnoncompliancetreatment-effectsselection-modelsATEepidemiology

Definition

Instrumental variables (IV) is a method for estimating causal effects when the variable of interest is endogenous — that is, when individuals who differ on that variable also differ on unobserved characteristics that affect outcomes, making ordinary least squares (OLS) invalid. An instrument is a variable that (1) is correlated with the endogenous variable (relevance) and (2) affects outcomes only through the endogenous variable, not directly, and is uncorrelated with unobserved outcome determinants (exclusion restriction). When these conditions hold, observed variation in outcomes across instrument values must be caused by the endogenous variable — the instrument creates quasi-experimental variation that mimics a randomized assignment.

Key Ideas

The Angrist, Imbens, and Rubin (AIR) (1996) Framework: Five Assumptions and Compliance Types

Angrist, Imbens, and Rubin (1996) embedded IV in the Rubin Causal Model and showed that five assumptions are jointly sufficient for a causal interpretation of the IV estimand as LATE:

  1. Stable Unit Treatment Value Assumption (SUTVA): No interference between units; a single well-defined treatment version.
  2. Ignorable assignment: Z is randomly (or conditionally ignorably) assigned.
  3. Exclusion restriction: Z affects Y only through D — any effect of Z on Y must operate via D. Untestable; must be argued substantively. The paper separates this cleanly from ignorable assignment, which the old econometric formulation bundled together as "zero correlation between instrument and disturbances."
  4. Instrument relevance: E[D(1)D(0)]0E[D(1)-D(0)] \neq 0 (the first stage is nonzero). Directly testable.
  5. Monotonicity: D(1)D(0)D(1) \geq D(0) for all units — no defiers (no one takes treatment precisely when assigned to control). Implied by designs that deny treatment to controls.

Compliance typology — for any binary instrument, the population partitions into four types:

D(0)=0D(0)=0 D(0)=1D(0)=1
D(1)=1D(1)=1 Complier — IV identifies ATE for this group Always-taker — exclusion: Z has no effect on Y
D(1)=0D(1)=0 Never-taker — exclusion: Z has no effect on Y Defier — ruled out by monotonicity

Under Assumptions 1–5, IV=LATE=E[Y(1)Y(0)complier]\text{IV} = \text{LATE} = E[Y(1)-Y(0) \mid \text{complier}]. The share of compliers equals the first-stage coefficient.

Sensitivity: Violations of the exclusion restriction bias IV by (noncomplier direct effect) ×\times (noncomplier odds); violations of monotonicity bias IV by (defier share) ×\times (complier-defier treatment effect gap). Both biases shrink as the instrument strengthens — weak instruments are fragile to both.

Heckman (1996) Critique of the AIR Framework

Heckman's comment in the same Journal of the American Statistical Association (JASA) issue raises four challenges:

  1. Priority: The potential outcomes framework is equivalent to the econometric switching regression model (Quandt 1958; Maddala and Nelson 1975) and Roy model (Heckman and Honoré 1990), which predate Rubin (1974). The novelty claim is overstated.
  2. Weaker conditions suffice for average treatment effect on the treated (ATT): The mean treatment effect on the treated — E[Y1Y0D=1,X]E[Y_1-Y_0 \mid D=1, X] — is identified under mean independence alone (A-1: ZE[Y0X]Z \perp E[Y_0|X]; A-2: ZE[D1D0X]Z \perp E[D_1-D_0|X]) without full distributional independence and without monotonicity. ATT targets an observable subpopulation (the treated); LATE targets an unobservable one (compliers). Heckman argues ATT is the more relevant evaluation parameter.
  3. Behavioral content of AIR assumptions: The independence conditions are not statistical — they impose Granger noncausality restrictions on the data-generating process. These are violated in the Roy model, competing risks, and Gronau-Heckman labor supply model whenever agents have private information about their own gains (essential heterogeneity). AIR assumptions are strong behavioral claims, not transparent identification conditions.
  4. Draft lottery instrument: High lottery numbers attract volunteers who selected in based on high perceived private gains (violating A-2); high numbers also lead to more job training, potentially creating a direct ZYZ \to Y path (exclusion restriction violation).

The fundamental disagreement is normative: AIR prioritizes transparent, partially testable assumptions without a structural model; Heckman prefers structurally interpretable parameters (ATT, marginal treatment effect (MTE), full distributions) under explicit behavioral assumptions. Both are valid approaches; the AIR framework became the dominant pedagogical tradition.

Robins and Greenland (1996) Comment

A complementary comment in the same JASA issue (pp. 456–458), from epidemiologists rather than econometricians. Key points:

AIR (1996) Rejoinder: Defense and Clarifications

Key clarifications from AIR's response to Heckman, Robins & Greenland, Moffitt, and Rosenbaum:

  1. Compliers are the only identified subpopulation: Always-takers are always treated; never-takers are never treated. The data simply cannot be informative about average treatment effects for these groups. Estimating ATT or ATE in IV contexts requires extrapolation beyond what the data directly support.
  2. Heckman's A-2' implicitly assumes always-takers = compliers: Under random assignment, exclusion, and monotonicity, Heckman's assumption (A-2') implies E[Y(1)Y(0)always-taker]=E[Y(1)Y(0)complier]E[Y(1)-Y(0)|\text{always-taker}] = E[Y(1)-Y(0)|\text{complier}]. ATT identification therefore requires assuming volunteers and draftees have the same average treatment effect — a strong, unstated assumption.
  3. Mean independence vs. full independence: If mean independence holds but full independence fails, instrument validity becomes tied to the functional form of the regression (valid for Y but not log(Y)). This reintroduces functional-form-dependence — exactly the approach AIR sought to avoid.
  4. One-sided noncompliance: LATE = ATT: In randomized eligibility designs where D(0)=0D(0)=0 for all (controls cannot take treatment), there are no always-takers and monotonicity holds automatically. In this case LATE = ATT exactly. Heckman (1995) acknowledged this, writing that such designs can be placed in an IV framework identifying E[ΔD=1,X]E[\Delta|D=1,X] — which is LATE.
  5. Bounds equivalence: Under monotonicity, the unknown components of ATE are Y(0)Y(0) for always-takers and E[Y(1)]E[Y(1)] for never-takers. Letting these range freely over YY's support gives sharp bounds — equivalent to both the Balke-Pearl and Robins-Manski bounds.
  6. Rosenbaum's contributions: Extended the AIR sensitivity analysis to nonignorable instrument assignment; proposed Hodges-Lehmann robust estimators for IV; showed the exclusion restriction for non-compliers need only require Y(0,D(0))=Y(1,D(1))Y(0,D(0)) = Y(1,D(1)) for units with D(0)=D(1)D(0)=D(1), without defining counterfactual potential outcomes for always-takers and never-takers.

Structural Model (Moffitt 2005)

Moffitt (2005) formalizes the endogeneity problem as a two-equation structural system:

OLS on the outcome equation is biased when TiT_i is correlated with εi\varepsilon_i (unobservable confounders) or with βi\beta_i (selection on heterogeneous treatment gains — individuals sort into treatment partly because they privately know their βi\beta_i is high; see Essential Heterogeneity). IV with a valid ZiZ_i recovers the average βi\beta_i for "switchers" — Moffitt's term for AIR's "compliers."

The area fixed-effects model extends this to panel/repeated cross-section data by first-differencing:

where ΔZi\Delta Z_i is the change in the area-level variable (policy, price, availability). Area fixed effects cancel in differencing; ΔZi\Delta Z_i instruments for ΔTi\Delta T_i because individual changes in T may still be endogenous. This is not the individual fixed-effects model, which estimates only the first equation and assumes differencing alone eliminates bias — an assumption that "leaves unspecified why individual changes in T occur" (Moffitt 2005). See Natural Experiments.

Two Types of Extrapolation Failure (Moffitt 2005)

Any IV estimate faces two distinct limits on external validity:

  1. Mechanism specificity: Each instrument Z represents one specific cause of T's variation. Whether all causes of T have the same average effect βi\beta_i is an empirical question the IV framework cannot answer internally — "the effect of T" is ill-defined without specifying the mechanism. This is distinct from the LATE/ATE problem; it concerns whether βi\beta_i itself varies with the cause that moved TiT_i.

  2. Range restriction: Z moves T only across a particular range. Extrapolating to the rest of the population requires additional assumptions. Instruments inducing larger first-stage variation are preferred for extrapolation but typically have weaker internal validity claims. This formalizes why examiner-IV estimates (Maestas et al. 2013; French and Song 2014) — which identify effects for marginal applicants at 10\approx 10 percentage-point allowance differentials — cannot be directly extended to all DI applicants or to policy counterfactuals of different magnitudes.

Four Instrument Types (Moffitt 2003)

  1. Ecological / area variables: Policies, prices, unemployment rates, or other area-level characteristics that affect individual behavior but are plausibly exogenous to individual outcomes. Examples: cigarette taxes to smoking; local unemployment rate to DI applications (Bartik IV in China Trade Shock). Objections: residential sorting, time-varying unobservables, lagged adjustment, ecological fallacy.
  2. Demographic difference-in-differences (DiD): Comparing outcome changes over time between demographic groups differentially exposed to a policy. Requires parallel trends assumption — that groups would have evolved identically absent the policy.
  3. Twin / sibling designs: Within-family variation in treatment status. Identifies effects for families where siblings differ on the treatment. Requires that within-family differences are not themselves endogenous to outcomes.
  4. Natural experiments: Narrow, sudden, quasi-random variation in treatment — a legal cutoff, a lottery assignment, a weather shock. Examples in this wiki: administrative law judge (ALJ) lottery (French and Song 2014), Disability Determination Services (DDS) examiner assignment (Maestas, Mullen, Strand 2013), field office closings (Deshpande and Li 2019), Regression Kink Design at vocational grid age cutoffs.

Internal vs. External Validity Tradeoff

How It Is Used in This Wiki

Weak Instruments (Bound, Jaeger, Baker 1995; Staiger and Stock 1997)

An instrument is weak when its correlation with the endogenous variable is small in finite samples, causing the IV estimator's bias to approach that of OLS even when the instrument is theoretically valid. The problem was established by Bound, Jaeger, and Baker (1995), who showed that replacing the Angrist-Krueger (1991) quarter-of-birth instrument with a randomly drawn quarter of birth (which is valid but irrelevant) yields essentially the same IV estimate as the original — because the first stage's explanatory power was very low, the IV estimator converged toward OLS regardless.

Formula for IV bias (van der Klaauw 2014, drawing on Hahn and Hausman 2005):

BiasIV#instruments×ρ(U,V)×(1Rpartial2)#observations×Rpartial2\text{Bias}_{IV} \approx \frac{\#\text{instruments} \times \rho(U,V) \times (1 - R^2_{\text{partial}})}{\#\text{observations} \times R^2_{\text{partial}}}

IV/OLS bias ratio:

BiasIVBiasOLS#instruments#observations×Rpartial2\frac{\text{Bias}_{IV}}{\text{Bias}_{OLS}} \approx \frac{\#\text{instruments}}{\#\text{observations} \times R^2_{\text{partial}}}

where Rpartial2R^2_{\text{partial}} is the contribution of the instruments to the first-stage R2R^2.

F-test rule of thumb (Staiger and Stock 1997; Staiger and Yogo 2005): In the simplest case (one endogenous regressor, one instrument, no other regressors), 1/F1/F approximates the IV/OLS bias ratio. Instruments are considered weak if the IV bias exceeds 10% of the OLS bias, which is approximately equivalent to a first-stage F-statistic below 10. This is now the standard applied practice. With multiple instruments, the Cragg-Donald (CD) statistic is preferred over the F-statistic.

Remedies: (1) Use a stronger first stage. (2) Use limited information maximum likelihood (LIML) or Fuller's modified LIML, which are median-unbiased under weak instruments and do not share the IV many-instrument bias. (3) Report weak-instrument robust confidence intervals (Anderson-Rubin, conditional likelihood-ratio).

LIML vs. 2SLS: Bayesian Foundations (Kleibergen and Zivot 2003)

The choice between LIML and two-stage least squares (2SLS) has a Bayesian interpretation. Kleibergen and Zivot (2003) show that four diffuse-prior Bayesian approaches to the IV model map onto classical estimators as follows:

Bayesian approach Classical analogue Ordering invariant? Robust to spurious instruments?
Jeffreys prior on Restricted Reduced Form (RRF) LIML Yes Yes
Drèze (1976) flat prior on Structural Form (SF) 2SLS (closer) No No
Bayesian two-stage (B2S) 2SLS No No
Flat prior on Unrestricted Reduced Form (URF) LIML-like Yes Moderate

Key result: The Jeffreys prior — the prior that encodes maximum ignorance in a parameter-invariant way — yields a posterior for structural parameters β\beta that is functionally identical to the exact finite-sample density of the LIML estimator. This gives LIML a principled Bayesian justification: it is the estimator that emerges from coherent prior ignorance.

Practical implication: Under weak instruments or many instruments, LIML (or Fuller's modified LIML with slightly modified tails) is preferred over 2SLS. The 2SLS estimator's sensitivity to spurious instruments and ordering non-invariance are not bugs to be patched — they are symptoms of an incoherent prior.

MTE Unification and the P1/P2/P3 Hierarchy (Heckman 2008)

Heckman and Vytlacil (1999, 2005) show that all standard estimators (OLS, IV, matching, DiD) identify different weighted averages of the Marginal Treatment Effect (MTE) — the average treatment effect at each threshold of the propensity score. The MTE framework reveals that the Imbens-Angrist/Heckman debate is not about which estimator is "right" but about which weighted average of MTE each design identifies and whether that weighted average corresponds to a policy-relevant parameter.

Heckman (2008) organizes policy questions into a hierarchy:

Marschak's Maxim: Use the minimum model needed for the policy question. For P1, statistical treatment effects implement the maxim. For P2/P3, structural identification of policy-invariant parameters is necessary, which requires more assumptions but is the only valid approach.

Convergence as Validation

The convergence of the MMS (2013) examiner-IV estimate (28\approx 28 pp LFP reduction from DI receipt) and the French and Song (2014) ALJ-lottery estimate (26\approx 26 pp) across two different instruments identifying different complier populations is the strongest available evidence for the causal work-disincentive. Different instruments with different biases pointing in the same direction constitutes informal external validity — a Moffitt-style synthesis argument.

Open Questions

Related

Sources