Definition
Instrumental variables (IV) is a method for estimating causal effects when the variable of interest is endogenous — that is, when individuals who differ on that variable also differ on unobserved characteristics that affect outcomes, making ordinary least squares (OLS) invalid. An instrument is a variable that (1) is correlated with the endogenous variable (relevance) and (2) affects outcomes only through the endogenous variable, not directly, and is uncorrelated with unobserved outcome determinants (exclusion restriction). When these conditions hold, observed variation in outcomes across instrument values must be caused by the endogenous variable — the instrument creates quasi-experimental variation that mimics a randomized assignment.
Key Ideas
- Potential outcomes framework: Every individual has two potential outcomes — Y1 (under treatment) and Y0 (without treatment). The causal effect β(i)=Y1−Y0 is unobservable because only one state is ever realized. Without additional assumptions, causal effects are not identified from observational data alone.
- Endogeneity: When the treatment variable D is correlated with unobservables (ε) in the outcome equation Y=α+βD+ε, OLS conflates the causal effect with selection bias. Typical examples: healthier people choose healthier behaviors; workers with higher ability select into higher-wage jobs; sicker workers apply to Disability Insurance (DI).
- Exclusion restriction: The instrument Z must affect D but have no direct path to Y and no correlation with ε. This assumption is untestable from the data alone; it must be justified on substantive grounds.
- LATE (Local Average Treatment Effect): An IV estimate identifies the average causal effect only for the subset of the population whose treatment status is actually changed by the instrument — the "compliers" or "switchers" (Imbens and Angrist 1994). If a price change moves the fraction of smokers from 30% to 40%, the IV estimate is the average causal effect for that specific 10% — not for the full population. Every IV result in this wiki refers to a LATE, not an average treatment effect (ATE).
- Essential heterogeneity (Heckman and Vytlacil; Cunha and Heckman 2007): When agents have private information about their own above-average returns and select into treatment accordingly, OLS overestimates the population ATE and IV can either exceed or fall below ATE depending on whether compliers have above- or below-average private returns. This is not a violation of instrument validity — it reflects rational selection on private information. Applied to schooling: students who attend college partly because they privately know their returns are high. An instrument like distance to college moves marginal students who may have lower private return signals, so IV ≤ ex ante ATE. The gap between ex ante returns (expected at decision time) and ex post returns (actually realized) measures the residual uncertainty agents face — Cunha and Heckman show roughly half of ex post return variance is unforeseeable at decision time. See Returns to Schooling.
- Reduced form vs. structural estimation: The reduced-form IV estimate (effect of Z on Y directly) is policy-relevant when the instrument is itself the policy lever. Structural estimation identifies the channel (effect of D on Y), which matters for generalization and mechanism testing.
The Angrist, Imbens, and Rubin (AIR) (1996) Framework: Five Assumptions and Compliance Types
Angrist, Imbens, and Rubin (1996) embedded IV in the Rubin Causal Model and showed that five assumptions are jointly sufficient for a causal interpretation of the IV estimand as LATE:
- Stable Unit Treatment Value Assumption (SUTVA): No interference between units; a single well-defined treatment version.
- Ignorable assignment: Z is randomly (or conditionally ignorably) assigned.
- Exclusion restriction: Z affects Y only through D — any effect of Z on Y must operate via D. Untestable; must be argued substantively. The paper separates this cleanly from ignorable assignment, which the old econometric formulation bundled together as "zero correlation between instrument and disturbances."
- Instrument relevance: E[D(1)−D(0)]=0 (the first stage is nonzero). Directly testable.
- Monotonicity: D(1)≥D(0) for all units — no defiers (no one takes treatment precisely when assigned to control). Implied by designs that deny treatment to controls.
Compliance typology — for any binary instrument, the population partitions into four types:
|
D(0)=0 |
D(0)=1 |
| D(1)=1 |
Complier — IV identifies ATE for this group |
Always-taker — exclusion: Z has no effect on Y |
| D(1)=0 |
Never-taker — exclusion: Z has no effect on Y |
Defier — ruled out by monotonicity |
Under Assumptions 1–5, IV=LATE=E[Y(1)−Y(0)∣complier]. The share of compliers equals the first-stage coefficient.
Sensitivity: Violations of the exclusion restriction bias IV by (noncomplier direct effect) × (noncomplier odds); violations of monotonicity bias IV by (defier share) × (complier-defier treatment effect gap). Both biases shrink as the instrument strengthens — weak instruments are fragile to both.
Heckman (1996) Critique of the AIR Framework
Heckman's comment in the same Journal of the American Statistical Association (JASA) issue raises four challenges:
- Priority: The potential outcomes framework is equivalent to the econometric switching regression model (Quandt 1958; Maddala and Nelson 1975) and Roy model (Heckman and Honoré 1990), which predate Rubin (1974). The novelty claim is overstated.
- Weaker conditions suffice for average treatment effect on the treated (ATT): The mean treatment effect on the treated — E[Y1−Y0∣D=1,X] — is identified under mean independence alone (A-1: Z⊥E[Y0∣X]; A-2: Z⊥E[D1−D0∣X]) without full distributional independence and without monotonicity. ATT targets an observable subpopulation (the treated); LATE targets an unobservable one (compliers). Heckman argues ATT is the more relevant evaluation parameter.
- Behavioral content of AIR assumptions: The independence conditions are not statistical — they impose Granger noncausality restrictions on the data-generating process. These are violated in the Roy model, competing risks, and Gronau-Heckman labor supply model whenever agents have private information about their own gains (essential heterogeneity). AIR assumptions are strong behavioral claims, not transparent identification conditions.
- Draft lottery instrument: High lottery numbers attract volunteers who selected in based on high perceived private gains (violating A-2); high numbers also lead to more job training, potentially creating a direct Z→Y path (exclusion restriction violation).
The fundamental disagreement is normative: AIR prioritizes transparent, partially testable assumptions without a structural model; Heckman prefers structurally interpretable parameters (ATT, marginal treatment effect (MTE), full distributions) under explicit behavioral assumptions. Both are valid approaches; the AIR framework became the dominant pedagogical tradition.
Robins and Greenland (1996) Comment
A complementary comment in the same JASA issue (pp. 456–458), from epidemiologists rather than econometricians. Key points:
- ATE over LATE: The average treatment effect in the full population — E[Y(1)−Y(0)] — is often the parameter of greatest public health interest, more so than LATE (compliers only) or intent-to-treat (ITT). Robins (1989) independently studied the same three AIR assumptions in the context of AIDS clinical trials.
- ATE is unidentified but boundable: Under AIR's assumptions alone, ATE is not point-identified — only LATE is. But the data places informative bounds on ATE (Robins-Manski bounds; sharpened by Balke and Pearl 1993). Wide bounds are useful: they make explicit how much conclusions depend on unverifiable prior beliefs.
- ATE is identifiable under SNMM without monotonicity: Robins' Assumption 6 — a structural nested mean model (SNMM) — implies ATE = IV estimand without requiring monotonicity and without assuming constant treatment effects. It requires only that within-arm average treatment effects are equal for treated and untreated subjects.
- Bioequivalence trials — LATE and ITT both fail: When comparing two active treatments, noncompliance means the sharp null of bioequivalence (Yi(1)=Yi(0) for all i) does not imply ITT=0. The ITT test may reject purely due to differential noncompliance rates. Neither LATE nor ITT is the right parameter; ATE requires methods (G-computation, inverse probability of censoring weighted (IPCW)) beyond the AIR framework.
AIR (1996) Rejoinder: Defense and Clarifications
Key clarifications from AIR's response to Heckman, Robins & Greenland, Moffitt, and Rosenbaum:
- Compliers are the only identified subpopulation: Always-takers are always treated; never-takers are never treated. The data simply cannot be informative about average treatment effects for these groups. Estimating ATT or ATE in IV contexts requires extrapolation beyond what the data directly support.
- Heckman's A-2' implicitly assumes always-takers = compliers: Under random assignment, exclusion, and monotonicity, Heckman's assumption (A-2') implies E[Y(1)−Y(0)∣always-taker]=E[Y(1)−Y(0)∣complier]. ATT identification therefore requires assuming volunteers and draftees have the same average treatment effect — a strong, unstated assumption.
- Mean independence vs. full independence: If mean independence holds but full independence fails, instrument validity becomes tied to the functional form of the regression (valid for Y but not log(Y)). This reintroduces functional-form-dependence — exactly the approach AIR sought to avoid.
- One-sided noncompliance: LATE = ATT: In randomized eligibility designs where D(0)=0 for all (controls cannot take treatment), there are no always-takers and monotonicity holds automatically. In this case LATE = ATT exactly. Heckman (1995) acknowledged this, writing that such designs can be placed in an IV framework identifying E[Δ∣D=1,X] — which is LATE.
- Bounds equivalence: Under monotonicity, the unknown components of ATE are Y(0) for always-takers and E[Y(1)] for never-takers. Letting these range freely over Y's support gives sharp bounds — equivalent to both the Balke-Pearl and Robins-Manski bounds.
- Rosenbaum's contributions: Extended the AIR sensitivity analysis to nonignorable instrument assignment; proposed Hodges-Lehmann robust estimators for IV; showed the exclusion restriction for non-compliers need only require Y(0,D(0))=Y(1,D(1)) for units with D(0)=D(1), without defining counterfactual potential outcomes for always-takers and never-takers.
Structural Model (Moffitt 2005)
Moffitt (2005) formalizes the endogeneity problem as a two-equation structural system:
- Yi=α+βiTi+γXi+εi (outcome equation)
- Ti=δ+θXi+ϕZi+υi (selection equation)
OLS on the outcome equation is biased when Ti is correlated with εi (unobservable confounders) or with βi (selection on heterogeneous treatment gains — individuals sort into treatment partly because they privately know their βi is high; see Essential Heterogeneity). IV with a valid Zi recovers the average βi for "switchers" — Moffitt's term for AIR's "compliers."
The area fixed-effects model extends this to panel/repeated cross-section data by first-differencing:
- ΔYi=α+βiΔTi+γΔXi+εi
- ΔTi=δ+θΔXi+ϕΔZi+υi
where ΔZi is the change in the area-level variable (policy, price, availability). Area fixed effects cancel in differencing; ΔZi instruments for ΔTi because individual changes in T may still be endogenous. This is not the individual fixed-effects model, which estimates only the first equation and assumes differencing alone eliminates bias — an assumption that "leaves unspecified why individual changes in T occur" (Moffitt 2005). See Natural Experiments.
Two Types of Extrapolation Failure (Moffitt 2005)
Any IV estimate faces two distinct limits on external validity:
Mechanism specificity: Each instrument Z represents one specific cause of T's variation. Whether all causes of T have the same average effect βi is an empirical question the IV framework cannot answer internally — "the effect of T" is ill-defined without specifying the mechanism. This is distinct from the LATE/ATE problem; it concerns whether βi itself varies with the cause that moved Ti.
Range restriction: Z moves T only across a particular range. Extrapolating to the rest of the population requires additional assumptions. Instruments inducing larger first-stage variation are preferred for extrapolation but typically have weaker internal validity claims. This formalizes why examiner-IV estimates (Maestas et al. 2013; French and Song 2014) — which identify effects for marginal applicants at ≈10 percentage-point allowance differentials — cannot be directly extended to all DI applicants or to policy counterfactuals of different magnitudes.
Four Instrument Types (Moffitt 2003)
- Ecological / area variables: Policies, prices, unemployment rates, or other area-level characteristics that affect individual behavior but are plausibly exogenous to individual outcomes. Examples: cigarette taxes to smoking; local unemployment rate to DI applications (Bartik IV in China Trade Shock). Objections: residential sorting, time-varying unobservables, lagged adjustment, ecological fallacy.
- Demographic difference-in-differences (DiD): Comparing outcome changes over time between demographic groups differentially exposed to a policy. Requires parallel trends assumption — that groups would have evolved identically absent the policy.
- Twin / sibling designs: Within-family variation in treatment status. Identifies effects for families where siblings differ on the treatment. Requires that within-family differences are not themselves endogenous to outcomes.
- Natural experiments: Narrow, sudden, quasi-random variation in treatment — a legal cutoff, a lottery assignment, a weather shock. Examples in this wiki: administrative law judge (ALJ) lottery (French and Song 2014), Disability Determination Services (DDS) examiner assignment (Maestas, Mullen, Strand 2013), field office closings (Deshpande and Li 2019), Regression Kink Design at vocational grid age cutoffs.
Internal vs. External Validity Tradeoff
- Internal validity: the instrument is genuinely exogenous, producing an unbiased estimate for the complier population. Natural experiments score highest.
- External validity: the estimate generalizes to the broader population or to other policy contexts. Population-level studies score highest; natural experiments often score lowest.
- The tradeoff is inherent: maximizing internal validity by using very narrow instruments (a lottery in two states for 4–6-year-olds) produces estimates that may not generalize to any broader population or policy. A collection of internally valid but externally narrow studies does not aggregate into general knowledge.
- The solution (Moffitt 2003) is synthesis: weight evidence from multiple studies with different IV types and different complier populations; look for convergence; use formal theory to extrapolate across contexts.
How It Is Used in This Wiki
- Examiner / judge lottery IV: DDS examiners are quasi-randomly assigned to DI applications. Examiner allowance rate variation instruments for actual DI receipt. Identifies a LATE for marginal applicants near the award threshold — ≈23% of applicants (Maestas, Mullen, Strand 2013). Causal labor force participation (LFP) reduction: ≈28 pp. French and Song (2014) use ALJ lottery at the appeals stage.
- Bartik IV / shift-share: Local industry employment shocks instrument for local labor demand faced by specific demographic groups. Used in China Trade Shock (Autor, Dorn, Hanson 2013/2019) to isolate the manufacturing decline effect on DI participation and family structure.
- Regression Kink Design (RKD): A kink (slope change) in an assignment rule instruments for benefit levels without a discontinuous jump. See Regression Kink Design for the DI income-mortality application at Average Indexed Monthly Earnings (AIME) bend points.
- Field office closings: Office closings instrument for application cost variation (Deshpande and Li 2019). Natural experiment demonstrating that cost increases deter approved-quality applicants. See DI Application Costs and Take-Up.
- Vocational grid age cutoffs: Discrete age thresholds (50, 55) at which DI eligibility rules change create sharp variation in award probability used as natural experiments. See Vocational Grid.
Weak Instruments (Bound, Jaeger, Baker 1995; Staiger and Stock 1997)
An instrument is weak when its correlation with the endogenous variable is small in finite samples, causing the IV estimator's bias to approach that of OLS even when the instrument is theoretically valid. The problem was established by Bound, Jaeger, and Baker (1995), who showed that replacing the Angrist-Krueger (1991) quarter-of-birth instrument with a randomly drawn quarter of birth (which is valid but irrelevant) yields essentially the same IV estimate as the original — because the first stage's explanatory power was very low, the IV estimator converged toward OLS regardless.
Formula for IV bias (van der Klaauw 2014, drawing on Hahn and Hausman 2005):
BiasIV≈#observations×Rpartial2#instruments×ρ(U,V)×(1−Rpartial2)
IV/OLS bias ratio:
BiasOLSBiasIV≈#observations×Rpartial2#instruments
where Rpartial2 is the contribution of the instruments to the first-stage R2.
F-test rule of thumb (Staiger and Stock 1997; Staiger and Yogo 2005): In the simplest case (one endogenous regressor, one instrument, no other regressors), 1/F approximates the IV/OLS bias ratio. Instruments are considered weak if the IV bias exceeds 10% of the OLS bias, which is approximately equivalent to a first-stage F-statistic below 10. This is now the standard applied practice. With multiple instruments, the Cragg-Donald (CD) statistic is preferred over the F-statistic.
Remedies: (1) Use a stronger first stage. (2) Use limited information maximum likelihood (LIML) or Fuller's modified LIML, which are median-unbiased under weak instruments and do not share the IV many-instrument bias. (3) Report weak-instrument robust confidence intervals (Anderson-Rubin, conditional likelihood-ratio).
LIML vs. 2SLS: Bayesian Foundations (Kleibergen and Zivot 2003)
The choice between LIML and two-stage least squares (2SLS) has a Bayesian interpretation. Kleibergen and Zivot (2003) show that four diffuse-prior Bayesian approaches to the IV model map onto classical estimators as follows:
| Bayesian approach |
Classical analogue |
Ordering invariant? |
Robust to spurious instruments? |
| Jeffreys prior on Restricted Reduced Form (RRF) |
LIML |
Yes |
Yes |
| Drèze (1976) flat prior on Structural Form (SF) |
2SLS (closer) |
No |
No |
| Bayesian two-stage (B2S) |
2SLS |
No |
No |
| Flat prior on Unrestricted Reduced Form (URF) |
LIML-like |
Yes |
Moderate |
Key result: The Jeffreys prior — the prior that encodes maximum ignorance in a parameter-invariant way — yields a posterior for structural parameters β that is functionally identical to the exact finite-sample density of the LIML estimator. This gives LIML a principled Bayesian justification: it is the estimator that emerges from coherent prior ignorance.
Practical implication: Under weak instruments or many instruments, LIML (or Fuller's modified LIML with slightly modified tails) is preferred over 2SLS. The 2SLS estimator's sensitivity to spurious instruments and ordering non-invariance are not bugs to be patched — they are symptoms of an incoherent prior.
MTE Unification and the P1/P2/P3 Hierarchy (Heckman 2008)
Heckman and Vytlacil (1999, 2005) show that all standard estimators (OLS, IV, matching, DiD) identify different weighted averages of the Marginal Treatment Effect (MTE) — the average treatment effect at each threshold of the propensity score. The MTE framework reveals that the Imbens-Angrist/Heckman debate is not about which estimator is "right" but about which weighted average of MTE each design identifies and whether that weighted average corresponds to a policy-relevant parameter.
Heckman (2008) organizes policy questions into a hierarchy:
- P1 (historical program evaluation): Evaluating whether a past program worked for the treated population. Statistical treatment-effects methods (LATE, ATT, matching) are adequate if the identifying instrument or comparison group mimics the program's assignment mechanism.
- P2 (new environments): Forecasting the effects of a known program in a new setting (different population, different macroeconomic conditions). Requires parameters that are invariant to the environmental change — a structural model component not recovered by reduced-form IV alone.
- P3 (new policies): Evaluating policies never experienced in any environment. Requires full structural identification of preferences, technology, and constraints — no extrapolation from treatment-effect estimates is valid without explicit invariance assumptions.
Marschak's Maxim: Use the minimum model needed for the policy question. For P1, statistical treatment effects implement the maxim. For P2/P3, structural identification of policy-invariant parameters is necessary, which requires more assumptions but is the only valid approach.
Convergence as Validation
The convergence of the MMS (2013) examiner-IV estimate (≈28 pp LFP reduction from DI receipt) and the French and Song (2014) ALJ-lottery estimate (≈26 pp) across two different instruments identifying different complier populations is the strongest available evidence for the causal work-disincentive. Different instruments with different biases pointing in the same direction constitutes informal external validity — a Moffitt-style synthesis argument.
Open Questions
- For any given IV result, who exactly are the compliers? How different are they from the full population of interest?
- How much do LATE estimates from examiner-IV (marginal DDS applicants, ≈23% of pool) generalize to the effects on the full DI population?
- The parallel trends assumption in DiD (including geographic DiD Bartik designs) is untestable — how should one assess plausibility?
Related
Sources