Definition
The return to schooling is the percentage increase in earnings from one additional year of education, the central parameter of the human capital model of educational investment (Becker 1964; Mincer 1974). In the standard Mincer equation, ln(w)=α+ρS+βX+γX2, the coefficient ρ≈0.07–0.10 in U.S. data — roughly a 7–10% wage premium per year of schooling. Cunha and Heckman (2007) introduce a fundamental refinement: because earnings are uncertain when the schooling decision is made, there are two distinct objects — the ex ante return (expected return at decision time) and the ex post return (realized return) — and they differ substantially.
Key Ideas
- Roy comparative advantage model (Willis and Rosen 1979): Educational choice is sector selection in a two-sector Roy economy. Individuals have two potential earnings paths — w1 (college) and w0 (non-college) — and attend college when expected lifetime return net of costs is positive. Selection is on the person-specific gain (w1−w0−c), not on absolute ability or earnings level alone. The critical empirical finding: college-goers are negatively selected on the no-college counterfactual (E[w0∣College]<E[w0∣No College]) — they would have earned below-average wages had they not attended college, reflecting complementarity between college-type human capital and college-sector earnings. This means the college and non-college populations sort on comparative advantage, not absolute advantage, and that ordinary least squares (OLS) comparison of mean wages across groups is an unreliable guide to causal returns for either group. See Sherwin Rosen, Robert J. Willis.
- OLS endogeneity: Higher-ability workers both earn more and choose more schooling. If ability is unobserved, OLS conflates the causal return with ability selection, typically biasing estimates upward ("ability bias"). However, credit constraints push low-ability/low-income students out of school, creating downward bias. OLS estimates in the range of 7–10% reflect these offsetting forces.
- OLS bias decomposition (Card 1999): plim(b^OLS)=βˉ+λ0+ϕ0Sˉ, where λ0 is the intercept-ability bias (correlation between schooling and the level of the earnings function) and ϕ0Sˉ is the slope-ability bias (correlation between schooling and the returns slope). Both components are positive. With measurement error attenuation R0≈0.90 partly offsetting, the net upward OLS bias is approximately 10% of the OLS coefficient.
- Instrumental Variable (IV) and Local Average Treatment Effect (LATE): Instruments for schooling (distance to college, compulsory schooling laws, tuition subsidies) identify a LATE — the return for the specific marginal students induced into more schooling by the instrument. Under Wooldridge (1997), IV estimates E[βg⋅ΔSg]/E[ΔSg], a weighted average of subgroup marginal returns. Because different instruments move different subgroups, different IVs yield different estimates without contradicting each other.
- IV > OLS puzzle and its resolution (Card 1999): IV estimates based on institutional features of the school system are systematically 20–40% above OLS. Card's preferred explanation: instruments like compulsory schooling or college proximity disproportionately affect disadvantaged subgroups who face higher marginal returns due to financial constraints (higher discount rates), not higher ability. If marginal returns decline with educational attainment and ability differences are not the dominant source of within-group schooling variation, then IV for less-educated compliers > average OLS.
- Family background as IV is upward-biased: Using parental or sibling education as an instrument typically yields estimates above OLS, not below. Family background correlates with own earnings independently of schooling (ability channel), violating the exclusion restriction. IV based on family background likely has more upward bias than OLS (Card 1999, Table 5).
- Measurement error compounds within-family estimates: Reliability of within-family difference in schooling: RA=R0(1−ρ)/(1−ρR0). For identical twins (R0≈0.9, schooling correlation ρ≈0.75), RA≈0.7 — a 30% attenuation bias. Naïve within-family estimates are substantially downward-biased by measurement error amplification, so within-family < OLS can reflect noise rather than ability bias elimination.
- Twins evidence: ability bias ≈10% (Ashenfelter and Rouse 1998; Rouse 1997): Measurement-error-corrected within-family IV for identical twins (Princeton Twins Survey) approximately equals the OLS estimate, implying that OLS ability bias is only ≈10%. This is the most credible available upper bound on OLS ability bias.
- IV < OLS: the GI Bill case (Angrist and Chen 2011): One of the rare well-documented cases where IV returns to schooling are below OLS (≈0.07 vs. ≈0.12). The Vietnam-era GI Bill provided a large college subsidy to veterans, inducing enrollment among men who otherwise would not have attended. These marginal veterans may have had below-average returns (consistent with selection into schooling by those with high private returns under normal conditions), or nonlinearities mean the college-specific return differs from the average return across education levels. The GI Bill setting is unusually clean because the subsidy mechanism is explicit and the instrument (draft lottery) directly changed the cost of schooling.
- Essential heterogeneity (Heckman and Vytlacil): When agents have private information about their own above-average returns and select into schooling accordingly, OLS overestimates the average treatment effect (ATE) for those not attending school, and standard IV estimates can exceed the ex ante ATE. The direction of IV bias relative to ATE depends on whether compliers have above- or below-average private returns. This is not a violation of IV assumptions — it is a consequence of rational selection under private information.
- Ex ante vs. ex post returns (Cunha and Heckman 2007): The schooling decision is made before earnings are realized. The ex ante return is what the agent expects given information available at decision time. The ex post return is the actual realized wage premium. Cunha and Heckman show that approximately half of ex post return variability was not predictable at decision time — agents face genuine earnings risk.
- Factor model identification: Cunha and Heckman use observable proxies for cognitive and non-cognitive skills to map the agent's information set at decision time. Variation in these proxies separates the ex ante signal from the ex post noise, allowing separate identification of the two distributions without assuming full information.
- College wage premium: The return to a college degree (relative to high school) doubled from roughly 50% to nearly 100% between 1979 and 2009 in the U.S. This rise is the primary explanation for growing earnings inequality in the Routine-Biased Technological Change literature.
How It Works
The Mincer equation is estimated on earnings and schooling data. OLS gives ρ^OLS, which conflates the causal return with selection. An instrument Z (e.g., distance to nearest 4-year college, tuition costs) that shifts schooling but has no direct effect on earnings identifies ρ^IV=Cov(Y,Z)/Cov(S,Z). Under homogeneous returns, IV=OLS=ATE. Under heterogeneous returns, IV=LATEZ, the average causal return for compliers — those whose schooling status was changed by Z. If compliers tend to have lower private information about their returns (because they needed the nudge of proximity to college), IV≤ATE. If compliers faced binding credit constraints that underinvestment corrected, IV>ATE.
Cunha and Heckman additionally use a factor model: let θ be unobserved skill (cognitive + non-cognitive). Observable proxies (test scores, grades) provide noisy signals of θ. Under the factor structure, one can separate θante (the signal available at decision time) from the component of θ only revealed ex post. This decomposition identifies the full distribution of ex ante returns and the distribution of ex post returns separately.
Why It Matters
Wage Inequality
Rising returns to schooling are the dominant explanation for U.S. earnings inequality growth since 1980 (alongside Routine-Biased Technological Change and Minimum Wage and Wage Inequality). The college wage premium doubling over 1979–2009 mechanically widens the earnings distribution even absent any change in the schooling distribution.
Education Policy
If ex ante and ex post returns are close, the education market is approximately efficient — students respond correctly to expected returns, and marginal interventions (e.g., information campaigns, small tuition subsidies) will have small effects. If they diverge substantially (as Cunha and Heckman show), there is scope for:
- Information provision: Providing students with better signals of their expected returns before the decision
- Consumption insurance: Allowing ex post smoothing of earnings risk, which is part of the case for income-contingent student loan repayment
Connection to DI and Labor Economics
The ex ante/ex post distinction applies broadly to any investment decision under uncertainty — not just schooling. The disability insurance (DI) context involves a similar structure: workers do not know at hire (or at skill investment) whether they will become disabled, and the ex ante expected earnings trajectory diverges from the ex post realized path when disability occurs. This is the structure underlying Match Quality and Marital Dissolution and the earnings-surprise mechanism in Disability and Marital Dissolution.
Open Questions
- How much does the ex ante/ex post gap vary by demographic group? If low-income students face greater earnings uncertainty (less family-network labor market information), the gap may be larger for them, amplifying the case for targeted information provision.
- How has the ex ante/ex post gap changed as labor markets became more volatile? Rising within-group earnings variance (Geweke and Keane 2000; Earnings Dynamics) may have widened the gap over time.
- Does essential heterogeneity explain the persistent IV-vs-OLS gap in cross-country returns-to-education estimates?
Related
Sources