An 8-year retrospective survey of the Lee-Carter (LC) model published in North American Actuarial Journal 4(1): 80–93. Lee reviews the original 1992 methodology, presents out-of-sample validation showing LC substantially outperformed the Social Security Administration (SSA) on life expectancy at birth (e0) gains 1989–1997, honestly catalogs six shortcomings, surveys extensions (sex disaggregation, cause disaggregation, jump-off fix, Wilmoth one-stage weighted singular value decomposition (SVD)), and discusses the Social Security finance application including the Congressional Budget Office's (CBO) adoption of the LC–Tuljapurkar stochastic framework. The paper does not introduce new empirical results; its contribution is synthesis, validation, and critical self-assessment.
Key Claims
Out-of-sample validation (1989–1997): LC forecast e0 gains of +1.30 years over this period; actual gain was +1.43 years. SSA's forecast over the same period implied only +0.73 years — roughly half the realized gain. LC's error was a small overshoot; SSA's was a large undershoot.
Jump-off problem: Standard LC forecasts from the fitted period average, causing a discontinuity at the jump-off year when recent rates deviate from the fitted ax. Recommended fix: initialize forecasts from the most recent observed rates rather than the fitted ax:
lnm(x,t+s)=lnm(x,t)+bx(kt+s−kt)
This eliminates the jump-off discontinuity and was adopted by Bell (1997) and others.
Second-stage k estimation: SVD minimizes squared log-rate errors, not life expectancy errors. Jensen's inequality causes the SVD kt to systematically misrepresent the observed e0. Lee and Carter (1992) correct this by re-estimating kt each year as the value that exactly matches observed e0 (or total deaths in some variants). Lee (2000) discusses variants that match total deaths rather than e0 as an alternative.
Six shortcomings of the LC model:
The linear decline in kt is an empirical regularity, not a structural result — no theoretical reason it must continue.
The bx profile (age-specific sensitivity) is assumed constant over the forecast horizon; historically it has changed, and there is no mechanism to prevent further change.
Confidence intervals (CIs) are too narrow (Alho 1990 critique): excluding the 1918 flu shock from kt's residual variance understates forecast uncertainty; correcting would widen CIs by ~57%.
Sex differentials in mortality are not directly modeled; male and female rates are forecast separately with different drift rates, causing male–female e0 gaps to diverge without bound — an implausible long-run behavior.
The single-factor structure cannot capture age-group divergences (e.g., young vs. old mortality moving in opposite directions).
Parameter uncertainty (in ax, bx, drift) is ignored; only kt forecast variance enters the CIs.
Extensions surveyed:
Sex disaggregation: Carter and Lee (1992) model male and female kt jointly using a bivariate autoregressive integrated moving average (ARIMA); Lee (2000) notes this improves long-run plausibility by constraining the sex gap.
Cause disaggregation: Modeling causes separately and summing always produces higher all-cause mortality than modeling all-cause directly, because cause-specific forecasts lose the negative cross-cause correlations (e.g., cardiovascular ↓ = cancer share ↑). Cause disaggregation is not recommended unless causes are the direct object of interest.
Latest-rates initialization (jump-off fix): Bell (1997) recommends using the most recent observed rates, not the fitted ax, as the jump-off — adopted in most subsequent implementations.
Wilmoth one-stage weighted SVD: Wilmoth (1993) proposes weighting the SVD by the square root of exposure (Ex,t), which downweights age-period cells with small populations (noisy rates). This is a one-stage estimator that simultaneously accounts for heteroskedasticity without requiring the two-stage SVD + re-estimation procedure.
Social Security finance application: Lee and Tuljapurkar (1994, 1998a, 1998b) applied the LC framework stochastically to project Social Security finances, generating distributions over Trust Fund solvency rather than point forecasts. CBO adopted this stochastic approach (Lee–Tuljapurkar 1996–1998 period). Under a 2 percentage point (pp) payroll tax increase scenario, the Trust Fund exhaustion date has a 75th percentile around 2070 in the Lee–Tuljapurkar distribution — the standard deterministic projection masks substantial uncertainty.
Discussant: Juha Alho (pp. 91–93): Confirms that the LC bilinear log model is an association model (a regression of log rates on latent index kt), not a structural model of mortality change. Recommends Poisson maximum likelihood (Poisson ML) as the proper fitting criterion given count-data structure; ordinary least squares (OLS) on log rates is suboptimal but the practical difference is modest; weighted SVD (Wilmoth 1993) resolves the heteroskedasticity issue. Alho applies LC to Finland as a demonstration.
"Although Lee-Carter confidence intervals are imprecise, the central forecast has performed well."
"Using the latest rates as the jump-off eliminates the discontinuity between historical and projected mortality and is now the standard recommendation."
"Disaggregating by cause of death always yields higher projected total mortality than modeling total mortality directly."
My Take
The validation finding is the most important result and deserves more weight than Lee gives it in the paper. LC missed by only +0.13 years over eight years while SSA missed by −0.70 years — a 5:1 accuracy ratio in favor of LC. The shortcomings catalog is refreshingly honest, particularly the sex-differential divergence problem and the parameter uncertainty exclusion. The cause-disaggregation warning is counterintuitive but correct and has direct policy relevance: agencies that model cause-specific mortality and sum will systematically overestimate total mortality, producing pessimistic life expectancy projections. The Alho discussant section independently confirms the two main technical critiques (OLS suboptimality, CI understatement) and provides Finnish validation — cross-country replication this early (2000) is notable. The gap between LC's CI narrowness and its point-forecast accuracy is the paper's central tension: it works better than it should given its known specification errors.