Carter and Lee (1986): The earliest direct ancestor of the LC architecture. Carter and Lee define a scalar marital fertility index g(t) and nuptiality index h(t), each convolved with a fixed demographic schedule to generate demographic event counts — the same "scalar index × fixed age schedule" structure that the Lee-Carter (LC) model applies to mortality six years later. The main difference is that LC uses singular value decomposition (SVD) to extract the scalar index k(t) from mortality data rather than constructing it as a ratio of observed to expected events. Carter and Lee also establish the autoregressive integrated moving average (ARIMA) framework for the scalar index and document the rapid growth of forecast uncertainty — themes that recur throughout the Lee-Carter literature. See Carter and Lee 1986 — Joint Forecasts of US Marital Fertility, Nuptiality, Births, and Marriages and Stochastic Fertility Forecasting.
Thompson et al. (1989): Thompson, Bell, Long, and Miller (1989) applied an identical dimensionality-reduction template to fertility: they fit a shifted gamma density to the U.S. age-specific fertility schedule each year, reducing 32 age-specific rates to three parameters (total fertility rate (TFR), mean age at childbearing (MACB), standard deviation of age at childbearing (SDACB)), then modeled those parameters jointly with a multivariate ARIMA. Lee and Carter (1992) cite this work explicitly and apply the same logic to mortality — compressing the age schedule to a single scalar index kt and modeling that index as a random walk. The key simplification LC makes relative to Thompson et al. is collapsing three fertility parameters to one mortality index, justified by the empirical finding (92.5% of variance explained by the first singular value pair) that mortality moves largely as a single factor.
Definition
The Lee-Carter (LC) model is a bilinear statistical model for forecasting age-specific mortality rates, introduced by Ronald Lee and Lawrence Carter in 1992. The model decomposes the log of the central death rate at age x in year t as:
lnm(x,t)=ax+bxkt+εx,t
where ax is the time-average log death rate at age x, bx is each age group's sensitivity to the overall mortality trend, kt is a latent scalar mortality index, and εx,t is an age- and year-specific error term. Forecasting mortality reduces to forecasting the single time series kt, which is modeled as an ARIMA(0,1,0) random walk with drift.
Key Ideas
Single-factor structure: All age-group mortality moves together via one common index kt. Cross-age correlation is captured entirely through the bx loadings; no residual cross-age dynamics are modeled.
bx profile: Empirically, bx is highest at infant and child ages (≈0.09–0.11) and lowest at oldest ages (≈0.029–0.033 at ages 85+), reflecting faster historical proportional improvement at young ages. As the bx pattern tilted toward older ages over the 20th century, e0 growth remained linear despite diminishing returns — the bx tilt mechanism. See Period Mortality. Rau et al. (2006) provide direct empirical corroboration of this age gradient in improvement rates: using Kannisto-Thatcher Data Bank (KTDB) data for 27 countries through 2000, they document that annual improvement rates are highest at ages 80–84 and decline monotonically toward 100+. This mirrors the bx profile and confirms it reflects a cross-national structural regularity — not a US-specific artifact or an approaching biological ceiling, since improvement at every age group rose monotonically across successive decades.
ax profile: The time-average log mortality profile follows Gompertz regularities — an elongated V-shape with minimum around ages 10–14 and near-linear increase thereafter.
Equivalence to independent rate extrapolation (McNown 1992): The LC model is algebraically equivalent to projecting each age-specific death rate at its own historical exponential trend rate independently. The bilinear structure is a compact representation of those independent extrapolations, not an additional identifying restriction.
How It Works
Step 1 — SVD estimation:
Compute the age-time matrix of demeaned log death rates, lnm(x,t)−aˉx. Take the singular value decomposition; the leading singular vectors define bx (right) and the unnormalized kt (left). The first singular value pair typically accounts for >90% of total log-rate variance (92.5% in the original U.S. fit over 1933–1987).
Step 2 — kt re-estimation:
SVD minimizes squared log-rate errors, not life expectancy errors. Jensen's inequality implies that the SVD kt does not reproduce observed e0 values. Lee and Carter re-estimate kt each year as the value that makes the model's implied life expectancy exactly match the observed e0, correcting this bias.
Step 3 — ARIMA(0,1,0) for kt:
Model the re-estimated kt as a random walk with drift:
kt=kt−1+d+et,et∼N(0,σ2)
where d is estimated from the data (approximately −0.365 per year in the original U.S. fit). A pandemic dummy (e.g., 1918 flu) may be included to absorb structural breaks without contaminating the drift estimate.
Step 4 — Forecasting:
Project kt forward using the ARIMA model. The uncertainty in kt at horizon h grows as σh, translating into an expanding confidence interval (CI) for e0. In the original 1992 fit, the 95% CI for e0 in 2065 (75-year horizon) had half-width 5.6 years.
State-Space Reformulation (Pedroza and King 2002)
Pedroza and King show that LC can be expressed as a unified state-space model:
LC as ordinary least squares (OLS): The SVD step is algebraically equivalent to maximum likelihood estimation (MLE) under homoskedastic Gaussian errors. The bilinear structure is not a special estimation trick — it is a statistical model.
Incorrect LC forecasts: In state-space form, the correct one-step forecast conditions on the last residual ϵaT. Standard LC ignores it. The bias in point predictions is small (the per-age LC drift δa=βaθ nearly equals the per-age random-walk drift ϕa, explaining why LC ≈ fitting a separate random walk to each age group).
Bayesian estimation: Gibbs sampling over all parameters jointly yields posteriors for αa, βa, κt, θ, σ2, and τ2 with no extra programming effort, plus automatic imputation of missing rates and posterior predictive checks.
Even the corrected Bayesian state-space version fails to forecast well out-of-sample (U.S. males, U.S. females, Japanese males 1991–2000), particularly at ages 15–30 where mortality deviated from historical trend. The failure is structural — it traces to the constant αa, βa assumption — not to the frequentist vs. Bayesian estimation choice. This motivates approaches that allow time-varying age sensitivity (e.g., Meseguer's Bayesian vector autoregression (BVAR)).
Out-of-Sample Validation and Jump-Off Fix
Validation (1989–1997, Lee 2000): In the 8 years after the 1992 publication, LC forecast e0 gains of +1.30 years; the actual gain was +1.43 years. Social Security Administration (SSA)'s forecasts over the same period implied only +0.73 years — roughly half the realized gain. LC's error was a small overshoot; SSA's was a large undershoot. This is the primary empirical case for LC over the Office of the Chief Actuary (OACT) judgmental approach.
Jump-off fix: Standard LC forecasts initialize from the fitted ax (the time-averaged log mortality profile), causing a discontinuity at the jump-off year when recent observed rates deviate from the fit. The recommended fix (Bell 1997; Lee 2000) initializes from the most recent observed rates:
lnm(x,t+s)=lnm(x,t)+bx(kt+s−kt)
This eliminates the jump-off discontinuity without changing the drift structure and is now the standard implementation.
Multi-country validation of jump-off bias (Booth, Tickle, and Smith 2005): A 10-country out-of-sample evaluation (1986–1996 forecast period) formally decomposed LC error into jump-off and model components across Australia, Canada, Denmark, England & Wales, Finland, France, Italy, Norway, Sweden, and Switzerland. Jump-off bias accounts for 57% of female and 42% of male total LC forecast error; for Booth-Maindonald-Smith (BMS) the corresponding shares fall to 11% and 4%. Both Lee-Miller (LM) and BMS avoid the majority of jump-off bias by initializing from observed rather than fitted rates at the jump-off year. BMS is marginally superior to LM in a majority of countries (8/10 females, 6/10 males by root mean square error (RMSE) of e0). See Booth Tickle and Smith 2005 — Evaluation of the Variants of the Lee-Carter Method of Forecasting Mortality A Multi-Country Comparison.
Cause disaggregation warning: Modeling cause-specific mortality and summing always produces higher all-cause projected mortality than modeling all-cause directly, because separate forecasts lose the negative cross-cause correlations. Wilmoth (1995) showed this is a mathematical result under standard assumptions: decomposed projections are always more pessimistic than aggregate projections. Agencies that disaggregate by cause will systematically overestimate total mortality and underestimate projected life expectancy.
International Generalization (Tuljapurkar, Li and Boe 2000)
Tuljapurkar, Li, and Boe applied the LC SVD decomposition to all seven Group of Seven (G7) countries (Canada, France, Germany, Italy, Japan, UK, US) using 1950–1994 data, establishing that the single-factor structure is not a US artifact. In every country the first singular value explains >94% of the temporal variance in log death rates (range 94.3%–97.5%), and k(t) declines linearly in all seven. This cross-national validation — across very different social, historical, and epidemiological contexts — is the strongest evidence that the LC bilinear structure captures a genuine empirical regularity of modern mortality change.
Key extensions and findings from Tuljapurkar, Li, and Boe (TLB):
Death-adjusted k^(t): TLB recompute k^(t) each year to match historical total deaths (analogous to LC's life-expectancy re-estimation), then model k^(t+1)=k^(t)−z+εt. Drift z ranges from 0.26/yr (US) to 0.79/yr (Japan).
Stochastic medians exceed official forecasts everywhere: In 2050, stochastic median e0 exceeds official central projections by 1.3 yr (UK) to 8.0 yr (Japan); US gap is 2.5 yr (82.9 vs. 80.5).
b(x) profiles cross-nationally: Broadly similar across G7 — all higher at younger ages, lower at old ages — but with social/historical variation. Japan has the highest drift z (fastest improvement) and the most accurate old-age data (no lumping at 85+).
Dependency ratio: Each 1-yr difference in e0 implies >5% difference in the dependency ratio (65+/20–64). TLB stochastic medians imply dependency ratios 6% (UK) to 40% (Japan) higher than official projections by 2050, with divergence growing beyond 2050 as all official scenarios assume eventual deceleration. See Long-Term Actuarial Balance.
LC Variants (Booth 2006 Taxonomy)
By 2005, four main variants of the LC model had been proposed and evaluated (Booth 2006):
Li and Lee (2005) — coherent multi-population: All populations in a related group share a common time trend (estimated jointly) with population-specific deviations that are mean-reverting. Prevents long-run divergence between related populations without the cointegration framework required by vector error correction model (VECM) approaches. Distinct from the Zhou et al. (2014) VECM approach in that coherence is imposed via a common factor structure rather than an error-correction term.
Hyndman and Ullah (2007) — functional data: Extends LC to multiple latent factors via principal component analysis of smooth functional representations of mortality curves. Addresses the single-factor limitation (shortcoming #5 below) by allowing young and old ages to evolve with different factor loadings. The functional data framework also handles missing ages and provides a principled smoothing step. Implemented in the demography R package (Hyndman 2006). See Rob J. Hyndman.
De Jong and Tickle (2006) — state space LC(smooth): Reformulates the LC model within a B-spline constrained state-space framework estimated via Kalman filtering; the design matrix X imposes smooth quadratic age profiles across age groups. Provides exact forecast error calculations and extends naturally to multi-population and multi-component modeling. See Piet de Jong, Leonie Tickle.
Multi-Population Extension (Zhou et al. 2014)
The single-population LC model cannot handle the joint modeling of two related populations without additional structure. Zhou et al. propose three approaches for the joint evolution of the two time-varying factors κ(1) and κ(2), all enforcing long-run non-divergence:
RWAR (Cairns et al. 2011 baseline): κ(1) as random walk, (κ(1)−κ(2)) as AR(1). Asymmetric — requires declaring one population "dominant." Empirically unstable: swapping the dominant population changes 50-year forecasts by ≈17%.
vector autoregression (VAR): Model (Δκ(1),Δκ(2)) jointly; no dominant-population assumption; captures cross-correlations. Non-divergence enforced as a parameter equality constraint.
VECM (preferred): Adds an error-correction term — the lagged disequilibrium κt−1(1)−κt−1(2) — to the VAR. Grounded in Engle-Granger cointegration: if κ(1) and κ(2) are cointegrated, non-divergence holds automatically. Wins on Bayesian information criterion (BIC) (174.6 vs. 176.4 vs. 179.2), residual adequacy, in-sample coverage, out-of-sample tracking, and robustness across calibration windows.
Key finding: in the England-Wales / UK insured-lives pair, the smaller population (260K exposures) leads the larger (4.7M), not vice versa — the standard dominance assumption is reversed. See Multi-Population Mortality Modeling.
Why It Matters
Standard reference model: The LC model is the most widely cited and implemented mortality forecasting method, used by SSA, the U.S. Census Bureau, and hundreds of national statistical agencies internationally.
Uncertainty quantification: Unlike OACT's judgmental approach (which produces no formal confidence intervals), LC generates explicit probability intervals for e0 and age-specific rates — enabling stochastic solvency analysis. The Congressional Budget Office (CBO) adopted the Lee-Tuljapurkar stochastic framework for long-run federal budget projections in 1996.
Social Security finance application: Lee and Tuljapurkar (1998) embedded LC mortality in the first fully probabilistic Old-Age, Survivors, and Disability Insurance (OASDI) finance model — 750 sample paths combining stochastic fertility and AR(1) productivity/interest rates. They found mean long-term actuarial balance (LTAB) =−3.3 pp vs. SSA's −2.2 pp, and that fertility — not mortality — is the dominant source of 75-year LTAB uncertainty, reversing SSA's own ranking. See Long-Term Actuarial Balance.
Empirical legacy: The original 1992 forecast of e0=86.05 in 2065 exceeded SSA's 1989 intermediate projection (80.45) by 5.6 years, with SSA's figure falling at the lower bound of the LC interval. Subsequent research (Lee and Miller 2001) confirmed that LC's trend-rate estimate is conservative — realized LE gains often exceeded LC projections.
Open Questions and Known Shortcomings
Lee (2000) catalogs six shortcomings the model's author acknowledges:
Linear k not structural: The ARIMA(0,1,0) drift in kt is an empirical pattern, not derived from any theory of mortality change. No reason it must continue indefinitely.
bx instability: The age-sensitivity profile is assumed constant over the forecast horizon; it has changed historically and may change again. Carter and Prskawetz (2001) provide the clearest empirical test: fitting 30 rolling 24-year SVD submatrices to Austrian data 1947–1999, they find that ax and kt are relatively stable across submatrices but bx is not — pre-1970 improvement is concentrated at young/middle ages, post-1970 increasingly at old ages, with a structural shift ~1968. The extended LC method using submatrix-specific bx reduces e0 fitting error by ≈0.3 years for females and substantially for e60; forecasts to 2050 are sensitive to base-period choice, especially for males (+3.5 years by 2050 from 1976–1999 vs. the full 1947–1999 series). See Carter and Prskawetz 2001 — Examining Structural Shifts in Mortality Using the Lee-Carter Method. Booth, Maindonald, and Smith (2002) deepen this finding for Australia: they show that bx instability and kt non-linearity are jointly caused by fitting over too long a period, and that both are substantially resolved once the fitting period is restricted to the structurally homogeneous 1968–1999 regime — identified objectively via their R(S)/RD(S) chi-squared ratio criterion. Critically, Lee and Miller's universal recommendation of 1950 as a starting year is among the worst choices for Australia. See Booth Maindonald and Smith 2002 — Age-Time Interactions in Mortality Projections Applying Lee-Carter to Australia.
Interval narrowness (Alho 1990): Excluding the 1918 flu shock from kt's residual variance understates forecast uncertainty by ≈57%. Meseguer (2008) empirically confirms that LC's 80% intervals achieve only 50–70% empirical coverage for e0 and near-zero coverage for dependency ratios at long horizons.
Sex differential divergence: Male and female kt are forecast with separate drift rates; in the long run, the male–female e0 gap grows without bound — an implausible structural feature. Carter and Lee (1992 IJF) address this with a bivariate ARIMA for joint male–female kt.
Single-factor limitation: All age-group trends must move in proportion to bx; if young-age and old-age mortality diverge (as during the opioid crisis), the model is misspecified.
Parameter uncertainty exclusion: Classical LC treats ax, bx, and the drift d as known constants. Only kt forecast variance contributes to interval width. Bayesian implementations (e.g., Meseguer 2006) show that including parameter uncertainty roughly doubles or triples interval widths at long horizons.
Old-age extrapolation: The model makes no structural assumptions about mortality at extreme ages; projections for ages 90+ rely on the same bx extrapolation as younger ages, potentially missing mortality deceleration or acceleration at the frailest ages.
Bongaarts (2004) shifting logistic critique: Bongaarts shows that the assumption of constant ρ(x,t) is violated empirically: improvement rates have declined at younger ages and risen at older ages across decades. Under the shifting logistic model, ρs(x,t)=e˙s(t)⋅ks(x,t) is inherently time- and age-varying. Because LC extrapolates each age group's rate at its own historical exponential rate (equivalent to fixing different drift rates per age group), differences in the b(x) profile eventually cause adjacent age groups' projected rates either to cross or to diverge without bound — producing implausible age patterns over multi-decade horizons. Bongaarts proposes the Shifting Logistic Model as a structurally grounded alternative that avoids this pathology while maintaining the same short-run accuracy.
Full system interval collapse (Meseguer 2010): When LC is extended to a full demographic system (mortality + fertility + population), its interval failure for elderly population is even more extreme than the out-of-sample dependency ratio results. Meseguer (2010) finds that LC's 90% confidence interval for female population aged 65+ does not cover the SSA Alternative 2 projection at any point from 2009 to 2088 — an 80-year non-overlap. The BVAR with an informative prior anchored to Alt 2 (Inf-BVAR) produces a population projection closely tracking the official forecast in expectation while generating appropriately wider uncertainty bands. See SSA Mortality Forecasting.
Method-Robustness Evidence (Denton, Feaver, and Spencer 2005)
Denton et al. (2005) estimated AR(2) and QVAR(1) system models of Canadian mortality rates (1926–2000, 38 age-sex groups) and tested three stochastic forecasting methods — nonparametric block bootstrap, partially parametric (AR(2) + bootstrap residuals), and fully parametric (AR(2) + multivariate normal errors). Key findings relevant to LC robustness:
50th percentile of mean e0 at age 0 in 2050: (a) nonparametric = 84.48 yr, (b) partially parametric = 84.76 yr, (c) fully parametric = 84.74 yr — all within 0.5–0.8 years of each other, and 0.5–0.8 years below Tuljapurkar et al. (2000)'s LC-based Canadian median of 85.26 yr.
Methods (b) and (c) yield nearly identical 90% probability intervals (3.43 vs. 3.39 years at age 0, 2050): assuming normality vs. drawing bootstrap residuals makes essentially no difference.
The non-LC 90% intervals (3.39–3.43 years) are wider than the Tuljapurkar et al. LC interval (2.78 years), suggesting LC may understate uncertainty for Canada.
These results support the view that the choice of stochastic method is less consequential than structural assumptions about long-run trend continuation.