Lee-Carter Model

mortality-forecastingdemographyactuarialSVDARIMALee-Carterlife-expectancystochasticstate-spaceBayesianG7international

Methodological Predecessors

Carter and Lee (1986): The earliest direct ancestor of the LC architecture. Carter and Lee define a scalar marital fertility index g(t)g(t) and nuptiality index h(t)h(t), each convolved with a fixed demographic schedule to generate demographic event counts — the same "scalar index × fixed age schedule" structure that the Lee-Carter (LC) model applies to mortality six years later. The main difference is that LC uses singular value decomposition (SVD) to extract the scalar index k(t)k(t) from mortality data rather than constructing it as a ratio of observed to expected events. Carter and Lee also establish the autoregressive integrated moving average (ARIMA) framework for the scalar index and document the rapid growth of forecast uncertainty — themes that recur throughout the Lee-Carter literature. See Carter and Lee 1986 — Joint Forecasts of US Marital Fertility, Nuptiality, Births, and Marriages and Stochastic Fertility Forecasting.

Thompson et al. (1989): Thompson, Bell, Long, and Miller (1989) applied an identical dimensionality-reduction template to fertility: they fit a shifted gamma density to the U.S. age-specific fertility schedule each year, reducing 32 age-specific rates to three parameters (total fertility rate (TFR), mean age at childbearing (MACB), standard deviation of age at childbearing (SDACB)), then modeled those parameters jointly with a multivariate ARIMA. Lee and Carter (1992) cite this work explicitly and apply the same logic to mortality — compressing the age schedule to a single scalar index ktk_t and modeling that index as a random walk. The key simplification LC makes relative to Thompson et al. is collapsing three fertility parameters to one mortality index, justified by the empirical finding (92.5%92.5\% of variance explained by the first singular value pair) that mortality moves largely as a single factor.

Definition

The Lee-Carter (LC) model is a bilinear statistical model for forecasting age-specific mortality rates, introduced by Ronald Lee and Lawrence Carter in 1992. The model decomposes the log of the central death rate at age xx in year tt as:

lnm(x,t)=ax+bxkt+εx,t\ln m(x,t) = a_x + b_x k_t + \varepsilon_{x,t}

where axa_x is the time-average log death rate at age xx, bxb_x is each age group's sensitivity to the overall mortality trend, ktk_t is a latent scalar mortality index, and εx,t\varepsilon_{x,t} is an age- and year-specific error term. Forecasting mortality reduces to forecasting the single time series ktk_t, which is modeled as an ARIMA(0,1,0) random walk with drift.

Key Ideas

How It Works

Step 1 — SVD estimation: Compute the age-time matrix of demeaned log death rates, lnm(x,t)aˉx\ln m(x,t) - \bar{a}_x. Take the singular value decomposition; the leading singular vectors define bxb_x (right) and the unnormalized ktk_t (left). The first singular value pair typically accounts for >90%>90\% of total log-rate variance (92.5%92.5\% in the original U.S. fit over 1933–1987).

Step 2 — ktk_t re-estimation: SVD minimizes squared log-rate errors, not life expectancy errors. Jensen's inequality implies that the SVD ktk_t does not reproduce observed e0e_0 values. Lee and Carter re-estimate ktk_t each year as the value that makes the model's implied life expectancy exactly match the observed e0e_0, correcting this bias.

Step 3 — ARIMA(0,1,0) for ktk_t: Model the re-estimated ktk_t as a random walk with drift: kt=kt1+d+et,etN(0,σ2)k_t = k_{t-1} + d + e_t, \quad e_t \sim N(0, \sigma^2) where dd is estimated from the data (approximately 0.365-0.365 per year in the original U.S. fit). A pandemic dummy (e.g., 1918 flu) may be included to absorb structural breaks without contaminating the drift estimate.

Step 4 — Forecasting: Project ktk_t forward using the ARIMA model. The uncertainty in ktk_t at horizon hh grows as σh\sigma \sqrt{h}, translating into an expanding confidence interval (CI) for e0e_0. In the original 1992 fit, the 95% CI for e0e_0 in 2065 (75-year horizon) had half-width 5.65.6 years.

State-Space Reformulation (Pedroza and King 2002)

Pedroza and King show that LC can be expressed as a unified state-space model:

This formulation has three implications:

  1. LC as ordinary least squares (OLS): The SVD step is algebraically equivalent to maximum likelihood estimation (MLE) under homoskedastic Gaussian errors. The bilinear structure is not a special estimation trick — it is a statistical model.
  2. Incorrect LC forecasts: In state-space form, the correct one-step forecast conditions on the last residual ϵaT\epsilon_{aT}. Standard LC ignores it. The bias in point predictions is small (the per-age LC drift δa=βaθ\delta_a = \beta_a \theta nearly equals the per-age random-walk drift ϕa\phi_a, explaining why LC \approx fitting a separate random walk to each age group).
  3. Bayesian estimation: Gibbs sampling over all parameters jointly yields posteriors for αa\alpha_a, βa\beta_a, κt\kappa_t, θ\theta, σ2\sigma^2, and τ2\tau^2 with no extra programming effort, plus automatic imputation of missing rates and posterior predictive checks.

Even the corrected Bayesian state-space version fails to forecast well out-of-sample (U.S. males, U.S. females, Japanese males 1991–2000), particularly at ages 15–30 where mortality deviated from historical trend. The failure is structural — it traces to the constant αa\alpha_a, βa\beta_a assumption — not to the frequentist vs. Bayesian estimation choice. This motivates approaches that allow time-varying age sensitivity (e.g., Meseguer's Bayesian vector autoregression (BVAR)).

Out-of-Sample Validation and Jump-Off Fix

Validation (1989–1997, Lee 2000): In the 8 years after the 1992 publication, LC forecast e0e_0 gains of +1.30+1.30 years; the actual gain was +1.43+1.43 years. Social Security Administration (SSA)'s forecasts over the same period implied only +0.73+0.73 years — roughly half the realized gain. LC's error was a small overshoot; SSA's was a large undershoot. This is the primary empirical case for LC over the Office of the Chief Actuary (OACT) judgmental approach.

Jump-off fix: Standard LC forecasts initialize from the fitted axa_x (the time-averaged log mortality profile), causing a discontinuity at the jump-off year when recent observed rates deviate from the fit. The recommended fix (Bell 1997; Lee 2000) initializes from the most recent observed rates: lnm(x,t+s)=lnm(x,t)+bx(kt+skt)\ln m(x, t+s) = \ln m(x, t) + b_x (k_{t+s} - k_t) This eliminates the jump-off discontinuity without changing the drift structure and is now the standard implementation.

Multi-country validation of jump-off bias (Booth, Tickle, and Smith 2005): A 10-country out-of-sample evaluation (1986–1996 forecast period) formally decomposed LC error into jump-off and model components across Australia, Canada, Denmark, England & Wales, Finland, France, Italy, Norway, Sweden, and Switzerland. Jump-off bias accounts for 57%57\% of female and 42%42\% of male total LC forecast error; for Booth-Maindonald-Smith (BMS) the corresponding shares fall to 11%11\% and 4%4\%. Both Lee-Miller (LM) and BMS avoid the majority of jump-off bias by initializing from observed rather than fitted rates at the jump-off year. BMS is marginally superior to LM in a majority of countries (8/10 females, 6/10 males by root mean square error (RMSE) of e0e_0). See Booth Tickle and Smith 2005 — Evaluation of the Variants of the Lee-Carter Method of Forecasting Mortality A Multi-Country Comparison.

Cause disaggregation warning: Modeling cause-specific mortality and summing always produces higher all-cause projected mortality than modeling all-cause directly, because separate forecasts lose the negative cross-cause correlations. Wilmoth (1995) showed this is a mathematical result under standard assumptions: decomposed projections are always more pessimistic than aggregate projections. Agencies that disaggregate by cause will systematically overestimate total mortality and underestimate projected life expectancy.

International Generalization (Tuljapurkar, Li and Boe 2000)

Tuljapurkar, Li, and Boe applied the LC SVD decomposition to all seven Group of Seven (G7) countries (Canada, France, Germany, Italy, Japan, UK, US) using 1950–1994 data, establishing that the single-factor structure is not a US artifact. In every country the first singular value explains >94%>94\% of the temporal variance in log death rates (range 94.3%94.3\%97.5%97.5\%), and k(t)k(t) declines linearly in all seven. This cross-national validation — across very different social, historical, and epidemiological contexts — is the strongest evidence that the LC bilinear structure captures a genuine empirical regularity of modern mortality change.

Key extensions and findings from Tuljapurkar, Li, and Boe (TLB):

LC Variants (Booth 2006 Taxonomy)

By 2005, four main variants of the LC model had been proposed and evaluated (Booth 2006):

Multi-Population Extension (Zhou et al. 2014)

The single-population LC model cannot handle the joint modeling of two related populations without additional structure. Zhou et al. propose three approaches for the joint evolution of the two time-varying factors κ(1)\kappa^{(1)} and κ(2)\kappa^{(2)}, all enforcing long-run non-divergence:

Key finding: in the England-Wales / UK insured-lives pair, the smaller population (260K exposures) leads the larger (4.7M), not vice versa — the standard dominance assumption is reversed. See Multi-Population Mortality Modeling.

Why It Matters

Open Questions and Known Shortcomings

Lee (2000) catalogs six shortcomings the model's author acknowledges:

  1. Linear kk not structural: The ARIMA(0,1,0) drift in ktk_t is an empirical pattern, not derived from any theory of mortality change. No reason it must continue indefinitely.
  2. bxb_x instability: The age-sensitivity profile is assumed constant over the forecast horizon; it has changed historically and may change again. Carter and Prskawetz (2001) provide the clearest empirical test: fitting 3030 rolling 24-year SVD submatrices to Austrian data 1947–1999, they find that axa_x and ktk_t are relatively stable across submatrices but bxb_x is not — pre-1970 improvement is concentrated at young/middle ages, post-1970 increasingly at old ages, with a structural shift ~1968. The extended LC method using submatrix-specific bxb_x reduces e0e_0 fitting error by 0.3\approx 0.3 years for females and substantially for e60e_{60}; forecasts to 2050 are sensitive to base-period choice, especially for males (+3.5+3.5 years by 2050 from 1976–1999 vs. the full 1947–1999 series). See Carter and Prskawetz 2001 — Examining Structural Shifts in Mortality Using the Lee-Carter Method. Booth, Maindonald, and Smith (2002) deepen this finding for Australia: they show that bxb_x instability and ktk_t non-linearity are jointly caused by fitting over too long a period, and that both are substantially resolved once the fitting period is restricted to the structurally homogeneous 1968–1999 regime — identified objectively via their R(S)/RD(S) chi-squared ratio criterion. Critically, Lee and Miller's universal recommendation of 1950 as a starting year is among the worst choices for Australia. See Booth Maindonald and Smith 2002 — Age-Time Interactions in Mortality Projections Applying Lee-Carter to Australia.
  3. Interval narrowness (Alho 1990): Excluding the 1918 flu shock from ktk_t's residual variance understates forecast uncertainty by 57%\approx 57\%. Meseguer (2008) empirically confirms that LC's 80%80\% intervals achieve only 505070%70\% empirical coverage for e0e_0 and near-zero coverage for dependency ratios at long horizons.
  4. Sex differential divergence: Male and female ktk_t are forecast with separate drift rates; in the long run, the male–female e0e_0 gap grows without bound — an implausible structural feature. Carter and Lee (1992 IJF) address this with a bivariate ARIMA for joint male–female ktk_t.
  5. Single-factor limitation: All age-group trends must move in proportion to bxb_x; if young-age and old-age mortality diverge (as during the opioid crisis), the model is misspecified.
  6. Parameter uncertainty exclusion: Classical LC treats axa_x, bxb_x, and the drift dd as known constants. Only ktk_t forecast variance contributes to interval width. Bayesian implementations (e.g., Meseguer 2006) show that including parameter uncertainty roughly doubles or triples interval widths at long horizons.

Old-age extrapolation: The model makes no structural assumptions about mortality at extreme ages; projections for ages 90+ rely on the same bxb_x extrapolation as younger ages, potentially missing mortality deceleration or acceleration at the frailest ages.

Bongaarts (2004) shifting logistic critique: Bongaarts shows that the assumption of constant ρ(x,t)\rho(x,t) is violated empirically: improvement rates have declined at younger ages and risen at older ages across decades. Under the shifting logistic model, ρs(x,t)=e˙s(t)ks(x,t)\rho_s(x,t) = \dot{e}_s(t) \cdot k_s(x,t) is inherently time- and age-varying. Because LC extrapolates each age group's rate at its own historical exponential rate (equivalent to fixing different drift rates per age group), differences in the b(x)b(x) profile eventually cause adjacent age groups' projected rates either to cross or to diverge without bound — producing implausible age patterns over multi-decade horizons. Bongaarts proposes the Shifting Logistic Model as a structurally grounded alternative that avoids this pathology while maintaining the same short-run accuracy.

Full system interval collapse (Meseguer 2010): When LC is extended to a full demographic system (mortality + fertility + population), its interval failure for elderly population is even more extreme than the out-of-sample dependency ratio results. Meseguer (2010) finds that LC's 90%90\% confidence interval for female population aged 65+ does not cover the SSA Alternative 2 projection at any point from 2009 to 2088 — an 80-year non-overlap. The BVAR with an informative prior anchored to Alt 2 (Inf-BVAR) produces a population projection closely tracking the official forecast in expectation while generating appropriately wider uncertainty bands. See SSA Mortality Forecasting.

Method-Robustness Evidence (Denton, Feaver, and Spencer 2005)

Denton et al. (2005) estimated AR(2) and QVAR(1) system models of Canadian mortality rates (1926–2000, 38 age-sex groups) and tested three stochastic forecasting methods — nonparametric block bootstrap, partially parametric (AR(2) + bootstrap residuals), and fully parametric (AR(2) + multivariate normal errors). Key findings relevant to LC robustness:

See Denton Feaver and Spencer 2005 — Time Series Analysis and Stochastic Forecasting.

Related

Sources