Meseguer (2008) — Office of Research, Evaluation, and Statistics (ORES) Working Paper No. 111 — compares the out-of-sample performance of two stochastic age-specific mortality forecasting methods: the bias-corrected Lee-Carter (LC) model and the first-order autoregressive (AR(1)) model of Denton, Feaver, and Spencer (DFS 2005). Using Human Mortality Database (HMD) data for 16 high-income countries, 21 age groups, a 1980 jump-off year, rolling-origin forecasting design, and 20,000 simulated paths per model, the paper finds that the two methods produce nearly identical point forecasts but sharply divergent interval forecasts. LC intervals are systematically too narrow — catastrophically so at retirement ages, with 0% empirical coverage in 8 of 16 countries at 16+ year horizons for dependency ratios. The AR(1) intervals are too wide but adequately calibrated. Because Old-Age, Survivors, and Disability Insurance (OASDI) solvency is sensitive to dependency ratios rather than life expectancy at birth (), and because coverage is a misleading aggregate (errors cancel across ages), the paper recommends the AR(1) model for policy purposes.
"The two methods produce virtually identical point forecast accuracy, but diverge markedly in interval coverage — the Lee-Carter method is consistently too narrow, especially at old ages and long horizons."
"Life expectancy at birth coverage is near 100% for both methods even when the age-specific interval coverage is essentially zero — a consequence of errors canceling across ages in the aggregation."
"For the purposes of Social Security solvency analysis, dependency ratios are the appropriate criterion, not life expectancy. On this criterion, the AR(1) model substantially outperforms the Lee-Carter model."
The paper's most important contribution is identifying the cancellation artifact as a methodological trap for mortality forecasting validation. Researchers who validate interval forecasts using alone can achieve near-perfect coverage scores while their age-specific interval forecasts are catastrophically miscalibrated — and it is precisely the age-specific failures that matter for dependency ratio uncertainty, which drives OASDI solvency risk. The retirement-age bias finding (97–99% of MSE concentrated at ages 65–95+) is also consequential for the Lee (2003) forecasting story: LC systematically underestimates the pace of old-age improvement, which is consistent with the tilt mechanism (improvement is migrating toward ages that LC models less precisely). The paper predates the Soneji-King (2012) Bayesian critique by four years and addresses a different failure mode — not subjective judgment vs. statistical rigor, but parametric interval miscalibration — making the two critiques complementary: Soneji-King shows SSA uses the wrong inputs; Meseguer shows the standard stochastic method uses the wrong covariance structure even if the inputs are correct.