SSA Mortality Forecasting

mortality-forecastingsocial-securitySSAactuarialOACTsolvencytrust-fundBayesiansmokingobesityLTABstochastic-forecastingcause-of-deathslope-endpointtechnical-panel

Definition

Social Security Administration (SSA) mortality forecasting refers to the process by which the SSA Office of the Chief Actuary (OACT) produces age- and sex-specific mortality projections 75 years into the future, used as the mortality input for annual Old-Age, Survivors, and Disability Insurance (OASDI) solvency projections mandated by Congress. The forecasts determine how long beneficiaries are expected to live — and therefore how many years of benefit payments the Trust Funds must support — making them one of the most consequential and uncertain components of Social Security financial planning. Soneji and King (2012) were the first to document the process in sufficient detail for external replication and to propose a formal statistical alternative.

How It Works (Current OACT Method)

The OACT's forecasting process involves five sequential steps, each offering opportunities for subjective judgment:

  1. Collect mortality data. Cause-specific death and population counts, 1980 to the most recent available year. Seven cause categories: heart disease, cancer, vascular disease, violence, respiratory diseases, diabetes mellitus, and all other causes. Population under 65 from Census intercensal data; 65+ from Medicare enrollment.

  2. Calculate central death rates. Age-group, sex, and cause-specific rates for ages 0–1, 1–4, and five-year groups thereafter, with the open interval beginning at 95.

  3. Set historical trend. Fit a least-squares line to logged central death rates as a function of year. The average annual percentage reduction is 1exp(slope)1 - \exp(\text{slope}).

  4. Select 70 "ultimate rates of decline." For five broad age groups (under 15, 15–49, 50–64, 65–84, 85+), two sexes, and seven causes of death, OACT selects the annual reduction rate to be achieved by 75 years after the projection start. These are determined by subjective actuarial judgment about future mortality trajectories and approved by the OASDI trustees (political appointees). The reasoning behind specific choices is not publicly available.

  5. Interpolate. For years 2007–2008, rates equal the 1980–2006 historical average. From 2009 to 2034, log mortality declines linearly from the historical average toward the chosen ultimate rate. After 2034, the ultimate rate applies indefinitely. Final cause-specific rates are summed to produce all-cause mortality inputs.

Known methodological problems:

The 2003 Technical Panel (TP2003) Methodology Audit (Wilmoth 2005)

John R. Wilmoth, a member of the 2003 Technical Panel on Assumptions and Methods, conducted the most technically precise external audit of OACT's projection process. His analysis — published in GENUS LXI(1) — identified five sequential methodological decision points where OACT's procedure was either empirically unjustifiable or systematically biased toward pessimism about longevity.

The Five Decision Points and Their Bias Direction

Decision point OACT (Trustees Report 2003, TR2003) choice Alternative (Panel) Bias direction
Historical baseline period 1979197920002000 (21 yr, stagnation-heavy) 1950195020002000 (50 yr, full trend) Pessimistic
Trend computation method Endpoint (ratio of final/initial rates) Slope (ordinary least squares (OLS) full period ≈ Lee-Carter) Pessimistic (see below)
Projection disaggregation By cause of death All-cause Pessimistic (see below)
Ultimate rate specification 0.68%0.68\%/yr (65+65{+}) vs. historical 0.78%0.78\%/yr Match historical average Pessimistic
Convergence trajectory Rapid convergence to lower ultimate Gradual or later convergence Pessimistic

"Every choice pessimistic" (Table 7 finding): TR2003's intermediate scenario selects the longevity-pessimistic option at each of these five decision points. Individually each choice might be defensible, but the consistent alignment across all five is implausible as a genuine central forecast. The result is an intermediate projection that is functionally closer to a pessimistic scenario than a true median.

Slope vs. Endpoint Method

The slope method fits an OLS regression of log death rate on year using all data in the base period — this is mathematically nearly equivalent to the Lee-Carter model applied to the same data. The endpoint method computes the ratio of rates in the final period to rates in the initial period.

For ages 65–79, both methods yield similar trend rates over 1950195020002000. For ages 80+, they diverge substantially because the US experienced a period of stagnating old-age mortality improvement during the 1990s:

OACT's piecewise approach is approximately equivalent to the endpoint method. The Lee-Carter model is approximately equivalent to the slope method. At the policy-critical ages above 80, the choice implies a 0.140.14 percentage points (pp) per year difference in projected improvement — compounding over 75 years into a substantial life expectancy gap.

Cause-of-Death Forecasting as Hidden Deceleration

When all-cause mortality is derived by aggregating cause-specific forecasts, a structural bias emerges: as projected deaths shift increasingly toward the most slowly declining cause (it gains share mechanically as faster-declining causes shrink), the total projected improvement rate decelerates automatically. This "slowest-cause convergence" is an artifact of the aggregation method, not a substantive assumption about future mortality dynamics.

Both the 2003 and 2007 Technical Panels recommended eliminating cause-specific forecasting from the central projection and retaining it only for analytical decomposition of historical trends. The 70 ultimate rates of decline — derived from 21 years of data for 55 age groups ×\times 22 sexes ×\times 77 causes — also introduce parameter instability. Example: OACT assumed cancer mortality declines 0.5%0.5\%/year, but observed 1979197920002000 rates were 0.06%-0.06\%/year for men and 1.13%-1.13\%/year for women.

The US as International Outlier

Using Human Mortality Database (HMD) data for 15 high-income countries over 1980198020002000, Wilmoth showed that 11 of 15 countries experienced accelerating rates of old-age mortality decline. The US was one of only four deviants showing stagnation. This international comparison provides strong grounds for projecting a US recovery toward the historical and cross-national trend rather than anchoring long-run forecasts to the anomalous 1990s period.

Rau et al. (2006) independently confirm and extend this finding using Kannisto-Thatcher Database (KTDB) data for 27 developed countries through 2000: 22 of 27 countries showed acceleration in the 19901990s relative to the 19801980s; the US was among only 5 that did not. Critically, the US had been the world's mortality leader — lowest old-age death rates of any country — through the mid-1980s. Its post-1985 stagnation is therefore doubly anomalous: it occurs from a position of prior leadership and persists while nations with initially worse outcomes rapidly improve. Rau et al. explicitly state the US stagnation "is not yet understood." The convergence of two independent datasets (HMD and KTDB) and two independent analyses (Wilmoth 2005; Rau et al. 2006) on the same qualitative finding substantially strengthens the case against anchoring OACT projections to the US's recent experience.

Quantitative Benchmarks

Projection e0e_0 (2070) Long-Term Actuarial Balance (LTAB)
TR2003 intermediate 82.882.8 years 1.92-1.92 pp
TP2003 (Panel alternative) 84.484.4 years 1.68\approx -1.68 pp
TP1999 (prior Panel) 85.285.2 years
Lee-Carter (LC) stochastic median 86+\approx 86{+} years 3.3\approx -3.3 pp (Lee-Tuljapurkar 1998)

OACT's assumed ultimate rate of decline for ages 65+65{+} (0.68%0.68\%/year) falls below the observed 20th-century average (0.78%0.78\%/year) — a conservative assumption that is hard to defend given the international record.

LTAB decomposition of Panel impact: The 2003 Technical Panel's recommended changes have approximately offsetting LTAB effects: the migration adjustment (more immigrants \to more working-age contributors) reduces the LTAB by 0.250.25 pp; the mortality adjustment (higher projected longevity \to more benefit years) increases the LTAB by 0.240.24 pp. The net impact is near-zero — meaning adopting more defensible mortality methodology does not substantially change the measured solvency gap, but it changes where the uncertainty is.

The Soneji–King Bayesian Alternative

Soneji and King (2012) propose a Bayesian hierarchical mortality forecasting model with two classes of information:

Demographic patterns as priors:

Risk factors as covariates:

Using a 25-year lag allows the model to forecast 25 years ahead via one-step-ahead projection without needing to forecast the covariates themselves.

Model dependence: Rather than reporting a single forecast with confidence intervals from estimation uncertainty (which assumes the model is true), Soneji and King use robust Bayesian analysis — a class of priors — to generate a range of forecasts reflecting model uncertainty. This is the dominant source of forecasting uncertainty in this domain.

The Smoking–Longevity–Cost Linkage

A key mechanism driving the solvency gap is counterintuitive from a public health standpoint:

Smokers are actuarially profitable for Social Security. They pay payroll taxes during their working years but die earlier than non-smokers, collecting fewer years of retirement benefits. Soneji and King (2012) cite estimates showing smokers represent a net financial gain for the Trust Funds.

Declining smoking rates are costly for Social Security. As fewer Americans smoke, the ex-smoker cohort lives longer, drawing OASI benefits for more years. SSA's intermediate-cost scenario does not fully capture this mechanism. The Bayesian model, incorporating cohort smoking prevalence lagged 25 years, does — and produces longevity forecasts closer to SSA's high-cost (worst-case solvency) scenario.

Cohort smoking as empirical foundation: Preston and Wang (2006) provide the demographic bedrock for the cohort smoking covariate. Analyzing U.S. sex mortality differences 1948194820032003, they show the male/female death rate ratio organizes along cohort diagonals driven by differential smoking histories. Controlling for cohort smoking raises the estimated underlying mortality decline from 48%48\% to 56%56\%88 percentage points of genuine improvement masked by men's rising smoking burden. This directly demonstrates the forecasting bias that omitting smoking creates: period-based mortality extrapolation without a smoking covariate systematically underestimates both the pace of historical improvement and the prospective longevity gains from declining smoking. See Period Mortality and Samuel H. Preston.

Obesity partially offsets: Rising obesity rates increase mortality risk, which would improve solvency. But the net effect of both trends is more favorable to longevity (worse for solvency) than SSA's intermediate assumptions assume.

Solvency Implications

Substituting Soneji/King mortality forecasts for SSA intermediate assumptions, holding all other inputs fixed at SSA intermediate levels (2008 Trustees Report):

Solvency Measure (2031) Soneji/King SSA Intermediate Gap
Annual net income $217-\$217 billion $128-\$128 billion $89-\$89 billion
Trust fund balance $1.87\$1.87 trillion $2.67\$2.67 trillion $800-\$800 billion
Cost rate (% of taxable payroll) 17.70%17.70\% 17.05%17.05\% +0.65+0.65 pp
Program costs (% of GDP) 6.70%6.70\% 6.45%6.45\% +0.25+0.25 pp

Net income turns negative 2233 years earlier: 2023202320242024 under Soneji/King vs. 2026202620272027 under SSA intermediate.

Trust fund peak occurs earlier and at a lower level: $3.44\approx \$3.44 trillion in 20202020 (Soneji/King) vs. $3.72\approx \$3.72 trillion in 20222022 (SSA).

The Lee-Carter Model (Lee and Carter 1992)

Lee and Carter (1992) introduced the foundational stochastic alternative to OACT's judgmental approach. Their bilinear model decomposes the log central death rate at age xx in year tt as:

lnm(x,t)=ax+bxkt+εx,t\ln m(x,t) = a_x + b_x k_t + \varepsilon_{x,t}

where axa_x is the time-averaged log death rate at age xx, bxb_x is each age's sensitivity to the common mortality trend, and ktk_t is a scalar latent mortality index estimated by singular value decomposition (SVD) of the demeaned log-rate matrix. After SVD, ktk_t is re-estimated each year to match observed e0e_0 (correcting Jensen's inequality bias from modeling log rates). The index ktk_t is then modeled as a random walk with drift — autoregressive integrated moving average (ARIMA)(0,1,0):

kt=kt10.365+5.24flu+et,R2=0.995,se=0.651k_t = k_{t-1} - 0.365 + 5.24 \cdot \text{flu} + e_t, \quad R^2 = 0.995, \quad \text{se} = 0.651

where the flu dummy absorbs the 1918 pandemic and 0.365-0.365 is the estimated annual drift over 1933193319871987.

Estimation results (Table 1): bxb_x is highest at infant and child ages (0.09\approx 0.090.110.11) and lowest at oldest ages (0.029\approx 0.0290.0330.033 for ages 85+85{+}), reflecting faster historical proportional improvement at young ages. The pattern captures the Period Mortality bxb_x tilt: as mortality improvement shifted toward older ages over the 20th century, each unit of ktk_t decline generated progressively larger e0e_0 gains, maintaining linear e0e_0 growth.

Forecast: k(2065)=38.80k(2065) = -38.80, implying e0=86.05e_0 = 86.05 in 2065. The 95% confidence intervals (CIs) — derived from the ARIMA variance for ktk_t at a 75-year horizon — have half-width 5.65.6 years: [80.45,91.65][80.45, 91.65]. SSA's 1989 intermediate projection of e0=80.45e_0 = 80.45 for 2065 falls exactly at the lower bound of the LC interval.

Out-of-sample validation 1989–1997 (Lee 2000): LC forecast +1.30+1.30 yr e0e_0 gain; actual gain +1.43+1.43 yr. SSA forecast implied only +0.73+0.73 yr — half the realized gain. LC's tracking error: +0.13+0.13 yr (small overshoot). SSA's tracking error: 0.70-0.70 yr (large undershoot). The 55:11 accuracy ratio in LC's favor is the primary empirical case against the judgmental approach.

McNown critique (1992): The LC model is algebraically equivalent to projecting each age-specific death rate at its own historical exponential trend independently. The bilinear structure is a compact notation, not a separate identifying assumption.

Alho critique (1992): The 95% CIs are understated. The flu dummy absorbs a large shock from ktk_t's level but is excluded from ktk_t's residual variance; correcting for this would widen the CI by roughly 57%57\%. Meseguer (2008) empirically confirmed this: LC's 80%80\% intervals achieve 505070%70\% empirical coverage for e0e_0 and near-zero coverage for dependency ratios at long horizons across 16 countries. See Lee-Carter Model.

Cross-National Validation (Tuljapurkar, Li and Boe 2000)

Tuljapurkar, Li, and Boe (TLB, 2000) extended the LC stochastic approach to all G7 countries, establishing that the single-factor structure is a cross-national empirical law, not a US artifact. In every G7 country, the first singular value of the SVD of log m(x,t)m(x,t) explains >94%>94\% of temporal variance in log death rates (range 94.3%94.3\%97.5%97.5\%), and k(t)k(t) declines linearly in all seven countries over 1950195019941994. For the US specifically, the stochastic median e0e_0 in 2050 is 82.982.9 years versus an official projection of 80.580.5 years — a gap of 2.52.5 years. Across all G7 countries, stochastic medians exceed official central projections by 1.31.3 yr (UK) to 8.08.0 yr (Japan). Each 1-year difference in e0e_0 corresponds to >5%>5\% difference in the dependency ratio (65+/2065{+}/206464); TLB's stochastic medians imply dependency ratios 6%6\% (UK) to 40%40\% (Japan) higher than official projections by 2050.

The cross-national validation is the strongest evidence that the LC bilinear structure captures a genuine empirical regularity of modern mortality change rather than a feature specific to US data. It also establishes that official mortality forecasting agencies across G7 countries share a common pattern of underestimating longevity gains — the US SSA undershoot documented in Lee (2000) and Lee (2003) is not exceptional. See Lee-Carter Model and Shripad Tuljapurkar.

Early International Evidence Against SSA Projections (Lee and Skinner 1999)

Four years before Lee (2003) and a decade before Soneji-King (2012), Lee and Skinner (1999) presented the same core critique using a different methodological approach: direct comparison of SSA's projected mortality decline rates with observed rates in peer countries.

Projected vs. observed mortality decline rates for the period 1975–89 across UK, France, Sweden, Netherlands, and Japan:

The comparison involves five countries that already had life expectancies equal to or higher than the US — so the argument that the US had less room for improvement does not apply. Indeed, Lee and Skinner note that under SSA's own projections, the US will not reach Japan's current (1996) LE of 80.4 until 2051 — a 55-year lag. The paper concludes: "In our view, the central Social Security Administration forecasts of mortality decline are far too low."

This 1999 paper is the earliest accessible English-language statement of the SSA pessimism argument at the population level, predating both Lee (2003) and Soneji and King (2012). See Jonathan Skinner.

The Lee (2003) Forecasting Gap

Nearly a decade before Soneji and King (2012), Ronald Lee (2003) documented a substantial discrepancy between empirical trend extrapolations and SSA's 2002 intermediate-cost mortality projection. Comparing four forecasting approaches for US e0e_0 at 2030:

Forecast Source Annual gain (years/yr) US e0e_0 in 2030
SSA 2002 intermediate 0.10\approx 0.10 79.5\approx 79.5
Lee-Carter (LC) model +0.144+0.144 81.3\approx 81.3
White (2002) high-income-country average +0.22+0.22 83.3\approx 83.3
Oeppen-Vaupel (2002) record extrapolation +0.23+0.230.240.24 83.3\approx 83.3

The White and Oeppen-Vaupel extrapolations imply a 2030 life expectancy approximately 3.83.8 years above SSA's intermediate projection. The LC model itself projects gains 44%44\% larger than SSA's implicit rate. This gap does not arise from modeling uncertainty within the LC framework — LC confidence intervals are wide but centered substantially above SSA's figure.

Lee's interpretation: empirical linear trend extrapolations from high-income countries represent a more reliable baseline than SSA's actuarial judgments about the future pace of improvement. The 2002 projections' low trajectory is consistent with a history of SSA underestimating longevity gains in post-projection reviews.

This predates the Soneji-King (2012) critique by nine years. Where Soneji and King demonstrate that the discrepancy operates through SSA's failure to incorporate smoking and obesity as covariates, Lee (2003) identifies the magnitude of the gap from a purely demographic perspective — comparing model-based and record-based extrapolation methods. The two critiques are complementary: Lee establishes that SSA is far below the empirical trend; Soneji and King identify a causal mechanism explaining why. See Ronald Lee and Period Mortality.

Olshansky et al. (2005): Obesity and the Diabetes Forecasting Critique

A third critique of SSA's mortality forecasting — focused on obesity and diabetes — appeared in Olshansky et al. (2005). Where Lee (2003) identifies the magnitude of the SSA-trend gap from a demographic vantage point, and Soneji-King (2012) traces it to missing behavioral covariates, Olshansky et al. surface a specific, verifiable directional error in SSA's cause-specific projections.

The diabetes critique: SSA's 2003 actuarial study projected diabetes death rates to decline 1.01.03.2%3.2\%/yr from 2010 onward. But observed diabetes death rates rose 2.8%2.8\%/yr (males) and 1.8%1.8\%/yr (females) from 19791979 to 19991999 — opposite to the SSA projection and to what the obesity epidemic would predict. SSA assumed improvement was coming precisely when the disease burden was worsening.

Historical forecasting record: Olshansky et al. document SSA's age-65 LE forecasts from actuarial studies published 1952195220032003. SSA consistently underestimated pre-1980 LE gains, then switched to optimistic extrapolation after 1980 — precisely when LE at age 65 for US women began to stall.

The three critiques are complementary and address distinct failure modes:

Critique Primary failure identified
Lee (2003) SSA's trend rate is well below empirical extrapolations from best-practice data
Olshansky et al. (2005) SSA's diabetes death-rate projection is directionally opposite to observed trend
Soneji-King (2012) SSA omits smoking and obesity as covariates, missing the longevity effect of declining smoking

Together they establish that SSA's intermediate-cost mortality projections were systematically optimistic across multiple independent analytical angles. See S. Jay Olshansky and Period Mortality.

Stochastic Model Interval Forecast Failure (Meseguer 2008)

Meseguer (2008) — ORES Working Paper No. 111 — provides a pre-Soneji-King audit of the two leading stochastic alternatives to OACT's judgmental approach: the bias-corrected Lee-Carter (LC) model and the autoregressive order-1 (AR(1)) model of Denton, Feaver, and Spencer (DFS, 2005). Using HMD data for 16 high-income countries, 2121 age groups, a 19801980 jump-off year, and 20,00020{,}000 simulated paths per model, the paper isolates a different failure mode from Soneji-King's critique: not subjective judgment, but parametric interval miscalibration.

The two-part verdict:

The e0e_0 cancellation trap: Both models achieve near-100%100\% e0e_0 coverage even when age-specific intervals fail catastrophically. Mortality overestimation at young ages (small variance contribution) and underestimation at old ages partially cancel in life expectancy aggregation. Validating stochastic models using e0e_0 coverage alone masks the failure precisely at the ages that matter most for OASDI.

OASDI policy implication: Dependency ratios determine the ratio of benefit claimants to payroll contributors and are the correct solvency metric. LC's near-zero dependency ratio coverage means it dramatically understates actuarial uncertainty and provides false confidence about Trust Fund sustainability. Meseguer recommends the AR(1) model for OASDI policy purposes.

Complementarity with Soneji-King (2012): The two critiques address distinct failure modes. Soneji and King show SSA's OACT uses the wrong inputs (no smoking/obesity covariates, subjective ultimate rates of decline). Meseguer shows that even if the standard stochastic method (LC) is used with correct inputs, its covariance structure is mis-specified — the random-walk-with-drift variance for ktk_t does not capture joint age-profile uncertainty adequately, while DFS's common residual covariance matrix Ω^=SS/T\hat{\Omega} = S'S/T does. Together, the critiques imply that OASDI solvency estimates understate uncertainty through both the central-tendency and the distributional channels. See Javier Meseguer.

BVAR as Alternative to Lee-Carter (Meseguer 2006)

Meseguer (2006) — an SSA Office of Policy working paper — proposes Bayesian Vector Autoregression (BVAR) with the Minnesota prior as an alternative to Lee-Carter. It serves as the theoretical foundation for the 2008 out-of-sample study: where 2008 demonstrates empirically that LC intervals fail, 2006 explains mechanically why they must.

The overfitting problem. A full vector autoregression (VAR) with m=21m=21 age groups requires 462462 autoregressive parameters at p=1p=1 — far more than the 7474-year (1928192820012001) sample can identify. The Minnesota prior solves this via shrinkage: four hyperparameters (λ1\lambda_1: overall tightness; λ2\lambda_2: cross-equation weight relative to λ1\lambda_1; λ3\lambda_3: lag-decay rate; λ4\lambda_4: intercept tightness) constrain coefficients toward a random-walk-with-drift specification, reducing the effective free parameters to a manageable handful.

Two specifications compared. BVAR(1)-I sets λ2(i,j)=0.001\lambda_2(i,j) = 0.001 for iji \neq j, effectively suppressing cross-equation information — the system reduces to a seemingly unrelated regression equations (SURE) system of univariate autoregressions. BVAR(1)-II calibrates λ2(i,j)=0.8×Ωi,j\lambda_2(i,j) = 0.8 \times \Omega_{i,j} from the empirical 21×2121 \times 21 cross-age correlation matrix of Δlog(mortality)\Delta\log(\text{mortality}): age groups that are more closely correlated (e.g., ages 30303434 and 35353939, r=0.88r=0.88) receive higher cross-equation weight than distant groups (e.g., age 00 and ages 95+95{+}, r=0.19r=0.19). Estimation uses 30,00030{,}000 Gibbs sampling draws after 2,0002{,}000 burn-in iterations.

Model selection. BVAR(1)-II dominates BVAR(1)-I by 202202 marginal log-likelihood units (4,478.74{,}478.7 vs. 4,276.54{,}276.5) — decisive by Bayesian standards. It also reduces in-sample SSR by 2222%22\% in 2020 of 2121 age groups, confirming that adjacent age groups carry genuine cross-predictive information. Both models achieve 90%\approx 90\% empirical coverage of their nominal 90%90\% credible intervals.

Point forecasts: convergence. Median LE at birth in 2075: BVAR(1)-II = 85.6485.64 years, Lee-Carter = 86.2286.22 years — a gap of 7\approx 7 months. Both methods agree on the point trajectory.

Interval widths: divergence. BVAR credible intervals are substantially wider than Lee-Carter's, and the gap widens with both horizon and age:

Quantity BVAR(1)-II 90% width Lee-Carter 90% width Ratio
LE at birth, 2075 9.679.67 years 5.265.26 years 1.84×1.84\times
LE at age 65, 2075 8.368.36 years 3.883.88 years 2.15×2.15\times
LE at age 80, 2075 8.088.08 years 2.692.69 years 3.00×3.00\times

The specific 2075 quantiles for LE at age 80 make the overconfidence concrete: BVAR 55th–9595th interval = 7.997.9916.0816.08 years; LC = 10.6410.6413.3313.33 years. The LC model projects old-age survival improvement with three times less uncertainty than BVAR assigns.

Why LC is too narrow: the parameter uncertainty gap. LC treats the fitted αi\alpha_i and βi\beta_i as known constants — only uncertainty in the projected ktk_t contributes to interval width. Bayesian inference treats all parameters as random variables, so both parameter and sampling uncertainty contribute to credible interval width. Meseguer (2006) identifies a conceptual obstacle to resolving this within the classical framework: the sampling distribution of a classical estimator has support on the sample space, not the parameter space, so there is no coherent way to propagate parameter uncertainty into interval forecasts without the Bayesian machinery.

Connection to Meseguer (2008). The 2006 in-sample result (BVAR intervals are 1.81.83×3\times wider because they incorporate parameter uncertainty) directly predicts the 2008 out-of-sample finding (LC intervals achieve 0%0\% empirical coverage of dependency ratios in 88 of 1616 countries at long horizons). The theoretical mechanism (parameter uncertainty exclusion) and the empirical symptom (interval miscalibration) are the same phenomenon seen from different angles. Together, they constitute a complete case that LC's narrow intervals are not a sampling artifact but a structural feature of the classical estimator. See Javier Meseguer.

BVAR Full Demographic System (Meseguer 2010)

Meseguer (2010) — an unpublished SSA manuscript — extends the mortality-only BVAR of Meseguer (2006) to a full demographic system forecasting mortality, fertility, and population jointly through 2100. Using OACT data (mortality 1928192820082008 for 2222 age groups; fertility 1917191720082008 for 77 age groups), the paper estimates gender-specific BVARs with the Minnesota prior and feeds the results into a Leslie-matrix population projection. Two variants are compared: Diff-BVAR (diffuse prior, data-driven) and Inf-BVAR (informative prior, with steady-state means anchored to the SSA Alternative 2 trajectory via the Villani 2006 mean-adjusted BVAR framework).

The LC elderly population interval collapse: The most striking finding is that the Lee-Carter 90%90\% confidence interval for the female population aged 65+65{+} does not cover the SSA Alternative 2 projection from 2009 all the way to 2088 — an 8080-year span. This goes beyond the Meseguer (2008) finding of near-zero empirical coverage for dependency ratios: the LC interval fails to even overlap with the official SSA forecast for the most policy-critical age group.

Comparative life expectancy at birth (year 2100):

Model Male e0e_0 Female e0e_0
SSA Alt 2 84.184.1 87.387.3
Lee-Carter 84.484.4 90.390.3
Diff-BVAR 84.684.6 (90%90\% CI: 79.379.389.989.9) 90.190.1
Inf-BVAR 83.183.1 87.587.5

LC and Diff-BVAR agree on the point forecast for male e0e_0; both project higher female e0e_0 than Alt 2. Inf-BVAR, constrained toward Alt 2's steady state, closely tracks the official projection in expectation.

Dependency ratios (2100): Aged (65+/2065{+}/206464) — Alt 2 = 0.4330.433, Inf-BVAR 0.437\approx 0.437, LC = 0.4660.466. Total — Alt 2 = 0.880.88, Inf-BVAR = 0.8910.891, LC = 0.9140.914, Diff-BVAR = 0.9160.916. The LC model's elevated dependency ratio point forecast (a recurring pattern from Meseguer 2008's out-of-sample results) persists in the full system.

Methodological contribution: The Villani (2006) mean-adjusted BVAR allows the prior on the unconditional mean of the process to be specified directly — encoding Alt 2's long-run trajectory as the steady-state prior rather than constraining AR coefficients. This gives the Inf-BVAR a natural interpretation: it is consistent with SSA's official projection in expectation while providing Bayesian uncertainty quantification that is far wider than LC's.

Policy implication: The Inf-BVAR is the most policy-relevant specification: it neither dismisses SSA's actuarial judgment (as Diff-BVAR does) nor treats it as a point truth (as the classical LC approach implicitly does). It provides a principled probability distribution centered near the official projection but with honest, wider uncertainty bounds — addressing both the bias problem identified in Lee (2003) and the interval failure documented in Meseguer (2008). See Javier Meseguer and Lee-Carter Model.

The Lee-Tuljapurkar Stochastic Finance Framework (Lee and Tuljapurkar 1998)

While the critiques above (Lee 2003, Olshansky et al. 2005, Soneji-King 2012) focus on SSA's mortality forecasting methodology, Lee and Tuljapurkar (1998) ask a different question: what does full stochastic uncertainty across all major input variables imply for the probability distribution of OASDI solvency outcomes?

Method: 750750 Monte Carlo sample paths from 1995 to 2070, combining (1) Lee-Carter stochastic mortality, (2) stochastic fertility with long-run mean fixed at SSA's 1.91.9 children/woman but variance allowed to spread to a 95%95\% confidence interval (CI) of 0.70.73.33.3, and (3) AR(1) processes for productivity growth and interest rates constrained to SSA's middle assumptions of 1%1\% and 2.3%2.3\%/year respectively. Net immigration is held fixed at SSA levels.

Key results:

Quantity Lee-Tuljapurkar Stochastic SSA Intermediate
Mean e0e_0 in 2070 8686 years 80\approx 80 years
Mean trust fund exhaustion 20262026 20292029
95%95\% CI for exhaustion year 2014201420372037 single scenario
Mean LTAB 3.3-3.3 pp 2.2-2.2 pp
95%95\% CI for LTAB 0.2-0.2 to 6.5-6.5 pp (width 6.36.3) +0.5+0.5 to 5.7-5.7 pp (width 6.16.1)
Median payroll tax in 2070 21%\approx 21\% 19%19\%
P97.5 payroll tax in 2070 34%\approx 34\% 28%28\% (high-cost)

Uncertainty decomposition: Over a 75-year horizon, the dominant sources of LTAB variance are: fertility >> productivity growth >> interest rates >> mortality. SSA's own sensitivity analysis (Board of Trustees 1996, pp. 132–134) ranks these in the exact reverse order: mortality first, fertility last. The reversal occurs partly because SSA holds all other variables fixed at middle values while varying one at a time — a method that ignores interactions — while the stochastic framework allows all variables to vary jointly.

Tax implications: Reducing trust fund exhaustion probability to 5%5\% requires an immediate +5+5 pp payroll tax increase. Even a +2+2 pp immediate increase leaves 74%74\% of sample paths still exhausting by 2070.

Policy significance: The finding that fertility dominates 75-year LTAB uncertainty suggests that immigration and workforce-participation reforms may have larger long-run solvency implications than mortality-targeted interventions — a reordering of policy priorities relative to SSA's scenario framing. The stochastic approach also shows the Long-Term Actuarial Balance (LTAB) metric to be a mean of a very wide distribution, not a point estimate — a communication problem that the single-number summary obscures.

Politicization Risk

Because the OASDI trustees — who are political appointees — must approve the 70 ultimate rates of decline, the process is structurally exposed to political influence on the most consequential assumptions. When actuarial optimism is convenient (e.g., to avoid triggering legislated benefit reductions or to support a political narrative about program sustainability), the opaque subjective choices provide cover. Formal statistical methods are not immune to this, but they make assumptions transparent and auditable: disagreements can be focused on specific prior choices or covariate specifications rather than on undocumented expert judgment.

Subpopulation Divergence and Multi-Population Modeling

All SSA mortality forecasting methods model the Social Security Area population as a single entity. But SSA must also project mortality for subpopulations — DI beneficiaries, early retirees, cohorts by earnings level — whose mortality trends may diverge from the general population. Single-population Lee-Carter applied to a subpopulation imposes the same drift and bxb_x profile as the general population, which may be wrong. Zhou et al. (2014) establish that even very large population size does not guarantee the larger population "leads" the smaller: UK insured lives (260260K) demonstrably lead the English/Welsh general population (4.74.7M) in mortality improvement timing. The vector error correction model (VECM) multi-population framework — which tests and models the lead-lag relationship rather than assuming it — is the appropriate tool for jointly forecasting SSA-covered workers and DI-beneficiary mortality, or for modeling the widening educational/earnings mortality gradient (Waldron 2007). Longevity basis risk in this context is primarily a capital adequacy concern: best-estimate benefit projections change little (<3%<3\%) when switching from single- to multi-population models, but tail-risk scenarios (the analogue of Solvency II's 99.599.5th-percentile Solvency Capital Requirement (SCR)) can change by 101047%47\%. See Multi-Population Mortality Modeling.

Open Questions

Related

Sources