Social Security Administration (SSA) mortality forecasting refers to the process by which the SSA Office of the Chief Actuary (OACT) produces age- and sex-specific mortality projections 75 years into the future, used as the mortality input for annual Old-Age, Survivors, and Disability Insurance (OASDI) solvency projections mandated by Congress. The forecasts determine how long beneficiaries are expected to live — and therefore how many years of benefit payments the Trust Funds must support — making them one of the most consequential and uncertain components of Social Security financial planning. Soneji and King (2012) were the first to document the process in sufficient detail for external replication and to propose a formal statistical alternative.
The OACT's forecasting process involves five sequential steps, each offering opportunities for subjective judgment:
Collect mortality data. Cause-specific death and population counts, 1980 to the most recent available year. Seven cause categories: heart disease, cancer, vascular disease, violence, respiratory diseases, diabetes mellitus, and all other causes. Population under 65 from Census intercensal data; 65+ from Medicare enrollment.
Calculate central death rates. Age-group, sex, and cause-specific rates for ages 0–1, 1–4, and five-year groups thereafter, with the open interval beginning at 95.
Set historical trend. Fit a least-squares line to logged central death rates as a function of year. The average annual percentage reduction is .
Select 70 "ultimate rates of decline." For five broad age groups (under 15, 15–49, 50–64, 65–84, 85+), two sexes, and seven causes of death, OACT selects the annual reduction rate to be achieved by 75 years after the projection start. These are determined by subjective actuarial judgment about future mortality trajectories and approved by the OASDI trustees (political appointees). The reasoning behind specific choices is not publicly available.
Interpolate. For years 2007–2008, rates equal the 1980–2006 historical average. From 2009 to 2034, log mortality declines linearly from the historical average toward the chosen ultimate rate. After 2034, the ultimate rate applies indefinitely. Final cause-specific rates are summed to produce all-cause mortality inputs.
Known methodological problems:
John R. Wilmoth, a member of the 2003 Technical Panel on Assumptions and Methods, conducted the most technically precise external audit of OACT's projection process. His analysis — published in GENUS LXI(1) — identified five sequential methodological decision points where OACT's procedure was either empirically unjustifiable or systematically biased toward pessimism about longevity.
| Decision point | OACT (Trustees Report 2003, TR2003) choice | Alternative (Panel) | Bias direction |
|---|---|---|---|
| Historical baseline period | – (21 yr, stagnation-heavy) | – (50 yr, full trend) | Pessimistic |
| Trend computation method | Endpoint (ratio of final/initial rates) | Slope (ordinary least squares (OLS) full period ≈ Lee-Carter) | Pessimistic (see below) |
| Projection disaggregation | By cause of death | All-cause | Pessimistic (see below) |
| Ultimate rate specification | /yr () vs. historical /yr | Match historical average | Pessimistic |
| Convergence trajectory | Rapid convergence to lower ultimate | Gradual or later convergence | Pessimistic |
"Every choice pessimistic" (Table 7 finding): TR2003's intermediate scenario selects the longevity-pessimistic option at each of these five decision points. Individually each choice might be defensible, but the consistent alignment across all five is implausible as a genuine central forecast. The result is an intermediate projection that is functionally closer to a pessimistic scenario than a true median.
The slope method fits an OLS regression of log death rate on year using all data in the base period — this is mathematically nearly equivalent to the Lee-Carter model applied to the same data. The endpoint method computes the ratio of rates in the final period to rates in the initial period.
For ages 65–79, both methods yield similar trend rates over –. For ages 80+, they diverge substantially because the US experienced a period of stagnating old-age mortality improvement during the 1990s:
OACT's piecewise approach is approximately equivalent to the endpoint method. The Lee-Carter model is approximately equivalent to the slope method. At the policy-critical ages above 80, the choice implies a percentage points (pp) per year difference in projected improvement — compounding over 75 years into a substantial life expectancy gap.
When all-cause mortality is derived by aggregating cause-specific forecasts, a structural bias emerges: as projected deaths shift increasingly toward the most slowly declining cause (it gains share mechanically as faster-declining causes shrink), the total projected improvement rate decelerates automatically. This "slowest-cause convergence" is an artifact of the aggregation method, not a substantive assumption about future mortality dynamics.
Both the 2003 and 2007 Technical Panels recommended eliminating cause-specific forecasting from the central projection and retaining it only for analytical decomposition of historical trends. The 70 ultimate rates of decline — derived from 21 years of data for age groups sexes causes — also introduce parameter instability. Example: OACT assumed cancer mortality declines /year, but observed – rates were /year for men and /year for women.
Using Human Mortality Database (HMD) data for 15 high-income countries over –, Wilmoth showed that 11 of 15 countries experienced accelerating rates of old-age mortality decline. The US was one of only four deviants showing stagnation. This international comparison provides strong grounds for projecting a US recovery toward the historical and cross-national trend rather than anchoring long-run forecasts to the anomalous 1990s period.
Rau et al. (2006) independently confirm and extend this finding using Kannisto-Thatcher Database (KTDB) data for 27 developed countries through 2000: 22 of 27 countries showed acceleration in the s relative to the s; the US was among only 5 that did not. Critically, the US had been the world's mortality leader — lowest old-age death rates of any country — through the mid-1980s. Its post-1985 stagnation is therefore doubly anomalous: it occurs from a position of prior leadership and persists while nations with initially worse outcomes rapidly improve. Rau et al. explicitly state the US stagnation "is not yet understood." The convergence of two independent datasets (HMD and KTDB) and two independent analyses (Wilmoth 2005; Rau et al. 2006) on the same qualitative finding substantially strengthens the case against anchoring OACT projections to the US's recent experience.
| Projection | (2070) | Long-Term Actuarial Balance (LTAB) |
|---|---|---|
| TR2003 intermediate | years | pp |
| TP2003 (Panel alternative) | years | pp |
| TP1999 (prior Panel) | years | — |
| Lee-Carter (LC) stochastic median | years | pp (Lee-Tuljapurkar 1998) |
OACT's assumed ultimate rate of decline for ages (/year) falls below the observed 20th-century average (/year) — a conservative assumption that is hard to defend given the international record.
LTAB decomposition of Panel impact: The 2003 Technical Panel's recommended changes have approximately offsetting LTAB effects: the migration adjustment (more immigrants more working-age contributors) reduces the LTAB by pp; the mortality adjustment (higher projected longevity more benefit years) increases the LTAB by pp. The net impact is near-zero — meaning adopting more defensible mortality methodology does not substantially change the measured solvency gap, but it changes where the uncertainty is.
Soneji and King (2012) propose a Bayesian hierarchical mortality forecasting model with two classes of information:
Demographic patterns as priors:
Risk factors as covariates:
Using a 25-year lag allows the model to forecast 25 years ahead via one-step-ahead projection without needing to forecast the covariates themselves.
Model dependence: Rather than reporting a single forecast with confidence intervals from estimation uncertainty (which assumes the model is true), Soneji and King use robust Bayesian analysis — a class of priors — to generate a range of forecasts reflecting model uncertainty. This is the dominant source of forecasting uncertainty in this domain.
A key mechanism driving the solvency gap is counterintuitive from a public health standpoint:
Smokers are actuarially profitable for Social Security. They pay payroll taxes during their working years but die earlier than non-smokers, collecting fewer years of retirement benefits. Soneji and King (2012) cite estimates showing smokers represent a net financial gain for the Trust Funds.
Declining smoking rates are costly for Social Security. As fewer Americans smoke, the ex-smoker cohort lives longer, drawing OASI benefits for more years. SSA's intermediate-cost scenario does not fully capture this mechanism. The Bayesian model, incorporating cohort smoking prevalence lagged 25 years, does — and produces longevity forecasts closer to SSA's high-cost (worst-case solvency) scenario.
Cohort smoking as empirical foundation: Preston and Wang (2006) provide the demographic bedrock for the cohort smoking covariate. Analyzing U.S. sex mortality differences –, they show the male/female death rate ratio organizes along cohort diagonals driven by differential smoking histories. Controlling for cohort smoking raises the estimated underlying mortality decline from to — percentage points of genuine improvement masked by men's rising smoking burden. This directly demonstrates the forecasting bias that omitting smoking creates: period-based mortality extrapolation without a smoking covariate systematically underestimates both the pace of historical improvement and the prospective longevity gains from declining smoking. See Period Mortality and Samuel H. Preston.
Obesity partially offsets: Rising obesity rates increase mortality risk, which would improve solvency. But the net effect of both trends is more favorable to longevity (worse for solvency) than SSA's intermediate assumptions assume.
Substituting Soneji/King mortality forecasts for SSA intermediate assumptions, holding all other inputs fixed at SSA intermediate levels (2008 Trustees Report):
| Solvency Measure (2031) | Soneji/King | SSA Intermediate | Gap |
|---|---|---|---|
| Annual net income | billion | billion | billion |
| Trust fund balance | trillion | trillion | billion |
| Cost rate (% of taxable payroll) | pp | ||
| Program costs (% of GDP) | pp |
Net income turns negative – years earlier: – under Soneji/King vs. – under SSA intermediate.
Trust fund peak occurs earlier and at a lower level: trillion in (Soneji/King) vs. trillion in (SSA).
Lee and Carter (1992) introduced the foundational stochastic alternative to OACT's judgmental approach. Their bilinear model decomposes the log central death rate at age in year as:
where is the time-averaged log death rate at age , is each age's sensitivity to the common mortality trend, and is a scalar latent mortality index estimated by singular value decomposition (SVD) of the demeaned log-rate matrix. After SVD, is re-estimated each year to match observed (correcting Jensen's inequality bias from modeling log rates). The index is then modeled as a random walk with drift — autoregressive integrated moving average (ARIMA)(0,1,0):
where the flu dummy absorbs the 1918 pandemic and is the estimated annual drift over –.
Estimation results (Table 1): is highest at infant and child ages (–) and lowest at oldest ages (– for ages ), reflecting faster historical proportional improvement at young ages. The pattern captures the Period Mortality tilt: as mortality improvement shifted toward older ages over the 20th century, each unit of decline generated progressively larger gains, maintaining linear growth.
Forecast: , implying in 2065. The 95% confidence intervals (CIs) — derived from the ARIMA variance for at a 75-year horizon — have half-width years: . SSA's 1989 intermediate projection of for 2065 falls exactly at the lower bound of the LC interval.
Out-of-sample validation 1989–1997 (Lee 2000): LC forecast yr gain; actual gain yr. SSA forecast implied only yr — half the realized gain. LC's tracking error: yr (small overshoot). SSA's tracking error: yr (large undershoot). The : accuracy ratio in LC's favor is the primary empirical case against the judgmental approach.
McNown critique (1992): The LC model is algebraically equivalent to projecting each age-specific death rate at its own historical exponential trend independently. The bilinear structure is a compact notation, not a separate identifying assumption.
Alho critique (1992): The 95% CIs are understated. The flu dummy absorbs a large shock from 's level but is excluded from 's residual variance; correcting for this would widen the CI by roughly . Meseguer (2008) empirically confirmed this: LC's intervals achieve – empirical coverage for and near-zero coverage for dependency ratios at long horizons across 16 countries. See Lee-Carter Model.
Tuljapurkar, Li, and Boe (TLB, 2000) extended the LC stochastic approach to all G7 countries, establishing that the single-factor structure is a cross-national empirical law, not a US artifact. In every G7 country, the first singular value of the SVD of log explains of temporal variance in log death rates (range –), and declines linearly in all seven countries over –. For the US specifically, the stochastic median in 2050 is years versus an official projection of years — a gap of years. Across all G7 countries, stochastic medians exceed official central projections by yr (UK) to yr (Japan). Each 1-year difference in corresponds to difference in the dependency ratio (–); TLB's stochastic medians imply dependency ratios (UK) to (Japan) higher than official projections by 2050.
The cross-national validation is the strongest evidence that the LC bilinear structure captures a genuine empirical regularity of modern mortality change rather than a feature specific to US data. It also establishes that official mortality forecasting agencies across G7 countries share a common pattern of underestimating longevity gains — the US SSA undershoot documented in Lee (2000) and Lee (2003) is not exceptional. See Lee-Carter Model and Shripad Tuljapurkar.
Four years before Lee (2003) and a decade before Soneji-King (2012), Lee and Skinner (1999) presented the same core critique using a different methodological approach: direct comparison of SSA's projected mortality decline rates with observed rates in peer countries.
Projected vs. observed mortality decline rates for the period 1975–89 across UK, France, Sweden, Netherlands, and Japan:
The comparison involves five countries that already had life expectancies equal to or higher than the US — so the argument that the US had less room for improvement does not apply. Indeed, Lee and Skinner note that under SSA's own projections, the US will not reach Japan's current (1996) LE of 80.4 until 2051 — a 55-year lag. The paper concludes: "In our view, the central Social Security Administration forecasts of mortality decline are far too low."
This 1999 paper is the earliest accessible English-language statement of the SSA pessimism argument at the population level, predating both Lee (2003) and Soneji and King (2012). See Jonathan Skinner.
Nearly a decade before Soneji and King (2012), Ronald Lee (2003) documented a substantial discrepancy between empirical trend extrapolations and SSA's 2002 intermediate-cost mortality projection. Comparing four forecasting approaches for US at 2030:
| Forecast Source | Annual gain (years/yr) | US in 2030 |
|---|---|---|
| SSA 2002 intermediate | ||
| Lee-Carter (LC) model | ||
| White (2002) high-income-country average | ||
| Oeppen-Vaupel (2002) record extrapolation | – |
The White and Oeppen-Vaupel extrapolations imply a 2030 life expectancy approximately years above SSA's intermediate projection. The LC model itself projects gains larger than SSA's implicit rate. This gap does not arise from modeling uncertainty within the LC framework — LC confidence intervals are wide but centered substantially above SSA's figure.
Lee's interpretation: empirical linear trend extrapolations from high-income countries represent a more reliable baseline than SSA's actuarial judgments about the future pace of improvement. The 2002 projections' low trajectory is consistent with a history of SSA underestimating longevity gains in post-projection reviews.
This predates the Soneji-King (2012) critique by nine years. Where Soneji and King demonstrate that the discrepancy operates through SSA's failure to incorporate smoking and obesity as covariates, Lee (2003) identifies the magnitude of the gap from a purely demographic perspective — comparing model-based and record-based extrapolation methods. The two critiques are complementary: Lee establishes that SSA is far below the empirical trend; Soneji and King identify a causal mechanism explaining why. See Ronald Lee and Period Mortality.
A third critique of SSA's mortality forecasting — focused on obesity and diabetes — appeared in Olshansky et al. (2005). Where Lee (2003) identifies the magnitude of the SSA-trend gap from a demographic vantage point, and Soneji-King (2012) traces it to missing behavioral covariates, Olshansky et al. surface a specific, verifiable directional error in SSA's cause-specific projections.
The diabetes critique: SSA's 2003 actuarial study projected diabetes death rates to decline –/yr from 2010 onward. But observed diabetes death rates rose /yr (males) and /yr (females) from to — opposite to the SSA projection and to what the obesity epidemic would predict. SSA assumed improvement was coming precisely when the disease burden was worsening.
Historical forecasting record: Olshansky et al. document SSA's age-65 LE forecasts from actuarial studies published –. SSA consistently underestimated pre-1980 LE gains, then switched to optimistic extrapolation after 1980 — precisely when LE at age 65 for US women began to stall.
The three critiques are complementary and address distinct failure modes:
| Critique | Primary failure identified |
|---|---|
| Lee (2003) | SSA's trend rate is well below empirical extrapolations from best-practice data |
| Olshansky et al. (2005) | SSA's diabetes death-rate projection is directionally opposite to observed trend |
| Soneji-King (2012) | SSA omits smoking and obesity as covariates, missing the longevity effect of declining smoking |
Together they establish that SSA's intermediate-cost mortality projections were systematically optimistic across multiple independent analytical angles. See S. Jay Olshansky and Period Mortality.
Meseguer (2008) — ORES Working Paper No. 111 — provides a pre-Soneji-King audit of the two leading stochastic alternatives to OACT's judgmental approach: the bias-corrected Lee-Carter (LC) model and the autoregressive order-1 (AR(1)) model of Denton, Feaver, and Spencer (DFS, 2005). Using HMD data for 16 high-income countries, age groups, a jump-off year, and simulated paths per model, the paper isolates a different failure mode from Soneji-King's critique: not subjective judgment, but parametric interval miscalibration.
The two-part verdict:
Point forecasts: Nearly identical (root mean squared error (RMSE) ratios LC/DFS across all countries and horizons). Both models systematically overestimate mortality at ages 65–95+, causing underestimation of and dependency ratios. Retirement ages account for – of total age-profile mean squared error (MSE).
Interval forecasts: Sharply divergent. LC's intervals achieve only – empirical probability content at long horizons; coverage for (old-age/total dependency ratio) reaches in of countries at year horizons. The DFS AR(1)'s intervals are too wide but adequately calibrated — coverage at or above nominal level.
The cancellation trap: Both models achieve near- coverage even when age-specific intervals fail catastrophically. Mortality overestimation at young ages (small variance contribution) and underestimation at old ages partially cancel in life expectancy aggregation. Validating stochastic models using coverage alone masks the failure precisely at the ages that matter most for OASDI.
OASDI policy implication: Dependency ratios determine the ratio of benefit claimants to payroll contributors and are the correct solvency metric. LC's near-zero dependency ratio coverage means it dramatically understates actuarial uncertainty and provides false confidence about Trust Fund sustainability. Meseguer recommends the AR(1) model for OASDI policy purposes.
Complementarity with Soneji-King (2012): The two critiques address distinct failure modes. Soneji and King show SSA's OACT uses the wrong inputs (no smoking/obesity covariates, subjective ultimate rates of decline). Meseguer shows that even if the standard stochastic method (LC) is used with correct inputs, its covariance structure is mis-specified — the random-walk-with-drift variance for does not capture joint age-profile uncertainty adequately, while DFS's common residual covariance matrix does. Together, the critiques imply that OASDI solvency estimates understate uncertainty through both the central-tendency and the distributional channels. See Javier Meseguer.
Meseguer (2006) — an SSA Office of Policy working paper — proposes Bayesian Vector Autoregression (BVAR) with the Minnesota prior as an alternative to Lee-Carter. It serves as the theoretical foundation for the 2008 out-of-sample study: where 2008 demonstrates empirically that LC intervals fail, 2006 explains mechanically why they must.
The overfitting problem. A full vector autoregression (VAR) with age groups requires autoregressive parameters at — far more than the -year (–) sample can identify. The Minnesota prior solves this via shrinkage: four hyperparameters (: overall tightness; : cross-equation weight relative to ; : lag-decay rate; : intercept tightness) constrain coefficients toward a random-walk-with-drift specification, reducing the effective free parameters to a manageable handful.
Two specifications compared. BVAR(1)-I sets for , effectively suppressing cross-equation information — the system reduces to a seemingly unrelated regression equations (SURE) system of univariate autoregressions. BVAR(1)-II calibrates from the empirical cross-age correlation matrix of : age groups that are more closely correlated (e.g., ages – and –, ) receive higher cross-equation weight than distant groups (e.g., age and ages , ). Estimation uses Gibbs sampling draws after burn-in iterations.
Model selection. BVAR(1)-II dominates BVAR(1)-I by marginal log-likelihood units ( vs. ) — decisive by Bayesian standards. It also reduces in-sample SSR by – in of age groups, confirming that adjacent age groups carry genuine cross-predictive information. Both models achieve empirical coverage of their nominal credible intervals.
Point forecasts: convergence. Median LE at birth in 2075: BVAR(1)-II = years, Lee-Carter = years — a gap of months. Both methods agree on the point trajectory.
Interval widths: divergence. BVAR credible intervals are substantially wider than Lee-Carter's, and the gap widens with both horizon and age:
| Quantity | BVAR(1)-II 90% width | Lee-Carter 90% width | Ratio |
|---|---|---|---|
| LE at birth, 2075 | years | years | |
| LE at age 65, 2075 | years | years | |
| LE at age 80, 2075 | years | years |
The specific 2075 quantiles for LE at age 80 make the overconfidence concrete: BVAR th–th interval = – years; LC = – years. The LC model projects old-age survival improvement with three times less uncertainty than BVAR assigns.
Why LC is too narrow: the parameter uncertainty gap. LC treats the fitted and as known constants — only uncertainty in the projected contributes to interval width. Bayesian inference treats all parameters as random variables, so both parameter and sampling uncertainty contribute to credible interval width. Meseguer (2006) identifies a conceptual obstacle to resolving this within the classical framework: the sampling distribution of a classical estimator has support on the sample space, not the parameter space, so there is no coherent way to propagate parameter uncertainty into interval forecasts without the Bayesian machinery.
Connection to Meseguer (2008). The 2006 in-sample result (BVAR intervals are – wider because they incorporate parameter uncertainty) directly predicts the 2008 out-of-sample finding (LC intervals achieve empirical coverage of dependency ratios in of countries at long horizons). The theoretical mechanism (parameter uncertainty exclusion) and the empirical symptom (interval miscalibration) are the same phenomenon seen from different angles. Together, they constitute a complete case that LC's narrow intervals are not a sampling artifact but a structural feature of the classical estimator. See Javier Meseguer.
Meseguer (2010) — an unpublished SSA manuscript — extends the mortality-only BVAR of Meseguer (2006) to a full demographic system forecasting mortality, fertility, and population jointly through 2100. Using OACT data (mortality – for age groups; fertility – for age groups), the paper estimates gender-specific BVARs with the Minnesota prior and feeds the results into a Leslie-matrix population projection. Two variants are compared: Diff-BVAR (diffuse prior, data-driven) and Inf-BVAR (informative prior, with steady-state means anchored to the SSA Alternative 2 trajectory via the Villani 2006 mean-adjusted BVAR framework).
The LC elderly population interval collapse: The most striking finding is that the Lee-Carter confidence interval for the female population aged does not cover the SSA Alternative 2 projection from 2009 all the way to 2088 — an -year span. This goes beyond the Meseguer (2008) finding of near-zero empirical coverage for dependency ratios: the LC interval fails to even overlap with the official SSA forecast for the most policy-critical age group.
Comparative life expectancy at birth (year 2100):
| Model | Male | Female |
|---|---|---|
| SSA Alt 2 | ||
| Lee-Carter | ||
| Diff-BVAR | ( CI: –) | |
| Inf-BVAR |
LC and Diff-BVAR agree on the point forecast for male ; both project higher female than Alt 2. Inf-BVAR, constrained toward Alt 2's steady state, closely tracks the official projection in expectation.
Dependency ratios (2100): Aged (–) — Alt 2 = , Inf-BVAR , LC = . Total — Alt 2 = , Inf-BVAR = , LC = , Diff-BVAR = . The LC model's elevated dependency ratio point forecast (a recurring pattern from Meseguer 2008's out-of-sample results) persists in the full system.
Methodological contribution: The Villani (2006) mean-adjusted BVAR allows the prior on the unconditional mean of the process to be specified directly — encoding Alt 2's long-run trajectory as the steady-state prior rather than constraining AR coefficients. This gives the Inf-BVAR a natural interpretation: it is consistent with SSA's official projection in expectation while providing Bayesian uncertainty quantification that is far wider than LC's.
Policy implication: The Inf-BVAR is the most policy-relevant specification: it neither dismisses SSA's actuarial judgment (as Diff-BVAR does) nor treats it as a point truth (as the classical LC approach implicitly does). It provides a principled probability distribution centered near the official projection but with honest, wider uncertainty bounds — addressing both the bias problem identified in Lee (2003) and the interval failure documented in Meseguer (2008). See Javier Meseguer and Lee-Carter Model.
While the critiques above (Lee 2003, Olshansky et al. 2005, Soneji-King 2012) focus on SSA's mortality forecasting methodology, Lee and Tuljapurkar (1998) ask a different question: what does full stochastic uncertainty across all major input variables imply for the probability distribution of OASDI solvency outcomes?
Method: Monte Carlo sample paths from 1995 to 2070, combining (1) Lee-Carter stochastic mortality, (2) stochastic fertility with long-run mean fixed at SSA's children/woman but variance allowed to spread to a confidence interval (CI) of –, and (3) AR(1) processes for productivity growth and interest rates constrained to SSA's middle assumptions of and /year respectively. Net immigration is held fixed at SSA levels.
Key results:
| Quantity | Lee-Tuljapurkar Stochastic | SSA Intermediate |
|---|---|---|
| Mean in 2070 | years | years |
| Mean trust fund exhaustion | ||
| CI for exhaustion year | – | single scenario |
| Mean LTAB | pp | pp |
| CI for LTAB | to pp (width ) | to pp (width ) |
| Median payroll tax in 2070 | ||
| P97.5 payroll tax in 2070 | (high-cost) |
Uncertainty decomposition: Over a 75-year horizon, the dominant sources of LTAB variance are: fertility productivity growth interest rates mortality. SSA's own sensitivity analysis (Board of Trustees 1996, pp. 132–134) ranks these in the exact reverse order: mortality first, fertility last. The reversal occurs partly because SSA holds all other variables fixed at middle values while varying one at a time — a method that ignores interactions — while the stochastic framework allows all variables to vary jointly.
Tax implications: Reducing trust fund exhaustion probability to requires an immediate pp payroll tax increase. Even a pp immediate increase leaves of sample paths still exhausting by 2070.
Policy significance: The finding that fertility dominates 75-year LTAB uncertainty suggests that immigration and workforce-participation reforms may have larger long-run solvency implications than mortality-targeted interventions — a reordering of policy priorities relative to SSA's scenario framing. The stochastic approach also shows the Long-Term Actuarial Balance (LTAB) metric to be a mean of a very wide distribution, not a point estimate — a communication problem that the single-number summary obscures.
Because the OASDI trustees — who are political appointees — must approve the 70 ultimate rates of decline, the process is structurally exposed to political influence on the most consequential assumptions. When actuarial optimism is convenient (e.g., to avoid triggering legislated benefit reductions or to support a political narrative about program sustainability), the opaque subjective choices provide cover. Formal statistical methods are not immune to this, but they make assumptions transparent and auditable: disagreements can be focused on specific prior choices or covariate specifications rather than on undocumented expert judgment.
All SSA mortality forecasting methods model the Social Security Area population as a single entity. But SSA must also project mortality for subpopulations — DI beneficiaries, early retirees, cohorts by earnings level — whose mortality trends may diverge from the general population. Single-population Lee-Carter applied to a subpopulation imposes the same drift and profile as the general population, which may be wrong. Zhou et al. (2014) establish that even very large population size does not guarantee the larger population "leads" the smaller: UK insured lives (K) demonstrably lead the English/Welsh general population (M) in mortality improvement timing. The vector error correction model (VECM) multi-population framework — which tests and models the lead-lag relationship rather than assuming it — is the appropriate tool for jointly forecasting SSA-covered workers and DI-beneficiary mortality, or for modeling the widening educational/earnings mortality gradient (Waldron 2007). Longevity basis risk in this context is primarily a capital adequacy concern: best-estimate benefit projections change little () when switching from single- to multi-population models, but tail-risk scenarios (the analogue of Solvency II's th-percentile Solvency Capital Requirement (SCR)) can change by –. See Multi-Population Mortality Modeling.