Summary
A comprehensive review of demographic forecasting methods published between 1980 and 2005, organized around three approaches (extrapolation, expectation, structural) and covering mortality, fertility, and full population projections. The central meta-finding, attributing Keilman (1997), is that despite twenty-five years of methodological development there is "no evidence that overall accuracy in mortality and fertility forecasting has improved over time." Probabilistic consistency — ensuring coherence between stochastic vital-rate forecasts and the population projection in which they are embedded — is identified as the field's major methodological advance of the period.
Key Claims
- Three-approach taxonomy: All demographic forecasting methods fall into one of three approaches: extrapolation (project past trends forward, no causal assumptions), expectation (individual surveys or expert judgment), or structural (causal modeling of drivers). Most practice is extrapolative; structural approaches face fundamental limits (Ashley's theorem).
- Factor hierarchy for mortality: Zero-factor models extrapolate age-specific rates independently; one-factor models (Lee-Carter, LC) capture the dominant trend via a single latent index; two-factor models allow for age-group divergence; three-factor models (age-period-cohort, APC) add cohort effects. Explanatory power and parsimony trade off across the hierarchy.
- Lee-Carter variants by 2005: Several variants had been proposed and formally evaluated:
- Lee-Miller (LM, 2001): the "standard" variant; uses fitting period beginning 1950 and initializes from observed rates at the jump-off year to eliminate the discontinuity identified by Bell (1997).
- Booth-Maindonald-Smith (BMS, 2002): data-driven fitting-period selection via the R(S)/RD(S) criterion; equally good as LM on average but situation-specific — superior in countries with recent structural breaks in k(t).
- Li and Lee (2005): coherent multi-population extension; all populations share a common time trend with population-specific deviations, preventing long-run divergence.
- Hyndman and Ullah (2007): functional data approach; extends LC to multiple latent factors via principal component analysis of smooth functional representations of mortality curves; handles the single-factor limitation.
- De Jong and Tickle (2006): state-space extension of the LC model; provides full probabilistic estimation of the joint evolution of age effects and the latent index; maximum-likelihood estimation avoids the two-stage singular value decomposition (SVD) procedure.
- Cause disaggregation always pessimistic: Disaggregating mortality by cause and summing the projections systematically produces higher all-cause mortality (lower projected life expectancy) than directly projecting all-cause mortality, because separate forecasts lose the negative cross-cause correlations. Wilmoth (1995) demonstrated this holds mathematically: decomposed projections are always higher than aggregate projections under standard assumptions.
- Assumption drag: Expert opinion on future mortality systematically underestimates the pace of improvement, a phenomenon called "assumption drag" (Alho and Spencer 1990). Expert panels anchor on recent experience and lag behind realized trends; confirmed for fertility as well (Lee 1999).
- Ashley's impossibility theorem (1983): Including an exogenous variable in a structural model improves forecast accuracy only if the model's forecast mean squared error (MSE) for that variable is less than the variable's unconditional variance. This condition becomes increasingly unlikely to hold as the forecast horizon lengthens, explaining why structural demographic models have not outperformed extrapolative ones at long horizons in practice.
- Accuracy non-improvement: "No evidence that overall accuracy in mortality and fertility forecasting has improved over time" (Keilman 1997). Despite major methodological advances, realized forecast errors have not diminished systematically across the 25-year review period.
- Three complementary uncertainty approaches: Three approaches to quantifying forecast uncertainty are now recognized as producing broadly consistent estimates: (1) ex ante model-based — stochastic time series generating prediction intervals analytically or by simulation (Lee-Carter, autoregressive moving average (ARMA)); (2) ex ante expert-based — structured expert elicitation as implemented in International Institute for Applied Systems Analysis (IIASA) world population projections (Lutz et al.); (3) ex post error-based — empirical distribution of historical forecast errors as an uncertainty benchmark (Keilman, Alho).
- Birth expectations ill-advised: Surveys of birth intentions and pregnancy expectations are unreliable inputs for fertility forecasting because they measure timing intentions (tempo effects) rather than quantum. Including birth expectations imports short-run timing fluctuations into long-run forecasts.
- Future directions (as of 2006): MicMac (multi-state micro-simulation) and agent-based models identified as methodological frontiers capable of linking individual-level demographic behavior to macro population outcomes.
Concepts Introduced or Extended
Entities Mentioned
Quotes
"No evidence that overall accuracy in mortality and fertility forecasting has improved over time." — Keilman (1997), as cited by Booth
My Take
This is a survey rather than an original empirical contribution; its value lies in the taxonomy and meta-findings, not new data. The three-approach taxonomy is a useful organizing frame, though the boundary between extrapolation and structural modeling is blurry in practice (Lee-Carter is extrapolative in form but appealed to structural regularities in its original framing). The accuracy non-improvement finding (Keilman 1997) is striking but should be interpreted carefully: a method that produces well-calibrated probability intervals can be better even if its point-forecast root mean squared error (RMSE) is unchanged. The most durable contribution is the synthesis of three uncertainty approaches as complementary — it legitimizes methodological pluralism without requiring convergence on a single standard. The Ashley theorem grounding for why structural models underperform is particularly useful for practitioners deciding whether to invest in cause-specific or socioeconomic driver models.