Stochastic fertility forecasting uses time-series models to generate probability distributions over future trajectories of the total fertility rate (TFR) and age-specific fertility rates, replacing scenario-based point forecasts with prediction intervals that explicitly quantify forecast uncertainty. Two distinct approaches exist: Thompson et al.'s (1989) parameterized schedule approach, which models the full age distribution of fertility through a gamma curve with three demographically interpretable parameters (TFR, mean age at childbearing [MACB], and standard deviation of age at childbearing [SDACB]) estimated jointly via a multivariate autoregressive integrated moving average (ARIMA); and Lee's constrained autoregressive moving average (ARMA) approach, which models TFR alone as an ARMA process mean-reverting toward a long-run average F∗, sacrificing age-schedule detail to focus entirely on long-run fertility-level uncertainty. The central methodological challenge in both approaches is that the long-run average TFR operates on 30–50 year generational timescales that cannot be reliably identified from available historical data, creating irreducible structural uncertainty on top of the statistical parameter uncertainty the ARIMA/ARMA model does estimate.
Key Ideas
Carter and Lee (1986) — index-based joint forecasting: Defined a scalar marital fertility index g(t) and nuptiality index h(t), each convolved with fixed demographic age/duration schedules (f(x) for fertility, n(a) for nuptiality) to generate births and first marriages. Five Box-Jenkins modeling strategies (Models A–E) were compared. For nuptiality, the univariate ARIMA (Model A) was optimal; the vector ARMA (Model D) was rejected for implying implausible negative nuptiality-fertility feedback. For fertility, a transfer function (Model B) outperformed the univariate ARIMA with slightly smaller root mean squared error (RMSE). The combined system implies an ultimate net reproduction rate (NRR) of 0.94 and unbounded forecast uncertainty: the 95% confidence interval (CI) for births spans ±100% of the forecast level by year 14, and the error correlation between birth and marriage forecasts grows to 0.86 by year 18. This "scalar index × fixed age schedule" architecture directly foreshadows the Lee-Carter Model, which applies the same logic to mortality via singular value decomposition (SVD). See Carter and Lee 1986 — Joint Forecasts of US Marital Fertility, Nuptiality, Births, and Marriages.
Thompson et al. (1989) — parameterized schedule approach: Fit a shifted gamma density to the age-specific fertility schedule each year; the density has three parameters — TFR (level), MACB (mean age at childbearing), and SDACB (standard deviation of age at childbearing). Model TFR, MACB, and SDACB jointly via a restricted multivariate ARIMA(0,1,0). Dimensionality falls from 32 age-specific series to 3. A bias adjustment (random-walk forecast of deviations of actual rates from the fitted curve) improves short-term accuracy but is negligible beyond ≈10 years. Bell (1997) confirms this cross-method: structured models that skip the bias adjustment lose to the age-specific random walk benchmark by 1.3–3×; Lee-Carter (LC) with bias adjustment ≈ random walk with drift (RWD) (mean absolute percentage error (MAPE) ratio 0.98–1.07). See Bell 1997 — Comparing and Assessing Time Series Methods for Forecasting Age-Specific Fertility and Mortality Rates. Key empirical finding: TFR is nearly independent of the age distribution (MACB and SDACB); the only strong relationship is the contemporaneous MACB–SDACB correlation. The approach was implemented in Census Bureau 1988 population projections by blending time-series short-term forecasts with judgmentally set ultimate levels — a model–judgment hybrid that later informed the Lee-Tuljapurkar framework. The template (compress age schedule to a few parameters; model those parameters) directly inspired Lee and Carter (1992)'s single-index mortality model. See Lee-Carter Model.
Lee's constrained ARMA model: (Ft−F∗)=a(Ft−1−F∗)+Ut+cUt−1. The TFR at time t is modeled as an ARMA process mean-reverting toward F∗, a fixed long-run average set a priori. For the US, the autoregressive coefficient a is close to 1 (near unit root), making short-run shocks persistent. The moving-average coefficient c is estimated from data.
Rapid uncertainty growth: 95% prediction intervals (PI) widen to a range of one child within 6 years of the launch date; asymptotic width exceeds two children in approximately ≈18 years, after which the interval width plateaus. Beyond 6 years, point predictions carry little information relative to the uncertainty.
The F∗ problem (structural vs. parameter uncertainty): F∗ cannot be reliably estimated from TFR data because the long-run generational dynamics of fertility operate on timescales (30–50 years) much longer than the available US record (1917–1992, 75 years). Time-series methods identify the short-run autocorrelation structure but cannot determine where fertility will average over the next century. Choosing F∗ far from the historical sample mean widens intervals via the (F∗−Fˉ)2 contribution to prediction variance.
No empirical attractor: A plot of US TFR at time t against TFR at time t−1 shows values clustered near the 45-degree line with no discernible attractor. There is no statistical evidence that fertility trends toward any particular level.
Expert myopia: Social Security Administration (SSA) and US Census Bureau anchor F∗ near recent replacement-level fertility (2.1), reflecting short-term experience rather than the longer historical range (US TFR ranged from 1.7 to 3.8 from 1917–1992; historical mean ≈2.6). This biases their long-run intervals downward and understates uncertainty over 75-year forecasting horizons.
Validation failures at short sample length: Historical launch experiments using 1917–1945 data with F∗=1.8 or 2.1 completely fail to capture the post-World War II (WWII) baby boom; only F∗=3.0 comes close. Later launches (1955 onward) perform well. This demonstrates that sample-based uncertainty is insufficient when subsequent variation substantially exceeds sample variation — a lesson with direct implications for current long-horizon Social Security forecasting.
Alho and Spencer (1985) — approximately linear propagation-of-error framework: The foundational US application of formal stochastic error analysis to official population forecasts. Key contributions: (a) the approximately linear modelV(t+1)=R(t)V(t) propagates errors through the full cohort-component system; (b) three additive error sources: model misspecification (bias η(t)), parameter estimation error, and random variation; (c) mixed estimation incorporates expert judgment as a pseudo-observation at time n+m+1, with weight κ controlling the blend between data-driven and judgmental priors — the direct precedent for Bayesian F∗ weighting; (d) fertility modeled as autoregressive order-1 (AR(1)) in log age-specific rates (autocorrelation ≈0.71, cross-correlations ≈0.88), using 1957-onward data to avoid the baby-boom structural break; (e) US Bureau of the Census (USBC) intervals too narrow: Alho-Spencer 67% PI upper/births forecast =1.096 at t=1, 1.315 at t=15 vs. USBC high/medium ratios of 1.010 and 1.143 — the official high-low is far too narrow for a 67% probabilistic interpretation; (f) fertility dominates total uncertainty at all horizons; migration is secondary at 15 years. See Alho and Spencer 1985 — Uncertain Population Forecasting.
Alho's naïve/base-line error approach (Alho 1990, 1997): An alternative, model-free way to bound forecast uncertainty. For each lead time t, make a naïve forecast (fertility = current TFR; mortality = constant decline rate) and compute the actual error. The distribution of these historical naïve errors gives a conservative upper bound on the ex ante error of any more sophisticated forecasting method that claims to be better than naïve. For log TFR in industrialized countries, errors are well approximated by et∼N(0,tσe2) with σe≈0.08, implying ≈5% expected error at t=1 and ≈40% at t=15. Alho applies this framework to evaluate official scenarios: the International Institute for Applied Systems Analysis (IIASA) 1994 world population high-low interval for 2030 translates to roughly an 85% prediction interval under the volatility-based model — a much stronger claim than official presentations imply. The United Nations (UN) 1993 narrower interval carries only ≈51% probability content unless one assumes large, implausible demographic policy success. See Alho 1997 — Scenarios Uncertainty and Conditional Forecasts of the World Population.
Non-Stationary Dynamic Factor Models (Ortega and Poncela, n.d.): An alternative forecasting approach that exploits cross-country commonality in TFR trends. Applied to four Southern European countries (Portugal, Spain, Greece, Italy), a Dynamic Factor Model (DFM) extracts one clearly non-stationary common factor (a random walk) capturing the shared post-1970 fertility decline, and a possible second borderline-stationary factor capturing timing differences. Joint modelling provides no RMSE gains at 1-year horizons (where ARIMA wins) but reduces RMSE by ≈1/3 vs. naive at 10-year horizons. ARIMA-with-trend (ARIT) collapses catastrophically at long horizons because linear fertility drift is unrealistic over a decade. The approach is restricted to TFR trends and does not incorporate cohort, parity, or postponement information — it is a complement to the Kohler-Ortega (K-O) tempo-adjustment framework, not a substitute.
Birth expectations ill-advised (Booth 2006): Surveys of birth intentions and pregnancy expectations are an unreliable input for fertility forecasting because they measure timing intentions (tempo effects) rather than quantum. Birth expectations capture when women plan to have children, not how many; including them tends to import short-run timing fluctuations into long-run forecasts and has consistently been found to add noise rather than signal.
Bayesian weighting of F∗:
Malthusian prior — flat weight on F∗∈[1.9,2.1] with exponentially declining weight above — produces nearly identical intervals to the F∗=2.1 benchmark.
Empirical prior — weights proportional to historical TFR frequency (mean ≈2.6) — produces intervals ≈15% wider and shifts the long-run forecast level up toward the historical mean.
Even a 15% widening from an extreme historical prior understates total structural uncertainty, since the empirical prior itself may not bound the plausible range of future long-run fertility.
How It Works
Model Estimation
Fix F∗ (from expert judgment or priors).
Detrend the TFR series: (Ft−F∗).
Estimate ARMA(1,1) coefficients a and c by maximum likelihood over the base period.
Estimate σu2, the innovation variance.
Simulate forward from the launch date to generate the predictive distribution.
Prediction Interval Width
The variance of the h-step-ahead forecast grows as a function of a, c, σu2, and (F∗−Fˉ)2. Because a≈1 (near unit root), variance grows rapidly in the first ≈18 years then asymptotes. The asymptotic width is approximately:
For a discrete set of F∗ values {f1,f2,…} with prior weights {p1,p2,…}, the overall forecast is:
P(TFR≤x at horizon h)=i∑pi⋅P(TFR≤x∣F∗=fi)
This mixes conditional predictive distributions rather than fitting a single model.
Why It Matters
Social Security solvency dominance: Mortality forecasting has a clear historical trend (Lee-Carter k(t) declining linearly) with bounded variance around a well-identified factor. Fertility forecasting inherits both ARMA variance and F∗ structural uncertainty. Over 75-year horizons, the structural F∗ uncertainty makes fertility the dominant source of Long-Term Actuarial Balance uncertainty — reversing SSA's own ordering (which ranks mortality first). See Lee and Tuljapurkar 1998 — Uncertain Demographic Futures and Social Security Finances.
Scenario-based forecasting is inferior: Scenarios specify a deterministic F∗ trajectory and ignore stochastic fluctuations entirely. Lee (1996) documents that Census Bureau long-term fertility levels are strongly anchored to recent experience, systematically understating long-run uncertainty. Stochastic intervals subsume the information in scenarios while also quantifying the probability of outcomes between and beyond scenario bounds.
Implications for population policy: If structural F∗ uncertainty is large and not reducible by more data, policies targeting fertility (immigration, pronatalism, work-family reconciliation) may have larger but more uncertain long-run Old-Age, Survivors, and Disability Insurance (OASDI) effects than policies targeting mortality (preventive care, chronic disease management), whose uncertainty is better bounded.
Open Questions
Can structural F∗ uncertainty be reduced by incorporating information from longer European or East Asian TFR records, where data extends further back and fertility has varied more?
Should the F∗ prior be updated as demographic and institutional conditions change (urbanization, women's labor force participation, immigration)? If so, how?
Does the Bayesian empirical prior approach systematically underestimate or overestimate future fertility range, given that the 1917–1992 sample includes a historically unusual baby boom that may not recur?
How should forecasters communicate structural vs. parameter uncertainty to non-technical policy audiences who may not distinguish types of uncertainty?