Summary
This paper evaluates Lee's constrained autoregressive moving average (ARMA) model for stochastic US fertility forecasting, addressing three questions: how to validate stochastic forecast models, how sensitive forecasts are to the assumed long-run total fertility rate (TFR) average (F∗), and how Bayesian probability-weighting over alternative F∗ values changes prediction uncertainty. The central finding is that F∗ — a subjectively set constraint that operates on 30–50 year timescales — is not identifiable from short historical time series, creating an irreducible source of structural uncertainty distinct from parameter uncertainty. This structural uncertainty partly explains why fertility dominates long-horizon long-term actuarial balance (LTAB) uncertainty: see Long-Term Actuarial Balance.
Key Claims
- Lee's constrained ARMA model: (Ft−F∗)=a(Ft−1−F∗)+Ut+cUt−1, where F∗ is a fixed long-run average TFR set a priori. For the 1917–1992 US series, 95% prediction intervals reach width of one child within 6 years of the launch date; asymptotic width exceeds two children in about 18 years, after which interval width plateaus.
- No empirical attractor for US TFR: A plot of TFR at time t against TFR at time t−1 shows values clustering near the 45-degree line with no discernible attractor — no level toward which fertility trends. Long-run dynamics are thus not recoverable from the short-run sample.
- F∗ operates on a different timescale: Long-run average fertility likely changes on 30–50 year generational timescales, not the short-run timescale captured by ARMA estimation. Specifying F∗ far from the sample mean widens intervals via the (F∗−Fˉ)2 contribution to prediction variance. This is structural uncertainty (Draper 1995), not parameter uncertainty.
- Validation experiments (launch years 1945, 1955, 1965, 1975; F∗∈{1.8,2.1,3.0}): The 1945-launch forecast completely fails to capture the baby boom when F∗=1.8 or 2.1; only F∗=3.0 comes close. Later launches (1955 onward) capture subsequent fertility well across all F∗ values. Conclusion: sample-based uncertainty is insufficient when subsequent variation substantially exceeds variation within the historical sample.
- Expert myopia: The Social Security Administration (SSA) and US Census Bureau anchor F∗ near replacement-level (2.1), matching recent fertility but discounting the empirical history of wide swings. Lee (1996) documents that Census long-run levels are strongly influenced by the most recent fertility experience.
- Bayesian priors over F∗: Two alternatives tested. (1) Malthusian prior — flat weight on F∗∈[1.9,2.1] with exponentially falling weight at higher values — produces nearly identical intervals to the F∗=2.1 benchmark. (2) Empirical prior — weights proportional to historical TFR frequency (mean 2.6) — produces intervals about 15% wider and shifts the long-run forecast level up toward the historical mean. The modest 15% widening from even an extreme historical prior understates the true structural uncertainty, which grows with the forecast horizon.
- Asymptotic uncertainty plateaus: Unlike a random walk, the constrained ARMA model produces prediction intervals that plateau in width once the forecast horizon exceeds ~18 years. This plateau is a feature of the constraint F∗, not a limitation of the data — it reflects what the model can and cannot say about the long run.
- Information value of stochastic forecasts: Stochastic prediction intervals are strictly more informative than scenario-based methods, which ignore historical variation and do not provide probability assessments. Point predictions become essentially useless after 6 years given a one-child interval width.
Concepts Introduced or Extended
Entities Mentioned
Quotes
"We argue that the long-run average of fertility, a key assumption of the model, operates on a time scale not probed by time-series methods."
"In short, we have no reliable guide to the long-run behavior of expected fertility."
"Sample-based uncertainty is not always sufficient to describe future uncertainty and the variance in the base period tends to increase with the length of the base period."
"Traditional scenario-based methods for forecasting fertility have performed poorly, in part because they do not incorporate the historical levels of variation in the series into an estimate of forecast uncertainty, and in part because they myopically focus on recent fertility trends while slighting the longer term behavior of the series."
My Take
This paper's core contribution is sharpening the distinction between parameter uncertainty (estimable from data) and structural uncertainty (arising from model assumptions unidentifiable from data). For fertility, the long-run average F∗ is the structural unknown. Tuljapurkar and Boe show that stochastic time-series methods, however sophisticated, cannot eliminate F∗ uncertainty — they merely quantify the ARMA component of variation while leaving the long-run level assumption to expert judgment. This has a direct implication for Social Security finance: mortality forecasting benefits from a clear Lee-Carter (LC) trend with bounded variance, while fertility forecasting inherits both ARMA variance and F∗ structural uncertainty, making fertility the dominant uncertainty source at 75-year horizons. The paper's validation exercises are particularly valuable: the 1945-launch failure under F∗=1.8 and 2.1 is a concrete demonstration that choosing a "reasonable" long-run level near replacement can systematically fail to capture historically plausible demographic events like the baby boom. Expert anchoring to recent levels is not conservatism — it is myopia with distributional consequences.