Summary
A 21-page UChicago working paper asking whether Bayesian shrinkage of individual parameters also improves estimation/prediction of their aggregate (sum or total). Motivated by Zellner-Chen (2001), which shrinks sector-level output forecasts to produce an aggregate GDP forecast, and a question by J.K. Ghosh (2001) about the resulting aggregate precision. The answer: under many conditions IT PAYS TO SHRINK for totals, but not always — the key distinction is whether shrinkage targets the overall level (Stein's zero-mean prior) or only cross-sectional deviations (Lindley's free-mean extension).
Key Claims
n-means problem (yi=θi+ui; total T=ι′θ; two loss functions: quadratic and balanced):
- Least squares (LS) / diffuse Bayes / Bayesian method of moments (BMOM), no conceptual sample: θ^=y, T^=ι′y — no shrinkage in either individuals or total; pathological "perfect fit" (u^=0) with n parameters and n observations.
- Stein's zero-mean prior (θi∼iid N(0,τ2)): posterior mean Eθ∣D=[1−1/(1+τ2)]y; posterior total ET∣D=[1−1/(1+τ2)]ι′y — total is also shrunk toward zero. Stein's estimator plugs in τ^2=y′y/n−1 to get θ^=(1−n/y′y)y.
- Lindley's free-mean extension (θi=μ+ei, μ uniform): posterior mean shrinks deviations from yˉ, i.e., Eθ∣D≈yˉι+[1−(n−3)/(y−yˉι)′(y−yˉι)](y−yˉι); but total T^=ι′y unshrunk — the level is not shrunk, only cross-sectional spread.
- BMOM with conceptual sample z=θ+v: Et′θ∣D=(1/2)(t′y+t′z)=[1−(1/2)(1−zˉ/yˉ)]ι′y — shrinkage estimate of total without likelihood or prior density, requiring only moment assumptions; optimal under quadratic loss.
- Model averaging (posterior odds K between Stein's H0 and Lindley's Ha): T^=[K/(1+K)]T^0+[1/(1+K)]T^a — weighted combination that interpolates between shrunk and unshrunk totals.
- Balanced loss function (L=w⋅fit+(1−w)⋅precision): optimal T^∗=wι′y+(1−w)ET∣D — total involves shrinkage whenever Eθ∣D=y.
Seemingly unrelated regression (SUR) extension (yα=Xαβα+uα, α=1,...,m; total T=J′β):
- Stein-type prior f(β)∼N(b,Ω−1): posterior mean Eβ∣D=(Z′Z+Ω)−1(Z′Zβ^+Ωb); total shrinks toward prior sum: J′Eβ∣D=J′b+J′(Z′Z+Ω)−1Z′Z(β^−b).
- BMOM with conceptual sample (g-prior Zc′Zc=gZ′Z): Eβ∣D=βˉc+[1/(1+g)](β^−βˉc) — shrinkage toward prior mean; predicting the future total Tf=ι′yf incorporates same shrinkage.
Risk dominance (Lemmas 1–2): If a shrinkage estimator θ^s uniformly dominates θ^ under quadratic loss, AND both mean squared error (MSE) matrices E(θ^s−θ)(θ^s−θ)′ and E(θ^−θ)(θ^−θ)′ are diagonal, then T^s=ι′θ^s uniformly dominates T^=ι′θ^. Same holds for prediction. Sufficient condition: independent estimation errors across equations.
Concepts Introduced or Extended
Entities Mentioned
Quotes
"As will be seen, for the loss functions employed, there is evidence that IT PAYS TO SHRINK in estimating totals in many circumstances, but not all."
"The result that shrinkage can affect the precision of forecasts of totals has implications for properties of 'consensus' forecasts of totals derived from surveys of individuals' anticipated values."
My Take
The Stein-vs-Lindley contrast is the most useful finding: shrinkage toward a fixed prior mean shrinks the total; shrinkage only of cross-sectional deviations (Lindley) leaves the total alone. This has an intuitive implication for macro aggregation: if you believe sector growth rates are exchangeable around a common unknown mean, bottom-up aggregation of shrinkage forecasts recovers the unshrunk total. The diagonality condition in Lemmas 1–2 is the key caveat — SUR-type cross-equation correlation can break the dominance result. Primarily of interest as a companion to the Zellner-Chen (2001) GDP forecasting application; the theoretical framework is clean but the empirical content is thin.