Afshartous-de Leeuw (2005) Prediction in Multilevel Models

multilevel-modelmixed-effectspredictionrandom-effectsbayesianshrinkagehierarchical-modelsimulationempirical-bayesrandom-coefficientsintraclass-correlation

Summary

Afshartous and de Leeuw (2005) fill a gap in the hierarchical linear model (HLM) literature by treating prediction of new individual observations yjy_{*j} as a problem distinct from estimation of random effects. They define three prediction rules — Ordinary Least Squares (OLS), prior, and multilevel (shrinkage) — compare them analytically and through a large simulation (30,000 datasets), and show that the multilevel predictor dominates in 24 of 25 design conditions. A key asymmetry emerges: increasing the within-group sample size njn_j benefits prediction more than increasing the number of groups JJ, while the opposite holds for parameter estimation.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"We define the multilevel predictor as the Bayes rule under squared error loss and show that it equals the best linear predictor."

"We find that the multilevel predictor outperforms the OLS and prior predictors in nearly all of the conditions investigated."

"Increasing the within-group sample size nn was more beneficial for prediction, while increasing the number of groups JJ was more beneficial for estimation."

My Take

The paper's main contribution is clarifying that prediction of new observations in hierarchical models is a different problem from estimating random effects — a distinction routinely glossed over in applied multilevel work. The Bayes-optimality proof (Proposition 1.1) is clean and important: it gives a theoretical basis for preferring the multilevel predictor that goes beyond simulation evidence. The study-design asymmetry (njn_j for prediction, JJ for estimation) is a practically useful rule of thumb. The main limitation is that all analytical results assume variances are known; estimation uncertainty in τ^2\hat\tau^2 is handled via the bias–shrinkage formula but is not propagated through a full Bayesian treatment. In practice, the bias formula shows the exposure is real: modest underestimation of Level 2 variance (common in restricted maximum likelihood in small JJ settings) can substantially over-shrink.