Blattberg-George (1991) Shrinkage Estimation of Price and Promotional Elasticities

shrinkagehierarchical-modelempirical-bayesgibbs-samplerseemingly-unrelated-regressionmarketingpoolingrandom-coefficient-modelforecasting

Summary

This paper tackles the "aggregation dilemma" faced when estimating price and promotional elasticities from supermarket scanner data at the chain-brand level: a single pooled model is too restrictive because elasticities genuinely vary across chains and brands, yet a separate ordinary least squares (OLS) model per chain-brand yields noisy, unstable, and frequently wrong-signed estimates. The authors resolve it with shrinkage estimators derived from a hierarchical Bayes model that borrows strength across chain-brands — pulling each unit's coefficients toward a common mean while still permitting genuine heterogeneity. Applied to a large bathroom-tissue scanner dataset, the shrinkage estimators deliver more reasonable (correctly signed) elasticities than OLS with no loss — indeed an improvement — in out-of-sample predictive power.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"By borrowing strength across chains and brands, these procedures reduce variability while providing flexibility that allows for separate elasticity estimates."

"Modeling sales at the chain-brand level often leads to counterintuitive and theoretically unreasonable estimates of separate elasticities… frequently having the wrong sign."

My Take

The paper is a clean, early applied demonstration that a hierarchical prior is the right cure for the bias-variance bind of panel-of-regressions estimation — exactly the random-coefficient / SUR setting where each unit has too little data to stand alone but the units are not identical. Its lasting interest is methodological rather than substantive: the same "shrink OLS toward a common mean, estimate the shrinkage from the data" logic underlies the Minnesota prior for VARs and the Black-Litterman treatment of return forecasts, and the paper is notable for using the then-new Gibbs sampler to make the fully hierarchical-Bayes version computable. The caveats are the usual empirical-Bayes ones — the variance-component estimates that drive the shrinkage are themselves noisy, and the wrong-sign diagnosis presumes the theory-predicted signs are correct.