Summary
Greenberg and Parks argue that model selection should be evaluated through its effect on the predictive distribution of observables rather than through significance tests on coefficients. They derive two diagnostic tools — a scaled mean shift and a generalized variance ratio (GVR) — comparing the predictive Student-t density of the full model versus a subset model, and apply them to the Fazzari-Hubbard-Petersen (FHP, 1988) investment model to show that cash flow (CF) is more central to predictions than Tobin's q, despite both being statistically significant.
Key Claims
- Predictive sufficiency over significance: A variable whose coefficient is highly significant in a large sample may shift predictions only trivially; comparing predictive densities directly provides an economically interpretable criterion.
- Predictive mean shift: Under a diffuse prior, the shift in predictive mean when X2 is dropped from y=X1β1+X2β2+ε is Λ′β^2 where Λ′=X02−X01δ^ measures how far the prediction point deviates from the collinearity hyperplane.
- Multicollinearity insight: Λ′=0 when the prediction point lies on the regression of X2 on X1, recovering the classical result that collinearity does not impair prediction if it persists into the forecast period.
- Generalized Variance Ratio: GVR = ∣Cov(y0∣X01)∣/∣Cov(y0∣X0)∣ factors into three terms: a degrees-of-freedom ratio (< 1), an R2-improvement term (expressible as the F-statistic, ≥ 1), and a prediction-location term (≤ 1). GVR ≈ 1 and a small mean shift together indicate X2 can be safely omitted.
- Overlap statistic: For in-sample diagnostics, computes the fraction of overlap between 95% highest posterior density (HPD) intervals under the two models; values near 1 signal predictive equivalence.
- Application: In the FHP investment model (443 firm-years), dropping CF shifts predictive means and raises GVR substantially more than dropping q; mean overlap when only CF retained: 0.87 vs. only q retained: 0.79. This supports cash flow as the more essential predictor.
- No prior model probabilities need to be specified, unlike Zellner's (1971) Bayesian model averaging.
Concepts Introduced or Extended
Entities Mentioned
Quotes
"there can be no practical value in adding a set of variables that leaves the predictive variance unchanged and changes the predictive mean only in the third decimal place of a variable that is observed to one decimal place, even if the regression coefficient is highly significant."
My Take
An elegant short paper that recasts the multicollinearity problem in Bayesian predictive terms. The key insight — that statistical significance and predictive relevance can sharply diverge — was already implicit in Bayesian reasoning but rarely stated so plainly. The GVR and overlap statistic are intuitive and computation is trivial (OLS plus quadratic forms). The main limitation is the dependence on diffuse priors; with informative priors the predictive covariance structure changes. Also: the approach works well for in-sample comparison at observed X values, but the choice of X0 for out-of-sample comparison still requires judgment.