The decision-theoretic companion to Brown-Vannucci-Fearn (1998). Instead of reporting the posterior over subsets, this paper chooses a subset of explanatory variables in multivariate regression by maximizing a utility that trades prediction accuracy against the cost of including variables — a multivariate extension of Lindley's (1968) decision approach under quadratic loss. It also replaces the pure natural-conjugate prior with a non-conjugate proper prior that adds a component of error "unexplainable by any number of predictors," deliberately avoiding the determinism (perfect prediction as ) that Dawid (1988) showed the conjugate prior implies. Simulated annealing with fast updating searches the huge subset space; the method is demonstrated on near-infrared spectroscopy with observations and predictors.
"Our approach balances prediction accuracy against costs attached to variables in a multivariate version of a decision theory approach pioneered by Lindley (1968)."
"[We employ] a non-conjugate proper prior distribution ... extending the standard normal-inverse Wishart by adding a component of error which is unexplainable by any number of predictor variables, thus avoiding the determinism identified by Dawid (1988)."
Where the 1998 paper gives you the posterior over subsets, this one answers the question a practitioner actually faces — which subset to use — by turning selection into a decision problem with an explicit cost for complexity, which is a cleaner justification for parsimony than a bare posterior mode. The genuinely interesting technical point is the Dawid determinism fix: the conjugate normal-inverse-Wishart prior secretly implies you can predict perfectly given enough predictors, an absurdity in chemometrics, and the remedy — an unexplainable-error component making the prior non-conjugate — is a good example of a modeling pathology visible only in high dimensions and worth flagging on the wiki's variable-selection page. The pair (1998 inference + 1999 decision, both with simulated-annealing/Metropolis search and fast updates) became a template for high-dimensional Bayesian regression in spectroscopy and later genomics, and its headline empirical lesson — that with a proper prior you can use more variables than observations and that jointly modeling correlated responses beats separate fits — is the enduring one.