Chen-Ibrahim-Yiannoutsos (1999) Prior Elicitation, Variable Selection and Bayesian Computation for Logistic Regression Models

variable-selectionlogistic-regressionbayesianpower-priormodel-selectiongibbs-samplermarginal-likelihoodinformative-priorbiostatistics

Summary

The paper develops a fully Bayesian approach to variable selection in logistic regression, tackling the three obstacles that make Bayesian model selection hard: (1) specifying priors for the regression coefficients in every candidate model, (2) specifying a prior over the model space, and (3) computing the analytically intractable prior and posterior model probabilities. For (1) it proposes a class of informative priors elicited semi-automatically from a prior prediction y0y_0 for the response and a scalar precision a0a_0 quantifying confidence in that prediction (a historical-data / power-prior-style construction). For (2) it gives a matching informative prior on the model space. For (3) it develops novel algorithms that compute all prior and posterior model probabilities from Gibbs samples of a single model, with theoretical properties derived. The methods are illustrated on cancer and AIDS clinical-trial data.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"We examine the problem of eliciting informative prior distributions for the regression parameters as well as the elicitation of an informative prior distribution for the model space for Bayesian selection of variables in logistic regression."

"Another major contribution of this paper is to develop novel computational algorithms for computing analytically intractable prior and posterior model probabilities. These algorithms are efficient and only require Gibbs samples from a single model."

My Take

This is the logistic-regression companion to the Ibrahim-Chen programme on informative priors from historical data, and it pairs naturally with the Chen-Ibrahim-Sinha (1999) cure-rate paper. The clever moves are (a) shifting elicitation from the opaque coefficient space to an interpretable prediction y0y_0 with a single confidence knob a0a_0, and (b) computing every submodel's marginal probability from one model's MCMC output — a big practical saving when the model space is large. It sits outside the wiki's time-series core but connects to the existing variable-selection (SSVS / George-McCulloch) and power-prior material, and reuses the same Gibbs/normalizing-constant machinery (Chib 1995, bridge/importance sampling) that the econometrics side relies on.