Summary
The paper develops a Bayesian model for toxicological multivariate mortality data with large families, where the discrete mortality (hazard) rate for each family at each time point depends on shared familial random effects and the toxicity level the family was exposed to. The model is a time-dependent random-effects logistic (discretized) hazard regression. Because a similar earlier experiment (with a related toxicant, NaSCN) exists, its data are used to build an informative power prior for the current (KSCN) study, with a weight a0 controlling how much the historical data count. Inference is by data augmentation / Gibbs sampling; the analysis incorporates Bayesian model diagnostics, variable subset selection (to decide the functional form of the time effect), and predictive-distribution model comparison. Applied to O'Hara Hines's potassium-thiocyanate (KSCN) trout-fish-egg tank data.
Key Claims
- Random-effects discretized-hazard model. For each family (tank) k, the number dying in interval j is modeled with a discrete (grouped) hazard through a logistic link, logit of the interval hazard =g(tj)+xk′β+bk, where g(t) is a flexible time effect, xk holds the toxicant level and water-hardening covariate, and bk is a familial random effect capturing within-family correlation / overdispersion among the many eggs in a tank (extending grouped-survival regression, Prentice-Gloeckler 1978, to correlated large families).
- Time-varying random effects. The familial effects are allowed to evolve over time, giving a time-dependent random-effects structure richer than a single static frailty.
- Historical-data (power) prior from the NaSCN study. The earlier NaSCN experiment (D0) supplies an informative prior via a power-prior construction: the historical likelihood is raised to a0∈[0,1], so a0 tunes how strongly the previous study informs the KSCN parameters — a natural use of genuinely relevant prior data.
- Computation via data augmentation. Latent variables render the complete-data likelihood tractable, and a Gibbs sampler simulates the posterior of the regression coefficients, random effects, and mortality rates, plus derived quantities of interest.
- Variable selection and predictive model comparison. A Bayesian subset-selection procedure decides the form of the time effect g(t) (e.g. how many knots), and predictive distributions (an L-measure-type criterion) compare several plausible models, verifying assumptions about time and family heterogeneity.
Concepts Introduced or Extended
- Data Augmentation — latent variables for the discretized-hazard logistic model's Gibbs sampler
- Power Prior — informative prior from the historical NaSCN study, weight a0
- Variable Selection — Bayesian subset selection for the time-effect specification
- Bayesian Model Selection — predictive-distribution (L-measure) comparison of candidate models
- Gibbs Sampler — posterior simulation of coefficients, random effects, and mortality rates
Entities Mentioned
Quotes
"The discrete mortality rate for each family of subjects at a given time depends on familial random effects and the toxicity level experienced by the family."
"A similar previous study (using sodium thiocyanate (NaSCN)) is used to construct a prior for the parameters in the current study."
My Take
The third of the Chen-Ibrahim/Dey/Sinha cluster ingested here, and a nice applied showcase of the same toolkit: a discretized-hazard logistic model made hierarchical with familial random effects, fit by data-augmented Gibbs, with the historical-data power prior doing real work (a genuinely comparable prior experiment, exactly the setting power priors are built for). It combines the threads of the other two papers — the survival/hazard modeling and the Bayesian variable selection — in one applied analysis. Outside the wiki's time-series core, but it reuses the shared data-augmentation, Gibbs, and power-prior machinery, and rounds out the biostatistics corner.