Summary
This is the May 1998 working paper manuscript that eventually appeared as Albert and Chib (2001) in Biometrics. Same title, same hospital length-of-stay application (N=1,000, J=12), but substantially longer (28 pages vs. 8) and containing material cut from the published version. Key additions over the published paper are: (i) a hierarchical sequential model (Algorithm 2) that shrinks the free cutpoints toward a quadratic polynomial surface; (ii) a 6-model comparison table with Weibull and log-logistic continuous-time survival models as additional alternatives; and (iii) full derivations of the cumulative ordinal Markov chain Monte Carlo (MCMC) algorithm (Algorithm 3) with the log-spacing reparameterization.
Key Claims
- Basic sequential model (Algorithm 1): Two-block Gibbs sampler identical to Albert-Chib (2001); the re-parameterized design vector xij′=(0,…,−1,0,…,xi′) absorbs cutpoints into β=(γ1,…,γJ−1,δ′)′; censored observations contribute only the "passed through" truncated-normal draws.
- Interaction model: Category-specific covariates xi(1) may have time-varying effects δj; their coefficients lie on a low-order polynomial δjl=ηl0+ηl1j+ηl2j2 via the Kronecker construction Aj=Im⊗(1jj2); the full model fits as Algorithm 1 with an expanded design matrix.
- Quadratic polynomial baseline (M4): Restricts γj=ϕ0+ϕ1j+ϕ2j2; parameters ϕ=(ϕ0,ϕ1,ϕ2) replace the J−1 free cutpoints; still fits via Algorithm 1 with modified Xi.
- Hierarchical sequential model (Algorithm 2, not in published paper): Cutpoints follow γj=ϕ0+ϕ1j+ϕ2j2+εj with εj∼iidN(0,τ2); τ2∼IG(c,d) (Inverse-Gamma); as τ2→0 collapses to M4, as τ2→∞ to basic M2; four-block Gibbs (z, β, β0, b); hierarchical posterior cutpoints interpolate between M2 and M4 values.
- Cumulative ordinal model (Algorithm 3): Applies the log-spacing reparameterization α1=logγ1c, αj=log(γjc−γj−1c) (Eq. 10) from Albert-Chib (1997b) to lift ordered cut-points to an unrestricted vector; multivariate-t Metropolis-Hastings (MH) proposal centred at mode α^ with dispersion τ2V.
- Weibull and log-logistic models (§4.2): h(t)=αλ(λt)α−1, S(t)=exp(−(λt)α) for Weibull; h(t)=αλ(λt)α−1/[1+(λt)α], S(t)=1/[1+(λt)α] for log-logistic; λi=exp(−xi′δ); posterior sampled via MH (Chib-Greenberg 1995).
- 6-model comparison (Table 2, training sample n0=200, evaluation n=800): M1 cumulative (lnm=−2172.8), M2 basic sequential (−2127.2), M3 sequential+comorbidity index (COMORB) interaction (−2145.0), M4 sequential+quadratic baseline (−2121.2), M5 Weibull (−2230.4), M6 log-logistic (−2255.6); Bayes factor BF21=exp(45.6)=6.4×1019; best model M4 reduced — race (RACE), private insurance (PRIVINS), COMORB, coronary artery bypass graft (CABG), percutaneous transluminal coronary angioplasty (PTCA) — lnm=−2117.4.
- Note on numbers: The manuscript's log marginal likelihoods differ from the published Albert-Chib (2001) values (e.g., M4 full: −2121.2 here vs. −2094.0 published) because the training/evaluation split differs between manuscript and published versions.
- Prior elicitation: Training sample of size n0=200 used to estimate hyperparameters (β0,B0) of the multivariate normal prior; remainder (n=800 for evaluation) used in model comparison; same training-sample approach as Albert-Chib (1997b).
Concepts Introduced or Extended
Entities Mentioned
Quotes
"The sequential model is useful in the analysis of discrete-time survival data." (p. 2)
"Although this was not discussed in the paper, [these methods] produced MCMC output that mixed remarkably well." (p. 26)
My Take
The main substantive addition over the published 2001 paper is the hierarchical sequential model of Section 3.2, which elegantly bridges the free-cutpoint model (M2) and the quadratic-polynomial model (M4) via a single variance hyperparameter τ2. This is genuinely useful when neither extreme seems right: the hierarchical shrinkage produces cutpoints that are smoother than M2 but not forced onto a rigid polynomial. The hierarchical posterior means in Table 4 confirm the interpolation is working as designed. The 6-model comparison table in the manuscript also gives clearer context than the published paper — seeing all six log marginal likelihoods side by side makes the hierarchy M4>M2≫M1>M3≫M5>M6 immediately legible. The differences in marginal likelihood values between the manuscript and published paper (due to different training/evaluation splits) are worth keeping in mind: the manuscript uses n0=200 out of 1,000, while the published paper apparently uses a different split.