McCulloch, Polson, and Rossi (2000) develop the "ID prior" — a prior placed directly on the identified parameter space of the multinomial probit (MNP) model — as an alternative to the McCulloch-Rossi (1994) "NID" approach, which works in the unidentified space and reports marginal posteriors. The key device is reparameterising the covariance matrix Σ so that σ11=1 with probability one, then assigning independent Normal and inverse-Wishart priors to the resulting free parameters. The posterior is computed by a four-block Gibbs sampler with no tuning parameters. The gain is interpretable, informative priors on the identified coefficients; the cost is higher chain autocorrelation relative to the NID sampler.
Key Claims
Identification: Parameters (β,Σ) are not identified because Y(cW)=Y(W) for all c>0; fixing σ11=1 achieves identification. NID approach (McCulloch-Rossi 1994) uses a proper prior on the full space and marginalises; ID approach enforces the constraint in the prior itself.
ID reparameterisation: Write ε=(U,Z′)′ where U=ε1, Z=(ε2,…,εp−1)′. Define γ=E(UZ) and Φ=ΣZ−γγ′/σ11 (conditional variance of Z∣U, the Schur complement). Then Σ=[1γγ′Φ+γγ′] and the one-to-one map (σ11,γ,Φ)↔Σ allows a prior on {Σ∣σ11=1} via γ∼N(γˉ,B−1) and Φ−1∼W(κ,C).
ID Gibbs sampler (four blocks, all conjugate, no tuning):
β∣γ,Φ,{Wi},D∼ multivariate Normal
Wij∣β,γ,Φ,{Wi(−j)},D∼ truncated univariate Normal
γ∣Φ,β,{Wi},D∼ Normal (regression of Z on U)
Φ∣γ,β,{Wi},D∼ inverted Wishart
Analytical result (NID prior, β): Under the NID prior β~∼N(bˉ,A−1), the identified β=β~/σ~11 follows v1.2χ⋅β~ where χ2∼χν−p+22; heavy-tailed when bˉ=0, skewed when bˉ=0. β and Σ are not independent under the NID prior.
Analytical result (NID prior, γ): Under the NID prior, the identified γ=σ21 follows a multivariate-t with ν−p+3 degrees of freedom, E(γ)=(V22)−1v21, Var(γ)=v1.2(ν−p+1)−1(V22)−1.
Analytical result (NID prior, Φ): Φ−1=W/ω where W∼Wishart(ν,(V22)−1) and ω∼v1.2χν−p+22, independent.
Prior assessment for E(Σ)=I: Setting E(γ)=0 and Δ=I reduces the ID prior to two scalar choices: κ (Wishart d.f.) and τ (variance of γ); large κ and small τ tighten the prior toward I. Recommended default: κ=p+2, τ=1/8.
Mixing trade-off: ID sampler autocorrelations are higher than NID. More diffuse Φ prior → slower mixing. In p=3 example, lag-10 autocorrelation of σ22 is 0.71 (proper ID) vs. 0.77 (improper ID) vs. lower for NID. Improper Φ prior can cause sampler to get stuck in p=6 case when likelihood is not informative enough.
Eigenvalue pathology: The improper (κ=0) Φ prior acts as a highly informative prior on the smallest eigenvalue of Σ, pushing it toward zero. Confirmed by simulating from N5(0,I) with n=7 observations: posterior of smallest eigenvalue is dominated by prior (peaks near 0.02) despite near-identity truth.
Advantages over NID: (a) Truly informative priors on β are straightforward (just N(βˉ,D−1)); (b) flat improper prior on β is allowed; (c) naturally accommodates hierarchical models βj∼N(βˉ,Vβ) since priors on identified parameters are interpretable; (d) analytical prior assessment results are available (unique to this parameterisation).
"Our new approach places a prior directly on the identified parameter space. The key is the specification of a prior on the covariance matrix so that the (1,1) element is fixed at 1 and it is possible to draw from the posterior using standard distributions." (Abstract)
"The more diffuse the prior on Σ is, the slower the autocorrelations die out." (p. 11)
My Take
The paper's main contribution is resolving a genuine awkwardness in McCulloch-Rossi (1994): when the prior is placed on the unidentified space, the induced prior on the identified parameters has an unusual form (χ2× Normal), making informative prior specification opaque. The ID approach trades this opacity for slower mixing — a reasonable engineering trade-off when the researcher has meaningful prior information (e.g., in a hierarchical model). The eigenvalue analysis in Section 7 is particularly useful: it explains concretely why the improper Φ prior is dangerous and provides a diagnostic (simulate with tiny n, inspect posterior of smallest eigenvalue). The paper has limited empirical content; the Imai-van Dyk (2005) marginal augmentation approach later showed that the NID chain can be dramatically accelerated, partially closing the gap.