Summary
Geweke and Keane (1999) generalise binary probit by replacing the N(0,1) disturbance with a finite mixture of m normals, p(εt)=∑j=1mpjhj1/2ϕ(hj1/2(εt−αj)), obtaining a model that is fully parametric and Bayesian yet near-semiparametric as m grows. Posterior simulation uses a six-block conjugate Gibbs sampler augmented with latent utility values and component-assignment indicators; marginal likelihoods are computed via the Gelfand-Dey (1994) estimator applied to the observed-data parameter space. Applied to female labor force participation (LFP) in the Panel Study of Income Dynamics (PSID; T=1,555), Bayes factors (BF) of order 107 favor the four-component scale mixture over standard probit, and the preferred model implies materially different effects of welfare benefits and spouse income.
Key Claims
- Model. y~t=β′xt+εt, dt=1(y~t>0); disturbance has the m-component mixture p(εt)=∑j=1mpjhj1/2ϕ(hj1/2(εt−αj)). Three identified variants: (a) full mixture — both αj and hj free, one hj∗=1, labeling restriction on either α or h; (b) scale mixture — αj=0, labeling restriction on h; (c) mean mixture — hj=1, labeling restriction on α (Section 2).
- Six-block conjugate Gibbs. Augment with latent utilities y~t and component assignments Lt∈{1,…,m}. Full conditionals: (1) y~t∣dt,Lt,β,α,h — truncated Normal on (0,∞) if dt=1, (−∞,0) if dt=0, mean αLt+β′xt, precision hLt; (2) Lt∣y~t,β,α,h,p — multinomial with P(Lt=j)∝pjexp[−0.5hj(y~t−αj−β′xt)2]; (3) β∣y~,L,α,h — Normal (precision-weighted regression eq. above eq. 2.6); (4) α∣y~,L,β,h — Normal with ordering inequality constraint, draws via Geweke (1991)/(1995) truncated-Normal algorithm; (5) hj∣y~,L,β,α for j=j∗ — sj2hj∼χ2(νj+Tj), sj2=sj,02+∑t∈Lj(y~t−αj−β′xt)2; (6) p∣L — Dirichlet with updated counts rj+Tj (eqs. 2.16–2.18, pp. 5–6).
- Marginal likelihood. Uses the Gelfand-Dey (1994) estimator extended by Geweke (1997b): M^−1=R−1∑rf(θ(r))/[p(d∣θ(r),X)π(θ(r))] where f(⋅) is a multivariate normal density centered at posterior mean, truncated to a highest-density region; applied to the observed-data space (β,α,h,p) with likelihood eq. (2.15) (marginalizing y~ and L analytically); normalizing constants for truncated labeling-restriction priors found by Monte Carlo (MC) (accuracy ∼10−3 in log units) (pp. 7–8).
- Artificial data (T=2,000, 5 data-generating processes, DGPs; Table 3): Conventional probit favored BF=16 over two-mixture for normal DGP; BF against conventional probit ≈1.45×1012 for scale-mixture DGP; BF ≈6.78×1033 for Cauchy DGP. Logit DGP indistinguishable from probit (BF ≈1/5): N(0,1) approximates the logistic density to within a Bayes factor below 10.
- PSID female LFP (T=1,555, 17 covariates, 80% participation; Tables 6–8): Conventional probit log-M=−566.6; best model (scale mixture of 4 normals) log-M=−548.3; ΔlogM=18.3 (BF≈107 against conventional probit); scale mixture dominates full mixture; covariate posteriors broadly similar across models but predictive probabilities differ substantially: Aid to Families with Dependent Children (AFDC)/food stamps raise nonparticipation probability by 13.1 percentage points, p.p. (probit) vs. 8.3 p.p. (scale mixture); participation probability as a function of spouse income declines faster in scale mixture at middle values.
Concepts Introduced or Extended
- Mixture of Normals — finite mixture of normals probit: component-assignment augmentation Lt as Gibbs block; scale and full mixture variants; scale mixture as finite approximation to Student-t (Geweke 1993)
- Binary Probit — non-normal disturbance extension: replaces fixed N(0,1) shock with a m-component mixture while retaining the Albert-Chib (1993b) latent utility augmentation for y~t
- Gibbs Sampler — new six-block conjugate sampler with dual data augmentation (latent utilities + component assignments); Gelfand-Dey marginal likelihood on observed-data parameter space
Entities Mentioned
Quotes
"In this application, Bayes factors strongly favor mixture of normals probit models over the conventional probit model, and the most favored models have mixtures of four normal distributions for the disturbance term." (Abstract)
"The data are sufficiently informative about the nature of the distribution of the shock that the most preferred model has four mixtures with six free parameters in the distribution." (p. 15)
My Take
The paper's key contribution is the combination of two devices: Albert-Chib (1993b) latent utility augmentation for the binary indicator, and Chib (1996)-style latent component-assignment augmentation Lt for the mixture — yielding a six-block fully conjugate sampler where every block is a standard distribution. The PSID application is convincing: 1,555 observations are enough to decisively reject standard probit in favor of a flexible four-component scale mixture. The finding that the scale mixture (symmetric, heavy-tailed) dominates the full mixture (asymmetric) suggests that the main failure of standard probit in female LFP is thin tails, not skewness. The Gelfand-Dey marginal likelihood computation is clever — marginalizing over the augmented latent variables to avoid the curse of dimensionality in the high-dimensional space — but requires an extra MC step for normalizing truncated-prior constants. The explicit connection to Geweke (1993) is an important intellectual lineage: the scale mixture probit (αj=0) with m→∞ converges to the Student-t probit, so this paper operationalizes a finite-dimensional parametric approximation to that limiting nonparametric case.
Originally circulated as Federal Reserve Bank of Minneapolis Research Department Staff Report 237 (1997); published as Geweke, J. and M. Keane (1999), "Mixture of Normals Probit Models," in C. Hsiao, K. Lahiri, L.-F. Lee, and M.H. Pesaran (eds.), Analysis of Panels and Limited Dependent Variable Models, Cambridge University Press, pp. 49–78 (the citation of record).