Antoniak (1974) Mixtures of Dirichlet Processes with Applications to Bayesian Nonparametric Problems

dirichlet-processnonparametric-bayesmixture-modelclusteringbayesianpolya-urnconcentration-parameterdensity-estimation

Summary

Antoniak extends Ferguson's (1973) Dirichlet process (DP) — a random probability measure whose sample paths are almost surely distributions, used as a prior over the space of distributions in Bayesian nonparametrics — to the case where the random measure is a mixing distribution for a parameter that governs the sampling distribution. In that setting the posterior for the random measure is no longer a single Dirichlet process but a mixture of Dirichlet processes (MDP). The paper defines MDPs formally, proves several properties (most importantly a closure property under conditioning on data), derives formulas for the posterior, and works applications in bio-assay, discrimination, regression, and estimation of a mixing distribution.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"This paper extends Ferguson's result to cases where the random measure is a mixing distribution for a parameter which determines the distribution from which observations are made."

"The conditional distribution of the random measure, given the observations, is no longer that of a simple Dirichlet process, but can be described as being a mixture of Dirichlet processes."

My Take

This paper is the theoretical hinge between Ferguson's (1973) Dirichlet process and the modern Dirichlet process mixture models that dominate Bayesian nonparametric clustering and density estimation. Two of its results do the heavy lifting downstream: the MDP closure property, which guarantees the posterior stays in a manageable class, and the distribution of the number of distinct clusters, which is exactly what lets practitioners reason about — and place a prior on — the concentration parameter MM that governs how many mixture components the data will support. Escobar–West's concentration-parameter sampler, MacEachern–Müller's and Neal's (2000) DP-mixture MCMC, and every "infinite mixture" application trace their tractability to the machinery Antoniak set down here. The paper is heavily measure-theoretic, but its practical residue — the Pólya-urn predictive rule and the E[k]Mlog(1+n/M)E[k]\approx M\log(1+n/M) scaling — is what working Bayesians actually use.