Carvalho-Polson-Scott (2010) The Horseshoe Estimator for Sparse Signals

horseshoeshrinkagesparsityscale-mixturebayesiannormal-meansrobustnessthresholding

Summary

This is the paper that introduced the horseshoe prior, a continuous global–local shrinkage prior for the sparse normal-means problem (yθN(θ,σ2I)y\mid\theta\sim\mathcal N(\theta,\sigma^2 I) with θ\theta believed sparse). Each mean gets its own local scale with a half-Cauchy hyperprior, giving a shrinkage profile that is spiked at zero (killing noise) and heavy-tailed (leaving genuine signals almost untouched). Carvalho, Polson, and Scott motivate the estimator, prove two theorems — one characterizing its tail robustness via a score-function representation, the other establishing a super-efficient convergence rate to the true sampling density in sparse settings — and show empirically that the horseshoe closely reproduces the answers of Bayesian model averaging under a "gold-standard" point-mass (spike-and-slab) mixture, but with a fully continuous, hyperparameter-free, easy-to-sample prior.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"This paper proposes a new approach to sparsity, called the horseshoe estimator, which arises from a prior based on multivariate-normal scale mixtures."

"The half-Cauchy prior on λi\lambda_i implies a horseshoe-shaped Be(1/2,1/2)\mathrm{Be}(1/2,1/2) prior for the shrinkage coefficient κi\kappa_i. The left side of the horseshoe, κi0\kappa_i\approx0, yields virtually no shrinkage… and describes signals. The right side… yields near-total shrinkage and describes noise."

My Take

The horseshoe's staying power comes from getting two things right at once that earlier shrinkage priors traded off: it shrinks noise as aggressively as a spike-and-slab (an unbounded spike at the origin) while shrinking true signals as gently as a flat prior (Cauchy tails), and it does so with a continuous, tuning-free prior that drops into a Gibbs or HMC sampler. The two theorems are the paper's spine — they turn "heavy tails and a spike at zero are good" into precise statements about the score function and the convergence rate — and they explain why the horseshoe became the default continuous competitor to the lasso and spike-and-slab. For the wiki this is the primary source the horseshoe-prior page was built around; it also connects the modern global–local literature back to the classical shrinkage tradition, now with a prior engineered specifically for sparsity rather than a common mean. The practical caveats it does not fully resolve — funnel geometry in sampling, no exact zeros, sensitivity to the global scale τ\tau — are what the later horseshoe+/regularized-horseshoe papers address.