This is the paper that introduced the horseshoe prior, a continuous global–local shrinkage prior for the sparse normal-means problem ( with believed sparse). Each mean gets its own local scale with a half-Cauchy hyperprior, giving a shrinkage profile that is spiked at zero (killing noise) and heavy-tailed (leaving genuine signals almost untouched). Carvalho, Polson, and Scott motivate the estimator, prove two theorems — one characterizing its tail robustness via a score-function representation, the other establishing a super-efficient convergence rate to the true sampling density in sparse settings — and show empirically that the horseshoe closely reproduces the answers of Bayesian model averaging under a "gold-standard" point-mass (spike-and-slab) mixture, but with a fully continuous, hyperparameter-free, easy-to-sample prior.
"This paper proposes a new approach to sparsity, called the horseshoe estimator, which arises from a prior based on multivariate-normal scale mixtures."
"The half-Cauchy prior on implies a horseshoe-shaped prior for the shrinkage coefficient . The left side of the horseshoe, , yields virtually no shrinkage… and describes signals. The right side… yields near-total shrinkage and describes noise."
The horseshoe's staying power comes from getting two things right at once that earlier shrinkage priors traded off: it shrinks noise as aggressively as a spike-and-slab (an unbounded spike at the origin) while shrinking true signals as gently as a flat prior (Cauchy tails), and it does so with a continuous, tuning-free prior that drops into a Gibbs or HMC sampler. The two theorems are the paper's spine — they turn "heavy tails and a spike at zero are good" into precise statements about the score function and the convergence rate — and they explain why the horseshoe became the default continuous competitor to the lasso and spike-and-slab. For the wiki this is the primary source the horseshoe-prior page was built around; it also connects the modern global–local literature back to the classical shrinkage tradition, now with a prior engineered specifically for sparsity rather than a common mean. The practical caveats it does not fully resolve — funnel geometry in sampling, no exact zeros, sensitivity to the global scale — are what the later horseshoe+/regularized-horseshoe papers address.