Summary
Gelman, Shor, Bafumi, and Park resolve a paradox in U.S. electoral geography: rich states vote Democratic while rich individuals vote Republican — both facts are simultaneously true and not contradictory. The resolution comes from a varying-intercept, varying-slope multilevel logistic regression: the income–vote slope is itself a function of average state income, being steepest in poor states (Mississippi) and near-zero in rich states (Connecticut). The pattern emerged after 1992 and is robust across multiple datasets and controls.
Key Claims
- The paradox: Aggregate (state-level) income → Democratic vote; individual-level income → Republican vote. The ecological fallacy makes these appear contradictory but they coexist in the two-level model.
- Varying-intercept model (eq. 1): Pr(yi=1)=logit−1(αs[i]+βxi) — constant income slope β across states, varying intercepts αs capture baseline state-level partisanship.
- Varying-intercept, varying-slope model (eq. 2): Pr(yi=1)=logit−1(αs[i]+βs[i]xi) with state-level regressions αs=aα+bαus+εsα, βs=aβ+bβus+εsβ. Key result: b^β<0 — income slope is steepest (most Republican) in poor states and near-zero in rich states.
- County-level varying-slope model (§2.4): yc=αs[c]+βs[c]xc+εc fits county-level data with state-level hierarchical priors on αs and βs; confirms the individual-level pattern using geographic aggregates.
- Time dynamics: The varying-slope pattern emerged circa 1992–1996 (Figures 14–16); it was essentially absent 1968–1988. The red-blue geographic divide is a post-Cold-War phenomenon.
- Data: National Election Studies (NES) 1952–2000 with 5-point income scale re-coded to {−2,−1,0,1,2} centered at state median; Annenberg 2000 survey (≈ 100,000 respondents) for precision; exit polls 2000 and 2004 for robustness.
- Race controls: Excluding African-Americans preserves the varying-slope pattern — racial composition does not explain the state-level income gradient.
- Gini index: Coefficient ≈0 — state-level income inequality does not explain between-state heterogeneity in voting behavior.
- "Secret weapon" (fn. 10): Repeated cross-sectional models stacked over election years; time-series plots of parameter estimates reveal structural change circa 1992.
- Cognitive psychology explanations: Typicality bias (perceiving one's state's median voter as more extreme than it is) and second-order availability bias (overweighting salient examples) as mechanisms for why individuals hold misperceptions about cross-state income–vote patterns.
- Estimation: WinBUGS via R (Spiegelhalter et al. 1994/2002); model checking via binned residual plots (Gelman et al. 2000 Applied Statistics 49: 247–268).
Concepts Introduced or Extended
- Bayesian Hierarchical Model — varying-intercept, varying-slope multilevel logistic regression applied to political science
- Random Coefficient Model — state-level income slopes βs modeled as a linear function of average state income us
- Ecological Fallacy — individual-level and aggregate-level regressions can have opposite signs without contradiction
Entities Mentioned
Quotes
"Rich individuals are more likely to vote Republican, but rich states are more likely to vote Democratic. We resolve this apparent paradox using multilevel modeling." (abstract, paraphrased)
"The income-voting relationship, though consistently positive at the individual level, is much stronger in poor states like Mississippi than in rich states like Connecticut." (§2.2, paraphrased)
My Take
The paper's key methodological contribution is using the varying-slope multilevel logistic model not just as a predictive device but as a diagnostic for ecological fallacies — showing explicitly that aggregate and individual regressions address different quantities. The time-dynamics finding (pattern emerging post-1992) is more interesting than the point estimate and deserves more attention than it gets. The cognitive psychology mechanism section is speculative and unfalsified. The "secret weapon" is actually a powerful general technique for any repeated cross-section with structural change. Connection to Random Coefficient Model literature is implicit but not cited.
Publication note: Published in Quarterly Journal of Political Science 2(4): 345–367 (2007); the working paper circulated in November 2005. The file slug retains -2005 for stability.
References Extracted
- Gelman, A., A. Jakulin, M.G. Pittau, and Y.-S. Su. (2008). "A Weakly Informative Default Prior Distribution for Logistic and Other Regression Models." Annals of Applied Statistics.
- Gelman, A. and J. Hill. (2007). Data Analysis Using Regression and Multilevel/Hierarchical Models. Cambridge: Cambridge University Press.
- Gelman, A., J.B. Carlin, H.S. Stern, and D.B. Rubin. (2003). Bayesian Data Analysis, 2nd ed. Boca Raton: Chapman & Hall/CRC.
- Gelman, A., G. King, and C. Liu. (2000). "Not Asked and Not Answered: Multiple Imputation for Multiple Surveys." In Applied Statistics 49: 247–268.
- McCarty, N., K.T. Poole, and H. Rosenthal. (2006). Polarized America: The Dance of Ideology and Unequal Riches. Cambridge: MIT Press.
- Raudenbush, S.W. and A.S. Bryk. (2002). Hierarchical Linear Models: Applications and Data Analysis Methods, 2nd ed. Thousand Oaks: Sage.
- Robinson, W.S. (1950). "Ecological Correlations and the Behavior of Individuals." American Sociological Review 15: 351–357.
- Spiegelhalter, D.J., A. Thomas, N.G. Best, and W.R. Gilks. (1994). BUGS: Bayesian Inference Using Gibbs Sampling. Version 0.30. Cambridge: MRC Biostatistics Unit.
- Tversky, A. and D. Kahneman. (1974). "Judgment under Uncertainty: Heuristics and Biases." Science 185: 1124–1131.