Definition
Systematic divergence between race and Hispanic-origin as recorded by funeral director proxies on death certificates (DC) (the numerator of all mortality rate calculations) and the race/ethnicity that the decedent would have self-reported. Because population denominators (census counts) are based on self-report while death certificate numerators rely on third-party observer coding, misclassification creates a numerator-denominator mismatch that biases group-specific mortality rates up or down depending on the direction of misclassification.
Key Ideas
- Three validation metrics: (1) Sensitivity — the fraction of people who self-identified as group G whose death certificates also code them as G; (2) Predictive Value Positive (PVP) — the fraction of death certificates coding group G whose decedents would have self-identified as G; (3) Classification Ratio (CR) — the ratio of self-identified population count to death-certificate-identified count (CR>1 means undercounting on the DC; CR<1 means overcounting).
- Correction formula: Corrected Age-Adjusted Death Rate (AADR): AADRi=AADRiraw×CRi. A CR of 1.30 for American Indian and Alaska Native (AIAN) means AIAN death rates should be multiplied by 1.30 — a 30% upward correction.
- White/Black: CR≈1.00; sensitivity ≈ 98–99%; no mortality bias. Race-specific mortality comparisons between these groups are not materially affected by misclassification.
- AIAN: CR=1.30; sensitivity ≈ 55% (most severe of all groups, persistent across decades). Correcting for misclassification reverses the apparent AIAN mortality advantage: raw AADR ratio vs. White =0.85 (apparent advantage) → corrected ratio =1.11 (real disadvantage). AIAN have the highest true mortality of any race-ethnicity (RE) group.
- Asian/Pacific Islander (API): CR=1.07; sensitivity = 90% and improving. Corrected AADR ratio vs. non-Hispanic White (NHW): 0.60→0.64 (small effect; API advantage is real).
- Hispanic: CR=1.05; corrected AADR = 83% of NHW (vs. 79% raw). The Hispanic mortality paradox is real and not a statistical artifact of DC misclassification.
- AIAN geography: CR ranges from 1.02 in high-coethnic-concentration states (AZ, NM, OK, MT, NE) to 1.63 in low-concentration states. Funeral directors rely on observable community cues; in states where AIAN residents are more visible, coding accuracy is higher.
- AIAN identity expansion paradox: Current Population Survey (CPS)/census AIAN self-identification grew >90% since 1960, driven by rising social acceptability of claiming mixed-ancestry AIAN identity — individuals who may not appear visibly AIAN to funeral directors. This inflates the CR by increasing the self-identified denominator without proportionally increasing DC-identified numerators.
- Hispanic nativity effect: Foreign-born CR = 1.02 vs. United States (US)-born CR = 1.07; funeral directors appear to use birthplace/accent/surname as Hispanic-identity cues, producing more accurate coding for the foreign-born.
- Hispanic subgroup denominator problem: The 2000 Census changed the Hispanic-origin question (removed country examples, replaced "origin" with "Latino"), causing 16% of Hispanics — disproportionately Central/South Americans (+34.4%) — to give only a generic term. This compromises denominator accuracy for Hispanic subgroup mortality rates regardless of DC coding quality.
How It Works
The National Longitudinal Mortality Study (NLMS) links CPS self-reports (2.3 million respondents, 1973–1998) to National Death Statistics System (NDSS) death certificates via probabilistic record linkage. The matched dataset (252,627 deaths) allows direct comparison of DC-coded race/ethnicity to prior self-report. Sensitivity, PVP, and CR are computed separately by group, sex, age, nativity, and geography. Corrected age-specific death rates are then computed as ASDRi×CRi, and direct standardization against the 2000 U.S. standard population yields corrected AADRs. The NLMS excludes some years and racial groups, and geographic variation in CR is estimated by stratifying the sample by coethnic state concentration.
Why It Matters
- AIAN invisibilization: The 55% AIAN sensitivity rate means nearly half of AIAN deaths are coded as something else (predominantly White). Standard mortality analyses that rely on raw death certificate data systematically understate AIAN mortality and mask what is actually the worst mortality trajectory among any U.S. racial group. See Mortality by Race and Ethnicity.
- Hispanic paradox validation: Because misclassification goes in the direction of overcounting Hispanic deaths relative to the population (CR>1, but only modestly), the paradox survives correction — it is a real biological/behavioral/social phenomenon, not a measurement artifact. This confirms the validity of healthy-immigrant and salmon-bias hypotheses as candidate explanations rather than artifacts.
- Policy/planning consequences: Programs targeted at AIAN health, or Social Security analyses that stratify mortality by race, use systematically wrong AIAN rates. The same applies to any analysis using Centers for Disease Control and Prevention (CDC) Wide-ranging Online Data for Epidemiologic Research (WONDER) without CR correction.
- Methodological lesson: Misclassification direction matters. CR>1 inflates apparent mortality advantage (AIAN looks healthier than it is); CR<1 would inflate apparent mortality disadvantage. The funeral-director proxy mechanism is not random error — it is structured by observability, creating systematic bias that varies predictably with geography and nativity.
Open Questions
- How has the AIAN CR changed since 1998 as National Health Interview Survey (NHIS)/Sexual Orientation and Gender Identity (SOGI) data collection methods evolved?
- Does the salmon-bias hypothesis affect the NLMS AIAN CR (return migration of dying AIAN individuals would change both numerator and denominator)?
- What is the CR for AIAN in the post-Affordable Care Act (ACA) era, when expanded coverage may have changed where deaths occur and who certifies them?
- Does CDC WONDER now apply any CR-based correction for AIAN rates, or is uncorrected data still the standard?
Related
Sources