The Social Security Administration (SSA) maintains five core administrative databases that collectively cover the entire U.S. disability program universe: all applicants and beneficiaries of DI (Social Security Disability Insurance) and SSI (Supplemental Security Income), all Disability Determination Services (DDS) disability determinations, the full earnings histories of every American worker since 1951, and identifying information for everyone who has ever held a Social Security Number (SSN). These systems, used individually or in matched combination with survey data, form the data infrastructure for virtually all longitudinal disability policy research.
Covers all OASDI (Old-Age, Survivors, and Disability Insurance) applicants and beneficiaries. Approximately 133 million person records grouped into ~93 million units by the primary wage earner's SSN. Each beneficiary in a unit has a separate SSN-linked record.
Key variables: SSN, gender, race, date of birth, primary insurance amount (PIA), average indexed monthly earnings (AIME), state/county code, date of entitlement, date of filing, type of claim, diagnosis code, dual-entitlement status.
Key extracts: 810/811 historical format (complete benefit amount history); monthly 1% and 10% snapshot extracts; MBR Universe file (entire ~133M cases, limited fields, updated every 6 months).
Covers all SSI applicants and recipients. Approximately 57 million person records grouped into ~42 million SSI units. One person may have multiple SSR records due to new applications or changes in household composition.
Key variables: SSN, date of birth, gender, race, noncitizen information, date of claim, primary and secondary disability diagnosis codes, state/county of residence, countable earned and unearned income, payment status, federal SSI benefit amount, state supplement amount, date of first payment.
Key extracts: Monthly CER (Characteristic Extract Record) — 10% sample by last three SSN digits, distinguishes between payment eligibility and actual payment received; 1% SSI Longitudinal Extract — updated every six months, contains monthly eligibility and payment history from January 1974 through present.
Records all disability decisions made by state Disability Determination Services. Covers initial applications, redeterminations, Continuing Disability Reviews (CDRs), and appeals, for both adult and child claimants.
Key variables: SSN, beneficiary identification code, filing date, type of claim, date of DDS/SSA decision, result of determination (allowed/denied), date of birth, primary and secondary impairment codes, date disability period began, gender, race.
Used primarily to supplement MBR and SSR with impairment codes and disability determination history.
Contains full (uncapped, not top-coded) annual earnings for workers based on their SSN — approximately 400 million earnings records covering 1951–present. Updated from Internal Revenue Service (IRS) records by November following each tax year.
Key variables: Annual Federal Insurance Contributions Act (FICA) earnings, Medicare earnings, gender, race, date of birth, date of death, first and last year of earnings.
Critical constraint: The MEF is legally owned by the IRS. SSA employees can use MEF internally, but SSA contractors and grantees cannot access it directly. Academic researchers typically access earnings data through alternative channels (Continuous Work History Sample extracts, or DER — Detailed Earnings Record — through IRS-SSA data agreements).
Contains information on all persons who have ever submitted an SSN application — approximately 689 million records for 389 million people. Applicants may have multiple records from SSN reissuances.
Key variables: Name, date of birth, city and state/country of birth, gender, race, mother's and father's names and SSNs, evidence used to establish citizenship, date of death.
Used as a demographic crosscheck, for citizenship/nativity verification, for death date verification, and as a sampling frame. The date-of-death field is an important supplement to vital statistics records.
A publicly distributed file (sold via the National Technical Information Service, NTIS) recording SSNs, names, dates of birth, dates of death, and states of SSN application for deceased individuals whose deaths were reported to SSA. Contains the majority of deceased Americans' records; coverage improved over time as state vital record linkage expanded.
Intended use: Fraud detection — allows lenders, employers, and agencies to identify attempts to use deceased persons' SSNs. Also widely used in research as a mortality follow-up source.
Research use: Standard linkage target when vital records are unavailable; used for mortality ascertainment in longitudinal DI/SSI studies.
Privacy risk (Acquisti and Gross 2009): Because the SSA's Enumeration at Birth program made SSN assignment highly correlated with birth date, DMF records (sorted by state and birth date) reveal the sequential SSN assignment scheme. An attacker can train on DMF patterns and then predict living individuals' SSNs from birth date and state alone — with 44% first-5-digit accuracy (single attempt) for post-1989 cohorts. This vulnerability directly prompted SSA to fully randomize SSN assignment in June 2011. Post-2011, public DMF access was restricted to a "Limited Access" version. See Death Master File.
The Survey of Income and Program Participation (SIPP) is the primary national survey matched to SSA administrative records for disability policy research. SIPP provides: monthly income detail (including source breakdowns), detailed asset information (from topical modules), health impairments and work limitations, and — crucially — information on non-participants who have never interacted with SSA programs.
By exactly matching SIPP respondents to SSA databases via SSN, researchers gain:
SIPP-SSA matched data enables policy simulation (e.g., effects of changing SSI income disregards) that neither source alone could support: SIPP provides the non-participant population needed to measure eligibility take-up; admin records provide the precise program rules and histories.
Admin records alone cannot answer questions about disability program participation in the general population, because they are truncated to program participants. All non-applicants are missing.
Survey data alone lacks the precision needed to simulate detailed eligibility rules (income and asset calculations), understates program participation (SSI/SS confusion), and cannot track participants longitudinally without expensive resurveys.
Matched data overcomes both limitations at the cost of matching complexity, privacy restrictions (SSN linkage), and the MEF access constraint.
Virtually every longitudinal DI/SSI study in this wiki uses some component of this infrastructure: