Summary
Rupp and Davies (2000) provide a concise primer on the Social Security Administration's administrative data infrastructure and the rationale for matching it with survey data (primarily the Survey of Income and Program Participation, SIPP) for disability policy research. The paper describes five core SSA databases (MBR, SSR, NDDS, MEF, NUMIDENT), discusses the relative strengths of administrative vs. survey data, and illustrates matched-data use through two case studies: the Project NetWork demonstration evaluation and the SIPP-based Supplemental Security Income (SSI) eligibility simulation. Publication date is uncertain (2000 based on the substantial gainful activity [SGA] level referenced); appears to be an SSA working paper or book chapter.
Key Claims
- MBR (Master Beneficiary Record): 133 million person records covering all Old-Age, Survivors, and Disability Insurance (OASDI) applicants and beneficiaries. Contains payment eligibility history, diagnosis codes, Average Indexed Monthly Earnings (AIME), state/county, dual-entitlement status. Key extracts: 810/811 historical format (full benefit history), monthly 1% and 10% snapshots.
- SSR (Supplemental Security Record): 57 million records for all SSI applicants and recipients. Monthly program participation and benefit history from January 1974 through present; includes diagnosis codes, countable income, state supplementation, payment status. Key extracts: 10% monthly CER (Characteristic Extract Record); 1% SSI Longitudinal Extract updated every six months.
- NDDS (National Disability Determination System): All Disability Determination Services (DDS) disability decisions — initial applications, continuing disability reviews (CDRs), redeterminations, appeals — for adults and children. Contains Social Security number (SSN), impairment codes, filing dates, DDS decision dates and results. Used to supplement demographic and diagnostic information from MBR/SSR.
- MEF (Master Earnings File): Full (uncapped, not top-coded) annual earnings for nearly 400 million workers, 1951–present, derived from IRS records. IRS ownership imposes strict confidentiality constraints: SSA employees can access MEF, but contractors and grantees cannot. Provides exact earnings histories for all covered workers, essential for measuring work activity, insured status, and SGA compliance.
- NUMIDENT: ~689 million records for ~389 million individuals who ever applied for an SSN. Contains name, date/place of birth, parents' names and SSNs, citizenship evidence, date of death. Used primarily as a demographic crosscheck and for death verification.
- Administrative records: strengths and limitations: Strengths — 100% coverage of program universe, no attrition, low marginal cost per additional observation, decades of longitudinal history (MEF back to 1951, SSR back to 1974). Limitations — rigid (not customizable for research), only cover program participants (truncated population), moral hazard in reporting (incentives to misreport to obtain favorable administrative outcomes), some fields overwritten or not updated.
- Survey data: complementary strengths: SIPP provides monthly income detail, asset information, health/disability topical modules, and — crucially — information on non-participants. Weaknesses: expensive, subject to attrition, sample sizes limit analysis of small subpopulations, self-report errors (respondents confuse SSI with Social Security, leading to participation misclassification).
- SIPP-SSA exact match creates a "superb" combined data source: By matching SIPP respondents to MBR/SSR/NDDS/MEF/NUMIDENT via SSN: (a) administrative participation flags replace survey self-reports, correcting the SSI/SS confusion problem; (b) longitudinal follow-up extends indefinitely (monthly SSR history from 1974); (c) DDS impairment codes improve disability classification; (d) concurrent SSI + Disability Insurance (DI) recipients can be identified; (e) MBR enables computation of DI insured status.
- Project NetWork case study: 146,861 cases identified by simulating eligibility from national administrative records (MBR, SSR, NDDS, MEF, NUMIDENT); random assignment off-site by independent contractor (Abt Associates); admin records tracked benefits and earnings outcomes without attrition for years post-randomization; two survey waves (3,439 baseline interviews, 87% response rate for participants / 50% for nonparticipants) added non-economic variables. Conclusion: admin records form the core of demonstration evaluations; surveys are supplementary.
- SIPP-based SSI eligibility simulation: Three-module model (categorical eligibility, assets, income) replicating SSI determination using SIPP data; administrative match then corrects participation flags and extends longitudinal histories. Enables policy simulation (e.g., "What would happen to SSI rolls if the unearned income disregard increased from $20 to $125?") that neither admin records alone nor SIPP alone could support.
Concepts Introduced or Extended
Entities Mentioned
Quotes
"The joint use of multiple records systems and associated survey data is especially intriguing in the case of the disabled segment of SSA's target populations."
"Administrative record data are at the core of SSA demonstration evaluations."
"By exact matching the SIPP to SSA administrative records data bases, we can refine and enhance the model, and expand the nature of policy questions the model can address."
My Take
This is the definitive technical reference for understanding why most DI/SSI econometric research uses SIPP-SSA matched data. The paper's value is entirely as a reference document — it clearly explains what each SSA database contains, what can and cannot be done with them, and why matching matters for disability research. Any paper in this wiki that uses "MBR," "SSR," "MEF," "NUMIDENT," "SIPP-SSA matched data," or "administrative records" is relying on the infrastructure this paper describes. One important limitation: the MEF restriction (SSA employees only, not contractors/grantees) means that many academic researchers using SSA data obtain earnings information through different channels (e.g., Continuous Work History Sample extracts, Detailed Earnings Record [DER] from IRS-SSA matching). The paper's discussion of SGA at $700 suggests it was written around 2000; no formal publication venue is given in the document.