Definition
Rational Unbiased Reporting (RUR) is the hypothesis that an individual's self-reported disability status is an unbiased predictor of the Social Security Administration (SSA)'s ultimate disability award decision. Formally, E[a~−d~∣x]=0, where a~ is the SSA's ultimate award indicator and d~ is the individual's self-reported disability status, conditional on a vector of covariates x. Equivalently, in a bivariate probit formulation, RUR requires βa=βd: the parameter vector governing the SSA's decision equals the vector governing the individual's self-report. The hypothesis does not claim individuals know the SSA's regulations; it claims that individuals' assessments of their own functional limitations align with the SSA's ultimate determinations in expectation.
Key Ideas
- Ultimate vs. first-stage award. RUR holds for the SSA's ultimate award decision (after all appeals) but is rejected for the SSA's first-stage Disability Determination Services (DDS) decision. Initial DDS screeners are systematically harsher than self-assessors; the appeal-and-review process restores unbiasedness at the final stage. This indicates the appeal process functions as a self-selection mechanism that corrects the initial denial overshoot.
hlimpw as approximate sufficient statistic. The Health and Retirement Study (HRS) binary variable "Does a health problem keep you from working altogether?" captures most of the predictive content of the full SSA determination process. About 20% of disability insurance (DI) recipients report no work limitation and 5% report earnings above Substantial Gainful Activity (SGA) — this is consistent with survey measurement noise, not systematic over-reporting of disability.
- Five-test battery. Benítez-Silva et al. (2004) test RUR with five progressively stringent methods: (1) unconditional mean test, (2) conditional moment restrictions, (3) ordinary least squares (OLS) projection, (4) Bierens (1990) consistent conditional moment test, (5) Horowitz-Spokoiny (2001) adaptive rate-optimal test. All five fail to reject RUR against ultimate award (p-values: 0.33,0.19,0.11,0.09,0.79). All five reject against first-stage DDS decisions (all p=0.00).
- Bivariate probit parametric test. The one-type bivariate probit likelihood ratio (LR) statistic (38.4) fails to reject βa=βd. The estimated ρ=0.206 indicates moderate positive correlation between SSA and individual unobservables, consistent with a common underlying health state affecting both. A two-type extension (60% "straightforward" cases, 40% "ambiguous") is rejected one-type vs. two-type (LR=75.66) but the restricted two-type model itself is only marginally rejected (LR=68.11, p=0.067).
- Kreider (1999) contrast. Prior parametric studies — notably Kreider (1999) — found large self-report biases. Benítez-Silva et al. (2004) show these biases disappear under nonparametric conditional moment tests. The prior findings were artifacts of imposed functional form restrictions, not genuine over-reporting.
- LFP timing evidence. Labor force participation (LFP) drops sharply — from ≈60% to ≈15% — in the month of reported disability onset, validating that
hlimpw onset dates correspond to genuine work-limiting events rather than strategic self-labeling.
How It Works
Let a~=I(x′βa+εa≥0) denote the SSA's (latent) disability decision and d~=I(x′βd+εd≥0) the individual's self-reported disability status. The RUR hypothesis is the restriction βa=βd. The nonparametric formulation requires only E[a~−d~∣x]=0, with no assumption on the functional form linking x to a~ or d~. The Bierens (1990) test evaluates this conditional expectation restriction via weighted integrals; the Horowitz-Spokoiny (2001) test adapts the bandwidth to local data density. Both are consistent against any fixed alternative in the class of conditional moment restrictions.
Why It Matters
- Validity of HRS-based disability research. If self-reported disability is unbiased for ultimate SSA determinations, it can be used as a proxy for "true" disability status in studies that lack administrative data linkage. This substantially expands the set of questions tractable with survey data.
- Audit critique. Researchers who find high DI award rates for "non-disabled" applicants (e.g., French and Song 2014) may be partly observing the first-stage denial overshoot that the appeal process corrects, not evidence that awarded beneficiaries are not genuinely disabled. RUR's distinction between ultimate vs. first-stage decisions is essential for interpreting these studies.
- Self-report bias as artifact. The paper is a methodological warning: strong parametric assumptions can generate apparent biases that vanish under nonparametric testing. Studies estimating behavioral responses to DI programs should be cautious about results that hinge on functional form.
- Design of disability surveys. If
hlimpw-type binary work-limitation questions are approximate sufficient statistics for SSA decisions, then simple survey measures may be adequate for most DI research purposes, without requiring complex self-assessment instruments.
Open Questions
- RUR holds in expectation (E[a~−d~∣x]=0), but variance structure matters too. Individuals may over-report in some dimensions and under-report in others, with the biases canceling on average.
- The result is conditional on the appeal process correcting the initial DDS denial overshoot. If the appeal mechanism were reformed (e.g., tighter administrative law judge (ALJ) discretion), RUR might fail for ultimate award as well.
- The 2004 paper uses HRS waves 1–3 with a small sample of confirmed DI applicants (n=356). Whether RUR holds in more recent HRS waves or among Supplemental Security Income (SSI) applicants is untested.
Related
Sources