Brooks Chapman and Schroeder 2018 — Understanding Treatment Effect Estimates When Treatment Effects Are Heterogeneous for More Than One Outcome

treatment-effect-heterogeneityinstrumental-variablesLATEATTessential-heterogeneitymultiple-outcomessimulationhealth-economicsexternal-validitytreatment-risk-paradox

Summary

Brooks, Chapman, and Schroeder (2018) extend the essential heterogeneity / local average treatment effect (LATE) framework from single-outcome to multi-outcome treatment choice. When patient–provider dyads make decisions by weighing expected benefits and detriments together — "sorting on the mix" — the true values of average treatment effect on the treated (ATT) and LATE for each outcome depend on the correlation structure of treatment effects across outcomes in the study population, even when the full distribution of benefit effects is identical across populations. Using five simulation scenarios (10001000 runs ×\times 50005000 patients each), they show that (1) estimator identification still holds (regression → ATT, instrumental variables (IV) → LATE, per outcome), but (2) these estimands vary across populations as the benefit/detriment correlation structure shifts, making external validity contingent on matching that structure.

Key Claims

Concepts Introduced or Extended

Entities Mentioned

Quotes

"If treatment choices reflect expected effects over more than one outcome, our simulation results showed that treatment effect estimates can provide evidence as to whether treatments were over or underused in the study population. We also showed that these estimates are very sensitive to the distributions of treatment effects across outcomes in each study population."

"Researchers and policy makers should be extremely cautious of generalizing estimates from a single study population to other patient populations."

My Take

The paper makes a genuine extension to the essential heterogeneity framework: treatment-risk paradox is reframed not as a clinical puzzle but as a predictable rational-choice consequence of multi-outcome sorting. The diagnostic logic — compare ATT and LATE jointly across outcomes to infer whether treatment is over/underused — is practical. The main limitation is that the paper uses binary outcomes with linear probability models; the extent to which the simulation results carry over to continuous outcomes or nonlinear models is not assessed. The paper is also primarily a methods contribution; it does not provide empirical evidence of any specific case where single-outcome interpretation led to a policy mistake, making the applied stakes somewhat abstract.