TY - RPRT
AU - Jaeger, David A.
TI - Robustness? Range Tests for Equality and Equivalence Across Specifications
PY - 2026/Aug/
PB - Institute of Labor Economics (IZA)
CY - Bonn
T2 - IZA Discussion Paper
IS - 18851
UR - https://www.iza.org/publications/dp18851
AB - Applied economists routinely compare estimates across specifications, observe that they are "similar," and conclude that their results are "robust." This common procedure makes an implicit inferential claim about the range of estimates, but usually does not account for their joint sampling distribution. I formalize informal practice with two bootstrap statistics. The minimum equivalence bound, R*1−α, is the smallest tolerance within which the estimates can be judged equivalent. The range-based equality p-value, pR, tests whether the estimates are statistically distinguishable. Together they distinguish failure to detect differences from affirmative evidence of agreement. Simulations show approximately correct size and coverage. Applications to five prominent papers validate some robustness claims while revealing cases in which apparent agreement reflects imprecision rather than stability. A survey of CEPR and NBER affiliates shows that expert judgments align with the framework in obvious cases but diverges in intermediate cases. I suggest that R*.95 and pR be reported whenever multiple specifications are presented as evidence of robustness.
KW - joint inference
KW - bootstrap inference
KW - equivalence testing
KW - specification sensitivity
KW - robustness
KW - model uncertainty
ER -