The effects of three amounts of DIF (10%, 15% and 30% of DIF-items), three test lengths (20-, 40-, and 60-items), and three test score (matching criterion) purification types (single-stage, two-stage, and iterative) on robustness and power of Mantel-Haenszel (MH) DIF detection procedures were studied. Item response data were generated under the three parameter logistic model (3PLM) for focal and reference group subjects, where the ability distributions of the two groups were equal. In the 10% DIF item conditions the three MH procedures are robust and have sufficient power, but in the 15% and 30% DIF item conditions robustness violation and insufficient powers occur. The influence of test length on power is rather modest. On the other hand, test score purification improve power, but the size of their effects is much larger in the 15% and 30% DIF item conditions than in the 10% DIF item conditions.
No takes yet. Share an insight, caveat, or question.
Gideon J. Mellenbergh (2000) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: