Measurement authorities propose that the reliability of multiple-choice tests will be enhanced if the distribution of item difficulties is concentrated at approximately S O . This recommendation is reinforced and extended in this article by viewing the 0/1 item scoring as a dichotomization of an underlying normally distributed ability score. If guessing is not a consideration, the dichotomous item scoring leads to maximum item and test reliability when all items have a difficulty of S O . When guessing is possible, an upward shift leads to an optimal difficulty that lies between .57 and .67. It is further shown, however, that the effect on test reliability is surprisingly small, even when the distribution of item p values is spread between .27 and .79 or centers around a mean difficulty as high as .73.
No takes yet. Share an insight, caveat, or question.
Leonard S. Feldt (1993) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: