Key points are not available for this paper at this time.
In recent years much attention has been given to criterion-referenced measures which relate test performance to absolute standards rather than to the performance of others. Popham and Husek (1969) provide a readable account of the differences between such measures and the more traditional norm-referenced tests. The purpose of this paper is to synthesize some of the literature on establishing standards and determining the number of items needed in criterion-referenced measures. This paper is written from the following perspective. A domain (i.e., population) of dichotomously scorable test items is concep-tualized. This population of items need not actually exist. What is important, though, is that it is described well enough so that a relatively high degree of agreement can be reached about which kinds of items are or are not members of the population. In practice, only a reasonably representative sample of items is required.2
Jason Millman (1973) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: