A distribution-free procedure for classifying a univariate random variable, z, into one of two populations on the basis of a sample of size N, in which m members are classified into one population and the remaining (N – m) into the other, is given as follows: Let t(z) = k(z) – h(z), where k(z) is the number of observations from the first population which are less than z and h(z) is similarly defined for the second population. If z ≦ ζ*, where ζ* is that value of z for which t(z) is a maximum, classify z into the first population, otherwise into the second. The probability of correct classification, and its estimate, [N – m + t(ζ*)]/N, both converge in probability to the maximum attainable probability of correct classification.
No takes yet. Share an insight, caveat, or question.
David S. Stoller (1954) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: