Two parallel forms of a test are considered “equated” when for a single group of examinees the standard deviations on the two forms are equal and the means on the two forms are equal. Since present test construction procedures are not sufficiently precise to produce forms with equivalent score scales, it is necessary to conduct an equating experiment for the purpose of collecting data for use in adjusting the initial raw score scales. Although the definition of equated score scales is given in terms of equivalence of the first two moments for a single group of examinees, the tests to be equated need not be administered to the same examinees. This paper considers two types of experiments in which the tests to be equated are administered to different examinees. Both types of experiments can be characterized as follows: If two tests (X and Y) are to be equated, X is administered to the first and Y to the second of a pair of samples. In addition, a third test (V) is administered to both samples. V may be either a separate test from. X and Y or a number of items contained in both X and Y. In either case, V may be considerably shorter than X and Y. The two types of experiments are identical with respect to these test administration procedures, but they differ with respect to the assumptions made about the characteristics of the two samples of examinees. The first type of experiment calls for the use of samples of examinees that are randomly drawn from, the same population, F. M. Lord has considered this general model for the case of equally reliable tests; in this paper a solution is given for the case when X and Y are unequally reliable. Using the statistics for the random samples, maximum likelihood estimates are made of the means and variances of true scores in the population from which the samples were drawn. Thus, the first two moments are estimated for a single group of examinees (the population). For this purpose it is necessary to assume that true scores on X, Y, and V intercorrelate unity. The second type of experiment allows the use of samples that are selected in a non‐random manner from a common population. In this case the equating of equally reliable X and Y as well as unequally reliable X and Y is considered. The equations are derived on the assumption that true scores on all tests intercorrelate unity; however, for the case of equally reliable tests a solution is possible so long as true scores on V correlate unity with true scores on either X or Y. It is shown that, if true scores on two tests correlate unity, for pairs of samples of different ability three quantities will remain invariant: 1) the ratio of standard deviations of true scores on the two tests; 2) the difference between the mean on one test and the product of the mean on the second and the ratio of true score standard deviations, and 3) the standard error of measurement. Data are presented which indicate that the equations derived for both types of experiments are sound.
No takes yet. Share an insight, caveat, or question.
Richard A. Levine (1955) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: