The application of any statistical technique to numerical data is strictly appropriate only when the data conform to the assumptions and requirements inherent in the development of that statistical method. In the analysis of variance, we require normality, additivity, independence, and homogeneity of variances of the data. When these conditions are not at least approximately satisfied, the analysis is suspect. In postulating the mathematical model for the rank analysis of incomplete block designs and more specifically for the method of paired comparisons [1] we chose the model because it seemed intuitively plausible, because it was related to the binomial model, and because it was mathematically workable. The major objective of this paper is to develop a procedure for testing the appropriateness of the model for the method of paired comparisons. The test is then applied to a variety of experiments involving taste, preference, and appearance judgments for a technique is only very useful when it is applicable to a considerable variety of data. In the reference cited above, three tests of significance were proposed. Two of them are tests of treatment effects, and the other is a test of agreement among judges or over groups of the data. The test of agreement could be regarded as a test of the need of the model that permits judges to differ on their judgments as opposed to the model in which it is postulated that all judges are coordinated in their judgments on the attributes under consideration. We could call the test of agreement a between-group test of goodness of fit. These comments will become clearer in view of the following summary of the mathematical model for paired comparisons.
No takes yet. Share an insight, caveat, or question.
Ralph A. Bradley (1954) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: