Los puntos clave no están disponibles para este artículo en este momento.
The problem of testing outlying observations, although an old one, is of considerable importance in applied statistics. Many and various types of significance tests have been proposed by statisticians interested in this field of application. In this connection, we bring out in the Histrical Comments notable advances toward a clear formulation of the problem and important points which should be considered in attempting a complete solution. In Section 4 we state some of the situations the experimental statistician will very likely encounter in practice, these considerations being based on experience. For testing the significance of the largest observation in a sample of size n from a normal population, we propose the statistic S²ₙS² = ^n-1₈=₁ (xᵢ - xₙ) ²ⁿ₈=₁ (xᵢ - x) ² where x₁ x₂ xₙ, xₙ = 1n - 1 ^n-1₈=₁ xᵢ and x = 1n^n₈=₁ xᵢ. A similar statistic, S²₁/S², can be used for testing whether the smallest observation is too low. It turns out that S²ₙS² = 1 - 1n - 1 (xₙ - xs) ² = 1 - 1n - 1 T²ₙ, where s² = 1n (xᵢ - x) ², and Tₙ is the studentized extreme deviation already suggested by E. Pearson and C. Chandra Sekar 1 for testing the significance of the largest observation. Based on previous work by W. R. Thompson 12, Pearson and Chandra Sekar were able to obtain certain percentage points of Tₙ without deriving the exact distribution of Tₙ. The exact distribution of S²ₙ/S² (or Tₙ) is apparently derived for the first time by the present author. For testing whether the two largest observations are too large we propose the statistic S²₍-₁, ₍S² = ^n-2₈=₁ (xᵢ - x₍-₁, ₍) ²ⁿ₈=₁ (xᵢ - x) ², x₍-₁, ₍ = 1n - 2 ^n-2₈=₁ xᵢ and a similar statistic, S²₁, ₂/S², can be used to test the significance of the two smallest observations. The probability distributions of the above sample statistics S² = ⁿ₈=₁ (xᵢ - x) ² where x = 1n ⁿ₈=₁ xᵢ S²ₙ = ^n-1₈=₁ (xᵢ - xₙ) ² where xₙ = 1n-1 ^n-1₈=₁ xᵢ S²₁ = ⁿ₈=₂ (xᵢ - x₁) ² where x₁ = 1n-1 ⁿ₈=₂ xᵢ are derived for a normal parent and tables of appropriate percentage points are given in this paper (Table I and Table V). Although the efficiencies of the above tests have not been completely investigated under various models for outlying observations, it is apparent that the proposed sample criteria have considerable intuitive appeal. In deriving the distributions of the sample statistics for testing the largest (or smallest) or the two largest (or two smallest) observations, it was first necessary to derive the distribution of the difference between the extreme observation and the sample mean in terms of the population. This probabilityX₁ x₂ x₃ xₙ s² = 1n ⁿ₈=₁ (xᵢ - x) ² x = 1n ⁿ₈=₁ xᵢ distribution was apparently derived first by A. T. McKay 11 who employed the method of characteristic functions. The author was not aware of the work of McKay when the simplified derivation for the distribution of xₙ - x outlined in Section 5 below was worked out by him in the spring of 1945, McKay's result being called to his attention by C. C. Craig. It has been noted also that K. R. Nair 20 worked out independently and published the same derivation of the distribution of the extreme minus the mean arrived at by the present author--see Biometrika, Vol. 35, May, 1948. We nevertheless include part of this derivation in Section 5 below as it was basic to the work in connection with the derivations given in Sections 8 and 9. Our table is considerably more extensive than Nair's table of the probability integral of the extreme deviation from the sample mean in normal samples, since Nair's table runs from n = 2 to n = 9, whereas our Table II is for n = 2 to n = 25. The present work is concluded with some examples.
Frank E. Grubbs (Wed,) studied this question.