Let X be a random vector, taking values in p-dimensional Euclidean space Eᵖ with density f(x; θ). The parameter θ belongs to a subset Θ of a Euclidean space Eq and is unkown. Let g be a function over the parameter space having continuous first partial derivatives and taking values in Eʳ (r q). To test the hypothesis g(θ) = 0 against the alternative g(θ) ≠ 0 using a sample of n independent observations of X, one frequently uses the Neyman-Pearson generalized likelihood ratio test statistic λₙ. The limiting distribution of -2lnλₙ under the null hypothesis, as n → ∞, was shown by Wilks (1938) to be chi-square with r degrees of freedom (assuming regularity conditions). If \θₙ\ is a sequence of alternatives converging to a point of the null hypothesis at the rate n1/2, the limiting distribution is noncentral chi-square with noncentrality parameter equal to the limit of n g(θₙ)' ∑⁻¹g (θₙ) g(θₙ), where ∑g(θ) is the asymptotic covariance matrix of the quantity n1/2 g(θ̂) - g(θ) as n → ∞ with θ fixed (θ̂ denoting the maximum-likelihood estimator of θ based on sample size n). This noncentral convergence was first proved by Wald (1943), along with a number of other results, on the basis of some rather severe uniformity conditions. Davidson and Lever (1970) have proved the result using more intuitive assumptions. Feder (1968) has obtained asymptotic noncentral chi-square for the case where both the hypothesis and alternative regions are cones in Θ; this is essentially a generalization of g(θ) = 0 versus g(θ) ≠ 0, since the hypothesis g(θ) = 0 is locally equivalent to a hyperplane and g(θ) ≠ 0 to its complement. Despite the generality, Feder's assumptions are quite mild compared with Wald's. The result appears in Wald's paper as a special case of a more general statement entitled "Theorem IX." This theorem states that for θ ∈ Θ and -∞ < t < ∞ the relationship {equation*}{1.1}P_θ -2 ln λ_n < t - P_θ K_n < t → 0{equation*} holds uniformly in t and θ, where Kₙ has a noncentral chi-square distribution with r degrees of freedom and noncentrality parameter equal to n g(θ)' ∑⁻¹g (θ) g(θ). This formulation of Wald is too strong. It will be shown by counterexample that, if θ is held fixed while n → ∞, relationship (1.1) fails to hold uniformly in t. The counterexample is that of testing the value of the mean of a normal distribution with unknown mean and variance. Wald's proof of Theorem IX treats two cases separately, case (i) where θₙ approaches the null hypothesis set at the rate n-1/2 or faster, and case (ii) where it does not. The proof of (1.1) in case (i) requires convergence of θₙ at the rate n-1/2 in order that the Taylor series expansion of the logarithm behave nicely. In case (ii) there is no reason at all to believe the distribution of Kₙ to be a good approximation to that of -2lnλₙ. From Wald's paper (page 480, line following (212)) one gets the impression that Wald felt that the statement of uniform convergence of (1.1) in case (ii) was trivial, since pointwise convergence is trivial (because both terms tend to zero for fixed t). But, since Kₙ does not converge in distribution to a random variable in case (ii), there is really no reason why pointwise convergence should imply uniform convergence. In the same paper, Wald (1943) also described a test procedure based only on the unrestricted maximum-likelihood estimator θ̂ₙ. This procedure rejects for large values of the statistic Qₙ = n g(θ̂ₙ)' ∑⁻¹g (θ̂ₙ) g(θ̂ₙ). Wald claimed in his paper that (1.1) again holds uniformly in t and θ if -2lnλₙ is replaced by Qₙ. This claim too is false, in the stated generality, as the same counterexample will demonstrate. Keeping θ as a fixed alternative while n → ∞ has the disadvantage that the limiting behavior of each of the quantities -2lnλₙ, Qₙ and Kₙ is degenerate in the sense that the probability mass moves out to infinity with increasing n. However, statement (1.1), uniform in t for fixed θ, has meaning here since both -2lnλₙ (or Qₙ) and Kₙ may be related to quantities with genuine limiting normal distributions which must be identical or at least very similar in order for (1.1) to be uniform in t. The precise result is embodied in a theorem presented in Section 2 of this paper. In Sections 3 and 4 we consider the case of X normally distributed with mean μ and variance σ², where -∞ < μ < ∞, 0 < σ₁ < σ < σ₂, and the hypothesis to be tested is μ = 0. It is shown in Sections 3 and 4, respectively, that for this problem the relationships P_θ Qₙ < t - P_θ Kₙ < t → 0 and P_θ -2lnλₙ < t - P_θ Kₙ < t → 0 fail to be uniform in t when θ = (μ, σ) is fixed and satisfies μ ≠ 0, σ₁² < σ² < σ₂² - μ². The space of values of σ has been truncated in order to satisfy Wald's regularity conditions. In the following section boldface letters denote vectors and matrices. The law of the random vector x is denoted throughout by L(x). In particular, N(μ, Σ) refers to a normal law with mean vector μ and covariance matrix Σ. By L(xₙ) → L(y) or L(xₙ) → N(μ, Σ) is meant, respectively, that the law of xₙ converges to the law of y or to the stated normal law, as n → ∞. The definitions of the Mann-Wald symbols Oₚ and oₚ may be found in Chernoff ((1956), Section 2), as may the statements of some basic results of large-sample theory which are used freely in the proof of the theorem.
No takes yet. Share an insight, caveat, or question.
T. W. F. Stroud (1972) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: