X(1), X(2), ⋯, X(m), Y(1), Y(2), ⋯, Y(n) are independent k-variate random variables. The distribution of $X(i)$ has pdf $f(x)$, say, where x denotes a k-dimensional vector throughout this paper, and the distribution of $Y(j)$ has pdf $g(x)$, say. We assume that $f(x)$ and $g(x)$ are piecewise continuous, and that each has a finite upper bound, which it is not necessary to specify. Denote by 2Rᵢ the distance from $X(i)$ to the nearest of the points X(1), ⋯, X(i - 1), X(i + 1), ⋯, X(m), and denote by Sᵢ the number of points Y(1), ⋯, Y(n) contained in the open sphere : | x - X(i) | < Rᵢ\. Clearly, the joint distribution of Sᵢ, Sⱼ is the same as the joint distribution of Si', Sj', for any subscripts with i ≠ j, i' ≠ j'. Let r be a non-negative integer, and α any fixed positive value. $Q(r)$ denotes the Lebesgue integral ∫Eₖ 2ᵏ α f² (x) g(x) ʳ g(x) + 2ᵏα f(x) r + 1 dx, where Eₖ denotes Euclidean k-space. We will show that limm → ∞, m/n = α Pm, n S₁ = s₁, S₂ = s₂ = Q(s₁)Q(s₂), for any non-negative integers s₁,s₂, the approach being uniform in s₁,s₂. Thus, in the limit S₁, S₂ are independently distributed, with limm → ∞, m/n = α Pm, n S₁ = s₁ = Q(s₁). In [1], which discussed the univariate case, Sᵢ was defined as the number of Y's closer to $X(i)$ than to any other X to their right. In the present paper, Sᵢ is defined as the number of Y's in another neighborhood of $X(i)$. Our present definition of Sᵢ does not become for $k = 1$ the same as the definition of Sᵢ in [1]. Rather, in the univariate case, our present definition of Sᵢ is the number of Y's lying within a distance Rᵢ on either side of $X(i)$. However, if limm → ∞, m/n = α Pm, n S₁, = s₁, S₂ = s₂ is computed for the univariate case using the definition of Sᵢ given in [1], the only way in which it differs from Q(s₁)Q(s₂) is that α is replaced by α/2. Thus it seems reasonable to treat the Sᵢ as defined here as k-dimensional analogues of the Sᵢ as defined in [1], at least for large samples. An intuitive reason for α being replaced by α/2 is that in our present case, ∑ᵐi = 1 Sᵢ may be less than n, whereas in [1] this sum must always equal n. Thus in our present case, we are in a sense discarding some of the Y's, which lowers n relative to m and thus raises α by a certain factor (2, as it happens). In our present case, ∑ Sᵢ may be less than n because the Rᵢ are chosen to make the spheres around the X's non-overlapping, thus simplifying the analysis. The Rᵢ were chosen to give the largest possible non-overlapping spheres because it would seem intuitively that the larger the spheres, the more rapid the approach of the probabilities to their limiting values.
No takes yet. Share an insight, caveat, or question.
Lionel Weiss (1960) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: