For the problem of inference about a real parameter μ on the basis of n independent observations x₁, ⋯, xₙ (or x) each distributed as N(μ, σ²) with σ² "unknown", it is commonly asserted, for example in [2] p. 465, that the Bayesian method is close to other forms of inference (significance tests, confidence and fiducial intervals) since it too may be based on sₙ₋₁(t), the probability density function (pdf) of Student's t with $n - 1$ degrees of freedom. The Bayesian role of sₙ₋₁(t) is that of the posterior pdf of t = n(n - 1)/S1/2 ( x - μ), where x = n⁻¹ ∑ xᵢ and S = ∑ (xᵢ - x)² are the sufficient statistics for μ and σ². It results from formal use in Bayes's Theorem of the improper prior pdf for μ and σ² described by "independence of μ and log σ and their uniform distributions on R¹". More convincing support for sₙ₋₁(t) as a posterior pdf could be obtained by detailed examination of the product space of proper (integrable) prior pdfs and ( x, S) and the determination of the essential features of the region where replacement of the posterior pdf of μ by that derived from sₙ₋₁(t) does not seriously affect inference about μ. In this note, attention will be confined to prior pdfs in the following class. Let ω denote the Fisher information σ⁻² and let I\ \ denote the 0-1 indicator function of a set. Consider prior pdfs for μ and ω drawn from the sequence {equation*}{1.1}p_α(μ, ω) ∝ ω⁻¹I\{μ, ωμ1α < μ < μ2α, ω1α < ω < ω2α\} α = 1, 2, ⋯.{equation*} For each member of this sequence, μ and ω are independent while μ and log ω (or log σ) have rectangular distributions (from which it is clear that the choice of (1.1) is motivated by the improper prior pdf for μ and σ² above). The posterior pdf of μ obtained by combining p_α(μ, ω) with the likelihood function p(xμ, ω) ∝ ω1/2n exp -1/2nω( x - μ)² - 1/2ω S is proportional to $∫^{ω2α}_{ω1α} ω1/2n-1 exp -1/2nω( x - μ)^2 - 1/2ω S dω· I\{μμ1α < μ < μ2α\}$ giving, with the change of variable $u = ω 1 + t^2/(n - 1) S$ {equation*}{1.2}p_α(t x) ∝ sₙ₋₁(t) ∫^{ 1+t^2/(n-1) Sω2α}_{ 1+t^2/(n-1) Sω1α} u1/2(n-2)e-1/2udu{equation*} I n(n - 1)/S1/2( x - μ2α) < t < n(n - 1)/S1/2( x - μ1α)\. To obtain sₙ₋₁(t), Jeffreys (p. 68 of [1]) uses a convergence argument which, in our specialisation, would involve letting {equation*}{1.3}μ1α → - ∞, μ2α → ∞, ω1α → 0, ω2α → ∞ as α → ∞{equation*} as necessary and sufficient conditions for {equation*}{1.4}lim p_α(t) ≡ sₙ₋₁(t){equation*} for all values of x. In (1.4), x is kept fixed. However, in changing α, we are changing the prior distribution used, so that keeping x fixed has no obvious relevance. To emphasize that a different x would normally be associated with a different prior pdf, we will, except in the proofs of Section 2, write x_α, x_α, S_α, t_α for the x, x, S, t associated with p_α(μ, ω). A radically different justification of sₙ₋₁(t) is provided as follows. Let us suppose that the person who is to make the inference about μ has the prior pdf pₛ(μ, ω) for s some positive integer, that is, a pdf that happens to be a member of the sequence (1.1). Examination of (1.2) shows that he can take sₙ₋₁(tₛ) as a good approximation to his posterior pdf provided {equation*}{1.5}S_sω₁ₛ 1, S_sω₂ₛ 1, S-1/2_s(μ₂ₛ - x_s) 1, S-1/2_s( x_s - μ₁ₛ) 1.{equation*} Now a person holding the prior pdf pₛ(μ, ω) would expect to obtain xₛ's according to the marginal pdf pₛ(xₛ) = ∫ ∫ p(xₛμ, ω)pₛ(μ, ω) dμ dω. The probability of (1.5) under pₛ(xₛ) is therefore the person's prior probability of being able to use sₙ₋₁(tₛ) as a basis for inference about μ. In the light of this, if, for the sequence (1.1), we were to have plim S_αω1α = 0, plim S_αω2α = ∞, {equation*}{1.6} plim S-1/2_α( x_α - μ1α) = ∞, plim S-1/2_α(x̄_α - μ1α) = ∞ {equation*} with the plim evaluated with respect to the sequence of marginal distributions p_α(x_α), we would, by proceeding down the sequence, be able to invest sₙ₋₁(t) with an asymptotic justification. (By plim z = ∞, we mean that lim Prob (z < K) = 0 for all K.) In Lemma 1 be (a) ρ2α/ρ1α→ ∞ {equation*}{1.7} (b) ρ2α → ∞ {equation*} (c) lim inf log ρ1α/log ρ2α 0 as α → ∞. Lemma 2 then shows that (1.6) is equivalent to {equation*}{1.8}plim p_α(t_α) ≡ sₙ₋₁(t){equation*} where the plim is again evaluated with respect to the sequence p_α(x_α), α → ∞. Hence (1.7) is necessary and sufficient for (1.8) which, since it allows direct comparison with the Jeffreys approach in (1.3) and (1.4), we state as the principal theorem. The interpretation of the conditions (1.3) is superficially straightforward; it is that the prior pdfs for μ and ω should (separately) approach conditions representing "complete ignorance". (1.7) is apparently more complex. In the requirement ρ2α/ρ1α → ∞, it agrees with (1.3); its principal divergence from (1.3) lies in the existence of the joint conditions, (b) and (c), on the developments of the prior pdfs of μ and ω. ρ1α and ρ2α may be regarded as measures of the information about μ in the least and most informative conditional distribution p(xμ, ω) allowed by p_α(μ, ω), relative to the prior information about μ measured by the quantity (μ2α - μ1α)⁻¹. (1.7) (c) requires that, although there is no necessity for ρ1α to approach zero at all, if it does so, it should not do so too rapidly that is, loosely speaking the least informative conditional distribution should not be too uninformative. For the case μ1α = -α, μ2α = α, ω1α = α^λ, ω2α = α, (1.3) requires -∞ < λ < 0, while (1.7) requires -2 λ < 1. The case μ1α = -1, μ2α = 1, ω1α = 1, ω2α = α satisfies (1.7) but not (1.3). The comparison of (1.3) and (1.7) is assisted by noting that t_α is invariant with respect to the simultaneous transformations of x and μ, x → a_α x + b_α, μ → a_αμ + b_α. We would therefore expect that any reasonable condition on the sequence (1.1) for the asymptotic relevance of sₙ₋₁(t) would be unaffected by these transformations, when coupled with ω → a⁻²_αω. (1.7) agrees with such expectation while (1.3) does not.
No takes yet. Share an insight, caveat, or question.
M. Stone (1963) studied this question.