Confidence Intervals for the Population Mean with Unknown Variance (t-Statistic) | 方差未知时总体均值的置信区间(t 统计量)

📚 Confidence Intervals for the Population Mean with Unknown Variance (t-Statistic) | 方差未知时总体均值的置信区间(t 统计量)

Suppose a nutritionist wants to estimate the mean caffeine content of a certain brand of energy drink. She takes a random sample of 12 cans and finds a sample mean of 82.5 mg and a sample standard deviation of 7.3 mg. She does not know the population standard deviation σ. How can she construct a 95% confidence interval for the true mean μ? This is exactly the scenario covered by the t-procedure, and it is one of the most frequently tested topics in IB Mathematics Analysis & Approaches HL, particularly in the Statistics and Probability option.

假设一位营养师想估计某品牌能量饮料的平均咖啡因含量。她随机抽取了 12 罐,发现样本均值为 82.5 mg,样本标准差为 7.3 mg。她并不知道总体标准差 σ。她该如何构建总体真实均值 μ 的 95% 置信区间呢?这正是 t 方法所处理的场景,也是 IB 数学分析与方法 HL 中,尤其是统计与概率模块中最常考查的知识点之一。


1. From Known Variance to Unknown Variance | 从已知方差到未知方差

When the population standard deviation σ is known, the quantity z = (x̄ − μ) / (σ/√n) follows a standard normal distribution N(0, 1). A 100(1 − α)% confidence interval is therefore x̄ ± z* · σ/√n, where z* is the upper α/2 critical point of N(0, 1).

当总体标准差 σ 已知时,统计量 z = (x̄ − μ) / (σ/√n) 服从标准正态分布 N(0, 1)。因此 100(1 − α)% 置信区间为 x̄ ± z* · σ/√n,其中 z* 是 N(0, 1) 的 α/2 上侧分位数。

In practice, however, σ is almost never known. If we blindly replace σ by the sample standard deviation s, the ratio (x̄ − μ) / (s/√n) no longer follows a normal distribution. Because s itself varies from sample to sample, the extra randomness makes the distribution more spread out than N(0, 1), especially for small sample sizes.

然而在现实中,σ 几乎从来都是未知的。如果盲目地用样本标准差 s 代替 σ,比例式 (x̄ − μ) / (s/√n) 就不再服从正态分布了。由于 s 本身随着样本的不同而波动,额外的随机性使该分布比 N(0, 1) 更加分散,在小样本时尤其明显。

The solution is to use the Student’s t-distribution, which was specifically designed to handle this situation. It adjusts the critical value so that the interval remains valid even when we have to estimate σ with s.

解决方案是使用学生 t 分布,它正是为处理这种情况而设计的。t 分布会调整临界值,使得即使我们不得不以 s 来估计 σ,所得到的区间仍然有效。


2. The Student’s t-Distribution | 学生 t 分布

The t-distribution, also called Student’s t-distribution, was developed by the English statistician William Sealy Gosset in 1908. Gosset worked at the Guinness Brewery in Dublin, where he was not allowed to publish under his own name. He consequently used the pseudonym “Student”.

t 分布,又称学生 t 分布,由英国统计学家威廉·西利·戈塞特于 1908 年提出。戈塞特在都柏林的健力士啤酒厂工作,当时他不能以本名发表文章,因此使用了笔名 “Student”(学生)。

Like the standard normal distribution, the t-distribution is symmetric, bell-shaped and centred at 0. Its key difference is that it has heavier tails: it assigns greater probability to extreme values, reflecting the additional uncertainty introduced by estimating σ using s.

与标准正态分布类似,t 分布是对称、钟形且以 0 为中心的。其关键区别在于它的尾部更厚:它赋予极端值更大的概率,这反映了用 s 估计 σ 所带来的额外不确定性。

The exact shape of the t-distribution is controlled by a parameter called the degrees of freedom (df). For one-sample t-procedures, df = n − 1. As the sample size n grows, the t-distribution approaches the standard normal distribution more and more closely.

t 分布的具体形状由一个称为自由度的参数控制。在单样本 t 方法中,df = n − 1。随着样本量 n 的增大,t 分布会越来越接近标准正态分布。


3. Degrees of Freedom | 自由度

The concept of degrees of freedom can be understood as the number of independent pieces of information available after estimating parameters. In the sample, we first estimate μ by x̄ and σ by s. Computing s requires the deviations xᵢ − x̄, and these n deviations always sum to zero. Consequently, knowing n − 1 of them determines the last one. Thus we have only n − 1 independent pieces of information, so df = n − 1.

自由度的概念可以理解为在估计参数之后剩余的独立信息量。在样本中,我们首先用 x̄ 估计 μ,用 s 估计 σ。计算 s 需要用到偏差 xᵢ − x̄,而这 n 个偏差之和恒为零。因此,只要知道其中 n − 1 个,最后一个就被完全确定了。所以我们只有 n − 1 个独立信息,故 df = n − 1。

For example, a sample of size 12 gives df = 11. In the t-table, we look along the row for df = 11 and the column for the desired tail probability. For a 95% confidence interval with two tails, the critical value is t* = 2.201. This is larger than the corresponding z* = 1.96, which is why the resulting interval is wider.

Published by TutorHao | IB Mathematics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading