Hypothesis Testing for the Difference Between Means | 两个均值差异的假设检验

📚 Hypothesis Testing for the Difference Between Means | 两个均值差异的假设检验

Hypothesis testing for the difference between means compares the population means of two independent groups. It allows you to decide whether an observed difference in sample means is statistically significant or can be explained by sampling variability.

两个均值差异的假设检验用于比较两个独立总体的均值。它帮助你判断样本均值之间的差异是否具有统计显著性,还是可以归因于抽样波动。


1. What a Two-Sample Mean Test Is | 什么是双样本均值检验

In a two-sample mean test, we take independent random samples from two populations and calculate the difference between the sample means. The test asks whether the true population means differ by a specified amount, usually zero. This is different from a one-sample test, which compares one sample mean to a known or claimed value.

在双样本均值检验中,我们从两个总体中分别抽取独立随机样本,并计算两个样本均值之差。检验的目的是判断两个总体均值是否存在指定大小的差异,通常是零。这与单样本检验不同,后者将单个样本均值与已知或声称的数值进行比较。


2. When to Use This Test | 何时使用该检验

Use a two-sample test when the data come from two independent groups. Examples include comparing the average scores of students taught by two methods, comparing the mean lifetimes of components from two factories, or comparing the mean blood pressure reduction from two drugs. If the data are paired, such as before-and-after measurements on the same individuals, you should use a paired test instead.

当数据来自两个独立组时,应使用双样本检验。例如,比较两种教学方法下学生的平均成绩、比较两个工厂生产的部件平均寿命,或比较两种药物对血压的平均降低效果。如果数据是配对的,例如同一组人的前后测量,则应使用配对检验。


3. Setting Up Null and Alternative Hypotheses | 建立零假设与备择假设

The null hypothesis usually states that there is no difference between the two population means. If a non-zero difference d is being tested, the null hypothesis is μ₁ – μ₂ = d. The alternative hypothesis can be one-tailed or two-tailed. For a two-tailed test, H₁: μ₁ – μ₂ ≠ d. For one-tailed tests, H₁: μ₁ – μ₂ > d or H₁: μ₁ – μ₂ < d.

零假设通常陈述两个总体均值之间没有差异。如果要检验非零差异 d,则零假设为 μ₁ – μ₂ = d。备择假设可以是单侧或双侧的。对于双侧检验,H₁: μ₁ – μ₂ ≠ d;对于单侧检验,H₁: μ₁ – μ₂ > d 或 H₁: μ₁ – μ₂ < d。

At A level, you will normally test μ₁ – μ₂ = 0. The significance level, usually 5% or 1%, must be stated before the test.

在 A level 中,通常检验 μ₁ – μ₂ = 0。显著性水平,通常为 5% 或 1%,必须在检验前说明。


4. Key Assumptions and Conditions | 关键假设与条件

For a valid test, the two samples must be independent random samples. Each population should be normally distributed, or the sample sizes should be large enough for the Central Limit Theorem to apply. If population variances are known, a z-test can be used. If population variances are unknown, a t-test is appropriate under certain conditions.

为了检验有效,两个样本必须是独立的随机样本。每个总体应服从正态分布,或者样本量足够大,使中心极限定理适用。如果总体方差已知,可使用 z 检验;如果总体方差未知,在满足一定条件时应使用 t 检验。


5. The Test Statistic for Known Variance | 已知方差下的检验统计量

When the population variances σ₁² and σ₂² are known, the test statistic is calculated from the difference in sample means x̄₁ – x̄₂. The standard error of the difference combines the two sampling variances. The formula is:

当总体方差 σ₁² 和 σ₂² 已知时,检验统计量由样本均值之差 x̄₁ – x̄₂ 计算。差值的标准误合并了两个抽样方差。公式为:

z = ( (x̄₁ – x̄₂) – d ) / √(σ₁²/n₁ + σ₂²/n₂)

Here d is the hypothesised difference, usually 0, and n₁ and n₂ are the sample sizes. Because the samples are independent, the variances add. You do not subtract the variances.

这里 d 是假设的差异,通常为 0,n₁ 和 n₂ 是样本容量。由于样本是独立的,方差相加,而不是相减。


6. Critical Regions and p-Values | 临界区域与 p 值

For a two-tailed test at significance level α, reject H₀ if z < -z(α/2) or z > z(α/2). For a one-tailed test with H₁: μ₁ – μ₂ > d, reject H₀ if z > z(α). With H₁: μ₁ – μ₂ < d, reject H₀ if z < -z(α). Alternatively, compare the p-value with α and reject H₀ when the p-value is less than α.

对于显著性水平为 α 的双侧检验,当 z < -z(α/2) 或 z > z(α/2) 时拒绝 H₀。对于 H₁: μ₁ – μ₂ > d 的单侧检验,当 z > z(α) 时拒绝 H₀;对于 H₁: μ₁ – μ₂ < d,当 z < -z(α) 时拒绝 H₀。也可以将 p 值与 α 比较,当 p 值小于 α 时拒绝 H₀。

The critical values come from the standard normal distribution. Common two-tailed critical values are 1.96 at the 5% level and 2.576 at the 1% level.

临界值来自标准正态分布。常用的双侧临界值为 5% 水平下的 1.96 和 1% 水平下的 2.576。


7. Worked Example 1: Known Variance | 示例 1:已知方差

A company compares the mean weight of two cereal filling lines. Sample 1 has n₁ = 40, x̄₁ = 82.4 g, σ₁ = 6.0 g. Sample 2 has n₂ = 45, x̄₂ = 79.1 g, σ₂ = 5.5 g. Test at the 5% significance level whether the population means differ. Let d = 0.

一家公司比较两条谷物灌装线的平均重量。样本 1:n₁ = 40,x̄₁ = 82.4 g,σ₁ = 6.0 g;样本 2:n₂ = 45,x̄₂ = 79.1 g,σ₂ = 5.5 g。在 5% 显著性水平下检验总体均值是否有差异,设 d = 0。

The standard error is √(6.0²/40 + 5.5²/45) = √(0.900 + 0.672) = √1.572 = 1.254. The test statistic is z = (82.4 – 79.1) / 1.254 = 2.632.

标准误为 √(6.0²/40 + 5.5²/45) = √(0.900 + 0.672) = √1.572 = 1.254。检验统计量为 z = (82.4 – 79.1) / 1.254 = 2.632。

The two-tailed critical value at 5% is 1.96. Since 2.632 > 1.96, we reject H₀. There is sufficient evidence at the 5% level to suggest the two lines have different mean filling weights.

5% 显著性水平下的双侧临界值为 1.96。因为 2.632 > 1.96,所以拒绝 H₀。在 5% 水平上有充分证据表明两条灌装线的平均灌装重量不同。


8. Unknown Variance and the Two-Sample t-Test | 未知方差与双样本 t 检验

When the population variances are unknown, the z procedure cannot be used directly. If it is reasonable to assume the two population variances are equal, we use a pooled two-sample t-test. The pooled estimate of variance is:

当总体方差未知时,不能直接使用 z 方法。如果可以合理假设两个总体方差相等,则使用合并的双样本 t 检验。合并方差估计为:

s²ₚ = [(n₁ – 1)s₁² + (n₂ – 1)s₂²] / (n₁ + n₂ – 2)

The test statistic is then:

检验统计量为:

t = ( (x̄₁ – x̄₂) – d ) / (sₚ √(1/n₁ + 1/n₂))

The degrees of freedom are n₁ + n₂ – 2. This t-test is reliable when both samples come from normal populations with equal variances. If variances are not equal, a separate variance approximation may be used, but this is beyond the usual A level requirement.

自由度为 n₁ + n₂ – 2。当两个样本来自方差相等的正态总体时,该 t 检验是可靠的。如果方差不相等,可以使用单独方差的近似方法,但这通常超出 A level 的要求范围。


9. Worked Example 2: Unknown Equal Variance | 示例 2:未知但相等的方差

Two revision apps are compared. Sample 1: n₁ = 12, x̄₁ = 54.6, s₁ = 4.2. Sample 2: n₂ = 14, x̄₂ = 50.1, s₂ = 3.9. Test at the 5% level whether the mean improvement differs. Assume equal population variances.

比较两个复习应用程序。样本 1:n₁ = 12,x̄₁ = 54.6,s₁ = 4.2;样本 2:n₂ = 14,x̄₂ = 50.1,s₂ = 3.9。在 5% 水平下检验平均改善是否有差异,假设总体方差相等。

The pooled variance is s²ₚ = [(11)(17.64) + (13)(15.21)] / 24 = (194.04 + 197.73) / 24 = 16.324. Thus sₚ = 4.040.

合并方差为 s²ₚ = [(11)(17.64) + (13)(15.21)] / 24 = (194.04 + 197.73) / 24 = 16.324,因此 sₚ = 4.040。

The standard error is 4.040 × √(1/12 + 1/14) = 4.040 × √0.15476 = 4.040 × 0.3934 = 1.589. The test statistic is t = (54.6 – 50.1) / 1.589 = 2.832.

标准误为 4.040 × √(1/12 + 1/14) = 4.040 × √0.15476 = 4.040 × 0.3934 = 1.589。检验统计量为 t = (54.6 – 50.1) / 1.589 = 2.832。

With 24 degrees of freedom, the two-tailed 5% critical value from t-tables is 2.064. Since 2.832 > 2.064, reject H₀. There is sufficient evidence at the 5% level to suggest the mean improvements are different.

自由度为 24,t 表中双侧 5% 临界值为 2.064。因为 2.832 > 2.064,所以拒绝 H₀。在 5% 水平上有充分证据表明平均改善不同。


10. Testing a Non-Zero Difference d | 检验非零差异 d

Sometimes the question specifies a non-zero difference. For example, H₀: μ₁ – μ₂ = 3 against H₁: μ₁ – μ₂ ≠ 3. The same formulas are used, but d is not zero. The hypothesised value d must be subtracted from x̄₁ – x̄₂ before dividing by the standard error.

有时题目会指定非零差异。例如,H₀: μ₁ – μ₂ = 3,备择假设 H₁: μ₁ – μ₂ ≠ 3。使用相同的公式,但 d 不是零。在除以标准误之前,必须从 x̄₁ – x̄₂ 中减去假设值 d。

Always check the wording carefully. If the question asks whether one mean exceeds another by more than a certain amount, set d equal to that amount.

始终仔细审题。如果题目问一个均值是否比另一个均值高出超过特定数值,则将 d 设定为该数值。


11. Interpreting Conclusions in Context | 在具体情境中解释结论

Your conclusion must be written in the context of the problem. Do not simply write ‘reject H₀’. For example, write: ‘At the 5% significance level, there is sufficient evidence to suggest that the mean weight produced by line A is greater than line B.’ If you do not reject H₀, write: ‘There is insufficient evidence to suggest a difference in means.’

结论必须结合问题情境来写。不要只写“拒绝 H₀”。例如,应写:“在 5% 显著性水平下,有充分证据表明 A 线生产的平均重量大于 B 线。”如果未拒绝 H₀,应写:“没有充分证据表明两个均值存在差异。”

The phrase ‘sufficient evidence’ is important in A level marking. Avoid saying ‘prove’ or ‘accept H₀’. You can only fail to reject H₀ when evidence is insufficient.

在 A level 评分中,“充分证据”这一措辞很重要。避免使用“证明”或“接受 H₀”。当证据不足时只能“不拒绝 H₀”。


12. Common Mistakes and Exam Tips | 常见错误与考试技巧

One common mistake is subtracting the two variances instead of adding them in the standard error. Remember that for independent samples, variances add. Another mistake is forgetting to subtract d when the hypothesized difference is not zero. Also, using a paired test for independent samples or a z-test for small samples with unknown variance can lead to incorrect conclusions.

一个常见错误是在标准误中把两个方差相减而不是相加。记住,对于独立样本,方差相加。另一个错误是当假设差异不为零时忘记减去 d。此外,对独立样本使用配对检验,或对方差未知的小样本使用 z 检验,都会导致错误的结论。

Before starting, state H₀, H₁ and the significance level. Check whether variances are known or unknown. Show your working clearly, including the standard error and test statistic. Finally, write your conclusion in context using the phrase ‘sufficient evidence at the 5% level’.

开始前,先写出 H₀、H₁ 和显著性水平。检查方差已知还是未知。清晰展示计算过程,包括标准误和检验统计量。最后,在具体情境中用“在 5% 水平上有充分证据”的句式写出结论。

Summary Table | 总结表

Known variances Use z-test: z = ((x̄₁ – x̄₂) – d) / √(σ₁²/n₁ + σ₂²/n₂)
Unknown but equal variances Use pooled t-test with df = n₁ + n₂ – 2
Independent samples Add variances in standard error
Paired samples Do not use this test; use paired t-test

Published by TutorHao | Mathematics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading