Regression, Correlation and Hypothesis Testing | 回归、相关与假设检验

📚 Regression, Correlation and Hypothesis Testing | 回归、相关与假设检验

Regression and correlation are fundamental statistical tools for exploring relationships between two variables. Combined with hypothesis testing, they allow us to make formal inferences about the strength and significance of these relationships. In the Edexcel A-Level Mathematics syllabus, this topic brings together scatter diagrams, the product moment correlation coefficient (PMCC), Spearman’s rank correlation coefficient, least squares regression, and tests for correlation using Student’s t-distribution.

回归和相关是探索两个变量之间关系的基本统计工具。结合假设检验,我们能够对这些关系的强度和显著性做出正式推断。在 Edexcel A-Level 数学大纲中,这一主题涵盖了散点图、积矩相关系数(PMCC)、斯皮尔曼等级相关系数、最小二乘回归以及利用 t 分布对相关性进行检验。


1. Scatter Diagrams and Types of Correlation | 散点图和相关类型

A scatter diagram displays paired data (x, y) as points on a Cartesian plane. It provides a visual first step in assessing the relationship between two variables. The pattern of points indicates whether there is correlation and, if so, its direction and form.

散点图将成对数据 (x, y) 以点的形式绘制在笛卡尔平面上,为判断两个变量之间的关系提供直观的第一步。点的分布形态揭示了是否存在相关性,以及相关的方向和形式。

Positive correlation occurs when y tends to increase as x increases, giving an overall upward slope. Negative correlation is seen when y tends to decrease as x increases, producing a downward slope. If no clear pattern emerges, the variables may be uncorrelated. The relationship may be linear, but scatter diagrams can also reveal non-linear (curvilinear) associations.

当 y 随 x 增大而增大时,整体呈向上趋势,称为正相关。当 y 随 x 增大而减小时,整体倾斜向下,为负相关。如果没有明显的规律,则变量可能不相关。相关关系可以是线性的,但散点图也可能显示出非线性(曲线)关联。


2. Product Moment Correlation Coefficient (PMCC) | 积矩相关系数

The PMCC, denoted by r, is a numerical measure of linear correlation between two variables. It is calculated from a sample and always lies between –1 and +1. The closer r is to +1, the stronger the positive linear correlation; the closer it is to –1, the stronger the negative linear correlation. An r value near 0 suggests very weak or no linear correlation.

积矩相关系数,记为 r,是衡量两个变量之间线性相关程度的数值指标。它由样本计算得出,始终介于 –1 和 +1 之间。r 越接近 +1,正线性相关越强;越接近 –1,负线性相关越强。r 接近 0 则表示线性相关很弱或没有。

It is important to remember that PMCC only measures linear association. A value of r close to zero does not necessarily mean the variables are independent; they could be related in a non-linear way.

必须牢记,PMCC 只度量线性关联。r 接近于零并不一定意味着变量相互独立;它们可能以非线性方式相关。


3. Calculating PMCC | 计算 PMCC

For a sample of n pairs (xi, yi), the PMCC is given by the formula

r = Sxy / √(Sxx × Syy)

对于样本容量为 n 的配对数据 (xi, yi),积矩相关系数的计算公式为

r = Sxy / √(Sxx × Syy)

where

Sxx = Σx² – (Σx)²/n,
Syy = Σy² – (Σy)²/n,
Sxy = Σxy – (Σx)(Σy)/n

其中

Sxx = Σx² – (Σx)²/n,
Syy = Σy² – (Σy)²/n,
Sxy = Σxy – (Σx)(Σy)/n

In exam problems, you are often given these sums or need to compute them precisely. Using a table to organise Σx, Σy, Σx², Σy² and Σxy helps reduce errors.

在考试题中,通常会直接给出这些总和,或要求准确计算。利用表格整理 Σx、Σy、Σx²、Σy² 和 Σxy 有助于减少计算错误。


4. Interpreting the Correlation Coefficient | 相关系数的解释

The value of r describes both the strength and the direction of a linear relationship. Informally, we may use thresholds such as 0.7 and above for strong positive correlation, but such rules are only rough guides. The interpretation must always be made in the context of the data.

相关系数 r 的取值同时描述了线性关系的强度和方向。非正式地,我们可能采用 0.7 及以上表示强正相关等经验阈值,但这些规则只是粗略指导。解读时必须始终结合数据背景。

A large |r| indicates that points lie close to a straight line; a small |r| indicates a large amount of scatter. However, a single outlier can greatly inflate or deflate r, so it is always wise to inspect the scatter diagram before trusting the numerical result.

较大的 |r| 表明数据点紧密分布在一条直线附近;较小的 |r| 则反映较大的离散程度。但单个异常值可能显著增大或减小 r,因此在相信数值结果之前,始终先检查散点图是明智的做法。


5. Spearman’s Rank Correlation Coefficient | 斯皮尔曼等级相关系数

When data are not linear or are based on ranks rather than actual measurements, Spearman’s rank correlation coefficient, rs, is more appropriate. It measures the strength of monotonic association. The formula for data with no tied ranks is

rs = 1 – 6Σd² / [n(n² – 1)]

当数据不是线性的,或基于秩次而非实际测量值时,斯皮尔曼等级相关系数 rs 更为合适。它衡量的是单调关联的强度。对于没有并列秩次的数据,公式为

rs = 1 – 6Σd² / [n(n² – 1)]

Here d is the difference between the ranks of each pair. rs also lies between –1 and +1 and is interpreted in a similar way to PMCC, though it does not require linearity.

其中 d 为每一对数据的秩次差。rs 同样落在 –1 到 +1 之间,其解释方式与 PMCC 类似,但不要求线性。

To calculate rs, rank each variable separately, compute differences d, square them, sum them, and substitute into the formula. When ties occur, other corrections are applied, though these are rarely examined at A-Level.

计算 rs 时,分别对每个变量排秩,计算差值 d,平方并求和,再代入公式。当出现并列秩次时,需使用修正公式,但在 A-Level 考试中较少涉及。


6. Least Squares Regression Line | 最小二乘回归线

A regression line describes how the response (dependent) variable y changes with the explanatory (independent) variable x. In A-Level, we model y on x using the least squares method, minimising the sum of squared vertical distances from points to the line.

回归线描述了响应变量(因变量)y 如何随解释变量(自变量)x 变化。在 A-Level 中,我们使用最小二乘法建立 y 关于 x 的回归模型,使各点到直线的竖直距离平方和最小。

The equation of the regression line is y = a + bx, where

b = Sxy / Sxx, a = ȳ – b x̄

回归线方程为 y = a + bx,其中

b = Sxy / Sxx, a = ȳ – b x̄

Here, b is the gradient (slope) and a is the y-intercept. It is essential to use the correct sums already calculated for the PMCC.

这里 b 是斜率,a 是 y 轴截距。必须使用先前计算 PMCC 时已求出的各项总和。


7. Interpreting the Regression Equation | 回归方程的解释

The gradient b gives the estimated change in y for a one-unit increase in x. The intercept a is the predicted value of y when x = 0, though this may not have a meaningful real-world interpretation if x = 0 is outside the observed range.

斜率 b 表示 x 每增加一个单位时 y 的估计变化量。截距 a 是 x = 0 时 y 的预测值,但如果 x = 0 超出观测范围,这一截距可能没有现实意义。

A regression line should only be used to predict y for values of x that lie within or very close to the original data range. Extrapolation far beyond the data can lead to unreliable or nonsensical predictions.

回归线只应用于预测位于原始数据范围内或非常接近该范围的 x 值对应的 y。进行远超出数据范围的外推可能导致不可靠或荒谬的预测。

The regression line always passes through the mean point (x̄, ȳ). This fact can be used as a quick check of calculations.

回归线总是通过均值点 (x̄, ȳ)。这一性质可用于快速验算。


8. Residuals and Goodness of Fit | 残差与拟合优度

A residual ei is the difference between an observed yi and the value predicted by the regression line: ei = yi – (a + b xi). Analysing residuals helps assess whether a linear model is appropriate.

残差 ei 是观测值 yi 与回归线预测值之间的差值:ei = yi – (a + b xi)。分析残差有助于判断线性模型是否合适。

If a scatter plot of residuals against x shows a random pattern with no clear trend, the linear model is likely suitable. A curved pattern or changing spread suggests a different model or transformation is needed.

若残差对 x 的散点图呈现随机分布,没有明显趋势,则线性模型很可能合适。若出现弯曲模式或离散程度变化,则提示需要不同的模型或数据变换。

The sum of residuals is always zero for the least squares line. Some exam questions may ask you to calculate residuals and interpret them.

对于最小二乘回归线,残差之和始终为零。某些考题可能会要求计算残差并加以解释。


9. Hypothesis Testing for Correlation (Testing ρ = 0) | 相关性假设检验(检验 ρ = 0)

We often wish to test whether a population correlation coefficient ρ is zero, using the sample PMCC r. For a bivariate Normal population, the test statistic under H0 : ρ = 0 is

t = r √(n – 2) / √(1 – r²)

我们常希望利用样本 PMCC r 检验总体相关系数 ρ 是否为零。对于二元正态总体,H0 : ρ = 0 下的检验统计量为

t = r √(n – 2) / √(1 – r²)

This t-statistic follows a t-distribution with (n – 2) degrees of freedom. The null hypothesis states there is no linear correlation in the population; the alternative hypothesis can be one-tailed (ρ > 0 or ρ < 0) or two-tailed (ρ ≠ 0).

该 t 统计量服从自由度为 (n – 2) 的 t 分布。零假设表示总体中不存在线性相关;备择假设可以是单尾的(ρ > 0 或 ρ < 0)或双尾的(ρ ≠ 0)。

A similar test can be performed for Spearman’s rank coefficient using the same t formula, as long as the sample is large enough or we use critical values from appropriate tables.

对斯皮尔曼等级相关系数也可以使用相同的 t 公式进行类似检验,前提是样本足够大,或直接查阅相应表格中的临界值。


10. One-Tailed and Two-Tailed Tests | 单尾与双尾检验

Choosing between a one-tailed and a two-tailed test depends on the research question. A two-tailed test is used when we simply want to detect any correlation, positive or negative. A one-tailed test is used when we have a specific direction in mind, such as ‘positive correlation exists’.

选择单尾还是双尾检验取决于研究问题。当我们只是想检测是否存在任何方向的相关性时,使用双尾检验。若事先有明确的方向性预期,例如“存在正相关”,则使用单尾检验。

In a one-tailed test for H1 : ρ > 0, the critical value tcrit is taken from the upper tail; for H1 : ρ < 0, the lower tail is used. In a two-tailed test, the significance level is split equally between both tails. The conclusion must be phrased in terms of the original problem.

对于 H1 : ρ > 0 的单尾检验,临界值 tcrit 取自上尾;H1 : ρ < 0 则取用下尾。双尾检验中将显著性水平等分至两侧尾部。结论必须用原问题的语境来表述。


11. Using p-Values and Critical Values | 使用 p 值与临界值

You may be asked to use either critical values or p-values to decide the outcome of the test. For a given significance level α, if |t| exceeds the critical value from t(n – 2), reject H0. Alternatively, if the p-value is less than α, reject H0.

题目可能会要求使用临界值或 p 值来判断检验结果。对于给定的显著性水平 α,若 |t| 超过来自 t(n – 2) 分布的临界值,则拒绝 H0。或者,若 p 值小于 α,则拒绝 H0。

Many A-Level questions provide a table of critical values for the PMCC. You can compare your calculated r directly with the critical value for the given n and α. If |r| exceeds the tabulated value, there is sufficient evidence to reject H0.

许多 A-Level 题目会提供 PMCC 的临界值表。此时可直接将计算出的 r 与给定 n 和 α 下的临界值比较。若 |r| 超过表中数值,则有足够证据拒绝 H0。


12. Assumptions and Cautions | 假设与注意事项

The test for ρ = 0 using the t-statistic assumes that the data come from a bivariate Normal distribution and that the observations are independent. While the test is reasonably robust to moderate departures from Normality, extreme skewness or outliers can undermine its validity.

使用 t 统计量检验 ρ = 0 的假设包括:数据来自二元正态分布,且观测值相互独立。虽然该检验对中等程度偏离正态性具有一定稳健性,但严重偏态或异常值会削弱其有效性。

An apparently significant correlation does not imply causation. A high r could be due to a lurking third variable or purely coincidental. Always combine statistical findings with the context of the data.

显著相关并不意味着因果关系。较高的 r 可能源于某个隐藏的第三变量,或纯属巧合。务必结合数据背景来解释统计学发现。

When using linear regression, check that the relationship is roughly linear and that points are not heavily clustered or influential. The presence of outliers can distort both the slope and the correlation coefficient.

使用线性回归时,应确认关系大致是线性的,且点没有过度聚集或存在强影响点。异常值的存在可能扭曲斜率和相关系数。


Published by TutorHao | Mathematics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version