Linear Regression | 线性回归

📚 Linear Regression | 线性回归

Linear regression is a fundamental statistical technique used to model the relationship between two quantitative variables. In A-Level Edexcel Mathematics, you will learn how to find the line of best fit using the least squares method, interpret its parameters, and assess the strength of the linear association. Mastery of this topic enables you to make predictions and test hypotheses about real-world data.

线性回归是用于建模两个定量变量之间关系的基本统计技术。在A-Level Edexcel数学中,你将学习如何使用最小二乘法找到最佳拟合直线,解释其参数,并评估线性关联的强度。掌握这一主题可以让你基于实际数据进行预测和假设检验。


1. Introduction to Linear Regression | 线性回归简介

Linear regression aims to describe how a response variable y depends on an explanatory variable x by fitting a straight line through a scatter of data points. The equation of the line is typically written as y = a + bx, where a is the intercept and b is the slope. The slope indicates how much y changes on average when x increases by one unit.

线性回归旨在通过拟合一条穿过数据散点的直线,来描述响应变量 y 如何依赖于解释变量 x。该直线的方程通常写作 y = a + bx,其中 a 是截距,b 是斜率。斜率表示当 x 增加一个单位时,y 平均变化多少。

In A-Level problems, you will often be given summary statistics such as Σx, Σy, Σx², Σy² and Σxy, and you must use these to construct the regression line. The most common model is the least squares regression line of y on x. It is important to distinguish this from the regression line of x on y, which is a different line used to predict x from y.

在A-Level问题中,你通常会得到诸如 Σx、Σy、Σx²、Σy² 和 Σxy 的汇总统计量,并需要利用这些量构造回归直线。最常见的模型是 y 关于 x 的最小二乘回归线。必须将其与 x 关于 y 的回归线区分开,后者是用于从 y 预测 x 的另一条直线。


2. The Least Squares Method | 最小二乘法

The least squares method finds the line that minimises the sum of the squared vertical distances from the data points to the line. These vertical distances are called residuals. By squaring them, we penalise larger deviations more heavily, and the resulting coefficients a and b have convenient mathematical properties.

最小二乘法通过最小化数据点到直线的垂直距离的平方和,来找到最佳直线。这些垂直距离被称为残差。通过对它们求平方,我们能更严厉地惩罚较大偏差,而且由此得到的系数 a 和 b 具有便利的数学性质。

The estimates for the slope b and intercept a are derived by solving two normal equations. However, in practice, you will directly use the formulas expressed in terms of sums of squares. This approach ensures that the line is unique and provides the best linear unbiased estimates under certain assumptions.

斜率 b 和截距 a 的估计值是通过求解两个正规方程推导出来的。但在实践中,你会直接使用以平方和形式表达的公式。这种方法确保了直线的唯一性,并在一定假设下提供了最佳线性无偏估计。


3. Calculating the Regression Line | 计算回归直线

To calculate the regression line of y on x from a set of n data pairs (xᵢ, yᵢ), first compute the means x̄ = Σxᵢ / n and ȳ = Σyᵢ / n. Then calculate the sums of squares:

要从一组 n 对数据 (xᵢ, yᵢ) 计算 y 关于 x 的回归直线,首先计算均值 x̄ = Σxᵢ / n 和 ȳ = Σyᵢ / n。然后计算平方和:

Sₓₓ = Σ(xᵢ – x̄)² = Σxᵢ² – (Σxᵢ)²/n

Sₓᵧ = Σ(xᵢ – x̄)(yᵢ – ȳ) = Σxᵢyᵢ – (Σxᵢ)(Σyᵢ)/n

Sᵧᵧ = Σ(yᵢ – ȳ)² = Σyᵢ² – (Σyᵢ)²/n

The slope is then b = Sₓᵧ / Sₓₓ and the intercept is a = ȳ – b x̄. The equation of the regression line is y = a + bx. Always remember to state the line with the correct context and appropriate rounding.

于是斜率 b = Sₓᵧ / Sₓₓ,截距 a = ȳ – b x̄。回归直线方程为 y = a + bx。务必在给出直线时提供正确的上下文和适当的四舍五入。

Many Edexcel exam questions give you the totals Σx, Σy, Σx², Σy² and Σxy directly. You should become fluent in substituting these into the simplified formulas. Using your calculator’s regression mode can also speed up calculations, but you must still be able to show the working for key steps.

许多Edexcel试题会直接给出 Σx、Σy、Σx²、Σy² 和 Σxy。你应当熟练地将它们代入简化公式。使用计算器的回归模式可以加快计算速度,但你必须仍能展示关键步骤的计算过程。


4. Interpreting Slope and Intercept | 解释斜率和截距

Interpreting the regression coefficients correctly is a common exam requirement. The intercept a is the predicted value of y when x = 0. This interpretation only makes sense if x = 0 lies within or near the observed data range; otherwise, a may be purely a mathematical anchor with no practical meaning.

正确解释回归系数是常见的考试要求。截距 a 是当 x = 0 时 y 的预测值。只有当 x = 0 位于或接近观测数据范围时,这种解释才有意义;否则,a 可能只是一个数学定位点,没有实际意义。

The slope b represents the estimated change in y for a one-unit increase in x. For example, if a line is y = 3.2 + 1.5x, then each additional unit of x is associated with an average increase of 1.5 units in y. Always state the units of both variables to give meaning to the numbers.

斜率 b 表示当 x 增加一个单位时 y 的估计变化量。例如,如果某条直线为 y = 3.2 + 1.5x,那么 x 每增加一个单位,y 平均增加 1.5 个单位。务必说明两个变量的单位,以赋予数字实际意义。

Be careful with the direction: a negative slope indicates that as x increases, y tends to decrease. The magnitude of b also tells you about the sensitivity of y to changes in x, but the strength of the relationship is better assessed by the correlation coefficient.

注意方向性:负斜率表示随着 x 增大,y 趋于减小。b 的大小也显示 y 对 x 变化的敏感度,但相关关系的强度更适合用相关系数来评估。


5. Residuals and Goodness of Fit | 残差与拟合优度

A residual is the difference between an observed y value and the value predicted by the regression line: eᵢ = yᵢ – ŷᵢ, where ŷᵢ = a + b xᵢ. Plotting residuals against x or against predicted values helps diagnose whether a linear model is appropriate. A random scatter of points around zero supports the linear model.

残差是观测到的 y 值与回归直线预测值之间的差值:eᵢ = yᵢ – ŷᵢ,其中 ŷᵢ = a + b xᵢ。将残差相对于 x 或相对于预测值绘图,有助于诊断线性模型是否合适。点围绕零随机散布就支持线性模型。

The coefficient of determination R² is the proportion of the variation in y that is explained by the regression on x. It is calculated as R² = (Sₓᵧ)² / (Sₓₓ × Sᵧᵧ) and equals the square of the Pearson correlation coefficient r. An R² close to 1 indicates a strong linear fit, while a value near 0 suggests a weak linear relationship.

决定系数 R² 是 y 的变异中被 x 的回归所解释的比例。计算公式为 R² = (Sₓᵧ)² / (Sₓₓ × Sᵧᵧ),且等于皮尔逊相关系数 r 的平方。R² 接近 1 表示线性拟合良好,接近 0 则表示线性关系很弱。

In Edexcel questions, you may be asked to compute R², interpret it, or comment on the validity of the linear model based on a residual plot. Remember that a high R² does not prove causation; it merely indicates how well the line fits the observed data.

在Edexcel题目中,你可能需要计算 R²、解释其含义,或基于残差图评论线性模型的有效性。请记住,高 R² 并不能证明因果关系;它仅仅表示直线对观测数据的拟合程度。


6. Making Predictions | 进行预测

Once you have a reliable regression line, you can use it to estimate the expected value of y for a given x. Simply substitute the x value into the equation y = a + bx. Predictions within the original range of x (interpolation) are generally more trustworthy than predictions outside that range (extrapolation).

一旦得到了可靠的回归直线,你就可以用它来估计给定 x 时期望的 y 值。只需将 x 值代入方程 y = a + bx 即可。在原 x 数据范围内的预测(内插)通常比范围外的预测(外推)更可靠。

Extrapolation can be dangerous because the linear trend observed in the data may not continue beyond the sample. Exam questions frequently ask you to state why a prediction might be unreliable, and the answer is often because it involves extrapolation or because the value of x lies far from the data range.

外推可能很危险,因为数据中观察到的线性趋势可能在样本范围之外不再继续。考题经常要求你陈述某个预测为何不可靠,答案通常是:该预测涉及外推,或者 x 的值远在数据范围之外。

When making multiple predictions, be aware of the uncertainty involved. The regression line gives a point estimate, but the actual value will vary due to random error. Confidence intervals for the mean response and prediction intervals for an individual value are more advanced, but understanding variability is important.

在进行多次预测时,要意识到其中包含的不确定性。回归直线给出的是点估计值,但实际值会因随机误差而变化。平均响应的置信区间和个体值的预测区间虽然更深入,但理解变异性非常重要。


7. Correlation and Regression | 相关性与回归

The Pearson product-moment correlation coefficient r measures the strength and direction of a linear relationship. It is defined as r = Sₓᵧ / √(Sₓₓ × Sᵧᵧ), and it always lies between -1 and 1. A positive r corresponds to a positive slope, and a negative r to a negative slope.

皮尔逊积矩相关系数 r 衡量线性关系的强度和方向。其定义为 r = Sₓᵧ / √(Sₓₓ × Sᵧᵧ),并且总是位于 -1 和 1 之间。正 r 对应正斜率,负 r 对应负斜率。

It is essential to understand that correlation does not imply causation. Two variables may move together because of a lurking third variable or by coincidence. In the exam you might be asked to explain why a strong correlation does not prove that x causes y.

必须理解相关性并不意味着因果关系。两个变量可能因为隐藏的第三变量或巧合而共同变动。考试中可能会要求你解释为什么强相关并不能证明 x 导致 y。

The slope b of the regression line is linked to r by the relationship b = r × (sᵧ / sₓ), where sₓ and sᵧ are the sample standard deviations. Thus, a zero correlation implies a slope of zero, meaning no linear relationship between the variables.

回归直线的斜率 b 与 r 的关系为 b = r × (sᵧ / sₓ),其中 sₓ 和 sᵧ 是样本标准差。因此,零相关意味着斜率为零,即变量之间没有线性关系。


8. Hypothesis Testing for Slope | 斜率假设检验

In A-Level Edexcel Statistics, you may carry out a t‑test to determine whether there is evidence of a linear relationship in the population. The null hypothesis is usually H₀: β = 0 (no linear relationship) against H₁: β ≠ 0. Here β represents the true population slope.

在A-Level Edexcel统计中,你可以进行 t 检验,以确定是否有证据表明总体中存在线性关系。原假设通常为 H₀: β = 0(无线性关系),备择假设为 H₁: β ≠ 0。这里 β 代表真实的总体斜率。

The test statistic is t = b / SE(b), where SE(b) is the standard error of the slope. The standard error is usually obtained from computer output or given in the exam. The degrees of freedom for the t‑distribution is n – 2.

检验统计量为 t = b / SE(b),其中 SE(b) 是斜率的标准误。标准误通常从计算机输出获得或在考试中直接给出。t 分布的自由度为 n – 2。

You compare the calculated t‑value with a critical value from the t‑table or use a given p‑value. If the p‑value is less than the significance level (commonly 5%), you reject H₀ and conclude there is sufficient evidence of a linear relationship.

你要将计算得到的 t 值与 t 分布表中的临界值进行比较,或使用给定的 p 值。如果 p 值小于显著性水平(通常为 5%),则拒绝 H₀,得出有充分证据表明存在线性关系的结论。

Sometimes the test is presented via an analysis of variance (ANOVA) table, but the underlying concept is the same: assessing whether the regression model explains a significant amount of the variation in y.

有时该检验通过方差分析(ANOVA)表呈现,但基本概念是相同的:评估回归模型是否解释了 y 变异的一个显著部分。


9. Assumptions of Linear Regression | 线性回归假设

For the least squares regression line to provide valid inference, several assumptions should be satisfied. The relationship between x and y should be approximately linear; this can be checked with a scatter plot. The residuals should be independent and normally distributed with a constant variance (homoscedasticity).

要使最小二乘回归线提供有效的推断,必须满足几个假设。x 和 y 之间的关系应近似线性;这可以用散点图来检查。残差应当是独立的、服从正态分布且具有恒定方差(同方差性)。

Independence means that one observation does not influence another; this is often ensured by good study design. Constant variance implies that the spread of residuals is similar across all levels of x. A funnel‑shaped residual plot indicates heteroscedasticity, which can make standard errors unreliable.

独立性意味着一个观测值不会影响另一个;这通常通过良好的研究设计来保证。恒定方差意味着残差的散布在所有 x 水平上都相似。呈漏斗形的残差图表明存在异方差性,这会使标准误不可靠。

Normality of residuals is important for small samples when conducting hypothesis tests or constructing confidence intervals. For large samples, the central limit theorem makes the t‑procedures robust to moderate departures from normality.

残差的正态性在进行假设检验或构造置信区间时对小样本很重要。对于大样本,中心极限定理使得 t 统计方法对中等程度的正态性偏离具有稳健性。


10. Using Technology Effectively | 有效使用技术

Modern calculators, such as the Casio fx‑991EX or TI‑84, can quickly compute the regression line, r and R² once you enter the data. Spreadsheet software like Excel also provides regression output with detailed statistics. In the exam, you are expected to use these tools to check your manual calculations and to interpret computer output.

现代计算器,如卡西欧 fx‑991EX 或 TI‑84,一旦输入数据,就可以快速计算回归直线、r 和 R²。像 Excel 这样的电子表格软件也能提供包含详细统计量的回归输出。在考试中,要求能够使用这些工具检查你的手动计算并解读计算机输出。

When using a calculator, always ensure the data are entered correctly and that you have selected the linear regression model. You should also understand the command words: ‘regression line of y on x’ tells you which variable is the response. If the question asks for the regression line of x on y, swap the roles.

使用计算器时,务必确保数据输入正确,并选择线性回归模型。你还应理解指令用语:’y 关于 x 的回归直线’ 告诉你哪个变量是响应变量。如果问题要求 x 关于 y 的回归直线,则交换角色。

Exam papers sometimes provide extracts from software output. You need to identify the slope, intercept, R², and the standard error or t‑statistic from that output and use them in the required context. Practise reading such tables under timed conditions.

试卷有时会提供软件输出的摘录。你需要从中识别斜率、截距、R² 以及标准误或 t 统计量,并在所要求的上下文中使用它们。建议在限时条件下练习阅读这类表格。


11. Common Mistakes and Exam Tips | 常见错误与考试技巧

One of the most frequent mistakes is confusing the regression line of y on x with that of x on y. Only use the correct formula based on which variable you want to predict. Also, do not forget to label axes and provide units when interpreting the slope and intercept in a real‑world context.

最常见的错误之一是混淆 y 关于 x 的回归线与 x 关于 y 的回归线。一定要根据你想预测哪个变量,使用正确的公式。此外,在现实情境中解释斜率和截距时,不要忘记标注坐标轴并给出单位。

Avoid making statements that imply causation when a regression only shows association. Always state that the slope indicates an estimated average change, not an exact deterministic change. When computing residuals, take care to subtract the predicted value from the observed value consistently.

避免在回归仅显示关联性时,做出暗示因果关系的表述。务必指出斜率表示的是估计的平均变化,而不是精确的确定性变化。计算残差时,注意始终用观测值减去预测值。

  • Double‑check the use of n: the denominator in Sₓₓ is n when using the simplified formula, but the formulas can be mixed up with population variance formulas.
  • 核对 n 的使用:使用简化公式时 Sₓₓ 的分母是 n,但这些公式可能会与总体方差公式混淆。
  • When reading from a residual plot, look for patterns that suggest non‑linearity or non‑constant variance, and mention them clearly.
  • 阅读残差图时,要寻找表明非线性或非恒定方差的模式,并清楚地指出。
  • If a question asks for a ‘reliable’ prediction, check the x‑value against the range of the data; if it lies outside, highlight extrapolation risk.
  • 如果问题要求“可靠”的预测,请检查 x 值是否在数据范围内;如果在范围外,要强调外推的风险。

Finally, present your answers clearly, showing the regression equation in the form y = a + bx, and include the correlation coefficient or R² when asked. Round coefficients sensibly — usually to three significant figures, unless the question specifies otherwise.

最后,清晰地呈现答案,以 y = a + bx 的形式展示回归方程,并在要求时附上相关系数或 R²。合理地对系数进行四舍五入——通常保留三位有效数字,除非题目另有要求。


12. Summary and Revision Focus | 总结与复习重点

Linear regression is a core topic that ties together scatter plots, correlation, least squares algebra, and inferential statistics. Key skills include computing the regression line, interpreting coefficients, checking model adequacy with residuals, and understanding the link between r and b.

线性回归是一个将散点图、相关性、最小二乘代数和推断统计结合在一起的核心主题。关键技能包括计算回归直线、解释系数、用残差检验模型适当性,以及理解 r 与 b 之间的联系。

For exam success, practise both manual calculations with summary statistics and interpreting software output. Be ready to comment on the reliability of predictions and to perform or interpret a t‑test for the slope. A solid grasp of this topic will also support further study in hypothesis testing and bivariate data analysis.

为了在考试中取得成功,既要练习使用汇总统计量进行手动计算,也要练习解读软件输出。准备好评论预测的可靠性,并进行或解释斜率的 t 检验。牢固掌握这一主题,也将为后续学习假设检验和双变量数据分析打下坚实基础。

Revisit the definitions of Sₓₓ, Sₓᵧ, Sᵧᵧ and the formulas for a, b, r and R². Make sure you can derive the regression line efficiently. With consistent practice, linear regression can become one of the most reliable marks in your A‑Level exam.

重温 Sₓₓ、Sₓᵧ、Sᵧᵧ 的定义以及 a、b、r 和 R² 的公式。确保自己能高效地推导出回归直线。通过持续练习,线性回归可以成为 A‑Level 考试中最稳定的得分点之一。

Published by TutorHao | Mathematics Revision Series | aleveler.com

Find Edexcel A Level Maths Textbooks on eBay UK

New, used and second-hand copies of textbooks and revision guides are often much cheaper than retail — check current listings and prices before you buy.

Browse on eBay UK →

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading