Mastering Edexcel S1 Statistics: Complete Guide | 精通 Edexcel S1 统计:完整指南

📚 Mastering Edexcel S1 Statistics: Complete Guide | 精通 Edexcel S1 统计:完整指南

This comprehensive guide covers the entire Edexcel S1 (Statistics 1) syllabus, often taken as part of A-Level Mathematics or Further Mathematics. We break down every chapter from the official Edexcel S1 textbook, providing clear explanations, key formulae, and bilingual study notes to help you master probability, data handling, distributions, correlation, and regression.

本完整指南涵盖整个 Edexcel S1(统计学 1)教学大纲,该模块常作为 A-Level 数学或进阶数学的一部分选考。我们按照 Edexcel S1 官方教材的章节逐一拆解,提供清晰的解释、关键公式和中英双语学习笔记,帮助你掌握概率、数据处理、分布、相关性和回归分析。

1. Mathematical Models in Probability and Statistics | 概率与统计中的数学模型

Statistics builds mathematical models to simplify and analyse real-world problems. A model describes a situation using assumptions, parameters, and random variables. For example, flipping a fair coin can be modelled by a binomial distribution B(1, 0.5) with the assumption of independence and constant probability. The modelling cycle involves: recognising a real-world problem, formulating a statistical model, collecting and analysing data, comparing observed outcomes with model predictions, and refining the model if necessary.

统计学通过建立数学模型来简化和分析现实问题。模型利用假设、参数和随机变量来描述情况。例如,抛掷一枚公平硬币可以用二项分布 B(1, 0.5) 建模,并假设各次抛掷独立且概率恒定。建模循环包括:识别现实问题、制定统计模型、收集与分析数据、将观测结果与模型预测进行比较,必要时对模型进行改进。

Key advantages of mathematical models are that they are quick, cheap, and allow us to make predictions. However, models are only approximations; if assumptions are violated, predictions become unreliable.

数学模型的主要优势在于快速、成本低且能做出预测。但模型只是近似,一旦假设不成立,预测就会变得不可靠。


2. Measures of Location and Spread | 位置度量与离散度量

Measures of location summarise the central tendency of data: mean (x̄ = Σx/n), median (middle value), and mode (most frequent). For grouped data, use midpoints to estimate the mean. The median for grouped data is interpolated using cumulative frequency. The mean is affected by outliers, whereas the median is robust.

位置度量总结数据的集中趋势:平均值(x̄ = Σx/n)、中位数(中间值)和众数(出现频率最高的值)。对于分组数据,使用组中点来估计平均值。分组数据的中位数通过累积频率进行插值求得。平均值受异常值影响,而中位数则具有稳健性。

Measures of spread describe variability: range, interquartile range (IQR = Q₃ – Q₁), variance (s² = Σ(x – x̄)² / n), and standard deviation (s = √variance). For grouped data, variance = Σf(x – x̄)² / Σf. The IQR is not affected by extreme values, making it ideal for skewed distributions.

离散度量描述变异性:极差、四分位距(IQR = Q₃ – Q₁)、方差(s² = Σ(x – x̄)² / n)和标准差(s = √方差)。分组数据的方差 = Σf(x – x̄)² / Σf。四分位距不受极端值影响,非常适合偏斜分布。


3. Representations of Data | 数据表示

Data can be displayed using box plots, histograms, cumulative frequency diagrams, and stem-and-leaf diagrams. A box plot shows minimum, Q₁, median, Q₃, maximum, and outliers (values > Q₃ + 1.5×IQR or < Q₁ – 1.5×IQR). Histograms use area proportional to frequency; for unequal class widths, frequency density = frequency / class width.

数据可通过箱形图、直方图、累积频率图和茎叶图来展示。箱形图显示最小值、Q₁、中位数、Q₃、最大值以及异常值(大于 Q₃ + 1.5×IQR 或小于 Q₁ – 1.5×IQR 的值)。直方图中面积与频率成正比;当组距不等时,频率密度 = 频率 / 组距。

Cumulative frequency curves (ogives) allow estimation of medians, quartiles, and percentiles. Stem-and-leaf diagrams preserve original data values and show distribution shape. Skewness can be identified visually or using measures: if mean > median > mode, the distribution is positively skewed; if mean < median < mode, it is negatively skewed.

累积频率曲线(拱形图)可用来估计中位数、四分位数和百分位数。茎叶图能保留原始数据值并展示分布形状。偏度可通过观察或度量来判断:如果平均值 > 中位数 > 众数,分布为正偏;如果平均值 < 中位数 < 众数,则为负偏。


4. Probability | 概率

Probability measures the chance of an event occurring, ranging from 0 (impossible) to 1 (certain). For a finite sample space with equally likely outcomes, P(A) = n(A) / n(S). The complement rule states P(A’) = 1 – P(A). The addition rule: P(A ∪ B) = P(A) + P(B) – P(A ∩ B). If A and B are mutually exclusive, P(A ∩ B) = 0.

概率衡量事件发生的可能性,取值范围从 0(不可能)到 1(必然)。对于等可能结果的有限样本空间,P(A) = n(A) / n(S)。互补规则为 P(A’) = 1 – P(A)。加法规则:P(A ∪ B) = P(A) + P(B) – P(A ∩ B)。若 A 与 B 互斥,则 P(A ∩ B) = 0。

Conditional probability: P(A|B) = P(A ∩ B) / P(B). Events A and B are independent if P(A ∩ B) = P(A)P(B), or equivalently P(A|B) = P(A). Tree diagrams are useful for sequential events; multiply along branches and add probabilities for combined outcomes. Be comfortable with Venn diagrams and two-way tables for solving problems.

条件概率:P(A|B) = P(A ∩ B) / P(B)。如果 P(A ∩ B) = P(A)P(B) 或等价地 P(A|B) = P(A),则事件 A 和 B 独立。树状图对序贯事件很有用;沿分支相乘,并将各路径概率相加得到组合结果。要熟练使用文氏图和双向表来解题。


5. Discrete Random Variables | 离散随机变量

A discrete random variable X takes a countable number of values with associated probabilities. The probability distribution lists each value x and its probability P(X = x), where ΣP(X = x) = 1. The expected value E(X) = Σ x·P(X = x) represents the long-run average. E(aX + b) = aE(X) + b.

离散随机变量 X 取可数个值,并带有相应概率。概率分布列出每个取值 x 及其概率 P(X = x),且满足 ΣP(X = x) = 1。期望值 E(X) = Σ x·P(X = x) 表示长期平均值。E(aX + b) = aE(X) + b。

Variance measures spread: Var(X) = E(X²) – [E(X)]², where E(X²) = Σ x²·P(X = x). For a linear transformation, Var(aX + b) = a² Var(X). The standard deviation is √Var(X). The cumulative distribution function F(x₀) = P(X ≤ x₀) is also tested.

方差衡量离散程度:Var(X) = E(X²) – [E(X)]²,其中 E(X²) = Σ x²·P(X = x)。对于线性变换,Var(aX + b) = a² Var(X)。标准差为 √Var(X)。累积分布函数 F(x₀) = P(X ≤ x₀) 也是考核内容。


6. The Binomial Distribution | 二项分布

If a fixed number of independent trials n have only two outcomes (success/failure) with constant probability of success p, then X ~ B(n, p). The probability of exactly r successes is given by P(X = r) = (nCr) pʳ (1 – p)ⁿ⁻ʳ, where nCr = n! / [r!(n – r)!].

若进行固定次数 n 次独立试验,每次只有两种结果(成功/失败),且成功概率 p 恒定,则 X ~ B(n, p)。恰好 r 次成功的概率为 P(X = r) = (nCr) pʳ (1 – p)ⁿ⁻ʳ,其中 nCr = n! / [r!(n – r)!]。

Mean of binomial: E(X) = np. Variance: Var(X) = np(1 – p). The distribution is symmetrical if p = 0.5, positively skewed if p < 0.5, and negatively skewed if p > 0.5. Tables or calculators can be used to find cumulative probabilities P(X ≤ k).

二项分布的均值:E(X) = np。方差:Var(X) = np(1 – p)。当 p = 0.5 时分布对称,p < 0.5 时正偏,p > 0.5 时负偏。可使用表格或计算器求累积概率 P(X ≤ k)。


7. The Normal Distribution | 正态分布

The normal distribution is a continuous symmetric bell-shaped curve defined by mean μ and standard deviation σ. The total area under the curve equals 1. The standard normal Z ~ N(0, 1) is obtained by standardising: Z = (X – μ) / σ. This transformation allows use of standard normal tables to find probabilities.

正态分布是由均值 μ 和标准差 σ 定义的连续对称钟形曲线。曲线下总面积为 1。通过标准化公式 Z = (X – μ) / σ 得到标准正态分布 Z ~ N(0, 1)。利用这一变换可以使用标准正态分布表求概率。

For any normal variable X ~ N(μ, σ²), P(X < a) = P(Z < (a – μ)/σ). To find an unknown mean or standard deviation, set up equations using given probabilities and inverse table readings. Remember that P(Z > z) = 1 – P(Z < z), and due to symmetry P(Z < –z) = P(Z > z).

对于任意正态变量 X ~ N(μ, σ²),P(X < a) = P(Z < (a – μ)/σ)。若要求未知的均值或标准差,可利用给定概率和逆查表建立方程。记住 P(Z > z) = 1 – P(Z < z),且根据对称性 P(Z < –z) = P(Z > z)。


8. Correlation | 相关性

Correlation measures the strength and direction of a linear relationship between two variables. The product moment correlation coefficient (PMCC) r is calculated using r = Sₓᵧ / √(Sₓₓ Sᵧᵧ), where Sₓₓ = Σx² – (Σx)²/n, Sᵧᵧ = Σy² – (Σy)²/n, Sₓᵧ = Σxy – (Σx)(Σy)/n. r always lies between –1 and 1; values close to 1 or –1 indicate strong linear correlation.

相关性衡量两个变量之间线性关系的强度和方向。乘积矩相关系数(PMCC) r 的计算公式为 r = Sₓᵧ / √(Sₓₓ Sᵧᵧ),其中 Sₓₓ = Σx² – (Σx)²/n, Sᵧᵧ = Σy² – (Σy)²/n, Sₓᵧ = Σxy – (Σx)(Σy)/n。r 的取值范围始终在 –1 到 1 之间;接近 1 或 –1 表明强线性相关。

Be aware that correlation does not imply causation. A scatter diagram should always be plotted to check for non-linear patterns or outliers that might distort r. For ranking data, Spearman’s rank correlation coefficient ρ = 1 – [6Σd² / n(n² – 1)] is used, where d is the difference in ranks.

注意相关性不能推导出因果关系。应始终绘制散点图以检查是否存在可能扭曲 r 的非线性模式或异常值。对于排序数据,使用斯皮尔曼等级相关系数 ρ = 1 – [6Σd² / n(n² – 1)],其中 d 为等级差。


9. Regression | 回归分析

Linear regression finds the line of best fit y = a + bx for predicting y from x. The least squares estimates are b = Sₓᵧ / Sₓₓ, and a = ȳ – b x̄. The regression line always passes through the mean point (x̄, ȳ). This line minimises the sum of squared vertical distances from the points to the line.

线性回归寻找用于根据 x 预测 y 的最佳拟合直线 y = a + bx。最小二乘估计值为 b = Sₓᵧ / Sₓₓ,a = ȳ – b x̄。回归直线总是通过均值点 (x̄, ȳ)。该直线最小化了各点到直线的垂直距离平方和。

Interpolation (predicting within the data range) is reliable provided the model fits well; extrapolation (predicting outside the data range) can be unreliable. Coding can simplify calculations: if u = (x – cₓ)/dₓ and v = (y – cᵧ)/dᵧ, then bₓᵧ = (dᵧ/dₓ) bᵤᵥ. Be familiar with coding effects on Sₓₓ, Sᵧᵧ, Sₓᵧ.

内插(在数据范围内预测)在模型拟合良好的前提下是可靠的;外推(在数据范围外预测)可能不可靠。编码可以简化计算:若 u = (x – cₓ)/dₓ 且 v = (y – cᵧ)/dᵧ,则 bₓᵧ = (dᵧ/dₓ) bᵤᵥ。需熟悉编码对 Sₓₓ、Sᵧᵧ、Sₓᵧ 的影响。


Published by TutorHao | Mathematics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading