📚 Key Statistics Revision for IB & CCEA Mathematics | IB 与 CCEA 数学统计考点精讲
Statistics in IB and CCEA mathematics builds a foundational toolkit for data analysis, probability, and inference. This article consolidates the essential topics you need to master, from descriptive statistics to hypothesis testing, with clear explanations and worked notation.
在 IB 和 CCEA 数学中,统计学构成了数据分析、概率与推断的基础工具箱。本文汇整了从描述性统计到假设检验的关键考点,配以清晰的解释与符号演示,帮助你系统复习。
1. Types of Data | 数据类型
Data is classified as qualitative (categorical) or quantitative (numerical). Quantitative data can be discrete (countable values) or continuous (measurements on a scale). Understanding data types determines which statistical methods and graphs are appropriate.
数据可分为定性(类别)数据和定量(数值)数据。定量数据又分为离散型(可数数值)和连续型(尺度测量值)。理解数据类型决定了应选用何种统计方法和图表。
Qualitative data is often summarised using frequency tables and bar charts, while quantitative data is displayed with histograms, box plots and cumulative frequency curves. For continuous data, class boundaries are used in histograms.
定性数据常使用频数表和条形图进行汇总,而定量数据则通过直方图、箱线图和累积频数曲线加以展示。对于连续数据,直方图中需使用组边界。
2. Measures of Central Tendency | 集中趋势的度量
The three main measures are mean, median and mode. The mean, x̄ = Σx/n for a sample, is sensitive to outliers. The median is the middle value when data is ordered, robust to skewness. The mode is the most frequent value.
三种主要的集中量数是平均数、中位数和众数。样本平均数 x̄ = Σx/n 对异常值敏感;中位数是排序后居中值,对偏态有稳健性;众数是出现频率最高的值。
For grouped data, the mean is estimated using Σ(f × midpoint) / Σf, and the median is found by interpolation within the median class using cumulative frequency.
对于分组数据,平均数通过 Σ(f × 组中值) / Σf 估算,中位数则需利用累积频数在中间组内进行插值求得。
3. Measures of Spread | 离散程度的度量
Range, interquartile range (IQR = Q₃ − Q₁), variance and standard deviation quantify variability. Sample variance s² = Σ(x − x̄)²/(n − 1), and standard deviation s = √s². For a population, divide by n.
全距、四分位距 (IQR = Q₃ − Q₁)、方差与标准差用于量化变异性。样本方差 s² = Σ(x − x̄)²/(n − 1),标准差 s = √s²;总体方差则除以 n。
Variance can also be computed using the formula s² = (Σx² − (Σx)²/n) / (n − 1). The standard deviation has the same units as the original data, making it more interpretable.
方差也可用公式 s² = (Σx² − (Σx)²/n) / (n − 1) 计算。标准差与原数据单位相同,更易于解读。
4. Probability Basics | 概率基础
Probability P(A) satisfies 0 ≤ P(A) ≤ 1. For equally likely outcomes, P(A) = number of favourable outcomes / total outcomes. The addition rule: P(A ∪ B) = P(A) + P(B) − P(A ∩ B). For mutually exclusive events, P(A ∩ B) = 0.
概率 P(A) 满足 0 ≤ P(A) ≤ 1。对于等可能结果,P(A) = 有利结果数 / 总结果数。加法法则:P(A ∪ B) = P(A) + P(B) − P(A ∩ B)。互斥事件时 P(A ∩ B) = 0。
Conditional probability P(A|B) = P(A ∩ B) / P(B). Events A and B are independent if P(A ∩ B) = P(A) × P(B), or equivalently P(A|B) = P(A). Tree diagrams help organise compound events.
条件概率 P(A|B) = P(A ∩ B) / P(B)。若 P(A ∩ B) = P(A) × P(B) 或 P(A|B) = P(A),则事件 A 与 B 独立。树状图有助于梳理复合事件。
5. Discrete Random Variables | 离散随机变量
A discrete random variable X takes countable values with a probability mass function P(X = x). The sum of all probabilities equals 1. The expected value E(X) = Σ x·P(X = x), and variance Var(X) = E(X²) − [E(X)]².
离散随机变量 X 取可数值,并有其概率质量函数 P(X = x)。所有概率之和为 1。期望值 E(X) = Σ x·P(X = x),方差 Var(X) = E(X²) − [E(X)]²。
For linear transformations, E(aX + b) = aE(X) + b and Var(aX + b) = a²Var(X). These properties simplify calculations in repeated games and scaling.
对于线性变换,有 E(aX + b) = aE(X) + b 和 Var(aX + b) = a²Var(X)。这些性质可简化重复博弈与缩放情形下的计算。
6. Binomial Distribution | 二项分布
A binomial distribution models the number of successes in n independent trials, each with success probability p. We write X ~ B(n, p). The probability of exactly k successes is P(X = k) = ⁿCₖ pᵏ (1 − p)ⁿ⁻ᵏ, where ⁿCₖ = n! / [k!(n−k)!].
二项分布用于描述 n 次独立试验中成功的次数,每次成功概率为 p。记作 X ~ B(n, p)。恰好有 k 次成功的概率为 P(X = k) = ⁿCₖ pᵏ (1 − p)ⁿ⁻ᵏ,其中 ⁿCₖ = n! / [k!(n−k)!]。
Mean: E(X) = np, variance: Var(X) = np(1 − p). The distribution is symmetric when p = 0.5, and skewed otherwise. Calculations can be done with tables or calculator functions.
期望值 E(X) = np,方差 Var(X) = np(1 − p)。当 p = 0.5 时分布对称,否则偏斜。计算时可使用表格或计算器函数。
7. Normal Distribution | 正态分布
A continuous random variable X follows a normal distribution with mean μ and standard deviation σ, written X ~ N(μ, σ²). The standard normal variable Z = (X − μ)/σ ~ N(0, 1) is used for probability calculations.
连续随机变量 X 服从均值为 μ、标准差为 σ 的正态分布,记作 X ~ N(μ, σ²)。通过标准正态变量 Z = (X − μ)/σ ~ N(0, 1) 进行概率计算。
Probabilities like P(X < a) are found by converting to Z-scores and using the standard normal table. The empirical rule states that roughly 68% of data lies within μ ± σ, 95% within μ ± 2σ, and 99.7% within μ ± 3σ.
求 P(X < a) 时,转换为 Z 值并查标准正态表。经验法则指出,约 68% 的数据落在 μ ± σ 内,95% 落在 μ ± 2σ 内,99.7% 落在 μ ± 3σ 内。
8. Sampling and Estimation | 抽样与估计
The sample mean x̄ is an unbiased estimator of the population mean μ. The distribution of x̄ for large samples is approximately normal with standard error σ/√n (Central Limit Theorem). When population standard deviation is unknown, the sample standard deviation s is used, and the t-distribution applies for small samples.
样本平均数 x̄ 是总体平均数 μ 的无偏估计量。大样本下,x̄ 的分布近似正态,标准误为 σ/√n(中心极限定理)。当总体标准差未知时,用样本标准差 s 替代,且小样本下使用 t 分布。
A confidence interval for μ (when σ is known) is x̄ ± z* × σ/√n, where z* is the critical value (e.g., 1.96 for 95% confidence). For unknown σ, replace σ with s and use t-critical value.
当 σ 已知时,μ 的置信区间为 x̄ ± z* × σ/√n,其中 z* 为临界值(例如 95% 置信度下为 1.96)。若 σ 未知,则以 s 代替 σ,并采用 t 临界值。
9. Hypothesis Testing | 假设检验
A hypothesis test assesses evidence against a null hypothesis H₀ in favour of an alternative H₁. The test statistic (e.g., Z or t) is computed from sample data, and the p-value is the probability of obtaining a result at least as extreme, assuming H₀ is true.
假设检验旨在评估反对原假设 H₀、支持备择假设 H₁ 的证据。根据样本数据计算检验统计量(如 Z 或 t),p 值是假定 H₀ 为真时获得至少如此极端结果的概率。
If p-value < significance level α (commonly 0.05), we reject H₀. Critical region approach: reject H₀ if test statistic falls in the critical region determined by α. One-tailed and two-tailed tests depend on the direction stated in H₁.
若 p 值 < 显著性水平 α(常取 0.05),则拒绝 H₀。临界域法:若检验统计量落入由 α 决定的临界域,则拒绝 H₀。单尾或双尾检验取决于 H₁ 中指定的方向。
For binomial tests, exact probabilities are used. For normal tests, Z-test is applied. When comparing two means, two-sample tests or paired tests are used based on the design.
二项检验使用精确概率;正态检验则应用 Z 检验。比较两个平均数时,需根据设计选用双样本检验或配对检验。
10. Correlation and Regression | 相关与回归
Pearson’s product-moment correlation coefficient r measures linear association between two variables. −1 ≤ r ≤ 1. r = 1 indicates perfect positive linear correlation, r = −1 perfect negative, and r = 0 no linear correlation.
皮尔逊积矩相关系数 r 衡量两变量间的线性关联。−1 ≤ r ≤ 1。r = 1 表示完全正线性相关,r = −1 表示完全负线性相关,r = 0 表示无线性相关。
The least-squares regression line of y on x is y = a + bx, where b = Sxy / Sxx and a = ȳ − bx̄. Here Sxy = Σ(x − x̄)(y − ȳ) and Sxx = Σ(x − x̄)². This line minimises the sum of squared residuals.
y 对 x 的最小二乘回归直线为 y = a + bx,其中 b = Sxy / Sxx,a = ȳ − bx̄。此处 Sxy = Σ(x − x̄)(y − ȳ),Sxx = Σ(x − x̄)²。该直线使残差平方和最小。
Interpolation (predicting within the data range) is reliable, but extrapolation (beyond the range) can be misleading. The coefficient of determination r² indicates the proportion of variance in y explained by x.
在数据范围内进行内插预测较为可靠,而外推(超出范围)则可能产生误导。决定系数 r² 指明 y 的变异中可由 x 解释的比例。
Published by TutorHao | Mathematics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply