📚 CCEA Year 12 Statistics Formula & Theorem Quick Reference | CCEA 12年级统计公式定理速查手册
This quick reference handbook brings together the essential formulae and theorems needed for the CCEA Year 12 Statistics course. It covers descriptive statistics, probability, discrete and continuous distributions, sampling, hypothesis testing, and correlation & regression. Use it as a portable revision companion to cement your understanding and practise applying these results under timed conditions.
这本速查手册汇集了CCEA 12年级统计课程所需的核心公式和定理。内容涵盖描述统计、概率、离散和连续分布、抽样、假设检验以及相关与回归。你可以把它当作便携式复习伙伴,用来巩固理解,并在限时条件下练习运用这些结果。
1. Measures of Central Tendency | 集中趋势的度量
The three principal measures of central tendency are the mean, median and mode. For ungrouped data, the sample mean is the sum of all observations divided by the number of observations.
三种主要的集中趋势度量是平均数、中位数和众数。对于未分组数据,样本平均数等于所有观测值之和除以观测值的个数。
Sample mean: x̄ = Σx / n
样本平均数: x̄ = Σx / n
When data are grouped into frequency tables, we use the midpoint of each class as an approximate value for the observations in that class. The mean is then given by Σfx / Σf.
当数据以频数表分组时,我们用每组的组中值作为该组观测值的近似。此时平均数由 Σfx / Σf 给出。
Grouped mean: x̄ = Σfx / Σf
分组平均数: x̄ = Σfx / Σf
The median is the middle value when data are ordered. For n observations, the position of the median is (n + 1)/2. If n is even it is the average of the two middle values. The mode is simply the value that occurs most frequently.
中位数是数据排序后的中间值。对于n个观测值,中位数的位置是(n + 1)/2。如果n为偶数,则为中间两个值的平均。众数就是出现频次最高的数值。
2. Measures of Dispersion | 离散程度的度量
Dispersion tells us how spread out the data are. The range, interquartile range (IQR), variance and standard deviation are the core measures. The range is simply the maximum minus the minimum. The IQR is Q₃ – Q₁, where Q₁ is the lower quartile and Q₃ is the upper quartile.
离散程度告诉我们数据的分散情况。极差、四分位距(IQR)、方差和标准差是核心度量。极差就是最大值减最小值。四分位距是Q₃ – Q₁,其中Q₁是下四分位数,Q₃是上四分位数。
For a population of size N, the variance is the average of the squared deviations from the mean. For a sample of size n, we divide by n – 1 to obtain an unbiased estimate.
对于大小为N的总体,方差是离均差平方的平均值。对于大小为n的样本,我们除以n – 1以获得无偏估计。
Population variance: σ² = Σ(x – μ)² / N
Sample variance: s² = Σ(x – x̄)² / (n – 1)
总体方差: σ² = Σ(x – μ)² / N
样本方差: s² = Σ(x – x̄)² / (n – 1)
An equivalent, often more convenient computational formula for sample variance is s² = [Σx² – (Σx)²/n] / (n – 1). The standard deviation is simply the positive square root of the variance: s = √s².
一个等价的、通常更便于计算的样本方差公式是 s² = [Σx² – (Σx)²/n] / (n – 1)。标准差就是方差的正平方根:s = √s²。
3. Probability Rules | 概率法则
Probability is a measure of the likelihood of an event, obeying 0 ≤ P(A) ≤ 1 for any event A. The complement rule states that P(not A) = 1 – P(A). For any two events A and B, the addition law links the probability of their union to the probability of their intersection.
概率是对事件可能性的度量,对于任何事件A,总有 0 ≤ P(A) ≤ 1。补集法则指出 P(非A) = 1 – P(A)。对于任意两个事件A和B,加法法则把它们的并集概率与交集概率联系起来。
P(A ∪ B) = P(A) + P(B) – P(A ∩ B)
P(A ∪ B) = P(A) + P(B) – P(A ∩ B)
If A and B are mutually exclusive, then P(A ∩ B) = 0, so P(A ∪ B) = P(A) + P(B).
如果A和B互斥,那么 P(A ∩ B) = 0,因此 P(A ∪ B) = P(A) + P(B)。
Conditional probability is defined as P(A|B) = P(A ∩ B) / P(B), provided P(B) > 0. Rearranged, this gives the multiplication rule: P(A ∩ B) = P(B) × P(A|B). If A and B are independent, then P(A|B) = P(A), so P(A ∩ B) = P(A) × P(B).
条件概率定义为 P(A|B) = P(A ∩ B) / P(B),其中 P(B) > 0。整理后得到乘法法则:P(A ∩ B) = P(B) × P(A|B)。如果A和B独立,则 P(A|B) = P(A),因此 P(A ∩ B) = P(A) × P(B)。
4. Discrete Random Variables | 离散随机变量
A discrete random variable X takes a countable set of values. Its probability distribution lists the values xᵢ and their probabilities pᵢ = P(X = xᵢ). The sum of all probabilities must equal 1. The expected value E(X) is the long-run average value of X.
离散随机变量X取可数个值。它的概率分布列出了取值xᵢ及其对应的概率 pᵢ = P(X = xᵢ)。所有概率之和必须等于1。期望值 E(X) 是X的长期平均值。
E(X) = Σ xᵢ pᵢ
E(X) = Σ xᵢ pᵢ
The variance of X, Var(X), measures the spread of the distribution and is computed as the expected squared deviation from the mean.
X的方差 Var(X) 衡量分布的分散程度,计算公式为离均差平方的期望。
Var(X) = Σ (xᵢ – μ)² pᵢ = E(X²) – [E(X)]²
Var(X) = Σ (xᵢ – μ)² pᵢ = E(X²) – [E(X)]²
For a linear transformation aX + b, the expectation and variance change in a predictable way: E(aX + b) = aE(X) + b, and Var(aX + b) = a² Var(X).
对于线性变换 aX + b,期望和方差按可预测的方式变化:E(aX + b) = aE(X) + b,Var(aX + b) = a² Var(X)。
5. Binomial Distribution | 二项分布
The binomial distribution models the number of successes in a fixed number n of independent trials, each with the same probability of success p. If X ~ B(n, p), then the probability of exactly x successes is given by the binomial probability function.
二项分布用于描述固定次数n的独立试验中成功的次数,每次试验成功的概率p相同。如果 X ~ B(n, p),那么恰好有x次成功的概率由二项概率函数给出。
P(X = x) = C(n, x) pˣ (1 – p)ⁿ⁻ˣ, for x = 0,1,2,…,n
P(X = x) = C(n, x) pˣ (1 – p)ⁿ⁻ˣ,x = 0,1,2,…,n
The binomial coefficient C(n, x) (also written as nCx or nCr) is the number of ways to choose x items from n, and equals n! / [x!(n – x)!]. The mean and variance of a binomial random variable are straightforward.
二项系数 C(n, x)(也写作 nCx 或 nCr)是从n个物品中选x个的方法数,等于 n! / [x!(n – x)!]。二项随机变量的均值和方差很简单。
E(X) = np
E(X) = np
Var(X) = np(1 – p)
Var(X) = np(1 – p)
The conditions for a binomial distribution are: a fixed number of trials, two possible outcomes per trial (success/failure), constant probability p, and independent trials.
适用二项分布的条件是:试验次数固定,每次试验只有两种可能结果(成功/失败),概率p恒定,且各次试验独立。
6. Poisson Distribution | 泊松分布
The Poisson distribution is used to model the number of events occurring in a fixed interval of time or space, when events occur independently at a constant average rate λ. If X ~ Po(λ), the probability of exactly x events is:
泊松分布用于模拟在固定时间或空间间隔内发生的事件数量,当事件之间相互独立,并以恒定的平均速率λ发生时。如果 X ~ Po(λ),恰好发生x个事件的概率为:
P(X = x) = (e⁻λ λˣ) / x!, for x = 0,1,2,…
P(X = x) = (e⁻λ λˣ) / x!,x = 0,1,2,…
A distinctive property of the Poisson distribution is that its mean and variance are equal.
泊松分布的一个显著特性是其均值和方差相等。
E(X) = λ
E(X) = λ
Var(X) = λ
Var(X) = λ
The Poisson distribution can also be used as an approximation to the binomial distribution B(n, p) when n is large and p is small, with λ = np. As a rule of thumb this works well when n ≥ 50 and np ≤ 5.
当n很大而p很小时,泊松分布可以用作二项分布 B(n, p) 的近似,此时令 λ = np。经验表明当 n ≥ 50 且 np ≤ 5 时这种近似效果良好。
7. Normal Distribution | 正态分布
The normal distribution is the most important continuous probability distribution, characterised by its bell-shaped curve. It is fully defined by two parameters: the population mean μ and standard deviation σ. We write X ~ N(μ, σ²).
正态分布是最重要的连续概率分布,其特征是钟形曲线。它由两个参数完全确定:总体均值μ和标准差σ。我们记作 X ~ N(μ, σ²)。
The probability density function is not directly used for calculation in Year 12; instead we standardise to the standard normal Z ~ N(0, 1²). The standardisation formula is:
在12年级,我们不直接使用概率密度函数进行计算;而是将其标准化为标准正态 Z ~ N(0, 1²)。标准化公式为:
Z = (X – μ) / σ
Z = (X – μ) / σ
Using standard normal tables we can find probabilities such as P(Z < z). The total area under the normal curve is 1, and it is symmetric about the mean, so P(Z < -z) = P(Z > z) = 1 – P(Z < z).
利用标准正态表我们可以求得 P(Z < z) 等概率。正态曲线下的总面积为1,且关于均值对称,因此 P(Z < -z) = P(Z > z) = 1 – P(Z < z)。
When using a continuous normal distribution to approximate a binomial distribution (continuity correction), we adjust by 0.5. For X ~ B(n, p) with large n, X ≈ N(np, np(1-p)), and we apply corrections like P(X ≤ k) ≈ P(Y < k + 0.5).
当用连续的正态分布近似二项分布时(连续性校正),需要加减0.5进行调整。对于 X ~ B(n, p) 且n足够大,有 X ≈ N(np, np(1-p)),并进行类似的校正:P(X ≤ k) ≈ P(Y < k + 0.5)。
8. Sampling & Central Limit Theorem | 抽样与中心极限定理
In statistics we often work with samples rather than populations. The sampling distribution of the sample mean is fundamental. If a population has mean μ and variance σ², and we take random samples of size n, then the sample mean x̄ has the following properties:
在统计学中,我们经常处理样本而不是总体。样本平均数的抽样分布是基础。如果总体均值为μ,方差为σ²,我们从中抽取大小为n的随机样本,那么样本平均数x̄具有以下性质:
E(x̄) = μ
E(x̄) = μ
Var(x̄) = σ² / n
Var(x̄) = σ² / n
The standard deviation of x̄, called the standard error, is σ/√n. The Central Limit Theorem (CLT) states that regardless of the shape of the population distribution, as n increases the distribution of x̄ approaches a normal distribution N(μ, σ²/n). For most practical purposes, n ≥ 30 is sufficient for the CLT to apply.
样本平均数的标准差称为标准误,等于 σ/√n。中心极限定理 (CLT) 指出,无论总体分布形状如何,当n增大时,x̄的分布趋近于正态分布 N(μ, σ²/n)。在实际应用中,n ≥ 30 通常足以使CLT适用。
9. Confidence Intervals | 置信区间
A confidence interval provides a range of plausible values for an unknown population parameter. For a population mean μ when the population standard deviation σ is known, a 95% confidence interval is constructed using the normal distribution.
置信区间给出了未知总体参数的一个合理取值范围。当总体标准差σ已知时,总体均值μ的95%置信区间用正态分布构造。
x̄ ± z* × (σ / √n), where z* = 1.96 for 95% confidence
x̄ ± z* × (σ / √n),其中对于95%置信度 z* = 1.96
When σ is unknown and the sample size is small, we use the t‑distribution with n – 1 degrees of freedom. The confidence interval then becomes x̄ ± t* × (s / √n), where s is the sample standard deviation.
当σ未知且样本量较小时,我们使用具有 n – 1 自由度的 t 分布。此时置信区间变为 x̄ ± t* × (s / √n),其中s为样本标准差。
10. Hypothesis Testing | 假设检验
Hypothesis testing is a formal method for making decisions about population parameters using sample data. We set up a null hypothesis H₀ and an alternative hypothesis H₁. The test statistic is calculated from the sample and compared with a critical value, or a p‑value is found. For a test on a population mean with known σ, the test statistic is:
假设检验是一种利用样本数据对总体参数做出决策的规范方法。我们建立原假设 H₀ 和备择假设 H₁。由样本计算检验统计量,并与临界值比较,或者求出 p 值。对于σ已知的总体均值检验,检验统计量为:
z = (x̄ – μ₀) / (σ / √n)
z = (x̄ – μ₀) / (σ / √n)
For a binomial proportion test, we compare the observed number of successes with the binomial distribution under H₀. The critical region is determined by the significance level α (usually 0.05). If the test statistic falls in the critical region we reject H₀ in favour of H₁. A p‑value less than α also leads to rejection.
对于二项比例检验,我们将观测到的成功次数与 H₀ 下的二项分布进行比较。临界域由显著性水平 α(通常取0.05)决定。如果检验统计量落入临界域,我们拒绝 H₀ 而支持 H₁。p 值小于 α 时也导致拒绝 H₀。
11. Correlation & Regression | 相关与回归
Correlation measures the strength and direction of a linear relationship between two variables. Pearson’s product-moment correlation coefficient r is given by a formula that standardises the covariance.
相关衡量两个变量之间线性关系的强度和方向。皮尔逊积矩相关系数 r 的公式将协方差标准化。
r = [n Σxy – (Σx)(Σy)] / √{[n Σx² – (Σx)²][n Σy² – (Σy)²]}
r = [n Σxy – (Σx)(Σy)] / √{[n Σx² – (Σx)²][n Σy² – (Σy)²]}
The value of r lies between -1 and 1. An r close to 1 indicates strong positive correlation, close to -1 indicates strong negative correlation, and near 0 suggests little or no linear correlation. The coefficient of determination r² tells us the proportion of variation in y explained by x.
r 的取值范围在 -1 和 1 之间。r 接近 1 表明强正相关,接近 -1 表明强负相关,接近 0 表明几乎没有线性相关。判定系数 r² 反映了 y 的变异中有多大比例可由 x 解释。
Linear regression models the relationship between an explanatory variable x and a response variable y with a straight line ŷ = a + bx. The slope b and intercept a are found using least squares.
线性回归用直线 ŷ = a + bx 来建模解释变量 x 与响应变量 y 之间的关系。斜率 b 和截距 a 通过最小二乘法求出。
b = [n Σxy – (Σx)(Σy)] / [n Σx² – (Σx)²]
b = [n Σxy – (Σx)(Σy)] / [n Σx² – (Σx)²]
a = ȳ – b x̄
a = ȳ – b x̄
12. Residuals & Outliers | 残差与异常值
A residual is the difference between an observed y‑value and the value predicted by the regression line: e = y – ŷ. Residuals are used to assess the goodness of fit. Unusually large residuals may indicate outliers or influential points that distort the regression model.
残差是观测到的 y 值与回归线预测值之间的差值:e = y – ŷ。残差用于评估拟合优度。异常大的残差可能表明存在离群值或强影响点,它们会扭曲回归模型。
Spearman’s rank correlation coefficient rs is a non‑parametric alternative that measures the strength of monotonic association between two ranked variables. It is calculated by replacing each variable’s values with their ranks and applying Pearson’s formula to the ranked data. It is particularly useful when the relationship is not strictly linear or when data are ordinal.
斯皮尔曼秩相关系数 rs 是一种非参数替代方法,衡量两个排序变量之间单调关联的强度。计算时将每个变量的数值替换为其秩,再对排名后的数据应用皮尔逊公式。当关系并非严格线性或数据为等级数据时尤其有用。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导