📚 Year 11 CCEA Statistics: Quick Reference Formula & Theorem Handbook | Year 11 CCEA 统计:公式定理速查手册
This handbook distills the essential formulas, theorems and key concepts for the Year 11 CCEA Statistics course into a single quick-reference document. Use it to check definitions, refresh your memory before an assessment, and ensure you can apply each relationship correctly in context.
本手册将 Year 11 CCEA 统计课程的核心公式、定理和关键概念提炼为一份速查文件。你可以用它核对定义、在评估前唤醒记忆,并确保你能在具体情境中正确运用每一个关系式。
1. Measures of Central Tendency | 集中趋势的度量
The arithmetic mean (often just called the average) is found by summing all data values and dividing by the number of values. For a set of n observations x₁, x₂, …, xₙ, the sample mean is denoted by x̄.
算术平均值(通常简称为平均数)通过将所有数据值相加再除以数据个数求得。对于 n 个观测值 x₁, x₂, …, xₙ,样本均值记为 x̄。
Mean = x̄ = Σx / n
The median is the middle value when the data are written in ascending order. If n is even, the median is the average of the two central numbers.
中位数是将数据按升序排列后位于正中间的值。若 n 为偶数,中位数则为中间两个数的平均值。
The mode (or modal class for grouped data) is the value or category that occurs most frequently. A data set may have no mode, one mode, or more than one mode.
众数(对于分组数据则是众数组)是出现频次最高的数值或类别。一组数据可能没有众数、有一个众数或多个众数。
2. Range & Interquartile Range (IQR) | 极差与四分位距
The range is the simplest measure of spread, calculated as the difference between the largest and smallest values in the data set.
极差是最简单的离散程度指标,等于数据集中最大值与最小值之差。
Range = Maximum value – Minimum value
The interquartile range (IQR) measures the spread of the middle 50% of the data and is unaffected by extreme values. It is the difference between the upper quartile (Q₃) and the lower quartile (Q₁).
四分位距 (IQR) 衡量中间 50% 数据的分散程度,不受极端值影响。它等于上四分位数 (Q₃) 与下四分位数 (Q₁) 的差值。
IQR = Q₃ – Q₁
The semi-interquartile range is half of the IQR, but IQR is the preferred measure in most CCEA contexts.
半四分位距是 IQR 的一半,但在 CCEA 的多数情境中 IQR 是首选度量。
3. Quartiles & Box Plots | 四分位数与箱线图
To find quartiles from a raw data list, first arrange the data in ascending order. The lower quartile Q₁ is the value one quarter of the way through the data; the median Q₂ is the halfway value; the upper quartile Q₃ is three-quarters of the way through.
要从原始数据列表中找出四分位数,首先将数据升序排列。下四分位数 Q₁ 是位于四分之一处的值;中位数 Q₂ 是位于一半处的值;上四分位数 Q₃ 是位于四分之三处的值。
When using cumulative frequency graphs, Q₁, Q₂ and Q₃ are read off at 25%, 50% and 75% of the total cumulative frequency.
当使用累积频数图时,Q₁、Q₂ 和 Q₃ 分别在总累积频数的 25%、50% 和 75% 处读取。
A box plot (or box-and-whisker diagram) displays the minimum, Q₁, median, Q₃ and maximum. It provides a clear visual summary of spread and skewness.
箱线图(或盒须图)展示最小值、Q₁、中位数、Q₃ 和最大值。它能清晰直观地概括数据的分散程度和偏态。
4. Standard Deviation | 标准差
Standard deviation quantifies the typical distance of data values from the mean. A larger standard deviation indicates greater spread. You must distinguish between a sample and a population.
标准差量化了数据值偏离均值的典型距离。标准差越大,表示数据越分散。必须区分样本与总体。
Sample standard deviation (used when you have a sample of size n):
样本标准差(当拥有容量为 n 的样本时):
s = √[ Σ (x – x̄)² / (n – 1) ]
Population standard deviation (used when you have data for every member of a population of size N):
总体标准差(当拥有总体中每个个体的数据且总体大小为 N 时):
σ = √[ Σ (x – μ)² / N ]
An alternative formula for the sample standard deviation avoids calculating every deviation from the mean explicitly: s = √{ [ Σx² – (Σx)²/n ] / (n – 1) }. This is convenient for large data sets.
样本标准差的另一种公式可避免显式计算每个值与均值的偏差:s = √{ [ Σx² – (Σx)²/n ] / (n – 1) },这在处理大数据集时很方便。
5. Probability Fundamentals | 概率基础
The probability of an event A, denoted P(A), is a measure of how likely the event is to occur, ranging from 0 (impossible) to 1 (certain). It can be expressed as a fraction, decimal or percentage.
事件 A 的概率记作 P(A),衡量该事件发生的可能性大小,介于 0(不可能)到 1(必然)之间。可用分数、小数或百分数表示。
P(A) = Number of favourable outcomes / Total number of possible outcomes
The complement rule states that the probability of an event not occurring is one minus the probability that it does occur.
互补规则指出,事件不发生的概率等于 1 减去它发生的概率。
P(not A) = 1 – P(A)
For mutually exclusive events A and B (they cannot occur simultaneously), the probability that A or B occurs is the sum of their individual probabilities.
对于互斥事件 A 与 B(两者不能同时发生),A 或 B 发生的概率等于它们各自概率之和。
P(A or B) = P(A) + P(B)
For non‑mutually exclusive events, we must subtract the overlap to avoid double‑counting.
对于非互斥事件,必须减去重叠部分以避免重复计算。
P(A or B) = P(A) + P(B) – P(A and B)
Expected frequency of an event in a large number of trials is obtained by multiplying the probability by the total number of trials.
大量试验中事件的期望频数通过将概率乘以总试验次数求得。
Expected frequency = n × P(A)
6. Conditional Probability & Tree Diagrams | 条件概率与树状图
Conditional probability deals with the likelihood of an event occurring given that another event has already occurred. The notation P(A|B) means “the probability of A given B”.
条件概率处理在另一事件已经发生的条件下某事件发生的可能性。记号 P(A|B) 表示“在 B 发生的条件下 A 发生的概率”。
P(A|B) = P(A ∩ B) / P(B), provided P(B) > 0
Tree diagrams help model multi‑stage experiments. Multiply probabilities along the branches to find the probability of a particular path; add path probabilities that correspond to the same final outcome.
树状图有助于为多阶段试验建模。沿分支将概率相乘可得到特定路径的概率;将对应同一最终结果的不同路径概率相加。
When drawing a tree diagram for events without replacement, the probabilities on the second set of branches depend on what happened first — these are conditional probabilities.
为不放回事件绘制树状图时,第二层分支上的概率取决于第一次发生的结果——这些就是条件概率。
To find the probability of an outcome described in words, identify all the paths that satisfy it, compute their individual probabilities and add them together.
要计算用文字描述的某个结果的概率,先找出所有满足该结果的路径,计算每条路径的概率,再将其相加。
7. Binomial Distribution | 二项分布
A binomial setting arises when there is a fixed number of independent trials, n, each with two possible outcomes (success and failure), and the probability of success, p, stays constant.
当试验满足以下条件时,就属于二项背景:固定次数 n 的独立试验,每次仅有两种可能结果(成功与失败),且成功的概率 p 保持不变。
The probability of obtaining exactly r successes in n trials is given by the binomial probability formula.
在 n 次试验中恰好获得 r 次成功的概率由二项概率公式给出。
P(X = r) = nCr pr (1 – p)n–r
Here nCr = n! / [r! (n – r)!] is the binomial coefficient, which counts how many ways r successes can be chosen from n trials.
此处 nCr = n! / [r! (n – r)!] 是二项式系数,表示从 n 次试验中选出 r 次成功的方式数。
The mean (expected value) and variance of a binomial random variable X ~ B(n, p) are:
二项随机变量 X ~ B(n, p) 的均值(期望值)和方差分别为:
E(X) = μ = np
Var(X) = σ² = np(1 – p)
These formulas allow you to predict the average number of successes and the variability around that average without listing all outcomes.
借助这些公式,你无需列出所有结果便可预测平均成功次数及其变动幅度。
8. Spearman’s Rank Correlation Coefficient | 斯皮尔曼等级相关系数
Spearman’s rank correlation coefficient, rs, measures the strength of a monotonic association between two sets of ranked data. It is appropriate when the relationship is non‑linear but still consistently increasing or decreasing.
斯皮尔曼等级相关系数 rs 衡量两组等级数据之间单调关联的强度。当关系是非线性但仍保持一致的上升或下降时,该系数非常适用。
rs = 1 – (6 Σ d²) / [n (n² – 1)]
Where d is the difference between the ranks of each pair, and n is the number of data pairs. Always rank the data within each variable separately before computing d.
其中 d 是每对数据的等级差,n 是数据对的个数。在计算 d 之前,务必先分别对每个变量进行排秩。
The value of rs lies between –1 and +1; +1 indicates perfect positive rank correlation, –1 indicates perfect negative rank correlation, and 0 suggests no monotonic trend.
rs 的值介于 –1 与 +1 之间;+1 表示完全正等级相关,–1 表示完全负等级相关,0 则表明没有单调趋势。
Tied ranks occur when two data values are identical; in that case, assign the average of the positions they occupy, and use those average ranks for all tied items.
当两个数据值相同时会出现并列等级;此时,给它们分配所占位置的平均值,并对所有并列项使用该平均等级。
9. Normal Distribution Properties | 正态分布的特性
The normal distribution is a continuous, bell‑shaped probability distribution that is symmetric about its mean. In a perfectly normal distribution, mean = median = mode.
正态分布是一种连续、钟形的概率分布,关于其均值对称。在完全正态分布中,均值 = 中位数 = 众数。
The total area under the normal curve equals 1, and probabilities correspond to areas under the curve. The spread is governed by the standard deviation σ.
正态曲线下的总面积为 1,概率对应于曲线下的面积。分散程度由标准差 σ 决定。
The empirical (68–95–99.7) rule states that roughly 68% of data lie within 1σ of the mean, about 95% within 2σ, and approximately 99.7% within 3σ.
经验法则 (68–95–99.7) 指出:约 68% 的数据落在均值 ±1σ 范围内,约 95% 落在 ±2σ 范围内,约 99.7% 落在 ±3σ 范围内。
In Year
Published by TutorHao | Year 11 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply