📚 Year 12 CAIE Statistics: Glossary and Memorisation Guide | Year 12 CAIE 统计:词汇术语速记指南
In CAIE Year 12 Statistics, mastering the terminology is half the battle. Whether you are interpreting a box‑and‑whisker plot, calculating a binomial probability or working with the normal distribution, precise language leads to accurate solutions. This bilingual guide pairs each key term with its Chinese equivalent and a short explanation, organised by topic for quick revision. Use it to strengthen your statistical vocabulary and avoid common misunderstandings in the exam.
在 CAIE 12 年级统计中,掌握术语是成功的一半。无论你在解读箱线图、计算二项概率还是处理正态分布,精准的语言都能带来准确的解答。这份双语指南将每个关键术语配以中文对应词和简短解释,按主题编排,便于快速复习。用它来强化统计词汇,避免考试中常见的理解偏差。
1. Descriptive Statistics: Data Types & Charts | 描述统计:数据类型与图表
Raw data – information collected before any processing, listed as individual numbers or categories.
原始数据 – 收集后未经任何处理的信息,以单个数字或类别列出。
Qualitative data – non‑numerical observations (e.g. colour, gender). Often summarised in bar charts or pie charts.
定性数据 – 非数值型的观测值(如颜色、性别),通常用条形图或饼图汇总。
Quantitative data – numerical observations that can be discrete (countable, e.g. number of cars) or continuous (measurable, e.g. height).
定量数据 – 数值型的观测值,可以是离散型(可数,如汽车数量)或连续型(可测量,如身高)。
Frequency – the number of times a data value appears in a dataset.
频数 – 某个数据值在数据集中出现的次数。
Cumulative frequency – the running total of frequencies up to a given value. Used to construct cumulative frequency curves (ogives).
累积频数 – 截至某一给定值的频数累加总数,用于绘制累积频数曲线(肩形图)。
Class interval – a group of data values defined by a lower and an upper boundary. For continuous data, class boundaries remove gaps between intervals.
组距 – 由下界和上界定义的一组数据值。对于连续数据,组界消除了组间的空隙。
Histogram – a chart where the area of each bar is proportional to the frequency. For unequal widths, frequency density = frequency ÷ class width.
直方图 – 一种图表,其中每个长方形的面积与频数成比例。当组距不等时,用频数密度 = 频数 ÷ 组距。
Box‑and‑whisker plot – displays the minimum, lower quartile (Q₁), median (Q₂), upper quartile (Q₃) and maximum. Outliers are shown as separate points.
箱线图 – 展示最小值、下四分位数(Q₁)、中位数(Q₂)、上四分位数(Q₃)和最大值。异常值单独标出。
2. Measures of Central Tendency | 集中趋势度量
Mean – the arithmetic average, x̄ = Σx / n for raw data, or Σfx / Σf for grouped data. Affected by extreme values.
平均数 – 算术平均值,原始数据用 x̄ = Σx / n,分组数据用 Σfx / Σf。受极端值影响。
Median – the middle value when data are ordered. For n observations, position = (n+1)/2. Not swayed by outliers.
中位数 – 数据排序后位于中间的值。n 个观测值时,位置为 (n+1)/2。不受异常值影响。
Mode – the most frequently occurring value. A dataset can have one mode, more than one mode (bimodal/multimodal) or no mode at all.
众数 – 出现次数最多的值。数据集可以有单众数、多个众数(双峰/多峰)或无众数。
Weighted mean – each value is assigned a weight reflecting its relative importance: x̄w = Σwx / Σw.
加权平均数 – 每个值被赋予一个权重以反映其相对重要性:x̄w = Σwx / Σw。
Mid‑range – (max + min)/2, a rough measure of centre that is extremely sensitive to outliers.
中程数 – (最大值 + 最小值)/2,一个粗略的中心度量,对异常值极为敏感。
3. Measures of Dispersion | 离散程度度量
Range – largest value minus smallest value. Simple but heavily influenced by outliers.
极差 – 最大值减最小值。简单但严重受异常值影响。
Interquartile range (IQR) – Q₃ − Q₁. It measures the spread of the middle 50% and is resistant to outliers.
四分位距(IQR) – Q₃ − Q₁。衡量中间 50% 数据的分散程度,能抵抗异常值的影响。
Variance – the average of the squared deviations from the mean. For a population: σ² = Σ(x − μ)² / N; for a sample: s² = Σ(x − x̄)² / (n − 1).
方差 – 各数据与平均数之差的平方的平均值。总体:σ² = Σ(x − μ)² / N;样本:s² = Σ(x − x̄)² / (n − 1)。
Standard deviation – the square root of variance, denoted σ or s. It shares the same units as the data.
标准差 – 方差的平方根,记作 σ 或 s,与数据单位相同。
Mean absolute deviation – average of the absolute deviations |x − x̄| / n. Less common but gives a linear spread measure.
平均绝对偏差 – 绝对偏差 |x − x̄| 的平均数 / n。不常用,但提供了线性离散度量。
Outlier – a value that lies well outside the overall pattern. Often defined as any point < Q₁ − 1.5×IQR or > Q₃ + 1.5×IQR.
异常值 – 远在总体模式之外的值。常定义为小于 Q₁ − 1.5×IQR 或大于 Q₃ + 1.5×IQR 的点。
4. Probability Foundations | 概率基础
Random experiment – a process whose outcome cannot be predicted with certainty (e.g. rolling a die).
随机试验 – 结果无法确切预知的过程(如掷骰子)。
Sample space (S) – the set of all possible outcomes of an experiment.
样本空间(S) – 试验所有可能结果的集合。
Event – a subset of the sample space, e.g. ‘getting an even number’ = {2, 4, 6}.
事件 – 样本空间的子集,例如“得到偶数” = {2, 4, 6}。
Probability of an event – a number between 0 and 1 that measures how likely the event is. P(A) = number of favourable outcomes / total number of outcomes, if all outcomes equally likely.
事件的概率 – 介于 0 和 1 之间的数,衡量事件发生的可能性。若所有结果等可能,P(A) = 有利结果数 / 总结果数。
Complement of A (A’) – the event that A does not occur. P(A’) = 1 − P(A).
A 的补事件(A’) – A 不发生的事件。P(A’) = 1 − P(A)。
Mutually exclusive events – two events that cannot happen at the same time. P(A ∩ B) = 0.
互斥事件 – 两个不可能同时发生的事件。P(A ∩ B) = 0。
Addition rule for mutually exclusive events – P(A ∪ B) = P(A) + P(B).
互斥事件的加法公式 – P(A ∪ B) = P(A) + P(B)。
General addition rule – P(A ∪ B) = P(A) + P(B) − P(A ∩ B).
一般加法公式 – P(A ∪ B) = P(A) + P(B) − P(A ∩ B)。
5. Conditional Probability & Independence | 条件概率与独立
Conditional probability – the probability that event A occurs given that event B has already occurred: P(A|B) = P(A ∩ B) / P(B), provided P(B) > 0.
条件概率 – 在事件 B 已发生的条件下,事件 A 发生的概率:P(A|B) = P(A ∩ B) / P(B),前提 P(B) > 0。
Independent events – events A and B are independent if the occurrence of one does not affect the probability of the other. Two conditions: P(A|B) = P(A) and P(B|A) = P(B), or equivalently P(A ∩ B) = P(A) × P(B).
独立事件 – 若一个事件的发生不影响另一个事件发生的概率,则 A 与 B 独立。两个条件:P(A|B) = P(A) 且 P(B|A) = P(B),或等价地 P(A ∩ B) = P(A) × P(B)。
Tree diagram – a branching diagram showing all possible outcomes and their probabilities, especially useful for sequential or conditional problems. Multiply along branches, add across final outcomes.
树形图 – 展示所有可能结果及其概率的分支图,特别适用于序贯或条件问题。沿分支相乘,最终结果相加。
Venn diagram – a diagram using circles to represent sets; helps visualise unions, intersections and complements.
韦恩图 – 用圆表示集合的图示,有助于可视化并、交和补集。
Two‑way table – a contingency table for calculating joint, marginal and conditional probabilities.
双向表 – 列联表,用于计算联合、边缘和条件概率。
6. Discrete Random Variables | 离散随机变量
Random variable (r.v.) – a variable whose value depends on the outcome of a random experiment. Usually denoted by an upper‑case letter, e.g. X.
随机变量(r.v.) – 其值取决于随机试验结果的变量,通常用大写字母表示,如 X。
Discrete random variable – takes a countable number of distinct values (e.g. 0, 1, 2, …). Described by a probability distribution that lists each possible value and its probability.
离散随机变量 – 取可数个互不相同的值(如 0, 1, 2, …)。通过概率分布描述,列出各可能值及其概率。
Probability distribution (probability mass function) – a table, formula or graph giving P(X = x) for every possible x. Must satisfy Σ P(X = x) = 1.
概率分布(概率质量函数) – 以表格、公式或图表给出每个可能 x 的 P(X = x)。必须满足 Σ P(X = x) = 1。
Cumulative distribution function F(x) – F(x) = P(X ≤ x). Built by adding probabilities in the distribution table.
累积分布函数 F(x) – F(x) = P(X ≤ x)。通过累加分布表中概率得到。
7. Expectation and Variance | 期望与方差
Expected value E(X) – the long‑run average of the random variable: E(X) = Σ [x · P(X = x)]. Often denoted by μ.
期望值 E(X) – 随机变量的长期平均值:E(X) = Σ [x · P(X = x)],常记作 μ。
Variance Var(X) – a measure of spread for a random variable: Var(X) = E[(X − μ)²] = E(X²) − [E(X)]².
方差 Var(X) – 衡量随机变量离散程度的量:Var(X) = E[(X − μ)²] = E(X²) − [E(X)]²。
Standard deviation of X – σ = √Var(X).
随机变量 X 的标准差 – σ = √Var(X)。
Linear transformations – For Y = aX + b, E(Y) = aE(X) + b and Var(Y) = a² Var(X). (Adding b shifts the centre but does not affect spread.)
线性变换 – 对于 Y = aX + b,E(Y) = aE(X) + b,Var(Y) = a² Var(X)。(加 b 平移中心但不影响离散度。)
8. Binomial Distribution | 二项分布
Conditions for a binomial model – a fixed number of trials n, each trial is independent, only two outcomes (success/failure), and the probability of success p is constant for every trial. Mnemonic: BINS (Binary, Independent, Number fixed, Same p).
二项模型的条件 – 固定试验次数 n,每次试验独立,只有两种结果(成功/失败),且每次试验成功的概率 p 不变。速记法:BINS(二元、独立、固定次数、相同 p)。
Binomial random variable – X ~ B(n, p) counts the number of successes in n trials. P(X = r) = C(n, r) × pʳ × (1 − p)ⁿ⁻ʳ, for r = 0,1,…, n.
二项随机变量 – X ~ B(n, p) 记录 n 次试验中的成功次数。P(X = r) = C(n, r) × pʳ × (1 − p)ⁿ⁻ʳ,r = 0,1,…, n。
nCr (binomial coefficient) – C(n, r) = n! / [r! (n − r)!], the number of ways to choose r items from n. Many calculators label it ‘nCr’.
nCr(二项式系数) – C(n, r) = n! / [r! (n − r)!],表示从 n 个物品中选出 r 个的方法数。许多计算器标注为“nCr”。
Expectation and variance for binomial – E(X) = np, Var(X) = np(1 − p).
二项分布的期望与方差 – E(X) = np,Var(X) = np(1 − p)。
Cumulative binomial probability – P(X ≤ r) can be found using tables or by summing individual terms. In CAIE exams, always read the inequality direction carefully.
累积二项概率 – P(X ≤ r) 可用表格或逐项求和求得。CAIE 考试中务必仔细阅读不等号方向。
9. Geometric Distribution | 几何分布
Conditions for a geometric model – repeated independent trials, each with constant p, until the first success occurs. The number of trials is not fixed; it depends on when success happens.
几何模型的条件 – 重复独立试验,每次 p 不变,直到首次成功为止。试验次数不固定,取决于成功发生的时刻。
Geometric random variable – X ~ Geo(p) can be defined in two ways: (1) the number of trials up to and including the first success: P(X = x) = (1 − p)ˣ⁻¹ p, x = 1,2,…; or (2) the number of failures before the first success. CAIE uses definition (1). Confirm which version your syllabus uses.
几何随机变量 – X ~ Geo(p) 有两种定义方式:(1)直到并包括首次成功时的试验次数:P(X = x) = (1 − p)ˣ⁻¹ p,x = 1,2,…;(2)首次成功前的失败次数。CAIE 使用定义 (1)。务必确认你的教学大纲采用的版本。
Expectation and variance for geometric – Using the ‘trials to first success’ definition: E(X) = 1/p, Var(X) = (1 − p) / p².
几何分布的期望与方差 – 采用“到首次成功的试验次数”定义:E(X) = 1/p,Var(X) = (1 − p) / p²。
‘At least’ and ‘more than’ in geometric distribution – P(X > k) = (1 − p)ᵏ, because the first k trials must all be failures. This property makes geometric distribution memoryless.
几何分布中的“至少”和“多于” – P(X > k) = (1 − p)ᵏ,因为前 k 次试验必须全为失败。此性质使几何分布具有无记忆性。
10. Normal Distribution | 正态分布
Normal random variable – X ~ N(μ, σ²). Bell‑shaped, symmetric about μ. The total area under the curve is 1.
正态随机变量 – X ~ N(μ, σ²)。呈钟形,关于 μ 对称。曲线下的总面积为 1。
Standardising – Z = (X − μ) / σ. This transforms any normal variable into the standard normal Z ~ N(0, 1²).
标准化 – Z = (X − μ) / σ,将任意正态变量转化为标准正态 Z ~ N(0, 1²)。
Standard normal table – provides cumulative probabilities Φ(z) = P(Z < z). For negative z, use symmetry: Φ(−z) = 1 − Φ(z).
标准正态表 – 提供累积概率 Φ(z) = P(Z < z)。对于负的 z,利用对称性:Φ(−z) = 1 − Φ(z)。
Finding an unknown μ or σ – set up the standardising equation and use the inverse normal (percentage points) to read the z‑value corresponding to a given probability.
求未知的 μ 或 σ – 建立标准化方程,并用逆正态(百分位点)查出对应给定概率的 z 值。
Normal approximation to binomial – when n is large and p is close to 0.5, X ~ B(n, p) can be approximated by N(np, np(1 − p)). Apply a continuity correction: adjust the discrete boundary by ± 0.5.
正态近似二项分布 – 当 n 较大且 p 接近 0.5 时,X ~ B(n, p) 可用 N(np, np(1 − p)) 近似。需使用连续性修正:将离散边界调整 ± 0.5。
Continuity correction formula – P(X ≥ r) ≈ P(Y > r − 0.5) where Y ~ N(np, np(1−p)). Always draw a sketch to decide whether to add or subtract 0.5.
连续性修正公式 – P(X ≥ r) ≈ P(Y > r − 0.5),式中 Y ~ N(np, np(1−p))。务必画草图以确定加或减 0.5。
11. Quick Memory Aids & Commonly Confused Terms | 速记技巧与常见混淆术语
BINS – the four conditions for a binomial distribution: Binary outcomes, Independent trials, Number fixed, Same probability p. Stick this acronym on your exam day revision card.
BINS – 二项分布的四个条件:二元结果、独立试验、固定次数、相同概率 p。把这条首字母缩写贴在考试当天的复习卡上。
Mean vs median in skew – In a right‑skewed distribution, mean > median; in a left‑skewed distribution, mean < median. For symmetric data, they coincide.
偏态中的平均数与中位数 – 右偏分布中,平均数 > 中位数;左偏分布中,平均数 < 中位数。对称数据中二者相等。
Standard deviation vs variance – variance uses squared units (awkward to interpret); standard deviation restores the original units. The formulae look similar – don’t grab the wrong one.
标准差与方差 – 方差使用平方单位(解读别扭);标准差还原为原始单位。公式看起来相似,别拿错。
Mutually exclusive ≠ independence – mutually exclusive events cannot happen together, so P(A ∩ B) = 0. Independent events can happen together; their joint probability is the product. The two concepts are completely different.
互斥 ≠ 独立 – 互斥事件不可能同时发生,故 P(A ∩ B) = 0。独立事件可以一同发生,其联合概率为乘积。两个概念截然不同。
P(X = x) vs P(X ≤ x) – many exam errors happen because students confuse individual and cumulative probabilities. Highlight the inequality symbol before using the binomial or geometric formula.
P(X = x) 与 P(X ≤ x) – 许多考试错误源于学生混淆单独概率与累积概率。在使用二项或几何公式前,用高亮标出不等号。
Continuity correction – ‘include or exclude the half?’ Write a simple rule: for ≥ r, use > r − 0.5; for ≤ r, use < r + 0.5. Verify with a quick diagram.
连续性修正 – “减半还是加半?”记住简单规则:≥ r 用 > r − 0.5;≤ r 用 < r + 0.5。用简图快速验证。
Geometric vs binomial – ‘How many trials until first success?’ ⇒ geometric. ‘How many successes in n trials?’ ⇒ binomial. The wording of the question is the key.
几何分布与二项分布 – “多少次试验才出现首次成功?”⇒ 几何。“在 n 次试验中有多少次成功?”⇒ 二项。题目措辞是关键。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply