Year 12 CAIE Statistics: Glossary and Memorisation Guide | Year 12 CAIE 统计:词汇术语速记指南

📚 Year 12 CAIE Statistics: Glossary and Memorisation Guide | Year 12 CAIE 统计:词汇术语速记指南

In CAIE Year 12 Statistics, mastering the terminology is half the battle. Whether you are interpreting a box‑and‑whisker plot, calculating a binomial probability or working with the normal distribution, precise language leads to accurate solutions. This bilingual guide pairs each key term with its Chinese equivalent and a short explanation, organised by topic for quick revision. Use it to strengthen your statistical vocabulary and avoid common misunderstandings in the exam.

在 CAIE 12 年级统计中,掌握术语是成功的一半。无论你在解读箱线图、计算二项概率还是处理正态分布,精准的语言都能带来准确的解答。这份双语指南将每个关键术语配以中文对应词和简短解释,按主题编排,便于快速复习。用它来强化统计词汇,避免考试中常见的理解偏差。


1. Descriptive Statistics: Data Types & Charts | 描述统计:数据类型与图表

Raw data – information collected before any processing, listed as individual numbers or categories.

原始数据 – 收集后未经任何处理的信息,以单个数字或类别列出。

Qualitative data – non‑numerical observations (e.g. colour, gender). Often summarised in bar charts or pie charts.

定性数据 – 非数值型的观测值(如颜色、性别),通常用条形图或饼图汇总。

Quantitative data – numerical observations that can be discrete (countable, e.g. number of cars) or continuous (measurable, e.g. height).

定量数据 – 数值型的观测值,可以是离散型(可数,如汽车数量)或连续型(可测量,如身高)。

Frequency – the number of times a data value appears in a dataset.

频数 – 某个数据值在数据集中出现的次数。

Cumulative frequency – the running total of frequencies up to a given value. Used to construct cumulative frequency curves (ogives).

累积频数 – 截至某一给定值的频数累加总数,用于绘制累积频数曲线(肩形图)。

Class interval – a group of data values defined by a lower and an upper boundary. For continuous data, class boundaries remove gaps between intervals.

组距 – 由下界和上界定义的一组数据值。对于连续数据,组界消除了组间的空隙。

Histogram – a chart where the area of each bar is proportional to the frequency. For unequal widths, frequency density = frequency ÷ class width.

直方图 – 一种图表,其中每个长方形的面积与频数成比例。当组距不等时,用频数密度 = 频数 ÷ 组距。

Box‑and‑whisker plot – displays the minimum, lower quartile (Q₁), median (Q₂), upper quartile (Q₃) and maximum. Outliers are shown as separate points.

箱线图 – 展示最小值、下四分位数(Q₁)、中位数(Q₂)、上四分位数(Q₃)和最大值。异常值单独标出。


2. Measures of Central Tendency | 集中趋势度量

Mean – the arithmetic average, x̄ = Σx / n for raw data, or Σfx / Σf for grouped data. Affected by extreme values.

平均数 – 算术平均值,原始数据用 x̄ = Σx / n,分组数据用 Σfx / Σf。受极端值影响。

Median – the middle value when data are ordered. For n observations, position = (n+1)/2. Not swayed by outliers.

中位数 – 数据排序后位于中间的值。n 个观测值时,位置为 (n+1)/2。不受异常值影响。

Mode – the most frequently occurring value. A dataset can have one mode, more than one mode (bimodal/multimodal) or no mode at all.

众数 – 出现次数最多的值。数据集可以有单众数、多个众数(双峰/多峰)或无众数。

Weighted mean – each value is assigned a weight reflecting its relative importance: x̄w = Σwx / Σw.

加权平均数 – 每个值被赋予一个权重以反映其相对重要性:x̄w = Σwx / Σw。

Mid‑range – (max + min)/2, a rough measure of centre that is extremely sensitive to outliers.

中程数 – (最大值 + 最小值)/2,一个粗略的中心度量,对异常值极为敏感。


3. Measures of Dispersion | 离散程度度量

Range – largest value minus smallest value. Simple but heavily influenced by outliers.

极差 – 最大值减最小值。简单但严重受异常值影响。

Interquartile range (IQR) – Q₃ − Q₁. It measures the spread of the middle 50% and is resistant to outliers.

四分位距(IQR) – Q₃ − Q₁。衡量中间 50% 数据的分散程度,能抵抗异常值的影响。

Variance – the average of the squared deviations from the mean. For a population: σ² = Σ(x − μ)² / N; for a sample: s² = Σ(x − x̄)² / (n − 1).

方差 – 各数据与平均数之差的平方的平均值。总体:σ² = Σ(x − μ)² / N;样本:s² = Σ(x − x̄)² / (n − 1)。

Standard deviation – the square root of variance, denoted σ or s. It shares the same units as the data.

标准差 – 方差的平方根,记作 σ 或 s,与数据单位相同。

Mean absolute deviation – average of the absolute deviations |x − x̄| / n. Less common but gives a linear spread measure.

平均绝对偏差 – 绝对偏差 |x − x̄| 的平均数 / n。不常用,但提供了线性离散度量。

Outlier – a value that lies well outside the overall pattern. Often defined as any point < Q₁ − 1.5×IQR or > Q₃ + 1.5×IQR.

异常值 – 远在总体模式之外的值。常定义为小于 Q₁ − 1.5×IQR 或大于 Q₃ + 1.5×IQR 的点。


4. Probability Foundations | 概率基础

Random experiment – a process whose outcome cannot be predicted with certainty (e.g. rolling a die).

随机试验 – 结果无法确切预知的过程(如掷骰子)。

Sample space (S) – the set of all possible outcomes of an experiment.

样本空间(S) – 试验所有可能结果的集合。

Event – a subset of the sample space, e.g. ‘getting an even number’ = {2, 4, 6}.

事件 – 样本空间的子集,例如“得到偶数” = {2, 4, 6}。

Probability of an event – a number between 0 and 1 that measures how likely the event is. P(A) = number of favourable outcomes / total number of outcomes, if all outcomes equally likely.

事件的概率 – 介于 0 和 1 之间的数,衡量事件发生的可能性。若所有结果等可能,P(A) = 有利结果数 / 总结果数。

Complement of A (A’) – the event that A does not occur. P(A’) = 1 − P(A).

A 的补事件(A’) – A 不发生的事件。P(A’) = 1 − P(A)。

Mutually exclusive events – two events that cannot happen at the same time. P(A ∩ B) = 0.

互斥事件 – 两个不可能同时发生的事件。P(A ∩ B) = 0。

Addition rule for mutually exclusive events – P(A ∪ B) = P(A) + P(B).

互斥事件的加法公式 – P(A ∪ B) = P(A) + P(B)。

General addition rule – P(A ∪ B) = P(A) + P(B) − P(A ∩ B).

一般加法公式 – P(A ∪ B) = P(A) + P(B) − P(A ∩ B)。


5. Conditional Probability & Independence | 条件概率与独立

Conditional probability – the probability that event A occurs given that event B has already occurred: P(A|B) = P(A ∩ B) / P(B), provided P(B) > 0.

条件概率 – 在事件 B 已发生的条件下,事件 A 发生的概率:P(A|B) = P(A ∩ B) / P(B),前提 P(B) > 0。

Independent events – events A and B are independent if the occurrence of one does not affect the probability of the other. Two conditions: P(A|B) = P(A) and P(B|A) = P(B), or equivalently P(A ∩ B) = P(A) × P(B).

独立事件 – 若一个事件的发生不影响另一个事件发生的概率,则 A 与 B 独立。两个条件:P(A|B) = P(A) 且 P(B|A) = P(B),或等价地 P(A ∩ B) = P(A) × P(B)。

Tree diagram – a branching diagram showing all possible outcomes and their probabilities, especially useful for sequential or conditional problems. Multiply along branches, add across final outcomes.

树形图 – 展示所有可能结果及其概率的分支图,特别适用于序贯或条件问题。沿分支相乘,最终结果相加。

Venn diagram – a diagram using circles to represent sets; helps visualise unions, intersections and complements.

韦恩图 – 用圆表示集合的图示,有助于可视化并、交和补集。

Two‑way table – a contingency table for calculating joint, marginal and conditional probabilities.

双向表 – 列联表,用于计算联合、边缘和条件概率。


6. Discrete Random Variables | 离散随机变量

Random variable (r.v.) – a variable whose value depends on the outcome of a random experiment. Usually denoted by an upper‑case letter, e.g. X.

随机变量(r.v.) – 其值取决于随机试验结果的变量,通常用大写字母表示,如 X。

Discrete random variable – takes a countable number of distinct values (e.g. 0, 1, 2, …). Described by a probability distribution that lists each possible value and its probability.

离散随机变量 – 取可数个互不相同的值(如 0, 1, 2, …)。通过概率分布描述,列出各可能值及其概率。

Probability distribution (probability mass function) – a table, formula or graph giving P(X = x) for every possible x. Must satisfy Σ P(X = x) = 1.

概率分布(概率质量函数) – 以表格、公式或图表给出每个可能 x 的 P(X = x)。必须满足 Σ P(X = x) = 1。

Cumulative distribution function F(x) – F(x) = P(X ≤ x). Built by adding probabilities in the distribution table.

累积分布函数 F(x) – F(x) = P(X ≤ x)。通过累加分布表中概率得到。


7. Expectation and Variance | 期望与方差

Expected value E(X) – the long‑run average of the random variable: E(X) = Σ [x · P(X = x)]. Often denoted by μ.

期望值 E(X) – 随机变量的长期平均值:E(X) = Σ [x · P(X = x)],常记作 μ。

Variance Var(X) – a measure of spread for a random variable: Var(X) = E[(X − μ)²] = E(X²) − [E(X)]².

方差 Var(X) – 衡量随机变量离散程度的量:Var(X) = E[(X − μ)²] = E(X²) − [E(X)]²。

Standard deviation of X – σ = √Var(X).

随机变量 X 的标准差 – σ = √Var(X)。

Linear transformations – For Y = aX + b, E(Y) = aE(X) + b and Var(Y) = a² Var(X). (Adding b shifts the centre but does not affect spread.)

线性变换 – 对于 Y = aX + b,E(Y) = aE(X) + b,Var(Y) = a² Var(X)。(加 b 平移中心但不影响离散度。)


8. Binomial Distribution | 二项分布

Conditions for a binomial model – a fixed number of trials n, each trial is independent, only two outcomes (success/failure), and the probability of success p is constant for every trial. Mnemonic: BINS (Binary, Independent, Number fixed, Same p).

二项模型的条件 – 固定试验次数 n,每次试验独立,只有两种结果(成功/失败),且每次试验成功的概率 p 不变。速记法:BINS(二元、独立、固定次数、相同 p)。

Binomial random variable – X ~ B(n, p) counts the number of successes in n trials. P(X = r) = C(n, r) × pʳ × (1 − p)ⁿ⁻ʳ, for r = 0,1,…, n.

二项随机变量 – X ~ B(n, p) 记录 n 次试验中的成功次数。P(X = r) = C(n, r) × pʳ × (1 − p)ⁿ⁻ʳ,r = 0,1,…, n。

nCr (binomial coefficient) – C(n, r) = n! / [r! (n − r)!], the number of ways to choose r items from n. Many calculators label it ‘nCr’.

nCr(二项式系数) – C(n, r) = n! / [r! (n − r)!],表示从 n 个物品中选出 r 个的方法数。许多计算器标注为“nCr”。

Expectation and variance for binomial – E(X) = np, Var(X) = np(1 − p).

二项分布的期望与方差 – E(X) = np,Var(X) = np(1 − p)。

Cumulative binomial probability – P(X ≤ r) can be found using tables or by summing individual terms. In CAIE exams, always read the inequality direction carefully.

累积二项概率 – P(X ≤ r) 可用表格或逐项求和求得。CAIE 考试中务必仔细阅读不等号方向。


9. Geometric Distribution | 几何分布

Conditions for a geometric model – repeated independent trials, each with constant p, until the first success occurs. The number of trials is not fixed; it depends on when success happens.

几何模型的条件 – 重复独立试验,每次 p 不变,直到首次成功为止。试验次数不固定,取决于成功发生的时刻。

Geometric random variable – X ~ Geo(p) can be defined in two ways: (1) the number of trials up to and including the first success: P(X = x) = (1 − p)ˣ⁻¹ p, x = 1,2,…; or (2) the number of failures before the first success. CAIE uses definition (1). Confirm which version your syllabus uses.

几何随机变量 – X ~ Geo(p) 有两种定义方式:(1)直到并包括首次成功时的试验次数:P(X = x) = (1 − p)ˣ⁻¹ p,x = 1,2,…;(2)首次成功前的失败次数。CAIE 使用定义 (1)。务必确认你的教学大纲采用的版本。

Expectation and variance for geometric – Using the ‘trials to first success’ definition: E(X) = 1/p, Var(X) = (1 − p) / p².

几何分布的期望与方差 – 采用“到首次成功的试验次数”定义:E(X) = 1/p,Var(X) = (1 − p) / p²。

‘At least’ and ‘more than’ in geometric distribution – P(X > k) = (1 − p)ᵏ, because the first k trials must all be failures. This property makes geometric distribution memoryless.

几何分布中的“至少”和“多于” – P(X > k) = (1 − p)ᵏ,因为前 k 次试验必须全为失败。此性质使几何分布具有无记忆性。


10. Normal Distribution | 正态分布

Normal random variable – X ~ N(μ, σ²). Bell‑shaped, symmetric about μ. The total area under the curve is 1.

正态随机变量 – X ~ N(μ, σ²)。呈钟形,关于 μ 对称。曲线下的总面积为 1。

Standardising – Z = (X − μ) / σ. This transforms any normal variable into the standard normal Z ~ N(0, 1²).

标准化 – Z = (X − μ) / σ,将任意正态变量转化为标准正态 Z ~ N(0, 1²)。

Standard normal table – provides cumulative probabilities Φ(z) = P(Z < z). For negative z, use symmetry: Φ(−z) = 1 − Φ(z).

标准正态表 – 提供累积概率 Φ(z) = P(Z < z)。对于负的 z,利用对称性:Φ(−z) = 1 − Φ(z)。

Finding an unknown μ or σ – set up the standardising equation and use the inverse normal (percentage points) to read the z‑value corresponding to a given probability.

求未知的 μ 或 σ – 建立标准化方程,并用逆正态(百分位点)查出对应给定概率的 z 值。

Normal approximation to binomial – when n is large and p is close to 0.5, X ~ B(n, p) can be approximated by N(np, np(1 − p)). Apply a continuity correction: adjust the discrete boundary by ± 0.5.

正态近似二项分布 – 当 n 较大且 p 接近 0.5 时,X ~ B(n, p) 可用 N(np, np(1 − p)) 近似。需使用连续性修正:将离散边界调整 ± 0.5。

Continuity correction formula – P(X ≥ r) ≈ P(Y > r − 0.5) where Y ~ N(np, np(1−p)). Always draw a sketch to decide whether to add or subtract 0.5.

连续性修正公式 – P(X ≥ r) ≈ P(Y > r − 0.5),式中 Y ~ N(np, np(1−p))。务必画草图以确定加或减 0.5。


11. Quick Memory Aids & Commonly Confused Terms | 速记技巧与常见混淆术语

BINS – the four conditions for a binomial distribution: Binary outcomes, Independent trials, Number fixed, Same probability p. Stick this acronym on your exam day revision card.

BINS – 二项分布的四个条件:二元结果、独立试验、固定次数、相同概率 p。把这条首字母缩写贴在考试当天的复习卡上。

Mean vs median in skew – In a right‑skewed distribution, mean > median; in a left‑skewed distribution, mean < median. For symmetric data, they coincide.

偏态中的平均数与中位数 – 右偏分布中,平均数 > 中位数;左偏分布中,平均数 < 中位数。对称数据中二者相等。

Standard deviation vs variance – variance uses squared units (awkward to interpret); standard deviation restores the original units. The formulae look similar – don’t grab the wrong one.

标准差与方差 – 方差使用平方单位(解读别扭);标准差还原为原始单位。公式看起来相似,别拿错。

Mutually exclusive ≠ independence – mutually exclusive events cannot happen together, so P(A ∩ B) = 0. Independent events can happen together; their joint probability is the product. The two concepts are completely different.

互斥 ≠ 独立 – 互斥事件不可能同时发生,故 P(A ∩ B) = 0。独立事件可以一同发生,其联合概率为乘积。两个概念截然不同。

P(X = x) vs P(X ≤ x) – many exam errors happen because students confuse individual and cumulative probabilities. Highlight the inequality symbol before using the binomial or geometric formula.

P(X = x) 与 P(X ≤ x) – 许多考试错误源于学生混淆单独概率与累积概率。在使用二项或几何公式前,用高亮标出不等号。

Continuity correction – ‘include or exclude the half?’ Write a simple rule: for ≥ r, use > r − 0.5; for ≤ r, use < r + 0.5. Verify with a quick diagram.

连续性修正 – “减半还是加半?”记住简单规则:≥ r 用 > r − 0.5;≤ r 用 < r + 0.5。用简图快速验证。

Geometric vs binomial – ‘How many trials until first success?’ ⇒ geometric. ‘How many successes in n trials?’ ⇒ binomial. The wording of the question is the key.

几何分布与二项分布 – “多少次试验才出现首次成功?”⇒ 几何。“在 n 次试验中有多少次成功?”⇒ 二项。题目措辞是关键。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version