A-Level CAIE Statistics: Quick Memorisation Guide to Key Terms | A-Level CAIE 统计:关键术语速记指南

📚 A-Level CAIE Statistics: Quick Memorisation Guide to Key Terms | A-Level CAIE 统计:关键术语速记指南

Mastering statistical terminology is essential for success in the CAIE A-Level Statistics exam. Terms like ‘discrete variable’, ‘hypothesis test’ or ‘correlation coefficient’ frequently appear in both multiple-choice and structured questions, and mixing them up can cost valuable marks. This guide pairs key concepts with bilingual explanations and memory tricks, helping you internalise the language of statistics quickly and confidently.

掌握统计术语是拿下 CAIE A-Level 统计考试的关键。’离散变量’、’假设检验’、’相关系数’等概念经常出现在选择题和综合题中,稍有不慎便可能混淆丢分。本指南为每个核心术语提供中英双语解释和速记技巧,帮助你快速自信地内化统计语言。

1. Types of Data | 数据类型

Data can be qualitative (descriptive) or quantitative (numerical). Quantitative data further splits into discrete and continuous types. Recognising the difference determines which charts and summary statistics to use.

数据可分为定性(描述性)和定量(数值型)。定量数据又分为离散型和连续型。区分它们决定了该使用哪种图表和汇总统计量。

Qualitative data: Non-numerical categories, e.g. eye colour, brand of phone. Think ‘quality’ not ‘quantity’.

定性数据:非数字类别,例如眼睛颜色、手机品牌。记法:’定性’想的是’性质’而非’数量’。

Discrete data: Countable number of outcomes, e.g. number of heads in 10 coin flips. Imagine whole, separate steps – like counting people, you cannot have 2.3 people.

离散数据:可数结果,例如投10次硬币正面的次数。想象一个个单独的台阶——就像数人数,你不能有2.3个人。

Continuous data: Measured on a scale, can take any value within a range, e.g. height, time. Memory tip: ‘Continuous’ = ‘continuum’, like the smooth flow of water.

连续数据:在某个范围内可以取任意值,例如身高、时间。记忆技巧:’连续’就像水流一样不间断的连续体。


2. Measures of Central Tendency | 集中趋势测量

The three musketeers of central tendency are mean, median and mode. Each gives a ‘typical’ value but behaves differently in skewed data. Using the wrong one can mislead conclusions.

集中趋势的三剑客是均值、中位数和众数。它们各自给出 ‘典型’值,但在偏态数据中表现不同,用错了会误导结论。

Mean (μ for population, x̄ for sample): The arithmetic average, sum of all values divided by the count. Mnemonic: ‘Mean is the balancing point’ – a see-saw balances at the mean.

均值(总体μ,样本x̄):算术平均数,所有数值之和除以个数。记忆法:’均值就像平衡点’——跷跷板在均值处平衡。

Median: The middle value when data is ordered. For grouped data use linear interpolation: Median = L + ( (n/2 – F) / f ) × w. Visualise a street where half the houses are on one side of the median.

中位数:排序后中间的值。分组数据用线性插值:中位数 = L + ( (n/2 – F) / f ) × w。想象一条街,中位数位于一半房屋的这一边。

Mode: The most frequent value. In a histogram, it is the highest bar. ‘Mode’ starts with ‘M’ for ‘Most’.

众数:出现频率最高的值。在直方图中是最高的柱子。’众’就是众多,最多的那个。

When data is symmetric, mean ≈ median. In right-skewed data, mean > median; left-skewed, mean < median. Remember the tail pulls the mean.

当数据对称,均值总位数。右偏时均值 > 中位数;左偏时均值 < 中位数。记住,尾巴把均值拉跑了。


3. Measures of Dispersion | 离散程度测量

Dispersion tells us how spread out the data are. Averages alone don’t capture variability – two classes can have the same mean score but very different consistency.

离散程度反映数据的分散情况。仅看平均值无法捕捉变异性——两个班级可以有相同的平均分,但成绩的稳定性截然不同。

Range = maximum − minimum. Quick but sensitive to outliers. Memory: ‘Range’ covers the whole ‘range of mountains’ from lowest to highest peak.

极差 = 最大值 − 最小值。计算快但易受异常值影响。记忆:’极差’就是山脉从最低到最高峰的全跨度。

Interquartile Range (IQR) = Q₃ − Q₁. The middle 50% of data; robust against outliers. Think of the middle half of a box plot.

四分位距 (IQR) = 上四分位数Q₃ − 下四分位数Q₁。覆盖中间50%数据,抵抗异常值。想象箱线图中箱子的长度。

Variance (σ² for population, s² for sample): The average of squared deviations from the mean. Formula: σ² = Σ(xᵢ – μ)² / N. It’s in squared units, so not directly comparable to original data.

方差(总体σ²,样本s²):离差平方的平均值。公式:σ² = Σ(xᵢ – μ)² / N。单位是原单位的平方,因此不能直接与原数据比较。

Standard deviation (σ or s): Square root of variance, restoring original units. Use ‘s = √(Σ(x – x̄)² / (n-1))’ for sample. Mnemonic: ‘Standard deviation is the ruler of spread’ – it tells you how far typical values stray from the mean.

标准差(σ 或 s):方差的平方根,恢复原单位。样本公式:s = √(Σ(x – x̄)² / (n-1))。记忆口诀:’标准差是离散的标尺’——衡量典型值偏离均值的距离。


4. Probability Fundamentals | 概率基础

Probability values range from 0 (impossible) to 1 (certain). Understanding sample spaces, events and set operations forms the bedrock for later topics like distributions and hypothesis testing.

概率值范围从0(不可能)到1(必然)。理解样本空间、事件和集合运算是分布、假设检验等后续内容的基石。

Sample space (S): The set of all possible outcomes. E.g. rolling a die, S = {1,2,3,4,5,6}. Visualise it as the ‘universe’ of your experiment.

样本空间 (S):所有可能结果的集合。例如掷骰子,S = {1,2,3,4,5,6}。把它想象成你实验的’宇宙’。

Mutually exclusive events cannot happen at the same time. P(A ∩ B) = 0. Mnemonic: ‘Mutually exclusive = Mutual exclusion’, like you can’t turn left and right simultaneously.

互斥事件不能同时发生。P(A ∩ B) = 0。记法:’互斥’就是互相排斥,就像无法同时向左转又向右转。

Independent events: The occurrence of one does not affect the probability of the other. P(A ∩ B) = P(A) × P(B). Test with P(A|B) = P(A). ‘Independent = Indifferent to each other’.

独立事件:一个事件的发生不影响另一个的概率。P(A ∩ B) = P(A) × P(B)。检验:P(A|B) = P(A)。’独立’就是彼此无关,互不影响。

Conditional probability: P(A|B) = P(A ∩ B) / P(B). ‘Given that B has happened, how likely is A?’ The vertical bar ‘|’ can be read as ‘given’.

条件概率:P(A|B) = P(A ∩ B) / P(B)。’已知B发生,A发生的概率是多少?’ 竖线 ‘|’ 读作’给定’。


5. Permutations and Combinations | 排列与组合

Counting principles underpin probability calculations. Permutations care about order, combinations do not. Mixing them up leads to wrong denominators in probability problems.

计数原理是概率计算的根基。排列关心顺序,组合不关心。两者混淆会让概率问题的分母出错。

Factorial n! = n × (n-1) × … × 1. It counts arrangements of n distinct items. 0! = 1 by definition. Think ‘!’ as a big surprise – the number gets huge quickly!

阶乘 n! = n × (n-1) × … × 1。计算n个不同物品的排列数。定义0! = 1。想象’!’像一个惊叹号,数字极速增长!

Permutation ⁿPᵣ = n! / (n-r)!: The number of ways to arrange r items out of n, order matters. E.g., race podium finishes (gold, silver, bronze). ‘Permutation = Position matters’.

排列 ⁿPᵣ = n! / (n-r)!:从n个物品中选r个排列,顺序重要。例如领奖台名次(金、银、铜)。’排列’的’排’就是排顺序。

Combination ⁿCᵣ = n! / (r!(n-r)!): Number of ways to choose r items from n, order doesn’t matter. E.g., selecting a committee. ‘Combination = Committee – no ranking’.

组合 ⁿCᵣ = n! / (r!(n-r)!):从n个中选r个,不计顺序。例如组建委员会。’组合’的’组’就像组队,不用排队。


6. Probability Distributions | 概率分布

Distributions model the probabilities of different outcomes. The binomial handles success/failure trials, while the normal describes continuous bell-shaped data. Master their parameters and conditions.

概率分布对不同结果赋予概率。二项分布处理成功/失败试验,正态分布描述连续钟形数据。掌握其参数和适用条件。

Binomial distribution B(n, p): Fixed number n of independent trials, each with same probability p of success. Mean = np, variance = np(1-p). Use BINS mnemonic: Binary, Independent, Number fixed, Same probability.

二项分布 B(n, p):固定次数n的独立试验,每次成功概率p相同。均值 = np,方差 = np(1-p)。用 BINS 口诀:二元结果、独立、次数固定、概率相同。

Normal distribution N(μ, σ²): Bell-shaped, symmetric about mean μ. 68% of data within μ ± σ, 95% within μ ± 2σ. Transform to standard normal Z = (X – μ) / σ. Think: ‘Normal is the natural bell curve, like heights of adults’.

正态分布 N(μ, σ²):钟形,关于均值μ对称。68%数据落在μ ± σ内,95%在μ ± 2σ内。转换为标准正态 Z = (X – μ) / σ。想象成年人的身高——天然钟形。

Expectation E(X) is the long-run average. For discrete, E(X) = Σ x·P(X=x). For continuous, integrate. ‘Expectation’ is what you ‘expect’ on average.

期望 E(X) 是长期平均值。离散型:E(X) = Σ x·P(X=x)。连续型用积分。’期望’就是你平均意义上’期望’得到的值。


7. Sampling and Estimation | 抽样与估计

We rarely have access to an entire population, so we take samples. A statistic from a sample (like x̄) estimates a population parameter (like μ). The sampling distribution tells us how much it varies.

我们很少能获取整个总体,因此抽样。样本计算出的统计量(如x̄)用来估计总体参数(如μ)。抽样分布告诉我们它有多大变异性。

Population vs sample: Population is the whole group; a sample is a subset. Parameter (e.g. μ) is fixed, while statistic (e.g. x̄) varies from sample to sample. ‘Parameter = Population; Statistic = Sample’ (both start with corresponding letters).

总体 vs 样本:总体是整个群体,样本是子集。参数(如μ)固定,统计量(如x̄)随样本而变。’Parameter’ 对应 ‘Population’,’Statistic’ 对应 ‘Sample’。

Random sample: Every member has an equal chance of being chosen. Simple random sampling avoids bias. ‘Lottery method’ is an easy mental image.

随机样本:每个成员被抽到的机会均等。简单随机抽样避免偏差。想象抽签方法。

Central Limit Theorem (CLT): For large enough sample size (n ≥ 30), the sampling distribution of the sample mean is approximately normal, with mean μ and standard error σ/√n. This miracle allows us to use normal techniques even when the population is not normal.

中心极限定理 (CLT):当样本量足够大(n ≥ 30),样本均值的抽样分布近似正态,均值为μ,标准误为σ/√n。这个’奇迹’使得即使总体非正态,也能使用正态方法。


8. Hypothesis Testing | 假设检验

Hypothesis testing is the backbone of statistical decision-making. We assume a null hypothesis, gather evidence, and decide whether to reject it. Misunderstanding p-values or errors is a common pitfall.

假设检验是统计决策的骨干。我们假设一个零假设,收集证据,并决定是否拒绝它。误解p值或两类错误是常见陷阱。

Null hypothesis H₀: The default assumption, usually ‘no effect’ or ‘no difference’. Alternative hypothesis H₁: The claim we seek evidence for. ‘Null means nothing new’.

零假设 H₀:默认假设,通常是 ‘无效果’ 或 ‘无差异’。备择假设 H₁:我们寻找证据支持的声明。’零’代表’归零’,即没有新东西。

Significance level α: The probability of rejecting H₀ when it is actually true (Type I error). Common α = 0.05. ‘α is the risk you’re willing to take for a false alarm’.

显著性水平 α:当 H₀ 为真时拒绝它的概率(第一类错误)。常用 α = 0.05。’α 是你愿意承担的假警报风险’。

p-value: Probability of obtaining a test statistic at least as extreme as the observed, assuming H₀ is true. If p-value < α, reject H₀. Memory: 'Small p-value, reject null; big p-value, fail to reject.'

p值:在 H₀ 为真的条件下,获得至少与观测值同样极端的检验统计量的概率。若 p值 < α,拒绝 H₀。口诀:'p小拒零,p大留零'。

Type I error: False positive – rejecting a true H₀. Type II error: False negative – failing to reject a false H₀. Use the medical analogy: Type I = diagnosing a healthy person as sick; Type II = missing a sick person.

第一类错误:假阳性——拒绝正确的 H₀。第二类错误:假阴性——未能拒绝错误的 H₀。用医学类比:第一类 = 把健康人误诊为生病;第二类 = 漏诊病人。

Type I error (α) Reject true H₀, false alarm 第一类错误 拒绝真H₀,假警报
Type II error (β) Fail to reject false H₀, missed detection 第二类错误 未拒绝假H₀,漏报

9. Correlation and Regression | 相关与回归

Correlation measures the strength and direction of a linear relationship; regression models that relationship to make predictions. Don’t confuse ‘correlation’ with ‘causation’ – CAIE examiners love to test this.

相关衡量线性关系的强度和方向;回归对关系建模并用于预测。不要把 ‘相关’ 与 ‘因果’ 混淆——考官喜欢考这点。

Product moment correlation coefficient (PMCC) r: Ranges from -1 to +1. r = 1: perfect positive linear correlation; r = -1: perfect negative; r = 0: none. Formula: r = S_{xy} / √(S_{xx} S_{yy}). Think ‘r = relationship strength’.

积矩相关系数 r:范围从 -1 到 +1。r = 1:完全正线性相关;r = -1:完全负相关;r = 0:无线性相关。公式:r = S_{xy} / √(S_{xx} S_{yy})。记住 ‘r 代表关系强度’。

Spearman’s rank correlation rₛ: Used when data is ordinal or not normally distributed. Based on ranks, not raw values. rₛ = 1 – (6Σd²) / (n(n²-1)), where d is the difference in ranks.

斯皮尔曼等级相关系数 rₛ:用于顺序数据或非正态数据。基于排位而非原始值。rₛ = 1 – (6Σd²) / (n(n²-1)),其中 d 是排位差。

Regression line (least squares): y = a + bx, where b = S_{xy} / S_{xx} and a = ȳ – b x̄. The line minimises the sum of squared residuals. ‘Least squares’ literally finds the smallest sum of squared vertical gaps.

回归直线(最小二乘法):y = a + bx,其中 b = S_{xy} / S_{xx},a = ȳ – b x̄。该直线最小化残差平方和。’最小二乘’就是最小化垂直距离的平方和。

Residual = observed y − predicted y. A residual plot should show random scatter; patterns suggest a poor model fit. Memory: ‘Residual = what’s left over after the model does its job’.

残差 = 观测值 y – 预测值 y。残差图应呈随机散落;出现模式说明模型拟合不佳。记忆:’残差就是模型干完活剩下的部分’。


10. Key Graphical Representations | 关键图形表示

Graphs turn raw data into visual stories. In CAIE exams, you need to interpret histograms, cumulative frequency curves and box plots, as well as identify skewness and outliers.

图表把原始数据变成可视化故事。CAIE 考试要求你解读直方图、累积频率曲线和箱线图,并识别偏态和异常值。

Histogram: For continuous grouped data, area ∝ frequency. If class widths are unequal, use frequency density = frequency / class width. ‘Histogram bars kiss but don’t overlap’ (no gaps).

直方图:用于连续分组数据,面积代表频率。若组距不等,使用频率密度 = 频率 ÷ 组距。’直方图的柱子紧紧挨着没有缝隙’。

Cumulative frequency curve (ogive): Plots running total against upper class boundary. Use it to estimate median, quartiles, percentiles. The steeper the curve, the more data is concentrated.

累积频率曲线(折线图):将累积频数对组上限描点。用于估计中位数、四分位数、百分位数。曲线越陡,数据越集中。

Box plot (box-and-whisker): Displays min, Q₁, median, Q₃, max. Outliers are plotted as individual points. A box plot instantly shows skewness: if median is closer to Q₁, data is right-skewed. ‘Box = middle 50%, whiskers = tails’.

箱线图:显示最小值、Q₁、中位数、Q₃、最大值。异常值作为单独的点标出。箱线图能一眼看出偏态:若中位数靠近 Q₁,数据右偏。’箱子是中间50%,须子是尾巴’。

Stem-and-leaf diagram: Retains original data while sorting. The ‘stem’ is the leading digit(s), ‘leaf’ the trailing digit. Great for finding median and mode quickly.

茎叶图:在排序的同时保留原始数据。’茎’是前导数字,’叶’是后随数字。能快速找到中位数和众数。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading