📚 Statistics Terminology Quick Memorisation Guide | 统计学术语速记指南
Mastering statistical vocabulary is the first step towards confident data analysis in Pre-U Edexcel Statistics. This guide pairs each key term with a memory hack, so you can recall definitions as fluently as you calculate. Use these pairings to build a mental map of the subject before you dive into past papers.
掌握统计学术语是自信应对 Pre-U Edexcel 统计学数据分析的第一步。本指南为每个关键术语搭配了一条记忆技巧,让你能像计算一样流畅地回忆定义。在刷真题之前,用这些配对构建你的学科心智地图。
1. Measures of Central Tendency | 集中趋势的度量
The mean (arithmetic average x̄) is the sum of all observations divided by the number of observations. Think of it as the “balancing point” of a dataset — if values were weights on a seesaw, the mean would be the fulcrum. Use the mnemonic “Mean is the Middle of the Maths” because it involves dividing the total after summing.
均值(算术平均数 x̄)是所有观测值之和除以观测值的个数。可以把它想象成数据集的“平衡点”——如果把数值看作跷跷板上的重量,均值就是支点。使用助记符 “Mean is the Middle of the Maths”,因为它需要在求和后除以个数。
The median is the middle value when data are ordered. Picture a road median strip splitting a highway into two equal halves — the median splits the dataset into two halves with equal count. For an odd number of data points, it is the central value; for an even number, it is the average of the two central values.
中位数是数据排序后位于中间的值。想象公路中间隔离带把马路分成相等的两半——中位数将数据集分成数量相等的两部分。数据个数为奇数时,它是正中间的值;为偶数时,是中间两个数的平均值。
The mode is the most frequently occurring value. Connect it to the French word “mode” meaning fashion — the mode is the “most fashionable” value, the trend that appears most often. A dataset can have one mode (unimodal), two modes (bimodal), or more.
众数是出现频率最高的值。把它和法语词“mode”(时尚)联系起来——众数就是“最时髦”的数值,最常出现的潮流。数据集可以有一个众数(单峰)、两个众数(双峰)或更多。
2. Measures of Spread | 离散程度的测量
The range = maximum − minimum. It is the simplest measure of spread. Remember it as “how far the arms reach” — from the smallest to the largest value. However, range is sensitive to outliers, so it is only a rough indicator.
极差 = 最大值 − 最小值。这是最简单的离散程度度量。记住它是“双臂能伸多远”——从最小到最大的跨度。但极差对异常值敏感,所以只是一个粗略指标。
The interquartile range (IQR) = Q₃ − Q₁. It measures the spread of the middle 50% of data, making it robust to outliers. Think of “IQR” as “I Quit Relying on extremes” — it ignores the tails and focuses on the core.
四分位距 (IQR) = Q₃ − Q₁。它度量中间 50% 数据的离散程度,因此不受异常值影响。把 “IQR” 想成 “I Quit Relying on extremes”——它忽略尾部,聚焦核心。
Variance σ² (population) or s² (sample) is the average of the squared deviations from the mean. Formula: σ² = Σ(x − μ)² / N for population; s² = Σ(x − x̄)²/(n−1) for sample. To remember why we square deviations, say “Squares Stop Sign problems” — squaring eliminates negative signs, treating positive and negative deviations equally.
方差 σ²(总体)或 s²(样本)是各数据与均值之差的平方的平均值。总体公式:σ² = Σ(x − μ)² / N;样本公式:s² = Σ(x − x̄)²/(n−1)。要记住为何要平方差值,就说“平方消除符号烦恼”——平方去掉了负号,使正负偏差被同等对待。
Standard deviation σ or s is the square root of variance. It rescales the spread to the original units. The symbol σ looks like a rounded ‘s’, and both start with “s”: “Standard deviation = Square root of Σ squared differences”.
标准差 σ 或 s 是方差的平方根,它将离散程度还原为原始单位。符号 σ 看起来像圆润的 ‘s’,两者都以 “s” 开头:“Standard deviation = Square root of Σ squared differences”。
3. Data Types and Classification | 数据类型与分类
Qualitative (categorical) data describe qualities or categories, e.g., eye colour, car brand. They are non-numerical. Think “Qualitative = Quality” — about characteristics, not quantities.
定性(分类)数据描述品质或类别,如眼睛颜色、汽车品牌。它们是非数值的。记住“Qualitative = Quality”——关于特征,而非数量。
Quantitative data are numerical, measuring quantities. Split these into discrete and continuous. “Quantitative = Quantity” — you can count or measure it.
定量数据是数值型,测量数量。细分为离散和连续。“Quantitative = Quantity”——可以计数或测量。
Discrete data can only take specific, separate values (often integers). Examples: number of students, dice rolls. The word “discrete” sounds like “discreet” (separate). Think of distinct steps like stairs — you can only stand on exact steps, not between them.
离散数据只能取特定、分离的值(通常是整数)。例如:学生人数、掷骰子点数。“离散”就像单独的台阶——你只能站在精确的台阶上,不能站在中间。
Continuous data can take any value within a range. Examples: height, time, temperature. Imagine a continuous line without breaks — the word “continuous” itself suggests unbroken flow.
连续数据可以在一个范围内取任意值。例如:身高、时间、温度。想象一条没有间断的连续线——“连续”这个词本身就暗示了不间断的流动。
4. Probability Essentials | 概率基础
A random experiment is a process leading to two or more possible outcomes, where the result cannot be predicted with certainty. Key phrase: “rAnDom = A Deterministic Result? No!” to emphasise unpredictability.
随机试验是产生两个或多个可能结果的过程,结果无法确定预测。助记:“随机 = 随缘机选”——强调不可预测性。
The sample space (S) is the set of all possible outcomes. Visualise it as a “space” that holds every outcome. “S = Space of Sample”. For two dice, S has 36 ordered pairs.
样本空间 (S) 是所有可能结果的集合。把它想象成一个容纳所有结局的“空间”。“S = 样本的空间”。掷两粒骰子时,S 有 36 个有序数对。
An event (E) is a subset of the sample space — a collection of outcomes of interest. Link event to “happening”: an event is “E-vent” — something that could happen.
事件 (E) 是样本空间的子集——一系列我们关心的结果。把事件和“发生”联系:一个事件就是可能“发生”的事。
Mutually exclusive events cannot occur simultaneously. P(A ∩ B) = 0. Think of “mutual exclusion” — like a light switch cannot be both on and off at the same time.
互斥事件不能同时发生。P(A ∩ B) = 0。想想“互相排斥”——就像电灯开关不能同时开与关。
Independent events have no influence on each other’s probabilities: P(A ∩ B) = P(A) × P(B). The word “independent” hints “not dependent” — the outcome of A does not change the probability of B.
独立事件互不影响概率:P(A ∩ B) = P(A) × P(B)。单词“独立”暗示“不依赖”——A 的结果不改变 B 的概率。
5. Probability Distributions | 概率分布
A probability distribution lists all possible values of a discrete random variable along with their probabilities. Visualise it as a “recipe” spreading total probability 1 across outcomes. The sum of all probabilities must equal 1: ΣP(X = x) = 1.
概率分布列出离散随机变量的所有可能取值及其概率。把它看作将一个总概率 1 分布到各个结果上的“配方”。所有概率之和必须等于 1:ΣP(X = x) = 1。
The binomial distribution B(n, p) models the number of successes in n independent trials, each with probability p of success. Think “Bi-nomial = two names (success/failure)”. Remember the assumptions using the mnemonic BINS: Binary, Independent, Number fixed, Same probability.
二项分布 B(n, p) 模型化 n 次独立试验中成功的次数,每次成功概率为 p。想想“二项 = 两种名称(成功/失败)”。用 BINS 记假设条件:二元结果、独立、试验次数固定、成功概率相同。
The normal distribution N(μ, σ²) is a continuous bell‑shaped curve, symmetric about the mean μ. “Normal” because many natural phenomena follow it. Key property: about 68% of data within 1σ of μ; 95% within 2σ; 99.7% within 3σ (the empirical rule). Use the “68‑95‑99.7” rhythm to recall it.
正态分布 N(μ, σ²) 是连续钟形曲线,关于均值 μ 对称。“正态”因为许多自然现象服从这个分布。关键性质:约 68% 数据在 μ±1σ 内;95% 在 μ±2σ 内;99.7% 在 μ±3σ 内(经验法则)。用“68‑95‑99.7”节奏记忆。
6. Sampling Methods | 抽样方法
Simple random sampling gives every member of the population an equal chance of selection, e.g., using a random number generator. Think of a lottery draw — “everyone has a ticket, one winner drawn by chance”. It avoids bias but can be impractical for large populations.
简单随机抽样让总体中每个成员被选中的机会相等,例如使用随机数生成器。想象抽奖——“人人都有一票,随机抽出胜者”。它避免偏差,但对大总体可能不切实际。
Stratified sampling divides the population into distinct groups (strata) and then takes a random sample from each stratum in proportion to its size. “Stratified” sounds like “strata” (layers). This ensures representation of all subgroups. Example: sampling 50% male, 50% female from a school where gender is equally split.
分层抽样先将总体分为不同的组(层),然后按比例从每一层中随机抽取样本。“分层”听起来像“层次”。这确保所有子群体都有代表。例:从男女比例各半的学校按 50:50 抽样。
Systematic sampling selects every kth individual from a list after a random start. “Systematic” suggests “system” — a regular pattern. Choose k = population size ÷ sample size, then pick a random starting point between 1 and k. It is easy to carry out but can introduce periodicity bias.
系统抽样从名单中随机起点后每隔 k 个抽取一个。“系统”暗示“固定模式”。计算 k = 总体大小 ÷ 样本大小,然后在 1 到 k 之间随机选起点。操作简便,但可能引入周期性偏差。
7. Correlation and Regression | 相关与回归
Correlation measures the strength and direction of a linear relationship between two variables. The product‑moment correlation coefficient r ranges from −1 to +1. Mnemonic: “r = relationship”. r close to +1 means strong positive correlation (as one goes up, the other tends to go up); r close to −1 means strong negative; r ≈ 0 means no linear correlation.
相关衡量两个变量之间线性关系的强度和方向。积矩相关系数 r 范围为 −1 到 +1。助记:“r = relationship(关系)”。r 接近 +1 表示强正相关(一个增加,另一个也增加);接近 −1 表示强负相关;r ≈ 0 表示无线行相关。
Regression aims to model the relationship with an equation, typically the least squares regression line y = a + bx. It minimises the sum of squared vertical distances from the data points to the line. Link “regression” to “prediction” — we can predict y for a given x. The line always passes through (x̄, ȳ).
回归旨在用方程建模关系,通常是最小二乘回归线 y = a + bx。它最小化数据点到直线的垂直距离平方和。把“回归”与“预测”联系起来——我们可以根据 x 预测 y。回归线总经过点 (x̄, ȳ)。
Interpolation is estimating within the range of observed data, which is reliable. Extrapolation is predicting outside the observed range, often unreliable. Think “inter- = internal, extra- = external”. Exercise: “Interpolation is inside your data’s interval; extrapolation is outside — risky!”
内插是在观测数据范围内估计,较为可靠。外推是预测观测范围之外的值,通常不可靠。记词根:“内 = 内部,外 = 外部”。口诀:“内插插在数据区间内;外推推到外面——冒险!”
8. Hypothesis Testing Key Terms | 假设检验关键术语
The null hypothesis H₀ is the default assumption, usually stating “no effect” or “no difference”. “Null” means zero, nothing — remember “H₀ = H‑zero, zero effect”. It is the hypothesis we assume to be true unless evidence compels its rejection.
零假设 H₀ 是默认假设,通常称“无效”或“无差异”。“Null”意为零、无——记住 “H₀ = H‑zero,零效果”。除非有充分证据,否则假定它为真。
The alternative hypothesis H₁ (or Hₐ) is what we believe might be true if H₀ is rejected. It can be one‑tailed (directional: “greater than” or “less than”) or two‑tailed (“not equal to”). Mnemonic: “When H₀ falls, H₁ rises.”
备择假设 H₁ (或 Hₐ) 是如果拒绝 H₀,我们可能相信为真的假设。可以是单尾(方向性:“大于”或“小于”)或双尾(“不等于”)。助记:“当 H₀ 倒下,H₁ 站起来。”
The significance level α is the probability of rejecting H₀ when it is actually true (Type I error). Commonly 0.05 (5%). “Alpha = Acceptable false‑alarm probability” — the risk you are willing to take.
显著性水平 α 是当 H₀ 为真时拒绝 H₀ 的概率(第一类错误)。通常为 0.05(5%)。“Alpha = Acceptable false‑alarm probability”——你愿意承担的误报风险。
The p‑value is the probability, assuming H₀ is true, of obtaining a test statistic at least as extreme as the one observed. If p ≤ α, reject H₀. Think “p‑value = probability value; p low, null must go.”
p 值 是在 H₀ 为真的前提下,观察到当前统计量或更极端值的概率。若 p ≤ α,拒绝 H₀。记住:“p 值 = 概率值;p 值小,零假设跑不了。”
9. Type I and Type II Errors | 第一类与第二类错误
A Type I error occurs when we reject a true null hypothesis — a false positive. Probability = α. Mnemonic: “Type I = falsely accuse an Innocent person” (beginning with ‘I’). You see an effect that isn’t there.
第一类错误发生在拒绝真实的零假设时——假阳性。概率 = α。助记:“第一类 = 错误地指控一个无辜 (Innocent) 的人”。你看到了不存在的效应。
A Type II error occurs when we fail to reject a false null hypothesis — a false negative. Probability = β. “Type II = Two little evidence leads to missing a real effect.” Think of a guilty person getting away.
第二类错误发生在未能拒绝错误的零假设时——假阴性。概率 = β。助记:“第二类 = 证据太少 (Two little evidence) 导致错失真效应”。想象罪犯逍遥法外。
The power of a test is 1 − β, the probability of correctly rejecting a false H₀. Larger sample size increases power. “Powerful test spots the difference.”
检验的功效 为 1 − β,即正确拒绝错误 H₀ 的概率。样本量越大,功效越大。“功效强的检验能发现问题。”
10. Confidence Intervals and Estimation | 置信区间与估计
A point estimate is a single value used to estimate a population parameter, e.g., sample mean x̄ estimates population mean μ. It is precise but gives no sense of uncertainty.
点估计是用于估计总体参数的单一值,例如样本均值 x̄ 估计总体均值 μ。它精确,但不反映不确定性。
A confidence interval gives a range of plausible values for the parameter. A 95% confidence interval means that if we repeated the sampling many times, 95% of the intervals would contain the true parameter. It is not a 95% chance that the parameter lies in a particular interval — the parameter is fixed; the interval is what varies.
置信区间给出参数的一个合理取值范围。95% 置信区间意味着,如果重复抽样多次,95% 的区间会包含真实参数。不是说该参数有 95% 的可能性落在某一特定区间内——参数是固定的,变化的是区间。
Formulas for common confidence intervals: for population mean μ with known σ: x̄ ± z × (σ/√n). For proportion p: p̂ ± z × √[p̂(1 − p̂)/n]. The margin of error is the half‑width of the interval.
常见置信区间公式:已知总体标准差 σ 的均值 μ 置信区间:x̄ ± z × (σ/√n)。比例 p 的置信区间:p̂ ± z × √[p̂(1 − p̂)/n]。误差范围是区间的一半宽度。
11. Experimental Design Terminology | 实验设计术语
Control group serves as a baseline, receiving no treatment or a placebo. This allows comparison against the treatment group. Think: “Control the conditions to isolate the effect.”
对照组作为基线,不接受处理或接受安慰剂。这样可以与处理组进行比较。记住:“控制条件以隔离效应。”
Randomisation assigns subjects to groups by chance, balancing out confounding variables. It’s the “great equaliser” — any lurking variable should be equally spread between groups. “Random = even distribution of the unknown.”
随机化通过随机方式分配受试者到各组,平衡混杂变量。它是“伟大的均衡器”——任何潜在变量应平等分配到各组。“随机 = 未知因素均匀分布。”
Replication means repeating the experiment or using sufficiently large sample sizes to obtain reliable results. “Replicate to validate” — one study is rarely enough.
重复指重复实验或使用足够大的样本来获得可靠结果。“重复以求证”——单一研究几乎不够。
Blocking is the arrangement of experimental units into similar groups (blocks) to reduce variation, e.g., grouping by age or gender before randomisation. Think of building blocks: arranging similar pieces together first.
区组是将实验单元安排到相似的组(区组)中以减少变异,例如随机化前按年龄或性别分组。想象搭积木:先把相似的积木归在一起。
12. Additional Rapid-Fire Terms | 附加速记术语
Parameter: a numerical summary of a population (μ, σ, p). Statistic: a numerical summary of a sample (x̄, s, p̂). “P for Parameter, P for Population; S for Statistic, S for Sample.”
参数:总体的数值概括(μ, σ, p)。统计量:样本的数值概括(x̄, s, p̂)。“参(P)对应总体(Population),统(S)对应样本(Sample)。”
Bias: systematic deviation of an estimate from the true value. Sampling bias, measurement bias, etc. “Bias bends results away from truth.”
偏差:估计值系统性地偏离真实值。抽样偏差、测量偏差等。“偏差让结果偏离真相。”
Outlier: an observation that lies far from the bulk of the data. Use 1.5 × IQR rule to identify: below Q₁ − 1.5×IQR or above Q₃ + 1.5×IQR. “Out‑liar — it lies outside the crowd.”
异常值:远离数据主体的观测值。用 1.5 × IQR 法则识别:低于 Q₁ − 1.5×IQR 或高于 Q₃ + 1.5×IQR。“异常值——在群体之外说谎”。
Central Limit Theorem (CLT): For large sample sizes, the sampling distribution of the sample mean is approximately normal, regardless of the population distribution. This is the engine of inference. “CLT = Counts Large Transcends distribution.” Mean of sample means = μ; standard error = σ/√n.
中心极限定理 (CLT):对于大样本,样本均值的抽样分布近似服从正态分布,无论总体分布如何。这是推断的引擎。“CLT = 大样本超越分布的界限。”样本均值的均值 = μ;标准误差 = σ/√n。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导