📚 Year 11 CCEA Statistics: Glossary Memorisation Guide | Year 11 CCEA 统计:词汇术语速记指南
Mastering the language of statistics is the first step towards confidently tackling any CCEA Year 11 exam question. This guide breaks down the essential vocabulary into logical clusters, pairing clear English explanations with Chinese translations and memory prompts. Use it alongside your revision to turn unfamiliar jargon into second nature.
掌握统计学的语言是自信应对CCEA 11年级考试的第一步。本指南将核心词汇按逻辑分组,配以清晰的英文解释、中文翻译和记忆提示。请结合复习使用,让陌生的术语变成你的本能。
1. Types of Data | 数据类型
Data can be qualitative (categorical, non-numerical) or quantitative (numerical). Quantitative data splits further into discrete data (countable, e.g. number of students) and continuous data (measurable, e.g. height). Understanding this hierarchy prevents mistakes when choosing charts and calculations.
数据可以是定性数据(分类的,非数值的)或定量数据(数值的)。定量数据又分为离散数据(可数的,如学生人数)和连续数据(可测量的,如身高)。理解这个层次可以避免在选择图表和计算时出错。
Qualitative data often describes attributes like colour, gender, or favourite subject. Quantitative data answers ‘how much’ or ‘how many’.
定性数据通常描述颜色、性别或最喜爱科目等属性。定量数据用来回答“多少”或“多大”。
A quick mnemonic: ‘Quality’ words are Qualitative; ‘Quantity’ numbers are Quantitative.
快速记忆法:与“质量/属性”相关的是定性数据;与“数量”相关的是定量数据。
2. Sampling Methods | 抽样方法
A population is the entire group you are interested in, while a sample is a subset used to draw conclusions. A sampling frame is a list of all individuals in the population, like a register.
总体是你感兴趣的整个群体,样本是用于得出结论的一个子集。抽样框是总体中所有个体的清单,比如一份名册。
Random sampling gives every member an equal chance of being selected, reducing bias. Stratified sampling divides the population into groups (strata) and samples proportionally from each, ensuring fair representation.
随机抽样让每个成员都有均等被选中的机会,从而减少偏差。分层抽样将总体分成若干组(层),然后按比例从每层抽样,确保公平代表性。
Systematic sampling picks every nth individual from a list. Quota sampling is non-random; interviewers fill quotas for certain characteristics, which can introduce bias but is cheap and quick.
系统抽样在清单上每隔n个个体抽取一个。配额抽样是非概率抽样;访问员按特定特征配额进行抽样,这容易引入偏差,但成本低、速度快。
Avoid convenience sampling for serious studies – it selects the easiest subjects and often produces unreliable results.
认真研究中请避免便利抽样——它选取最容易获得的样本,常产生不可靠的结果。
3. Charts and Diagrams | 图表
Bar charts display categorical data with gaps between bars; height shows frequency. Pie charts show proportions of a whole in slices, good for visual comparisons of a few categories.
条形图用带间隙的条形展示分类数据,条形高度表示频数。饼图用扇形展示整体的各个部分比例,适合对少量类别进行直观比较。
Histograms are for grouped continuous data: bars touch, and area represents frequency. If class widths are unequal, use frequency density = frequency ÷ class width.
直方图用于分组连续数据:条形紧贴,面积代表频数。若组距不相等,要用频率密度 = 频数 ÷ 组距。
Cumulative frequency diagrams (ogives) plot running totals against upper class boundaries, allowing you to find medians, quartiles, and percentiles smoothly.
累积频数图(拱形图)将累计总数绘制在上组界上,可以用它平滑地找出中位数、四分位数和百分位数。
Box plots (box-and-whisker) summarise data using minimum, lower quartile, median, upper quartile, and maximum – ideal for comparing distributions.
箱线图(盒须图)用最小值、下四分位数、中位数、上四分位数和最大值概括数据,非常适合对比分布。
4. Measures of Central Tendency | 集中趋势度量
The mean (x̅) is the arithmetic average: sum all values and divide by the number of values. It uses all data but is affected by extreme outliers.
平均数(x̅)是算术平均值:将所有值相加再除以值的个数。它使用了所有数据,但会受极端异常值影响。
The median is the middle value when data is ordered. It splits the distribution in half and is resistant to outliers. If n is even, average the two middle numbers.
中位数是将数据排序后的中间值。它将分布一分为二,而且不受异常值影响。如果n为偶数,则取中间两个数的平均值。
The mode is the most frequent value. A data set can have no mode, one mode (unimodal), or several modes (multimodal). Modes work with qualitative data.
众数是出现频率最高的值。一组数据可能没有众数、有一个众数(单峰)或有多个众数(多峰)。众数可用于定性数据。
In a symmetrical distribution, mean ≈ median ≈ mode. In a positively skewed set (tail to the right), mean > median > mode.
在对称分布中,平均数 ≈ 中位数 ≈ 众数。在正偏态分布中(尾部向右),平均数 > 中位数 > 众数。
5. Measures of Spread | 离散度度量
Range = maximum value – minimum value. It is simple but ignores the shape of the data and is easily distorted by outliers.
极差 = 最大值 − 最小值。它很简单,但忽略了数据的分布形态,且容易被异常值扭曲。
Interquartile range (IQR) = upper quartile (Q₃) – lower quartile (Q₁). It describes the spread of the middle 50% of data, making it robust against extremes.
四分位距 (IQR) = 上四分位数 (Q₃) − 下四分位数 (Q₁)。它描述中间50%数据的分散程度,对极端值具有稳健性。
Variance and standard deviation measure how far values typically deviate from the mean. Standard deviation s = √(Σ(x – x̅)² / (n – 1)) for a sample. The squared version is variance.
方差和标准差衡量数值通常偏离平均数的程度。样本标准差 s = √(Σ(x – x̅)² / (n – 1))。其平方就是方差。
Small standard deviation means data clusters tightly around the mean; large standard deviation indicates wide dispersion.
小标准差意味着数据紧密聚集在平均数周围;大标准差表示离散程度高。
6. Probability Basics | 概率基础
Probability P(A) = number of favourable outcomes / total number of equally likely outcomes. It always lies between 0 (impossible) and 1 (certain).
概率 P(A) = 有利结果数 / 所有等可能结果总数。它总是介于0(不可能)到1(必然)之间。
The complement of event A, written A’, means ‘not A’. P(A’) = 1 – P(A). This is extremely useful when calculating ‘at least one’ problems.
事件A的补集,写作A’,表示“非A”。P(A’) = 1 – P(A)。在计算“至少一次”类问题时非常有用。
Mutually exclusive events cannot happen at the same time. The addition rule for them: P(A or B) = P(A) + P(B). Independent events do not influence each other; the multiplication rule: P(A and B) = P(A) × P(B).
互斥事件不能同时发生。其加法法则为:P(A或B) = P(A) + P(B)。独立事件互不影响;乘法法则为:P(A且B) = P(A) × P(B)。
Expected frequency = probability × number of trials. In a fair coin tossed 200 times, expect 100 heads.
期望频数 = 概率 × 试验次数。抛掷200次公平硬币,期望出现100次正面。
7. Venn Diagrams and Set Notation | 维恩图与集合符号
Venn diagrams use overlapping circles to represent sets. The universal set ε contains everything under consideration. Intersection A ∩ B is ‘both A and B’; union A ∪ B is ‘A or B or both’.
维恩图用重叠的圆圈表示集合。全集 ε 包含考虑的所有元素。交集 A ∩ B 表示“A且B”;并集 A ∪ B 表示“A或B或两者”。
The notation n(A) gives the number of elements in set A. n(A ∪ B) = n(A) + n(B) – n(A ∩ B) prevents double-counting.
符号 n(A) 表示集合A中的元素个数。n(A ∪ B) = n(A) + n(B) – n(A ∩ B) 可避免重复计算。
Conditional probability P(A|B) is the probability of A given that B has occurred. Formula: P(A|B) = P(A ∩ B) / P(B). Always restrict the sample space to B.
条件概率 P(A|B) 是在B发生的条件下A发生的概率。公式:P(A|B) = P(A ∩ B) / P(B)。始终将样本空间限定为B。
Two events are independent if P(A ∩ B) = P(A) × P(B), or equivalently P(A|B) = P(A). Check this in exam questions involving tree diagrams.
若 P(A ∩ B) = P(A) × P(B),或等价的 P(A|B) = P(A),则两事件独立。在涉及树形图的考题中请检查这一点。
8. Correlation and Regression | 相关与回归
Scatter graphs reveal the relationship between two variables. Correlation describes the strength and direction: positive (as x rises, y rises), negative (as x rises, y falls), or zero (no pattern).
散点图揭示两个变量之间的关系。相关性描述强度和方向:正相关(x增大,y增大)、负相关(x增大,y减小)或无相关(无章可循)。
Causal relationship means change in one variable directly causes change in another. Correlation does not imply causation – a third lurking variable may explain the link.
因果关系指一个变量的变化直接引起另一个变量的变化。相关性并不意味着因果关系——可能存在第三个潜在变量导致两者关联。
The line of best fit (regression line) is drawn through the scatter points to model the trend. It should pass close to as many points as possible, with roughly equal numbers above and below.
最佳拟合线(回归线)穿过散点以模拟趋势。它应尽可能靠近较多的点,上方和下方的点数应大致相等。
Use the line for interpolation (predicting within the data range) but be cautious with extrapolation (predicting beyond the range) – trends may not continue.
用这条线做内插预测(在数据范围内预测)是可行的,但外推(超出范围预测)要谨慎——趋势可能不会持续。
9. Time Series and Moving Averages | 时间序列与移动平均
A time series records data at regular intervals (e.g. quarterly sales). It usually contains a trend and seasonal variation. Plotting helps spot patterns over time.
时间序列以固定间隔记录数据(如季度销售)。它通常包含趋势和季节性变动。绘图有助于发现随时间变化的规律。
Moving averages smooth out short-term fluctuations to reveal the underlying trend. To calculate, find the mean of consecutive time periods. For quarterly data, a 4-point moving average then centre it if necessary.
移动平均可以平滑短期波动以揭示潜在趋势。计算方法是取连续时段的平均值。对于季度数据,可计算4点移动平均并酌情居中。
Seasonal variation = actual value – trend (additive model). A positive seasonal effect means the value is above trend. Averaging seasonal effects across years gives reliable forecasts.
季节性变动 = 实际值 − 趋势(加法模型)。正的季节性效应意味着数值高于趋势。对多年的季节性效应取平均可得出可靠的预测。
Forecasting: predicted value = trend projection + average seasonal effect for that season. State assumptions clearly in exam answers.
预测:预测值 = 趋势预测值 + 该季节的平均季节性效应。在答题时请清晰说明假设。
10. Index Numbers | 指数
Index numbers compare values relative to a base period. The base year index is usually 100. A price index of 120 means a 20% increase from the base year.
指数将数值与基期进行比较。基年指数通常设为100。价格指数为120表示相比基年上涨了20%。
Simple index = (current value / base value) × 100. Composite indices like the Consumer Price Index (CPI) weigh multiple items according to their importance.
简单指数 = (当前值 / 基期值) × 100。综合指数如消费者价格指数(CPI)根据各项目的重要性赋予权重。
Weighted aggregate index = Σ(price relative × weight) / Σ(weights). Weights reflect consumption patterns. Always check the base period when interpreting indices.
加权综合指数 = Σ(价格相对数 × 权重) / Σ(权重)。权重反映消费模式。解读指数时务必确认基期。
Use index numbers to deflate nominal values to real values: real value = (nominal value / price index) × 100. This removes inflation effects.
使用指数可将名义值缩减为实际值:实际值 = (名义值 / 价格指数) × 100。这消除了通货膨胀的影响。
11. Misleading Statistics and Critiquing Data | 误导性统计与数据批判
Statistics can be misrepresented through truncated axes, misleading scales, or 3D effects that distort proportions. Always read axis labels and check the origin.
统计可能通过截断坐标轴、误导性比例或扭曲比例的3D效果来误传信息。请务必阅读坐标轴标签并检查原点。
Selective reporting (cherry-picking) only shows favourable data. Small sample sizes lead to unreliable conclusions; look for sample size n in reports.
选择性报告(摘樱桃)只展示有利的数据。小样本量导致结论不可靠;请留意报告中样本量n。
Questionable surveys: leading questions, non-response bias, or sampling from a self-selected group (e.g. online polls) undermine validity. Good statistics demand transparency.
有问题的调查:诱导性问题、无应答偏差或从自选群体抽样(如网络投票)都会损害有效性。好的统计需要透明度。
Critiquing a study involves commenting on sampling method, sample size, potential bias, data representation, and whether conclusions are justified by the evidence.
评论一项研究要评价抽样方法、样本量、潜在偏差、数据呈现方式以及结论是否得到证据支持。
12. Key Formulae Recap | 关键公式回顾
Keep these core expressions handy. Mean x̅ = Σx / n; Median position = (n + 1)/2; IQR = Q₃ – Q₁.
记住这些核心表达式:平均数 x̅ = Σx / n;中位数位置 = (n + 1)/2;IQR = Q₃ – Q₁。
Standard deviation s = √[Σ(x – x̅)² / (n – 1)] for a sample; alternatively s = √[(Σx² – (Σx)²/n) / (n – 1)].
样本标准差 s = √[Σ(x – x̅)² / (n – 1)],或者 s = √[(Σx² – (Σx)²/n) / (n – 1)]。
Probability: P(A|B) = P(A ∩ B) / P(B); P(A ∪ B) = P(A) + P(B) – P(A ∩ B); for independent events P(A ∩ B) = P(A)P(B).
概率:P(A|B) = P(A ∩ B) / P(B);P(A ∪ B) = P(A) + P(B) – P(A ∩ B);独立事件 P(A ∩ B) = P(A)P(B)。
Frequency density = frequency / class width. Index = (value / base value) × 100. These appear throughout CCEA papers, so practise using them with different scenarios.
频率密度 = 频数 / 组距。指数 = (数值 / 基值) × 100。这些公式在CCEA试卷中随处可见,请在不同情境下反复练习。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导