📚 Year 11 CAIE Statistics: Terminology Quick-Memorisation Guide | Year 11 CAIE 统计:词汇术语速记指南
Mastering statistical vocabulary is the first step to excelling in your CAIE Statistics exam. This guide pairs essential terms with simple explanations, vivid associations, and cross-language cues to help you internalise concepts quickly and confidently.
掌握统计词汇是你在 CAIE 统计考试中脱颖而出的第一步。本指南将核心术语与简洁解释、生动联想和双语提示相结合,助你快速、扎实地内化每一个概念。
1. What is Statistics? | 什么是统计学?
Statistics is the science of collecting, organising, analysing, and interpreting numerical data to make informed decisions.
统计学是一门收集、整理、分析和解释数值数据,以便做出明智决策的科学。
Think of it as a detective toolkit: you gather clues (data), spot patterns (analysis), and draw conclusions (inference) without jumping to false assumptions.
把它想成一套侦探工具箱:你收集线索(数据)、发现规律(分析)、得出结论(推断),而不会草率地做出错误假设。
The two main branches are descriptive statistics (summarising data) and inferential statistics (drawing conclusions from samples).
两大分支是描述统计学(概括数据)和推断统计学(从样本中得出结论)。
2. Types of Data | 数据类型
Qualitative data describes qualities or categories (e.g., eye colour, brand names) and is non-numerical.
定性数据描述性质或类别(如眼睛颜色、品牌名称),是非数值的。
Quantitative data is numerical and can be discrete (countable, like number of students) or continuous (measurable, like height).
定量数据是数值型的,可以是离散的(可数的,如学生人数)或连续的(可测量的,如身高)。
Memory hook: ‘Quality-No-numbers’ vs ‘Quantity-Numbers’. Also recall that discrete data often comes from counting, while continuous data comes from measuring.
记忆钩子:「定性没数字」对「定量有数字」。同时记住离散数据通常来自计数,连续数据来自测量。
3. Measures of Central Tendency | 集中趋势的度量
The mean (x̄) is the arithmetic average: sum all values and divide by the number of values. It uses every data point but is sensitive to outliers.
平均数(x̄)是算术平均值:将所有数值相加再除以数值的个数。它使用了每一个数据点,但对异常值很敏感。
The median is the middle value when data is ordered. It is robust to outliers and often preferred for skewed distributions.
中位数是数据排序后位于中间的值。它对异常值不敏感,通常适用于偏态分布。
The mode is the most frequently occurring value. A data set can have no mode, one mode, or multiple modes.
众数是出现次数最多的值。一组数据可以没有众数、有一个众数或多个众数。
Mnemonics: ‘Mean is mean to outliers’, ‘Median stays in the middle’, ‘Mode is the most popular’.
助记口诀:「平均数对异常值很‘刻薄’」、「中位数永远居中」、「众数是人气王」。
4. Measures of Spread | 离散程度的度量
Range = highest value – lowest value. It gives a rough sense of spread but ignores the distribution’s shape.
极差 = 最大值 – 最小值。它给出数据分散的粗略感觉,但忽略了分布形态。
Interquartile range (IQR) = upper quartile (Q₃) – lower quartile (Q₁). It focuses on the middle 50% and resists outliers.
四分位距 (IQR) = 上四分位数 (Q₃) – 下四分位数 (Q₁)。它专注于中间 50% 的数据,并能抵抗异常值。
Variance (σ²) and standard deviation (σ) measure average distance of each data point from the mean. A small σ means data clumps near the mean; a large σ means it is widely spread.
方差 (σ²) 和标准差 (σ) 衡量每个数据点与平均数之间的平均距离。σ 小意味着数据聚集在均值附近;σ 大意味着数据分布广泛。
Quick trick: IQR is like the ‘core’ of your data—ignore the extremes. Standard deviation is the ‘typical’ distance from the centre.
快速技巧:IQR 就像数据的「核心」——忽略两端的极端值。标准差就是距中心的「典型」距离。
5. Quartiles and Box Plots | 四分位数与箱线图
Lower quartile (Q₁) is the median of the lower half of data; upper quartile (Q₃) is the median of the upper half.
下四分位数 (Q₁) 是数据下半部分的中位数;上四分位数 (Q₃) 是上半部分的中位数。
A box-and-whisker plot displays the five-number summary: minimum, Q₁, median, Q₃, maximum. The box spans the IQR.
箱线图展示了五数汇总:最小值、Q₁、中位数、Q₃、最大值。箱体覆盖了 IQR。
Outliers are usually shown as separate dots beyond the whiskers, typically defined as points more than 1.5 × IQR beyond Q₁ or Q₃.
异常值通常显示在须线之外的独立圆点上,一般定义为超出 Q₁ 或 Q₃ 超过 1.5 × IQR 的点。
Visualise the box as the ‘secure middle ground’ and whiskers as ‘cautious arms’ reaching out to the extreme but not outlier values.
将箱体想象成「安全的中间地带」,而须线则是向极值(但不含异常值)伸出的「谨慎的手臂」。
6. Cumulative Frequency | 累积频数
Cumulative frequency is the running total of frequencies up to a certain boundary. It helps estimate medians, quartiles, and percentiles from grouped data.
累积频数是截至某一界限的频数累计总和。它有助于从分组数据中估算中位数、四分位数和百分位数。
A cumulative frequency curve (or ogive) plots cumulative frequency against the upper class boundary. The steeper the curve, the higher the frequency density in that interval.
累积频数曲线(累积频数图)将累积频数与上组界绘图。曲线越陡峭,该区间的频率密度越高。
To find the median from the curve, go to half the total frequency on the vertical axis, draw across to the curve, then drop down to read the value.
要从曲线上找中位数,先在纵轴上找到总频数的一半,水平画到曲线,再垂直向下读出对应值。
Remember: ‘Cumulative’ means ‘adding as you go’. The curve never decreases.
记住:「累积」就是「边走边加」。该曲线永远不会下降。
7. Histograms and Frequency Density | 直方图与频率密度
A histogram displays grouped continuous data with bars that touch, where the area of each bar represents frequency, not just its height.
直方图展示分组的连续数据,条形之间相互接触,每个条形的面积(而非仅高度)表示频数。
Frequency density = frequency ÷ class width. This is crucial when class intervals are unequal. Always label the vertical axis as frequency density.
频率密度 = 频数 ÷ 组距。当组距不相等时,这一点至关重要。纵轴必须标记为频率密度。
Common pitfall: confusing a histogram with a bar chart. Bar charts have gaps and are for categorical data; histograms have no gaps and are for continuous data.
常见误区:混淆直方图和条形图。条形图之间有间隙,适用于分类数据;直方图无间隙,适用于连续数据。
Image cue: ‘Histograms are solid blocks, bar charts are separate towers.’
图形提示:「直方图是连体块,条形图是独立塔」。
8. Scatter Diagrams and Correlation | 散点图与相关性
A scatter diagram plots pairs of numerical data to show relationships between two variables.
散点图将成对的数值数据绘制出来,显示两个变量之间的关系。
Correlation describes the direction and strength of a linear relationship. Positive correlation: as x increases, y tends to increase. Negative correlation: as x increases, y tends to decrease.
相关性描述线性关系的方向和强度。正相关:x 增加时 y 趋于增加。负相关:x 增加时 y 趋于减少。
Correlation does not imply causation. A strong correlation could be accidental or influenced by a lurking third variable.
相关性并不意味着因果关系。强相关可能是偶然的,也可能受到隐藏的第三变量的影响。
Line of best fit (trend line) can be drawn by eye, passing as close to as many points as possible, with roughly equal points above and below the line.
最佳拟合线(趋势线)可以通过目测绘制,穿过尽可能多的点,并让线上和下两侧的点的数量大致相等。
9. Probability Fundamentals | 概率基础
Probability measures the chance of an event occurring, ranging from 0 (impossible) to 1 (certain).
概率度量一个事件发生的可能性,范围从 0(不可能)到 1(必然)。
For equally likely outcomes, P(event) = number of favourable outcomes ÷ total number of possible outcomes.
对于等可能的结果,P(事件) = 有利结果的数量 ÷ 可能结果的总数。
The sample space is the set of all possible outcomes. The complement of event A is ‘not A’, and P(not A) = 1 – P(A).
样本空间是所有可能结果的集合。事件 A 的补集是「非 A」,且 P(非 A) = 1 – P(A)。
Mutually exclusive events cannot happen together. Independent events are those where the outcome of one does not affect the probability of the other.
互斥事件不能同时发生。独立事件是指一个事件的结果不影响另一个事件发生的概率。
Memory aid: ‘Exclusive means exclude each other; independent means I don’t care what you did.’
记忆辅助:「互斥就是相互排斥;独立就是我不在乎你做了什么。」
10. Expected Value and Risk | 期望值与风险
The expected value of a random variable is the long-run average outcome if the experiment is repeated many times. It is calculated as Σ [x · P(x)].
随机变量的期望值是如果试验重复很多次时的长期平均结果。计算方式为 Σ [x · P(x)]。
Expected value does not guarantee a single outcome; it is a weighted mean. In games of chance, it tells you whether the game is ‘fair’ or biased in favour of the house.
期望值并不能保证某一次的结果;它是一个加权平均数。在机遇游戏中,它能告诉你游戏是否「公平」或偏向庄家。
A simple example: rolling a fair die gives expected value = 3.5, even though you can never roll a 3.5.
一个简单的例子:掷一个公平的骰子,期望值 = 3.5,尽管你永远掷不出 3.5。
Think of expected value as the ‘centre of mass’ if probability were weight. It helps assess long-term financial or game outcomes.
将期望值想象成以概率为权重的「质心」。它有助于评估长期财务或游戏结果。
11. Sampling and Bias | 抽样与偏差
A population is the entire group you want to study; a sample is a subset selected from it. A good sample represents the population fairly.
总体是你想要研究的整个群体;样本是从中选出的一个子集。一个好的样本能够公平地代表总体。
Random sampling gives every member an equal chance of being selected, reducing selection bias. Stratified sampling divides the population into groups and samples proportionally from each.
随机抽样让每个成员都有相等的被选中的机会,从而减少选择偏差。分层抽样将总体分组,并按比例从每组抽样。
Bias is a systematic error that distorts results. Common types: selection bias, non-response bias, measurement bias.
偏差是一种系统性错误,会扭曲结果。常见类型:选择偏差、无回应偏差、测量偏差。
Visual clue: ‘Sample is a miniature of the population. If your sample is crooked, your conclusions will be crooked too.’
视觉提示:「样本是总体的微缩模型。如果你的样本歪了,你的结论也会歪。」
12. Correlation Coefficients and Regression | 相关系数与回归
The product-moment correlation coefficient (r) measures the strength and direction of a linear relationship. r = +1 is perfect positive, r = -1 is perfect negative, r = 0 indicates no linear correlation.
积差相关系数 (r) 衡量线性关系的强度和方向。r = +1 是完全正相关,r = -1 是完全负相关,r = 0 表示没有线性相关。
A regression line (y = a + bx) predicts the dependent variable y from the independent variable x. It is different from a trend line because it is precisely calculated using the least squares method.
回归线 (y = a + bx) 根据自变量 x 预测因变量 y。它不同于趋势线,因为它采用最小二乘法精确计算得出。
Key caution: extrapolation (predicting far beyond the data range) can be highly unreliable. Only interpolate within the observed range.
关键警示:外推(预测远超过数据范围)可能非常不可靠。只能在观测范围内进行内插。
Remember: ‘Correlation is a dance, regression gives you the steps.’
记法:「相关性是共舞,回归给你舞步。」
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导