📚 IGCSE Cambridge Statistics: Glossary and Mnemonic Speed Guide | IGCSE 剑桥统计:词汇术语速记指南
This revision guide sorts the most important IGCSE Statistics terms into logical groups and pairs each idea with a quick memory hook. Use it to check definitions, avoid confusion in written answers, and revise efficiently before the Cambridge exam.
本速记指南将 IGCSE 统计最重要的术语分组整理,并为每个概念配上快速记忆线索。你可以用它检查定义、避免在书面作答中混淆概念,并在剑桥考试前高效复习。
1. Why Terminology Matters | 为什么要重视术语
IGCSE Statistics rewards precise language. Many questions do not ask for long calculations; they ask you to define, state, identify or explain a statistical term. If two words such as ‘discrete’ and ‘continuous’ are mixed up, marks can be lost quickly.
IGCSE 统计考试非常看重语言精确。许多题目并不要求冗长计算,而是要求你定义、陈述、识别或解释某个统计术语。如果把 ‘discrete’ 和 ‘continuous’ 这类词混淆,很容易丢分。
A practical strategy is to learn terms in pairs or small clusters, not as isolated words. For example, learn ‘population’ together with ‘sample’, and ‘range’ together with ‘interquartile range’. That makes differences clearer.
一个实用的策略是成对或分组学习术语,而不是孤立背单词。例如把 ‘population’ 和 ‘sample’ 一起学,把 ‘range’ 和 ‘interquartile range’ 一起学,差别会更清楚。
The most common command words are define, state, calculate, compare, describe, explain and interpret. Each one requires a different type of answer, so the same term may appear in different question styles.
最常见的指令词是 define、state、calculate、compare、describe、explain 和 interpret。每个词要求的作答方式不同,因此同一个术语可能以不同题型出现。
2. Data Types: Qualitative, Quantitative, Discrete, Continuous | 数据类型:定性、定量、离散、连续
Data can first be classified as qualitative or quantitative. Qualitative data are non-numerical descriptions, such as eye colour, favourite sport or type of transport. Quantitative data are numerical measurements or counts.
数据首先可分为定性数据和定量数据。定性数据是非数字描述,例如眼睛颜色、最喜欢的运动或交通方式。定量数据是数字测量或计数。
Quantitative data are then split into discrete and continuous. Discrete data can only take exact, separate values, usually counted. Examples are number of children, shoe size or goals scored. Continuous data can take any value in a range, usually measured. Examples are height, mass, time or temperature.
定量数据又分为离散数据和连续数据。离散数据只能取精确的、分开的值,通常是数出来的,例如孩子数量、鞋码或进球数。连续数据在一个区间内可以取任意值,通常是测量出来的,例如身高、质量、时间或温度。
Quick memory hook: ‘Discrete = counted; Continuous = measured.’ If you can ask ‘How many?’ it is probably discrete. If you can ask ‘How much?’ and use a scale, it is continuous.
快速记忆:Discrete 是数出来的,Continuous 是量出来的。如果能问 ‘How many?’,通常是离散数据;如果能问 ‘How much?’ 并需要用刻度测量,通常是连续数据。
Do not confuse ‘qualitative’ with ‘discrete’: qualitative data may use numbers as labels, such as jersey number 7, but those numbers cannot be used in calculations meaningfully.
不要把 ‘qualitative’ 和 ‘discrete’ 混淆:定性数据有时用数字做标签,例如 7 号球衣,但这些数字不能进行有意义的计算。
3. Population, Sample and Census | 总体、样本与普查
The population is the entire group being studied, such as all students in a school or all cars in a city. A sample is a subset of the population selected to represent it. A census collects data from every member of the population.
总体是被研究的整个群体,例如一所学校的所有学生或一座城市的所有汽车。样本是从总体中选出的一部分,用来代表总体。普查则从总体的每一个成员收集数据。
Memory hook: think of the population as a whole pizza and the sample as one slice. A census means you examine every slice and the plate too.
记忆线索:把总体想象成一整个披萨,样本是其中一块。普查意味着你把每一块甚至盘子都检查一遍。
A sample is often used instead of a census because a census can be expensive, time-consuming or impossible. However, a sample can be biased if it does not represent the population fairly.
通常使用样本而不是普查,因为普查可能昂贵、耗时或根本不可行。但如果样本不能公平地代表总体,就会产生偏差。
A sampling frame is the list of all members from which a sample is drawn. Bias is any systematic error that makes the sample differ from the population in a particular direction.
抽样框是用来抽取样本的全体成员名单。偏差是任何系统性的误差,它使样本在某个方向上与总体不一致。
4. Sampling Methods: Random, Stratified, Systematic, Quota | 抽样方法:随机、分层、系统、配额
In random sampling, every member of the population has an equal chance of being selected. It reduces bias but requires a complete sampling frame and may be impractical for large populations.
在随机抽样中,总体中每个成员被抽中的机会相等。它可以减少偏差,但需要完整的抽样框,对大规模总体可能不太实际。
Stratified sampling divides the population into groups called strata, such as age groups or year groups, then takes a random sample from each group in proportion to its size. The formula for the number selected from one group is:
分层抽样先把总体分成若干层,例如年龄组或年级组,然后按照各层在总体中的比例从每层随机抽取。从某一层抽取的数量公式为:
Number from group = (Group size ÷ Total population) × Total sample size
Stratified sampling is useful because it guarantees representation from each subgroup. The memory hook is ‘Stratified = layers’.
分层抽样很有用,因为它保证每个子群体都有代表。记忆线索是 ‘Stratified = 层次’。
Systematic sampling selects every nth member from a list after a random start. For example, choose every 20th student in a register. It is quick, but if there is a hidden pattern in the list, bias can occur.
系统抽样在随机起点之后,从名单中每隔 n 个成员抽取一个。例如每隔 20 名学生抽一个。它速度快,但如果名单中存在隐藏规律,就可能产生偏差。
Quota sampling is non-random. The interviewer is given fixed numbers, or quotas, of people to select from different groups. It is cheap and fast, but it can be biased because the selection within each quota is not random.
配额抽样是非随机抽样。调查员被要求从不同群体中按固定数量抽取人员。它成本低、速度快,但每个配额内的选择不是随机的,因此可能产生偏差。
Convenience sampling selects people who are easy to reach. It is rarely representative, so it is weak for serious statistical conclusions. Memory hook: ‘Random = lottery; Stratified = layers; Systematic = steps; Quota = quantities.’
便利抽样选择容易接触到的人。它很少具有代表性,因此不适合得出严肃的统计结论。记忆线索:Random 是抽签,Stratified 是分层,Systematic 是隔步,Quota 是定量。
5. Frequency Tables and Diagrams | 频数表与图表
A frequency table records how often each value or class occurs. A class interval is a range such as 10 ≤ x < 20. The class width is the size of that interval, and the class midpoint is the centre of the interval.
频数表记录每个数值或组出现的次数。组区间是一个范围,例如 10 ≤ x < 20。组距是这个区间的大小,组中值是区间的中心。
Cumulative frequency is the running total of frequencies. You find it by adding each frequency to the sum of the previous frequencies. It is used to draw an ogive and to locate medians and quartiles.
累积频数是频数的累加总和。把每个频数与之前频数之和相加即可得到。它用于绘制累积频数曲线,并寻找中位数和四分位数。
For a histogram with unequal class widths, the vertical axis is frequency density, not frequency. The formula is:
当组距不相等时,直方图的纵轴是频数密度,而不是频数。公式为:
Frequency density = Frequency ÷ Class width
Memory hook: ‘FD = F over W’ means frequency density is frequency spread over the width. A bar with a narrow width but a large frequency will be very tall.
记忆线索:’FD = F over W’ 表示频数密度等于频数除以组距。组距窄而频数大的条形会非常高。
Always check whether the question gives equal or unequal class widths. If widths are equal, a bar chart may be acceptable, but a histogram with frequency density is expected when widths differ.
务必检查题目给出的是相等组距还是不相等组距。如果组距相等,条形图可能可用;如果组距不同,则应该使用以频数密度为纵轴的直方图。
6. Averages: Mean, Median, Mode | 平均数:均值、中位数、众数
An average is a measure of central tendency. The three most common averages are the mean, median and mode. Each one summarises a data set in a different way.
平均数是一种集中趋势的度量。最常见的三种平均数是均值、中位数和众数。它们以不同方式概括一组数据。
The mean is found by adding all values and dividing by the number of values:
均值是将所有数值相加后除以数值的个数:
Mean = Σx ÷ n
The mean uses every value, so it is sensitive to outliers. If one extreme value changes, the mean changes noticeably.
均值使用每一个数据,因此对异常值很敏感。只要一个极端值改变,均值就会明显改变。
The median is the middle value when the data are arranged in order. If there are two middle values, take their mean. The median is not strongly affected by outliers, so it is useful for skewed data.
中位数是数据按顺序排列后的中间值。如果有两个中间值,就取它们的均值。中位数不容易受异常值影响,因此适用于偏斜数据。
The mode is the most frequent value or class. A data set can have one mode, more than one mode, or no mode. The mode is the only average that can be used for qualitative data.
众数是出现次数最多的数值或组。一组数据可以有一个众数、多个众数,或者没有众数。众数是唯一可用于定性数据的平均数。
Memory hook: ‘Mean is the balance point; median is the middle; mode is the most.’ Another useful phrase is ‘Mean is mean to outliers.’
记忆线索:均值是平衡点,中位数是中间值,众数是最常见值。另一个有用的说法是 ‘Mean is mean to outliers’,意思是均值对异常值很敏感。
7. Spread: Range, Quartiles, IQR and Standard Deviation | 离散程度:极差、四分位数、四分位距与标准差
A measure of spread tells you how spread out the data are. The simplest is the range:
离散程度的度量告诉你数据分散到什么程度。最简单的是极差:
Range = Maximum value − Minimum value
The range is easy to compute, but it only uses the two most extreme values and can be distorted by outliers.
极差容易计算,但它只使用两个最极端的值,容易被异常值扭曲。
Quartiles divide an ordered data set into four equal parts. The lower quartile Q₁ is the first quarter, the median Q₂ is the second quarter, and the upper quartile Q₃ is the third quarter.
四分位数把有序数据分成四个相等的部分。下四分位数 Q₁ 是第一个四分位,中位数 Q₂ 是第二个四分位,上四分位数 Q₃ 是第三个四分位。
The interquartile range (IQR) is the difference between the upper and lower quartiles:
四分位距是上四分位数与下四分位数之差:
IQR = Q₃ − Q₁
The IQR measures the spread of the middle 50% of the data. It is more resistant to outliers than the range. Memory hook: ‘IQR ignores the extremes.’
四分位距度量中间 50% 数据的分散程度。它比极差更能抵抗异常值。记忆线索:’IQR ignores the extremes’,即四分位距忽略极端值。
Standard deviation measures how far values are from the mean on average. A small standard deviation means the data are clustered around the mean; a large one means they are widely spread. The population standard deviation formula is:
标准差度量数据值平均离均值多远。标准差小说明数据集中在均值附近,标准差大说明数据分布很分散。总体标准差公式为:
σ = √(Σ(x − μ)² ÷ n)
Variance is the square of the standard deviation. In calculator work, know whether the data are a population or a sample, because sample standard deviation uses n − 1 in the denominator.
方差是标准差的平方。在使用计算器时,要清楚数据是总体还是样本,因为样本标准差的分母用 n − 1。
8. Charts and Graphs: Histogram, Ogive, Box Plot, Scatter Diagram | 统计图表:直方图、累积频数曲线、箱线图、散点图
A histogram displays frequency density for continuous or grouped data. The area of each bar is proportional to the frequency. If class widths are equal, the heights are also proportional to frequencies.
直方图用频数密度表示连续或分组数据。每个条形的面积与频数成正比。如果组距相等,条形高度也与频数成正比。
A frequency polygon is made by joining the midpoints of the tops of histogram bars. It helps show the shape of the distribution, such as symmetrical, positively skewed or negatively skewed.
频数折线图是通过连接直方图条形顶部中点得到的。它有助于显示分布形状,例如对称、正偏斜或负偏斜。
A cumulative frequency curve, or ogive, plots the upper class boundary against cumulative frequency. Use it to estimate the median by finding the value at half the total frequency, and the quartiles at one quarter and three quarters of the total.
累积频数曲线是把组的上限对累积频数作图。利用它可以估计中位数,即总频数一半对应的值;四分位数则对应总频数的四分之一和四分之三。
A box-and-whisker plot shows the five-number summary: minimum, Q₁, median, Q₃ and maximum. It is excellent for comparing two data sets and for showing skewness.
箱线图显示五数综合:最小值、Q₁、中位数、Q₃ 和最大值。它非常适合比较两组数据以及展示偏斜情况。
A scatter diagram shows bivariate data: pairs of values for two variables. Each point represents one item. The pattern tells you whether there is correlation, and a line of best fit can model the relationship.
散点图显示双变量数据,即两个变量的一对值。每个点代表一个个体。点的分布模式能说明是否存在相关关系,最佳拟合线则可以建立模型。
9. Correlation and Regression | 相关与回归
Correlation describes the strength and direction of a linear relationship between two variables. Positive correlation means as one variable increases, the other tends to increase. Negative correlation means as one variable increases, the other tends to decrease.
相关描述两个变量之间线性关系的强度和方向。正相关表示一个变量增大时,另一个变量通常也增大。负相关表示一个变量增大时,另一个变量通常减小。
Correlation can be strong, weak or zero. The correlation coefficient r ranges from −1 to +1. A value close to +1 is strong positive, close to −1 is strong negative, and close to 0 is weak or no linear correlation.
相关可以是强、弱或零相关。相关系数 r 的范围是 −1 到 +1。接近 +1 为强正相关,接近 −1 为强负相关,接近 0 为弱相关或无线性相关。
A line of best fit is drawn through the points on a scatter diagram. It should pass close to the mean point, and it can be used to estimate one variable from another.
最佳拟合线穿过散点图中的点,应该接近均值点,可以用来由一个变量估计另一个变量。
Interpolation is estimating a value inside the range of the data. Extrapolation is estimating a value outside the range. Extrapolation is less reliable because the trend may not continue beyond the observed data.
内插是在数据范围内进行估计。外推是在数据范围之外进行估计。外推不太可靠,因为趋势不一定在观测数据之外继续保持不变。
The most important exam phrase is: ‘Correlation does not imply causation.’ Two variables may move together because both depend on a third variable, or simply by chance.
考试中最重要的一句话是:相关不代表因果。两个变量可能因为都依赖于第三个变量而一起变化,也可能只是偶然。
10. Probability Essentials | 概率基础
Probability measures how likely an event is to occur. It is always between 0 and 1. A probability of 0 means impossible; a probability of 1 means certain.
概率度量事件发生的可能性。它的取值总是在 0 到 1 之间。概率为 0 表示不可能发生,概率为 1 表示必然发生。
The basic probability formula is:
基本概率公式为:
P(Event) = Number of favourable outcomes ÷ Total number of outcomes
A sample space is the set of all possible outcomes. An event is a subset of the sample space. For example, rolling a die has sample space {1, 2, 3, 4, 5, 6}, and ‘rolling an even number’ is an event.
样本空间是所有可能结果的集合。事件是样本空间的一个子集。例如掷一个骰子的样本空间是 {1, 2, 3, 4, 5, 6},’掷出偶数’ 就是一个事件。
Mutually exclusive events cannot happen at the same time. For mutually exclusive events:
互斥事件不能同时发生。对于互斥事件:
P(A or B) = P(A) + P(B)
Independent events are those where one event does not affect the probability of the other. For independent events:
独立事件是指一个事件的发生不影响另一个事件发生的概率。对于独立事件:
P(A and B) = P(A) × P(B)
Conditional probability is the probability of event A given that event B has already happened. The formula is:
条件概率是在事件 B 已经发生的情况下事件 A 发生的概率。公式为:
P(A|B) = P(A and B) ÷ P(B)
Expected frequency is the number of times an event is expected to happen in repeated trials. It is found by multiplying the probability by the number of trials:
期望频数是在重复试验中某事件预计发生的次数。它等于概率乘以试验次数:
Expected frequency = n × P(A)
11. Exam Command Words and Quick Mnemonics | 考试指令词与快速记忆法
Command
Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导