IGCSE Cambridge Statistics: Vocabulary Mnemonics and Quick Reference | IGCSE 剑桥统计:词汇术语速记指南

📚 IGCSE Cambridge Statistics: Vocabulary Mnemonics and Quick Reference | IGCSE 剑桥统计:词汇术语速记指南

Mastering statistics begins with a clear understanding of its terminology. This guide pairs every key IGCSE Cambridge Statistics term with a concise definition and memory aid, helping you recall concepts faster in revision and exams. Each section groups related words so you can build your statistical vocabulary logically.

掌握统计始于透彻理解其术语。本指南将每个 IGCSE 剑桥统计关键术语与简明定义及记忆法配对,帮助你在复习和考试中更快回忆概念。每一节将相关词汇归类,让你能有逻辑地构建统计词汇体系。


1. Core Concepts: Population, Sample, Variable, Data | 核心概念:总体、样本、变量、数据

Population: The complete set of individuals, objects, or measurements you want to study. Think ‘pop’ as in everything pops up together. A census measures the whole population.

总体:你所希望研究的全部个体、对象或测量值的完整集合。记忆:总体如同人口普查,涵盖所有成员,一个都不少。

Sample: A subset of the population selected to represent it. If the selection is not careful, the sample can be biased. A reliable sample allows you to make inferences about the population.

样本:从总体中选出的一个子集,用以代表总体。若选取不慎,样本会带有偏差。可靠的样本让你能对总体做出推断。

Variable: Any characteristic that can take different values from one individual to another, e.g., height, test score, eye colour. Variables can be classified as qualitative or quantitative.

变量:在不同个体间可取不同值的任何特征,如身高、考试分数、眼睛颜色。变量可分为定性变量和定量变量。

Data: The actual observations you collect for a variable. ‘Data’ is plural (datum is singular). Data can be raw, grouped, primary (collected yourself) or secondary (from someone else).

数据:你为某一变量收集的实际观测值。’Data’ 是复数形式。数据可以是原始数据、分组数据、一手数据(自己收集)或二手数据(他人收集)。


2. Types of Data: Qualitative, Quantitative, Discrete, Continuous | 数据类型:定性、定量、离散、连续

Qualitative data: Non‑numerical categories, e.g., favourite colour, type of car. Mnemonic: Qualitative = Quality, it describes a property. Also called categorical data.

定性数据:非数值的类别,如最喜欢的颜色、汽车类型。记忆:定性关乎品质描述,亦称分类数据。

Quantitative data: Numerical measurements or counts. Quantitative = Quantity, it tells you how much or how many. It is further split into discrete and continuous types.

定量数据:数值的测量或计数。定量关乎数量大小,告诉你有多少或多高。定量数据进一步分为离散型和连续型。

Discrete data: Values that can only take distinct, separate numbers, usually whole counts (number of students, goals scored). You can count them; there are ‘jumps’ between values.

离散数据:只能取彼此分离的值的数值,通常为整数计数(学生人数、进球数)。可数,数值之间存在“跳跃”。

Continuous data: Can take any value within a range and is measured, not counted. Height, time, temperature are continuous. Precision depends on the measuring instrument.

连续数据:可在某个区间内取任意值,是通过测量而非计数得到的。身高、时间、温度属连续数据。精度取决于测量工具。

Feature Discrete Continuous
Nature Counted, whole numbers Measured, any value in an interval
Example Shoe sizes (if only whole sizes) Length of a leaf (cm)
Graph Bar chart with gaps Histogram with adjoining bars

3. Graphical Displays: Bar Chart, Pie Chart, Histogram, Frequency Polygon | 图表展示:条形图、饼图、直方图、频率多边形

Bar chart: Uses bars of equal width with gaps to show frequencies for categorical or discrete data. Height equals frequency. The categories can be reordered without losing meaning.

条形图:用等宽且有间隙的条形表示分类数据或离散数据的频数。高度代表频数。类别可重新排序而不影响含义。

Pie chart: A circular chart divided into sectors; sector angle = (frequency ÷ total frequency) × 360°. Ideal for showing proportions of a whole. Too many slices make it hard to read.

饼图:将圆形划分为扇形,扇区角度 = (频数 ÷ 总频数) × 360°。适合展示各部分占整体的比例。切片过多时不易阅读。

Histogram: For continuous (or grouped) data. Bars touch, and the area of each bar is proportional to frequency. The vertical axis is frequency density = frequency / class width. No gaps between bars.

直方图:用于连续数据(或分组数据)。条形之间紧贴,每个条形的面积与频数成比例。纵轴为频率密度 = 频数 ÷ 组距。条形之间无间隙。

Frequency polygon: A line graph formed by joining the midpoints of the tops of histogram bars (or class midpoints). Often used to compare two distributions on the same axes.

频率多边形:通过连接直方图各条形顶部中点(或组中值)形成的折线图。常用于在同一坐标轴上比较两个分布。


4. Cumulative Frequency and Box Plots | 累积频率与箱线图

Cumulative frequency: A running total of frequencies. Adding a ‘cumulative frequency’ column to a table helps you plot an ogive. Used to estimate medians and quartiles quickly.

累积频率:频数的累计总和。在表格中加入累积频率列可帮助你绘制累积频率曲线(ogive),便于快速估算中位数和四分位数。

Cumulative frequency graph (ogive): Plot the upper class boundary against cumulative frequency. A smooth S‑shaped curve is drawn. The median is found at half the total cumulative frequency.

累积频率曲线图(ogive):以组上界为横轴、累积频率为纵轴描点,连成平滑S形曲线。在累积频率一半处读取中位数。

Box‑and‑whisker plot (box plot): A diagram showing the five‑number summary: minimum, lower quartile (Q1), median (Q2), upper quartile (Q3), and maximum. It clearly reveals spread and skewness. Whiskers extend from the box to the extreme non‑outlier values.

箱线图(盒须图):展示五数概括:最小值、下四分位数 (Q1)、中位数 (Q2)、上四分位数 (Q3)、最大值。能清晰反映离散程度和偏态。须从箱体延伸至非异常极值。


5. Measures of Central Tendency: Mean, Median, Mode | 集中趋势量数:平均数、中位数、众数

Mean (arithmetic mean): The sum of all values divided by the number of values. Formula: x̄ = Σx/n for a sample. It is sensitive to extreme values (outliers). Think ‘mean’ as a fair share.

平均数(算术平均数):所有数值之和除以数值个数。样本平均数公式:x̄ = Σx/n。它对极端值(离群值)敏感。可记忆为“平均”即公平分配。

Median: The middle value when the data are ordered. For n odd, it is the (n+1)/2 th value; for n even, the average of the two middle values. Unaffected by outliers – use when data are skewed.

中位数:将数据排序后居中的数值。n为奇数时,第 (n+1)/2 个值;n为偶数时,中间两数的平均。不受离群值影响,数据偏斜时宜用中位数。

Mode: The most frequently occurring value. A data set may have one mode (unimodal), two modes (bimodal) or no mode. For grouped data, use the modal class – the class with the highest frequency density.

众数:出现频次最高的值。数据集可有单众数、双众数或无众数。对分组数据,使用众数组——频率密度最高的组。


6. Measures of Dispersion: Range, IQR, Standard Deviation | 离散程度量数:极差、四分位距、标准差

Range: maximum value − minimum value. Simplest measure of spread, but totally depends on the extremes. Quick to compute, but can be misleading if outliers exist.

极差(全距):最大值 − 最小值。最简单的离散程度量数,但完全取决于极端值。计算快捷,但有离群值时可能产生误导。

Quartiles and Interquartile Range (IQR): Q1 (lower quartile) is the median of the lower half, Q3 (upper quartile) is the median of the upper half. IQR = Q3 − Q1. The IQR measures the middle 50% spread and is resistant to outliers.

四分位数与四分位距 (IQR):Q1(下四分位数)是下半部分的中位数,Q3(上四分位数)是上半部分的中位数。IQR = Q3 − Q1。IQR 测量中间 50% 的离散程度,抗离群值。

Standard deviation (σ for population, s for sample): sqrt[ Σ(x − mean)² / n ] (or n−1 for sample). It measures how far, on average, each observation is from the mean. A small s means data cluster tightly around the mean.

标准差(σ 表示总体,s 表示样本):sqrt[ Σ(x − 平均数)² / n ](样本分母用 n−1)。它度量各观测值平均偏离平均数的距离。s 小则数据紧密聚集在平均数周围。

Variance: The square of the standard deviation (σ² or s²). It is used in further calculations but harder to interpret in original units.

方差:标准差的平方 (σ² 或 s²)。用于进一步计算,但以原始单位平方解释较困难。


7. Probability Vocabulary: Experiment, Event, Independence, Mutual Exclusivity | 概率术语:试验、事件、独立、互斥

Experiment / Trial: An action with a set of possible results, e.g., rolling a die. Outcome is a single possible result. Sample space is the set of all possible outcomes.

试验:产生一系列可能结果的行为,如掷一枚骰子。结果是单个可能的结果。样本空间是所有可能结果的集合。

Event: Any subset of the sample space, e.g., ‘rolling an even number’. A complementary event (A’) is everything not in A. Probability of an event is between 0 (impossible) and 1 (certain).

事件:样本空间的任意子集,如“掷出偶数”。互补事件 (A’) 是 A 不发生的所有情况。事件的概率介于 0(不可能)和 1(必然)之间。

Mutually exclusive events: Two events that cannot occur simultaneously. For mutually exclusive A and B, P(A or B) = P(A) + P(B). Mnemonic: ‘cannot share an outcome’.

互斥事件:不可能同时发生的两个事件。对于互斥事件 A 与 B,P(A 或 B) = P(A) + P(B)。记忆:无法共享任何结果。

Independent events: The occurrence of one does not affect the probability of the other. For independent A and B, P(A and B) = P(A) × P(B). Tree diagrams often model independence.

独立事件:一个事件的发生不影响另一个发生的概率。对于独立事件,P(A 且 B) = P(A) × P(B)。树状图常用于建模独立事件。


8. Correlation and Regression: Scatter Diagrams, Line of Best Fit | 相关与回归:散点图、最佳拟合线

Scatter diagram (scatter plot): A graph showing the relationship between two quantitative variables. Each point represents a pair (x, y). The pattern reveals correlation.

散点图:展示两个定量变量间关系的图表。每一点代表一对 (x, y) 坐标。点的分布形态揭示了相关性。

Correlation: Describes the direction and strength of a linear relationship. Positive: as x increases, y tends to increase. Negative: as x increases, y tends to decrease. Zero: no linear pattern. Correlation is not causation.

相关性:描述线性关系的方向与强度。正相关:x 增加,y 也趋于增加。负相关:x 增加,y 趋于减小。零相关:无线性模式。相关不蕴含因果。

Line of best fit: A straight line drawn through a scatter diagram to model the relationship. Drawn ‘by eye’ or using least squares regression. Used for prediction. Be careful with extrapolation (predicting outside the data range).

最佳拟合线:穿过散点图的一条直线,用于建模变量关系。可通过目测或最小二乘法绘制。用于预测。外推(超出数据范围的预测)须谨慎。


9. Sampling Methods: Random, Stratified, Systematic | 抽样方法:随机、分层、系统

Simple random sampling: Every member of the population has an equal chance of being selected. Usually uses random number generators or drawing lots. Avoids bias but can be impractical for large populations.

简单随机抽样:总体中每个成员被选中的机会均等。常使用随机数生成器或抽签。可避免偏差,但总体巨大时实施困难。

Stratified sampling: The population is divided into distinct strata (groups), and a random sample is taken from each stratum in proportion to its size. Guarantees representation of all key subgroups.

分层抽样:将总体分成不同的层(组),然后按各层大小比例从每层随机抽样。能确保所有关键子群均有代表。

Systematic sampling: Choose a starting point at random, then select every k-th member. Simple to use, but can introduce periodicity bias if the list has a hidden pattern.

系统抽样:随机选择一个起点,然后每隔 k 个抽取一个成员。使用简便,但若名单存在隐藏周期性,可能引入偏差。

Bias: A systematic error that causes the sample to misrepresent the population. Selection bias, non‑response bias and measurement bias are common types. Always check how the sample was obtained.

偏差:导致样本无法代表总体的系统性误差。常见类型有选择偏差、无回应偏差和测量偏差。务必核查样本的获取方式。


10. Distribution Shape: Symmetrical, Skewed, Normal | 分布形态:对称、偏斜、正态

Symmetrical distribution: The left and right sides are mirror images. In a perfectly symmetrical unimodal distribution, mean = median = mode. Bell‑shaped symmetry often suggests a normal distribution.

对称分布:左右两侧互为镜像。在完美的单峰对称分布中,平均数 = 中位数 = 众数。钟形对称常暗示正态分布。

Positively skewed (right‑skewed): The tail extends to the right. The mean is pulled towards the tail, so typically mean > median > mode. Income data often show positive skew.

正偏态(右偏):尾部向右延伸。平均数被拉向尾部,故通常 平均数 > 中位数 > 众数。收入数据常呈正偏态。

Negatively skewed (left‑skewed): The tail extends to the left. The mean is lower than the median, so mean < median < mode. Exam scores with a ceiling effect may be negatively skewed.

负偏态(左偏):尾部向左延伸。平均数低于中位数,故 平均数 < 中位数 < 众数。有天花板效应的考试成绩可能呈负偏态。

Normal distribution: A continuous, symmetric, bell‑shaped curve defined by its mean and standard deviation. Approximately 68% of data lie within ±1σ, 95% within ±2σ. Many natural measurements roughly follow a normal model.

正态分布:由均值和标准差决定的连续对称钟形曲线。约 68% 的数据落在 ±1σ 内,95% 落在 ±2σ 内。许多自然测量值近似正态分布。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version