📚 IGCSE Cambridge Statistics: Vocabulary Mnemonics and Quick Reference | IGCSE 剑桥统计:词汇术语速记指南
Mastering statistics begins with a clear understanding of its terminology. This guide pairs every key IGCSE Cambridge Statistics term with a concise definition and memory aid, helping you recall concepts faster in revision and exams. Each section groups related words so you can build your statistical vocabulary logically.
掌握统计始于透彻理解其术语。本指南将每个 IGCSE 剑桥统计关键术语与简明定义及记忆法配对,帮助你在复习和考试中更快回忆概念。每一节将相关词汇归类,让你能有逻辑地构建统计词汇体系。
1. Core Concepts: Population, Sample, Variable, Data | 核心概念:总体、样本、变量、数据
Population: The complete set of individuals, objects, or measurements you want to study. Think ‘pop’ as in everything pops up together. A census measures the whole population.
总体:你所希望研究的全部个体、对象或测量值的完整集合。记忆:总体如同人口普查,涵盖所有成员,一个都不少。
Sample: A subset of the population selected to represent it. If the selection is not careful, the sample can be biased. A reliable sample allows you to make inferences about the population.
样本:从总体中选出的一个子集,用以代表总体。若选取不慎,样本会带有偏差。可靠的样本让你能对总体做出推断。
Variable: Any characteristic that can take different values from one individual to another, e.g., height, test score, eye colour. Variables can be classified as qualitative or quantitative.
变量:在不同个体间可取不同值的任何特征,如身高、考试分数、眼睛颜色。变量可分为定性变量和定量变量。
Data: The actual observations you collect for a variable. ‘Data’ is plural (datum is singular). Data can be raw, grouped, primary (collected yourself) or secondary (from someone else).
数据:你为某一变量收集的实际观测值。’Data’ 是复数形式。数据可以是原始数据、分组数据、一手数据(自己收集)或二手数据(他人收集)。
2. Types of Data: Qualitative, Quantitative, Discrete, Continuous | 数据类型:定性、定量、离散、连续
Qualitative data: Non‑numerical categories, e.g., favourite colour, type of car. Mnemonic: Qualitative = Quality, it describes a property. Also called categorical data.
定性数据:非数值的类别,如最喜欢的颜色、汽车类型。记忆:定性关乎品质描述,亦称分类数据。
Quantitative data: Numerical measurements or counts. Quantitative = Quantity, it tells you how much or how many. It is further split into discrete and continuous types.
定量数据:数值的测量或计数。定量关乎数量大小,告诉你有多少或多高。定量数据进一步分为离散型和连续型。
Discrete data: Values that can only take distinct, separate numbers, usually whole counts (number of students, goals scored). You can count them; there are ‘jumps’ between values.
离散数据:只能取彼此分离的值的数值,通常为整数计数(学生人数、进球数)。可数,数值之间存在“跳跃”。
Continuous data: Can take any value within a range and is measured, not counted. Height, time, temperature are continuous. Precision depends on the measuring instrument.
连续数据:可在某个区间内取任意值,是通过测量而非计数得到的。身高、时间、温度属连续数据。精度取决于测量工具。
| Feature | Discrete | Continuous |
|---|---|---|
| Nature | Counted, whole numbers | Measured, any value in an interval |
| Example | Shoe sizes (if only whole sizes) | Length of a leaf (cm) |
| Graph | Bar chart with gaps | Histogram with adjoining bars |
3. Graphical Displays: Bar Chart, Pie Chart, Histogram, Frequency Polygon | 图表展示:条形图、饼图、直方图、频率多边形
Bar chart: Uses bars of equal width with gaps to show frequencies for categorical or discrete data. Height equals frequency. The categories can be reordered without losing meaning.
条形图:用等宽且有间隙的条形表示分类数据或离散数据的频数。高度代表频数。类别可重新排序而不影响含义。
Pie chart: A circular chart divided into sectors; sector angle = (frequency ÷ total frequency) × 360°. Ideal for showing proportions of a whole. Too many slices make it hard to read.
饼图:将圆形划分为扇形,扇区角度 = (频数 ÷ 总频数) × 360°。适合展示各部分占整体的比例。切片过多时不易阅读。
Histogram: For continuous (or grouped) data. Bars touch, and the area of each bar is proportional to frequency. The vertical axis is frequency density = frequency / class width. No gaps between bars.
直方图:用于连续数据(或分组数据)。条形之间紧贴,每个条形的面积与频数成比例。纵轴为频率密度 = 频数 ÷ 组距。条形之间无间隙。
Frequency polygon: A line graph formed by joining the midpoints of the tops of histogram bars (or class midpoints). Often used to compare two distributions on the same axes.
频率多边形:通过连接直方图各条形顶部中点(或组中值)形成的折线图。常用于在同一坐标轴上比较两个分布。
4. Cumulative Frequency and Box Plots | 累积频率与箱线图
Cumulative frequency: A running total of frequencies. Adding a ‘cumulative frequency’ column to a table helps you plot an ogive. Used to estimate medians and quartiles quickly.
累积频率:频数的累计总和。在表格中加入累积频率列可帮助你绘制累积频率曲线(ogive),便于快速估算中位数和四分位数。
Cumulative frequency graph (ogive): Plot the upper class boundary against cumulative frequency. A smooth S‑shaped curve is drawn. The median is found at half the total cumulative frequency.
累积频率曲线图(ogive):以组上界为横轴、累积频率为纵轴描点,连成平滑S形曲线。在累积频率一半处读取中位数。
Box‑and‑whisker plot (box plot): A diagram showing the five‑number summary: minimum, lower quartile (Q1), median (Q2), upper quartile (Q3), and maximum. It clearly reveals spread and skewness. Whiskers extend from the box to the extreme non‑outlier values.
箱线图(盒须图):展示五数概括:最小值、下四分位数 (Q1)、中位数 (Q2)、上四分位数 (Q3)、最大值。能清晰反映离散程度和偏态。须从箱体延伸至非异常极值。
5. Measures of Central Tendency: Mean, Median, Mode | 集中趋势量数:平均数、中位数、众数
Mean (arithmetic mean): The sum of all values divided by the number of values. Formula: x̄ = Σx/n for a sample. It is sensitive to extreme values (outliers). Think ‘mean’ as a fair share.
平均数(算术平均数):所有数值之和除以数值个数。样本平均数公式:x̄ = Σx/n。它对极端值(离群值)敏感。可记忆为“平均”即公平分配。
Median: The middle value when the data are ordered. For n odd, it is the (n+1)/2 th value; for n even, the average of the two middle values. Unaffected by outliers – use when data are skewed.
中位数:将数据排序后居中的数值。n为奇数时,第 (n+1)/2 个值;n为偶数时,中间两数的平均。不受离群值影响,数据偏斜时宜用中位数。
Mode: The most frequently occurring value. A data set may have one mode (unimodal), two modes (bimodal) or no mode. For grouped data, use the modal class – the class with the highest frequency density.
众数:出现频次最高的值。数据集可有单众数、双众数或无众数。对分组数据,使用众数组——频率密度最高的组。
6. Measures of Dispersion: Range, IQR, Standard Deviation | 离散程度量数:极差、四分位距、标准差
Range: maximum value − minimum value. Simplest measure of spread, but totally depends on the extremes. Quick to compute, but can be misleading if outliers exist.
极差(全距):最大值 − 最小值。最简单的离散程度量数,但完全取决于极端值。计算快捷,但有离群值时可能产生误导。
Quartiles and Interquartile Range (IQR): Q1 (lower quartile) is the median of the lower half, Q3 (upper quartile) is the median of the upper half. IQR = Q3 − Q1. The IQR measures the middle 50% spread and is resistant to outliers.
四分位数与四分位距 (IQR):Q1(下四分位数)是下半部分的中位数,Q3(上四分位数)是上半部分的中位数。IQR = Q3 − Q1。IQR 测量中间 50% 的离散程度,抗离群值。
Standard deviation (σ for population, s for sample): sqrt[ Σ(x − mean)² / n ] (or n−1 for sample). It measures how far, on average, each observation is from the mean. A small s means data cluster tightly around the mean.
标准差(σ 表示总体,s 表示样本):sqrt[ Σ(x − 平均数)² / n ](样本分母用 n−1)。它度量各观测值平均偏离平均数的距离。s 小则数据紧密聚集在平均数周围。
Variance: The square of the standard deviation (σ² or s²). It is used in further calculations but harder to interpret in original units.
方差:标准差的平方 (σ² 或 s²)。用于进一步计算,但以原始单位平方解释较困难。
7. Probability Vocabulary: Experiment, Event, Independence, Mutual Exclusivity | 概率术语:试验、事件、独立、互斥
Experiment / Trial: An action with a set of possible results, e.g., rolling a die. Outcome is a single possible result. Sample space is the set of all possible outcomes.
试验:产生一系列可能结果的行为,如掷一枚骰子。结果是单个可能的结果。样本空间是所有可能结果的集合。
Event: Any subset of the sample space, e.g., ‘rolling an even number’. A complementary event (A’) is everything not in A. Probability of an event is between 0 (impossible) and 1 (certain).
事件:样本空间的任意子集,如“掷出偶数”。互补事件 (A’) 是 A 不发生的所有情况。事件的概率介于 0(不可能)和 1(必然)之间。
Mutually exclusive events: Two events that cannot occur simultaneously. For mutually exclusive A and B, P(A or B) = P(A) + P(B). Mnemonic: ‘cannot share an outcome’.
互斥事件:不可能同时发生的两个事件。对于互斥事件 A 与 B,P(A 或 B) = P(A) + P(B)。记忆:无法共享任何结果。
Independent events: The occurrence of one does not affect the probability of the other. For independent A and B, P(A and B) = P(A) × P(B). Tree diagrams often model independence.
独立事件:一个事件的发生不影响另一个发生的概率。对于独立事件,P(A 且 B) = P(A) × P(B)。树状图常用于建模独立事件。
8. Correlation and Regression: Scatter Diagrams, Line of Best Fit | 相关与回归:散点图、最佳拟合线
Scatter diagram (scatter plot): A graph showing the relationship between two quantitative variables. Each point represents a pair (x, y). The pattern reveals correlation.
散点图:展示两个定量变量间关系的图表。每一点代表一对 (x, y) 坐标。点的分布形态揭示了相关性。
Correlation: Describes the direction and strength of a linear relationship. Positive: as x increases, y tends to increase. Negative: as x increases, y tends to decrease. Zero: no linear pattern. Correlation is not causation.
相关性:描述线性关系的方向与强度。正相关:x 增加,y 也趋于增加。负相关:x 增加,y 趋于减小。零相关:无线性模式。相关不蕴含因果。
Line of best fit: A straight line drawn through a scatter diagram to model the relationship. Drawn ‘by eye’ or using least squares regression. Used for prediction. Be careful with extrapolation (predicting outside the data range).
最佳拟合线:穿过散点图的一条直线,用于建模变量关系。可通过目测或最小二乘法绘制。用于预测。外推(超出数据范围的预测)须谨慎。
9. Sampling Methods: Random, Stratified, Systematic | 抽样方法:随机、分层、系统
Simple random sampling: Every member of the population has an equal chance of being selected. Usually uses random number generators or drawing lots. Avoids bias but can be impractical for large populations.
简单随机抽样:总体中每个成员被选中的机会均等。常使用随机数生成器或抽签。可避免偏差,但总体巨大时实施困难。
Stratified sampling: The population is divided into distinct strata (groups), and a random sample is taken from each stratum in proportion to its size. Guarantees representation of all key subgroups.
分层抽样:将总体分成不同的层(组),然后按各层大小比例从每层随机抽样。能确保所有关键子群均有代表。
Systematic sampling: Choose a starting point at random, then select every k-th member. Simple to use, but can introduce periodicity bias if the list has a hidden pattern.
系统抽样:随机选择一个起点,然后每隔 k 个抽取一个成员。使用简便,但若名单存在隐藏周期性,可能引入偏差。
Bias: A systematic error that causes the sample to misrepresent the population. Selection bias, non‑response bias and measurement bias are common types. Always check how the sample was obtained.
偏差:导致样本无法代表总体的系统性误差。常见类型有选择偏差、无回应偏差和测量偏差。务必核查样本的获取方式。
10. Distribution Shape: Symmetrical, Skewed, Normal | 分布形态:对称、偏斜、正态
Symmetrical distribution: The left and right sides are mirror images. In a perfectly symmetrical unimodal distribution, mean = median = mode. Bell‑shaped symmetry often suggests a normal distribution.
对称分布:左右两侧互为镜像。在完美的单峰对称分布中,平均数 = 中位数 = 众数。钟形对称常暗示正态分布。
Positively skewed (right‑skewed): The tail extends to the right. The mean is pulled towards the tail, so typically mean > median > mode. Income data often show positive skew.
正偏态(右偏):尾部向右延伸。平均数被拉向尾部,故通常 平均数 > 中位数 > 众数。收入数据常呈正偏态。
Negatively skewed (left‑skewed): The tail extends to the left. The mean is lower than the median, so mean < median < mode. Exam scores with a ceiling effect may be negatively skewed.
负偏态(左偏):尾部向左延伸。平均数低于中位数,故 平均数 < 中位数 < 众数。有天花板效应的考试成绩可能呈负偏态。
Normal distribution: A continuous, symmetric, bell‑shaped curve defined by its mean and standard deviation. Approximately 68% of data lie within ±1σ, 95% within ±2σ. Many natural measurements roughly follow a normal model.
正态分布:由均值和标准差决定的连续对称钟形曲线。约 68% 的数据落在 ±1σ 内,95% 落在 ±2σ 内。许多自然测量值近似正态分布。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply