📚 IGCSE CIE Statistics: Vocabulary & Terminology Quick-Reference Guide | IGCSE CIE 统计:词汇术语速记指南
Mastering statistical terminology is the first step to excelling in the IGCSE CIE Statistics examination. Whether you are preparing for the core or extended paper, a clear understanding of key terms such as mean, median, interquartile range, and correlation will help you interpret data correctly and avoid common pitfalls. This guide is designed to provide bilingual definitions and concise explanations, making it easier to memorise essential vocabulary.
掌握统计术语是IGCSE CIE 统计考试取得高分的第一步。无论你备考核心卷还是扩展卷,清晰理解平均值、中位数、四分位距、相关性等关键术语,能帮助你正确解读数据并避免常见错误。本指南提供双语定义和简明解释,助力速记必备词汇。
1. Data Classification | 数据分类
Qualitative data (or categorical data) describes qualities or groups and cannot be measured with numbers directly, for example eye colour or type of car.
定性数据(分类数据)描述属性或类别,无法直接用数字测量,例如眼睛颜色或汽车类型。
Quantitative data (or numerical data) consists of numbers that represent counts or measurements, such as height in cm or number of students.
定量数据(数值数据)由表示计数或测量的数字组成,例如身高(厘米)或学生人数。
Discrete data can only take specific, separate values, usually whole numbers, like the number of pets in a household.
离散数据只能取特定的、分立的值,通常是整数,例如家庭宠物数量。
Continuous data can take any value within a given range and is measured, such as mass, time or temperature.
连续数据可以在给定范围内取任意值,通过测量得到,例如质量、时间或温度。
Nominal data labels categories with no intrinsic order, e.g. favourite colour: red, blue, green.
名义数据用标签分类且无固有顺序,例如最喜欢的颜色:红、蓝、绿。
Ordinal data categories have a meaningful order but no precise numerical differences, like survey ratings: poor, good, excellent.
有序数据类别具有有意义的顺序但没有精确的数值差异,例如调查评级:差、好、优秀。
2. Key Sampling Terms | 关键抽样术语
Population refers to the entire set of individuals or items the study aims to understand.
总体指研究想要了解的全部个体或项目的集合。
A sample is a subset of the population selected to represent the whole for data collection.
样本是从总体中选出的子集,用于代表整体进行数据收集。
In a random sample, every member of the population has an equal chance of being chosen, reducing selection bias.
在随机样本中,总体的每个成员被选中的机会均等,以减少选择偏差。
Stratified sampling divides the population into distinct subgroups (strata) and then randomly samples proportionally from each stratum.
分层抽样将总体分为不同的子组(层),然后按比例从每层随机抽样。
Systematic sampling selects individuals at regular intervals from an ordered list, e.g. every 10th name.
系统抽样从有序列表中每隔固定间隔选取个体,例如每第10个名字。
Bias occurs when some members of the population are consistently more likely to be selected, distorting the results.
当总体的某些成员持续更有机会被选中时,就会产生偏差,扭曲结果。
3. Graphical Representations | 图形表示
A bar chart uses separate bars of equal width to display the frequency of categorical data; the height shows frequency.
条形图使用等宽的独立条形来显示分类数据的频数;条形高度表示频数。
A pie chart is a circular graph divided into sectors whose angles are proportional to the frequency of each category.
饼图是一种圆形图表,被分成扇形,每个扇形的角度与对应类别的频数成比例。
A histogram represents grouped continuous data with bars that touch; the area of each bar is proportional to the frequency, and the height represents frequency density.
直方图用相邻的条形表示分组连续数据;每个条形的面积与频数成比例,高度表示频率密度。
Frequency Density = Frequency ÷ Class Width
频率密度 = 频数 ÷ 组距
A frequency polygon joins the midpoints of the tops of histogram bars with straight lines, often extended to show trend.
频率多边形用直线连接直方图各条形顶部的中点,通常延伸以显示趋势。
A cumulative frequency curve (ogive) plots cumulative frequency against the upper class boundary, useful for estimating medians and quartiles.
累积频数曲线(肩形图)将累积频数相对于组上界作图,用于估计中位数和四分位数。
A box-and-whisker plot displays the five‑number summary: minimum, lower quartile (Q₁), median, upper quartile (Q₃), and maximum, highlighting spread and outliers.
箱线图显示五数概括:最小值、下四分位数(Q₁)、中位数、上四分位数(Q₃)和最大值,突出散布和异常值。
A stem‑and‑leaf diagram lists data values by splitting each number into a stem (leading digit(s)) and a leaf (final digit), preserving original data.
茎叶图通过将每个数值拆分为茎(前导数字)和叶(末位数字)来列出数据,保留了原始数据。
4. Measures of Central Tendency | 集中趋势度量
The mean is the arithmetic average, calculated by summing all values and dividing by the number of values. For grouped data, use midpoints: mean = Σ(f × x) ÷ Σf.
平均值是算术平均数,由所有数值之和除以数值个数算出。对于分组数据,使用组中值:平均值 = Σ(f × x) ÷ Σf。
The median is the middle value when data are ordered. If n is even, it is the mean of the two middle values.
中位数是数据排序后的中间值。若n为偶数,则是中间两个数的平均值。
The mode is the value that occurs most frequently. A data set can have one mode, more than one mode, or no mode.
众数是出现频数最高的数值。一组数据可以有一个众数、多于一个众数或无众数。
The modal class is the class interval with the highest frequency in grouped data, pinpointing the most common range.
众数组是分组数据中频数最高的组距,指出最常见的范围。
The mid-range is the average of the smallest and largest values, easily influenced by extreme values.
中程数是最小值与最大值的平均数,容易受极端值影响。
5. Measures of Dispersion | 离散程度度量
The range is the difference between the largest and smallest values, giving a simple measure of spread.
极差是最大值与最小值的差值,提供简单的离散度量。
The interquartile range (IQR) is the range of the middle 50% of the data: IQR = Q₃ − Q₁. It is resistant to outliers.
四分位距(IQR)是中间50%数据的范围:IQR = Q₃ − Q₁,对异常值不敏感。
The lower quartile (Q₁) is the median of the lower half of the data; the upper quartile (Q₃) is the median of the upper half.
下四分位数(Q₁)是数据下半部分的中位数;上四分位数(Q₃)是上半部分的中位数。
Variance measures the average squared deviation from the mean. For a population: σ² = Σ(x − μ)² ÷ N.
方差衡量数据与平均值偏差平方的平均数。对于总体:σ² = Σ(x − μ)² ÷ N。
σ² = Σ(x − μ)² ÷ N
Standard deviation (σ) is the square root of the variance, giving spread in original units. σ = √[Σ(x − μ)² ÷ N].
标准差(σ)是方差的平方根,以原始单位表示离散程度。σ = √[Σ(x − μ)² ÷ N]。
σ = √[Σ(x − μ)² ÷ N]
An outlier is a value lying far outside the main body of data, often defined as more than 1.5 × IQR below Q₁ or above Q₃.
异常值是远离数据主体的值,通常定义为低于 Q₁ − 1.5 × IQR 或高于 Q₃ + 1.5 × IQR。
6. Probability Terminology | 概率术语
A random experiment is a process that leads to a well‑defined set of possible outcomes, such as rolling a die.
随机试验是产生一组明确定义的可能结果的过程,例如掷骰子。
The sample space (S) is the set of all possible outcomes of an experiment, often listed or shown in a table.
样本空间(S)是试验所有可能结果的集合,常用列表或表格表示。
An event is any subset of the sample space; probability of event A is P(A) = number of favourable outcomes ÷ total number of outcomes.
事件是样本空间的任意子集;事件A的概率为 P(A) = 有利结果数量 ÷ 总结果数量。
P(A) = n(A) ÷ n(S)
Experimental probability (relative frequency) is based on actual trials: P(event) = number of successful trials ÷ total trials.
实验概率(相对频率)基于实际试验:P(事件) = 成功试验次数 ÷ 总试验次数。
Mutually exclusive events cannot happen at the same time, so P(A or B) = P(A) + P(B).
互斥事件不能同时发生,因此 P(A 或 B) = P(A) + P(B)。
Independent events do not affect each other’s probabilities; for independent events, P(A and B) = P(A) × P(B).
独立事件互不影响对方的概率;对于独立事件,P(A 和 B) = P(A) × P(B)。
A tree diagram shows all possible outcomes of a compound event with branches labelled by probabilities, helping to calculate ‘and’ and ‘or’ scenarios.
树状图显示复合事件所有可能的结果,分支上标有概率,有助于计算“且”与“或”的情景。
7. Correlation and Regression | 相关与回归
A scatter diagram plots pairs of values for two variables on a coordinate grid, revealing possible relationship.
散点图在坐标网格上绘制两个变量的成对值,揭示可能的关系。
Positive correlation means as one variable increases, the other tends to increase; the points slope upward.
正相关表示当一个变量增加时,另一个变量也趋向增加;点向上倾斜。
Negative correlation means as one variable increases, the other tends to decrease; the points slope downward.
负相关表示当一个变量增加时,另一个变量趋向减少;点向下倾斜。
No correlation indicates no clear linear pattern between the two variables.
无相关表明两个变量之间没有明显的线性模式。
The line of best fit is drawn through the scatter points to model the trend and enable prediction. It can be drawn by eye.
最佳拟合线穿过散点以模拟趋势并进行预测,可以目估绘制。
Interpolation uses the line of best fit to estimate a value inside the range of given data, which is generally reliable.
内插使用最佳拟合线估计已知数据范围内的值,通常较为可靠。
Extrapolation extends the line beyond the data range to predict unknown values, which can be risky and less accurate.
外推将线延伸到数据范围之外以预测未知值,这可能有风险且准确性较低。
Spearman’s rank correlation coefficient (rₛ) measures the strength of monotonic relationship between two ranked variables.
斯皮尔曼等级相关系数(rₛ)度量两个排序变量之间单调关系的强度。
rₛ = 1 − (6 Σd²) ÷ [n(n² − 1)]
8. Data Integrity and Misleading Statistics | 数据可靠性与误导性统计
A misleading graph may use truncated axes, distorted scales, or inappropriate chart types to exaggerate or hide trends.
误导性图表可能通过截断坐标轴、扭曲比例或不恰当的图表类型来夸大或隐藏趋势。
Sampling bias happens when the sample is not representative, e.g. surveying only a convenience group.
抽样偏差发生于样本不具代表性时,例如只调查方便获取的群体。
Leading questions in surveys can influence answers, for instance ‘Don’t you agree that…?’
调查中的引导性问题会影响回答,例如“难道您不同意……吗?”
Accuracy refers to how close a measured value is to the true value, while reliability refers to consistency of repeated measurements.
准确性指测量值接近真实值的程度,而可靠性指重复测量的一致程度。
A census collects data from every member of the population, eliminating sampling error but often costly and time‑consuming.
普查收集总体中每个成员的数据,消除了抽样误差,但通常成本高且耗时。
9. Quick‑Reference Glossary | 快速词汇表
Raw data: unprocessed information collected directly from observation or survey.
原始数据:从观察或调查中直接收集的未经处理的信息。
Class interval: the width of a group in a frequency table, e.g. 10 ≤ x < 20.
组距:频数表中一组的宽度,例如 10 ≤ x < 20。
Cumulative frequency: the running total of frequencies up to the end of a given class interval.
累积频数:截至给定组距末尾的频数累计和。
Percentile: the value below which a given percentage of observations fall; the median is the 50th percentile.
百分位数:给定百分比的观测值落在其下的数值;中位数是第50百分位数。
Symmetry: a distribution where the left and right sides are mirror images; mean and median coincide.
对称:分布左右侧互为镜像,平均值与中位数重合。
Skewness: describes the lack of symmetry; positive skew means the tail extends to the right, negative skew to the left.
偏度:描述不对称性;正偏表示尾部向右延伸,负偏表示向左延伸。
Bivariate data: data involving two variables, often analysed for correlation.
双变量数据:涉及两个变量的数据,常用于相关分析。
Frequency density: a calculated value (frequency ÷ class width) used to plot the height of bars in a histogram.
频率密度:计算值(频数 ÷ 组距),用于确定直方图中条形的高度。
Outlier boundary: fences defined by Q₁ − 1.5×IQR and Q₃ + 1.5×IQR to identify extreme values.
异常值界限:由 Q₁ − 1.5×IQR 和 Q₃ + 1.5×IQR 定义的围栏,用于识别极端值。
Line of best fit equation: for a linear model y = a + bx, b is the gradient and a is the y‑intercept, often found from a scatter graph.
最佳拟合线方程:对于线性模型 y = a + bx,b 是斜率,a 是 y
Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导