📚 IGCSE CCEA Statistics: Essential Vocabulary Quick-Memorisation Guide | IGCSE CCEA 统计:核心词汇速记指南
The building blocks of statistical understanding are the words we use. In the IGCSE CCEA Statistics course, you must not only perform calculations but also interpret problems, explain findings, and justify choices. A strong command of terminology turns a jumble of numbers into a clear story. This guide groups essential terms into ten thematic sections, providing direct English–Chinese pairings and memory hooks to help you master them quickly and confidently.
统计理解的基石是我们使用的词汇。在 IGCSE CCEA 统计课程中,你不仅要完成计算,还要能解释问题、阐述发现、论证选择。扎实的术语功底能把一堆杂乱的数据变成清晰的故事。本指南将核心词汇归入十个主题小节,提供直接的英文–中文配对和记忆钩子,助你快速、自信地掌握它们。
1. Populations and Samples | 总体与样本
A population is the entire group of individuals or items that we want to investigate. Think of it as all possible members that fit a description, such as all students in a school or every lightbulb from a production line.
总体(population)是我们想研究的全部个体或事物。可以把它想象成符合描述的所有成员,比如一所学校的所有学生或生产线上出的每一个灯泡。
A sample is a smaller group selected from the population. We use a sample to draw conclusions about the population when it is impractical or impossible to measure every member. The key is that the sample must be representative – mirroring the characteristics of the population.
样本(sample)是从总体中选出的较小群体。当测量每个成员不现实或不可能时,我们就用样本来推断总体。关键在于样本必须具有代表性(representative)——能反映总体的特征。
A sampling frame is a list of all members of the population from which a sample is actually drawn. Without a clear frame, bias can creep in.
抽样框(sampling frame)是总体中所有成员的名单,样本实际就是从这份名单中抽选的。没有清晰的抽样框,就容易引入偏差。
The census is when data is collected from every single member of the population. It gives the true values but is often costly and time-consuming.
普查(census)是指从总体的每一个成员那里收集数据。它能给出真实值,但通常成本高、耗时长。
2. Types of Data: Qualitative vs Quantitative | 数据类型:定性数据与定量数据
Qualitative data describes qualities or categories that cannot be measured with numbers in a natural way. It is often called categorical data. Examples include eye colour, type of car, or favourite subject.
定性数据(qualitative data)描述的是无法用数字自然度量的品质或类别。它常被称作分类数据。例如眼睛的颜色、汽车类型或最喜欢的科目。
Quantitative data consists of numerical measurements or counts. It answers ‘how much’ or ‘how many’. Height, test scores, and the number of pets are quantitative.
定量数据(quantitative data)由数值测量或计数组成。它回答“有多少”或“多少量”。身高、测验分数和宠物数量都属于定量数据。
Within quantitative data, we further split into discrete and continuous. Discrete data can only take certain values – usually whole numbers, like the number of students in a class. Continuous data can take any value within a range, like weight, time, or temperature.
在定量数据内部,我们进一步分为离散数据(discrete data)和连续数据(continuous data)。离散数据只能取特定的值——通常是整数,比如班上的学生人数。连续数据可以在一个范围内取任意值,比如体重、时间或温度。
A quick memory hook: if you would count it, it’s discrete; if you would measure it with a ruler or scale, it’s continuous.
记忆小窍门:如果是数出来的,就是离散数据;如果是用尺或秤测出来的,就是连续数据。
3. Discrete and Continuous Data in Depth | 离散数据与连续数据的深入辨析
When you see a table of shoe sizes displayed as 6, 7, 8, remember that although shoe size looks like a number, it is often treated as discrete because only half and full sizes exist. However, the actual length of a foot is continuous.
当你看到鞋码显示为 6、7、8 这样的表格时,请记住,尽管鞋码看起来像数字,但由于只有半码和整码存在,它通常被视为离散数据。然而,脚的实际长度是连续数据。
In charts, discrete data is best shown with bar charts or frequency diagrams where gaps appear between bars. Continuous data is grouped into class intervals and displayed using histograms with no gaps.
在图表中,离散数据最适合用条形图或有间隔的频率图表示。连续数据则被分入组距(class intervals),用没有间隔的直方图(histogram)展示。
A class interval for continuous data must be written clearly, for example 0 ≤ x < 10. The boundaries are often called lower class boundary and upper class boundary, and the midpoint is used for calculations.
连续数据的组距(class interval)必须书写清晰,例如 0 ≤ x < 10。其边界常被称为下组界(lower class boundary)和上组界(upper class boundary),计算时用组中值。
When estimating the mean from grouped data, we multiply each midpoint by its frequency, sum those products, and divide by total frequency. This is just an estimate because we don’t know the exact individual values.
根据分组数据估算平均数时,我们用每个组中值乘以其频数,将这些乘积相加后再除以总频数。这只是估算值,因为我们不知道每个个体的确切值。
4. Measures of Central Tendency | 集中趋势的度量
The mean is found by adding all data values and dividing by the number of values. It is the arithmetic average and makes use of every piece of data, which means it is sensitive to extreme values (outliers).
平均数(mean)是将所有数据值相加再除以数据个数得到的。它是算术平均值,利用了每一个数据点,因此对极端值(异常值,outliers)敏感。
The median is the middle value when the data is ordered from smallest to largest. If there is an even number of values, it is the mean of the two central numbers. The median is not affected by outliers, making it a better average for skewed distributions.
中位数(median)是将数据从小到大排序后位于中间的值。如果数据个数是偶数,则为中间两个数的平均数。中位数不受异常值影响,因此在偏态分布中是更好的平均值指标。
The mode is the value that occurs most frequently. A data set can have one mode (unimodal), two modes (bimodal), or more. The mode is the only average suitable for qualitative data.
众数(mode)是出现次数最多的值。一个数据集可以有一个众数(单峰)、两个众数(双峰)或更多。众数是唯一适用于定性数据的平均数。
Remember: Mean uses all data, median marks the centre, mode picks the most popular. For skewed data, the median is often the most honest summary.
记忆口诀:平均数用了所有数据,中位数标记中心,众数选出最热门。对于偏态数据,中位数通常是最诚实的概括。
5. Measures of Spread | 离散程度的度量
The range is simply the largest value minus the smallest value. It gives a quick sense of variability but depends only on two extreme points, so one unusual value can distort it dramatically.
极差(range)就是最大值减去最小值。它让人快速了解数据的变异性,但仅依赖于两个极端点,因此一个异常值就可能严重扭曲它。
The interquartile range (IQR) is the difference between the upper quartile (Q3) and the lower quartile (Q1). It measures the spread of the middle 50% of the data and is resistant to outliers.
四分位距(interquartile range,IQR)是上四分位数(Q3)与下四分位数(Q1)之差。它衡量中间 50% 数据的散布情况,并且抗异常值干扰。
The quartiles split the ordered data into four equal parts. The lower quartile is the median of the lower half, and the upper quartile is the median of the upper half. The IQR is used to construct box-and-whisker plots, showing the five-number summary: minimum, Q1, median, Q3, maximum.
四分位数将有序数据分成四等份。下四分位数是下半部分的中位数,上四分位数是上半部分的中位数。IQR 被用来绘制箱线图(box-and-whisker plot),展示五数概括:最小值、Q1、中位数、Q3、最大值。
The standard deviation measures how far, on average, each data point deviates from the mean. A low standard deviation means data are clustered tightly around the mean; a high one indicates wide scatter.
标准差(standard deviation)衡量每个数据点平均偏离平均数的程度。标准差小意味着数据紧密聚集在平均数周围;标准差大则表示散布很广。
For IGCSE, you are often given the formula: σ = √[Σ(x – x̄)² / n] or using the alternative formula. Use it carefully, always rounding final answers to a sensible degree of accuracy.
在 IGCSE 考试中,你通常会被给到公式:σ = √[Σ(x – x̄)² / n] 或使用其变形公式。仔细计算,最终答案要舍入到合理的精确度。
6. Frequency Distributions and Tables | 频数分布与频数表
A frequency is simply a count of how many times something occurs. A frequency table organizes raw data, showing each value or category alongside its frequency.
频数(frequency)就是某事物出现的次数。频数表(frequency table)是对原始数据的整理,将每个值或类别及其对应的频数列出来。
The cumulative frequency is a running total of frequencies. It helps us quickly find how many data values lie below a certain threshold. Cumulative frequency diagrams are used to estimate the median and quartiles.
累积频数(cumulative frequency)是频数的累加总和。它帮助我们快速找出有多少数据值低于某个界限。累积频数图用于估算中位数和四分位数。
A relative frequency is the proportion of the total frequency that falls in a category. It is calculated as frequency divided by total frequency and can be expressed as a fraction, decimal, or percentage.
相对频数(relative frequency)是某类别频数占总频数的比例。它由频数除以总频数计算得出,可以用分数、小数或百分数表示。
When working with grouped frequency tables for continuous data, always use the mid-interval values as an approximation. Also ensure that the class intervals cover all possible values without overlap.
处理连续数据的分组频数表时,始终使用组中值(mid-interval values)作为近似。还要确保组距覆盖所有可能的值且不重叠。
7. Charts and Graphs | 图表与图形
A bar chart represents discrete or categorical data with rectangular bars. The height (or length) of each bar is proportional to the frequency it represents, and bars should be equally spaced with gaps in between.
条形图(bar chart)用矩形条表示离散或分类数据。每条的高度(或长度)与其代表的频数成比例,各条间应等距且有间隔。
A histogram looks similar but is used for continuous data. Here, there are no gaps between bars, and the area of each bar (not the height alone) is proportional to frequency if class intervals are unequal. The frequency density is calculated as frequency ÷ class width, and this is plotted on the vertical axis.
直方图(histogram)看起来类似,但用于连续数据。其条形之间没有间隔,且如果组距不等,每条的面积(而非仅高度)才与频数成比例。频率密度(frequency density)计算公式为频数 ÷ 组距宽度,并绘制在纵轴上。
A pie chart displays categories as sectors of a circle. Each sector’s angle is 360° multiplied by the category’s fraction of the total. Pie charts are excellent for showing proportions at a glance.
饼图(pie chart)将各类别显示为圆的扇形。每个扇形的角度等于 360° 乘以该类别占总数的比例。饼图非常适合一目了然地展示占比。
A scatter graph (or scatter plot) shows the relationship between two quantitative variables. We look for correlation – positive, negative, or none – and can draw a line of best fit to make predictions.
散点图(scatter graph)展示两个定量变量之间的关系。我们观察相关性(correlation)——正相关、负相关或零相关——并可以画出最佳拟合线进行预测。
Finally, time-series graphs plot data values against time. They reveal trends, seasonal patterns, and help in forecasting future values using moving averages.
最后,时间序列图(time-series graph)将数据值按时间绘制。它们揭示趋势、季节性模式,并借助移动平均数帮助预测未来值。
8. Probability Basics | 概率基础
Probability is a measure of the chance that an event will occur. It is always a number between 0 and 1, where 0 means impossible and 1 means certain. Often it is expressed as a fraction, decimal, or percentage.
概率(probability)是对某个事件发生机会的度量。它总是一个介于 0 和 1 之间的数,0 表示不可能,1 表示必然。通常用分数、小数或百分数表达。
An experiment is any repeatable procedure with a set of possible outcomes. Tossing a coin, rolling a die, and drawing a card are all experiments.
试验(experiment)是任何可重复进行的、有一组可能结果的过程。抛硬币、掷骰子和抽纸牌都是试验。
The sample space is the set of all possible outcomes. For a fair coin, the sample space is {heads, tails}. For a fair six-sided die, it is {1, 2, 3, 4, 5, 6}. Listing the sample space systematically helps calculate probabilities correctly.
样本空间(sample space)是所有可能结果的集合。对于一枚公平的硬币,样本空间是 {正面,反面}。对于公平的六面骰子,它是 {1, 2, 3, 4, 5, 6}。系统性地列出样本空间有助于正确计算概率。
Probability of an event A, written P(A) = number of outcomes in A / total number of outcomes, provided all outcomes are equally likely. The complement of A, written P(not A) = 1 – P(A).
事件 A 的概率,写作 P(A) = A 中结果的数目 / 所有可能结果的数目,前提是所有结果是等可能出现的。A 的补集(complement),P(非 A) = 1 – P(A)。
Two events are mutually exclusive if they cannot both happen at the same time. The addition rule states P(A or B) = P(A) + P(B). If they are not mutually exclusive, subtract P(A and B).
两个事件若不能同时发生,则是互斥事件(mutually exclusive)。加法法则为 P(A 或 B) = P(A) + P(B)。如果它们不互斥,则要减去 P(A 且 B)。
For independent events, the outcome of one has no effect on the other. The multiplication rule states P(A and B) = P(A) × P(B). Tree diagrams can beautifully organise sequential events and show conditional probabilities.
对于独立事件(independent events),一个事件的结果不影响另一个。乘法法则为 P(A 且 B) = P(A) × P(B)。树形图可以优雅地组织序列事件并展示条件概率。
9. Bias and Fairness in Data Collection | 数据收集中的偏差与公平性
Bias is any systematic error that makes a sample unrepresentative of the population. If a survey only gathers responses from a sports club, it is biased towards active people and excludes sedentary individuals.
偏差(bias)是指任何使样本不能代表总体的系统性错误。如果一项调查只从体育俱乐部收集回答,那么它就偏向于喜爱运动的人,而排除了久坐的人群。
A random sample gives every member of the population an equal chance of being selected. Simple random sampling can be achieved using random number tables or a lottery method. It is the gold standard but not always practical.
随机样本(random sample)让总体中每个成员都有同等被选中的机会。简单随机抽样可以使用随机数表或抽签方法实现。它是黄金标准,但并非总是切实可行。
Stratified sampling divides the population into groups (strata) that share a common characteristic, then randomly samples from each group in proportion to its size. This ensures all important subgroups are represented.
分层抽样(stratified sampling)将总体划分为具有共同特征的小组(层),然后按比例从每个组中随机抽样。这确保所有重要子群都能得到代表。
Systematic sampling chooses every k-th person from a list after a random start. It is easy and spreads the sample evenly, but dangers arise if the list has a hidden pattern.
系统抽样(systematic sampling)是从名单上随机起点后,每隔 k 个人选一人。该方法简单且能将样本均匀分布,但如果名单具有隐藏的规律,就会出现风险。
Quota sampling involves setting quotas for certain categories and then selecting respondents non-randomly to fill them. It is cheaper and faster but can easily introduce interviewer bias.
配额抽样(quota sampling)是先为某些类别设定配额,再通过非随机方式选取受访者填满名额。它成本低、速度快,但容易引入调查员偏差。
Always read exam questions carefully: if they ask which sampling method is best, think about whether a sampling frame exists, if you need proportional representation, and how much resource is available.
考试答题时一定要仔细阅读:如果问你哪种抽样方法最好,要思考是否存在抽样框、是否需要比例代表、以及有多少可用资源。
10. Correlation and Regression Terminology | 相关与回归术语
Correlation describes the strength and direction of a linear relationship between two variables. It is not about cause and effect; just because there is a strong correlation between ice cream sales and drowning incidents, it does not mean one causes the other.
相关性(correlation)描述两个变量之间线性关系的强度和方向。它不涉及因果关系;冰淇淋销量与溺水事件之间的强相关,并不意味着一个引起另一个。
The correlation coefficient, usually denoted by r, ranges from -1 to +1. r = +1 means perfect positive linear correlation; r = -1 means perfect negative linear correlation; r = 0 indicates no linear correlation.
相关系数(correlation coefficient)通常记作 r,取值范围从 -1 到 +1。r = +1 表示完全正线性相关;r = -1 表示完全负线性相关;r = 0 表示无线性相关。
A line of best fit (also called a regression line) can be drawn through a scatter plot to model the relationship. It should pass as close as possible to all points, with roughly equal points above and below. The equation y = mx + c can then be used to predict values, but predictions beyond the data range (extrapolation) are unreliable.
最佳拟合线(line of best fit)(又称回归线)可以穿过散点图来模拟这种关系。它应尽可能靠近所有点,点上下的点数大致相等。然后可以用方程 y = mx + c 预测数值,但超出数据范围的预测(外推法)不可靠。
The slope of the line indicates the change in y for a unit increase in x. The intercept is the y-value when x = 0. In IGCSE, you might be asked to draw the line by eye or calculate the equation from given summary statistics.
直线的斜率表示 x 每增加一个单位时 y 的变化量。截距是当 x = 0 时的 y 值。在 IGCSE 中,你可能需要凭眼力画出直线,或根据给定的汇总统计量计算方程。
An outlier is an observation that lies an abnormal distance from other values. Outliers can significantly affect the mean and correlation coefficient, so they need to be identified and considered carefully.
异常值(outlier)是与其他数值距离异常远的观测值。异常值会显著影响平均数和相关系数,因此需要识别并仔细处理。
11. Types of Errors and Reliability | 误差类型与可靠性
Random errors arise from unpredictable fluctuations in readings. Repeating measurements and calculating an average can reduce their impact. Random errors affect precision but not necessarily accuracy.
随机误差(random errors)来自于读数时无法预测的波动。重复测量并计算平均值可以减小其影响。随机误差影响精密度,但不一定影响准确度。
Systematic errors are consistent mistakes, such as using a scale that always reads 0.5 kg too high. They make measurements inaccurate but often still precise. Careful calibration and checking equipment can eliminate these.
系统误差(systematic errors)是一贯性的错误,比如一台秤总是指示偏高 0.5 kg。它们使测量不准确,但往往仍是精密的。仔细校准和检查设备可以消除这些误差。
Reliability refers to consistency – if you measure the same object several times, do you get similar results? Validity and accuracy refer to how close the measurement is to the true value. A survey could be reliable (consistent results) but biased (not truly capturing the population).
信度(reliability)指一致性——如果你多次测量同一个对象,会得到相似的结果吗?效度(validity)和准确度(accuracy)指测量值接近真实值的程度。一项调查可能信度高(结果一致),但存在偏差(未能真实反映总体)。
In statistical reports, always ask: who collected the data? How? When? and for what purpose? This critical eye will help you spot potential bias and evaluate the quality of evidence.
在统计报告中,永远要问:谁收集了这些数据?怎样收集的?何时收集的?以及出于什么目的?这种批判性的眼光能帮助你发现潜在偏差并评估证据的质量。
12. Data Transformations and Index Numbers | 数据转换与指数
Sometimes we adjust data to make it easier to compare. A percentage change is calculated as (new value – old value) / old value × 100%. This allows us to see relative growth regardless of base size.
有时我们会调整数据以便于比较。百分比变化(percentage change)的计算公式为(新值 – 旧值)/ 旧值 × 100%。这使我们能看到相对增长,不受基数大小的影响。
An index number selects a base period and sets its value to 100. Then other values are expressed as a percentage of the base value. For example, if the base price is £2 and the current price is £2.50, the index is (2.50 / 2.00) × 100 = 125. Index numbers simplify comparisons across time.
指数(index number)选择基期并将其数值设定为 100。然后用基期数值的百分数来表示其他值。例如,若基期价格为 £2,当前价格为 £2.50,则指数为 (2.50 / 2.00) × 100 = 125。指数简化了跨时间的比较。
The consumer price index (CPI) and retail price index (RPI) are common real-world examples. They track the cost of a basket of goods and services to measure inflation.
消费者价格指数(CPI)和零售价格指数(RPI)是常见的现实例子。它们追踪一篮子商品和服务的成本,以衡量通货膨胀。
When variables are not linearly related, we may apply a transformation such as taking logarithms or reciprocals to create a more linear pattern for analysis. IGCSE CCEA asks you to recognise these cases but not perform complex transformations.
当变量之间的关系非线性时,我们可以应用转换,比如取对数或取倒数,以创造更线性的模式进行分析。IGCSE CCEA 要求你能识别这些情况,但不要求执行复杂的转换。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导