IGCSE Edexcel Statistics: Quick Memorisation Guide to Key Terms | IGCSE Edexcel 统计:关键术语速记指南

📚 IGCSE Edexcel Statistics: Quick Memorisation Guide to Key Terms | IGCSE Edexcel 统计:关键术语速记指南

Mastering statistical vocabulary is the first step to excelling in IGCSE Edexcel Statistics. Many marks are lost simply because a term is confused or a definition is not precise enough. This guide organises all the essential terminology you need to know by topic, with clear twin-language explanations to make memorisation fast and reliable. Use it alongside your revision to ensure you can define each term accurately in the exam.

掌握统计学术语是 IGCSE Edexcel 统计取得高分的第一步。很多分数只是因为概念混淆或定义不够准确而丢失。本指南按主题整理了你需要掌握的所有基本术语,并配有清晰的双语解释,让记忆变得快速而可靠。结合复习使用,确保在考试中能准确给出每个术语的定义。

1. Types of Data | 数据类型

Data can be classified in two main ways: by its nature (qualitative or quantitative) and by the type of values it takes (discrete or continuous). Getting these right underpins every statistical calculation and graph choice.

数据有两种主要分类方式:按性质(定性或定量)以及按取值的类型(离散或连续)。正确区分它们是所有统计计算和图表选择的基础。

Qualitative data describes qualities or categories that cannot be measured numerically, like eye colour or car brands. Quantitative data is numerical and can be measured or counted, such as height, weight, or number of students.

定性数据描述无法用数字测量的属性或类别,如眼睛颜色或汽车品牌。定量数据是数值型的,可以被测量或计数,例如身高、体重或学生人数。

Quantitative data is further split into discrete data, which can only take certain distinct values (often whole numbers) like shoe size or number of goals, and continuous data, which can take any value within a range, like time or temperature.

定量数据又可细分为离散数据(只取特定、互不连续的值,通常是整数)如鞋码或进球数,以及连续数据(可在某一区间内取任意值)如时间或温度。

A quick memory trick: If you can count it using whole numbers without fractions, it is likely discrete; if you need a measuring instrument that gives readings on a continuous scale, it is continuous.

一个快速记忆技巧:如果可以用整数计数而不出现分数,很可能是离散的;如果需要使用连续标尺的测量工具获得读数,则为连续数据。


2. Sampling Methods | 抽样方法

In real-world data collection, we rarely measure an entire population. We select a sample. The method of selection determines whether the sample is representative and free from bias.

在现实数据收集中,我们很少测量整个总体。我们选取一个样本。选择的方法决定了样本是否具有代表性且无偏。

A random sample is one where every member of the population has an equal chance of being chosen, often using a random number generator. This avoids selection bias. In systematic sampling, we choose every nth item from a list after a random starting point. It is quick but can introduce bias if there is a hidden pattern in the list.

随机样本是指总体中每个成员被选中的机会均等,通常使用随机数生成器。这避免了选择偏差。在系统抽样中,从列表中随机起点开始,每隔一定间隔抽取一个项目。它速度快,但如果列表中存在隐藏模式,可能会引入偏差。

Stratified sampling divides the population into distinct groups (strata), then takes a random sample from each group in proportion to its size. This guarantees representation of each subgroup. Quota sampling is similar but non-random: the interviewer picks people within each stratum until a quota is filled, which can lead to bias.

分层抽样将总体分成不同的组(层),然后按比例从每组中随机抽取样本。这能保证每个子群体都有代表。配额抽样类似但非随机:访员在每一层内挑选人员直到满足配额,这可能导致偏差。

Other terms to know: sampling frame (the list from which the sample is drawn), bias (systematic error that distorts results), and pilot survey (a small trial run to test the questionnaire).

其他需要了解的术语:抽样框(用于抽取样本的名单)、偏差(扭曲结果的系统性误差)和试点调查(测试问卷的小范围试运行)。


3. Measures of Central Tendency | 集中趋势的度量

The three main averages summarise a dataset with a single representative value. Understanding when to use each is a favourite exam topic.

三种主要的平均数用单一代表值概括数据集。理解何时使用哪种是考试中的常见考点。

The mean is the sum of all values divided by the number of values. It is affected by outliers. The median is the middle value when data is ordered; it resists outliers. The mode is the most frequent value; it is the only average suitable for qualitative data.

平均数(均值)是所有数值之和除以数值个数,易受异常值影响。中位数是数据排序后的中间值,能抵抗异常值。众数是出现频率最高的值,是唯一适用于定性数据的平均数。

For a grouped frequency table, we estimate the mean using midpoints. The modal class is the class interval with the highest frequency, and the median class is found using cumulative frequency.

对于分组频率表,我们使用组中点来估计平均数。众数组是频率最高的组区间,而中位数组须通过累积频率确定。

Remember: if the data is skewed, the median gives a better sense of ‘typical’ than the mean. For symmetric data, mean ≈ median ≈ mode.

记住:如果数据偏斜,中位数比均值更能体现“典型”。对于对称数据,平均数 ≈ 中位数 ≈ 众数。


4. Measures of Dispersion | 离散程度的度量

Averages alone do not tell the whole story; how spread out the data is matters just as much. Dispersion measures describe variability.

只有平均数并不完整;数据的分散程度同样重要。离散度量描述变异性。

The range is the difference between the largest and smallest values. It is simple but easily distorted by outliers. The interquartile range (IQR) is the difference between the upper quartile (Q3) and lower quartile (Q1), giving the spread of the middle 50% of data. It is robust to outliers.

极差是最小值和最大值之差,简单但容易被异常值扭曲。四分位距(IQR)是上四分位数(Q₃)与下四分位数(Q₁)之差,反映中间50%数据的散布,对异常值稳健。

Variance measures the average squared deviation from the mean, while standard deviation is the square root of variance, returning to the original units. Both reflect how much data points typically differ from the mean.

方差衡量数据与均值之差的平方的平均值,标准差是方差的平方根,恢复到原单位。两者都反映数据点通常偏离均值的程度。

When comparing two datasets, always mention both an average and a measure of spread. A higher standard deviation means data is more spread out and less consistent.

在比较两个数据集时,务必同时提及平均数和离散度量。标准差越大,数据越分散,一致性越低。


5. Data Representation & Charts | 数据表示与图表

IGCSE Edexcel Statistics demands fluency in choosing and interpreting diagrams. Each chart serves a specific purpose.

IGCSE Edexcel 统计要求熟练选择和解读图表。每种图表都有其特定用途。

A bar chart is for categorical data, with gaps between bars. A histogram is for continuous grouped data, with no gaps and area proportional to frequency. In a histogram, frequency density = frequency ÷ class width. A pie chart shows proportions of a whole.

条形图用于分类数据,条间有间隔。直方图用于连续分组数据,条间无间隔,面积与频率成正比。在直方图中,频率密度 = 频率 ÷ 组宽。饼图显示各部分占整体的比例。

A frequency polygon plots midpoints of class intervals against frequency; a cumulative frequency curve (ogive) plots running total against upper class boundaries, useful for finding medians and quartiles. Box plots (box-and-whisker) display the five-number summary: minimum, Q1, median, Q3, maximum. They are excellent for comparing distributions.

频率多边形将组中点与频率绘点连线;累积频率曲线(形似 S 的曲线)以累积频数对上组界,用于求中位数和四分位数。箱形图(箱须图)展示五数综合:最小值、Q₁、中位数、Q₃、最大值,非常适合比较分布。

For bivariate data, a scatter graph is used to identify correlation. Always label axes clearly and choose an appropriate scale.

对于双变量数据,用散点图来识别相关性。务必清晰标注坐标轴并选择合适刻度。


6. Probability Vocabulary | 概率术语

Probability is the foundation of statistical inference. The language must be precise to avoid losing marks.

概率是统计推断的基础。语言必须准确,以免失分。

An experiment is a repeatable process that gives outcomes. The sample space is the set of all possible outcomes. An event is one or more outcomes of interest. For a fair, unbiased scenario, equally likely outcomes have the same chance of occurring.

试验是可重复产生结果的过程。样本空间是所有可能结果的集合。事件是我们感兴趣的一个或多个结果。对于公平、无偏的情形,等可能结果具有相同发生机会。

Probability of an event = number of favourable outcomes ÷ total number of possible outcomes, assuming equally likely. The complement of event A, written A’, has probability P(A’) = 1 − P(A). Mutually exclusive events cannot happen at the same time. Independent events do not influence each other: P(A and B)=P(A)×P(B).

事件概率 = 有利结果数 ÷ 所有可能结果数,假设等可能。事件A的补事件,记作 A’,概率 P(A’) = 1 − P(A)。互斥事件不可能同时发生。独立事件彼此不影响:P(A 且 B) = P(A) × P(B)。

Conditional probability is the chance of event B occurring given that A has already happened, written P(B|A). Use tree diagrams to handle multi-stage experiments and remember to multiply along branches and add across final outcomes for ‘or’ scenarios.

条件概率是在事件A已发生的条件下B发生的概率,记作 P(B|A)。使用树图处理多阶段试验,记得沿分支相乘,对于“或”的情况,将最终结果的概率相加。


7. Correlation and Regression | 相关与回归

When investigating relationships between two variables, correlation measures strength and direction, while regression provides a model for prediction.

在研究两个变量之间的关系时,相关衡量强度和方向,而回归提供预测模型。

Correlation can be positive (as one variable increases, so does the other), negative (one increases, the other decreases), or zero (no linear relationship). It is described as strong, moderate, or weak depending on how closely points follow a straight line. The correlation coefficient r ranges from −1 to 1.

相关可以是正相关(一变量增加,另一也增加)、负相关(一变量增加,另一减少)或零相关(无线性关系)。根据点靠近直线的紧密程度,可描述为强、中等或弱。相关系数 r 介于 −1 到 1 之间。

Regression line is the line of best fit, often drawn by eye or calculated using the least squares method. The equation is y = a + bx, where b is the gradient and a is the y-intercept. Interpolation (predicting within the data range) is reliable; extrapolation (predicting beyond the range) can be unreliable.

回归线是最佳拟合线,常通过目测或用最小二乘法计算得到。方程为 y = a + bx,其中 b 是斜率,a 是 y 轴截距。内插(在数据范围之内预测)可靠;外推(超出范围预测)可能不可靠。

Remember: correlation does not imply causation. Two variables may move together due to a third hidden factor or just by chance.

记住:相关并不表明因果关系。两个变量一起变动可能是由于第三个隐藏因素或纯属偶然。


8. Index Numbers and Time Series | 指数与时间序列

Index numbers compare values over time relative to a base period. They are widely used in economics and business statistics.

指数用于比较数值在一段时间内相对于基期的变化,广泛用于经济和商业统计。

For a simple index number: Index = (value ÷ base year value) × 100. The base year always has an index of 100. A weighted index accounts for the relative importance of items, such as a weighted price index where weights reflect quantities consumed.

简单指数:指数 = (数值 ÷ 基年数值) × 100。基年的指数总是 100。加权指数考虑各项的相对重要性,例如加权价格指数用权重反映消费数量。

A time series is a sequence of data points collected over regular time intervals. It can be decomposed into four components: trend (long-term movement), seasonal variation (regular pattern within a year), cyclical variation (longer-term cycles), and random variation (irregular fluctuations).

时间序列是按固定时间间隔收集的一列数据点。可分解为四个成分:趋势(长期走向)、季节变动(一年内规律模式)、周期变动(更长周期的循环)和随机变动(不规则波动)。

Moving averages smooth out short-term fluctuations to reveal the trend. For quarterly data, a 4-point moving average is typical; for monthly data, a 12-point moving average is used.

移动平均能平滑短期波动以揭示趋势。对于季度数据,通常用4点移动平均;月度数据用12点移动平均。


9. Key Terms for Statistical Enquiry | 统计调查关键术语

The statistical enquiry cycle—hypothesis, data collection, analysis, conclusion—relies on precise language. Examiners expect you to use these terms confidently.

统计调查循环——假设、数据收集、分析、结论——依赖于准确的语言。考官希望你自信地使用这些术语。

A hypothesis is a testable statement about a population, often expressed as a null hypothesis (H₀) and alternative hypothesis (H₁). A variable is any characteristic that can vary, while an attribute is a specific quality.

假设是关于总体的可检验陈述,通常表示为原假设(H₀)和备择假设(H₁)。变量是任何可变特征,属性则是特定品质。

Primary data is collected first-hand by the researcher (e.g., surveys). Secondary data is data that already exists (e.g., government statistics). Raw data is unprocessed data as originally collected. A census attempts to measure every member of a population, whereas a sample surveys only a subset.

第一手数据是研究者直接收集的(如调查问卷)。第二手数据是已经存在的数据(如政府统计数据)。原始数据是最初收集未经处理的数据。普查试图测量总体的每个成员,而样本只调查其中一部分。

Other important terms: outlier (an extreme value that does not fit the pattern), grouping (organising raw data into class intervals), and frequency (the number of times a value or category occurs).

其他重要术语:异常值(不符合模式的极端值)、分组(将原始数据归类到组区间)、频数(某个值或类别出现的次数)。


10. Quick Memorisation Table | 速记对比表

The table below provides a side-by-side review of frequently confused concepts. Use it for last-minute revision.

下表提供了易混淆概念的并排复习,适合考前快速浏览。

Term (English) 术语 (中文) Key Distinction
Population vs Sample 总体 vs 样本 Population: all members; Sample: a subset
Parameter vs Statistic 参数 vs 统计量 Parameter: describes population; Statistic: describes sample
Discrete vs Continuous 离散 vs 连续 Discrete: countable (e.g. 0,1,2…); Continuous: measurable (e.g. 1.67 m)
Mean vs Median 平均数 vs 中位数 Mean uses all values; Median uses middle rank
Range vs IQR 极差 vs 四分位距 Range: max−min; IQR: Q₃−Q₁, spreads middle 50%
Histogram vs Bar Chart 直方图 vs 条形图 Histogram: area = frequency; Bar chart: height = frequency, gaps present
Correlation vs Causation 相关 vs 因果 Correlation describes association; Causation means one variable influences the other
Trend vs Seasonal Variation 趋势 vs 季节变动 Trend: general direction; Seasonal: regular periodic pattern
Primary vs Secondary Data 第一手数据 vs 第二手数据 Primary: collected by user; Secondary: already available

Test yourself: cover the English or Chinese column and try to recall the definition. Write short, accurate sentences as you would in an exam.

自测方法:遮挡英文或中文列,尝试回忆定义。像在考试中一样写出简短而准确的句子。


Published by TutorHao | Statistics Revision Series | aleveler.com

Find Edexcel IGCSE Statistics Textbooks on eBay UK

New, used and second-hand copies of textbooks and revision guides are often much cheaper than retail — check current listings and prices before you buy.

Browse on eBay UK →

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading