Year 11 CIE Statistics: Essential Vocabulary Quick-Reference Guide | CIE 十一年级统计:核心术语速记指南

📚 Year 11 CIE Statistics: Essential Vocabulary Quick-Reference Guide | CIE 十一年级统计:核心术语速记指南

Mastering statistical terminology is the foundation for success in CIE IGCSE Statistics. Being able to recall and correctly apply terms like ‘stratified sampling’ or ‘interquartile range’ not only boosts your confidence but directly improves your marks on written and data‑response questions. This quick‑reference guide breaks down the key vocabulary into themed sections, each with bilingual explanations and memory hooks to help Year 11 students learn faster and remember longer.

掌握统计学术语是在 CIE IGCSE 统计考试中取得优异成绩的基石。能够迅速回忆起并准确运用“分层抽样”或“四分位距”等词汇,不仅能提升你的自信,还能直接提高你在书面题和数据分析题上的得分。这份速记指南将核心术语按主题分节,提供中英双语解释和记忆技巧,帮助十一年级的学生学得更快、记得更牢。


1. Data Types and Collection | 数据类型与收集

Population: The complete collection of all individuals, items, or data points that you wish to study. Think of it as the ‘whole group’.

总体:你所希望研究的所有个体、项目或数据点的完整集合。可以记忆为“全体”。

Sample: A subset of the population selected for the actual investigation. A good sample is representative, meaning it mirrors the characteristics of the population.

样本:从总体中选出用于实际研究的子集。一个好的样本应当具有代表性,即能反映出总体的特征。

Discrete data: Numerical data that can only take specific, separate values. Examples include the number of students in a class or the score on a die. Obtain it by counting.

离散数据:只能取特定、分离数值的数据。例如班级学生人数或骰子点数。通过计数获得。记忆:“可数、跳变”。

Continuous data: Numerical data that can take any value within a given range. Height, mass, and time are continuous. Obtain it by measuring.

连续数据:在给定范围内可以取任意值的数据。身高、质量和时间都是连续的。通过测量获得。记忆:“可测、连续”。

Categorical (qualitative) data: Data that describes qualities or categories. Eye colour, type of car, or favourite subject are categorical. This can be further grouped into nominal (no order, e.g. colours) and ordinal (natural order, e.g. exam grades).

分类(定性)数据:描述性质或类别的数据。如眼睛颜色、车辆类型或最喜欢的科目。可进一步分为名义分类(无顺序,如颜色)和有序分类(有自然顺序,如考试等级)。记忆:“性质而非数量”。

Primary data: Data collected directly by the researcher for the specific purpose at hand, e.g. through surveys or experiments.

一手数据:研究者为当前特定目的直接收集的数据,例如通过问卷或实验获得的数据。

Secondary data: Data that has already been collected by someone else for another purpose, such as government statistics or historical records.

二手数据:已被他人为其他目的收集好的数据,如政府统计数据或历史记录。记忆:“别人家的老数据”。


2. Sampling Techniques | 抽样方法

Simple random sampling: Every member of the population has an equal chance of being selected. Methods include drawing names from a hat or using a random number generator. It minimises bias.

简单随机抽样:总体中每个成员被选中的机会均等,例如抽签或使用随机数生成器。能最大限度减少偏差。

Stratified sampling: The population is divided into distinct subgroups (strata) sharing similar characteristics, and a random sample is taken from each stratum in proportion to its size. This guarantees representation of all key groups.

分层抽样:将总体按相似特征分成若干层,然后按各层大小比例从每层中随机抽取样本。确保所有重要群体都被代表。

Systematic sampling: Individuals are selected from a list at regular intervals after a random start. For example, choosing every 10th name on a register. It is simple but can introduce bias if a hidden pattern exists.

系统抽样:在随机起始点之后,按照固定间隔从名单中选取个体。比如从花名册上每隔9人抽取一人。操作简单,但如果存在隐藏规律可能引入偏差。

Quota sampling: The researcher selects a predetermined number of individuals with specific characteristics. It is non‑random and can easily lead to interviewer bias, but it is cheap and quick.

配额抽样:研究者选取具有特定特征的预定数量的个体。这是一种非随机方法,容易产生访问者偏差,但成本低、速度快。

Opportunity (convenience) sampling: Selecting individuals who are easily available, e.g. the first 20 students you meet in the corridor. It is the least reliable because the sample is unlikely to be representative.

方便抽样:选取最容易获得的个体,如你在走廊遇到的前20名学生。可靠性最低,因为样本不太可能具有代表性。记忆:“图方便,不精准”。


3. Measures of Central Tendency | 集中趋势的量度

Mean: The arithmetic average, calculated by summing all values and dividing by the number of values. For a data set x₁, x₂, … xₙ, the formula is

平均数:算术平均值,将所有数值相加再除以数值的个数。数据集 x₁, x₂, … xₙ 的公式为

x̄ = Σx ÷ n

The mean uses every data point, so it is sensitive to outliers. Memorise as “sum divided by count”.

均值使用了每一个数据点,因此对异常值敏感。记忆为“总和除以个数”。

Median: The middle value when the data set is ordered from smallest to largest. If there is an even number of observations, the median is the average of the two middle numbers. It is unaffected by extreme values.

中位数:将数据集从小到大排序后位于中间位置的数值。如果数据个数为偶数,则为中间两个数的平均值。不受极端值影响。记忆:“排好队,站中间”。

Mode: The value that occurs most frequently. A data set can have one mode (unimodal), two modes (bimodal), or more. The mode is the only measure of central tendency applicable to non‑numerical categorical data.

众数:出现频率最高的值。一组数据可以有一个众数、两个众数或多个众数。众数是唯一可用于非数值分类数据的集中趋势量度。记忆:“出现最多是众数”。

Weighted mean: Used when some values contribute more than others. Each value is multiplied by its weight, and the sum of these products is divided by the total weight.

加权平均数:当某些数值的重要性不同时使用。每个数值乘以其权重,所有乘积之和再除以总权重。记忆:“重要程度不同,乘上权重再平均”。


4. Measures of Dispersion | 离散程度的量度

Range: The difference between the largest and smallest values. It is quick to compute but heavily affected by outliers.

极差:最大值与最小值之差。计算迅速但极易受异常值影响。

Quartiles and interquartile range (IQR): The lower quartile (Q₁) is the median of the lower half of the data, the upper quartile (Q₃) is the median of the upper half. IQR = Q₃ − Q₁. The IQR measures the spread of the middle 50% of data and is robust against outliers.

四分位数与四分位距 (IQR):下四分位数 (Q₁) 是数据下半部分的中位数,上四分位数 (Q₃) 是数据上半部分的中位数。IQR = Q₃ − Q₁。它衡量中间50%数据的离散程度,对异常值稳健。记忆:“箱形图的身体宽度”。

Percentiles: The k‑th percentile is the value below which k% of the observations fall. The median is the 50th percentile.

百分位数:第 k 百分位数是有 k% 的观测值落在其下方的数值。中位数就是第50百分位数。记忆:“百份排名”。

Variance and standard deviation: Variance measures the average squared difference from the mean. For a sample, the variance s² is

s² = Σ(x − x̄)² ÷ (n − 1)

方差与标准差:方差衡量各数据与均值之差的平方的平均值。样本方差 s² 为

s² = Σ(x − x̄)² ÷ (n − 1)

The standard deviation s is the square root of the variance. It has the same units as the original data and is the most widely used measure of spread. Reminder: “SD = √variance”.

标准差 s 是方差的平方根,单位和原始数据相同,是最常用的离散量度。记忆:“方差开平方,单位回原家”。


5. Data Representation Charts | 数据图示

Bar chart: Uses rectangular bars of equal width whose lengths represent frequency or quantity. Bars are separated to show distinct categories. Height of bar ∝ frequency.

条形图:用宽度相等的矩形条表示频数或数量,条与条之间有间隙,用于展示不同类别。条形高度与频数成正比。

Pie chart: A circular chart divided into sectors, each sector angle proportional to the frequency of the category. Angle = (frequency ÷ total) × 360°.

饼图:将圆形分成若干个扇形,每个扇形的圆心角与对应类别的频数成比例。角度 = (频数 ÷ 总数) × 360°。

Histogram: A diagram for continuous data in which the area of each bar is proportional to the frequency. Bars touch, and frequency density = frequency ÷ class width is used when class widths are unequal.

直方图:用于连续数据的图形,每块矩形面积与频数成比例。矩形之间没有空隙。当组距不等时,需使用频率密度 = 频数 ÷ 组距。记忆:“面积代表频数”。

Frequency polygon: A line graph formed by joining the mid‑points of the tops of histogram bars (or class midpoints at their respective frequencies). Often used to compare distributions.

频率多边形:将直方图各矩形上端中点(或各组中值对应频数点)连接而成的折线图。常用于比较分布。

Stem‑and‑leaf diagram: A display that retains original data values while showing the shape of the distribution. The ‘stem’ represents leading digits, and the ‘leaf’ represents the final digit.

茎叶图:既能保留原始数据值又能展示分布形态的图示。“茎”代表前导数位,“叶”代表最后一位数字。记忆:“数据枝叶看得见”。


6. Cumulative Frequency and Box Plots | 累积频率与箱形图

Cumulative frequency: The running total of frequencies up to the end of a given interval. A cumulative frequency curve (ogive) is used to estimate medians, quartiles, and percentiles.

累积频率:截至某个区间末尾的频数累计总和。累积频率曲线(形如“S”形)用于估算中位数、四分位数和百分位数。

Box‑and‑whisker plot (box plot): A graphical summary of the five‑number summary: minimum, Q₁, median, Q₃, maximum. The box shows the IQR, and whiskers extend to the minimum and maximum (or to 1.5 × IQR to identify outliers). Outliers can be plotted as individual points.

箱形图(盒须图):五数概括的图形展示:最小值、Q₁、中位数、Q₃、最大值。箱体表示 IQR,触须延伸至最小值和最大值(或延伸至 1.5 × IQR 识别异常值)。异常值可单独绘制。记忆:“五数一箱,须看两端”。

Interpolation from cumulative frequency graphs: To find a percentile, locate the appropriate position on the cumulative frequency axis, draw a horizontal line to the curve, and then a vertical line down to the data axis.

从累积频率曲线插值:先在累积频率轴上找到对应百分比位置,画水平线交于曲线,再垂直向下读数据轴。记忆:“一横一竖,插出分位”。


7. Probability Fundamentals | 概率基础

Probability of an event A, written P(A): If all outcomes are equally likely, P(A) = number of favourable outcomes ÷ total number of outcomes. Probability always lies between 0 (impossible) and 1 (certain).

事件 A 的概率,记作 P(A):若所有结果等可能,则 P(A) = 有利结果数 ÷ 所有可能结果总数。概率总是在 0(不可能)和 1(必然)之间。

Sample space: The set of all possible outcomes of an experiment. Represent it using a list, a two‑way table, or a tree diagram.

样本空间:试验所有可能结果的集合。可用列表、双向表或树状图表示。

Mutually exclusive events: Events that cannot happen at the same time. If A and B are mutually exclusive, P(A or B) = P(A) + P(B).

互斥事件:不能同时发生的事件。若 A 和 B 互斥,则 P(A 或 B) = P(A) + P(B)。记忆:“有你无我”。

Independent events: The outcome of one event does not affect the probability of the other. For independent events A and B, P(A and B) = P(A) × P(B). On a tree diagram, probabilities along the branches remain unchanged.

独立事件:一个事件的发生不影响另一事件发生的概率。对于独立事件 A 和 B,P(A 且 B) = P(A) × P(B)。在树状图上,各分支概率保持不变。记忆:“各走各路,概率相乘”。

Conditional probability: The probability of event A occurring given that B has already occurred, denoted P(A|B) = P(A and B) ÷ P(B). Often calculated from a two‑way table or a restricted sample space.

条件概率:在事件 B 已发生的条件下,事件 A 发生的概率,记作 P(A|B) = P(A 且 B) ÷ P(B)。常通过双向表或缩小的样本空间计算。记忆:“给定条件,缩小范围再算”。


8. Correlation and Scatter Graphs | 相关与散点图

Scatter diagram (scatter plot): A graph showing paired data points for two variables, used to investigate the relationship between them. One variable on each axis.

散点图:展示两个变量成对数据点的图形,用于探究变量间的关系。两轴各代表一个变量。

Correlation: Describes the strength and direction of the linear relationship between two variables. Positive correlation: as one variable increases, the other tends to increase. Negative correlation: as one variable increases, the other tends to decrease. No correlation: no apparent linear pattern.

相关性:描述两个变量之间线性关系的强度和方向。正相关:一个变量增大,另一个也趋向增大。负相关:一个变量增大,另一个趋向减小。零相关:无明显线性模式。

Line of best fit (regression line): A straight line drawn through a scatter plot that best represents the trend of the data. It should pass close to as many points as possible, with roughly equal numbers of points above and below the line.

最佳拟合线(回归线):在散点图中画出的、最能代表数据趋势的直线。应尽可能贴近多数点,且线上下方点数大致相等。记忆:“穿云而过,左右平衡”。

Causal relationship: Remember that correlation does not imply causation. Two variables may show a strong correlation simply because they are both influenced by a third hidden factor.

因果关系:记住相关性不意味着因果关系。两个变量可能只是因同时受第三个隐藏因素影响而呈现强相关。记忆:“相关非因果”。


9. Time Series and Moving Averages | 时间序列与移动平均

Time series graph: A graph in which data points are plotted in time order, with time on the horizontal axis. It reveals overall trends, seasonal variations, and irregular fluctuations.

时间序列图:按时间顺序绘制数据点的图形,时间放在横轴。可展示整体趋势、季节性波动和不规则变动。

Trend: The long‑term general movement of the data over time, ignoring short‑term ups and downs.

趋势:数据随时间变化的长期总体走向,忽略短期起伏。记忆:“大方向”。

Seasonal variation: A repeating pattern that occurs at fixed time intervals, e.g. higher ice‑cream sales in summer every year.

季节性变动:在固定时间间隔内重复出现的规律性模式,如冰淇淋销量每年夏季都更高。记忆:“日历规律”。

Moving average: A series of averages calculated from successive groups of data to smooth out short‑term fluctuations and highlight the trend. For quarterly data, a four‑point moving average is common.

移动平均:依次计算连续若干组数据的平均值,以平滑短期波动、突显趋势。对于季度数据,常使用四点移动平均。记忆:“移动窗口,平滑曲线”。


10. Interpolation and Extrapolation | 内插与外推

Interpolation: Estimating a value within the range of the original data. When using a line of best fit, interpolation is relatively reliable because it is based on known data points.

内插法:在原始数据范围内进行估值。利用最佳拟合线进行内插相对可靠,因为它基于已知数据点。记忆:“域内预测,比较靠谱”。

Extrapolation: Estimating a value outside the range of the original data by extending the line of best fit. Extrapolation can be unreliable because the established trend may not continue beyond the observed data.

外推法:将最佳拟合线延伸,对原始数据范围外的值进行估算。外推可能不可靠,因为趋势在观测范围外未必保持。记忆:“域外预测,风险自担”。

Danger of extrapolation: Always treat extrapolated values with caution and state that they are estimates only. Real‑world changes may break the pattern.

外推风险:务必谨慎对待外推值,并声明这只是估算。现实世界的变化可能打破原有模式。考试中常常要求指出“趋势未必持续”。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading