📚 Statistics Vocabulary & Terminology Quick Reference Guide | 统计词汇术语速记指南
Mastering the language of statistics is the first step towards interpreting data with confidence. In Year 9 CAIE Statistics, a strong grasp of key terms helps you describe, summarise and display information clearly. This guide pairs each English term with its Chinese equivalent and provides simple definitions, memory tips and examples, so you can build your statistical vocabulary quickly and effectively.
掌握统计学的语言是自信解读数据的第一步。在 Year 9 CAIE 统计课程中,牢固掌握关键术语有助于你清晰地描述、总结和呈现信息。本指南将每个英文术语与其对应的中文术语配对,并提供简单的定义、记忆技巧和示例,帮助你快速有效地积累统计词汇。
1. Data Types | 数据类型
Data is divided into two broad types: qualitative and quantitative. Qualitative data describes qualities or categories, while quantitative data measures quantities using numbers. Recognising the type of data helps you choose the right graph and summary statistic.
数据分为两大类:定性数据和定量数据。定性数据描述性质或类别,而定量数据则用数字衡量数量。识别数据类型有助于你选择合适的图表和汇总统计量。
Qualitative data can be nominal (no natural order, like eye colour: blue, brown, green) or ordinal (a natural order, like satisfaction rating: poor, satisfactory, good). Quantitative data is either discrete (countable, whole numbers, e.g. number of students) or continuous (measurable, any value in a range, e.g. height in cm).
定性数据可以是名义数据(无自然顺序,如水颜色:蓝色、棕色、绿色)或有序数据(有自然顺序,如满意度评分:差、满意、好)。定量数据可以是离散型(可数的整数,如学生人数)或连续型(可测量的,区间内的任意值,如整厘米的身高)。
| English Term | 中文术语 | Definition / 定义 |
|---|---|---|
| Qualitative (categorical) | 定性数据(分类数据) | Non-numerical descriptions / 非数值描述 |
| Quantitative | 定量数据 | Numerical measurements / 数值测量 |
| Nominal | 名义数据 | Categories with no order / 无顺序的类别 |
| Ordinal | 有序数据 | Categories with a clear order / 有明确顺序的类别 |
| Discrete | 离散型 | Countable values, usually integers / 可数的值,通常是整数 |
| Continuous | 连续型 | Measurable values, infinite possibilities / 可测量的值,无限可能 |
2. Measures of Central Tendency | 集中趋势的度量
Measures of central tendency describe the ‘centre’ of a data set. The three most common are the mean, median and mode. Each gives a different perspective on a typical value, so it is important to understand when to use each one.
集中趋势的度量描述数据集的“中心”。最常见的是平均数、中位数和众数。它们各自从不同角度反映了典型值,因此了解何时使用每种度量很重要。
The mean is the arithmetic average. The median is the middle value when data is ordered. The mode is the most frequently occurring value. For symmetrical data, the mean and median are close; for skewed data, the median is often more representative.
平均数是算术平均值。中位数是将数据排序后位于中间的值。众数是出现次数最多的值。对于对称数据,平均数和中位数接近;对于偏斜数据,中位数往往更有代表性。
| English Term | 中文术语 | Memory Tip (记忆窍门) |
|---|---|---|
| Mean | 平均数 | ‘Mean’ sounds like ‘average machine’ |
| Median | 中位数 | ‘Median’ reminds us of ‘middle’ like on a road |
| Mode | 众数 | ‘Mode’ starts with ‘mo’ as in ‘most’ |
3. Mean Calculation | 平均数的计算
The mean is found by adding all the data values together and dividing by the number of values. For a population, the mean is denoted by the Greek letter μ (mu); for a sample, it is denoted by x̄ (x-bar). In Year 9, we focus on the sample mean formula.
平均数的计算方法是将所有数据值相加,再除以数据的个数。对于总体,平均数用希腊字母 μ 表示;对于样本,则用 x̄ 表示。在 Year 9 阶段,我们主要学习样本平均数公式。
Mean = (Σx) / n
Here, Σ (capital sigma) means ‘sum of’, x stands for each data value, and n is the total number of values. For example, the scores 5, 7, 8, 10 have a sum of 30 and n = 4, so the mean is 30 ÷ 4 = 7.5.
这里,Σ(大写西格玛)表示“求和”,x 代表每个数据值,n 是数据的总个数。例如,分数 5、7、8、10 的总和为 30,n = 4,因此平均数为 30 ÷ 4 = 7.5。
When data is presented in a frequency table, the mean is calculated using Σ(fx) / Σf, where f is the frequency and x is the corresponding value. Weighted means are an extension of this idea.
当数据以频数表形式呈现时,平均数使用公式 Σ(fx) / Σf 计算,其中 f 是频数,x 是相应的数值。加权平均数就是这一概念的延伸。
4. Median and Mode | 中位数与众数
The median is the value that splits an ordered data set into two equal halves. To find it, arrange the data from smallest to largest and pick the middle one. If there are two middle numbers, the median is their average.
中位数是将排序后的数据集分成相等的两半的数值。要找到中位数,需将数据从小到大排列,并找出正中间的那个。如果中间有两个数,中位数就是它们的平均数。
The mode is the data value that appears most often. A data set can have one mode (unimodal), two modes (bimodal) or more. If no value repeats, there is no mode. The mode is the only measure of central tendency that can be used with nominal qualitative data.
众数是出现次数最多的数据值。数据集可以有一个众数(单峰)、两个众数(双峰)或更多。如果没有值重复,就没有众数。众数是唯一可用于名义定性数据的集中趋势度量。
In Year 9, we also meet the modal class for grouped data, which is the class interval with the highest frequency. Unlike a single mode, the modal class is a range, not a precise value.
在 Year 9 中,我们还会遇到分组数据的众数区间,即频数最高的组距。与单一的众数不同,众数区间是一个范围,而不是精确值。
5. Measures of Spread | 离散程度的度量
Measures of spread describe how data is scattered around the centre. The simplest measure is the range, which is the difference between the largest and smallest values. It gives a quick sense of variability but is sensitive to outliers.
离散程度的度量描述数据在中心周围如何散布。最简单的度量是极差,即最大值与最小值之差。它能快速反映变异性,但对异常值敏感。
Quartiles divide ordered data into four equal parts. The lower quartile (Q₁) is the median of the lower half, and the upper quartile (Q₃) is the median of the upper half. The interquartile range (IQR = Q₃ – Q₁) measures the spread of the middle 50% and is not affected by extreme values.
四分位数将排序后的数据分成四等份。下四分位数(Q₁)是下半部分数据的中位数,上四分位数(Q₃)是上半部分数据的中位数。四分位距(IQR = Q₃ – Q₁)衡量中间 50% 数据的散布情况,不受极端值影响。
For a quick reminder: ‘range is how far from end to end, IQR covers the middle blend’. The IQR is used to identify outliers, usually any value below Q₁ – 1.5×IQR or above Q₃ + 1.5×IQR.
速记窍门:“极差是从头到尾的距离,四分位距覆盖中间区域。” IQR 常用来识别异常值,通常任何小于 Q₁ – 1.5×IQR 或大于 Q₃ + 1.5×IQR 的值都被视为异常值。
6. Frequency Distributions | 频数分布
A frequency distribution shows how often each value or group of values occurs. A frequency table lists data values alongside their frequencies. For large or continuous data, we use grouped frequency tables with class intervals.
频数分布显示每个值或每组值出现的次数。频数表列出数据值及其对应的频数。对于大量或连续的数据,我们使用带组距的分组频数表。
Key terms include class width (the range of each interval), class boundaries (the precise limits between classes) and midpoint (the centre of a class interval). When calculating the mean from grouped data, we use the midpoint of each interval as an estimate.
关键术语包括组距宽度(每个区间的范围)、组界(类别之间的精确界限)和中点(组区间的中心)。当通过分组数据计算平均数时,我们用每个区间的中点作为估计值。
Relative frequency is the proportion of total observations that fall into a class, calculated as frequency ÷ total frequency. It can be expressed as a fraction, decimal or percentage and helps compare distributions of different sizes.
相对频数是观测值落入某个类别的比例,计算公式为 频数 ÷ 总频数。它可以用分数、小数或百分比表示,有助于比较不同大小的分布。
7. Charts and Graphs | 统计图表
Choosing the right visual display is a crucial statistical skill. Bar charts are used for categorical data and have spaces between bars. Histograms are for continuous or grouped data and bars touch each other to show the continuous nature.
选择合适的可视化展示是一项重要的统计技能。条形图用于分类数据,条与条之间有间隙。直方图用于连续或分组数据,条之间互相接触以显示连续性。
Pie charts show proportions as slices of a circle, where the angle of each slice = (category frequency / total frequency) × 360°. They are best for displaying percentages of a whole when there are only a few categories.
饼图以圆的扇形展示比例,每个扇形的角度 = (类别频数 / 总频数)× 360°。当只有少数几个类别时,饼图最适合展示整体中的百分比。
Line graphs and time series plots show how a variable changes over time. Stem-and-leaf diagrams preserve the original data while showing shape, making it easy to find the median and quartiles directly.
折线图和时间序列图展示某个变量如何随时间变化。茎叶图在保留原始数据的同时显示分布形态,便于直接找出中位数和四分位数。
| Graph Type | 图表类型 | Best for / 最适合 |
|---|---|---|
| Bar chart | 条形图 | Categories / 类别比较 |
| Histogram | 直方图 | Grouped continuous data / 分组连续数据 |
| Pie chart | 饼图 | Proportions of a whole / 整体比例 |
| Stem-and-leaf | 茎叶图 | Small data sets, finding quartiles / 小数据集,找四分位数 |
8. Cumulative Frequency | 累积频数
Cumulative frequency is the running total of frequencies up to a certain point. It allows us to construct a cumulative frequency curve (or ogive), which helps estimate the median, quartiles and percentiles directly from the graph.
累积频数是截至某一点的频数累加总和。它可以用来绘制累积频数曲线(或称尖形图),从而直接从图上估计中位数、四分位数和百分位数。
To draw the curve, plot each upper class boundary against its cumulative frequency and join the points with a smooth curve. The median corresponds to the 50th percentile; reading from the graph, it is the value at half the total cumulative frequency.
绘制曲线时,将每个组的上限与对应的累积频数描点,并用平滑曲线连接。中位数对应第 50 百分位数;从图上读取,它是累积频数一半处的值。
The interquartile range can be estimated by reading Q₁ at 25% and Q₃ at 75% of the total cumulative frequency. This non-calculator method is a key Year 9 skill for grouped data.
四分位距可以通过在总累积频数的 25% 和 75% 处读取 Q₁ 和 Q₃ 来估算。这种无需计算器的方法是 Year 9 处理分组数据的关键技能。
9. Scatter Graphs and Correlation | 散点图与相关性
A scatter graph displays the relationship between two quantitative variables. Each point on the graph represents a pair of values (x, y). The pattern of the points helps us identify correlation: positive, negative or none.
散点图展示两个定量变量之间的关系。图上每个点代表一对数值 (x, y)。点的分布模式有助于我们识别相关性:正相关、负相关或无相关。
Positive correlation means as one variable increases, the other tends to increase. Negative correlation means as one increases, the other decreases. No correlation means there is no clear pattern. Correlation does not imply causation; a strong relationship might be due to chance or a third factor.
正相关意味着一个变量增加时,另一个也倾向于增加。负相关意味着一个增加时,另一个减少。无相关表示没有清晰的模式。相关性并不意味着因果关系;强相关可能是偶然或由第三个因素引起。
The line of best fit (trend line) is drawn through the points to summarise the relationship. It should follow the general trend and have roughly equal numbers of points above and below. We use it to estimate unknown values (interpolation within the data range, extrapolation outside the range).
最佳拟合线(趋势线)穿过点群,总结两者关系。它应遵循总体趋势,且线上下的点数大致相等。我们用它来估计未知值(数据范围内的插值,范围外的外推)。
10. Basic Probability Terms | 基本概率术语
Probability measures how likely an event is, on a scale from 0 (impossible) to 1 (certain). It can be expressed as a fraction, decimal or percentage. A probability experiment is a trial, and the outcomes are all possible results.
概率衡量事件发生的可能性,范围从 0(不可能)到 1(一定)。可以用分数、小数或百分比表示。概率实验是一次试验,结果就是所有可能的结果。
The sample space is the set of all possible outcomes. An event is a specific set of outcomes we are interested in. For equally likely outcomes, probability = (number of favourable outcomes) / (total number of outcomes).
样本空间是所有可能结果的集合。事件是我们感兴趣的特定结果集。对于等可能的结果,概率 = (有利结果的数量)/(结果的总数量)。
Key vocabulary includes: mutually exclusive events (cannot happen at the same time), exhaustive events (together cover all outcomes), and relative frequency as an experimental estimate of probability. The experimental probability gets closer to the theoretical probability with more trials (law of large numbers).
关键术语包括:互斥事件(不能同时发生)、穷举事件(共同覆盖所有结果),以及作为实验概率估计值的相对频数。随着试验次数增加,实验概率会趋近理论概率(大数定律)。
The complement of an event A, written A’ or not A, has probability 1 – P(A). For rolling a fair die, P(not rolling a 6) = 1 – 1/6 = 5/6. Remember: ‘Probability from 0 to 1, sample space all outcomes sun.’
事件 A 的补集,写作 A’ 或非 A,其概率为 1 – P(A)。掷一枚均匀骰子,P(未掷出 6) = 1 – 1/6 = 5/6。记住口诀:“概率在 0 到 1 之间,样本空间囊括万全。”
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导