📚 SQA Statistics Year 9: Vocabulary & Terminology Quick-Reference Guide | SQA 统计 Year 9:词汇术语速记指南
Welcome to your one-stop glossary for the SQA Year 9 Statistics course. Mastering the language of data is the first step to interpreting graphs, calculating averages, and understanding probability. This guide pairs every English term with its Chinese equivalent, clear definitions, and real examples to help you build confidence fast. Keep it bookmarked for quick revision before classroom tests or end-of-year assessments.
欢迎来到 SQA Year 9 统计课程的一站式词汇表。掌握数据的语言是解读图表、计算平均值和理解概率的第一步。本指南将每个英文术语与其中文对应词、清晰的定义和真实示例配对,帮助你快速建立信心。请收藏本页,以便在课堂测验或年终评估前快速复习。
1. Population and Sample | 总体与样本
A population is the entire group of individuals or items that you want to study. For example, all S2 pupils in Scotland would be the population if you were investigating lunch preferences across the country.
总体 是指你想要研究的全部个体或项目组成的完整集合。例如,如果你在调查全苏格兰的午餐偏好,那么所有 S2 学生就是总体。
A sample is a smaller group selected from the population to represent it. Sampling saves time and money, but the sample must be chosen carefully to avoid bias. A sample of 100 randomly selected S2 pupils from five different schools could represent the whole country’s preference reasonably well.
样本 是从总体中选出的一个较小的群体,用以代表总体。抽样可以节省时间和金钱,但样本必须谨慎选择以避免偏差。从五所不同学校随机抽取的 100 名 S2 学生样本,就可以相当好地代表整个国家的偏好。
2. Variable Types: Qualitative and Quantitative | 变量类型:定性变量与定量变量
A qualitative variable describes a quality or category that cannot be measured with numbers in a meaningful way. Examples include eye colour (blue, brown, green), favourite subject (Maths, English, Art), or type of pet (dog, cat, hamster).
定性变量 描述的是无法用数字进行有意义测量的属性或类别。例子包括眼睛颜色(蓝色、棕色、绿色)、最喜欢的科目(数学、英语、美术)或宠物类型(狗、猫、仓鼠)。
A quantitative variable is numerical and can be counted or measured. These split further into discrete and continuous. The number of siblings you have is discrete because you count whole numbers (0, 1, 2, 3…). Your height in centimetres (e.g. 152.4 cm) is continuous because it can take any value within a range.
定量变量 是数值型的,可以被计数或测量。它们进一步分为离散型和连续型。你拥有的兄弟姐妹数量是离散的,因为你用整数来计数(0, 1, 2, 3…)。你的身高(以厘米为单位,例如 152.4 cm)是连续的,因为它可以取某个范围内的任何值。
3. Measures of Central Tendency: Mean, Median, Mode | 集中趋势的度量:平均数、中位数、众数
The mean is calculated by adding all values together and dividing by how many values there are. If five friends have pocket money of £5, £7, £8, £6, and £24, the mean is (£5 + £7 + £8 + £6 + £24) ÷ 5 = £10. Notice how one unusually high value pulls the mean upwards.
平均数 的计算方法是将所有数值相加,然后除以数值的个数。如果五个朋友的零花钱分别是 £5、£7、£8、£6 和 £24,那么平均数是 (£5 + £7 + £8 + £6 + £24) ÷ 5 = £10。请注意,一个异常高的值会拉高平均数。
The median is the middle value when the data is placed in order. For the same dataset (5, 6, 7, 8, 24), the median is £7. The median is less affected by extreme values, which makes it useful for skewed data like house prices or incomes.
中位数 是将数据排序后位于中间的值。对于同一数据集(5, 6, 7, 8, 24),中位数是 £7。中位数受极端值的影响较小,这使得它对于房价或收入等偏态数据非常有用。
The mode is simply the most frequently occurring value. In a survey of favourite colours, if ‘blue’ appears 15 times and all other colours fewer, blue is the mode. A dataset can have no mode, one mode (unimodal), or more than one mode (bimodal).
众数 就是出现频率最高的值。在一项关于最喜欢的颜色的调查中,如果“蓝色”出现了 15 次而其他颜色出现次数都更少,那么蓝色就是众数。一个数据集可以没有众数,有一个众数(单峰),或者有多个众数(双峰)。
4. Measures of Spread: Range and Interquartile Range | 离散度的度量:极差与四分位距
Range measures how spread out the data is from lowest to highest. It is found by subtracting the smallest value from the largest. If the highest test score is 95% and the lowest is 42%, the range is 95 − 42 = 53 percentage points. A larger range indicates greater variability.
极差 衡量数据从最低到最高的分散程度。它的计算方法是最大值减去最小值。如果最高测试分数是 95%,最低是 42%,那么极差是 95 − 42 = 53 个百分点。极差越大表明变异性越大。
The interquartile range (IQR) looks at the spread of the middle 50% of data, ignoring the top and bottom quarters. First find the lower quartile (Q₁, the median of the lower half) and the upper quartile (Q₃, the median of the upper half). Then IQR = Q₃ − Q₁. This measure resists the influence of outliers and is preferred when comparing box plots.
四分位距 (IQR) 关注的是中间 50% 数据的分散程度,忽略顶部和底部的四分之一。首先找到下四分位数(Q₁,下半部分的中位数)和上四分位数(Q₃,上半部分的中位数)。然后 IQR = Q₃ − Q₁。这个度量能抵抗异常值的影响,在比较箱线图时更为可取。
5. Frequency Distributions and Grouped Data | 频率分布与分组数据
A frequency tells you how many times something occurs. A frequency distribution organises data into a table that lists each category or group alongside its count. For continuous data, we often use class intervals like 0 ≤ h < 10, 10 ≤ h < 20, and so on, where h represents height in centimetres.
频率 告诉你某事物发生了多少次。频率分布 将数据组织成一个表格,列出每个类别或组别及其计数。对于连续数据,我们通常使用组距,例如 0 ≤ h < 10、10 ≤ h < 20 等,其中 h 代表身高(厘米)。
When working with grouped data, we cannot find the exact mean because individual values are unknown. Instead we estimate the mean by using the midpoint of each interval. Multiply each midpoint by its frequency, sum these products, and divide by the total frequency. This gives a reasonable estimate even without raw data.
在处理分组数据时,由于不知道具体的个体数值,我们无法求得精确的平均数。我们可以通过使用每个组距的中点值来估算平均数。将每个中点值乘以其频率,将这些乘积相加,然后除以总频率。即使在缺少原始数据的情况下,这也能给出一个合理的估计值。
6. Charts for Displaying Data: Bar, Pie, and Line Graphs | 数据展示图表:条形图、饼图与折线图
A bar chart uses rectangular bars with heights proportional to the frequency. Bars are separated by gaps for categorical data. Use bar charts for comparing discrete categories like transport types (walk, bus, cycle) or favourite crisp flavours.
条形图 使用矩形条,其高度与频率成正比。对于分类数据,条与条之间留有间隙。使用条形图来比较离散的类别,例如交通方式(步行、公交、自行车)或最喜欢的薯片口味。
A pie chart shows proportions as slices of a circle. The angle of each slice equals (frequency ÷ total) × 360°. Pie charts are excellent for displaying relative shares, such as how a monthly budget is divided into rent, food, and entertainment, but they become difficult to read when there are too many small slices.
饼图 以圆形切片的方式展示比例。每个扇形的角度等于(频率 ÷ 总频率)× 360°。饼图非常适合展示相对份额,例如月度预算如何划分为房租、食物和娱乐,但当细小扇形过多时会变得难以阅读。
Line graphs connect data points with straight lines and are mainly used to show trends over time. Plotting temperature at noon each day for a week creates a line graph that reveals whether it is warming up or cooling down.
折线图 用直线连接数据点,主要用于显示随时间变化的趋势。绘制一周中每天正午的温度就能创建一幅折线图,揭示天气是在变暖还是变冷。
7. Stem-and-Leaf and Scatter Diagrams | 茎叶图与散点图
A stem-and-leaf diagram keeps the original data visible while showing the shape of the distribution. The ‘stem’ represents the leading digit(s) and the ‘leaf’ shows the final digit. For numbers 23, 25, 31, 34, the stem 2 has leaves 3 and 5; stem 3 has leaves 1 and 4. Always include a key, e.g. ‘2|3 means 23’.
茎叶图 在展示分布形状的同时保留原始数据。’茎’代表前导数字,’叶’代表末位数字。对于数字 23、25、31、34,茎 2 有叶 3 和 5;茎 3 有叶 1 和 4。务必附带图例,例如’2|3 表示 23’。
A scatter diagram plots paired numerical data on an x‑y grid to examine whether a relationship exists. Each point represents one individual or item. Plotting hours of revision against test scores might show a positive correlation: as revision hours increase, scores tend to rise. If points form a downward trend, the correlation is negative.
散点图 将成对的数值数据绘制在 x-y 网格上,以检验是否存在某种关系。每个点代表一个个体或项目。将复习小时数与测试分数绘制成图,可能会显示出正相关:随着复习时间增加,分数往往会上升。如果各点形成下降趋势,则相关性为负。
8. Correlation and the Line of Best Fit | 相关性与最佳拟合线
Correlation describes the strength and direction of a linear relationship between two variables. It can be positive (both increase together), negative (one increases while the other decreases), or zero (no linear pattern). Remember: correlation does not imply causation. Ice cream sales and drowning incidents both rise in summer, but buying ice cream does not cause drowning.
相关性 描述了两个变量之间线性关系的强度和方向。它可以是正相关(两者同时增加)、负相关(一个增加而另一个减少)或零相关(没有线性模式)。请记住:相关性并不意味着因果关系。冰淇淋销量和溺水事件在夏季都会上升,但购买冰淇淋并不会导致溺水。
The line of best fit is a straight line drawn through a scatter plot to model the trend. It should pass as close to as many points as possible, with roughly equal numbers of points above and below the line. You can use the line to estimate values: reading within the data range is interpolation (reliable); reading beyond it is extrapolation (less trustworthy).
最佳拟合线 是一条穿过散点图的直线,用于模拟趋势。它应该尽可能靠近尽可能多的点,线上方和线下方的点数大致相等。你可以使用这条线来估计数值:在数据范围内读取叫内插(可靠);超出范围读取叫外推(可信度较低)。
9. Probability Vocabulary: Events, Outcomes, and Sample Space | 概率词汇:事件、结果与样本空间
An outcome is a single possible result of an experiment. When you roll a fair six‑sided die, the outcomes are 1, 2, 3, 4, 5, 6. The sample space is the set of all possible outcomes, often listed inside curly brackets: S = {1, 2, 3, 4, 5, 6}.
结果 是实验的一个可能结果。当你掷一枚公平的六面骰子时,结果是 1、2、3、4、5、6。样本空间 是所有可能结果的集合,通常列在大括号内:S = {1, 2, 3, 4, 5, 6}。
An event is a collection of one or more outcomes. ‘Rolling an even number’ is an event that includes outcomes 2, 4, and 6. The probability of an event = (number of favourable outcomes) ÷ (total number of outcomes), provided all outcomes are equally likely. Probability values always lie between 0 (impossible) and 1 (certain).
事件 是一个或多个结果的集合。’掷出偶数’是一个事件,包含结果 2、4 和 6。事件的概率 =(有利结果的数量)÷(所有结果的总数),前提是所有结果都是等可能的。概率值总是在 0(不可能)和 1(必然)之间。
10. Experimental and Theoretical Probability | 实验概率与理论概率
Theoretical probability is what we expect to happen based on equally likely outcomes without actually doing the experiment. The theoretical probability of getting tails on a fair coin is ½ or 0.5. It rests on the assumption that the coin is perfectly balanced.
理论概率 是我们基于等可能结果而期望发生的事情,无需实际进行实验。一枚公平硬币掷出反面的理论概率是 ½ 或 0.5。它建立在硬币完全平衡的假设之上。
Experimental probability (or relative frequency) comes from actually performing the trial. If you toss a coin 100 times and get tails 47 times, the experimental probability is 47/100 = 0.47. With more trials, experimental probability usually gets closer to theoretical probability – this idea is called the Law of Large Numbers.
实验概率(或称相对频率)来自实际执行试验。如果你抛硬币 100 次,得到 47 次反面,那么实验概率是 47/100 = 0.47。随着试验次数增多,实验概率通常会趋近理论概率——这个思想被称为大数定律。
11. Bias, Randomness, and Fairness | 偏差、随机性与公平性
Bias occurs when a sample or experiment systematically favours certain outcomes. A survey about internet usage conducted only via an online form excludes people without internet access, creating a biased result. In statistics, we must actively identify and minimise sampling bias.
偏差 发生在样本或实验系统性地偏向某些结果时。一项仅通过网络表格进行的互联网使用调查会将无法上网的人排除在外,从而产生有偏差的结果。在统计中,我们必须主动识别并尽量减少抽样偏差。
Randomness means each member of the population has an equal chance of being selected, or each outcome has an equal chance in a probability experiment. A fair die or coin is one where all outcomes are equally likely. Random number generators or drawing names from a hat help achieve simple random sampling in classroom investigations.
随机性 意味着总体中的每个成员都有同等机会被选中,或者在概率实验中每个结果都有同等机会。一个公平的骰子或硬币是指所有结果等可能出现的物品。随机数生成器或从帽子里抽名字有助于在课堂调查中实现简单随机抽样。
12. Key Statistical Notation and Symbols | 关键统计符号与记法
Understanding common symbols speeds up your reading of questions and mark schemes. The symbol n usually stands for the number of data values (sample size). The Greek letter Σ (sigma) means ‘sum of’. So Σx means ‘sum all the x values’. x̄ (x‑bar) represents the sample mean.
理解常见符号能加快你阅读题目和评分方案的速度。符号 n 通常代表数据值的数量(样本大小)。希腊字母 Σ (sigma) 表示’…之和’。所以 Σx 表示’将所有 x 值求和’。x̄ (x-bar) 代表样本平均数。
For quartiles, Q₁ is the lower quartile, Q₂ is the median (second quartile), and Q₃ is the upper quartile. The interquartile range is often abbreviated as IQR. For probability, P(A) reads as ‘the probability of event A occurring’. A’ (A prime or A complement) means ‘not A’.
对于四分位数,Q₁ 是下四分位数,Q₂ 是中位数(第二四分位数),Q₃ 是上四分位数。四分位距通常缩写为 IQR。在概率中,P(A) 读作’事件 A 发生的概率’。A’(A撇 或 A的补集)表示’非A’。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导