📚 Year 8 Edexcel Statistics: A Quick Guide to Vocabulary and Terminology | 词汇术语速记指南
Statistics is all about collecting, presenting, and interpreting data. To speak the language of statistics, you need to learn key words and phrases. This guide will help you quickly remember essential statistical vocabulary for the Year 8 Edexcel course, with clear definitions and examples in both English and Chinese.
统计学的核心是收集、展示和解读数据。要掌握统计这门语言,你需要学会关键术语。本指南将帮助你快速记住 Year 8 Edexcel 课程中的核心统计词汇,并配有清晰的定义和中英双语示例。
1. Types of Data | 数据类型
Data can be sorted into different types depending on what the values represent and how they can be measured. The first major split is between categorical (qualitative) and numerical (quantitative) data. Categorical data describes qualities or groups that cannot be measured with numbers, like favourite colour or type of pet. Numerical data is made up of numbers that you can count or measure, such as test scores or heights.
数据可以根据其代表的含义和测量方式分为不同类型。最主要的划分是分类(定性)数据和数值(定量)数据。分类数据描述不能用数字衡量的品质或类别,例如最喜欢的颜色或宠物种类。数值数据由可以计数或测量的数字组成,例如考试分数或身高。
Numerical data can be further divided into discrete and continuous. Discrete data can only take specific, separate values – often whole numbers – and there are gaps between the values. For example, the number of siblings a person has can be 0, 1, 2, 3… but never 2.5. Continuous data can take any value within a given range and can be measured to different levels of precision. Height, time, mass and temperature are continuous because you can get values like 1.63 m, 3.2 s or 20.5 °C.
数值数据可以进一步分为离散数据和连续数据。离散数据只能取特定的、分离的值——通常是整数——且数值之间存在间隔。例如,一个人的兄弟姐妹数量可以是 0, 1, 2, 3……但不能是 2.5。连续数据可以在给定范围内取任意值,并可以测量到不同的精确度。身高、时间、质量和温度都是连续数据,因为你可以得到像 1.63 m、3.2 s 或 20.5 °C 这样的值。
Quick summary table:
快速总结表:
| Type | Description | Examples |
| Categorical | Groups or labels, not numbers | Eye colour, car brand, month born |
| Numerical / Discrete | Counted numbers, separate steps | Number of books, shoe size (UK) |
| Numerical / Continuous | Measured numbers, any value in a range | Height, weight, time taken |
2. Averages: Mean, Median, Mode | 平均数:均值、中位数、众数
An average is a single value that summarises a whole set of numbers by identifying a ‘typical’ or central value. The three most common averages are mean, median and mode – each gives a different view of the data centre.
平均数是总结一组数字的单一数值,通过找出“典型”或中心值来概括整个数据集。最常见的三种平均数是均值、中位数和众数——每一种都从不同角度展现数据中心。
The mean is what most people call the ‘average’. You find it by adding all the values together and then dividing by how many values there are.
Mean = sum of all values ÷ number of values
均值是大多数人所说的“平均数”。计算方法是将所有数值相加,然后除以数值的个数。
均值 = 所有值的总和 ÷ 值的个数
The median is the middle value when the data is ordered from smallest to largest. If there are two middle numbers, the median is the mean of those two. The median is not affected by extremely high or low values, which makes it useful for skewed data.
中位数是将数据从小到大排列后处于中间位置的值。如果有两个中间数,则中位数为这两个数的均值。中位数不受极高或极低值的影响,因此对于偏斜数据很有用。
The mode (or modal value) is the value that appears most often. A data set can have one mode, more than one mode (bimodal, multimodal) or no mode if all values appear the same number of times. The mode is especially useful for categorical data – for example, the most popular pizza topping.
众数是出现次数最多的值。一个数据集可以有一个众数、多个众数(双峰、多峰),或者如果所有值出现次数相同则没有众数。众数对于分类数据特别有用——例如,最受欢迎的披萨配料。
3. Range and Spread | 极差与离散程度
While averages tell you about the centre of data, measures of spread tell you how spread out the data is. The simplest measure of spread is the range. The range is the difference between the largest and smallest values. It is quick to calculate, but it can be heavily affected by outliers (extreme values).
平均数告诉你数据的中心在哪里,而离散程度的度量则告诉你数据分散到什么程度。最简单的离散度量是极差。极差是最大值与最小值之间的差值。它计算简便,但容易受到异常值(极端值)的极大影响。
Range = maximum value – minimum value
极差 = 最大值 – 最小值
A small range means the data values are quite close together, while a large range suggests they are more spread out. In Year 8, you will also begin to think about why two data sets with the same mean can look very different because of their spread. Comparing ranges alongside averages gives a fuller picture.
极差小意味着数据值比较集中,而极差大则表示数据分布更分散。在 Year 8,你也会开始思考为什么两个均值相同的数据集会因其离散程度不同而看起来差异很大。将极差与平均数一起比较能得出更全面的情况。
4. Frequency and Tally | 频数与划记
Frequency is simply the number of times something occurs. When you collect data, you often organise it into a frequency table. A frequency table lists each category or value and its frequency. It makes patterns much easier to spot.
频数就是某事物出现的次数。当你收集数据时,常常会将其整理成频数表。频数表列出每个类别或数值及其频数。这样更容易发现模式。
A tally is a way of recording frequencies as you gather data. You draw vertical marks, and every fifth mark is drawn diagonally across the previous four to make a group of five – this looks like a gate. Tally marks help you count quickly without losing track.
划记是在收集数据时记录频数的一种方式。你画竖线标记,每第五个标记斜着穿过前四个,形成一组五个——看起来像一扇门。划记符号有助于快速点数而不乱。
| Favourite Fruit | Tally | Frequency |
| Apple | |||| | 4 |
| Banana | |||| |||| | 10 |
When using a tally table, each full gate represents 5. For the banana example, you have two gates, so 5 + 5 = 10. This method reduces counting mistakes.
使用划记表时,每个完整的“门”代表 5。以香蕉为例,有两个门,所以 5 + 5 = 10。这种方法可以减少计数错误。
5. Bar Charts and Pictograms | 条形图与象形图
A bar chart displays categorical data using rectangular bars. The bars are usually drawn with gaps between them to show the categories are separate. The height (or length, for horizontal bars) represents the frequency or value. Bar charts make it easy to compare categories at a glance.
条形图使用矩形条来显示分类数据。条形之间通常会留有空隙,以显示类别是独立的。条的高度(或水平条的长度)代表频数或数值。条形图便于一目了然地比较各类别。
When drawing a bar chart yourself, remember to label both axes, give the chart a title, and use a sensible scale so the tallest bar fits neatly. Always use a ruler and keep the bar widths equal.
当你自己绘制条形图时,记得标注两个轴、给图表加上标题,并使用合理的刻度使最高的条能整齐地放入。始终使用直尺并保持条宽相等。
A pictogram is a chart where small pictures or symbols represent data. A pictogram must have a key that tells you what one symbol stands for. Symbols can be cut in half to show fractions of that value.
象形图是一种用小图片或符号来表示数据的图表。象形图必须有一个图例,说明一个符号代表什么。符号可以切分成一半来表示该值的分数。
For instance, if one smiley face represents 2 pupils, half a face represents 1 pupil. Pictograms are visually attractive but can be less precise than bar charts.
例如,如果一个笑脸代表 2 名学生,那么半个笑脸代表 1 名学生。象形图视觉上很吸引人,但可能不如条形图精确。
6. Pie Charts | 饼图
A pie chart is a circular graph divided into sectors (slices). Each sector shows the proportion of the whole that belongs to a category. The entire circle (360°) represents all the data. The angle of each sector is proportional to its frequency.
饼图是一种被分割成扇形(切片)的圆形图表。每个扇形显示一个类别在整体中所占的比例。整个圆(360°)代表全部数据。每个扇形的角度与其频数成比例。
Sector angle = (frequency ÷ total frequency) × 360°
扇形角度 = (频数 ÷ 总频数) × 360°
Pie charts are excellent for showing relative sizes, but they become hard to read when there are too many small slices. They do not show actual frequencies directly unless they are labelled.
饼图非常适合显示相对大小,但如果小切片太多,阅读起来就会困难。除非标注,它们不会直接显示实际频数。
When constructing a pie chart, first calculate the angle for each category, then use a protractor to measure and draw each sector. Always label the sectors or provide a legend.
制作饼图时,先计算每个类别的角度,然后用量角器测量并画出每个扇形。始终给扇形加上标签或提供图例。
7. Line Graphs and Time Series | 折线图与时间序列
A line graph is used to show how data changes over a continuous interval, most commonly time. Data points are plotted and joined with straight lines. Line graphs help you see trends clearly – for example, a rising line shows an increase.
折线图用于展示数据如何随连续区间(最常见的是时间)变化。数据点被标出并用直线连接。折线图帮助你看清趋势——例如,上升的线条表示增加。
Time series data is a collection of observations taken at regular time intervals, such as daily temperatures or monthly sales. A time series graph is a line graph where the horizontal axis is always time. It can reveal patterns like seasonal swings or a long-term upward trend.
时间序列数据是按固定时间间隔收集的观测值集合,例如每日温度或每月销售额。时间序列图是一种折线图,其水平轴始终是时间。它能揭示季节性波动或长期上升趋势等模式。
Remember to label the axes ‘Time’ and the variable being measured. Use equally spaced intervals on the time axis and join the points with ruled, straight lines – do not draw a smooth curve unless you are sure the real change is smooth.
记住标注轴为“时间”和被测量的变量。时间轴上使用等间距间隔,并用直尺以直线连接各点——除非确定实际变化是平滑的,否则不要画平滑曲线。
8. Scatter Graphs and Correlation | 散点图与相关性
A scatter graph (or scatter plot) is used to investigate whether there is a relationship between two numerical variables. Each point on the graph represents a pair of values – one reading for each variable. The pattern of points suggests the type of relationship.
散点图(或称散点图)用于探究两个数值变量之间是否存在关系。图上的每个点代表一对数值——每个变量一个读数。点的分布模式表明关系的类型。
We describe the relationship using the word correlation. There are three main types:
- Positive correlation: as one variable increases, the other also tends to increase (e.g. temperature and ice cream sales).
- Negative correlation: as one variable increases, the other tends to decrease (e.g. age of a car and its value).
- No correlation: there is no clear pattern; the points are scattered randomly.
我们用相关性来描述这种关系。主要有三种类型:
- 正相关:当一个变量增加时,另一个也倾向于增加(例如温度与冰淇淋销量)。
- 负相关:当一个变量增加时,另一个趋于减少(例如车龄与汽车价值)。
- 无相关:没有明显模式;点随机分布。
Correlation does not imply causation – just because two things are linked does not mean one causes the other. A scatter graph may show a strong correlation, but that could be due to a third factor.
相关性并不意味着因果关系——两个事物有关联并不代表一个导致另一个。散点图可能显示出强相关,但这可能是由第三个因素引起的。
9. Probability Basics | 概率基础
Probability is the branch of maths that deals with how likely events are to happen. The probability scale goes from 0 to 1, where 0 means an event is impossible and 1 means it is certain. Probabilities can be written as fractions, decimals or percentages.
概率是数学中处理事件发生可能性的分支。概率尺度的范围从 0 到 1,0 表示事件不可能发生,1 表示必然发生。概率可以写成分数、小数或百分比。
Probability of an event = number of favourable outcomes ÷ total number of possible outcomes
事件概率 = 有利结果的数量 ÷ 所有可能结果的总数
Key vocabulary: an outcome is a possible result of an experiment, and an event is a set of one or more outcomes. The sample space is the list of all possible outcomes. For example, when rolling a fair six-sided die, the sample space is {1, 2, 3, 4, 5, 6} and the probability of rolling an even number is 3/6 = 1/2.
关键术语:结果是试验的一个可能结果,而事件是一个或多个结果的集合。样本空间是所有可能结果的列表。例如,抛一枚公平六面骰子时,样本空间为{1, 2, 3, 4, 5, 6},掷出偶数的概率为 3/6 = 1/2。
Words like ‘likely’, ‘unlikely’, ‘even chance’ and ‘certain’ are also used to describe probability informally, but in Year 8 you will increasingly work with numerical probabilities.
像“可能”、“不太可能”、“等概率”和“必然”等词语也用于非正式地描述概率,但在 Year 8 你会越来越多地用数值概率来解决问题。
10. Sampling and Bias | 抽样与偏差
In statistics, the population is the whole set of people or things you want to find out about. It is often too large or expensive to survey the entire population, so you select a sample – a smaller, manageable group that should represent the population fairly.
在统计学中,总体是你要研究的全部人或事物的集合。通常对整个总体进行调查规模太大或成本太高,因此你会选择一个样本——一个较小、可管理的群体,它应该公平地代表总体。
To get reliable results, a sample must be random (every member of the population has an equal chance of being chosen) and representative (it mirrors the characteristics of the whole population). If a sample is not random, it may introduce bias, making the conclusions unreliable.
为了得到可靠的结果,样本必须是随机的(总体中每个成员有均等被选中的机会)且具有代表性(它反映整个总体的特征)。如果样本不是随机的,可能会引入偏差,导致结论不可靠。
For example, if you want to know the favourite sport of all Year 8 students in a school and you only ask the boys’ football team, your sample is biased – it does not represent the whole population. A better sample would randomly pick students from every form group.
例如,如果你想了解某校所有 Year 8 学生最喜欢的运动,而你只询问了男生足球队,那么你的样本是有偏差的——它不代表整个总体。更好的样本是从每个班级随机挑选学生。
Bias can also come from the way questions are asked in a survey. Leading questions like ‘Don’t you agree that maths is the best subject?’ push people towards a particular answer. A fair question should be neutral, such as ‘Which subject do you prefer: maths or English?’
偏差也可能来自调查中提问的方式。像“你不同意数学是最好的科目吗?”这样的引导性问题会将人们推向特定答案。公正的问题应该是中性的,例如“你更喜欢哪个科目:数学还是英语?”
11. Primary and Secondary Data | 一手数据与二手数据
Primary data is information you collect yourself, specifically for your investigation. It can be gathered through experiments, questionnaires, interviews or direct observations. The main advantage is that you know exactly how it was obtained and can trust its reliability, but it can be time-consuming to collect.
一手数据是你自己为特定调查收集的信息。它可以通过实验、问卷、访谈或直接观察获得。主要优点是你确切知道数据的获取方式,可以相信其可靠性,但收集起来可能耗时。
Secondary data is data that someone else has already collected for a different purpose. You might find it in books, on reputable websites, in newspapers or in government reports. It is often quicker and cheaper to use, but you must evaluate its source carefully to check for bias or inaccuracy.
二手数据是别人已出于不同目的收集好的数据。你可以在书籍、可靠网站、报纸或政府报告中找到。使用二手数据通常更快、更省钱,但你必须仔细评估其来源,以检查是否存在偏差或不准确。
Both types of data are useful. In Year 8 projects, you may design a short questionnaire to gather primary data and then compare your findings with secondary data from an official source, like the Office for National Statistics.
两类数据都有用。在 Year 8 项目中,你可以设计一份简短的问卷来收集一手数据,然后将其与来自官方来源(如国家统计局)的二手数据进行比较。
12. Designing a Good Statistical Enquiry | 设计一个好的统计调查
Every statistical investigation follows a cycle: posing a question, collecting data, analysing the data and drawing conclusions. The vocabulary you have learnt fits into each stage. A clear, well-worded question is the starting point. For example, ‘What is the most common method of travelling to school among Year 8 pupils?’ is a focused question.
每个统计调查都遵循一个循环:提出问题、收集数据、分析数据并得出结论。你所学的词汇适用于每个阶段。一个清晰、措辞得当的问题是起点。例如,“Year 8 学生最常用的上学方式是什么?”就是一个聚焦的问题。
When designing a data collection sheet or questionnaire, make sure response boxes do not overlap and that categories cover all possibilities. Always include an ‘Other’ option where necessary. Use the language of frequency, tally
Published by TutorHao | Year 8 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply