📚 Year 8 CAIE Statistics: Core Knowledge Points Review | Year 8 CAIE 统计:核心知识点梳理
Welcome to the Year 8 CAIE Statistics revision guide. In this article, we will systematically organise the key concepts you need to master in statistics at this level. From understanding different data types to calculating averages and interpreting probabilities, each section pairs English explanations with Chinese translations to support bilingual learning. Whether you are preparing for a checkpoint test or building a strong foundation for IGCSE, this article will clarify the core ideas and common question types you will face.
欢迎来到 Year 8 CAIE 统计复习指南。在这篇文章中,我们将系统梳理这个阶段你需要掌握的统计核心概念。从理解不同的数据类型,到计算各种平均数,再到解释概率的含义,每一个小节都包含英文解释和对应的中文翻译,帮助你进行双语学习。无论你是在为期中测验做准备,还是在为 IGCSE 打下扎实基础,这篇文章都能帮你理清核心知识点和常见题型。
1. Types of Data | 数据类型
In statistics, data can be divided into two broad categories: qualitative data and quantitative data. Qualitative data describes qualities or characteristics that cannot be measured with numbers, such as eye colour, favourite food, or types of transport. Quantitative data, on the other hand, consists of numerical values that can be counted or measured, such as height, test scores, or number of siblings.
在统计学中,数据可以分为两大类:定性数据和定量数据。定性数据描述的是无法用数字衡量的性质或特征,比如眼睛的颜色、最喜爱的食物、交通工具的类型。定量数据则是由可以计数或测量的数值构成,例如身高、测验分数或兄弟姐妹的数量。
Quantitative data is further split into discrete and continuous data. Discrete data can only take specific, separate values – often whole numbers – such as the number of students in a class (you cannot have half a student). Continuous data can take any value within a given range and is usually obtained by measuring; examples include time, weight, and temperature, where readings like 36.2°C or 1.65 m are possible.
定量数据又可细分为离散数据和连续数据。离散数据只能取特定的、分开的数值——通常是整数——比如班级里的学生人数(不可能有半个学生)。连续数据则可以取某一范围内的任意值,通常通过测量得到;例子包括时间、体重和温度,像 36.2°C 或 1.65 米这样的读数都是可能的。
2. Collecting Data | 数据收集
Data collection is the first step in any statistical investigation. Primary data is information you gather yourself through surveys, experiments, or observations. For example, if you ask your classmates how they travel to school and record the answers, you are collecting primary data. Secondary data comes from sources that already exist, such as websites, books, or government statistics. For instance, using rainfall records from a weather website is using secondary data.
数据收集是任何统计调查的第一步。一手数据是你通过调查、实验或观察自己收集到的信息。例如,如果你询问班上同学如何来上学并记录答案,你就是在收集一手数据。二手数据则来自已有的资料,比如网站、书籍或政府统计数据。比如,使用天气网站上的降雨量记录就是使用二手数据。
A well-designed data collection sheet or questionnaire is essential. Questions should be clear and unbiased, and they must gather the right type of data for the problem. When designing a questionnaire, avoid leading questions such as ‘Don’t you agree that maths is the best subject?’ and instead use neutral wording like ‘What is your favourite subject?’.
设计良好的数据收集表或问卷至关重要。问题应该清晰、无偏向性,并且必须针对问题收集正确类型的数据。在设计问卷时,要避免引导性问题,比如“你不觉得数学是最好的科目吗?”,而应该使用中性的措辞,比如“你最喜爱的科目是什么?”。
3. Frequency Tables | 频数表
Once raw data is collected, it is often organised into a frequency table to make it easier to read and interpret. A frequency table lists each distinct outcome or category alongside how many times it occurs (its frequency). For example, a frequency table showing the outcome of rolling a die 30 times might list each number from 1 to 6 and the corresponding count.
收集到原始数据后,通常会将其整理成频数表,以便于阅读和解释。频数表列出了每一个不同的结果或类别,以及它们出现的次数(即频数)。例如,一张显示掷骰子30次结果的频数表,可能会列出从1到6的每个数字及其对应的出现次数。
Tally marks are a quick way to record frequencies while collecting data. Each vertical line represents one observation, and every fifth line is drawn diagonally across the previous four to make counting in groups of five easier. When grouping continuous data, we use class intervals such as 0 ≤ x < 10, 10 ≤ x < 20, and we must take care that intervals do not overlap.
划记是在收集数据时快速记录频数的一种方法。每条竖线代表一次观测,每第五个记号以对角线划过前四竖,这样便于五个一组地计数。当对连续数据进行分组时,我们使用组区间,比如 0 ≤ x < 10,10 ≤ x < 20,并且必须注意区间不要互相重叠。
| Score | Tally | Frequency |
|---|---|---|
| 0–9 | |||| | 4 |
| 10–19 | |||| || | 7 |
| 20–29 | |||| | 4 |
4. Bar Charts and Pictograms | 条形图与象形图
Bar charts are used to display and compare the frequencies of categorical or discrete data. The height of each bar represents the frequency, and all bars must have equal width with gaps between them, unless the data is continuous. The categories are usually placed on the horizontal axis, while frequency is on the vertical axis.
条形图用于显示和比较分类或离散数据的频数。每个直条的高度代表频数,并且所有直条必须宽度相同,直条之间留有间隙,除非数据是连续的。类别通常置于水平轴,而频数置于垂直轴。
A pictogram uses small pictures or symbols to represent a certain number of items. For example, one picture of a car might represent 10 cars. A key must always be provided to show the scale. When a frequency is not an exact multiple of the symbol, a part of the symbol is drawn. Pictograms are visually appealing but can be less precise than bar charts.
象形图使用小图片或符号来代表一定数量的项目。例如,一个小汽车的图片可能代表 10 辆汽车。必须提供一个图例来说明比例。当频数不是符号的整数倍时,画出一部分符号。象形图直观吸引人,但可能不如条形图精确。
5. Pie Charts | 饼图
Pie charts represent data as sectors of a circle, where the angle of each sector is proportional to the frequency of that category. The whole circle represents the total, which is 360°. To find the angle for a category, use the formula:
Angle = (Frequency of category ÷ Total frequency) × 360°
饼图将数据表示为圆的扇形,每个扇形的角度与该类别的频数成比例。整个圆代表总数,即 360°。要计算某个类别的角度,使用公式:
角度 =(类别的频数 ÷ 总频数)× 360°
When drawing a pie chart, you must first calculate the angle for each category using the formula, then measure and draw each sector accurately with a protractor. Label each sector or provide a legend. Pie charts are excellent for showing proportions out of a whole, but they become difficult to read when there are too many categories or when frequencies are very similar.
绘制饼图时,你必须先用公式计算出每个类别的角度,然后用量角器准确地测量并画出每个扇形。给每个扇形做标注,或者提供图例。饼图非常适合展示整体中的比例,但当类别太多或频数非常接近时,阅读起来就比较困难。
6. Line Graphs and Scatter Graphs | 折线图与散点图
Line graphs are used to display data that changes over time, such as temperature readings throughout the day or a plant’s height over several weeks. Points are plotted and connected in order with straight lines. The horizontal axis often represents time, while the vertical axis shows the variable being measured.
折线图用于显示随时间变化的数据,比如一天中气温的读数,或者几周内植物的高度。将各点标出,并按顺序用直线连接。水平轴通常代表时间,垂直轴则显示被测量的变量。
Scatter graphs (or scatter plots) are used to explore the relationship between two sets of quantitative data. Each point on the graph represents a pair of values, such as the number of hours studied and the test score achieved. If the points tend to rise together, there is a positive correlation; if one falls as the other rises, there is a negative correlation. If no pattern is visible, there is no correlation. Scatter graphs help us decide whether two variables might be linked.
散点图(或称散点图)用于探索两组定量数据之间的关系。图上的每个点代表一对数值,例如学习的小时数和取得的测验分数。如果点的走势一起上升,则为正相关;如果一个下降而另一个上升,则为负相关;如果没有可见的模式,则为零相关。散点图帮助我们判断两个变量之间是否存在关联。
7. Mean, Median, Mode | 平均数、中位数、众数
Measures of central tendency help us find a typical or average value in a data set. The three main measures are the mean, median, and mode.
集中趋势的度量帮助我们找到数据集中的典型值或平均值。三个主要的度量是平均数、中位数和众数。
The mean is calculated by adding all the values together and dividing by the number of values.
Mean = (Sum of all values) ÷ (Number of values)
平均数是将所有数值相加,再除以数值的个数。
平均数 =(所有数值的总和)÷(数值的个数)
The median is the middle value when the data is arranged in ascending order. If there is an odd number of values, the median is the central number. If there is an even number of values, the median is the mean of the two middle numbers.
中位数是将数据按升序排列后位于中间的值。如果有奇数个数值,中位数就是中间的那个数。如果有偶数个数值,中位数则是中间两个数的平均数。
The mode is the value that appears most frequently. A data set can have one mode, more than one mode (bimodal or multimodal), or no mode at all if no value repeats.
众数是出现次数最多的值。一个数据集可以有一个众数、多个众数(双众数或多众数),或者如果没有重复的数值,则没有众数。
8. Range | 极差
Range is a measure of spread that tells us how widely data values are distributed. It is the difference between the largest value and the smallest value in the data set.
Range = Largest value – Smallest value
极差是一种衡量离散程度的指标,它告诉我们数据值的分布有多广。它是数据集中最大值与最小值的差。
极差 = 最大值 – 最小值
Although the range is easy to calculate, it can be heavily influenced by extreme values (outliers). A single unusually high or low value can make the range very large, even if the rest of the data is closely grouped. That is why we often use the range alongside the mean or median to get a fuller picture of a data set.
虽然极差很容易计算,但它极易受到极端值(异常值)的影响。即使其他数据都很集中,一个异常的高值或低值就能让极差变得非常大。因此,我们常常将极差与平均数或中位数一起使用,以便更全面地了解数据集的情况。
9. Introduction to Probability | 概率入门
Probability measures how likely an event is to occur. It is always expressed as a number between 0 and 1 (inclusive), or equivalently as a percentage between 0% and 100%. A probability of 0 means the event is impossible, while a probability of 1 means it is certain to happen.
概率衡量一个事件发生的可能性大小。它总是用一个介于 0 到 1(含)之间的数字来表示,或者等价地用 0% 到 100% 的百分数来表示。概率为 0 表示事件不可能发生,概率为 1 表示事件必然发生。
The probability of an event is calculated using the formula:
Probability of event = (Number of favourable outcomes) ÷ (Total number of possible outcomes)
事件的概率用以下公式计算:
事件的概率 =(有利结果的数目)÷(所有可能结果的总数)
This formula assumes that all outcomes are equally likely, as with a fair coin or a fair six‑sided die. For example, the probability of rolling a 4 on a fair die is 1 ÷ 6, because there is one favourable outcome and six possible outcomes.
这个公式假设所有结果发生的可能性相等,就如同抛一枚公平的硬币或掷一个公平的六面骰子一样。例如,掷一个公平骰子得到 4 的概率是 1 ÷ 6,因为有一个有利结果和六个可能结果。
10. Probability Scale and Simple Events | 概率尺度与简单事件
The probability scale is a useful tool for visualising likelihood. It is a line from 0 to 1 with labels such as impossible, unlikely, evens, likely, and certain. Placing events on this scale helps compare how probable different outcomes are. For instance, ‘the sun will rise tomorrow’ is a certain event (probability 1), while ‘rolling a 7 on a standard six‑sided die’ is impossible (probability 0).
概率尺度是一个用于将可能性可视化的实用工具。这是一条从 0 到 1 的线段,标有“不可能”、“不太可能”、“均等可能”、“很可能”和“必然”等标签。将事件放在这个尺度上,有助于比较不同结果发生的可能性。例如,“太阳明天会升起”是必然事件(概率为 1),而“在一个标准六面骰子上掷出 7”是不可能事件(概率为 0)。
When working with simple events, it is vital to list all possible outcomes systematically, often using a sample space diagram, a list, or a two‑way table. For two‑step events, such as flipping a coin and rolling a die, the total number of outcomes is found by multiplying the number of outcomes for each step (the multiplication principle). Understanding sample spaces sets the stage for more advanced probability topics in later years.
在处理简单事件时,系统性地列出所有可能的结果至关重要,通常会使用样本空间图、清单或双向表格。对于两步事件,比如抛一枚硬币并掷一个骰子,总的结果数可由每一步的结果数相乘得到(乘法原理)。理解样本空间能为今后学习更高阶的概率主题打下基础。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导