📚 Year 9 CAIE Statistics: Core Topics Overview | Year 9 CAIE 统计:核心知识点梳理
Welcome to the Year 9 CAIE Statistics revision guide. This article brings together the key concepts you need to master — from collecting and displaying data to measuring averages, spread, and probability. Whether you are preparing for a test or catching up on missed work, this overview will help you connect the main topics in a straightforward, step‑by‑step way.
欢迎来到 Year 9 CAIE 统计复习指南。本文汇集了你需要掌握的核心概念——从数据收集与展示到集中趋势、离散程度和概率。无论你是在准备考试还是补习遗漏的内容,这篇梳理都将以清晰、循序渐进的方式帮助你串联起主要的知识点。
1. Types of Data and Data Collection | 数据类型与数据收集
Data can be split into two broad categories: qualitative (categorical) and quantitative (numerical). Qualitative data describes qualities or categories, such as favourite colour or type of pet. Quantitative data is numerical and can be either discrete (countable, like number of siblings) or continuous (measurable, like height or time).
数据可以分为两大类:定性数据(分类数据)和定量数据(数值数据)。定性数据描述品质或类别,例如最爱的颜色或宠物的种类。定量数据是数值型的,可以是离散的(可数的,如兄弟姐妹的个数)或连续的(可测量的,如身高或时间)。
Primary data is collected directly by you for a specific purpose, for example through a survey or an experiment. Secondary data is obtained from existing sources, such as websites, books, or government records. Knowing the source helps you judge the reliability of the data.
一手数据是你为特定目的直接收集的,例如通过问卷调查或实验。二手数据来自已有的来源,例如网站、书籍或政府记录。了解数据的来源有助于判断数据的可靠性。
When planning data collection, you must decide on the population (the whole group of interest) and, if it is too large, choose a sample. A random sample gives each member an equal chance of being selected, which helps avoid bias.
在设计数据收集方案时,你必须确定总体(你关心的整个群体),如果总体太大,就选取一个样本。随机样本让每个成员被选中的机会均等,这有助于避免偏差。
2. Organising and Stem‑and‑Leaf Diagrams | 数据整理与茎叶图
Once data is collected, it needs to be organised so patterns can be seen. A stem‑and‑leaf diagram keeps all original values while showing the shape of the distribution. The stem represents the leading digit(s) and the leaf is the final digit.
收集到数据后,需要对其进行整理以便发现规律。茎叶图既能保留所有原始数值,又能展示分布的形态。茎代表前一位(或多位)数字,叶则是最后一位数字。
For example, the numbers 23, 25, 31, 33, 33, 38 can be drawn as:
Stem 2 | Leaf 3 5
Stem 3 | Leaf 1 3 3 8
Always include a key, e.g. ‘2 | 3 means 23’. The diagram quickly reveals the minimum, maximum, mode, and whether data is clustered.
例如,数字 23, 25, 31, 33, 33, 38 可以画成:
茎 2 | 叶 3 5
茎 3 | 叶 1 3 3 8
记得加上图例,例如“2 | 3 表示 23”。该图能快速揭示最小值、最大值、众数以及数据是否集中。
Ordered stem‑and‑leaf diagrams sort the leaves in each row, making it even easier to find the median. For large data sets, a back‑to‑back stem‑and‑leaf plot can compare two related distributions.
有序茎叶图将每一行的叶按大小排列,使得寻找中位数更加容易。对于大型数据集,背靠背茎叶图可以比较两个相关的分布。
3. Mean, Median, and Mode | 平均数、中位数与众数
The three main averages summarise the centre of a data set. The mode is the value that appears most often. There can be one mode, more than one (bimodal or multimodal), or none at all.
三种主要的平均数概括了数据集的中心。众数是出现次数最多的值。众数可以有一个、多个(双众数或多众数),也可能没有。
The median is the middle value when data is arranged in order. For an odd number of values, it is the centre one; for an even number, it is the mean of the two middle values. The median is not affected by extreme values (outliers).
中位数是将数据按顺序排列后位于中间的值。观测值个数为奇数时,中位数是正中间的那个数;个数为偶数时,它是中间两个数的平均数。中位数不受极端值(异常值)的影响。
The mean is calculated by adding all values and dividing by the number of values. We often write this as:
Mean = (Σx) / n
The mean uses every piece of data, so it is sensitive to outliers.
平均数(均值)通过将所有数值相加再除以数值的个数来计算。我们通常写作:
平均数 = (Σx) / n
平均数用到了每一个数据,因此对异常值敏感。
For data like shoe sizes or exam scores, choose the average that best represents the typical value. If data is skewed, the median is often more reliable than the mean.
对于鞋子尺码或考试分数这类数据,要选择最能代表典型值的那一种平均数。如果数据偏态分布,中位数通常比均值更可靠。
4. Range and Quartiles | 极差与四分位数
Measure of spread describes how spread out the data is. The simplest measure is the range:
Range = maximum value – minimum value
It tells you the width of the whole data set, but says nothing about the shape in between.
离散程度的度量描述数据的分散情况。最简单的度量是极差:
极差 = 最大值 – 最小值
它告诉你整个数据集的宽度,但不能说明中间部分的形态。
Quartiles divide ordered data into four equal parts. The lower quartile (Q₁) is the median of the lower half of the data. The upper quartile (Q₃) is the median of the upper half. The interquartile range (IQR) = Q₃ – Q₁, and it measures the spread of the middle 50% of the data.
四分位数将有序数据分成四个相等的部分。下四分位数(Q₁)是数据下半部分的中位数。上四分位数(Q₃)是数据上半部分的中位数。四分位距(IQR)= Q₃ – Q₁,它衡量了中间 50% 数据的分散程度。
Quartiles help construct box‑and‑whisker plots, which display the minimum, Q₁, median, Q₃, and maximum. Box plots are excellent for comparing distributions side by side and for spotting outliers defined as values less than Q₁ – 1.5×IQR or greater than Q₃ + 1.5×IQR.
四分位数可用于构建箱线图,它呈现最小值、Q₁、中位数、Q₃ 和最大值。箱线图非常适合并列比较分布,也可用于识别异常值——异常值定义为小于 Q₁ – 1.5×IQR 或大于 Q₃ + 1.5×IQR 的值。
5. Frequency Tables and Grouped Data | 频率表与分组数据
When data is too numerous to list individually, a frequency table organises values and shows how often each one occurs. For discrete data, list each value; for continuous data or a large range, group data into class intervals.
当数据太多无法逐一列出时,频率表可以整理数值并显示每个值出现的频率。对于离散数据,列出每一个值;对于连续数据或范围很大的数据,则将数据分组成组距。
Grouped frequency tables require you to know the class boundaries, class width, and midpoints. The midpoint of a class = (lower boundary + upper boundary) ÷ 2. You use midpoints to estimate the mean of grouped data:
Estimated mean = Σ(f × midpoint) / Σf
where f is the frequency of each class.
分组频率表要求你了解组限、组距和组中值。组中值 = (下限 + 上限) ÷ 2。利用组中值可以估计分组数据的平均数:
估计平均数 = Σ(f × 组中值) / Σf
其中 f 是每组的频数。
The modal class is the class interval with the highest frequency. The median class can be found by working out cumulative frequency and locating the position of the median (n/2).
众数组是频数最高的那个组距。中位数组可以通过计算累积频率并找出中位数所在位置(n/2)来确定。
6. Bar Charts and Histograms | 条形图与直方图
Bar charts are used for categorical or discrete data. Each category gets a bar of equal width, with height proportional to its frequency. Bars do not touch, highlighting that the categories are separate.
条形图用于分类数据或离散数据。每个类别用一个等宽的条形表示,条形的高度与频数成正比。条形之间不相连,以突出类别是分开的。
Histograms look similar but are used for grouped continuous data. In a histogram, the bars touch to show that the data is continuous. The area of each bar represents frequency. When class widths are unequal, use frequency density:
Frequency density = frequency / class width
Then plot frequency density on the vertical axis.
直方图看起来类似,但用于分组连续数据。直方图中条形相连,表示数据是连续的。每个条形的面积代表频数。当组距不相等时,使用频数密度:
频数密度 = 频数 / 组距
然后在纵轴上绘制频数密度。
Labelling axes, choosing sensible scales, and giving a title are essential for both bar charts and histograms. Always check that the vertical axis starts at zero to avoid misleading impressions.
为坐标轴加标签、选择合理的刻度并添加标题,对条形图和直方图来说都至关重要。务必检查纵轴是否从零开始,以免造成误导。
7. Pie Charts and Pictograms | 饼图与象形图
A pie chart displays data as slices of a circle, where the angle of each slice is proportional to the frequency. To find the angle:
Angle = (frequency / total frequency) × 360°
Pie charts are especially useful for showing proportions and comparing parts of a whole.
饼图以圆形的扇区来呈现数据,每个扇区的角度与频数成正比。求角度的公式:
角度 = (频数 / 总频数) × 360°
饼图在显示比例和比较整体各部分时特别有用。
Pictograms use small icons or pictures to represent a certain number of items. They make data visually engaging. Always include a key that states what one symbol stands for. When a frequency is not a multiple of the symbol’s value, a fraction of the symbol is used.
象形图使用小图标或图片来表示一定数量的对象,使数据更具视觉吸引力。务必添加图例,说明一个符号代表什么。当频数不是符号值的整数倍时,可以使用符号的一小部分。
Both pie charts and pictograms should have clear titles and labels. They work well for illustrating survey results about favourite sports, colours, or modes of transport.
饼图和象形图都应有清晰的标题和标签。它们在展示关于最爱的运动、颜色或交通方式的调查结果时效果很好。
8. Scatter Graphs and Correlation | 散点图与相关性
A scatter graph plots paired numerical data points to see if there is a relationship between two variables. Each point represents one pair of values (x, y). The independent variable goes on the horizontal axis; the dependent variable on the vertical axis.
散点图将成对的数值数据点绘制出来,以观察两个变量之间是否存在关系。每个点代表一对数值(x, y)。自变量位于横轴,因变量位于纵轴。
Correlation describes the direction and strength of the relationship. Positive correlation means as x increases, y tends to increase. Negative correlation means as x increases, y tends to decrease. No correlation means there is no clear pattern.
相关性描述关系的方向和强弱。正相关意味着随着 x 增大,y 也倾向于增大。负相关意味着随着 x 增大,y 倾向于减小。无相关意味着没有清晰的模式。
Strength can be described as strong, moderate, or weak by looking at how tightly the points cluster around an imaginary straight line. Outliers that lie far from the general trend should be noted separately.
通过观察数据点围绕一条虚拟直线的紧密程度,可以将相关强度描述为强、中等或弱。远离总体趋势的异常值应单独标注。
A line of best fit can be drawn by eye through the middle of the points. It should have roughly equal numbers of points above and below the line. This line can be used to estimate unknown values — interpolation within the data range is reliable; extrapolation outside the range is less certain.
可以通过目测穿过数据点中部画一条最佳拟合线。该线上下方的数据点数量应大致相等。这条线可以用来估计未知数值——在数据范围内插值是可靠的;超出范围的外推则不太可靠。
9. Introduction to Probability | 概率入门
Probability measures how likely an event is to happen. It is always a number between 0 and 1 inclusive. A probability of 0 means impossible; 1 means certain. The probability scale can also be expressed in words (impossible, unlikely, even chance, likely, certain).
概率衡量一个事件发生的可能性大小。它总是介于 0 和 1 之间(包括 0 和 1)。概率为 0 意味着不可能发生;1 意味着必然发生。概率尺度也可以用文字表示(不可能、不太可能、均等机会、很可能、必然)。
The basic formula for equally likely outcomes is:
P(event) = number of favourable outcomes / total number of possible outcomes
For example, when rolling a fair six‑sided die, P(rolling a 4) = 1/6.
等可能结果的基本公式为:
P(事件) = 有利结果的数量 / 所有可能结果的总数
例如,掷一个均匀的六面骰子,P(掷出 4) = 1/6。
Probabilities can be expressed as fractions, decimals, or percentages. The sum of probabilities of all mutually exclusive outcomes in a sample space equals 1. Complementary events: P(event not happening) = 1 – P(event happening).
概率可以用分数、小数或百分数表示。样本空间中所有互斥结果的概率之和等于 1。互补事件:P(事件不发生) = 1 – P(事件发生)。
10. Venn Diagrams and Tree Diagrams | 维恩图与树状图
Venn diagrams are used to show sets and their overlaps. A rectangle represents the universal set, and circles inside represent subsets. The overlap (intersection) contains elements that belong to both sets. The union contains elements in either set.
维恩图用于展示集合及其重叠关系。长方形代表全集,内部的圆代表子集。重叠部分(交集)包含属于两个集合的元素。并集包含属于任意一个集合的元素。
When solving probability problems, Venn diagrams help organise information. For two events A and B, P(A ∪ B) = P(A) + P(B) – P(A ∩ B). This formula avoids double‑counting the intersection.
在解决概率问题时,维恩图有助于整理信息。对于两个事件 A 和 B,P(A ∪ B) = P(A) + P(B) – P(A ∩ B)。这个公式避免了重复计算交集。
Tree diagrams show the outcomes of two or more simple events in sequence. Each branch is labelled with a probability. To find the probability of a combined path, multiply the probabilities along the branches. If there is more than one path to an outcome, add the path probabilities.
树状图按顺序展示两个或多个简单事件的结果。每条分支都标有概率。要计算一个组合路径的概率,就沿分支将概率相乘。如果到达某个结果有多条路径,就把各路径的概率相加。
For independent events, the outcome of one does not affect the other. Tree diagrams also handle conditional probability, where the probability on the second branch depends on what happened first. At Year 9 level, you usually see replacement and non‑replacement situations.
对于独立事件,一个事件的结果不会影响另一个事件。树状图也能处理条件概率,即第二条分支上的概率取决于第一步发生的事情。在 Year 9 阶段,通常会看到有放回和不放回的情况。
11. Designing and Critiquing Statistical Investigations | 统计调查的设计与评价
A statistical investigation follows the statistical enquiry cycle: pose a question, collect data, analyse data, and interpret results. The question should be clear, measurable, and focused, for example ‘How many hours per week do Year 9 students spend on homework?’
统计调查遵循统计探究循环:提出问题、收集数据、分析数据、解释结果。问题必须清晰、可测量且聚焦,例如“Year 9 学生每周花多少小时做家庭作业?”。
When designing a survey, avoid leading questions, ambiguous wording, and restricted response options that introduce bias. Pilot surveys can help refine the questionnaire. Sampling methods, such as random, systematic, or stratified sampling, affect how representative the results are.
在设计调查问卷时,要避免引导性问题、模糊的措辞和限制性的选项,这些都会引入偏差。可以先做试点调查来完善问卷。抽样方法(如随机抽样、系统抽样或分层抽样)会影响结果的代表性。
Critiquing a statistical claim involves checking the source, the sample size, and the presentation of data. Misleading graphs — with broken axes, unusual scales, or 3D effects — can distort the truth. Always ask: Is this fair? Could there be another explanation?
评价一项统计主张需要检查来源、样本量以及数据呈现方式。误导性的图表——例如带有断裂坐标轴、异常比例尺或三维效果的图表——可能扭曲事实。始终要问:这公平吗?是否有其他解释?
12. Bringing It All Together | 核心知识点整合
Year 9 CAIE Statistics connects many topics that build a foundation for later study. From collecting data responsibly to displaying it with clear graphs and summarising it with averages and spread, each skill reinforces the others. Probability and diagrams provide tools for reasoning about uncertainty and logical structure.
Year 9 CAIE 统计将许多主题联系在一起,为以后的学习打下基础。从负责任地收集数据,到用清晰的图形展示数据,再到用平均数和离散程度进行概括,每一项技能都相互促进。概率和各种图形则为不确定性和逻辑结构的推理提供了工具。
When you revise, practise reading different charts, calculating summary statistics, and interpreting what they tell you. Always check your answers make sense in the real‑world context. Master these core ideas and you will be well‑prepared for IGCSE Statistics and beyond.
在复习时,要多练习阅读不同的图表、计算汇总统计量并解释它们传达的信息。务必检查你的答案在现实情境中是否合理。掌握这些核心思想,你将为 IGCSE 统计及更高级别的学习做好充分准备。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply