📚 Year 9 Cambridge Statistics: Core Concepts Overview | Year 9 剑桥统计:核心知识点梳理
Welcome to your Year 9 Cambridge Statistics revision guide. This article covers the essential topics you need to master, from collecting and organising data to calculating probabilities. Each section presents a key concept with clear explanations and practical examples, helping you build a solid foundation for IGCSE Mathematics. Whether you are reviewing types of data, constructing stem-and-leaf diagrams, or understanding how to compare two data sets using averages and range, you will find all the core knowledge in one place.
欢迎来到 Year 9 剑桥统计复习指南。本文涵盖了你需要掌握的核心主题,从收集和整理数据到计算概率。每个部分都用清晰的解释和实用的例子展示关键概念,帮助你为 IGCSE 数学打下扎实的基础。无论你是在复习数据类型、构建茎叶图,还是理解如何用平均数和极差比较两组数据,这里都能找到全部核心知识点。
1. Types of Data and Data Collection | 数据类型与数据收集
Data can be classified as qualitative or quantitative. Qualitative data describes qualities and is non-numerical, such as favourite colours or types of fruit. Quantitative data involves numbers and can be further divided into discrete and continuous data. Discrete data can only take specific, separate values, like the number of students in a class or shoe sizes. Continuous data can take any value within a range, such as height, weight, or time.
数据可以分为定性数据和定量数据。定性数据描述品质,是非数值型的,比如最喜欢的颜色或水果种类。定量数据涉及数字,可进一步分为离散数据和连续数据。离散数据只能取特定的、分开的值,比如班级学生人数或鞋码。连续数据可以在一个范围内取任意值,比如身高、体重或时间。
When collecting data, we must ensure it is reliable. This can be done through questionnaires, experiments, or observations. A sample is a smaller group chosen from a population. To avoid bias, the sample should be random and representative. Common methods include simple random sampling and stratified sampling, where the population is divided into groups and a fixed number is taken from each group proportionally.
收集数据时,我们必须确保它是可靠的。这可以通过问卷、实验或观察来完成。样本是从总体中选出的较小群体。为了避免偏差,样本应该是随机且有代表性的。常用的方法包括简单随机抽样和分层抽样,后者将总体分为若干组,并按比例从每组中抽取固定数量的个体。
2. Frequency Tables and Grouped Data | 频数表与分组数据
A frequency table organises raw data by listing each value alongside how many times it occurs. Tally marks are often used during data collection to keep a running count. From a frequency table, we can quickly see the mode – the value that appears most often.
频数表通过列出每个数值及其出现的次数来整理原始数据。在收集数据时常用画“正”字计数来记录。从频数表中,我们可以快速看出众数——出现次数最多的数值。
When data has a wide range of values, it is convenient to group it into class intervals. For example, test scores of 0–9, 10–19, and so on. In a grouped frequency table, we lose exact individual values but gain a clearer picture of the distribution. Inequalities are used to define boundaries clearly: 0 ≤ score < 10, 10 ≤ score < 20, etc.
当数据值范围很广时,将其分组为组距很方便。例如,考试分数 0–9、10–19 等等。在分组频数表中,我们丢失了精确的单个值,但能更清晰地看到分布情况。用不等式来明确定义边界:0 ≤ 分数 < 10,10 ≤ 分数 < 20 等等。
| Class Interval | Frequency |
| 0–9 | 5 |
| 10–19 | 12 |
| 20–29 | 8 |
The modal class interval is the one with the highest frequency, here 10–19. With grouped data, we cannot find the exact mean without the raw data; instead we use midpoints of intervals to estimate it.
众数所在的组距是频率最高的那个区间,这里是 10–19。对于分组数据,没有原始数据时无法求出精确的均值,我们使用区间的中点来估算均值。
3. Bar Charts, Pictograms, and Pie Charts | 条形图、象形图与饼图
Bar charts are used to display categorical or discrete data. The height of each bar represents the frequency. Bars should be of equal width and separated by gaps, unless it is a histogram for continuous data. Always label the axes and give the chart a title.
条形图用于展示分类数据或离散数据。每个条形的高度表示频数。条形应等宽且之间留有间隙,除非是用于连续数据的直方图。一定要给坐标轴加上标签,并给图表加上标题。
Pictograms use symbols or pictures to represent a certain number of items. A key must be provided to show what one symbol stands for. For instance, one book icon could represent 10 students. Part of a symbol represents a fraction of that unit.
象形图使用符号或图片来表示一定数量的项目。必须提供图例来说明一个符号代表什么。例如,一个书本图标可以代表 10 名学生。部分符号则表示该单位的几分之一。
Pie charts show proportions of a whole. Each sector’s angle is calculated using: angle = (frequency ÷ total frequency) × 360°. Always use a protractor for accuracy and label each sector or include a legend.
饼图显示整体中各部分的比例。每个扇形的角度这样计算:角度 = (频数 ÷ 总频数) × 360°。务必用量角器准确绘制,并给每个扇形加标签或图例。
4. Stem-and-Leaf Diagrams | 茎叶图
A stem-and-leaf diagram is a way of ordering and displaying numerical data in its original form. Each value is split into a ‘stem’ (the leading digit or digits) and a ‘leaf’ (the final digit). For example, the number 48 has stem 4 and leaf 8. The stems are written in a vertical column in order, and the leaves are listed horizontally next to the appropriate stem, usually in ascending order.
茎叶图是一种以原始形式排序并展示数值数据的方法。每个数值被分为“茎”(前一位或多位数字)和“叶”(最后一位数字)。例如,48 的茎是 4,叶是 8。茎按顺序排成一列,叶在对应茎的右侧横向列出,通常按升序排列。
A key must always be included to show how to read the diagram, e.g. ‘4 | 8 means 48’. Stem-and-leaf diagrams have the advantage of retaining the exact data values, unlike bar charts. They also make it easy to find the median and mode.
必须包含图例来说明如何解读图表,例如 “4 | 8 表示 48”。与条形图不同,茎叶图保留了精确的数据值,这一点很有优势。它也便于找出中位数和众数。
For a two-digit number like 105, the stem could be 10 and leaf 5, depending on the context. Always decide the split carefully and state it in the key.
对于像 105 这样的三位数,茎可以是 10,叶是 5,取决于具体情况。务必仔细决定拆分方式并在图例中说明。
5. Back-to-Back Stem-and-Leaf Diagrams | 背靠背茎叶图
When we need to compare two related data sets, a back-to-back stem-and-leaf diagram is very effective. A common stem is placed in the centre, with leaves for one data set extending to the left and leaves for the other extending to the right. Both sides should be ordered numerically away from the stem, usually with the smallest leaves closest to the stem.
当我们需要比较两组相关的数据集时,背靠背茎叶图非常有效。将共用的茎放在中间,一组数据的叶向左伸展,另一组数据的叶向右伸展。两边都应按照离茎的远近进行数值排序,通常最小的叶离茎最近。
| Leaf (Class A) | Stem | Leaf (Class B) |
| 8, 5, 2 | 1 | 3, 6, 9 |
| 9, 4, 1 | 2 | 0, 7, 8 |
Key: 1 | 2 means 12 for Class A (left); 1 | 3 means 13 for Class B (right). This diagram allows direct visual comparison of medians, ranges, and overall shape of the two distributions.
图例:1 | 2 表示 A 班的 12(左侧);1 | 3 表示 B 班的 13(右侧)。这种图可以直接从视觉上比较两组数据的中位数、极差和分布的整体形态。
6. Measures of Central Tendency: Mean, Median, Mode | 集中趋势的度量:均值、中位数、众数
Three common averages summarise the centre of a data set. The mode is the value that appears most frequently. A set can have one mode, more than one mode (bimodal or multimodal), or no mode at all.
三种常用的平均数概括数据集的中心。众数是出现最频繁的数值。一个数据集可能有一个众数、多个众数(双峰或多峰),或者没有众数。
The median is the middle value when the data is arranged in order. For an odd number of values, it is the exact centre. For an even number, it is the mean of the two middle values. The median is not affected by extreme outliers.
中位数是将数据按顺序排列后位于中间的数值。数据个数为奇数时,它就是正中间的那个值;为偶数时,则是中间两个值的平均数。中位数不受极端异常值的影响。
The mean is calculated by summing all values and dividing by the number of values: mean = sum of all values ÷ number of values. It uses every data point, so it is influenced by outliers. For grouped data, use midpoints: estimated mean = Σ(f × midpoint) ÷ Σf.
均值的计算是将所有数值相加再除以数据的个数:均值 = 所有数值之和 ÷ 数据个数。均值用到了每个数据点,因此容易受异常值的影响。对于分组数据,使用中点估算:估计均值 = Σ(f × 中点) ÷ Σf。
Choosing the best average depends on the data and what you want to show. The mean is useful for further calculations, while the median gives a better idea of the ‘typical’ value when outliers exist.
选用哪个平均数最好取决于数据和你想展示的内容。均值便于进行进一步计算,而当存在异常值时,中位数更能体现“典型”值。
7. Range and Comparing Data Sets | 极差与数据集的比较
The range is a simple measure of spread: range = largest value – smallest value. It tells us how spread out the data is. A small range indicates consistency, while a large range suggests variability.
极差是衡量数据离散程度的一个简单指标:极差 = 最大值 – 最小值。它告诉我们数据的分散程度。极差小表示数据一致性好,极差大则说明变化大。
When comparing two data sets, always refer to both a measure of central tendency (mean or median) and a measure of spread (range). For example, ‘Data set A has a higher median, so on average values are larger. Its range is also larger, meaning the data is more spread out than data set B.’ Avoid saying one is ‘better’ without linking it to context.
比较两组数据时,务必同时提及集中趋势的度量(均值或中位数)和离散程度的度量(极差)。例如,“数据集 A 的中位数更高,因此平均而言数值更大。它的极差也更大,意味着数据比数据集 B 更分散。” 避免在没有上下文的情况下简单说某一个“更好”。
8. Introduction to Probability | 概率简介
Probability measures how likely an event is to happen. It can be written as a fraction, decimal, or percentage. The probability scale runs from 0 (impossible) to 1 (certain). The probability of an event not happening is 1 minus the probability that it does happen.
概率衡量事件发生的可能性大小。它可以写成分数、小数或百分数。概率的取值从 0(不可能)到 1(必然)。事件不发生的概率等于 1 减去它发生的概率。
For equally likely outcomes, probability = (number of favourable outcomes) ÷ (total number of outcomes). This is the theoretical probability. Example: rolling a 5 on a fair six-sided die, P(5) = 1/6.
对于等可能的结果,概率 = (有利结果的数量) ÷ (所有可能结果的总数)。这就是理论概率。例如:掷一个均匀的六面骰子得到 5 的概率,P(5) = 1/6。
The sum of probabilities of all possible outcomes of an experiment is always 1. If the probability of rain tomorrow is 0.3, the probability of no rain is 1 – 0.3 = 0.7.
一个试验中所有可能结果的概率之和总是 1。如果明天下雨的概率是 0.3,那么不下雨的概率就是 1 – 0.3 = 0.7。
9. Experimental Probability and Relative Frequency | 实验概率与相对频率
When we perform an actual experiment, the estimated probability is called relative frequency: relative frequency = (number of times event occurs) ÷ (total number of trials). As the number of trials increases, the relative frequency usually gets closer to the theoretical probability. This is known as the law of large numbers.
当我们真正进行实验时,估算出的概率叫做相对频率:相对频率 = (事件发生的次数) ÷ (试验总次数)。随着试验次数的增加,相对频率通常会越来越接近理论概率。这就是大数定律。
For example, if you toss a coin 50 times and get heads 28 times, the relative frequency of heads is 28/50 = 0.56. If you increase the number of tosses to 500, the value tends to move towards 0.5. Experimental probability is useful when theoretical probability is difficult to calculate, such as finding the probability that a drawing pin lands point up.
例如,如果抛一枚硬币 50 次,得到 28 次正面,正面的相对频率就是 28/50 = 0.56。如果把抛掷次数增加到 500 次,这个值往往会趋向 0.5。当理论概率难以计算时,实验概率就很有用,比如求图钉落地时针尖朝上的概率。
10. Venn Diagrams and Sample Spaces | 韦恩图与样本空间
A Venn diagram shows relationships between sets. A rectangle represents the universal set, and circles represent subsets. Overlapping regions indicate elements belonging to more than one set. Numbers inside regions represent frequency or probability.
韦恩图展示集合之间的关系。矩形表示全集,圆圈表示子集。重叠区域表示属于多个集合的元素。区域内的数字表示频数或概率。
In probability, Venn diagrams help visualise ‘AND’ (intersection) and ‘OR’ (union). For two events A and B, P(A ∪ B) = P(A) + P(B) – P(A ∩ B). Be careful not to double-count the intersection when adding probabilities.
在概率中,韦恩图可以帮助直观理解“与”(交集)和“或”(并集)。对于两个事件 A 和 B,P(A ∪ B) = P(A) + P(B) – P(A ∩ B)。相加概率时注意不要重复计算重叠部分。
Sample space is the set of all possible outcomes. You can list outcomes in a table or as ordered pairs. For example, when throwing two dice, the sample space can be shown as a 6×6 grid of 36 equally likely outcomes. This makes calculating combined events straightforward.
样本空间是所有可能结果的集合。可以用表格或有序数对列出结果。例如,抛两个骰子时,样本空间可以显示为一个 6×6 的网格,包含 36 个等可能的结果。这使计算组合事件变得简单直接。
11. Misleading Graphs | 误导性图表
Graphs can be drawn to mislead the reader, either deliberately or accidentally. Common tricks include starting the vertical axis at a value other than zero, which exaggerates small differences between bars. Using unequal bar widths without adjusting frequency density, or using three-dimensional effects that distort proportions, can also be misleading.
图表可能被画成误导读者,无论是有意还是无意。常见的伎俩包括让纵轴不从零开始,这会夸大条形之间的微小差异。使用不等宽的条形却不调整频率密度,或者使用扭曲比例的三维效果,也可能产生误导。
Always check the scale, labels, and whether the areas of pictures in pictograms are proportional to the frequencies they represent. Look critically at any graph and ask: does the visual impression match the numbers?
一定要检查刻度、标签,以及象形图中图形的面积是否与它们所代表的频数成比例。要批判性地审视任何图表,并问自己:视觉印象与数字是否匹配?
When creating your own graphs, always use a broken axis symbol if you must truncate the scale, and ensure your representations are honest and clear. A good statistician presents data fairly.
在制作自己的图表时,如果必须截断坐标轴,请始终使用断裂轴符号,并确保你的表示方式是诚实且清晰的。优秀的统计学家会公正地展示数据。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导