Year 9 CIE Statistics: Case Study Practical Drills | 9年级 CIE 统计:案例分析实战演练

📚 Year 9 CIE Statistics: Case Study Practical Drills | 9年级 CIE 统计:案例分析实战演练

In the CIE Year 9 Statistics syllabus, one of the most valuable ways to build confidence is through real-world case studies. This article walks you through a complete statistical investigation, from collecting raw data to drawing conclusions and applying basic probability. Using a practical example – weekly reading hours of Year 9 students in two classes – you will see how each topic (tables, charts, averages, spread, and probability) links together in a single coherent analysis. By the end, you will be ready to tackle your own data investigation with clarity and precision.

在 CIE 9年级统计课程中,通过真实案例建立信心是最宝贵的学习方式之一。本文带你完成一个完整的统计调查,从收集原始数据到得出结论并应用基础概率。我们将以两个班级九年级学生每周课外阅读时数为实际例子,展示每一个知识点(表格、图表、平均数、离散程度和概率)如何在一个连贯的分析中串联起来。读完本文后,你将能够清晰、准确地独立完成自己的数据调查。


1. Case Introduction and Data Collection | 案例导入与数据收集

Our investigation begins with a simple question: ‘How many hours per week do Year 9 students spend reading for pleasure?’ To obtain representative data, we surveyed 15 students from Class A and another 15 from Class B. The data collection method was a short questionnaire, ensuring each response was a whole number of hours. Below is the raw data recorded for each group.

我们的调查从一个简单的问题开始:“九年级学生每周课外阅读多少小时?”为了获得有代表性的数据,我们调查了A班的15名学生和B班的15名学生。数据收集方法是一份简短的问卷,确保每个回答都以整小时数记录。下面是每组记录的原始数据。

Class A (hours): 0, 1, 2, 2, 3, 3, 3, 4, 4, 5, 5, 6, 7, 8, 10

A班(小时): 0, 1, 2, 2, 3, 3, 3, 4, 4, 5, 5, 6, 7, 8, 10

Class B (hours): 1, 1, 2, 2, 3, 4, 4, 5, 5, 6, 7, 8, 9, 10, 12

B班(小时): 1, 1, 2, 2, 3, 4, 4, 5, 5, 6, 7, 8, 9, 10, 12

Remember that clear data collection is the first step to reliable statistics. Always note the unit (hours) and ensure the values are recorded accurately before any analysis begins.

请记住,清晰的数据收集是可靠统计的第一步。在开始任何分析之前,务必注明单位(小时)并确保数值记录准确。


2. Organising Data: Sorting and Stem-and-Leaf Diagram | 整理数据:排序与茎叶图

Once raw data is collected, the next step is to organise it. Sorting the values in ascending order helps us quickly identify the minimum, maximum, and centre. We then display the data using a stem‑and‑leaf diagram, which preserves every original value while showing the distribution shape.

收集到原始数据后,下一步是整理数据。将数值按升序排列有助于快速识别最小值、最大值和中心位置。然后我们使用茎叶图来展示数据,它既能保留每个原始数值,又能显示分布形态。

Sorted Class A: 0, 1, 2, 2, 3, 3, 3, 4, 4, 5, 5, 6, 7, 8, 10

排序后的A班: 0, 1, 2, 2, 3, 3, 3, 4, 4, 5, 5, 6, 7, 8, 10

Sorted Class B: 1, 1, 2, 2, 3, 4, 4, 5, 5, 6, 7, 8, 9, 10, 12

排序后的B班: 1, 1, 2, 2, 3, 4, 4, 5, 5, 6, 7, 8, 9, 10, 12

The stem‑and‑leaf diagram for each class uses the tens digit as the stem and the units digit as the leaf. A key is provided to avoid confusion.

每个班级的茎叶图以十位数字为茎,个位数字为叶。我们提供图例以避免混淆。

Class A Stem Leaf Class B Stem Leaf
0 0 1 2 2 3 3 3 4 4 5 5 6 7 8 0 1 1 2 2 3 4 4 5 5 6 7 8 9
1 0 1 0 2

Key for Class A: 0|3 means 3 hours. Key for Class B: 1|2 means 12 hours.

A班图例:0|3 表示 3 小时。B班图例:1|2 表示 12 小时。

From the diagrams we can already see that Class B has more higher values (leaves in stem 1) and no zero‑hour entry, whereas Class A has one student who reads 0 hours and a more spread‑out distribution in stem 0.

从图中我们已经可以看到,B班有更多高数值(茎1中的叶),并且没有零小时读数,而A班有一个学生阅读0小时,茎0中的分布也更分散。


3. Frequency Distribution Table | 频率分布表

To make the data even clearer, we group the hours into intervals. For continuous‑like data such as time, it is common to use equal class intervals. Here we choose intervals 0–2, 3–5, 6–8 and 9–12 (where boundaries are set so each value falls into only one group). A frequency table counts how many students fall into each interval.

为了让数据更加清晰,我们将小时数分组。对于时间这类连续型数据,通常使用相等的组距。这里我们选择区间 0–2, 3–5, 6–8 和 9–12(边界设定确保每个值只落入一个组)。频率表统计每个区间内的学生人数。

Hours (interval) Class A frequency Class B frequency
0–2 4 4
3–5 7 5
6–8 3 4
9–12 1 2

This grouped frequency table immediately reveals that Class A has more students in the 3–5 hour bracket, whereas Class B is slightly more evenly spread across the higher intervals. Such tables form the foundation for drawing frequency polygons and histograms.

这个分组频率表立刻显示出,A班在 3–5 小时区间有更多学生,而 B 班在较高区间分布更均匀。这类表格是绘制频数多边形和直方图的基础。


4. Data Visualisation: Bar Chart and Frequency Polygon | 数据可视化:条形图和频数多边形

Visual displays help us compare the two groups at a glance. Because our intervals are of equal width, we can use a multiple bar chart to show both classes side by side. Alternatively, a frequency polygon (using the midpoint of each interval) is excellent for comparing the shapes of the two distributions.

可视化展示帮助我们将两组数据一目了然地比较。由于我们的区间宽度相等,我们可以使用复式条形图并列显示两个班级。或者,频数多边形(使用每个区间的中点)非常适合比较两个分布的形状。

The midpoints for the intervals are: 0–2 midpoint = 1; 3–5 midpoint = 4; 6–8 midpoint = 7; 9–12 midpoint = 10.5.

各区间的中点分别为:0–2 中点 = 1;3–5 中点 = 4;6–8 中点 = 7;9–12 中点 = 10.5。

If we were to sketch the frequency polygon, we would plot the midpoints on the horizontal axis and the frequency on the vertical axis, connecting the points with straight lines. Class A’s polygon peaks at midpoint 4, whereas Class B’s polygon rises again at midpoint 10.5, indicating a second smaller peak. Always label axes clearly: ‘Hours (midpoint)’ and ‘Frequency’.

如果我们绘制频数多边形,我们会将中点放在横轴上,频率放在纵轴上,并用直线连接各点。A班的多边形在中点4处达到峰值,而B班的多边形在中点10.5处再次上升,显示出第二个较小的峰值。始终清晰地标记坐标轴:“小时(中点)”和“频率”。

Choosing the right visual display depends on your data type. For discrete or categorical data a bar chart is suitable; for continuous data grouped into intervals, a frequency polygon or histogram is preferred.

选择合适的可视化方式取决于数据类型。对于离散或分类数据,条形图较合适;对于分组为区间的连续数据,优选频数多边形或直方图。


5. Measures of Central Tendency: Mode, Median, and Mean | 集中趋势的度量:众数、中位数、均值

To describe the ‘typical’ reading hours we calculate the three main averages: mode (most frequent), median (middle value), and mean (arithmetic average). These summaries condense the data into single representative values.

为了描述“典型”的阅读小时数,我们计算三种主要的平均数:众数(出现最多的值)、中位数(中间值)和均值(算术平均值)。这些汇总统计量将数据浓缩为单一的代表值。

For Class A (ordered data): 0,1,2,2,3,3,3,4,4,5,5,6,7,8,10.

A班(排序数据):0,1,2,2,3,3,3,4,4,5,5,6,7,8,10。

The mode is 3 hours, as it appears three times. The median is the 8th value (with n=15): that is 4 hours. The mean is calculated as:

众数为 3 小时,因为它出现了三次。中位数为第8个值(n=15):即 4 小时。均值计算如下:

Meanₐ = (0+1+2+2+3+3+3+4+4+5+5+6+7+8+10) ÷ 15 = 63 ÷ 15 = 4.2 hours

A班均值 = (0+1+2+2+3+3+3+4+4+5+5+6+7+8+10) ÷ 15 = 63 ÷ 15 = 4.2 小时

For Class B (ordered): 1,1,2,2,3,4,4,5,5,6,7,8,9,10,12.

B班(排序):1,1,2,2,3,4,4,5,5,6,7,8,9,10,12。

Class B has multiple modes: 1, 2, 4 and 5 each appear twice, so we say it has no unique mode. The median (8th value) is 5 hours. The mean is:

B班有多个众数:1, 2, 4, 5 均出现两次,因此我们说没有唯一的众数。中位数(第8个值)为 5 小时。均值为:

Mean_b = (1+1+2+2+3+4+4+5+5+6+7+8+9+10+12) ÷ 15 = 79 ÷ 15 ≈ 5.27 hours

B班均值 = (1+1+2+2+3+4+4+5+5+6+7+8+9+10+12) ÷ 15 = 79 ÷ 15 ≈ 5.27 小时

Notice that the mean for Class B is higher than Class A’s mean, and the medians follow the same pattern. However, the presence of a very low value (0) in Class A pulls its mean down. Always consider which average best represents your data.

注意,B班的均值高于A班,中位数也是如此。然而,A班中极低值(0)的存在拉低了其均值。始终要考虑哪个平均数最能代表你的数据。


6. Measures of Spread: Range and Interquartile Range | 离散程度的度量:极差和四分位数范围

Averages alone do not tell the whole story; we need to measure how spread out the data are. The simplest spread measure is the range: maximum minus minimum. For a more robust view, we use the interquartile range (IQR), which captures the middle 50% of data.

仅靠平均数并不能说明全部问题;我们需要衡量数据的离散程度。最简单的离散度量是极差:最大值减去最小值。为了得到更稳健的视角,我们使用四分位数范围(IQR),它捕捉了中间50%的数据。

For Class A: Minimum = 0, Maximum = 10, so Range = 10 – 0 = 10 hours. To find the IQR: Q₁ is the 4th value = 2 hours; Q₃ is the 12th value = 6 hours. Therefore IQR = Q₃ – Q₁ = 6 – 2 = 4 hours.

A班:最小值 = 0,最大值 = 10,因此极差 = 10 – 0 = 10 小时。求 IQR:Q₁ 为第4个值 = 2 小时;Q₃ 为第12个值 = 6 小时。因此 IQR = Q₃ – Q₁ = 6 – 2 = 4 小时。

For Class B: Minimum = 1, Maximum = 12, so Range = 12 – 1 = 11 hours. Q₁ (4th value) = 2 hours; Q₃ (12th value) = 8 hours. IQR = 8 – 2 = 6 hours.

B班:最小值 = 1,最大值 = 12,因此极差 = 12 – 1 = 11 小时。Q₁(第4个值)= 2 小时;Q₃(第12个值)= 8 小时。IQR = 8 – 2 = 6 小时。

The larger IQR for Class B tells us that the middle half of Class B students have more variable reading habits compared to Class A. Range alone can be affected by outliers, but IQR is resistant to them.

B班较大的 IQR 告诉我们,B班中间一半学生的阅读习惯比A班更分散。仅凭极差可能受异常值影响,但IQR对此具有抵抗力。


7. Comparing Two Data Sets: Box-and-Whisker Plots | 比较两组数据:箱线图

A box-and-whisker plot (box plot) uses the five‑number summary – minimum, Q₁, median, Q₃, maximum – to create a compact visual comparison. Drawing both box plots on the same scale makes differences in centre, spread, and skewness immediately visible.

箱线图(盒须图)使用五数概括——最小值、Q₁、中位数、Q₃、最大值——创建一个紧凑的可视化对比。在同一尺度上绘制两个箱线图,可以立即看出中心、离散度和偏斜度的差异。

Class A five‑number summary: Min=0, Q₁=2, Median=4, Q₃=6, Max=10.

A班五数概括:最小值=0,Q₁=2,中位数=4,Q₃=6,最大值=10。

Class B five‑number summary: Min=1, Q₁=2, Median=5

Published by TutorHao | Year 9 统计 Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading