📚 Year 10 Cambridge Statistics: Case Study Practice | 剑桥十年级统计:案例分析实战演练
Real-world data analysis is a key skill in Cambridge Year 10 Statistics. This case study practises the full statistical enquiry cycle: from designing a study and collecting data, to summarising, visualising, interpreting, and evaluating findings. You will work through a realistic example on study time and test performance, reinforcing core topics such as frequency tables, averages, spread, charts, correlation, and probability.
真实世界的数据分析是剑桥十年级统计课程的关键技能。本次案例分析练习完整的统计探究循环:从设计研究、收集数据,到汇总、可视化、解释和评估结果。你将通过一个关于学习时间与测试成绩的实际例子,强化频数表、平均数、离散程度、图表、相关性和概率等核心主题。
1. Introducing the Case Study | 案例介绍
A group of Year 10 students wanted to investigate whether the number of hours spent studying maths each week is related to their maths test scores. They surveyed 30 classmates, recording each student’s weekly study time (hours) and most recent test percentage.
一组十年级学生想调查每周学习数学的小时数是否与数学测试成绩相关。他们调查了 30 位同学,记录每位学生每周的学习时间(小时)和最近一次测试的百分比成绩。
This dataset will be used throughout the case study to practise key statistical techniques. We will organise, display, and analyse the data, then draw evidence-based conclusions.
整个案例分析将使用这个数据集来练习关键的统计技术。我们将整理、展示和分析数据,然后得出有依据的结论。
2. Data Collection and Types | 数据收集与类型
The students collected primary data through a questionnaire. They decided to record study time as a continuous numerical variable (hours, measured to the nearest half-hour) and test score as a discrete numerical variable (percentage). They also noted gender, but only study time and score will be analysed here.
学生们通过问卷收集了原始数据。他们决定将学习时间记录为连续数值变量(小时,精确到半小时),将测试成绩记录为离散数值变量(百分比)。他们还记录了性别,但这里只分析学习时间和成绩。
It is important to identify the variable types: study time is continuous because it can take any value in an interval; test score is discrete if only whole percentages are recorded, but can be treated as continuous for grouping.
识别变量类型很重要:学习时间是连续的,因为它可以取区间内的任何值;测试成绩如果只记录整数百分比则为离散变量,但在分组时可视为连续变量。
3. Organising Data: Frequency Tables | 整理数据:频数表
To summarise the study time data, the students created a grouped frequency table. They chose class intervals of 0 ≤ t < 2, 2 ≤ t < 4, 4 ≤ t < 6, 6 ≤ t < 8, and 8 ≤ t ≤ 10 hours. The frequencies are shown below.
为了汇总学习时间数据,学生们制作了一个分组频数表。他们选择的组距是 0 ≤ t < 2、2 ≤ t < 4、4 ≤ t < 6、6 ≤ t < 8 和 8 ≤ t ≤ 10 小时。频数如下表所示。
| Study time (hours) | Frequency |
|---|---|
| 0 ≤ t < 2 | 4 |
| 2 ≤ t < 4 | 8 |
| 4 ≤ t < 6 | 10 |
| 6 ≤ t < 8 | 6 |
| 8 ≤ t ≤ 10 | 2 |
A similar table was made for test scores using intervals of width 10%. Grouping helps to see patterns and calculate statistics from large datasets.
测试成绩也使用组距为 10% 的区间制作了类似的表格。分组有助于发现模式,并计算大数据集的统计量。
4. Visualising Data: Histogram | 数据可视化:直方图
A histogram of the study time data was drawn with class boundaries on the horizontal axis and frequency density on the vertical axis, because the class widths are equal (2 hours). Frequency density = frequency ÷ class width, so here the heights are just the frequencies.
绘制学习时间直方图时,横轴为组界,纵轴为频数密度,因为组距相等(2 小时)。频数密度 = 频数 ÷ 组距,所以此处的高度就是频数。
The histogram shows that the modal class is 4–6 hours, and the distribution is roughly symmetric with a slight positive skew. For test scores, a histogram revealed a roughly bell-shaped distribution.
直方图显示众数所在组为 4–6 小时,分布大致对称,略有正偏态。测试成绩的直方图呈现大致钟形分布。
5. Measures of Central Tendency | 集中趋势的度量
Using the grouped frequency table, we estimated the mean study time by finding the mid-point (x) of each class, multiplying by frequency (f), summing, then dividing by total frequency (n=30).
使用分组频数表,我们通过找到每组的组中值 (x),乘以频数 (f),求和后除以总频数 (n=30),来估算平均学习时间。
Mean ≈ Σ(fx) / n
The mid-points are 1, 3, 5, 7, 9. Σ(fx) = 4×1 + 8×3 + 10×5 + 6×7 + 2×9 = 4+24+50+42+18 = 138. Mean ≈ 138 / 30 = 4.6 hours.
组中值为 1, 3, 5, 7, 9。Σ(fx) = 4×1 + 8×3 + 10×5 + 6×7 + 2×9 = 4+24+50+42+18 = 138。均值 ≈ 138 / 30 = 4.6 小时。
The median class is the one containing the 15th and 16th values. Cumulative frequencies: 4, 12, 22, 28, 30. The median falls in the 4–6 hours class. We can estimate the median using interpolation, giving approximately 4.9 hours.
中位数所在组是包含第 15 和第 16 个值的组。累计频数:4、12、22、28、30。中位数落在 4–6 小时组。通过插值估算中位数约为 4.9 小时。
The modal class is 4–6 hours, the class with the highest frequency (10).
众数所在组为 4–6 小时,该组频数最高 (10)。
6. Measures of Spread | 离散程度的度量
Spread tells us how varied the data are. The range is the difference between the highest and lowest values: maximum 10, minimum 0, so range = 10 − 0 = 10 hours.
离散程度反映数据的变异程度。极差是最大值与最小值的差:最大值 10,最小值 0,极差 = 10 − 0 = 10 小时。
Quartiles divide the sorted data into four equal parts. The lower quartile (Q₁) is the ¼(n+1)th value, median (Q₂) ½(n+1)th, upper quartile (Q₃) ¾(n+1)th. For n=30, Q₁ is at position 7.75, which lies in the 2–4 class; Q₃ at 23.25, in the 6–8 class. Interpolation gives Q₁ ≈ 2.9 h, Q₃ ≈ 6.3 h.
四分位数将排序后的数据分成四等份。下四分位数 (Q₁) 为第 ¼(n+1) 个值,中位数 (Q₂) 为第 ½(n+1) 个,上四分位数 (Q₃) 为第 ¾(n+1) 个。n=30 时,Q₁ 的位置是 7.75,落在 2–4 组;Q₃ 的位置是 23.25,落在 6–8 组。插值得出 Q₁ ≈ 2.9 h,Q₃ ≈ 6.3 h。
The interquartile range (IQR) = Q₃ − Q₁ ≈ 6.3 − 2.9 = 3.4 hours, showing the spread of the middle 50% of students.
四分位距 (IQR) = Q₃ − Q₁ ≈ 6.3 − 2.9 = 3.4 小时,显示了中间 50% 学生的分布范围。
7. Box-and-Whisker Plot | 箱线图
A box-and-whisker plot (box plot) is constructed using the five-number summary: minimum (0), Q₁ (2.9), median (4.9), Q₃ (6.3), and maximum (10). The box spans IQR, and whiskers extend to the minimum and maximum, assuming no outliers.
箱线图利用五数概括法绘制:最小值 (0)、
Published by TutorHao | Year 10 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导