📚 Case Study Showdown: Applying GCSE Statistics to a School Survey | 案例分析实战:学校调查中的 GCSE 统计应用
Welcome to this real-world GCSE Statistics case study. We will follow a Year 10 investigation into the relationship between students’ sleep duration and their test scores, applying the core concepts you need for Edexcel. By working through data collection, graphical displays, averages, measures of spread, correlation and basic probability, you will see exactly how statistical techniques are used to make sense of everyday information.
欢迎来到这个真实的 GCSE 统计案例研究。我们将跟随一项 Year 10 的调查,探究学生睡眠时长与考试成绩之间的关系,运用你为 Edexcel 考试所需的核心概念。通过处理数据收集、图表展示、平均数、离散度量、相关性和基本概率,你将清晰地看到如何运用统计方法理解日常信息。
1. Setting the Scene: The School Sleep & Score Survey | 场景设定:学校睡眠与成绩调查
A group of Year 10 students wondered whether getting more sleep the night before a test could improve performance. They decided to carry out a statistical investigation. The question was: “Is there a relationship between the average number of hours of sleep a student gets each night and their most recent Mathematics test score (out of 100)?” They planned to collect data from 20 students in their year group.
一群 Year 10 学生想知道考试前一晚睡得更久是否有助于提高成绩。他们决定开展一项统计调查。问题是:“学生每晚平均睡眠时长与他们最近一次的数学测验成绩(百分制)之间是否存在关系?”他们计划从同年级中收集 20 名学生的数据。
To keep the project manageable, they defined “average sleep” as the typical number of hours per night over the previous two weeks, self-reported to the nearest half hour. The test score was the most recent end-of-topic assessment. They also recorded the year group and whether the student was in set 1, 2 or 3 for Mathematics, to ensure a fair mix.
为使项目可控,他们将“平均睡眠”定义为过去两周每晚的典型睡眠时长,以半小时为精度自报。测验成绩是最近一次单元结束测评。他们还记录了年级以及学生在数学课归属的 1、2 或 3 组,以保证混合均匀。
2. Data Collection and Sampling | 数据收集与抽样
The students used a stratified sampling method. They divided the Year 10 cohort into three strata according to Mathematics sets, then randomly selected a proportional number of pupils from each set. This ensured the sample represented different ability levels. Each selected student completed a short, anonymised questionnaire.
学生们采用了分层抽样方法。他们将 Year 10 群体按照数学组别分为三个层,然后从每组中按比例随机抽取学生。这确保样本代表了不同的能力水平。每位被选中的学生完成了一份简短的匿名问卷。
Here are the raw data they collected. Take a moment to scan the table: Student ID, Sleep (hours), and Test Score (marks out of 100).
以下为他们收集的原始数据。请浏览表格:学生编号、睡眠时长(小时)和测验成绩(百分制)。
| ID | Sleep (h) | Score (marks) |
|---|---|---|
| 1 | 7.5 | 68 |
| 2 | 6.0 | 54 |
| 3 | 8.0 | 82 |
| 4 | 5.5 | 48 |
| 5 | 9.0 | 95 |
| 6 | 7.0 | 71 |
| 7 | 6.5 | 59 |
| 8 | 8.5 | 88 |
| 9 | 5.0 | 40 |
| 10 | 7.0 | 69 |
| 11 | 8.0 | 84 |
| 12 | 6.0 | 52 |
| 13 | 9.5 | 91 |
| 14 | 7.5 | 73 |
| 15 | 6.5 | 56 |
| 16 | 8.5 | 86 |
| 17 | 7.0 | 66 |
| 18 | 5.5 | 45 |
| 19 | 9.0 | 90 |
| 20 | 8.0 | 80 |
The students checked for any obvious errors or missing values – luckily all 20 responses were complete. This data set is now ready for analysis.
学生们检查了是否存在明显错误或缺失值——幸运的是所有 20 份回复均完整。这份数据集已准备好进行分析。
3. Organising Data: Frequency Tables | 数据整理:频率表
Before drawing graphs, the group decided to organise the sleep data into a grouped frequency table. They chose class intervals of equal width: 5.0 ≤ h < 6.0, 6.0 ≤ h < 7.0, and so on up to 9.0 ≤ h < 10.0. The tally and frequencies are shown below.
在绘制图表之前,小组决定将睡眠数据整理成组距式频率表。他们选择了等宽的组距:5.0 ≤ h < 6.0、6.0 ≤ h < 7.0,依此类推直至 9.0 ≤ h < 10.0。下面是划记与频数。
| Sleep, h (hours) | Tally | Frequency |
|---|---|---|
| 5.0 ≤ h < 6.0 | || | 2 |
| 6.0 ≤ h < 7.0 | |||| | 4 |
| 7.0 ≤ h < 8.0 | |||| || | 7 |
| 8.0 ≤ h < 9.0 | |||| | 4 |
| 9.0 ≤ h < 10.0 | ||| | 3 |
Grouping helps us see the distribution. Most students sleep between 7.0 and 8.0 hours per night. We can also add a cumulative frequency column for later use, but first let’s visualise the data.
分组有助于我们看清分布。大多数学生每晚睡眠在 7.0 到 8.0 小时之间。我们还可以添加累积频率列以备后用,但首先让我们将数据可视化。
4. Graphical Displays: Histograms and Box Plots | 图形展示:直方图与箱线图
Because the class intervals are of equal width, a straightforward histogram can be drawn. The horizontal axis shows sleep hours, and the vertical axis shows frequency. Each bar is drawn with height equal to the frequency. The histogram would show a peak at the 7.0–8.0 h interval, with the data slightly skewed to the left (more students sleeping longer).
由于组距等宽,可直接绘制直方图。横轴表示睡眠时长,纵轴表示频数。每一条形的高度等于对应的频数。直方图会显示 7.0–8.0 小时区间出现峰值,数据略微负偏态(更多学生睡眠时间更长)。
For the test scores, the group decided to construct a box plot using the five-number summary. They sorted the 20 scores: 40, 45, 48, 52, 54, 56, 59, 66, 68, 69, 71, 73, 80, 82, 84, 86, 88, 90, 91, 95. The minimum is 40, the lower quartile Q₁ is the median of the first half (54), the median Q₂ is (69+71)/2 = 70, the upper quartile Q₃ is 86, and the maximum is 95. The interquartile range (IQR) = 86 – 54 = 32.
对于测验分数,小组决定利用五数概括绘制箱线图。他们将 20 个分数排序:40, 45, 48, 52, 54, 56, 59, 66, 68, 69, 71, 73, 80, 82, 84, 86, 88, 90, 91, 95。最小值为 40,下四分位数 Q₁ 为前一半的中位数(54),中位数 Q₂ 为 (69+71)/2 = 70,上四分位数 Q₃ 为 86,最大值为 95。四分位距 (IQR) = 86 – 54 = 32。
The box plot shows a fairly symmetrical distribution with no extreme outliers (any value below Q₁ – 1.5×IQR = 6 or above Q₃ + 1.5×IQR = 134 would be outliers; none exist). Thus the spread of scores is consistent.
箱线图显示分布大致对称,没有极端离群值(低于 Q₁ – 1.5×IQR = 6 或高于 Q₃ + 1.5×IQR = 134 的值才算离群值;此处不存在)。因此分数的离散情况较为一致。
5. Measuring Central Tendency | 集中趋势测量
Calculating the mean, median and mode for both variables gives us a clearer picture. For sleep hours: the raw data sum = 7.5+6.0+8.0+5.5+9.0+7.0+6.5+8.5+5.0+7.0+8.0+6.0+9.5+7.5+6.5+8.5+7.0+5.5+9.0+8.0 = 146.5 hours. The mean sleep = 146.5 ÷ 20 = 7.325 hours. The median sleep (using the ordered list) is the average of the 10th and 11th values: both are 7.5, so median = 7.5 h. The modal class is 7.0 ≤ h < 8.0 h, and the mode of raw data is 7.0 h (occurs 3 times).
计算两个变量的均值、中位数和众数可使我们更清晰。睡眠时长:原始数据总和 = 7.5+6.0+8.0+5.5+9.0+7.0+6.5+8.5+5.0+7.0+8.0+6.0+9.5+7.5+6.5+8.5+7.0+5.5+9.0+8.0 = 146.5 小时。睡眠均值 = 146.5 ÷ 20 = 7.325 小时。睡眠中位数(使用排序列表)为第 10 和第 11 个值的平均数:两者均为 7.5,因此中位数 = 7.5 小时。众数所在组为 7.0 ≤ h < 8.0 小时,原始数据的众数是 7.0 小时(出现 3 次)。
For test scores: sum = 68+54+82+48+95+71+59+88+40+69+84+52+91+73+56+86+66+45+90+80 = 1397. Mean = 1397 ÷ 20 = 69.85 marks. The median is 70 marks. Because the mean and median are very close, the distribution is roughly symmetric.
测验成绩:总和 = 68+54+82+48+95+71+59+88+40+69+84+52+91+73+56+86+66+45+90+80 = 1397。均值 = 1397 ÷ 20 = 69.85 分。中位数为 70 分。由于均值和中位数非常接近,分布大致对称。
Which measure is better? The mean uses all values, but the median is robust to extreme scores. Here both tell a similar story.
哪一种度量更好?均值使用了所有数值,但中位数对极端分数具有稳健性。这里两者讲述的故事相似。
6. Measuring Spread: Range, IQR and Standard Deviation | 离散度测量:极差、四分位距与标准差
Spread tells us how consistent the data are. For sleep hours: range = 9.5 – 5.0 = 4.5 h. But range is sensitive to extremes. The IQR, as computed from the sleep data, requires ordering. The ordered sleep list: 5.0, 5.5, 5.5, 6.0, 6.0, 6.5, 6.5, 7.0, 7.0, 7.0, 7.5, 7.5, 8.0, 8.0, 8.0, 8.5, 8.5, 9.0, 9.0, 9.5. Q₁ = median of first 10 values = (6.0+6.5)/2 = 6.25 h; Q₃ = median of last 10 values = (8.0+8.5)/2 = 8.25 h. So IQR = 8.25 – 6.25 = 2.0 h.
离散度告诉我们数据的一致性。睡眠时长:极差 = 9.5 – 5.0 = 4.5 小时。但极差对极端值敏感。根据睡眠数据计算四分位距需要排序。排序后睡眠列表:5.0, 5.5, 5.5, 6.0, 6.0, 6.5, 6.5, 7.0, 7.0, 7.0, 7.5, 7.5, 8.0, 8.0, 8.0, 8.5, 8.5, 9.0, 9.0, 9.5。Q₁ = 前 10 个值的中位数 = (6.0+6.5)/2 = 6.25 小时;Q₃ = 后 10 个值的中位数 = (8.0+8.5)/2 = 8.25 小时。因此 IQR = 8.25 – 6.25 = 2.0 小时。
The standard deviation is the most comprehensive measure. We use the sample standard deviation formula s = √[ Σ(x – x̄)²/(n-1) ]. For sleep: x̄ = 7.325, n = 20. The deviations and squared deviations can be computed (or using a calculator). The sum of squared deviations from the mean works out to be 30.5375 (approx). Dividing by 19 gives ≈1.607, and √1.607 ≈ 1.268 h. So the sleep hours typically vary by about 1.27 h from the mean.
标准差是最全面的度量。我们使用样本标准差公式 s = √[ Σ(x – x̄)²/(n-1) ]。对于睡眠数据:x̄ = 7.325,n = 20。可以计算离差和离差平方(或使用计算器)。离均差平方和约为 30.5375。除以 19 得到约 1.607,开方得 √1.607 ≈ 1.268 小时。因此睡眠时长通常与均值相差约 1.27 小时。
For scores: range = 95 – 40 = 55 marks. Ordered scores Q₁=54, Q₃=86, IQR=32. Standard deviation (sample) can be found similarly; sum of squared deviations ≈ 4725.55, s = √(4725.55/19) ≈ √248.71 ≈ 15.77 marks. So scores are spread out by about 16 marks on average.
对于分数:极差 = 95 – 40 = 55 分。排序后 Q₁=54,Q₃=86,IQR=32。同样可求出样本标准差;离差平方和约为 4725.55,s = √(4725.55/19) ≈ √248.71 ≈ 15.77 分。因此分数平均分散约 16 分。
7. Cumulative Frequency and Percentiles | 累积频率与百分位数
A cumulative frequency graph is useful for estimating medians and percentiles. Using the grouped sleep data, we add a cumulative frequency column. The upper class boundaries are 6.0, 7.0, 8.0, 9.0, 10.0. The cumulative frequencies: 2, 2+4=6, 6+7=13, 13+4=17, 17+3=20. Plotting these points and drawing a smooth curve, we can read off the median (50th percentile) at the 10.5th value, which lies around 7.4 h. The 25th percentile (Q₁) is at the 5.25th observation: about 6.4 h, and the 75th percentile (Q₃) at the 15.75th observation: about 8.5 h. These slightly differ from the raw data quartiles because of grouping.
累积频率图有助于估计中位数和百分位数。使用分组的睡眠数据,我们添加一列累积频率。上组界为 6.0、7.0、8.0、9.0、10.0。累积频率:2,2+4=6,6+7=13,13+4=17,17+3=20。绘制这些点并画出平滑曲线,可以读取中位数(第 50 百分位数)在第 10.5 个值处,约 7.4 小时。第 25 百分位数(Q₁)在第 5.25 个观测值处:约 6.4 小时,第 75 百分位数(Q₃)在第 15.75 个观测值处:约 8.5 小时。这些与原始数据四分位数略有不同,因为分组造成的。
The technique reinforces how grouped data slightly smooths the detail. For the test scores (ungrouped), we can use the sorted list to say that a student scoring 73 is at the 60th percentile (since 12 out of 20 scores are ≤73, giving a percentile rank of (12/20)×100 = 60).
这一技巧强化了分组数据如何略微平滑细节。对于测验分数(未分组),可以利用排序列表说明:得分为 73 的学生处于第 60 百分位数(因为 20 个分数中有 12 个 ≤73,百分位等级 = (12/20)×100 = 60)。
8. Scatter Diagrams and Correlation | 散点图与相关性
To investigate the link between sleep and scores, the students plotted a scatter graph with sleep on the x-axis and test score on the y-axis. Each point represents one student. By inspecting the pattern, they noticed a fairly strong positive correlation: as sleep increases, scores tend to go up. The points cluster along an upward-sloping imaginary line.
为探究睡眠与分数之间的联系,学生们绘制了散点图,以睡眠时长为 x 轴、测验成绩为 y 轴。每一点代表一名学生。通过观察分布模式,他们注意到存在相当强的正相关:睡眠越多,分数往往越高。各点沿着一条上倾的假想线聚集。
To quantify this, Edexcel GCSE Statistics introduces Spearman’s rank correlation coefficient (rₛ). First rank sleep and score separately (tied ranks receive the average). Then compute the difference d between ranks for each student, square them, and use the formula:
为了量化,Edexcel GCSE 统计引入了斯皮尔曼等级相关系数 (rₛ)。首先分别对睡眠和分数排序(相同数据取平均秩次)。然后计算每个学生秩次差 d,将其平方,并代入公式:
rₛ = 1 – (6 Σ d²) / [n(n² – 1)]
Let’s do a quick calculation. For simplicity, using the data: rank 1 for shortest sleep (5.0 h), rank 20 for 9.5 h. After ranking scores similarly, we get Σ d² ≈ 22.5 (actual value depending on tie handling). Then with n=20, rₛ = 1 – (6×22.5) / (20×399) = 1 – 135/7980 ≈ 1 – 0.0169 ≈ 0.983. This indicates a very strong, near-perfect positive correlation. (The high value is partly due to the constructed dataset; in reality, caution is needed.)
我们快速计算一下。为简便起见,使用数据:睡眠最短(5.0 小时)秩次为 1,9.5 小时秩次为 20。分数类似排序后,得到 Σ d² ≈ 22.5(实际值取决于并列处理)。然后 n=20,rₛ = 1 – (6×22.5) / (20×399) = 1 – 135/7980 ≈ 1 – 0.0169 ≈ 0.983。这表明存在极强的、近乎完美的正相关。(高数值部分源于构造的数据集;现实中需谨慎。)
A scatter graph also lets us draw a line of best fit by eye. While a full regression line is beyond GCSE, we can predict that a student with 7.0 h sleep might score around 70 marks.
散点图还允许我们通过目测画出最佳拟合线。尽管完整回归线超出 GCSE 范围,但我们可以预测睡眠 7.0 小时的学生得分大约在 70 分。
9. Basic Probability from Data | 从数据中计算基本概率
We can use the sample data to estimate probabilities, assuming the 20 students are representative of the whole Year 10. For example, what is the probability that a randomly chosen Year 10 student sleeps less than 7 hours? From the table, 2 + 4 = 6 students sleep under 7 h. So P(sleep < 7) = 6/20 = 0.3.
我们可以用样本数据估计概率,假定这 20 名学生能代表整个 Year 10。例如,随机选取一名 Year 10 学生,其睡眠不足 7 小时的概率是多少?根据表格,2 + 4 = 6 名学生睡眠少于 7 小时。因此 P(睡眠 < 7) = 6/20 = 0.3
Published by TutorHao | Year 10 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导