📚 Case Study Practical Exercises | 案例分析实战演练
Welcome to this CIE Statistics case study walkthrough. We will examine a realistic scenario: investigating the relationship between the daily screen time of Year 11 students and their performance in a recent mathematics test. By working through this exercise step by step, you will see how to collect data, summarise it with graphs and numbers, handle probability questions, and explore correlation and regression — all essential skills for the CIE Statistics examination.
欢迎来到这篇CIE统计案例分析实战演练。我们将研究一个真实场景:调查11年级学生每日屏幕使用时间与近期数学测试成绩之间的关系。通过逐步完成这个练习,你将学会如何收集数据、用图表和数字进行概括、处理概率问题、探索相关与回归——这些都是CIE统计考试的核心技能。
1. Understanding the Context | 理解背景
Our school wants to find out whether there is a link between how much time students spend on their phones and their academic results. A random sample of 30 Year 11 students was selected. Each student reported their average daily screen time in hours (from a phone usage tracker) and provided their most recent maths test score out of 50. The goal is to analyse this sample and draw conclusions about the whole year group, while reflecting on the limitations of the data.
学校想知道学生使用手机的时间与学业成绩之间是否存在联系。我们从11年级随机抽取了30名学生。每个学生提供了手机使用追踪器记录的每日平均屏幕时间(小时)和最近一次数学测试的成绩(满分50分)。我们的目标是分析这个样本,对整个年级的情况做出推断,并反思数据的局限性。
2. Data Collection and Sampling Method | 数据收集与抽样方法
The students were chosen using a simple random sampling technique: each of the 200 Year 11 students was assigned a number from 1 to 200, and 30 numbers were drawn using a random number generator. This method reduces bias and gives every student an equal chance of being selected. However, the data on screen time is self‑reported, which could lead to measurement bias if some students understate their actual phone usage.
选择学生时采用了简单随机抽样方法:将200名11年级学生从1到200编号,使用随机数生成器抽取30个号码。这种方法减少了偏差,让每个学生都有均等被选中的机会。然而,屏幕使用时间的数据是学生自己报告的,如果有些人低报了实际使用时间,就可能导致测量偏差。
3. Organising the Raw Data | 整理原始数据
The collected data for screen time (hours per day) and test scores (out of 50) are shown in the table below. To make analysis easier, we will construct frequency tables for both variables separately, grouping screen time into intervals of width 1 hour and test scores into intervals of width 10 marks.
收集到的屏幕使用时间(小时/天)和测试成绩(满分50分)数据如下表所示。为了便于分析,我们将为两个变量分别制作频率表,把屏幕时间按宽度1小时分组,测试成绩按宽度10分分组。
| Screen time (h) | 4.2 | 5.6 | 3.1 | 6.8 | 4.9 | 2.3 | 5.1 | 7.2 | 3.8 | 6.0 |
| Test score | 32 | 28 | 45 | 22 | 36 | 48 | 29 | 18 | 40 | 24 |
(The full dataset for all 30 students continues in the same pattern; for brevity we show only 10 rows here. In a real exam you would be given all the values.)
(所有30名学生的完整数据集按相同模式继续;为简洁起见这里只展示了10行。在实际考试中你会得到全部数据。)
We group the screen time data into classes: 2.0–2.9, 3.0–3.9, 4.0–4.9, etc., and the test scores into 0–9, 10–19, 20–29, 30–39, 40–50. Then we count the frequencies.
我们将屏幕时间数据分组:2.0–2.9、3.0–3.9、4.0–4.9等,将测试成绩分为0–9、10–19、20–29、30–39、40–50。然后计算频数。
4. Displaying Data with Charts | 用图表展示数据
For the grouped screen time data, a histogram is appropriate because the classes have equal widths. The horizontal axis shows screen time intervals, and the vertical axis shows frequency density (since frequency equals density when widths are equal). For the test scores we can draw a cumulative frequency curve to estimate medians and quartiles.
对于分组的屏幕时间数据,适合使用直方图,因为组距相等。横轴表示屏幕时间区间,纵轴表示频数密度(由于组距相等,频数密度就是频数)。对于测试成绩,我们可以绘制累积频数曲线,用来估计中位数和四分位数。
In our histogram, the modal class for screen time is 5.0–5.9 hours, while the cumulative frequency graph for test scores shows that the median score is 31 out of 50. The interquartile range gives a measure of spread.
在我们的直方图中,屏幕时间的众数组是5.0–5.9小时,而测试成绩的累积频数图显示中位数为31分(满分50分)。四分位距提供了离散程度的度量。
5. Measures of Central Tendency | 集中趋势的度量
For ungrouped data, the mean screen time x̄ = Σx / n. Calculating from the full dataset gives a mean of 5.04 hours per day. The median screen time (middle value when sorted) is 5.05 hours. The mode is 5.2 hours, but since it is grouped, we use the modal class. For test scores, the mean is 31.3 marks, the median is 31, and the mode is 36.
对于未分组数据,平均屏幕时间 x̄ = Σx / n。根据完整数据集计算得出均值为每天5.04小时。屏幕时间的中位数(排序后的中间值)为5.05小时。众数是5.2小时,但因为数据被分组,我们使用众数组。对于测试成绩,均值为31.3分,中位数为31分,众数为36分。
Comparing these measures helps to identify the shape of the distribution. Both screen time and test scores appear roughly symmetric since the mean, median and mode are close to each other.
比较这些度量有助于判断分布形状。屏幕时间和测试成绩的平均数、中位数和众数都很接近,表明分布大致对称。
6. Measures of Dispersion | 离散程度的度量
To understand the spread, we compute the range, interquartile range (IQR), and standard deviation. The range of screen time is 7.2 – 2.3 = 4.9 hours. The lower quartile Q₁ is 3.8 h, the upper quartile Q₃ is 6.2 h, so IQR = Q₃ – Q₁ = 2.4 h. This tells us the middle 50% of students spend between 3.8 and 6.2 hours on their phones daily.
为了了解离散程度,我们计算了极差、四分位距(IQR)和标准差。屏幕时间的极差为7.2 – 2.3 = 4.9小时。下四分位数 Q₁ 为3.8小时,上四分位数 Q₃ 为6.2小时,因此 IQR = Q₃ – Q₁ = 2.4小时。这表明中间50%的学生每天使用手机的时间在3.8到6.2小时之间。
The standard deviation for screen time is approximately 1.42 h. For test scores the standard deviation is 8.9 marks, showing a wider spread relative to the mean. These figures are key when we later build a regression model.
屏幕时间的标准差约为1.42小时。测试成绩的标准差为8.9分,相对于均值而言离散程度更大。这些数字在我们后续建立回归模型时很关键。
7. Probability from the Data | 基于数据的概率
If we select one student at random from the sample, what is the probability that they spend more than 6 hours on their phone? Counting from the frequency table, 9 out of 30 students exceed 6 hours. So P(screen time > 6) = 9/30 = 0.3. Similarly, the probability that a randomly chosen student scores above 40 is 7/30 ≈ 0.233.
如果从样本中随机抽取一名学生,其屏幕使用时间超过6小时的概率是多少?根据频数表统计,30名学生中有9人超过6小时,因此 P(屏幕时间 > 6) = 9/30 = 0.3。类似地,随机选出一名学生其成绩高于40分的概率为7/30 ≈ 0.233。
We can also estimate combined probabilities. Assuming the events are independent (which may not be true in reality), P(screen time > 6 AND score > 40) = 0.3 × 0.233 = 0.07. Later, we will check whether the data actually supports independence.
我们也可以估计联合概率。假设事件相互独立(现实中可能不成立),P(屏幕时间 > 6 且 成绩 > 40) = 0.3 × 0.233 = 0.07。稍后我们将检验数据是否真的支持独立性。
8. Correlation Analysis | 相关性分析
A scatter diagram of screen time against test score reveals a downward pattern: as screen time increases, test scores tend to decrease. To quantify this, we calculate Pearson’s product‑moment correlation coefficient, r.
屏幕时间与测试成绩的散点图显示了一种向下的模式:随着屏幕时间增加,测试成绩往往下降。为了量化这种关系,我们计算皮尔逊积矩相关系数 r。
Using the formula r = Σ(x – x̄)(y – ȳ) / √[Σ(x – x̄)² Σ(y – ȳ)²], we obtain r ≈ –0.64. This indicates a moderately strong negative linear correlation. The value r² = 0.4096 suggests that about 41% of the variation in test scores can be explained by the variation in screen time.
使用公式 r = Σ(x – x̄)(y – ȳ) / √[Σ(x – x̄)² Σ(y – ȳ)²],我们得到 r ≈ –0.64。这表明存在中等偏强的负线性相关。r² = 0.4096 意味着测试成绩中大约41%的变异可以由屏幕时间的变化来解释。
9. Regression Line and Predictions | 回归线与预测
The equation of the least squares regression line of test score (y) on screen time (x) is found using y – ȳ = r (s_y / s_x) (x – x̄). Plugging in the values: ȳ = 31.3, x̄ = 5.04, s_y = 8.9, s_x = 1.42, r = –0.64.
测试成绩(y)对屏幕时间(x)的最小二乘回归线方程由 y – ȳ = r (s_y / s_x) (x – x̄) 求得。代入数值:ȳ = 31.3,x̄ = 5.04,s_y = 8.9,s_x = 1.42,r = –0.64。
The gradient b = r × (s_y / s_x) = –0.64 × (8.9 / 1.42) ≈ –4.01. The intercept a = ȳ – b x̄ = 31.3 – (–4.01)×5.04 ≈ 51.5. So the regression equation is y ≈ 51.5 – 4.01x.
斜率 b = r × (s_y / s_x) = –0.64 × (8.9 / 1.42) ≈ –4.01。截距 a = ȳ – b x̄ = 31.3 – (–4.01)×5.04 ≈ 51.5。因此回归方程为 y ≈ 51.5 – 4.01x。
This model predicts that for a student spending 5 hours on their phone, the expected maths score is 51.5 – 4.01×5 = 31.45 marks. However, predictions beyond the data range (extrapolation) would be unreliable.
该模型预测,一个每天使用手机5小时的学生,其数学期望成绩为51.5 – 4.01×5 = 31.45分。然而,超出数据范围的预测(外推)是不可靠的。
10. Drawing Conclusions and Evaluating the Study | 得出结论并评估研究
Based on the sample, there is evidence of a negative association between screen time and maths performance. However, correlation does not imply causation. Other factors — such as study habits, sleep, or prior ability — could be influencing both variables. The sample size of 30, though larger than the minimum often used in CIE problems, limits the generalisability of the conclusion to the entire Year 11 group.
根据样本,有证据表明屏幕时间与数学成绩之间存在负相关。然而,相关关系并不意味着因果关系。其他因素——如学习习惯、睡眠或先前能力——可能在同时影响两个变量。样本量只有30,尽管大于CIE问题中常用的最小样本量,但它限制了结论对整个11年级的普适性。
To improve the study, we could increase the sample size, collect screen time data objectively through apps, and include a wider range of control variables. In the CIE examination, you should be ready to comment on whether the method of data collection was suitable, the type of correlation, the interpretation of r and the regression line, and the limitations of the findings.
要改进这项研究,我们可以扩大样本量,通过应用程序客观收集屏幕时间数据,并纳入更广泛的控制变量。在CIE考试中,你需要准备好评论数据收集方法是否合适、相关性的类型、r和回归线的解释以及研究结果的局限性。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply