📚 Case Study in Statistics: A Practical Walkthrough | 统计学案例分析实战演练
Statistics is not just a collection of formulas and numbers; it is a powerful tool for making sense of the world around us. In CCEA GCSE Statistics, the ability to carry out a case study from start to finish demonstrates real statistical competence. This article will walk you through a complete practical example, showing how to plan, collect, analyse, interpret, and evaluate data. We will use a fictional investigation into the relationship between daily screen time and sleep quality among Year 11 students, applying the key statistical skills required for your coursework and examinations.
统计学不仅仅是一堆公式和数字,它更是理解我们周围世界的有力工具。在 CCEA GCSE 统计课程中,从头到尾完成一个案例分析的能力展示了真正的统计素养。本文将通过一个完整的实战演练,向你展示如何规划、收集、分析、解释和评价数据。我们将以一个虚构的调查为例,探究 11 年级学生每天屏幕使用时间与睡眠质量之间的关系,并运用课程和考试中所需的关键统计技能。
1. Introduction to Case Studies in Statistics | 统计学案例分析简介
A statistical case study is a structured investigation that uses real or realistic data to answer a research question. It involves moving through the statistical enquiry cycle: posing a problem, planning data collection, gathering and processing data, performing calculations, and drawing conclusions. This hands‑on approach deepens understanding and prepares you for the problem‑solving style of CCEA exam questions.
统计案例分析是一种结构化的调查,它使用真实或仿真数据来回答研究问题。这需要经过统计调研周期:提出问题、规划数据收集、采集和处理数据、进行计算以及得出结论。这种动手实践的方法可以加深理解,并为 CCEA 考试中解决问题类型的题目做好准备。
2. Defining the Problem and Hypotheses | 定义问题与假设
Every case study begins with a clear research question. For our investigation we ask: ‘Is there a relationship between the number of hours of screen time per day and the quality of sleep reported by Year 11 students?’ We can then state a null hypothesis H₀: there is no correlation between screen time and sleep quality in the population. The alternative hypothesis H₁ states that there is a negative correlation — as screen time increases, sleep quality tends to decrease.
每个案例分析都始于一个明确的研究问题。在我们的调查中,我们问:“11 年级学生每天屏幕使用时间与睡眠质量之间是否存在关系?” 随后,我们可以提出原假设 H₀:总体中屏幕时间与睡眠质量之间没有相关关系。备择假设 H₁ 则表明存在负相关——随着屏幕时间增加,睡眠质量倾向于下降。
Clearly defined hypotheses give direction to the whole study and determine which statistical tests are appropriate. In this case, we plan to collect bivariate numerical data and examine correlation and regression.
明确界定的假设为整个研究指明了方向,并决定了适用哪些统计检验。在本例中,我们计划收集双变量数值数据,并考察相关与回归。
3. Data Collection Methods | 数据收集方法
To obtain reliable data, we design a simple questionnaire that asks participants to record their average daily screen time (in hours, to the nearest 0.5) and to rate their sleep quality on a scale from 1 (very poor) to 10 (excellent). We use a random sampling method — such as giving every Year 11 student in a school a number and generating random numbers — to select 20 students. This helps avoid bias and makes the sample more representative.
为了获得可靠的数据,我们设计了一份简单的问卷,要求参与者记录他们每日平均屏幕使用时间(以小时为单位,精确到 0.5),并对睡眠质量进行评分,从 1(非常差)到 10(非常好)。我们使用随机抽样方法——例如给学校里每位 11 年级学生编号并生成随机数——选出 20 名学生。这有助于避免偏差,使样本更具代表性。
Ethical considerations are also important: we make sure the responses are anonymous and that participants know they can withdraw at any time. The data are collected during a typical school week to reflect a normal routine.
伦理考量也很重要:我们确保回答为匿名,并且参与者知晓他们可以随时退出。数据在正常的教学周内收集,以反映日常生活规律。
4. Organising and Cleaning Data | 数据整理与清洗
Once collected, the raw data must be organised into a clear table. Below is the dataset for our 20 participants. Cleaning involves checking for missing or unusual values; here all responses are complete and within plausible ranges.
数据收集完毕后,必须整理成清晰的表格。下面是 20 名参与者的数据集。清洗包括检查缺失或异常值;此处的所有回答都完整且在合理范围内。
| ID | Screen Time (h) | Sleep Quality (1–10) |
|---|---|---|
| 1 | 2.5 | 8 |
| 2 | 3.0 | 7 |
| 3 | 4.5 | 6 |
| 4 | 1.5 | 9 |
| 5 | 5.0 | 5 |
| 6 | 2.0 | 9 |
| 7 | 3.5 | 7 |
| 8 | 4.0 | 6 |
| 9 | 5.5 | 4 |
| 10 | 1.0 | 10 |
| 11 | 2.5 | 8 |
| 12 | 3.0 | 8 |
| 13 | 4.0 | 7 |
| 14 | 4.5 | 5 |
| 15 | 6.0 | 3 |
| 16 | 2.0 | 9 |
| 17 | 3.5 | 6 |
| 18 | 5.0 | 5 |
| 19 | 1.5 | 9 |
| 20 | 3.0 | 7 |
Having data in a structured format makes it easy to compute statistics and identify patterns. Always double-check that entries make sense — a screen time of 85 hours would be an outlier or a typing error.
将数据整理成结构化的格式,便于计算统计量和识别模式。务必再次检查输入是否合理——比如 85 小时的屏幕时间可能是异常值或录入错误。
5. Visualising Data with Charts | 用图表可视化数据
A scatter graph is the best way to display the relationship between two numerical variables. Plot screen time on the horizontal axis and sleep quality on the vertical axis. The points should show a clear downward trend: students with lower screen times tend to have higher sleep quality scores, while those spending more hours in front of screens score lower.
散点图是展示两个数值变量之间关系的最佳方式。将屏幕时间置于横轴,睡眠质量置于纵轴。各点应呈现出明显的下降趋势:屏幕时间较少的学生睡眠质量评分较高,而花更多时间看屏幕的学生评分则较低。
In your case study, you would draw the scatter plot neatly and label all axes. You can comment on the strength (the points are fairly close to an imaginary straight line), direction (negative), and any outliers. This visual inspection supports the idea of a negative correlation before we even calculate numbers.
在你的案例分析中,你需要整洁地绘制散点图,并标注所有坐标轴。你可以评价其强度(各点相当接近一条假想直线)、方向(负)以及任何异常值。这种视觉检查甚至在计算数值之前就支持了存在负相关的设想。
6. Calculating Averages and Spread | 计算平均数与离散程度
Before exploring the relationship, it is useful to summarise each variable individually. For screen time (x), we compute the mean, median, range and interquartile range. Using the dataset, the mean screen time is x̄ = (2.5 + 3.0 + 4.5 + … + 3.0) ÷ 20 = 67.5 ÷ 20 = 3.375 hours. The median lies between the 10th and 11th ordered values: after sorting we find the median is 3.25 hours (the midpoint of 3.0 and 3.5). The range is 6.0 − 1.0 = 5.0 hours.
在探究关系之前,分别总结每个变量是很有用的。对于屏幕时间 (x),我们计算平均数、中位数、极差和四分位距。使用数据集,平均屏幕时间为 x̄ = (2.5 + 3.0 + 4.5 + … + 3.0) ÷ 20 = 67.5 ÷ 20 = 3.375 小时。中位数位于排序后第 10 和第 11 个值之间:排序后中位数为 3.25 小时(3.0 与 3.5 的中点)。极差为 6.0 − 1.0 = 5.0 小时。
For sleep quality (y), the mean is ȳ = 138 ÷ 20 = 6.9. The median is 7 (between the 10th and 11th sorted values of 7 and 7). The range is 10 − 3 = 7. In a full analysis you would also calculate quartiles and the interquartile range to describe the spread more robustly.
对于睡眠质量 (y),平均数为 ȳ = 138 ÷ 20 = 6.9。中位数为 7(处于排序后第 10 和第 11 个值 7 与 7 之间)。极差为 10 − 3 = 7。在完整的分析中,你还可以计算四分位数和四分位距,以更稳健地描述离散程度。
7. Exploring Relationships: Correlation | 探索关系:相关分析
The Pearson product‑moment correlation coefficient (r) measures the strength and direction of a linear relationship. The formula is:
皮尔逊积矩相关系数 (r) 衡量线性关系的强度和方向。公式为:
r = Σ(xᵢ − x̄)(yᵢ − ȳ) / √[Σ(xᵢ − x̄)² Σ(yᵢ − ȳ)²]
Using a calculator or statistical software, we find r ≈ −0.92 for our data. This value is close to −1, indicating a strong negative correlation. The coefficient of determination, r² ≈ 0.85, tells us that about 85% of the variation in sleep quality can be explained by the variation in screen time within this sample.
使用计算器或统计软件,我们得出这组数据的 r ≈ −0.92。该值接近 −1,表明存在强负相关。决定系数 r² ≈ 0.85 告诉我们,在这个样本中,睡眠质量约 85% 的变异可以用屏幕时间的变异来解释。
Always interpret r in context, and remember that correlation does not imply causation. A strong correlation supports our alternative hypothesis but does not prove that extra screen time causes poor sleep.
务必结合背景解释 r,并记住相关关系并不意味着因果关系。强相关支持了备择假设,但不能证明额外的屏幕时间导致睡眠质量下降。
8. Simple Linear Regression | 简单线性回归
When a linear relationship appears strong, we can model it with a regression line of the form y = a + bx. The slope b is given by b = r × (s_y / s_x), and the intercept a = ȳ − b x̄. For our data, the calculations yield b ≈ −1.48 and a ≈ 11.9, giving the equation:
当线性关系看起来很强时,我们可以用形如 y = a + bx 的回归直线来建模。斜率 b 由 b = r × (s_y / s_x) 给出,截距 a = ȳ − b x̄。对于我们的数据,计算得到 b ≈ −1.48,a ≈ 11.9,方程为:
Sleep Quality = 11.9 − 1.48 × Screen Time
This line can be drawn on the scatter graph and used to make predictions within the data range. For instance, a student with 4 hours of screen time is predicted to have a sleep quality of about 11.9 − 1.48(4) = 6.0. The negative slope tells us that each extra hour of screen time is associated with a drop of approximately 1.5 points in sleep quality score.
这条直线可以画在散点图上,并用于在数据范围内进行预测。例如,屏幕时间为 4 小时的学生预计睡眠质量约为 11.9 − 1.48(4) = 6.0。负斜率告诉我们,每增加一小时屏幕时间,睡眠质量评分大约下降 1.5 分。
9. Probability and Risk | 概率与风险
Probability can add another layer to a case study. Using our modelled relationship, we can estimate the risk that a student has ‘low’ sleep quality (say, below 5). By solving the regression equation, a predicted score of 5 occurs when 11.9 − 1.48x = 5, giving x ≈ 4.66 hours. This suggests that students exceeding about 4.7 hours of daily screen time are more likely to report very poor sleep.
概率可以为案例研究增添另一个维度。利用我们建模的关系,我们可以估计学生睡眠质量“低”(比如低于 5 分)的风险。通过解回归方程,当 11.9 − 1.48x = 5 时,得到 x ≈ 4.66 小时。这表明每日屏幕时间超过约 4.7 小时的学生更有可能报告很差的睡眠质量。
In your analysis you could count how many students in the sample fell into this high‑risk category and compare it with the actual number scoring below 5. This moves from pure description towards practical decision‑making.
在分析中,你可以统计样本中有多少学生落入了这个高风险类别,并与实际得分低于 5 的人数进行比较。这就从纯粹的描述转向了实际的决策。
10. Interpreting Findings and Drawing Conclusions | 解释发现并得出结论
Our statistical evidence indicates a strong negative association between screen time and sleep quality in this sample of Year 11 students. The correlation is significant and the regression model fits reasonably well. We therefore reject the null hypothesis and conclude that there is a negative correlation in the population, consistent with the idea that more screen time is linked to poorer sleep.
我们的统计证据表明,在这个 11 年级学生样本中,屏幕时间与睡眠质量存在强烈的负相关。相关性显著,回归模型拟合得较好。因此,我们拒绝原假设,并得出结论:总体中存在负相关,这与屏幕时间越长、睡眠越差的观点一致。
However, we must be cautious: the study is observational and cannot prove that reducing screen time will improve sleep. Confounding variables — such as stress, caffeine intake, or bedtime routines — may also play a role.
然而,我们必须谨慎:这项研究是观察性的,无法证明减少屏幕时间会改善睡眠。混杂变量——如压力、咖啡因摄入或就寝习惯——也可能起作用。
11. Evaluating Limitations and Reliability | 评估局限性与可靠性
Every case study has limitations. Our sample size of 20 is small, which affects the reliability of the correlation coefficient. A larger random sample would give more trustworthy results. Self‑reported screen time may be inaccurate, and the sleep quality scale is subjective. Moreover, the data come from a single school, so findings may not apply to all Year 11 students.
每个案例分析都有局限性。我们的样本量只有 20,这会影响相关系数的可靠性。更大的随机样本会给出更可信的结果。自我报告的屏幕时间可能不准确,睡眠质量评分量表也带有主观性。此外,数据仅来自一所学校,因此结果可能不适用于所有 11 年级学生。
Despite these limitations, the study is a valid exercise in applying statistical techniques. Acknowledging weaknesses shows a mature understanding of the statistical process and can earn higher marks in your CCEA coursework.
尽管存在这些局限,这项研究仍是应用统计技术的有效练习。承认不足表明你对统计过程有成熟的理解,可以在 CCEA 课程作业中获得更高分数。
12. Presenting Your Case Study | 展示你的案例分析
A well‑structured report should contain: a title and introduction, a description of sampling and data collection, raw data tables, graphs (scatter plot with line of best fit), calculations and statistical measures, interpretation of findings, and a final evaluation. Use clear headings and refer to every chart and figure in your text.
一份结构良好的报告应包含:标题与引言、抽样与数据收集描述、原始数据表格、图表(散点图与最佳拟合线)、计算与统计量、结果解释以及最终评估。使用清晰的标题,并在正文中引用每个图表和图形。
In CCEA exams, you may be asked to critique a given case study or suggest improvements. Practise evaluating your own work using the criteria of validity, reliability, and representativeness. This reflective approach prepares you for both the written paper and the controlled assessment.
在 CCEA 考试中,你可能被要求评论一个给定的案例研究或提出改进建议。使用有效性、可靠性和代表性等标准来练习评价自己的工作。这种
Published by TutorHao | Year 11 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导