📚 Year 11 Edexcel Statistics: Case Study Practical Walkthrough | Edexcel 统计:案例分析实战演练
Mastering case study questions in Edexcel Statistics is essential for achieving the top grades. These extended problems test your ability to plan an investigation, collect and present data, choose appropriate statistical measures and draw valid conclusions in context. This article takes you through a realistic scenario while highlighting key techniques and common pitfalls.
在 Edexcel 统计学中攻克案例分析题是取得高分的关键。这类综合题全面考查你从制定调查计划、收集与呈现数据,到选用统计量并在情境中得出有效结论的能力。本文将带领你拆解一个贴近真实的案例,串讲核心方法,并指出常见的失分点。
1. Case Study Scenario and Objectives | 案例背景与目标
A mathematics teacher suspects that the amount of time students spend on their phones each day may be linked to their recent maths test scores. The objective is to investigate whether there is a linear relationship between daily phone use (in hours) and the test score (out of 100), and if so, whether phone use can be used to predict performance.
一位数学老师猜测学生每天花在手机上的时间可能与最近一次数学测验的成绩有关。研究目标就是考察每日手机使用时长(小时)与测验分数(满分 100 分)之间是否存在线性关系;如果存在,能否用手机使用时间来预测成绩。
A full statistical enquiry involves formulating a hypothesis, designing a data collection method, summarising the data numerically and graphically, calculating correlation and regression, and finally interpreting the results in the context of the original problem.
一次完整的统计探究包括提出假设、设计数据收集方式、用数字和图表汇总数据、计算相关性与回归方程,最后在原问题的情境下解读结果。
2. Designing the Investigation and Sampling Method | 调查设计与抽样方法
This is an observational study, not an experiment, because we simply record existing phone habits and test scores without imposing any treatment. The target population is all Year 11 students in the school, but it is impractical to survey everyone, so a sample of 20 students is selected.
这是一项观察性研究而非实验,因为我们只是记录已有的手机习惯和测验分数,并未施加任何处理。目标总体是全校所有十一年级学生,但调查全体并不现实,所以从中抽取 20 名学生作为样本。
To obtain a representative sample, a stratified sampling method can be used: the Year 11 cohort is first split into strata by tutor group or gender, and then simple random sampling is applied within each stratum. This helps reduce selection bias and ensures important subgroups are proportionally represented.
为了获得有代表性的样本,可以采用分层抽样:先将十一年级按导师组或性别分层,然后在各层内进行简单随机抽样。这有助于减少选择偏差,并确保重要子群体按比例覆盖。
Each selected student gives informed consent, and the data is anonymised to protect privacy. The variables recorded are ‘Daily Phone Use (hours)’ and ‘Maths Test Score’.
每位被选中的学生都知情同意,且数据匿名化处理以保护隐私。记录的两个变量是“每日手机使用(小时)”和“数学测验分数”。
3. Data Collection and Initial Organisation | 数据收集与初步整理
The table below shows the first 10 observations from the sample. The full data set contains 20 paired values, which have been entered into a spreadsheet for analysis.
下表展示了样本中前 10 组观测值。完整数据集包含 20 组成对数据,已录入电子表格以备分析。
| Student | Daily Phone Use (hours) | Maths Score (out of 100) |
|---|---|---|
| 1 | 1.0 | 95 |
| 2 | 1.5 | 92 |
| 3 | 2.0 | 88 |
| 4 | 2.0 | 90 |
| 5 | 2.5 | 85 |
| 6 | 3.0 | 82 |
| 7 | 3.0 | 84 |
| 8 | 3.5 | 78 |
| 9 | 4.0 | 80 |
| 10 | 4.0 | 76 |
Before any calculation, it is good practice to check for obvious data entry errors, such as a score above 100 or negative phone hours. The data should also be cleaned for any missing values.
在任何计算开始之前,最好检查是否存在明显的录入错误,例如分数超过 100 或手机使用时长为负数。同时还要清理缺失值。
4. Descriptive Statistics: Central Tendency and Spread | 描述性统计:集中趋势与离散程度
For the variable ‘Daily Phone Use’ (x), the sample statistics from the full 20 students are: mean x̄ = 4.2 hours, median = 3.75 hours, standard deviation sₓ = 2.1 hours. The median is lower than the mean, indicating a slight positive skew in phone use.
对于变量“每日手机使用”(x),由 20 名学生计算出的样本统计量为:均值 x̄ = 4.2 小时,中位数 = 3.75 小时,标准差 sₓ = 2.1 小时。中位数低于均值,说明手机使用数据略呈正偏态。
For ‘Maths Score’ (y), the mean ȳ = 74.5, median = 76, and standard deviation sᵧ = 14.8. The median and mean are close, so the score distribution is roughly symmetric.
对于“数学分数”(y),均值 ȳ = 74.5,中位数 = 76,标准差 sᵧ = 14.8。中位数与均值接近,分数分布大致对称。
x̄ = Σx/n sₓ = √( Σ(x − x̄)² / (n−1) )
The standard deviation tells us that typical phone use varies by about 2.1 hours from the mean, while scores vary by about 14.8 marks. This variation is what we hope to explain through correlation and regression.
标准差告诉我们,手机使用时长的典型波动约为均值上下 2.1 小时,而分数的典型波动约为 14.8 分。我们正希望用相关与回归来解释这种变异性。
5. Visualising the Data: Scatter Plot and Box Plots | 数据可视化:散点图与箱线图
A scatter plot is the most natural way to assess the relationship between two numerical variables. On the horizontal axis we place ‘Daily Phone Use’ and on the vertical axis ‘Maths Score’. Each point represents one student.
散点图是考察两个数值变量关系最自然的方式。横轴为“每日手机使用”,纵轴为“数学分数”,每个点代表一名学生。
The plot reveals a clear downward trend: as phone use increases, the maths score tends to decrease. The points cluster around an invisible line, suggesting a strong negative linear association.
散点图显示出明显的下降趋势:手机使用时间越长,数学分数往往越低。数据点密集地分布在一个假想直线的周围,暗示存在较强的负线性关联。
Side-by-side box plots can also be drawn by grouping phone use into categories (e.g. low, medium, high) and comparing the spread and medians of the scores. This reinforces the observation that higher phone use is associated with lower median achievement.
也可将手机使用时长分组(如低、中、高),绘制并列箱线图,比较各组成绩的分散程度和中位数。这进一步印证了手机使用越多,成绩中位数越低的趋势。
6. Measuring Correlation: Pearson’s Product-Moment Correlation Coefficient | 相关性度量:皮尔逊积矩相关系数
The strength and direction of a linear relationship are quantified by Pearson’s correlation coefficient r. Its formula is:
线性关系的强度和方向用皮尔逊相关系数 r 来量化。公式如下:
r = ( Σxy − (Σx)(Σy)/n ) / √( [Σx² − (Σx)²/n] × [Σy² − (Σy)²/n] )
Calculated for our 20 observations, we obtain r = −0.94. This value is close to −1, confirming a very strong negative linear correlation. A negative r means that as one variable increases, the other tends to decrease.
代入全部 20 组数据计算,得到 r = −0.94。该值接近 −1,证实存在极强的负线性相关。r 为负意味着一个变量增大时,另一个变量倾向于减小。
It is important to remember that correlation does not imply causation. A high r only indicates that the two variables move together linearly; it does not prove that higher phone use causes lower scores.
必须牢记,相关性并不意味着因果关系。高 r 仅表明两个变量呈线性共变,不能证明手机用得多导致分数下降。
7. Building a Linear Regression Model | 建立线性回归模型
Since the scatter plot is roughly linear and r is very strong, we can fit a least-squares regression line of the form y = a + bx, where x is phone use and y is predicted maths score.
既然散点图近似线性且 r 非常强,我们就可以拟合一条最小二乘回归直线,形式为 y = a + bx,其中 x 为手机使用时间,y 为预测的数学分数。
b = r × (sᵧ / sₓ) a = ȳ − b × x̄
Using our summary statistics: b = −0.94 × (14.8 / 2.1) ≈ −6.62, and a = 74.5 − (−6.62)×4.2 ≈ 102.3. The regression equation is therefore:
代入汇总统计量:b = −0.94 × (14.8 / 2.1) ≈ −6.62,a = 74.5 − (−6.62)×4.2 ≈ 102.3。因此回归方程为:
y = 102.3 − 6.6x
The slope b = −6.6 means that for every additional hour of daily phone use, the maths score is predicted to drop by about 6.6 marks, on average. The intercept a = 102.3 would be the predicted score for a student who does not use a phone at all – a value that should be interpreted cautiously because it lies outside the range of observed data.
斜率 b = −6.6 的含义是:日均手机使用每增加 1 小时,预计数学分数平均下降约 6.6 分。截距 a = 102.3 是对从不使用手机的学生的预测分数——但解读时需谨慎,因为该值已超出观测数据的范围。
8. Residual Analysis and Model Assessment | 残差分析与模型评估
A residual is the difference between an actual y value and the value predicted by the regression line: e = y − ŷ. Plotting residuals against x helps check whether a linear model is appropriate.
残差是实际 y 值与回归线预测值之间的差值:e = y − ŷ。将残差对 x 做图,有助于判断线性模型是否合适。
In our residual plot, the points are scattered randomly above and below the zero line with no obvious pattern, and the spread appears roughly constant. This ‘random scatter’ supports the use of a linear model.
在我们的残差图中,数据点随机分布在零线的上下,没有明显的规律,分散程度大致恒定。这种“随机散落”形态支持使用线性模型。
If the residual plot had shown a curved pattern or increasing spread, we would need to consider a non-linear model or a data transformation. For this case, the linear model passes the residual check satisfactorily.
假如残差图呈现出弯曲形态或喇叭状扩散,我们就需要考虑非线性模型或数据变换。在本案例中,线性模型顺利通过残差检验。
9. Interpretation and Conclusions: Statistical versus Practical Significance | 解释与结论:统计意义与实际意义
The analysis demonstrates a strong negative linear relationship between daily phone use and maths test scores for this sample. The regression equation suggests that limiting phone time could be associated with higher predicted scores.
分析表明,在该样本中,每日手机使用时间与数学测试成绩之间存在很强的负线性关系。回归方程提示,限制手机使用时间可能与较高的预测成绩相关联。
We can also quote the coefficient of determination, R² = r² = 0.88, meaning that 88% of the variation in maths scores can be explained by the linear relationship with phone use. The remaining 12% is due to other factors or random variation.
我们还可引用决定系数 R² = r² = 0.88,这意味着数学分数中 88% 的变异可以通过与手机使用时长的线性关系来解释,剩下的 12% 归因于其他因素或随机波动。
However, we must stress that the study only shows association. Other variables, such as hours of study, sleep quality or prior attainment, could confound the relationship. The practical significance is meaningful – a difference of 6.6 marks per hour is substantial – but more controlled research would be needed to establish causation.
然而,必须强调该研究仅显示关联关系。其他变量,如学习时间、睡眠质量或先前成绩,都可能成为混杂因素。实际意义上,每小时 6.6 分的差异相当可观,但要确定因果关系仍需更严格的研究。
10. Critical Reflection and Limitations | 批判性反思与局限性
The sample size of 20 students is relatively small, which makes the correlation estimate less reliable and the regression line sensitive to outliers. Extrapolating beyond the range of phone use (0–9 hours) would be unwise because the relationship may not remain linear outside this span.
样本容量 20 人相对较小,这使得相关系数估计不够稳定,回归线对异常值也比较敏感。在手机使用范围(0–9 小时)之外进行外推是不明智的,因为该范围之外的关系可能不再保持线性。
All data were collected from one school, so the findings cannot be generalised to all Year 11 students. Self-reported phone use might also be subject to recall bias or under-reporting. Future studies could use app-tracked screen time for greater accuracy.
所有数据
Published by TutorHao | Year 11 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导