📚 Year 10 Eduqas Statistics: Case Study Drill | Year 10 Eduqas 统计:案例分析实战演练
Welcome to this comprehensive case study drill designed for Year 10 students following the Eduqas GCSE Statistics specification. In this article, we will work through a realistic data investigation, from planning and data collection to analysis and evaluation, using a dataset on teenagers’ weekly exercise hours and their self-reported well-being scores. This step-by-step walkthrough will help you master the statistical enquiry cycle and apply your skills to GCSE-style questions.
欢迎参加这个为 Eduqas GCSE 统计学课程设计的综合案例演练。在本文中,我们将通过一个真实的数据调查,从规划和数据收集到分析和评估,使用关于青少年每周锻炼小时数和自报幸福感评分的数据集,逐步完成统计探究。这个循序渐进的讲解将帮助你掌握统计调查周期,并将你的技能应用到 GCSE 风格的考题中。
1. Defining the Problem and Setting Hypotheses | 定义问题与设定假设
Before any data is collected, a statistician must clearly state the aim of the investigation. For this case study, our research question is: ‘Is there a relationship between the number of hours teenagers spend on physical exercise per week and their self-assessed well-being score on a scale from 1 to 10?’ We suspect that as exercise hours increase, well-being scores also tend to increase. Therefore, our hypothesis is that there is a positive correlation between the two variables.
在收集任何数据之前,统计人员必须清楚地陈述调查目的。在这个案例研究中,我们的研究问题是:“青少年每周用于体育锻炼的小时数与他们在1至10分制上的自评幸福感评分之间是否存在关系?”我们推测,随着锻炼时间的增加,幸福感评分也倾向于提高。因此,我们的假设是这两个变量之间存在正相关关系。
It is important to remember that a hypothesis should be testable. In GCSE Statistics, you will often be asked to suggest a suitable null hypothesis and an alternative hypothesis. The null hypothesis (H₀) would state that there is no correlation between exercise hours and well-being score, while the alternative hypothesis (H₁) would suggest a positive (or negative) correlation. In this case, we are investigating a one-tailed positive correlation.
重要的是要记住,假设应该是可检验的。在 GCSE 统计学中,你经常会被要求提出一个合适的零假设和备择假设。零假设(H₀)会说明锻炼小时数与幸福感评分之间没有相关关系,而备择假设(H₁)则暗示存在正(或负)相关。在本案例中,我们正在研究单尾正相关。
2. Planning the Survey: Population, Sample and Data Types | 调查规划:总体、样本与数据类型
The target population for this study is all 14- to 16-year-old students in a particular secondary school. Because it is impractical to survey every student, we decide to take a representative sample. We could use stratified sampling by year group and gender to ensure the sample mirrors the school’s composition, or a simple random sample drawn from the school register. For this drill, we assume a simple random sample of 30 students has been selected.
这项研究的目标总体是某所中学所有14至16岁的学生。由于调查每名学生不切实际,我们决定抽取有代表性的样本。我们可以采用按年级和性别分层抽样的方法,确保样本反映学校的构成,或者从学校名册中抽取简单随机样本。在本次演练中,我们假设已选取了30名学生的简单随机样本。
We need to collect two pieces of data from each student: the number of hours of physical exercise per week (to the nearest half hour) and their overall well-being score on a scale of 1 to 10, where 1 means ‘very unhappy’ and 10 means ‘extremely happy’. The exercise variable is continuous (measured on a ratio scale), while the well-being score is discrete but treated as ordinal or scale data for analysis. A well-designed questionnaire would ask for these values in a consistent manner, perhaps with a visual analogue scale for well-being.
我们需要从每名学生那里收集两项数据:每周体育锻炼的小时数(精确到半小时)和他们在1至10分制上的整体幸福感评分,其中1代表“非常不快乐”,10代表“极其快乐”。锻炼变量是连续的(在比率尺度上测量),而幸福感评分是离散的,但分析时可作为定序或尺度数据处理。一份设计良好的问卷应以一致的方式询问这些数值,或许用视觉模拟标尺进行幸福感评分。
Secondary data could also have been used, but for this case study we simulate primary data collection to practise the full cycle. Always remember to consider ethical issues: students’ anonymity is protected, and participation is voluntary.
也可以使用二手数据,但在本案例研究中我们模拟一手数据收集,以便练习整个统计周期。要始终记得考虑伦理问题:保护学生的匿名性,且参与是自愿的。
3. Collecting and Recording the Data | 收集与记录数据
After administering the questionnaire, we obtain numerical data from 30 students. The dataset is recorded in a table with three columns: Student ID, Exercise (hours/week), and Well-being Score (1-10). For convenience, we only display the first 10 rows below, but the entire dataset will be used for calculations.
在发放问卷后,我们从30名学生那里获得了数值数据。数据集被记录在一个三列表中:学生编号、锻炼(小时/周)和幸福感评分(1-10)。为方便起见,我们在下面仅显示前10行,但整个数据集将用于计算。
| Student ID | Exercise (hours) | Well-being Score |
|---|---|---|
| 1 | 0.5 | 2.0 |
| 2 | 1.0 | 2.5 |
| 3 | 1.5 | 3.0 |
| 4 | 2.0 | 3.8 |
| 5 | 2.5 | 4.2 |
| 6 | 3.0 | 4.5 |
| 7 | 3.5 | 5.1 |
| 8 | 4.0 | 5.5 |
| 9 | 4.5 | 6.0 |
| 10 | 5.0 | 6.4 |
The full dataset of 30 observations is used for all subsequent calculations. The raw data should be checked for obvious errors or outliers – none were found in this exercise.
全部30个观测值的完整数据集将用于后续所有的计算。应检查原始数据是否存在明显错误或异常值——本次练习中未发现任何异常。
4. Univariate Analysis: Exercise Hours | 单变量分析:锻炼小时数
We begin by summarising the exercise hours (x) on their own. The mean number of hours per week is calculated as:
我们首先单独对锻炼小时数(x)进行汇总。每周平均小时数计算如下:
x̄ = Σxᵢ / n = 147.5 / 30 ≈ 4.92 hours
x̄ = Σxᵢ / n = 147.5 / 30 ≈ 4.92 小时
The median (Q₂) is the middle value when the data are sorted. After ordering the 30 observations, the median lies between the 15th and 16th values, giving 5.0 hours. This suggests a roughly symmetric distribution, as mean ≈ median.
中位数(Q₂)是数据排序后的中间值。将30个观测值排序后,中位数位于第15和第16个值之间,得出5.0小时。这表明分布大致对称,因为均值≈中位数。
The five-number summary helps build a box plot:
五数概括有助于构建箱线图:
Minimum = 0.0 hrs, Q₁ = 2.5 hrs, Median = 5.0 hrs, Q₃ = 7.5 hrs, Maximum = 10.0 hrs. Interquartile range (IQR) = Q₃ − Q₁ = 5.0 hrs.
最小值 = 0.0小时,第一四分位数 Q₁ = 2.5小时,中位数 = 5.0小时,第三四分位数 Q₃ = 7.5小时,最大值 = 10.0小时。四分位距(IQR)= Q₃ − Q₁ = 5.0小时。
The standard deviation (s) measures the spread about the mean. For the sample data, s is found using:
标准差(s)衡量数据围绕均值的离散程度。对于样本数据,s 通过以下公式求得:
s = √[ Σ(xᵢ − x̄)² / (n − 1) ]
s = √[ Σ(xᵢ − x̄)² / (n − 1) ]
Using the full dataset, Σ(xᵢ − x̄)² = 219.64, so s = √(219.64 / 29) ≈ √7.57 ≈ 2.75 hours. Since the standard deviation is moderately large relative to the mean, there is considerable variation in exercise habits among students.
使用完整数据集,Σ(xᵢ − x̄)² = 219.64,因此 s = √(219.64 / 29) ≈ √7.57 ≈ 2.75小时。由于标准差相对于均值较大,说明学生们的锻炼习惯存在相当大的差异。
A box plot would show the box from 2.5 to 7.5 hours with the median line at 5.0, and whiskers extending to 0 and 10. There are no extreme outliers according to the 1.5 × IQR rule.
箱线图将显示箱子从2.5到7.5小时,中位线在5.0处,须线延伸至0和10。根据 1.5 × IQR 规则,没有极端异常值。
5. Univariate Analysis: Well-being Score | 单变量分析:幸福感评分
Now we turn to the well-being score (y). The summary statistics are:
现在我们来看幸福感评分(y)。汇总统计量如下:
ȳ = Σyᵢ / n = 193.8 / 30 ≈ 6.46
ȳ = Σyᵢ / n = 193.8 / 30 ≈ 6.46
Minimum = 1.5, Q₁ = 4.2, Median = 6.8, Q₃ = 8.5, Maximum = 10.0. The IQR = 4.3. The mean is slightly lower than the median, suggesting a slight negative
Published by TutorHao | Year 10 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导