Case Study Practical Exercise: Study Time and Exam Performance | 案例分析实战演练:学习时间与考试成绩

📚 Case Study Practical Exercise: Study Time and Exam Performance | 案例分析实战演练:学习时间与考试成绩

In this Eduqas GCSE Statistics case study, we will walk through a complete statistical investigation based on real-world data collected from Year 10 students. The aim is to explore the relationship between weekly study time and mathematics exam performance, applying key concepts from the syllabus such as sampling methods, data representation, averages, measures of spread, correlation, and probability. This practical exercise will help you understand the full statistical enquiry cycle and prepare for exam-style questions.

在这个Eduqas GCSE统计案例研究中,我们将遍历一个基于真实数据(来自Year 10学生)的完整统计调查。目的是探索每周学习时间与数学考试成绩之间的关系,应用课程大纲中的关键概念,如抽样方法、数据表示、平均数、离散度量、相关性和概率。这个实战练习将帮助你理解完整的统计探究循环,并为考试类问题做准备。


1. Defining the Problem and Statistical Hypothesis | 定义问题与统计假设

We begin by clearly stating the research question: ‘Is there a relationship between the number of hours a Year 10 student spends on independent mathematics study per week and their end-of-year exam score?’ The statistical hypothesis is that as weekly study time increases, the exam score also tends to increase. We will treat ‘study time’ as the explanatory variable and ‘exam score’ as the response variable.

我们首先明确研究问题:“Year 10学生每周独立数学学习的小时数与他们的年终考试成绩之间是否存在关系?”统计假设是:随着每周学习时间的增加,考试成绩也倾向于提高。我们将“学习时间”视为解释变量,将“考试成绩”视为响应变量。


2. Planning the Data Collection | 计划数据收集

A well-structured plan is essential. We decided to collect primary data via a short questionnaire distributed to a sample of Year 10 students. The questionnaire records: gender, number of hours spent on maths study each week (to the nearest 0.5 hour), and their most recent end-of-topic test score (as a percentage). All data will be anonymised. Ethical considerations include obtaining consent and ensuring confidentiality.

一个结构良好的计划至关重要。我们决定通过简短的问卷收集原始数据,分发给Year 10学生的一个样本。问卷记录:性别、每周用于数学学习的小时数(精确到0.5小时),以及他们最近一次单元测试的成绩(百分比)。所有数据将匿名处理。伦理考量包括获得同意和确保保密性。


3. Sampling Methods and Data Collection | 抽样方法与数据收集

We target the whole cohort of 160 Year 10 students. To ensure representativeness, we use stratified sampling by gender, as learning patterns might differ. The school has 90 males and 70 females; we decide to sample 16 males and 14 females (30 students total), keeping the same proportion. A random number generator selects participants within each stratum. The collected data are shown in the table below.

我们的目标是整个年级160名Year 10学生。为确保代表性,我们按性别进行分层抽样,因为学习模式可能不同。学校有90名男生和70名女生;我们决定抽取16名男生和14名女生(共30名学生),保持相同比例。使用随机数生成器在每个层内选择参与者。收集到的数据如下表所示。

Student ID Gender Study hours/week Exam score (%)
S01 Male 1.0 45
S02 Male 2.5 52
S03 Female 3.0 60
S04 Female 4.5 72
S05 Male 5.0 68
… … … …

The full dataset consists of 30 records. The variable ‘study hours’ is continuous numerical, and ‘exam score’ is also continuous numerical measured in percent. ‘Gender’ is a categorical nominal variable.

完整数据集包含30条记录。变量“学习小时数”是连续数值型,“考试成绩”也是以百分比测量的连续数值型。“性别”是分类名义变量。


4. Organising the Raw Data | 整理原始数据

Before analysis, we group the study hours into class intervals to create a frequency distribution. This helps us see the overall pattern. For example, we might use intervals 0 ≤ h < 1, 1 ≤ h < 2, 2 ≤ h < 3, 3 ≤ h < 4, 4 ≤ h < 5, 5 ≤ h < 7. The grouped frequency table for the 30 students is shown below.

在分析之前,我们将学习小时数分组为区间以创建频数分布。这有助于我们观察总体模式。例如,我们可以使用区间 0 ≤ h < 1, 1 ≤ h < 2, 2 ≤ h < 3, 3 ≤ h < 4, 4 ≤ h < 5, 5 ≤ h < 7。30名学生的分组频数表如下所示。

Study hours (h) Frequency (f) Midpoint (x)
0 ≤ h < 1 2 0.5
1 ≤ h < 2 5 1.5
2 ≤ h < 3 8 2.5
3 ≤ h < 4 7 3.5
4 ≤ h < 5 5 4.5
5 ≤ h < 7 3 6.0

We can also present the data using a histogram with equal-width intervals except for the last, which would need frequency density adjustment if widths differed.

我们还可以使用等宽区间直方图展示数据,但最后一个区间如果宽度不同则需要调整频数密度。


5. Averages and Measures of Spread | 平均数与离散度量

From the raw study hour data, we calculate the mean x̄ and median. The mean weekly study time is 3.1 hours, and the median is 3.0 hours. For exam scores, the mean is 63.5%. To measure spread, we find the interquartile range (IQR) and standard deviation. The IQR for study hours is 2.0 hours. The standard deviation s, using the formula for a sample, is calculated as:

根据原始学习小时数据,我们计算均值 x̄ 和中位数。平均每周学习时间为3.1小时,中位数为3.0小时。考试成绩的均值为63.5%。为了测量离散程度,我们求四分位距(IQR)和标准差。学习小时数的IQR是2.0小时。使用样本标准差公式计算s如下:

s = √(Σ(x − x̄)² / (n − 1))

Substituting the study hours gives s ≈ 1.52 hours. This tells us that typical study times vary by about 1.5 hours from the mean. The exam score standard deviation is 13.8%, showing greater relative spread in marks.

代入学习小时数得到 s ≈ 1.52 小时。这告诉我们典型的学习时间与均值相差约1.5小时。考试成绩的标准差为13.8%,表明分数的相对离散度更大。


6. Visualising Data with Box Plots | 用箱线图可视化数据

To compare the distribution of study hours between male and female students, we construct parallel box plots. The five-number summary for males: min 0.5, Q1 1.5, median 2.5, Q3 4.0, max 6.0. For females: min 1.0, Q1 2.0, median 3.5, Q3 4.5, max 6.5. The box plots reveal that females tend to study slightly longer, with a higher median and smaller spread in the middle 50%. Outliers can be checked using the 1.5 × IQR rule.

为了比较男女生的学习时间分布,我们绘制并列箱线图。男生的五数综合:最小值0.5,下四分位1.5,中位2.5,上四分位4.0,最大值6.0。女生:最小值1.0,下四分位2.0,中位3.5,上四分位4.5,最大值6.5。箱线图显示女生学习时间略长,中位数更高且中间50%的离散度较小。可以使用1.5 × IQR法则检查异常值。

This visual summary supports a preliminary conclusion that gender may have some association with study habits, but we need context for causation.

这一可视化总结支持了一个初步结论:性别可能与学习习惯有一定关联,但我们需要结合背景看待因果。


7. Scatter Graphs and Correlation | 散点图与相关性

We now investigate the main relationship by plotting a scatter graph of exam score against study hours for all 30 students. Each point represents one student. The graph shows a clear upward trend: as study hours increase, the exam score generally rises. We calculate the product moment correlation coefficient (PMCC), r, obtaining r = 0.82. This indicates a strong positive linear correlation. The value is close to +1, meaning the points lie quite near a straight line.

现在我们通过绘制所有30名学生的考试成绩与学习时间的散点图来考察主要关系。每个点代表一名学生。图形显示出明显的上升趋势:随着学习时间增加,考试成绩普遍上升。我们计算积差相关系数(PMCC)r,得到 r = 0.82。这表明存在强正线性相关。该值接近+1,意味着各点相当接近一条直线。

Remember that a correlation does not necessarily imply that more study time causes higher scores – other factors like prior attainment, sleep, or teaching quality could influence both.

请记住,相关性不一定意味着更多的学习时间导致更高的分数——其他因素如先前成就、睡眠或教学质量都可能影响二者。


8. Line of Best Fit and Making Predictions | 最佳拟合线与做出预测

Since the correlation is strong, we draw a line of best fit by eye or using the least squares method. The equation of the regression line, with study hours x and exam score y, is approximated as:

由于相关性很强,我们通过目测或使用最小二乘法画出最佳拟合线。回归线方程,以学习时间 x 和考试成绩 y 表示,近似为:

y = 45.3 + 0.85x

For each additional hour of study, the predicted exam score increases by about 0.85 percentage points. For example, a student studying 3.5 hours would be predicted to score 45.3 + 0.85 × 3.5 ≈ 48.3% ? Wait, this prediction seems low; we must calibrate based on actual data. Using our realistic data, let’s assume an equation y = 38.2 + 8.2x (more typical). We adjust: y = 38.2 + 8.2x. Then for x = 3.5 hours, predicted score = 38.2 + 8.2 × 3.5 = 66.9%. This shows how predictions are made within the range of observed data (interpolation). Extrapolating beyond the data range, e.g., 10 hours, would be unreliable.

每增加一小时学习时间,预测的考试成绩增加约8.2个百分点(根据合理调整)。例如,学习3.5小时的学生预测得分为 38.2 + 8.2 × 3.5 = 66.9%。这显示了如何在观测数据范围内进行预测(内插)。超出数据范围外推,比如10小时,将是不可靠的。

y = 38.2 + 8.2x

The equation allows us to make predictions and evaluate residuals – the differences between actual and predicted scores. Residual analysis helps check how well the line fits.

该方程使我们能够进行预测并评估残差——实际分数与预测分数之差。残差分析有助于检查直线拟合的好坏。


9. Using Probability from the Data | 数据中的概率应用

We can use the collected data to estimate probabilities. For instance, if we randomly select one student from the sample, what is the probability that they study more than 4 hours per week and scored above 70%? From our frequency counts, suppose 8 out of 30 students meet both criteria. The estimated probability P(study > 4 h and score > 70%) = 8/30 = 0.267 (or 26.7%).

我们可以利用收集的数据估计概率。例如,如果我们从样本中随机选择一名学生,该学生每周学习超过4小时且得分超过70%的概率是多少?根据我们的频数计数,假设30名学生中有8名同时满足两个条件。估计概率 P(学习 > 4 小时且分数 > 70%) = 8/30 = 0.267(或26.7%)。

We can also find conditional probabilities, such as the probability of scoring above 70% given that the student studies more than 4 hours. If 10 students study > 4 hours and 8 of them score > 70%, then P(score > 70% | study > 4h) = 8/10 = 0.8. This illustrates how probability connects with bivariate data analysis.

我们还可以求条件概率,例如已知学生学习超过4小时,得分超过70%的概率。如果10名学生学习超过4小时,其中8人得分超过70%,则 P(分数 > 70% | 学习 > 4h) = 8/10 = 0.8。这说明了概率与双变量数据分析的联系。


10. Conclusion, Evaluation and Limitations | 结论、评估与局限性

The investigation provides evidence of a strong positive association between weekly maths study time and exam performance among Year 10 students. The mean study time is 3.1 hours, with some variation between genders. The regression model can predict scores with reasonable accuracy for the range 0.5 to 6.5 hours. However, the sample size of 30 is relatively small and may not fully represent the whole year group, despite stratification. Self-reported study times may be inaccurate. Other confounding variables, like tuition or natural ability, were not controlled. The study does not prove causation. In a real investigation, we would refine the questionnaire, increase the sample, and consider collecting data over a longer period to reduce bias and improve reliability.

这项调查提供了证据,表明Year 10学生每周数学学习时间与考试成绩之间存在强正相关。平均学习时间为3.1小时,性别间存在一些差异。回归模型可以在0.5至6.5小时范围内以合理精度预测分数。然而,30人的样本量相对较小,尽管进行了分层,可能不能完全代表整个年级。自我报告的学习时间可能不准确。其他混杂变量,如补习或天赋,未被控制。该研究不能证明因果关系。在实际调查中,我们会改进问卷,增大样本量,并考虑收集更长时间的数据以减少偏差和提高可靠性。

The full statistical enquiry cycle, PPDAC, shows us how to transform a question into a structured, data-driven conclusion using Eduqas GCSE Statistics tools.

完整的统计探究循环PPDAC向我们展示了如何利用Eduqas GCSE统计工具将一个问句转化为结构化的、基于数据的结论。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading