GCSE WJEC Statistics: Case Study Practical Exercises | GCSE WJEC 统计:案例分析实战演练

📚 GCSE WJEC Statistics: Case Study Practical Exercises | GCSE WJEC 统计:案例分析实战演练

In the WJEC GCSE Statistics specification, the ability to apply statistical techniques to real-world scenarios is essential. Case studies allow you to bring together data collection, representation, analysis, probability and hypothesis testing into one coherent investigation. This article walks you through a full practical exercise using a fictional case: investigating the relationship between students’ daily social media use and their academic performance. Every step mirrors the kind of task you might meet in your examination or controlled assessment, strengthening both your understanding and your confidence.

在 WJEC GCSE 统计学大纲中,将统计技术应用于真实情境的能力至关重要。案例分析能够让你把数据收集、展示、分析、概率和假设检验融合到一个连贯的调查中。本文将通过一个虚构案例——探究学生每日社交媒体使用时长与学业成绩之间的关系——带你完成一次完整的实战演练。每一步都与你可能在考试或课程作业中遇到的任务相似,从而加深理解并增强信心。

1. Introducing the Case Study: Social Media and Academic Performance | 案例引入:社交媒体与学业表现

Our scenario centres on a school where teachers suspect that heavy social media use might be linked to lower test scores. A student researcher decides to investigate by surveying 30 Year 11 pupils. The aim is to collect data on daily social media hours and recent mock exam percentages, then use WJEC statistical tools to explore any association and test a claim about the proportion of heavy users. This practical exercise will cover questionnaire design, sampling, data handling, graphical methods, summary statistics, correlation, regression, probability models and hypothesis testing.

我们的情境围绕一所学校展开,老师们怀疑频繁使用社交媒体可能与较低的考试成绩有关。一名学生研究员决定调查此事,对 30 名十一年级学生进行问卷调查。目标是收集每日社交媒体使用时长和近期模拟考试百分比的数据,然后运用 WJEC 统计工具探索关联性,并检验关于重度使用者比例的某个说法。本次实战演练将涵盖问卷设计、抽样、数据处理、图形方法、汇总统计、相关性、回归、概率模型和假设检验。


2. Designing a Questionnaire and Sampling Strategy | 问卷设计与抽样策略

Good data starts with a well-designed questionnaire. Our student creates questions that are clear, unbiased and produce quantitative answers. The first question asks: ‘On average, how many hours per day do you spend on social media?’ with a blank for a numerical answer. The second asks for the student’s most recent mock exam result as a percentage. To ensure a representative sample, she decides to use stratified sampling by gender, selecting 15 males and 15 females from the Year 11 register using random numbers. This minimises selection bias and makes the findings more generalisable to the whole year group.

好的数据源自一份精心设计的问卷。我们的学生设计的问题清晰、无偏且能产生定量答案。第一个问题是:‘你平均每天在社交媒体上花费多少小时?’并留出空白填写数字。第二个问题询问学生最近一次模拟考试的百分比成绩。为了保证样本具有代表性,她决定按性别进行分层抽样,使用随机数从十一年级名单中抽取 15 名男生和 15 名女生。这能最大限度地减少选择偏差,让结果更能代表整个年级。


3. Data Collection and Ethical Considerations | 数据收集与伦理考量

Before collecting data, the researcher obtains permission from the head of year and ensures all responses are anonymous. Pupils are informed that participation is voluntary and their answers will only be used for this statistics project. Each consenting pupil fills in the questionnaire during tutor time. The completed sheets are stored securely. Ethical practice is a key part of the WJEC specification, and you should always mention steps taken to protect participants’ identities and data confidentiality.

在收集数据之前,调查者获得了年级组长的许可,并确保所有回答都是匿名的。学生被告知参与是自愿的,他们的回答仅用于此统计项目。每位同意的学生在辅导时间填写问卷。完成的问卷被安全存放。伦理实践是 WJEC 大纲的重要组成部分,你应始终提及为保护参与者身份和数据机密性而采取的措施。


4. Organising Raw Data: Tally Charts and Frequency Tables | 整理原始数据:计数表和频数表

Once the questionnaires are collected, the raw data must be organised. The social media hours are recorded to one decimal place. Table 1 shows the first five rows of the data set. To analyse the screen time variable, the researcher creates a grouped frequency table with classes 0–<1 h, 1–<2 h, up to 5–<6 h. Tally marks are used to count how many pupils fall into each class. The frequency table gives a clear overview of the distribution and will be used later to draw a histogram and calculate summary statistics for grouped data.

一旦回收问卷,就需要整理原始数据。社交媒体使用时长被记录到一位小数。表 1 显示了数据集的前五行。为了分析屏幕时间变量,研究者编制了一个分组频数表,组距为 0–<1 小时、1–<2 小时,直至 5–<6 小时。使用画记符号计算每个组别中的学生人数。频数表能清晰地展示分布概况,后续将用于绘制直方图和计算分组数据的汇总统计量。

Pupil Social media (h/day) Mock exam (%)
1 2.3 76
2 4.1 58
3 0.5 92
4 3.0 68
5 1.8 81

Table 1: Extract of raw data for the first five students.

表 1:前五名学生的原始数据摘录。


5. Visualising Data: Bar Charts, Pie Charts and Histograms | 数据可视化:条形图、饼图与直方图

Appropriate graphs depend on the data type. For the social media hours, which are continuous, a histogram is the correct choice. The student plots frequency density on the vertical axis, where frequency density = frequency ÷ class width. Since all class widths are equal to 1 hour, the bar heights are simply the frequencies. This reveals a roughly symmetric but slightly right‑skewed distribution with a peak at 2–<3 h. She also creates a bar chart to display preferred social media platform (a categorical variable collected by a third question) and a pie chart to show the proportion of students exceeding 3 h per day. Each graph is fully labelled and includes a title.

合适的图表取决于数据类型。对于连续型的社交媒体使用时长,正确的选择是绘制直方图。学生在纵轴上标出频率密度,频率密度 = 频数 ÷ 组距宽度。由于所有组距宽度均为 1 小时,因此条形高度就是频数。直方图显示出大致对称但略微右偏的分布,峰值出现在 2–<3 小时组。她还制作了条形图来展示偏好的社交媒体平台(由第三个问题收集的分类变量),以及一幅饼图来显示每天使用超过 3 小时的学生比例。每张图都配有完整的标签和标题。


6. Measures of Central Tendency and Spread | 集中趋势与离散程度的度量

For the ungrouped social media hours data (n = 30), the student calculates the mean as Σx ÷ n. She obtains x̄ = 2.47 h. The median, found by locating the 15.5th ordered value, is 2.4 h. The mode is 2.3 h, occurring four times. To describe spread, she computes the range (maximum minus minimum) = 5.2 – 0.3 = 4.9 h and the interquartile range (IQR) = Q₃ – Q₁ = 3.4 – 1.6 = 1.8 h. She also uses the formula for sample standard deviation: s = √[ Σ(x – x̄)² ÷ (n – 1) ]. A calculator gives s ≈ 1.34 h. These measures indicate that typical daily social media use lies around 2.5 h with moderate variability.

对于未分组的社会媒体时长数据(n = 30),学生计算均值为 Σx ÷ n。她得到 x̄ = 2.47 小时。中位数通过找出第 15.5 个排序值得到,为 2.4 小时。众数是 2.3 小时,出现了四次。为了描述离散程度,她计算了极差(最大值减最小值)= 5.2 – 0.3 = 4.9 小时,以及四分位距 IQR = Q₃ – Q₁ = 3.4 – 1.6 = 1.8 小时。她还使用了样本标准差公式:s = √[ Σ(x – x̄)² ÷ (n – 1) ]。计算器给出 s ≈ 1.34 小时。这些度量表明,学生每日社交媒体使用时长通常集中在 2.5 小时左右,且具有中等变异程度。


7. Cumulative Frequency Curves and Box Plots | 累积频率曲线与箱形图

To further analyse the distribution, the student constructs a cumulative frequency table from the grouped data. She plots the upper class boundary against cumulative frequency and draws a smooth curve. From the curve, she estimates the median at around 2.5 h, the lower quartile at 1.7 h and the upper quartile at 3.3 h, giving an IQR of 1.6 h. These values are slightly different from the ungrouped values because of grouping. She then draws a box plot using these five‑number summaries (minimum, Q₁, median, Q₃, maximum) to compare the social media habits of males and females. The box plots show that females tend to spend slightly more time on social media, with a higher median and a narrower IQR.

为了进一步分析分布,学生根据分组数据构建了累积频率表。她以上组界为横轴、累积频率为纵轴描点,并画出平滑曲线。根据曲线,她估计中位数约为 2.5 小时,下四分位数约为 1.7 小时,上四分位数约为 3.3 小时,四分位距为 1.6 小时。由于数据分组,这些值与未分组数据的结果略有不同。随后,她利用五数概括(最小值、Q₁、中位数、Q₃、最大值)绘制箱形图,以比较男生和女生的社交媒体使用习惯。箱形图显示,女生花在社交媒体上的时间略多,中位数更高且四分位距更窄。


8. Scatter Diagrams and Correlation Analysis | 散点图与相关性分析

Now the researcher turns to the relationship between social media hours (x) and mock exam percentage (y). She plots a scatter diagram for all 30 pupils. The points show a negative trend: higher social media use seems to correspond to lower exam scores. To quantify this, she calculates the product moment correlation coefficient (PMCC). Using the formula r = Σ[(x – x̄)(y – ȳ)] / √[ Σ(x – x̄)² Σ(y – ȳ)² ] or a calculator’s statistical function, she finds r = –0.62. A value of –0.62 suggests a moderate negative linear correlation. The student notes that correlation does not imply causation and other factors like study time may be involved.

现在,研究者转而探讨社交媒体使用时长(x)与模拟考试百分比(y)之间的关系。她为全部 30 名学生绘制散点图。各点呈现出负向趋势:较高的社交媒体使用时长似乎与较低的考试分数相关。为了量化这一关系,她计算了积差相关系数(PMCC)。利用公式 r = Σ[(x – x̄)(y – ȳ)] / √[ Σ(x – x̄)² Σ(y – ȳ)² ] 或计算器的统计功能,她得到 r = –0.62。数值 –0.62 表明存在中等程度的负线性相关。学生注意到,相关性并不意味着因果关系,其他因素(如学习时间)也可能在起作用。


9. Regression Lines and Predictions | 回归线与预测

Because the scatter diagram suggests a linear relationship, the student fits a regression line of y on x to predict exam performance from social media hours. She computes the slope b = Σ[(x – x̄)(y – ȳ)] / Σ(x – x̄)² and the intercept a = ȳ – b x̄. With x̄ = 2.47 h, ȳ = 72.3%, she obtains the equation y = 86.5 – 5.74x. The negative slope indicates that, on average, each extra hour of daily social media use is associated with a drop of about 5.7 percentage points in the mock exam. She uses this line to predict that a student using social media for 4 h per day would score approximately 86.5 – 5.74×4 = 63.5%, but cautions that extrapolation beyond the data range could be unreliable.

由于散点图显示出线性关系,学生拟合了 y 对 x 的回归线,以便根据社交媒体使用时长预测考试成绩。她计算了斜率 b = Σ[(x – x̄)(y – ȳ)] / Σ(x – x̄)² 和截距 a = ȳ – b x̄。代入 x̄ = 2.47 小时,ȳ = 72.3%,她得到方程 y = 86.5 – 5.74x。负斜率表明,平均而言,每天额外使用一小时社交媒体,模拟考试成绩下降约 5.7 个百分点。她利用这条线预测,每天使用 4 小时社交媒体的学生得分约为 86.5 – 5.74×4 = 63.5%,但同时告诫,超出数据范围的推论可能不可靠。


10. Probability Models: Binomial Distribution for Survey Responses | 概率模型:调查反应的二项分布

A third question on the questionnaire asked: ‘Do you use social media for more than 3 hours per day?’ with a Yes/No response. In the sample of 30, 18 students answered Yes. The student models the number of Yes answers as a binomial random variable X ~ B(30, p), where p is the true proportion of Year 11 pupils who are heavy users. From the data, an estimate of p is p̂ = 18/30 = 0.6. She then calculates various binomial probabilities using the formula P(X = k) = ⁿCₖ pᵏ (1 – p)ⁿ⁻ᵏ or a calculator. For instance, if p = 0.5, the probability of getting exactly 18 heavy users is P(X = 18) = ³⁰C₁₈ (0.5)¹⁸ (0.5)¹² ≈ 0.0794.

问卷上的第三个问题是:‘你每天使用社交媒体的时间是否超过 3 小时?’回答为‘是’或‘否’。在 30 名学生样本中,有 18 人回答‘是’。这名学生将‘是’的数量建模为二项随机变量 X ~ B(30, p),其中 p 是十一年级学生中重度使用者的真实比例。根据数据,p 的估计值为 p̂ = 18/30 = 0.6。随后,她使用公式 P(X = k) = ⁿCₖ pᵏ (1 – p)ⁿ⁻ᵏ 或计算器计算各种二项概率。例如,如果 p = 0.5,则恰好得到 18 名重度使用者的概率为 P(X = 18) = ³⁰C₁₈ (0.5)¹⁸ (0.5)¹² ≈ 0.0794。


11. Hypothesis Testing Using the Binomial Distribution | 使用二项分布进行假设检验

The school newsletter once claimed that ‘fewer than half of Year 11 students spend more than 3 hours a day on social media’. The researcher decides to test this claim at the 5% significance level. She sets up the null hypothesis H₀: p = 0.5 and the alternative hypothesis H₁: p < 0.5, where p is the proportion of heavy users. The test statistic is the number of Yes responses in a sample of 30. From the data, X = 18. However, since H₁ is a lower‑tail test, she needs to find the critical region or compute the p‑value. The p‑value is P(X ≤ 18 | p = 0.5). Using binomial tables or a calculator, P(X ≤ 18) = 0.8998. This is far above 0.05, so there is insufficient evidence to reject H₀. The data do not support the newsletter's claim. In fact, the sample proportion 0.6 suggests the opposite might be true, but a two‑tailed test could be considered.

学校通讯此前声称“不到一半的十一年级学生每天使用社交媒体超过 3 小时”。研究者决定在 5% 的显著性水平下检验这一说法。她建立原假设 H₀: p = 0.5,备择假设 H₁: p < 0.5,其中 p 是重度使用者的比例。检验统计量为样本容量 30 中回答‘是’的个数。数据给出 X = 18。然而,由于 H₁ 是下尾检验,她需要找出临界域或计算 p 值。p 值为 P(X ≤ 18 | p = 0.5)。利用二项分布表或计算器,P(X ≤ 18) = 0.8998。该值远大于 0.05,因此没有足够证据拒绝 H₀。数据并不支持通讯上的说法。实际上,样本比例 0.6 暗示情况可能相反,但也可以考虑进行双尾检验。


12. Interpreting Results and Drawing Conclusions | 结果解读与结论得出

Bringing together all the analyses, the student writes a final report. She summarises that the typical daily social media use among the sampled Year 11 pupils is around 2.5 hours, with a moderate spread. The scatter plot and PMCC indicate a moderate negative correlation with exam performance, and the regression model suggests a drop of roughly 5.7% in marks per extra hour of use. The hypothesis test provided no evidence that the proportion of heavy users is below 50%. She acknowledges limitations: the sample size is small, data rely on self‑reporting, and correlation does not prove causality. To improve the investigation, she suggests collecting data on study hours and running a larger stratified sample. This reflective evaluation is exactly what WJEC examiners look for in high‑level responses.

综合所有分析,学生撰写了一份最终报告。她总结道,受调查的十一年级学生每日社交媒体使用时长的典型值约为 2.5 小时,离散程度适中。散点图和积差相关系数表明,社交媒体使用时长与考试成绩之间存在中等程度的负线性相关,回归模型预测每额外使用一小时,成绩下降约 5.7 个百分点。假设检验没有提供证据表明重度使用者的比例低于 50%。她也认识到研究的局限性:样本量较小,数据依赖自我报告,相关性并不能证明因果关系。为了改进调查,她建议收集学习时长数据并进行更大规模的分层抽样。这种反思性评价正是 WJEC 考官在高水平答案中所看重的。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading