📚 Case Study Practical Exercises | 案例分析实战演练
Welcome to this WJEC Statistics case study, which brings together key topics from the Year 11 syllabus in one realistic investigation. We follow a school survey of study habits, sleep, and exam performance, applying statistical techniques from data collection through to hypothesis testing. Each section presents a scenario, shows you how to work through the problem, and gives you the chance to practise similar skills on your own.
欢迎来到本次 WJEC 统计学案例研究,它将 Year 11 课程中的关键主题整合到一项真实的调查中。我们跟随一项关于学习习惯、睡眠和考试成绩的学校调查,应用从数据收集到假设检验的统计技术。每个部分都呈现一个情境,向你展示如何解决问题,并让你有机会自己练习类似的技能。
1. Data Collection and Sampling Methods | 数据收集与抽样方法
A school wants to investigate the study habits of its 200 Year 11 students. The target population includes all students in this year group. Using a simple random sample, each student would have an equal chance of being selected, but this might not represent subgroups like gender or class grouping. A stratified sample was chosen instead: the population was divided into strata by gender (100 male, 100 female) and a random sample of 25 males and 25 females was taken. This ensures proportional representation and improves the reliability of comparisons between genders.
一所学校希望调查其 200 名 Year 11 学生的学习习惯。目标总体包含该年级的所有学生。若采用简单随机抽样,每位学生被选中的机会相等,但这可能无法代表不同亚组,如性别或班级分组。因此选择了分层抽样:按性别将总体分层(100 名男生,100 名女生),并随机抽取 25 名男生和 25 名女生。这确保了比例代表性,并提高了性别间比较的可靠性。
Data were collected via a questionnaire asking for the number of hours spent studying per day, hours of sleep per night, and their most recent mock exam score (out of 100). A pilot survey was carried out to refine the questions, avoiding ambiguity and leading questions.
数据通过问卷收集,询问每天的学习小时数、每晚睡眠小时数,以及最近的模拟考试分数(满分 100)。在正式调查前进行了一次试点调查,以优化问题,避免模糊和诱导性问题。
2. Organising and Displaying Data | 数据整理与展示
The raw data for daily study hours (to the nearest half-hour) for the 25 male students were: 1.0, 1.5, 2.0, 2.0, 2.5, 3.0, 3.0, 3.5, 4.0, 4.0, 4.5, 4.5, 5.0, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, 9.0, 9.5, 10.0, 10.5. A back-to-back stem-and-leaf diagram is useful to compare male and female distributions. For instance, using stems 1 to 10, leaves represent the tenths. The male plot shows a relatively symmetrical spread, while a similar female dataset might reveal a higher concentration around 4–6 hours.
25 名男生每天学习时间的原始数据(精确到半小时)为:1.0, 1.5, 2.0, 2.0, 2.5, 3.0, 3.0, 3.5, 4.0, 4.0, 4.5, 4.5, 5.0, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, 9.0, 9.5, 10.0, 10.5。背靠背茎叶图有助于比较男生和女生的分布。例如,使用茎 1 至 10,叶代表十分位。男生图显示出相对对称的分布,而相似的女生数据集可能显示数据更集中于 4–6 小时。
A box plot can also be constructed using the five-number summary. For males: minimum = 1.0, Q₁ = 3.0, median = 5.0, Q₃ = 7.5, maximum = 10.5. The interquartile range (IQR) is 4.5 hours. Any value below Q₁ − 1.5 × IQR = −3.75 or above Q₃ + 1.5 × IQR = 14.25 would be an outlier; none are present. The whiskers extend to the minimum and maximum. A parallel box plot for females would allow direct visual comparison of location and spread.
箱线图也可以使用五数概括来构建。男生数据:最小值 = 1.0,Q₁ = 3.0,中位数 = 5.0,Q₃ = 7.5,最大值 = 10.5。四分位距 (IQR) 为 4.5 小时。任何低于 Q₁ − 1.5×IQR = −3.75 或高于 Q₃ + 1.5×IQR = 14.25 的值将被视为异常值;此处没有异常值。触须延伸至最小值和最大值。绘制女生的平行箱线图可以直接在视觉上比较位置和离散程度。
3. Measures of Central Tendency and Dispersion | 集中趋势与离散度量
For the male study hours, the mean is calculated as x̄ = Σx/n = (sum of all 25 values)/25. The sum is 137.5, so x̄ = 5.5 hours. The mode is 4.0 and 4.5 (bimodal). The sample standard deviation s = √[Σ(x − x̄)²/(n−1)] = 2.74 hours. This relatively large standard deviation shows considerable variation in study habits among males. For the female sample (not fully shown), suppose the mean is 4.8 hours and s = 1.9 hours. The smaller standard deviation indicates that females’ study hours are more consistent.
对于男生的学习时间,均值计算为 x̄ = Σx/n =(所有 25 个值的总和)/25。总和为 137.5,因此 x̄ = 5.5 小时。众数为 4.0 和 4.5(双峰)。样本标准差 s = √[Σ(x − x̄)²/(n−1)] = 2.74 小时。这个较大的标准差表明男生的学习习惯差异显著。对于女生样本(未完全展示),假设均值为 4.8 小时,s = 1.9 小时。较小的标准差表明女生的学习时间更为一致。
The IQR for males is 7.5 − 3.0 = 4.5 hours, which is a resistant measure of spread. If we compare the coefficients of variation (CV = s/x̄ × 100%), males have CV ≈ 49.8% and females ≈ 39.6%, confirming greater relative variability in males.
男生数据的四分位距为 7.5 − 3.0 = 4.5 小时,这是一种耐抗的离散度量。若比较变异系数 (CV = s/x̄ × 100%),男生 CV ≈ 49.8%,女生 ≈ 39.6%,证实男生的相对变异性更大。
4. Probability and Tree Diagrams | 概率与树形图
Based on the survey, the probability that a randomly chosen student studies more than 4 hours per day is 0.3, studies 2–4 hours is 0.5, and studies less than 2 hours is 0.2. The mock exam results are categorized as ‘high achievers’ (A/B). The conditional probabilities are: P(high | >4 hrs) = 0.7, P(high | 2–4 hrs) = 0.4, P(high | <2 hrs) = 0.1. A tree diagram helps visualise the problem: first branch for study time, second branch for exam outcome. The joint probabilities are obtained by multiplication.
根据调查,随机选择一名学生每天学习超过 4 小时的概率为 0.3,学习 2–4 小时为 0.5,少于 2 小时为 0.2。模拟考试结果分为“高分者”(A/B)。条件概率为:P(高分 | >4 小时) = 0.7,P(高分 | 2–4 小时) = 0.4,P(高分 | <2 小时) = 0.1。树形图有助于可视化问题:第一分支为学习时间,第二分支为考试结果。联合概率通过乘法得到。
The total probability that a student is a high achiever is P(high) = 0.3×0.7 + 0.5×0.4 + 0.2×0.1 = 0.21 + 0.20 + 0.02 = 0.43. Using Bayes’ theorem, the probability that a high achiever studied more than 4 hours is P(>4 hrs | high) = (0.21)/(0.43) ≈ 0.488, or about 48.8%.
学生成为高分者的总概率为 P(高分) = 0.3×0.7 + 0.5×0.4 + 0.2×0.1 = 0.21 + 0.20 + 0.02 = 0.43。利用贝叶斯定理,一名高分者学习超过 4 小时的概率为 P(>4 小时 | 高分) = (0.21)/(0.43) ≈ 0.488,约 48.8%。
5. Binomial Distribution: Modelling Successes | 二项分布:成功次数建模
A teacher claims that 80% of Year 11 students will achieve a grade C or above in the final exam. If we take a random sample of 10 students, and assuming each student’s outcome is independent with constant probability p = 0.8, we can model the number of successes using a binomial distribution X ~ B(10, 0.8). The probability of exactly 8 successes is P(X=8) = ¹⁰C₈ × (0.8)⁸ × (0.2)². ¹⁰C₈ = 45, so P(X=8) = 45 × 0.1678 × 0.04 ≈ 0.3020.
一位老师声称 80% 的 Year 11 学生将在期末考试中取得 C 级或以上成绩。若随机抽取 10 名学生,并假设每位学生的结果相互独立且具有恒定概率 p = 0.8,我们可以使用二项分布 X ~ B(10, 0.8) 来建模成功次数。恰好有 8 人成功的概率为 P(X=8) = ¹⁰C₈ × (0.8)⁸ × (0.2)²。¹⁰C₈ = 45,因此 P(X=8) = 45 × 0.1678 × 0.04 ≈ 0.3020。
To find P(X ≥ 9), we calculate P(X=9) + P(X=10). P(X=9) = ¹⁰C₉ × 0.8⁹ × 0.2¹ ≈ 10 × 0.1342 × 0.2 ≈ 0.2684. P(X=10) = 0.8¹⁰ ≈ 0.1074. Summing gives 0.3758. The expected number of successes is np = 8, so on average we would expect 8 out of 10 to pass.
求 P(X ≥ 9),我们计算 P(X=9) + P(X=10)。P(X=9) = ¹⁰C₉ × 0.8⁹ × 0.2¹ ≈ 10 × 0.1342 × 0.2 ≈ 0.2684。P(X=10) = 0.8¹⁰ ≈ 0.1074。总和为 0.3758。成功次数的期望值为 np = 8,因此平均而言,10 人中预计有 8 人通过。
6. Normal Distribution and Probability Calculations | 正态分布与概率计算
Mock exam scores across the whole year group are found to follow a normal distribution with mean μ = 60 and standard deviation σ = 12. To find the proportion of students scoring above 75, we standardise: z = (x − μ)/σ = (75 − 60)/12 = 1.25. Using standard normal tables, Φ(1.25) = 0.8944, so the upper tail probability is 1 − 0.8944 = 0.1056. About 10.6% of students score above 75.
模拟考试分数在整个年级组中服从均值为 μ = 60、标准差为 σ = 12 的正态分布。要找出分数高于 75 的比例,我们进行标准化:z = (x − μ)/σ = (75 − 60)/12 = 1.25。使用标准正态表,Φ(1.25) = 0.8944,因此上尾概率为 1 − 0.8944 = 0.1056。约 10.6% 的学生分数超过 75。
To find the symmetric interval containing the middle 80% of scores, we need the z-values that cut off 10% in each tail. Φ(z) = 0.9 gives z ≈ 1.2816. The lower bound is μ − zσ = 60 − 1.2816×12 ≈ 44.62, and the upper bound is 60 + 15.38 ≈ 75.38. So, the middle 80% of scores lie between about 45 and 75.
要找出包含中间 80% 分数的对称区间,我们需要截断每尾 10% 的 z 值。Φ(z) = 0.9 对应 z ≈ 1.2816。下限为 μ − zσ = 60 − 1.2816×12 ≈ 44.62,上限为 60 + 15.38 ≈ 75.38。因此,中间 80% 的分数介于约 45 到 75 之间。
7. Scatter Graphs and Correlation | 散点图与相关性
The school wants to investigate the linear relationship between hours of study (x) and mock exam score (y). Data from 10 randomly chosen students are recorded: (x,y) = (1,42), (2,48), (3,53), (4,59), (5,65), (6,66), (7,72), (8,76), (9,82), (10,88). The product moment correlation coefficient r measures the strength of the linear association. First, calculate sums: Σx = 55, Σy = 651, Σx² = 385, Σy² = 44043, Σxy = 3987. Then Sxx = Σx² − (Σx)²/n = 385 − 3025/10 = 82.5, Syy = 44043 − 651²/10 = 44043 − 42380.1 = 1662.9, Sxy = 3987 − (55×651)/10 = 3987 − 3580.5 = 406.5. Thus r = Sxy/√(Sxx Syy) = 406.5/√(82.5×1662.9) ≈ 406.5/370.2 ≈ 0.98. This indicates a very strong positive correlation.
学校希望研究学习小时数 (x) 与模拟考试分数 (y) 之间的线性关系。记录了 10 名随机学生的数据:(x,y) = (1,42), (2,48), (3,53), (4,59), (5,65), (6,66), (7,72), (8,76), (9,82), (10,88)。积矩相关系数 r 度量线性关联的强度。首先计算和:Σx = 55,Σy = 651,Σx² = 385,Σy² = 44043,Σxy = 3987。然后 Sxx = Σx² − (Σx)²/n = 385 − 3025/10 = 82.5,Syy = 44043 − 651²/10 = 44043 − 42380.1 = 1662.9,Sxy = 3987 − (55×651)/10 = 3987 − 3580.5 = 406.5。因此 r = Sxy/√(Sxx Syy) = 406.5/√(82.5×1662.9) ≈ 406.5/370.2 ≈ 0.98。这表明存在极强的正相关。
8. Linear Regression | 线性回归
Using the same data, we find the equation of the regression line of y on x: y = a + bx. The slope b = Sxy/Sxx = 406.5/82.5 = 4.927. The intercept a = ȳ − b x̄. x̄ = 55/10 = 5.5, ȳ = 651/10 = 65.1. So a = 65.1 − 4.927×5.5 ≈ 65.1 − 27.1 = 38.0. The regression equation is y = 38.0 + 4.93x (to 3 s.f.). This means that for each additional hour studied, the predicted exam score increases by about 4.93 marks.
利用相同的数据,我们求 y 对 x 的回归线方程:y = a + bx。斜率 b = Sxy/Sxx = 406.5/82.5 = 4.927。截距 a = ȳ − b x̄。x̄ = 55/10 = 5.5,ȳ = 651/10 = 65.1。因此 a = 65.1 − 4.927×5.5 ≈ 65.1 − 27.1 = 38.0。回归方程为 y = 38.0 + 4.93x(保留三位有效数字)。这意味着每多学习一个小时,预测的考试分数增加约 4.93 分。
We can use the equation to predict the score for a student who studies 5.5 hours: y = 38.0 + 4.93×5.5 ≈ 65.1, which is exactly the mean. Prediction for 12 hours would be 38.0 + 59.16 = 97.16, but this is outside the range of observed x-values, so its reliability is questionable.
我们可以用该方程预测一名学习 5.5 小时的学生的分数:y = 38.0 + 4.93×5.5 ≈ 65.1,正好是均值。预测 12 小时的分数为 38.0 + 59.16 = 97.16,但这超出了 x 的观测范围,因此其可靠性值得怀疑。
9. Time Series Analysis and Moving Averages | 时间序列分析与移动平均
A student recorded their weekly study hours over an 8-week term: Week 1: 20, Week 2: 18, Week 3: 22, Week 4: 19, Week 5: 21, Week 6: 17, Week 7: 23, Week 8: 16. To smooth out the erratic fluctuations and reveal the underlying trend, we calculate a 4-point moving average. The first average is (20+18+22+19)/4 = 19.75, centred at week 2.5; the next is (18+22+19+21)/4 = 20.0, centred at week 3.5; continuing gives averages: 20.0 (wk3.5), 19.75 (wk4.5), 20.0 (wk5.5), 19.25 (wk6.5). Because the number of points is even, centring requires averaging adjacent moving averages. The centred moving averages are: (19.75+20.0)/2=19.875 (wk3), (20.0+19.75)/2=19.875 (wk4), (19.75+20.0)/2=19.875 (wk5), (20.0+19.25)/2=19.625 (wk6). The trend appears fairly flat around 19.8, suggesting little overall change during the middle weeks.
一名学生记录了 8 周学期中每周的学习小时数:第 1 周:20,第 2 周:18,第 3 周:22,第 4 周:19,第 5 周:21,第 6 周:17,第 7 周:23,第 8 周:16。为了平滑无规律的波动并揭示基本趋势,我们计算 4 点移动平均。第一个平均值为 (20+18+22+19)/4 = 19.75,中心位于第 2.5 周;下一个为 (18+22+19+21)/4 = 20.0,中心位于第 3.5 周;继续得到平均值:20.0(第
Published by TutorHao | Year 11 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导