Case Study Practical Exercises | 案例分析实战演练

📚 Case Study Practical Exercises | 案例分析实战演练

Welcome to this WJEC Statistics case study, which brings together key topics from the Year 11 syllabus in one realistic investigation. We follow a school survey of study habits, sleep, and exam performance, applying statistical techniques from data collection through to hypothesis testing. Each section presents a scenario, shows you how to work through the problem, and gives you the chance to practise similar skills on your own.

欢迎来到本次 WJEC 统计学案例研究,它将 Year 11 课程中的关键主题整合到一项真实的调查中。我们跟随一项关于学习习惯、睡眠和考试成绩的学校调查,应用从数据收集到假设检验的统计技术。每个部分都呈现一个情境,向你展示如何解决问题,并让你有机会自己练习类似的技能。


1. Data Collection and Sampling Methods | 数据收集与抽样方法

A school wants to investigate the study habits of its 200 Year 11 students. The target population includes all students in this year group. Using a simple random sample, each student would have an equal chance of being selected, but this might not represent subgroups like gender or class grouping. A stratified sample was chosen instead: the population was divided into strata by gender (100 male, 100 female) and a random sample of 25 males and 25 females was taken. This ensures proportional representation and improves the reliability of comparisons between genders.

一所学校希望调查其 200 名 Year 11 学生的学习习惯。目标总体包含该年级的所有学生。若采用简单随机抽样,每位学生被选中的机会相等,但这可能无法代表不同亚组,如性别或班级分组。因此选择了分层抽样:按性别将总体分层(100 名男生,100 名女生),并随机抽取 25 名男生和 25 名女生。这确保了比例代表性,并提高了性别间比较的可靠性。

Data were collected via a questionnaire asking for the number of hours spent studying per day, hours of sleep per night, and their most recent mock exam score (out of 100). A pilot survey was carried out to refine the questions, avoiding ambiguity and leading questions.

数据通过问卷收集,询问每天的学习小时数、每晚睡眠小时数,以及最近的模拟考试分数(满分 100)。在正式调查前进行了一次试点调查,以优化问题,避免模糊和诱导性问题。


2. Organising and Displaying Data | 数据整理与展示

The raw data for daily study hours (to the nearest half-hour) for the 25 male students were: 1.0, 1.5, 2.0, 2.0, 2.5, 3.0, 3.0, 3.5, 4.0, 4.0, 4.5, 4.5, 5.0, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, 9.0, 9.5, 10.0, 10.5. A back-to-back stem-and-leaf diagram is useful to compare male and female distributions. For instance, using stems 1 to 10, leaves represent the tenths. The male plot shows a relatively symmetrical spread, while a similar female dataset might reveal a higher concentration around 4–6 hours.

25 名男生每天学习时间的原始数据(精确到半小时)为:1.0, 1.5, 2.0, 2.0, 2.5, 3.0, 3.0, 3.5, 4.0, 4.0, 4.5, 4.5, 5.0, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, 9.0, 9.5, 10.0, 10.5。背靠背茎叶图有助于比较男生和女生的分布。例如,使用茎 1 至 10,叶代表十分位。男生图显示出相对对称的分布,而相似的女生数据集可能显示数据更集中于 4–6 小时。

A box plot can also be constructed using the five-number summary. For males: minimum = 1.0, Q₁ = 3.0, median = 5.0, Q₃ = 7.5, maximum = 10.5. The interquartile range (IQR) is 4.5 hours. Any value below Q₁ − 1.5 × IQR = −3.75 or above Q₃ + 1.5 × IQR = 14.25 would be an outlier; none are present. The whiskers extend to the minimum and maximum. A parallel box plot for females would allow direct visual comparison of location and spread.

箱线图也可以使用五数概括来构建。男生数据:最小值 = 1.0,Q₁ = 3.0,中位数 = 5.0,Q₃ = 7.5,最大值 = 10.5。四分位距 (IQR) 为 4.5 小时。任何低于 Q₁ − 1.5×IQR = −3.75 或高于 Q₃ + 1.5×IQR = 14.25 的值将被视为异常值;此处没有异常值。触须延伸至最小值和最大值。绘制女生的平行箱线图可以直接在视觉上比较位置和离散程度。


3. Measures of Central Tendency and Dispersion | 集中趋势与离散度量

For the male study hours, the mean is calculated as x̄ = Σx/n = (sum of all 25 values)/25. The sum is 137.5, so x̄ = 5.5 hours. The mode is 4.0 and 4.5 (bimodal). The sample standard deviation s = √[Σ(x − x̄)²/(n−1)] = 2.74 hours. This relatively large standard deviation shows considerable variation in study habits among males. For the female sample (not fully shown), suppose the mean is 4.8 hours and s = 1.9 hours. The smaller standard deviation indicates that females’ study hours are more consistent.

对于男生的学习时间,均值计算为 x̄ = Σx/n =(所有 25 个值的总和)/25。总和为 137.5,因此 x̄ = 5.5 小时。众数为 4.0 和 4.5(双峰)。样本标准差 s = √[Σ(x − x̄)²/(n−1)] = 2.74 小时。这个较大的标准差表明男生的学习习惯差异显著。对于女生样本(未完全展示),假设均值为 4.8 小时,s = 1.9 小时。较小的标准差表明女生的学习时间更为一致。

The IQR for males is 7.5 − 3.0 = 4.5 hours, which is a resistant measure of spread. If we compare the coefficients of variation (CV = s/x̄ × 100%), males have CV ≈ 49.8% and females ≈ 39.6%, confirming greater relative variability in males.

男生数据的四分位距为 7.5 − 3.0 = 4.5 小时,这是一种耐抗的离散度量。若比较变异系数 (CV = s/x̄ × 100%),男生 CV ≈ 49.8%,女生 ≈ 39.6%,证实男生的相对变异性更大。


4. Probability and Tree Diagrams | 概率与树形图

Based on the survey, the probability that a randomly chosen student studies more than 4 hours per day is 0.3, studies 2–4 hours is 0.5, and studies less than 2 hours is 0.2. The mock exam results are categorized as ‘high achievers’ (A/B). The conditional probabilities are: P(high | >4 hrs) = 0.7, P(high | 2–4 hrs) = 0.4, P(high | <2 hrs) = 0.1. A tree diagram helps visualise the problem: first branch for study time, second branch for exam outcome. The joint probabilities are obtained by multiplication.

根据调查,随机选择一名学生每天学习超过 4 小时的概率为 0.3,学习 2–4 小时为 0.5,少于 2 小时为 0.2。模拟考试结果分为“高分者”(A/B)。条件概率为:P(高分 | >4 小时) = 0.7,P(高分 | 2–4 小时) = 0.4,P(高分 | <2 小时) = 0.1。树形图有助于可视化问题:第一分支为学习时间,第二分支为考试结果。联合概率通过乘法得到。

The total probability that a student is a high achiever is P(high) = 0.3×0.7 + 0.5×0.4 + 0.2×0.1 = 0.21 + 0.20 + 0.02 = 0.43. Using Bayes’ theorem, the probability that a high achiever studied more than 4 hours is P(>4 hrs | high) = (0.21)/(0.43) ≈ 0.488, or about 48.8%.

学生成为高分者的总概率为 P(高分) = 0.3×0.7 + 0.5×0.4 + 0.2×0.1 = 0.21 + 0.20 + 0.02 = 0.43。利用贝叶斯定理,一名高分者学习超过 4 小时的概率为 P(>4 小时 | 高分) = (0.21)/(0.43) ≈ 0.488,约 48.8%。


5. Binomial Distribution: Modelling Successes | 二项分布:成功次数建模

A teacher claims that 80% of Year 11 students will achieve a grade C or above in the final exam. If we take a random sample of 10 students, and assuming each student’s outcome is independent with constant probability p = 0.8, we can model the number of successes using a binomial distribution X ~ B(10, 0.8). The probability of exactly 8 successes is P(X=8) = ¹⁰C₈ × (0.8)⁸ × (0.2)². ¹⁰C₈ = 45, so P(X=8) = 45 × 0.1678 × 0.04 ≈ 0.3020.

一位老师声称 80% 的 Year 11 学生将在期末考试中取得 C 级或以上成绩。若随机抽取 10 名学生,并假设每位学生的结果相互独立且具有恒定概率 p = 0.8,我们可以使用二项分布 X ~ B(10, 0.8) 来建模成功次数。恰好有 8 人成功的概率为 P(X=8) = ¹⁰C₈ × (0.8)⁸ × (0.2)²。¹⁰C₈ = 45,因此 P(X=8) = 45 × 0.1678 × 0.04 ≈ 0.3020。

To find P(X ≥ 9), we calculate P(X=9) + P(X=10). P(X=9) = ¹⁰C₉ × 0.8⁹ × 0.2¹ ≈ 10 × 0.1342 × 0.2 ≈ 0.2684. P(X=10) = 0.8¹⁰ ≈ 0.1074. Summing gives 0.3758. The expected number of successes is np = 8, so on average we would expect 8 out of 10 to pass.

求 P(X ≥ 9),我们计算 P(X=9) + P(X=10)。P(X=9) = ¹⁰C₉ × 0.8⁹ × 0.2¹ ≈ 10 × 0.1342 × 0.2 ≈ 0.2684。P(X=10) = 0.8¹⁰ ≈ 0.1074。总和为 0.3758。成功次数的期望值为 np = 8,因此平均而言,10 人中预计有 8 人通过。


6. Normal Distribution and Probability Calculations | 正态分布与概率计算

Mock exam scores across the whole year group are found to follow a normal distribution with mean μ = 60 and standard deviation σ = 12. To find the proportion of students scoring above 75, we standardise: z = (x − μ)/σ = (75 − 60)/12 = 1.25. Using standard normal tables, Φ(1.25) = 0.8944, so the upper tail probability is 1 − 0.8944 = 0.1056. About 10.6% of students score above 75.

模拟考试分数在整个年级组中服从均值为 μ = 60、标准差为 σ = 12 的正态分布。要找出分数高于 75 的比例,我们进行标准化:z = (x − μ)/σ = (75 − 60)/12 = 1.25。使用标准正态表,Φ(1.25) = 0.8944,因此上尾概率为 1 − 0.8944 = 0.1056。约 10.6% 的学生分数超过 75。

To find the symmetric interval containing the middle 80% of scores, we need the z-values that cut off 10% in each tail. Φ(z) = 0.9 gives z ≈ 1.2816. The lower bound is μ − zσ = 60 − 1.2816×12 ≈ 44.62, and the upper bound is 60 + 15.38 ≈ 75.38. So, the middle 80% of scores lie between about 45 and 75.

要找出包含中间 80% 分数的对称区间,我们需要截断每尾 10% 的 z 值。Φ(z) = 0.9 对应 z ≈ 1.2816。下限为 μ − zσ = 60 − 1.2816×12 ≈ 44.62,上限为 60 + 15.38 ≈ 75.38。因此,中间 80% 的分数介于约 45 到 75 之间。


7. Scatter Graphs and Correlation | 散点图与相关性

The school wants to investigate the linear relationship between hours of study (x) and mock exam score (y). Data from 10 randomly chosen students are recorded: (x,y) = (1,42), (2,48), (3,53), (4,59), (5,65), (6,66), (7,72), (8,76), (9,82), (10,88). The product moment correlation coefficient r measures the strength of the linear association. First, calculate sums: Σx = 55, Σy = 651, Σx² = 385, Σy² = 44043, Σxy = 3987. Then Sxx = Σx² − (Σx)²/n = 385 − 3025/10 = 82.5, Syy = 44043 − 651²/10 = 44043 − 42380.1 = 1662.9, Sxy = 3987 − (55×651)/10 = 3987 − 3580.5 = 406.5. Thus r = Sxy/√(Sxx Syy) = 406.5/√(82.5×1662.9) ≈ 406.5/370.2 ≈ 0.98. This indicates a very strong positive correlation.

学校希望研究学习小时数 (x) 与模拟考试分数 (y) 之间的线性关系。记录了 10 名随机学生的数据:(x,y) = (1,42), (2,48), (3,53), (4,59), (5,65), (6,66), (7,72), (8,76), (9,82), (10,88)。积矩相关系数 r 度量线性关联的强度。首先计算和:Σx = 55,Σy = 651,Σx² = 385,Σy² = 44043,Σxy = 3987。然后 Sxx = Σx² − (Σx)²/n = 385 − 3025/10 = 82.5,Syy = 44043 − 651²/10 = 44043 − 42380.1 = 1662.9,Sxy = 3987 − (55×651)/10 = 3987 − 3580.5 = 406.5。因此 r = Sxy/√(Sxx Syy) = 406.5/√(82.5×1662.9) ≈ 406.5/370.2 ≈ 0.98。这表明存在极强的正相关。


8. Linear Regression | 线性回归

Using the same data, we find the equation of the regression line of y on x: y = a + bx. The slope b = Sxy/Sxx = 406.5/82.5 = 4.927. The intercept a = ȳ − b x̄. x̄ = 55/10 = 5.5, ȳ = 651/10 = 65.1. So a = 65.1 − 4.927×5.5 ≈ 65.1 − 27.1 = 38.0. The regression equation is y = 38.0 + 4.93x (to 3 s.f.). This means that for each additional hour studied, the predicted exam score increases by about 4.93 marks.

利用相同的数据,我们求 y 对 x 的回归线方程:y = a + bx。斜率 b = Sxy/Sxx = 406.5/82.5 = 4.927。截距 a = ȳ − b x̄。x̄ = 55/10 = 5.5,ȳ = 651/10 = 65.1。因此 a = 65.1 − 4.927×5.5 ≈ 65.1 − 27.1 = 38.0。回归方程为 y = 38.0 + 4.93x(保留三位有效数字)。这意味着每多学习一个小时,预测的考试分数增加约 4.93 分。

We can use the equation to predict the score for a student who studies 5.5 hours: y = 38.0 + 4.93×5.5 ≈ 65.1, which is exactly the mean. Prediction for 12 hours would be 38.0 + 59.16 = 97.16, but this is outside the range of observed x-values, so its reliability is questionable.

我们可以用该方程预测一名学习 5.5 小时的学生的分数:y = 38.0 + 4.93×5.5 ≈ 65.1,正好是均值。预测 12 小时的分数为 38.0 + 59.16 = 97.16,但这超出了 x 的观测范围,因此其可靠性值得怀疑。


9. Time Series Analysis and Moving Averages | 时间序列分析与移动平均

A student recorded their weekly study hours over an 8-week term: Week 1: 20, Week 2: 18, Week 3: 22, Week 4: 19, Week 5: 21, Week 6: 17, Week 7: 23, Week 8: 16. To smooth out the erratic fluctuations and reveal the underlying trend, we calculate a 4-point moving average. The first average is (20+18+22+19)/4 = 19.75, centred at week 2.5; the next is (18+22+19+21)/4 = 20.0, centred at week 3.5; continuing gives averages: 20.0 (wk3.5), 19.75 (wk4.5), 20.0 (wk5.5), 19.25 (wk6.5). Because the number of points is even, centring requires averaging adjacent moving averages. The centred moving averages are: (19.75+20.0)/2=19.875 (wk3), (20.0+19.75)/2=19.875 (wk4), (19.75+20.0)/2=19.875 (wk5), (20.0+19.25)/2=19.625 (wk6). The trend appears fairly flat around 19.8, suggesting little overall change during the middle weeks.

一名学生记录了 8 周学期中每周的学习小时数:第 1 周:20,第 2 周:18,第 3 周:22,第 4 周:19,第 5 周:21,第 6 周:17,第 7 周:23,第 8 周:16。为了平滑无规律的波动并揭示基本趋势,我们计算 4 点移动平均。第一个平均值为 (20+18+22+19)/4 = 19.75,中心位于第 2.5 周;下一个为 (18+22+19+21)/4 = 20.0,中心位于第 3.5 周;继续得到平均值:20.0(第

Published by TutorHao | Year 11 统计 Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading