📚 Case Study Drill: Applying IGCSE Statistics to Real Data | 案例分析实战演练:将IGCSE统计学应用于真实数据
In IGCSE Statistics, mastering calculations is only half the story. The exam will often present you with a real-world scenario and ask you to select appropriate methods, interpret results, and critique the process. This case study walks you through a complete investigation, from data collection to evaluation, mimicking the style of Edexcel exam tasks. You will encounter frequency tables, visual representations, measures of centre and spread, cumulative frequency, correlation, probability, and sampling considerations – all woven into one coherent story.
在IGCSE统计学中,掌握计算只是成功的一半。考试通常会给你一个真实场景,要求你选择合适的统计方法、解释结果并评判整个过程。本案例分析将带你走过一个完整的调查研究,从数据收集到评估,模拟Edexcel考试任务的风格。你将接触频数表、可视化呈现、中心与离散量数、累积频数、相关性、概率以及抽样注意事项——所有这些都编织成一个连贯的故事。
1. Setting the Scene: The ‘Boost Revision’ Programme | 场景设定:’Boost Revision’复习计划
Greenwood International School launched a revision initiative called ‘Boost Revision’ for its Year 11 students preparing for the IGCSE Mathematics exam. The programme involved targeted small-group tutoring and daily retrieval practice over six weeks. The school wants to know if the programme led to a genuine improvement in mock exam scores. Your role as the statistician is to design the study, collect data, analyse the results, and present evidence-based conclusions.
格林伍德国际学校为准备IGCSE数学考试的11年级学生发起了一项名为’Boost Revision’的复习计划。该计划包括为期六周的有针对性小组辅导和每日提取练习。学校想知道该计划是否真正提高了模拟考试成绩。作为统计专家,你的任务是设计研究方案、收集数据、分析结果并提出基于证据的结论。
Before-and-after mock scores (out of 100) were recorded for a random sample of 20 students who participated. Here are the scores:
我们记录了参与计划的20名随机抽样学生模拟考试的前后成绩(满分100)。得分如下:
| Student | Before (x) | After (y) |
|---|---|---|
| 1 | 48 | 56 |
| 2 | 62 | 70 |
| 3 | 55 | 60 |
| 4 | 70 | 74 |
| 5 | 38 | 52 |
| 6 | 75 | 81 |
| 7 | 42 | 48 |
| 8 | 51 | 58 |
| 9 | 80 | 85 |
| 10 | 65 | 72 |
| 11 | 44 | 50 |
| 12 | 58 | 63 |
| 13 | 69 | 75 |
| 14 | 47 | 53 |
| 15 | 72 | 79 |
| 16 | 36 | 46 |
| 17 | 63 | 68 |
| 18 | 53 | 59 |
| 19 | 78 | 84 |
| 20 | 57 | 61 |
2. Data Collection Methods | 数据收集方法
The ‘before’ scores were obtained from the mock exam sat in January, while the ‘after’ scores came from a second mock in April, after the programme ended. The data were collected via the school’s assessment database, ensuring accuracy. Using a random sample of 20 out of 80 participants reduces selection bias and allows us to apply statistical inference, though a larger sample would have provided more precise estimates.
‘之前’成绩来自1月的模拟考试,’之后’成绩来自计划结束后的4月第二次模拟考试。数据通过学校的评估数据库收集,确保了准确性。从80名参与者中随机抽取20人减少了选择偏差,并允许我们进行统计推断,尽管更大的样本会提供更精确的估计。
Note that the same students are measured before and after, so this is a paired or dependent design. This is essential when assessing change because it controls for individual aptitude differences. Data like these are continuous (scores out of 100), but can be grouped for frequency distributions.
请注意,相同的学生在前后两个时间点被测量,因此这是一个配对或相关设计。在评估变化时这一点至关重要,因为它控制了个人能力差异。这类数据是连续的(满分100分),但可以分组形成频数分布。
3. Organising Raw Data into Frequency Tables | 将原始数据整理为频数表
To start the analysis, we group the ‘Before’ scores into classes. For IGCSE, a sensible number of groups is often 5 to 8. With scores ranging from 36 to 80, class intervals of width 10 are convenient.
为了开始分析,我们将’之前’成绩分组。在IGCSE中,合理的组数通常是5到8组。分数范围从36到80,组距为10很方便。
| Score Interval (Before) | Frequency (f) |
|---|---|
| 30 ≤ x < 40 | 2 |
| 40 ≤ x < 50 | 4 |
| 50 ≤ x < 60 | 5 |
| 60 ≤ x < 70 | 4 |
| 70 ≤ x < 80 | 4 |
| 80 ≤ x < 90 | 1 |
Always check the total frequency equals the sample size (n = 20). The grouped ‘After’ scores can be tabulated similarly, and comparing the shapes of the two distributions will be our first visual insight.
务必检查总频数等于样本量(n = 20)。分组的’之后’成绩可以类似地制成表格,比较两个分布的形状将是我们第一个可视化洞察。
4. Visualising Data: Bar Charts and Histograms | 数据可视化:条形图与直方图
For discrete or categorical data we use bar charts, but for this continuous grouped data, the correct graph is a histogram. Since class intervals are equal, the height of each bar represents frequency. Plotting both ‘Before’ and ‘After’ histograms on the same axes reveals a rightward shift in the distribution after the programme, suggesting improvement.
对于离散或分类数据我们使用条形图,但对于这些连续分组数据,正确的图形是直方图。由于组距相等,每个条形的高度代表频数。将’之前’和’之后’的直方图绘制在同一坐标轴上,显示出计划后分布向右移动,表明有提高。
A frequency polygon can also be overlaid, joining the midpoints of the tops of the bars. The modal class before was 50 ≤ x < 60, whereas after it shifted to 70 ≤ x < 80, a clear positive impact. In an exam, you would be expected to describe the shape (skewness) and centre.
还可以叠加频数多边形,连接条形顶点的中点。之前的众数所在组是50 ≤ x < 60,而之后转变为70 ≤ x < 80,有明显的正面影响。在考试中,你需要描述形状(偏态)和中心。
5. Measures of Central Tendency: Mean, Median, Mode | 集中趋势量数:平均值、中位数、众数
From the raw ‘Before’ data, the mean is calculated by summing all scores and dividing by 20:
从原始’之前’数据中,平均值通过将所有分数相加并除以20来计算:
Sum of before = 48+62+55+70+38+75+42+51+80+65+44+58+69+47+72+36+63+53+78+57 = 1,163
Mean before = 1163 ÷ 20 = 58.15
Similarly, for ‘After’ data: sum = 56+70+60+74+52+81+48+58+85+72+50+63+75+53+79+46+68+59+84+61 = 1,294.
类似地,’之后’数据:总和 = 1294。
Mean after = 1294 ÷ 20 = 64.7
The mean increased by 6.55 points. To find the median, order the 20 values and average the 10th and 11th. For ‘Before’, ordered values: 36,38,42,44,47,48,51,53,55,57,58,62,63,65,69,70,72,75,78,80. The median is (57+58)/2 = 57.5. For ‘After’, the ordered list gives median (63+68)/2 = 65.5. Thus all central measures point to an upward shift.
平均值提高了6.55分。为求中位数,将20个值排序并取第10和第11个的平均数。’之前’排序后中位数为(57+58)/2 = 57.5。’之后’中位数为(63+68)/2 = 65.5。因此所有中心量数都指向上升趋势。
The mode can be identified from the highest frequency in a stem-and-leaf diagram or frequency table. Before, the mode is around 57-58; after, it’s higher.
众数可以通过茎叶图或频数表中的最高频数来确定。之前众数约在57-58;之后更高。
6. Measures of Spread: Range, Interquartile Range, Standard Deviation | 离散程度量数:极差、四分位距、标准差
Range before = 80 – 36 = 44; Range after = 85 – 46 = 39. The slight reduction in range suggests scores became more consistent after the programme. The interquartile range (IQR) is more robust. Using the ordered list, find lower quartile (median of first 10 values) and upper quartile (median of last 10).
之前极差 = 80 – 36 = 44;之后极差 = 85 – 46 = 39。极差略微缩小,表明计划后成绩变得更一致。四分位距(IQR)更为稳健。使用排序列表,求下四分位数(前10个的中位数)和上四分位数(后10个的中位数)。
Before: first 10 = 36,38,42,44,47,48,51,53,55,57 → Q₁ = (47+48)/2 = 47.5; last 10 = 58,62,63,65,69,70,72,75,78,80 → Q₃ = (69+70)/2 = 69.5. IQR = 69.5 – 47.5 = 22.0. After: Q₁ = 53.5; Q₃ = 76.0; IQR = 22.5 – very similar spread.
之前:Q₁ = 47.5;Q₃ = 69.5;IQR = 22.0。之后:Q₁ = 53.5;Q₃ = 76.0;IQR = 22.5 – 非常相近的离散程度。
Standard deviation (population estimate, s) uses the formula: s = √(Σ(x – x̄)²/(n-1)). For ‘Before’, computing squared deviations from 58.15 gives Σ(x – x̄)² ≈ 2862.55. Then s_before = √(2862.55/19) ≈ √150.66 ≈ 12.27. For ‘After’, mean = 64.7, Σ(y – ȳ)² ≈ 2538.2, s_after = √(2538.2/19) ≈ √133.59 ≈ 11.56. The standard deviation has decreased slightly, supporting the range finding of improved consistency.
标准差(总体估计,s)使用公式:s = √(Σ(x – x̄)²/(n-1))。对于’之前’,计算与58.15的偏差平方和得到约2862.55。则s之前 ≈ 12.27。对于’之后’,s之后 ≈ 11.56。标准差略微下降,支持了成绩变得更一致的极差发现。
7. Cumulative Frequency and Box Plots | 累积频数与箱线图
Cumulative frequency allows us to find percentiles and draw box plots. Let’s construct the cumulative frequency table for ‘After’ scores using the same groupings:
累积频数使我们能够找到百分位数并绘制箱线图。我们为’之后’成绩构建累积频数表,使用相同的分组:
| Interval | f | Cumulative f |
|---|---|---|
| 30 ≤ y < 40 | 0 | 0 |
| 40 ≤ y < 50 | 2 | 2 |
| 50 ≤ y < 60 | 4 | 6 |
| 60 ≤ y < 70 | 4 | 10 |
| 70 ≤ y < 80 | 7 | 17 |
| 80 ≤ y < 90 | 3 | 20 |
Plot the cumulative frequency curve (ogive) by joining points (upper class boundary, cumulative frequency). From the curve, we can read off median ≈ 66, Q₁ ≈ 53, Q₃ ≈ 78. A box plot for both datasets side by side confirms the programme shifted the entire score distribution upward while preserving similar variability.
通过点(上限,累积频数)绘制累积频数曲线(肩形图)。从曲线上我们可以读出中位数≈66,Q₁≈53,Q₃≈78。将两个数据集的箱线图并排比较,证实该计划使整个分数分布上移,同时保持了相似的变异性。
8. Scatter Graphs and Correlation | 散点图与相关性
A fascinating question: is the improvement related to the starting score? Plot ‘Before’ on the x-axis and ‘After’ on the y-axis. You would see a strong, positive linear relationship – students who scored higher before tended to score higher after. The correlation coefficient r can be calculated; here r ≈ 0.97, indicating very strong positive correlation.
一个有趣的问题:进步幅度与起始分数有关吗?将’之前’放在x轴,’之后’放在y轴。你会看到强正线性关系——之前得分较高的学生在之后也往往得分较高。相关系数r可以计算出来;这里r≈0.97,表明非常强的正相关。
However, to assess the programme’s effect, we should look at the improvement (difference = after – before). The differences are: 8,8,5,4,14,6,6,7,5,7,6,5,6,6,7,10,5,6,6,4. The mean improvement is 6.45 with a small variation. Plotting ‘Before’ against the difference shows no clear pattern – the programme benefitted students across the ability spectrum. This is an important statistical insight: correlation vs. causal effect.
然而,为评估计划效果,我们应该看进步幅度(差值=之后 – 之前)。差值序列如上,平均进步6.45,变异很小。将’之前’与差值作图没有清晰模式——该计划使各能力层次的学生都受益。这是一个重要的统计学洞察:相关性与因果效应。
9. Probability from Two-Way Tables | 来自双向表的概率
We can categorise students based on their ‘Before’ performance, e.g., ‘At or above median’ vs ‘Below median’. Create a two-way table for before median (57.5) and after median (65.5):
我们可以根据“之前”表现对学生进行分类,例如“在中位数及以上”与“低于中位数”。创建一个关于之前中位数(57.5)和之后中位数(65.5)的双向表:
| After ≥ 65.5 | After < 65.5 | Total | |
|---|---|---|---|
| Before ≥ 57.5 | 9 | 1 | 10 |
| Before < 57.5 | 1 | 9 | 10 |
| Total | 10 | 10 | 20 |
From this table, the probability that a randomly selected student was at or above the after-median, given they were below the before-median, is 1/10 = 0.1. In contrast, P(at or above after-median | at or above before-median) = 9/10 = 0.9. This suggests that the programme did not completely overcome initial gaps, but many below-median students did improve – a nuanced conclusion.
从此表可知,随机选择一名学生,已知其之前低于中位数,其后达到或超过之后中位数的概率是1/10 = 0.1。而P(达到或超过之后中位数 | 之前达到或超过中位数) = 9/10 = 0.9。这表明该计划并未完全消除初始差距,但许多低于中位数的学生确实取得了进步——一个微妙的结论。
10. Sampling and Bias Discussion | 抽样与偏差讨论
The sample of 20 was randomly selected from the 80 participants, which helps reduce selection bias. However, the study only includes volunteer participants who chose to attend the revision programme. They may be more motivated than non-participants, introducing volunteer bias. Also, the ‘after’ mock exam may have been easier than the ‘before’ mock, introducing a possible confounding variable. A controlled experiment with a matched control group would strengthen causal claims.
20人的样本是从80名参与者中随机抽取的,这有助于减少选择偏差。然而,研究只包括自愿参加复习计划的学生。他们可能比未参加者更有动力,从而引入志愿者偏差。此外,“之后”的模拟考试可能比“之前”的考试更容易,引入了一个可能的混杂变量。一个有配对对照组的受控实验将加强因果推断。
Measurement error could occur if different markers graded the mocks with different standards. The school ensured both mocks were marked by the same team using identical mark schemes, enhancing reliability.
如果不同的评分者以不同的标准批改模拟试卷,测量误差可能发生。学校确保两次模拟考试由同一团队使用相同的评分方案批改,提高了可靠性。
11. Drawing Conclusions and Evaluating the Programme | 得出结论并评估计划
The evidence strongly suggests that the ‘Boost Revision’ programme was associated with an average score increase of about 6.5 marks. The shift is visible in all measures: mean, median, box plots. Spread remained roughly constant, so the programme affected all types of students. However, from the two-way table, we note that a student’s prior attainment still strongly predicts their later attainment.
证据强烈表明,’Boost Revision’复习计划与分数平均提高约6.5分相关。这一变化在所有量数中均可见:平均值、中位数、箱线图。离散程度基本保持不变,因此该计划影响了各类学生。然而,从双向表我们注意到,学生之前的成绩仍然强烈预测其之后的成绩。
Critically, we cannot definitively claim causation because this is an observational study with no control group of non-participating students. The school should consider a follow-up study comparing participants with a matched sample of non-participants to isolate the programme’s effect.
关键的是,我们不能明确声称因果关系,因为这是一项没有非参与对照组学生的观察性研究。学校应考虑进行一项后续研究,将参与者与匹配的非参与者样本进行比较,以分离出该计划的效果。
12. Exam-Style Case Study Questions | 考试风格案例研究问题
To prepare for the IGCSE exam, practise answering questions that mimic the Edexcel style:
为了准备IGCSE考试,请练习回答模仿Edexcel风格的问题:
Question 1: Using the raw data, construct a stem-and-leaf diagram for the ‘After’ scores and find the median and interquartile range.
问题1:使用原始数据,为’之后’成绩构建茎叶图,并求中位数和四分位距。
Question 2: Compare the ‘Before’ and ‘After’ distributions by referring to measures of central tendency and spread. What does this tell you about the revision programme?
问题2:参考集中趋势和离散程度量数,比较’之前’和’之后’的分布情况。这告诉你关于复习计划的什么信息?
Question 3: Explain one possible source of bias in this study and suggest how it could be reduced.
问题3:解释本研究中一个可能的偏差来源,并建议如何减少该偏差。
Question 4: A student claims, ‘Because the correlation coefficient is 0.97, the programme caused the improvement.’ Critically evaluate this claim.
问题4:一名学生声称,’因为相关系数为0.97,所以该计划导致了进步。’批判性地评价这一说法。
By working through such applied tasks, you develop the analytical skills essential for high marks in the IGCSE Statistics exam.
通过完成这类应用任务,你培养了在IGCSE统计学考试中取得高分所必需的分析技能。
Published by TutorHao | Statistics Revision Series | aleveler.com
Find Edexcel IGCSE Statistics Textbooks on eBay UK
New, used and second-hand copies of textbooks and revision guides are often much cheaper than retail — check current listings and prices before you buy.
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导