📚 IGCSE CCEA Statistics: Case Study Practice | IGCSE CCEA 统计:案例分析实战演练
In this revision guide, we work through a complete statistical case study based on a survey of 50 students at a secondary school. The data set includes categorical and numerical variables, which allows us to practise data collection methods, frequency tables, graphical representation, measures of central tendency and spread, cumulative frequency, scatter graphs, basic correlation, and probability. The step‑by‑step approach mirrors the style of CCEA IGCSE Statistics exam questions and helps you build confidence in applying statistical techniques to real data.
在这份复习指南中,我们将完成一个基于对某中学50名学生调查的完整统计案例研究。数据集包含分类变量和数值变量,让我们能够练习数据收集方法、频数表、图表表示、集中趋势和离散程度的度量、累积频率图、散点图、基本相关性和概率等内容。这种分步骤的方法与 CCEA IGCSE 统计考试试题的风格一致,有助于你将统计方法自信地应用到真实数据中。
1. Stating the Case Study Scenario | 案例情景说明
A statistics class surveyed 50 pupils in Year 11. They recorded each student’s age (15 or 16), height (cm), percentage score in a recent mathematics test, favourite sport (football, basketball, tennis, other) and the number of hours of physical exercise per week. The aim was to investigate any relationships between physical activity and test performance, as well as to describe the typical profile of a Year 11 student.
一个统计班对50名11年级学生进行了调查。他们记录了每名学生的年龄(15或16岁)、身高(厘米)、最近一次数学测验的百分制得分、最喜欢的运动(足球、篮球、网球、其它)以及每周体育锻炼的小时数。目的是探究体育活动与考试成绩之间的关系,并描述11年级学生的典型特征。
Having clear research questions before collecting data prevents ‘data dredging’. Two specific questions were set: ‘Is there an association between hours of exercise and maths score?’ and ‘What is the distribution of favourite sports among Year 11 students?’
在收集数据之前明确研究问题可以防止“数据挖掘”。设定了两个具体问题:“锻炼时间与数学成绩之间是否存在关联?”以及“11年级学生最喜爱的运动分布是怎样的?”
2. Data Collection and Sampling Methods | 数据收集与抽样方法
The class decided to use a stratified random sample of 50 students. Strata were based on tutor groups to ensure each group was represented in proportion to its size. This reduces sampling bias compared with convenience sampling, where the researcher might just ask friends.
该班级决定采用分层随机抽样,抽取50名学生。分层以辅导组为基础,确保每个辅导组按其人数比例被抽取。与可能只询问朋友的便利抽样相比,分层抽样可以减少抽样偏差。
A simple random sample within each tutor group was obtained by assigning numbers to students and using a random number generator. This gave every student in the group an equal chance of being selected. The data collection instrument was a short questionnaire that recorded age, height (self‑reported, but measured by researchers for accuracy), maths test score (from school records with permission), favourite sport from a list of options, and the usual weekly exercise hours.
在各个辅导组内部,通过给每个学生编号并使用随机数生成器得到了简单随机样本。这使组内每名学生都有均等机会被选中。数据收集工具是一份简短的问卷,记录了年龄、身高(学生自报,但研究人员为确保准确性进行了测量)、数学测验成绩(经许可从学校记录中获取)、从选项清单中选择的最喜爱运动,以及通常的每周锻炼小时数。
3. Organising Data: Frequency Tables | 数据整理:频数表
First we organise the categorical variable ‘favourite sport’. The raw responses are summarised in a frequency table.
我们首先整理分类变量“最喜爱的运动”。原始回答汇总在一张频数表中。
| Favourite sport | Frequency |
|---|---|
| Football | 14 |
| Basketball | 12 |
| Tennis | 10 |
| Other | 14 |
For continuous data like height, we group the values into class intervals. The raw heights ranged from 142 cm to 188 cm, so we use intervals of 10 cm starting at 140‑149.
对于身高这样的连续数据,我们把数值分组为组距。原始身高范围从 142 cm 到 188 cm,因此我们使用从 140–149 开始的 10 cm 组距。
| Height (cm) | Frequency |
|---|---|
| 140 – 149 | 3 |
| 150 – 159 | 8 |
| 160 – 169 | 17 |
| 170 – 179 | 15 |
| 180 – 189 | 7 |
Grouping sacrifices some detail but reveals the shape of the distribution, which is roughly symmetric with a peak in the 160–169 cm class.
分组会牺牲一些细节,但能揭示分布的形状,该分布大致对称,峰值在 160–169 cm 组。
4. Data Visualisation: Bar Chart and Pie Chart | 数据可视化:条形图与饼图
A bar chart is appropriate for the sport frequency data because the categories are separate. The height of each bar represents the frequency. We include a title, labelled axes, and equal bar widths to avoid misleading the reader.
条形图适用于运动频数数据,因为各个类别是分开的。每根条的高度代表频数。我们加上标题、标记坐标轴并使用等宽的条形,以免误导读者。
For the same categorical data, a pie chart shows proportions of the whole. The angle for football is (14/50)×360°=100.8°. Similarly, basketball gets 86.4°, tennis 72°, and other 100.8°. Always check that angles sum to 360°.
对于相同的分类数据,饼图展示各部分在整体中所占的比例。足球的角度为 (14/50)×360°=100.8°。同样地,篮球为 86.4°,网球为 72°,其它为 100.8°。务必检查角度之和为 360°。
We use the bar chart to compare the frequencies directly, and the pie chart to emphasise that football and ‘other’ each account for about a quarter of the students. CCEA exam questions often ask you to interpret the same data presented in two different graphs.
我们用条形图直接比较频数,而用饼图强调足球和“其它”各占学生总数的大约四分之一。CCEA 考题常要求你解释用两种不同图形表示的同一组数据。
5. Central Tendency: Mean, Median and Mode | 集中趋势:平均数、中位数和众数
For the mathematics test scores, the raw data (sorted) are: 42, 45, 48, 50, 52, 55, 56, 58, 60, 61, … up to 98. The sum of all 50 scores is 3450, so the mean = 3450÷50 = 69.0%. The mean is widely used but can be influenced by extreme values.
数学测验的原始成绩(排序后)为:42, 45, 48, 50, 52, 55, 56, 58, 60, 61, … 直至 98。50 个成绩的总和为 3450,因此平均数 = 3450÷50 = 69.0%。平均数被广泛使用,但可能受极端值影响。
To find the median we locate the (50+1)/2 = 25.5th value. The 25th score is 68 and the 26th is 70, so the median is (68+70)/2 = 69%. The median is robust to outliers. The mode in this data set is 72, which occurs three times; however, with continuous-like scores the mode is less informative.
为找中位数,我们定位第 (50+1)/2 = 25.5 个值。第 25 个成绩是 68,第 26 个是 70,因此中位数为 (68+70)/2 = 69%。中位数不易受异常值影响。数据中众数是 72,出现了三次;但由于成绩类似连续变量,众数提供的信息较少。
For grouped height data, we estimate the mean using midpoints. The midpoint for 140–149 is 144.5. Estimated mean = (3×144.5 + 8×154.5 + 17×164.5 + 15×174.5 + 7×184.5) ÷ 50 = 167.1 cm. The modal class is 160–169 cm with frequency 17.
对于分组的身高数据,我们使用组中值估计平均数。140–149 的组中值为 144.5。估计平均数 = (3×144.5 + 8×154.5 + 17×164.5 + 15×174.5 + 7×184.5) ÷ 50 = 167.1 cm。众数组是 160–169 cm,频数为 17。
6. Measures of Dispersion: Range, Quartiles and Interquartile Range | 离散程度:极差、四分位数和四分位距
The range of the test scores is 98 – 42 = 56, but this only uses the two extreme values. The interquartile range (IQR) gives a better picture of spread for the middle 50% of the data.
测验成绩的极差为 98 – 42 = 56,但这只用到了两个极端值。四分位距 (IQR) 可以更好地反映中间 50% 数据的分布情况。
To find quartiles for the 50 scores, Q₁ is the median of the lower half (the first 25 values): the 13th value is 56. Q₃ is the median of the upper half: the 38th value is 80. Therefore IQR = 80 – 56 = 24. This tells us that the middle half of the scores are spread over 24 percentage points.
要对 50 个成绩求四分位数,Q₁ 是下半部分(前 25 个值)的中位数:第 13 个值为 56。Q₃ 是上半部分的中位数:第 38 个值为 80。因此 IQR = 80 – 56 = 24。这说明中间一半的成绩散布在 24 个百分点之内。
For the grouped height data, we use cumulative frequencies to estimate quartiles, which is covered in the next section.
对于分组的身高数据,我们使用累积频数来估计四分位数,这将在下一节介绍。
7. Box‑and‑Whisker Plots | 箱线图
A box plot (box‑and‑whisker diagram) displays the minimum, Q₁, median, Q₃ and maximum. It is excellent for comparing distributions. For the test scores, the five‑number summary is: minimum 42, Q₁=56, median=69, Q₃=80, maximum 98.
箱线图(盒须图)展示了最小值、Q₁、中位数、Q₃ 和最大值。它非常适合比较分布。测验成绩的五数综合是:最小值 42,Q₁=56,中位数=69,Q₃=80,最大值 98。
The box is drawn from Q₁ to Q₃ with a line at the median. Whiskers extend to the minimum and maximum, provided there are no outliers. An outlier can be defined as a value more than 1.5×IQR below Q₁ or above Q₃. Here, lower fence = 56 – 1.5×24 = 20, upper fence = 80 + 36 = 116, so no outliers exist.
箱子从 Q₁ 画到 Q₃,并在中位数处画一条线。须线延伸到最小值和最大值,前提是没有异常值。异常值可定义为低于 Q₁ – 1.5×IQR 或高于 Q₃ + 1.5×IQR 的值。此处下限 = 56 – 1.5×24 = 20,上限 = 80 + 36 = 116,因此没有异常值。
If we also plot the box plot for another subject’s scores, we could compare the medians and the spreads. A longer box or longer whiskers indicate greater variability.
如果我们同时画出另一门学科成绩的箱线图,就可以比较中位数和分布情况。箱子或须线越长,表示变异性越大。
8. Cumulative Frequency and the Cumulative Frequency Curve | 累积频数与累积频数曲线
We construct a cumulative frequency table for the grouped heights by adding frequencies downwards.
我们通过向下累加频数,为分组身高构建累积频数表。
| Height (cm) | Upper class boundary | Frequency | Cumulative frequency |
|---|---|---|---|
| 140 – 149 | 149.5 | 3 | 3 |
| 150 – 159 | 159.5 | 8 | 11 |
| 160 – 169 | 169.5 | 17 | 28 |
| 170 – 179 | 179.5 | 15 | 43 |
| 180 – 189 | 189.5 | 7 | 50 |
Plot the points (149.5, 3), (159.5, 11), (169.5, 28), (179.5, 43), (189.5, 50) on a graph with upper class boundaries on the horizontal axis and cumulative frequency on the vertical. Join the points with a smooth curve, starting at (139.5, 0).
在图上标出点 (149.5, 3)、(159.5, 11)、(169.5, 28)、(179.5, 43)、(189.5, 50),横轴为上限组界,纵轴为累积频数。用一条平滑曲线连接各点,并从 (139.5, 0) 开始。
From the curve, the median corresponds to a cumulative frequency of 25. Drawing a horizontal line at 25 and dropping to the axis gives an estimated median height of about 166 cm. Q₁ (cumulative frequency 12.5) gives ≈ 160 cm, and Q₃ (37.5) gives ≈ 175 cm, leading to an IQR of roughly 15 cm.
从曲线上,中位数对应累积频数 25。在 25 处画水平线并向下作垂线,得到身高中位数的估计值约为 166 cm。Q₁(累积频数 12.5)约为 160 cm,Q₃(37.5)约为 175 cm,由此得到 IQR 大约为 15 cm。
9. Scatter Graphs and Correlation | 散点图与相关性
To investigate the relationship between hours of exercise and maths test scores, we plot a scatter graph with exercise hours on the x‑axis and test score on the y‑axis. Each of the 50 students contributes one point.
为了研究锻炼小时数与数学测验成绩之间的关系,我们绘制散点图,以锻炼小时数为 x 轴,测验成绩为 y 轴。50 名学生每人都贡献一个点。
The plotted points show a slight upward trend: students who do more exercise tend to have slightly higher maths scores, but the relationship is weak. There is considerable scatter, and some students with low exercise hours still achieve high scores. The correlation appears positive but not strong.
描出的点显示略微上升的趋势:锻炼较多的学生往往数学成绩稍高,但关系较弱。点的分散程度较大,一些锻炼时间少的学生仍然获得高分。相关性似乎为正向但不强。
We should not conclude causation: an association does not mean that more exercise causes higher test scores. A third factor, such as overall motivation, might influence both variables. When interpreting scatter graphs, look for clusters, outliers, and possible non‑linear patterns.
我们不应该得出因果结论:关联并不意味着锻炼越多会导致考试成绩越高。可能第三个因素,如整体积极性,对两个变量都有影响。解读散点图时,要寻找聚集、异常值和可能的非线性模式。
10. Spearman’s Rank Correlation Coefficient | 斯皮尔曼等级相关系数
CCEA IGCSE Statistics may require calculation of Spearman’s rank correlation coefficient (rₛ). We rank the 50 students for both exercise hours and maths scores, giving rank 1 to the highest value. Tied ranks are handled by assigning the average rank.
CCEA IGCSE 统计可能要求计算斯皮尔曼等级相关系数 (rₛ)。我们对 50 名学生分别按锻炼小时数和数学成绩排列名次,最高值排第 1。对于并列值,赋予平均名次。
The formula is rₛ = 1 – (6Σd²) / [n(n² – 1)], where d is the difference between the two ranks for each student. After calculating the differences, squaring them, and summing, suppose Σd² = 8520. Then rₛ = 1 – (6×8520) / [50×(2500 – 1)] = 1 – (51120 / 124950) ≈ 1 – 0.409 = 0.591.
公式为 rₛ = 1 – (6Σd²) / [n(n² – 1)],其中 d 是每名学生的两个名次之差。计算出差值、平方并求和,假设 Σd² = 8520。那么 rₛ = 1 – (6×8520) / [50×(2500 – 1)] = 1 – (51120 / 124950) ≈ 1 – 0.409 = 0.591。
This value indicates a moderate positive correlation. For n=50, a critical value from tables might be around 0.279 at the 5% significance level, so the correlation is statistically significant, meaning we reject the null hypothesis of no association. However, the result still needs to be interpreted in context.
该值表明存在中度正向相关。对 n=50,5% 显著性水平下的临界值约在 0.279 左右,因此该相关在统计上是显著的,这意味着我们拒绝无关联的零假设。但这一结果仍需结合具体情境加以解释。
11. Probability from the Data | 基于数据的概率
We can estimate probabilities using relative frequencies from the sample. For instance, if we randomly select one student, P(favourite sport is football) = 14/50 = 0.28. P(height ≥ 170 cm) can be estimated by adding the frequencies of the 170–179 and 180–189 classes, giving (15+7)/50 = 22/50 = 0.44.
我们可以使用样本中的相对频数来估计概率。例如,若随机抽取一名学生,P(最喜爱的运动为足球)= 14/50 = 0.28。P(身高 ≥ 170 cm)可以通过把 170–179 和 180–189 组的频数相加来估计,得到 (15+7)/50 = 22/50 = 0.44。
If we select two students without replacement, the probability that both like tennis is (10/50) × (9/49) = 90/2450 ≈ 0.0367. Such calculations assume random selection and are fundamental to understanding how samples reflect the population.
如果不放回地选取两名学生,两人都喜欢网球的概率为 (10/50) × (9/49) = 90/2450 ≈ 0.0367。这类计算假定随机选取,对于理解样本如何反映总体至关重要。
We can also look at combined events. P(student exercises more than 5 hours and scores above 70%) is found by counting students satisfying both conditions, say 12 students, so frequency leads to 12/50 = 0.24. This conditional reasoning underpins more advanced topics.
我们还可以考察组合事件。P(学生锻炼超过 5 小时且得分高于 70%)可通过统计同时满足两个条件的学生人数得到,比如 12 人,那么按频数得出的概率为 12/50 = 0.24。这种条件推理是更深入专题的基础。
12. Drawing Conclusions and Evaluating the Study | 得出结论与评估研究
From our analysis we learn: the typical Year 11 student in this sample has a mean maths score of 69% with IQR 24%; the most common favourite sports are football and other sports (each 28%); the median height is about 166 cm; and a moderate positive rank correlation exists between exercise and test scores.
通过分析我们了解到:此样本中典型的 11 年级学生数学平均成绩为 69%,IQR 为 24%;最受喜爱的运动是足球和其它运动(各占 28%);身高约在 166 cm;锻炼与考试成绩之间存在中等的正向等级相关。
We must evaluate the reliability of our findings. The sample size of 50 is reasonable but still subject to sampling error. Self‑reported exercise hours might be inaccurate. The correlation only describes this particular school; we cannot generalise to all schools without careful thought. A randomised experiment would be needed to test causation.
我们必须评估研究结果的可靠性。50 人的样本量尚可,但仍存在抽样误差。学生自报的锻炼小时数可能不准确。该相关性仅描述这所特定的学校;若不仔细思考,我们不能将其推广至所有学校。要检验因果关系,需要随机对照实验。
Reflecting on the process, we see that every step—from planning the data collection to interpreting probabilities—requires clear statistical thinking. CCEA examiners reward candidates who can critically evaluate data, recognise limitations, and communicate findings clearly using appropriate statistical language.
回顾整个过程,我们看到从规划数据收集到解读概率的每一步都需要清晰的统计思维。CCEA 考官会奖励那些能够批判性地评估数据、认识到局限性并用恰当的统计语言清晰表述研究结果的考生。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply