📚 GCSE CAIE Statistics: Case Study Practical Exercises | GCSE CAIE 统计:案例分析实战演练
In this article, we will work through a complete statistical investigation based on a real-world scenario. We will follow a case study on a school’s canteen satisfaction survey, tackling data collection, organisation, representation, analysis and interpretation. Each step is designed to reinforce the core skills required for the CAIE GCSE Statistics syllabus and to help you feel confident when faced with exam-style practical tasks.
本文将通过一个真实场景的完整统计调查,带领大家进行一次实战演练。我们将围绕一所学校食堂满意度调查的案例,逐步处理数据收集、整理、展示、分析与解读。每个环节都紧扣 CAIE GCSE 统计课程的核心技能,帮助你自信应对考试中的实践题型。
1. Understanding the Case Study Context | 理解案例背景
A survey was conducted at Greenfield Secondary School to evaluate student satisfaction with the new canteen menu. A random sample of 120 students from Years 7 to 11 was asked to rate their overall satisfaction on a scale from 1 (very dissatisfied) to 10 (very satisfied). Additional information collected included daily spending (£), number of canteen visits per week, gender and year group. The goal is to use statistical methods to summarise and analyse the data, and to draw conclusions about student opinion.
格林菲尔德中学进行了一项调查,以评估学生对新食堂菜单的满意度。从7至11年级随机抽取了120名学生,要求他们对总体满意度进行评分,分值从1分(非常不满意)到10分(非常满意)。调查还收集了每日消费金额(英镑)、每周光顾食堂次数、性别和年级等信息。目标是运用统计方法汇总和分析数据,并得出有关学生意见的结论。
2. Data Collection Methods | 数据收集方法
The school used a stratified sampling method, dividing the student population by year group and then selecting a proportional number of participants from each stratum. This ensured that each year group was fairly represented. Paper questionnaires were handed out during form time, and responses were anonymised to encourage honesty. The questionnaire included one quantitative continuous variable (daily spend), one quantitative discrete variable (visits per week), one rating scale variable (satisfaction) and two categorical variables (gender and year).
学校采用了分层抽样方法,先将学生按年级分层,再从各层中按比例选取参与者。这确保了每个年级都得到公平的代表。纸质问卷在晨会时间发放,问卷采用匿名方式,以鼓励学生真实作答。问卷中包含一个连续定量变量(每日消费金额)、一个离散定量变量(每周次数)、一个评分量表变量(满意度)以及两个分类变量(性别和年级)。
Before analysing, the raw data were coded and entered into a spreadsheet. Non-responses and outliers, such as a daily spend of £20 when most were under £5, were flagged for investigation. Data cleaning is a vital first step in any statistical investigation.
在分析之前,原始数据被编码并录入电子表格。无应答数据以及异常值(例如大多数每日消费金额低于5英镑,却出现20英镑)均被标记以待核查。在任何统计调查中,数据清理都是至关重要的第一步。
3. Organising Data and Frequency Tables | 数据整理与频数分布表
Satisfaction scores (out of 10) for the 120 students were grouped into intervals: 1-2, 3-4, 5-6, 7-8, 9-10. A grouped frequency table was constructed, which also included cumulative frequency to help find medians and quartiles later.
120名学生的满意度评分(满分10分)被分为若干区间:1-2, 3-4, 5-6, 7-8, 9-10。我们构建了分组频数表,并加入了累积频数,便于随后计算中位数和四分位数。
| Score Interval | Frequency (f) | Cumulative Frequency |
|---|---|---|
| 1-2 | 10 | 10 |
| 3-4 | 24 | 34 |
| 5-6 | 38 | 72 |
| 7-8 | 32 | 104 |
| 9-10 | 16 | 120 |
For daily spending, a stem-and-leaf diagram was used instead of a grouped table because it preserves exact values while showing the shape of the distribution.
每日消费金额则使用茎叶图代替分组表格,因为茎叶图既能保留原始数值,又能展示分布形态。
4. Displaying Data with Charts | 利用图表展示数据
A bar chart was created to display the frequency of each satisfaction interval. The horizontal axis showed score intervals and the vertical axis showed frequency. This visual immediately highlighted that the modal class was 5-6 and that the distribution was roughly symmetric.
我们绘制了条形图来展示各个满意度区间的频数。横轴为评分区间,纵轴为频数。该图表直观地显示出众数所在组为5-6,且分布大致对称。
A pie chart compared the responses by gender: 45% male, 52% female and 3% who preferred not to say. This gave a quick snapshot of the sample composition.
饼状图比较了不同性别的作答情况:男性占45%,女性占52%,不愿透露性别者占3%。这快速呈现了样本构成。
To examine the relationship between daily spend and satisfaction, a scatter graph was drawn. Each point represented a student, with daily spend on the x-axis and satisfaction on the y-axis. The points showed a weak positive trend, suggesting that students who spent more in the canteen tended to rate satisfaction slightly higher.
为探究每日消费金额与满意度之间的关系,我们绘制了散点图。每个点代表一名学生,x轴为每日消费金额,y轴为满意度。散点呈现微弱的正向趋势,表明在食堂消费较多的学生满意度评分略高。
5. Measures of Central Tendency | 集中趋势测度
Using the raw satisfaction scores, the mean was calculated as the sum of all scores divided by 120. The result was approximately 5.8, indicating that on average students were moderately satisfied.
利用原始满意度评分,计算平均值为所有评分之和除以120。结果约为5.8,说明学生平均满意度中等。
The median satisfaction score was found by identifying the 60th and 61st values in the ordered list; both fell in the 5-6 interval, giving a median of 5.5. The mode was the score 6, which occurred most frequently. Because the mean, median and mode were close, the distribution was fairly symmetrical.
将数据排序后,找到第60和第61个值确定中位数;两者均落在5-6区间,因此中位数为5.5。众数为6分,出现次数最多。平均值、中位数和众数接近,显示分布相当对称。
6. Measures of Dispersion | 离散度测度
The range of satisfaction scores was 10 – 1 = 9, but this is easily affected by extremes. A more robust measure, the interquartile range (IQR), was determined from the cumulative frequency graph: Q1 ≈ 4, Q3 ≈ 7.5, so IQR ≈ 3.5. This tells us the middle 50% of satisfaction scores fell within 3.5 points.
满意度评分的极差为 10 – 1 = 9,但易受极端值影响。更稳健的度量——四分位距 (IQR) 通过累积频数图确定:下四分位数 Q1 ≈ 4,上四分位数 Q3 ≈ 7.5,因此 IQR ≈ 3.5。这表明中间50%的满意度评分集中在3.5分的范围内。
The standard deviation was calculated for the daily spending data. Using the formula:
s = √[ Σ(x – x̄)² / (n-1) ]
where x̄ was £3.20, the standard deviation came to £1.15. This shows that individual daily spends typically varied by about £1.15 from the mean.
我们还计算了每日消费金额的标准差。运用公式:
s = √[ Σ(x – x̄)² / (n-1) ]
其中 x̄ = £3.20,得到标准差为 £1.15。这表明个人每日消费金额平均偏离均值约1.15英镑。
7. Probability Analysis | 概率分析
We can explore probability using the survey data. For instance, what is the probability that a randomly selected student gave a satisfaction rating of 7 or above? From the frequency table, 32 + 16 = 48 students scored 7-10, so P(high) = 48/120 = 0.4.
我们可以运用调查数据探索概率问题。例如,随机抽取一名学生,其满意度评分在7分及以上的概率是多少?根据频数表,评分7-10的有32+16=48人,因此 P(高满意度) = 48/120 = 0.4。
If two students are selected without replacement, the probability that both rated satisfaction 5 or less can be found using a tree diagram. The probability the first student is ‘low’ is (10+24+38)/120 = 72/120 = 0.6. After removing one low-rated student, the probability the second is also low becomes 71/119. Hence P(both low) = 0.6 × (71/119) ≈ 0.358.
若不放回抽取两名学生,两人均评分5分及以下的概率可用树状图求得。第一名学生为“低分”的概率为 (10+24+38)/120 = 72/120 = 0.6。移除一名低分学生后,第二名也是低分的概率变为 71/119。因此 P(两人均低分) = 0.6 × (71/119) ≈ 0.358。
Conditional probability can also be addressed: Given a student is female, what is the probability she is highly satisfied? If 30 out of 62 females rated 7+, then P(high|female) = 30/62 ≈ 0.484.
条件概率同样可以计算:已知某学生为女性,她属于高满意度的概率是多少?若62名女生中有30人评分7+,则 P(高满意度|女性) = 30/62 ≈ 0.484。
8. Correlation and Regression | 相关与回归
To examine the relationship between daily spend (x) and satisfaction (y), we calculated the product moment correlation coefficient (PMCC). The formula used was:
r = [nΣxy – (Σx)(Σy)] / √[ (nΣx² – (Σx)²)(nΣy² – (Σy)²) ]
For our data, r ≈ 0.32, indicating a weak positive correlation. This suggests that as spending increases, satisfaction tends to rise slightly, but the relationship is not strong.
为探究每日消费金额 (x) 与满意度 (y) 的关系,我们计算了积矩相关系数 (PMCC)。所用公式为:
r = [nΣxy – (Σx)(Σy)] / √[ (nΣx² – (Σx)²)(nΣy² – (Σy)²) ]
计算得 r ≈ 0.32,表明存在弱正相关。这意味着随着消费金额增加,满意度有轻微上升趋势,但关系并不显著。
The equation of the regression line was then determined: y = 4.2 + 0.5x. This can be used to predict satisfaction for a given daily spend. For example, if a student spends £4.50, predicted satisfaction = 4.2 + 0.5×4.5 = 6.45. However, because the correlation is weak, such predictions should be treated with caution.
随后计算出回归直线方程:y = 4.2 + 0.5x。该方程可用于根据每日消费金额预测满意度。例如,若某学生消费 £4.50,预测满意度 = 4.2 + 0.5×4.5 = 6.45。但由于相关性较弱,此类预测应谨慎对待。
9. Statistical Inference and Decision Making | 统计推断与决策
Based on the analysis, the school’s catering manager can infer that the overall satisfaction is moderate but not outstanding. The IQR of 3.5 suggests that while most students are clustered around scores 4 to 7.5, there is a notable minority who are very dissatisfied or very satisfied. The weak positive correlation with spend may indicate that students who use the canteen more are slightly more content, possibly because they are familiar with the menu or it meets their dietary needs better.
基于以上分析,学校餐饮经理可以推断,总体满意度处于中等水平,但并非十分突出。四分位距为3.5表明,虽然大多数学生的评分集中在4至7.5之间,但仍有少数学生非常不满意或非常满意。满意度与消费金额之间的弱正相关可能意味着,更常光顾食堂的学生满意度稍高,或许是因为他们对菜单更为熟悉,或食堂更能满足其饮食需求。
Recommendations might include gathering further data through focus groups to understand what the lowest-scoring students dislike, and trialling a different menu on a small scale before implementing changes. The data provide evidence to support decisions, but must be combined with contextual knowledge.
建议可能包括通过焦点小组收集更多数据,以了解评分最低的学生不满意的原因,并在全面调整菜单前进行小规模试点。数据为决策提供了证据,但必须结合背景知识进行判断。
10. Using Technology Tools | 使用技术工具
Throughout this case study, spreadsheets and statistical software played a key role. Functions like AVERAGE, MEDIAN, MODE, STDEV.P, and QUARTILE were used for quick calculations. A scatter plot with trendline feature generated the regression equation automatically. However, CAIE exams often require candidates to demonstrate manual calculation steps, so understanding the underlying processes remains essential.
在整个案例分析中,电子表格和统计软件发挥了关键作用。我们使用了 AVERAGE、MEDIAN、MODE、STDEV.P 和 QUARTILE 等函数进行快速计算。通过散点图及趋势线功能自动生成了回归方程。然而,CAIE 考试往往要求考生展示手动计算步骤,因此理解底层计算过程仍然至关重要。
When using technology, always check that data ranges are correct and that any diagrams produced are accurately labelled. A graph without proper titles and axis labels will lose marks even if the data are right.
在使用技术工具时,务必检查数据范围是否正确,并确保生成的图表有准确的标注。即使数据正确,缺少合适标题和轴标签的图表也会失分。
11. Common Mistakes and How to Avoid Them | 常见错误与避免方法
Many students confuse the range and interquartile range. Remember that the range uses only the extreme values, while the IQR focuses on the middle 50% and is less affected by outliers. When calculating the mean from grouped data, always use the midpoint of each interval. Using the interval boundaries directly is a frequent error.
许多学生容易混淆极差和四分位距。请记住,极差仅使用极端值,而四分位距着眼于中间50%的数据,受异常值影响较小。从分组数据计算平均数时,务必使用各组的组中值,直接使用区间边界是一个常见错误。
In probability problems with and without replacement, check whether the denominator changes. For scatter graphs, do not assume a correlation is strong just because the points seem to follow a vague direction. Always calculate r to confirm. Finally, always write a sentence interpreting your statistical findings in context; a bare numerical answer is rarely sufficient.
在处理有放回和无放回的概率问题时,要检查分母是否发生变化。对于散点图,不要因为点阵看似遵循某个方向就假设相关性强;务必计算 r 加以确认。最后,一定要结合背景写一句对统计结果的解释;仅给出数字答案往往是不够的。
12. Bringing It All Together: Writing a Statistical Report | 综合整理:撰写统计报告
A complete statistical report should include: an introduction stating the purpose and sampling method, a description of data collection, appropriate tables and graphs, calculations of central tendency and dispersion, any correlation or probability findings, and a conclusion with recommendations. By following a structured approach, you demonstrate the ability to handle a statistical investigation from start to finish — exactly what CAIE exam tasks require.
一份完整的统计报告应包括:说明调查目的和抽样方法的引言、数据收集的描述、恰当的表格和图表、集中趋势与离散度的计算、任何相关或概率的发现,以及提出建议的结论。遵循结构化的方法,你便展示了从头到尾完成一项统计调查的能力——这正是 CAIE 考试任务所要求的。
Practice with different scenarios, such as surveys on sports preferences, environmental awareness or revision hours versus test scores. The more you apply statistical thinking to real data, the more intuitive and manageable the exam will become.
尝试练习不同的情境,如体育偏好调查、环保意识调查,或复习时长与测验成绩的关系。你对真实数据运用统计思维的次数越多,考试就会变得越直观、越得心应手。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply