📚 GCSE CIE Statistics: Case Study Practical Walkthrough | GCSE CIE 统计:案例分析实战演练
In this article, we will work through a full statistical case study designed to mirror the style of CIE IGCSE Statistics Paper 2. You will see how to plan, collect, analyse and interpret data, using a realistic scenario based on a cafe owner investigating customer spending. Every stage includes paired English and Chinese explanations to reinforce key concepts and exam techniques.
本文将通过一个完整的统计案例演练,还原 CIE IGCSE 统计卷二的实战风格。案例围绕咖啡馆店主调查顾客消费情况展开,你将看到如何规划、收集、分析并解读数据。每个环节都配有中英对照讲解,帮助你巩固核心概念与考试技巧。
1. Scenario and Objectives | 情景与目标
A cafe owner believes that a recent modern refurbishment has increased the average amount customers spend per visit. Before the change, the long‑term mean spend was μ = ₤4.50 with a population standard deviation σ = ₤1.20. The owner wants to use statistical methods to test this claim, assess spending patterns by gender and age, and evaluate whether a newly introduced loyalty card influences drink choices. The investigation is designed as an IGCSE Statistics project, requiring a clear hypothesis, sampling method, data handling and formal conclusions.
一家咖啡馆的店主认为,近期现代风格的重新装修提高了顾客每次到店的平均消费额。装修前的长期人均消费为 μ = ₤4.50,总体标准差 σ = ₤1.20。店主希望用统计方法检验这一说法,分析性别与年龄对消费的影响,并评估新推出的积分卡是否影响饮品选择。本调查按 IGCSE 统计项目设计,要求提出明确假设、选择合适的抽样方法并完成数据处理与正式结论。
2. Planning Data Collection | 规划数据收集
We decide to collect a simple random sample of 50 recent transactions from the till system for the post‑refurbishment period. A random number generator is used to select transaction IDs, avoiding any human bias. We record each customer’s gender, age (to the nearest year), total spend (₤) and whether they purchased a speciality drink (latte, cappuccino or americano). Ethical considerations: all data are anonymised and no personal identities are stored.
我们决定从收银系统中随机抽取 50 笔近期装修后的交易记录,作为简单随机样本。使用随机数生成器选择交易 ID 以避免人为偏差。我们记录每位顾客的性别、年龄(精确到岁)、消费总额(英镑)以及是否购买了特色饮品(拿铁、卡布奇诺或美式)。道德考量:所有数据均匿名化处理,不储存任何个人身份信息。
3. Collecting and Recording Data | 收集与记录数据
After selecting 50 transactions, a structured table is prepared. The first five rows are shown below as an example:
| ID | Gender | Age | Spend(₤) | Speciality |
| 1 | F | 23 | 5.20 | Yes |
| 2 | M | 31 | 3.80 | No |
| 3 | F | 45 | 6.00 | Yes |
| 4 | M | 19 | 4.10 | No |
| 5 | F | 29 | 5.75 | Yes |
All 50 records are entered into a spreadsheet for further processing. Checks are performed to ensure there are no unrealistic values, such as a spend of ₤0 or an age of 150.
在选出 50 笔交易后,我们准备了一张结构化的表格。以下展示前五行示例:
| 编号 | 性别 | 年龄 | 花费(₤) | 特色饮品 |
| 1 | 女 | 23 | 5.20 | 是 |
| 2 | 男 | 31 | 3.80 | 否 |
| 3 | 女 | 45 | 6.00 | 是 |
| 4 | 男 | 19 | 4.10 | 否 |
| 5 | 女 | 29 | 5.75 | 是 |
全部 50 条记录录入电子表格以待进一步处理。我们会检查是否存在不合理数值,比如花费为 ₤0 或年龄为 150 岁。
4. Organising and Visualising Data | 整理与可视化数据
For categorical variables such as ‘Speciality drink’, a bar chart is produced. Frequencies: Yes = 32, No = 18. For the continuous ‘Spend’ variable, a histogram with intervals ₤2.50–₤3.50, ₤3.50–₤4.50, ₤4.50–₤5.50, ₤5.50–₤6.50 shows the distribution is roughly symmetric, peaking around ₤4.50–₤5.50. A box‑and‑whisker plot is also constructed: minimum = ₤2.80, Q₁ = ₤3.90, median = ₤4.70, Q₃ = ₤5.40, maximum = ₤6.80.
对于“特色饮品”这类分类变量,我们绘制柱状图。频数:是 = 32,否 = 18。对于连续变量“消费额”,绘制直方图,组距为 ₤2.50–₤3.50、₤3.50–₤4.50、₤4.50–₤5.50、₤5.50–₤6.50,分布大致对称,峰值在 ₤4.50–₤5.50 附近。同时还构建了箱线图:最小值 = ₤2.80,Q₁ = ₤3.90,中位数 = ₤4.70,Q₃ = ₤5.40,最大值 = ₤6.80。
5. Descriptive Statistics: Central Tendency and Spread | 描述统计:集中趋势与离散度
From the 50 post‑refurbishment spends, the sample mean is x̄ = ₤4.80 and the sample standard deviation is s = ₤1.05. The median is ₤4.70, very close to the mean, suggesting little skew. The range is ₤4.00 and the interquartile range (IQR) is Q₃ − Q₁ = ₤1.50. These measures indicate that typical customer spending is slightly higher than the pre‑refurbishment mean of ₤4.50, but we need formal inference to decide if the difference is statistically significant.
从 50 笔装修后消费额计算得到样本均值 x̄ = ₤4.80,样本标准差 s = ₤1.05。中位数为 ₤4.70,与均值非常接近,说明偏态很小。全距为 ₤4.00,四分位距 IQR = Q₃ − Q₁ = ₤1.50。这些指标显示典型顾客消费略高于装修前的均值 ₤4.50,但我们需要通过正式推断来判断该差异是否具有统计显著性。
6. Probability and Combined Events | 概率与组合事件
The cafe offers three speciality drinks: Latte, Cappuccino and Americano. Among the 32 speciality‑drink customers, 14 chose Latte, 10 Cappuccino and 8 Americano. If we randomly select one speciality customer, the probability they chose Latte is P(L) = 14/32 = 0.4375, and Cappuccino is P(C) = 10/32 = 0.3125. Since a customer can only choose one drink, the events are mutually exclusive. The probability of choosing either Latte or Cappuccino is P(L ∪ C) = P(L) + P(C) = 24/32 = 0.75. These simple probabilities help the owner understand preference shares.
咖啡馆提供三种特色饮品:拿铁、卡布奇诺和美式。32 位购买特色饮品的顾客中,14 人选拿铁,10 人选卡布奇诺,8 人选美式。如果随机选取一位特色饮品顾客,其选择拿铁的概率为 P(L) = 14/32 = 0.4375,卡布奇诺的概率为 P(C) = 10/32 = 0.3125。由于顾客只能选择一种饮品,这些事件互斥。选择拿铁或卡布奇诺的概率为 P(L ∪ C) = P(L) + P(C) = 24/32 = 0.75。这些简单概率能帮助店主了解偏好占比。
7. The Normal Distribution and Standardisation | 正态分布与标准化
Assuming the pre‑refurbishment spend is normally distributed with μ = 4.50 and σ = 1.20, we can calculate the probability that a randomly chosen old‑style customer spent more than ₤5.50. First, compute the z‑score: z = (X − μ)/σ = (5.50 − 4.50)/1.20 = 0.8333. Using the standard normal table, P(Z > 0.833) = 1 − 0.7977 = 0.2023. So about 20% of customers spent over ₤5.50 before the change.
假设装修前消费额服从正态分布, μ = 4.50,σ = 1.20。我们可以计算随机选取一位原有顾客其消费超过 ₤5.50 的概率。首先计算 z 分数:z = (X − μ)/σ = (5.50 − 4.50)/1.20 = 0.8333。查标准正态表得 P(Z > 0.833) = 1 − 0.7977 = 0.2023。因此装修前约有 20% 的顾客花费超过 ₤5.50。
8. Confidence Intervals for the Mean | 均值的置信区间
We construct a 95% confidence interval for the true mean spend after refurbishment. Since n = 50 is large, we can use the z‑interval even if σ is estimated by s (₤1.05). The formula is: x̄ ± z × (σ / √n). Using the known pre‑refurbishment σ = 1.20 as a conservative estimate gives a slightly wider interval: 4.80 ± 1.96 × (1.20 / √50) = 4.80 ± 0.332, i.e. (₤4.47, ₤5.13). Using s = 1.05 yields 4.80 ± 1.96 × (1.05 / √50) = 4.80 ± 0.291, i.e. (₤4.51, ₤5.09). Both intervals lie partly above ₤4.50, hinting at a possible increase.
我们为装修后的真实平均消费构建 95% 置信区间。由于 n = 50 足够大,即使用 s = ₤1.05 估计 σ,也可以使用 z 区间。公式为:x̄ ± z × (σ / √n)。采用保守的装修前 σ = 1.20 会得到略宽的区间:4.80 ± 1.96 × (1.20 / √50) = 4.80 ± 0.332,即 (₤4.47, ₤5.13)。使用 s = 1.05 则得 4.80 ± 1.96 × (1.05 / √50) = 4.80 ± 0.291,即 (₤4.51, ₤5.09)。两个区间都部分高于 ₤4.50,暗示可能有所增加。
9. Hypothesis Testing for a Mean | 均值的假设检验
We now formally test whether the population mean has changed. Let μ be the mean spend after refurbishment. Hypotheses: H₀: μ = 4.50, H₁: μ ≠ 4.50 (two‑tailed). Significance level α = 0.05, so the critical z‑values are ±1.96. Using the known σ = 1.20, the test statistic is z = (x̄ − μ₀) / (σ/√n) = (4.80 − 4.50) / (1.20/√50) = 0.30 / 0.1697 ≈ 1.768. Since |1.768| < 1.96, we do not reject H₀. There is insufficient evidence at the 5% level to conclude that the refurbishment has changed the mean spend.
现在我们正式检验总体均值是否发生变化。设 μ 为装修后平均消费。假设为 H₀: μ = 4.50,H₁: μ ≠ 4.50(双尾检验)。显著性水平 α = 0.05,临界 z 值为 ±1.96。使用已知 σ = 1.20,检验统计量为 z = (x̄ − μ₀) / (σ/√n) = (4.80 − 4.50) / (1.20/√50) = 0.30 / 0.1697 ≈ 1.768。由于 |1.768| < 1.96,我们不拒绝 H₀。在 5% 显著性水平下,没有足够证据表明装修改变了平均消费额。
10. Bivariate Data: Scatter Diagrams, Correlation and Regression | 双变量数据:散点图、相关与回归
The cafe owner also suspects that older customers tend to spend more. We plot age (x) against spend (y) for the 50 customers. The scatter diagram shows a weak positive linear pattern. Pearson’s product‑moment correlation coefficient is calculated as r = 0.38. The least squares regression line is y = 3.20 + 0.045x, meaning each additional year of age is associated with an estimated extra ₤0.045 in spend. The coefficient of determination r² = 0.144 tells us that only about 14.4% of the variation in spend is explained by age.
店主还推测年长顾客往往花费更多。我们绘制 50 位顾客年龄 (x) 对消费额 (y) 的散点图。图形显示微弱的正向线性趋势。积差相关系数经计算为 r = 0.38。最小二乘回归直线为 y = 3.20 + 0.045x,即年龄每增加一岁,消费额估计增加 ₤0.045。决定系数 r² = 0.144 表明消费额变异中仅有约 14.4% 可由年龄解释。
11. Chi‑Squared Test for Association | 卡方独立性检验
We now investigate whether gender and speciality‑drink choice are associated. Observed frequencies:
| Latte | Cappuccino | Americano | Total | |
| Male | 6 | 7 | 5 | 18 |
| Female | 8 | 3 | 3 | 14 |
| Total | 14 | 10 | 8 | 32 |
Expected frequencies are calculated as (row total × column total)/grand total. The χ² statistic is found to be 2.07 with (2−1)×(3−1) = 2 degrees of freedom. The critical value at 5% level is 5.991. Since 2.07 < 5.991, we do not reject H₀: there is no significant association between gender and drink choice.
现在我们考察性别与特色饮品选择是否有关联。观测频数表:
| 拿铁 | 卡布奇诺 | 更多咨询请联系16621398022(同微信)
CommentsMore posts |
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply