GCSE CIE Statistics: Case Study Practical Walkthrough | GCSE CIE 统计:案例分析实战演练

📚 GCSE CIE Statistics: Case Study Practical Walkthrough | GCSE CIE 统计:案例分析实战演练

In this article, we will work through a full statistical case study designed to mirror the style of CIE IGCSE Statistics Paper 2. You will see how to plan, collect, analyse and interpret data, using a realistic scenario based on a cafe owner investigating customer spending. Every stage includes paired English and Chinese explanations to reinforce key concepts and exam techniques.

本文将通过一个完整的统计案例演练,还原 CIE IGCSE 统计卷二的实战风格。案例围绕咖啡馆店主调查顾客消费情况展开,你将看到如何规划、收集、分析并解读数据。每个环节都配有中英对照讲解,帮助你巩固核心概念与考试技巧。

1. Scenario and Objectives | 情景与目标

A cafe owner believes that a recent modern refurbishment has increased the average amount customers spend per visit. Before the change, the long‑term mean spend was μ = ₤4.50 with a population standard deviation σ = ₤1.20. The owner wants to use statistical methods to test this claim, assess spending patterns by gender and age, and evaluate whether a newly introduced loyalty card influences drink choices. The investigation is designed as an IGCSE Statistics project, requiring a clear hypothesis, sampling method, data handling and formal conclusions.

一家咖啡馆的店主认为,近期现代风格的重新装修提高了顾客每次到店的平均消费额。装修前的长期人均消费为 μ = ₤4.50,总体标准差 σ = ₤1.20。店主希望用统计方法检验这一说法,分析性别与年龄对消费的影响,并评估新推出的积分卡是否影响饮品选择。本调查按 IGCSE 统计项目设计,要求提出明确假设、选择合适的抽样方法并完成数据处理与正式结论。


2. Planning Data Collection | 规划数据收集

We decide to collect a simple random sample of 50 recent transactions from the till system for the post‑refurbishment period. A random number generator is used to select transaction IDs, avoiding any human bias. We record each customer’s gender, age (to the nearest year), total spend (₤) and whether they purchased a speciality drink (latte, cappuccino or americano). Ethical considerations: all data are anonymised and no personal identities are stored.

我们决定从收银系统中随机抽取 50 笔近期装修后的交易记录,作为简单随机样本。使用随机数生成器选择交易 ID 以避免人为偏差。我们记录每位顾客的性别、年龄(精确到岁)、消费总额(英镑)以及是否购买了特色饮品(拿铁、卡布奇诺或美式)。道德考量:所有数据均匿名化处理,不储存任何个人身份信息。


3. Collecting and Recording Data | 收集与记录数据

After selecting 50 transactions, a structured table is prepared. The first five rows are shown below as an example:

ID Gender Age Spend(₤) Speciality
1 F 23 5.20 Yes
2 M 31 3.80 No
3 F 45 6.00 Yes
4 M 19 4.10 No
5 F 29 5.75 Yes

All 50 records are entered into a spreadsheet for further processing. Checks are performed to ensure there are no unrealistic values, such as a spend of ₤0 or an age of 150.

在选出 50 笔交易后,我们准备了一张结构化的表格。以下展示前五行示例:

编号 性别 年龄 花费(₤) 特色饮品
1 23 5.20
2 31 3.80
3 45 6.00
4 19 4.10
5 29 5.75

全部 50 条记录录入电子表格以待进一步处理。我们会检查是否存在不合理数值,比如花费为 ₤0 或年龄为 150 岁。


4. Organising and Visualising Data | 整理与可视化数据

For categorical variables such as ‘Speciality drink’, a bar chart is produced. Frequencies: Yes = 32, No = 18. For the continuous ‘Spend’ variable, a histogram with intervals ₤2.50–₤3.50, ₤3.50–₤4.50, ₤4.50–₤5.50, ₤5.50–₤6.50 shows the distribution is roughly symmetric, peaking around ₤4.50–₤5.50. A box‑and‑whisker plot is also constructed: minimum = ₤2.80, Q₁ = ₤3.90, median = ₤4.70, Q₃ = ₤5.40, maximum = ₤6.80.

对于“特色饮品”这类分类变量,我们绘制柱状图。频数:是 = 32,否 = 18。对于连续变量“消费额”,绘制直方图,组距为 ₤2.50–₤3.50、₤3.50–₤4.50、₤4.50–₤5.50、₤5.50–₤6.50,分布大致对称,峰值在 ₤4.50–₤5.50 附近。同时还构建了箱线图:最小值 = ₤2.80,Q₁ = ₤3.90,中位数 = ₤4.70,Q₃ = ₤5.40,最大值 = ₤6.80。


5. Descriptive Statistics: Central Tendency and Spread | 描述统计:集中趋势与离散度

From the 50 post‑refurbishment spends, the sample mean is x̄ = ₤4.80 and the sample standard deviation is s = ₤1.05. The median is ₤4.70, very close to the mean, suggesting little skew. The range is ₤4.00 and the interquartile range (IQR) is Q₃ − Q₁ = ₤1.50. These measures indicate that typical customer spending is slightly higher than the pre‑refurbishment mean of ₤4.50, but we need formal inference to decide if the difference is statistically significant.

从 50 笔装修后消费额计算得到样本均值 x̄ = ₤4.80,样本标准差 s = ₤1.05。中位数为 ₤4.70,与均值非常接近,说明偏态很小。全距为 ₤4.00,四分位距 IQR = Q₃ − Q₁ = ₤1.50。这些指标显示典型顾客消费略高于装修前的均值 ₤4.50,但我们需要通过正式推断来判断该差异是否具有统计显著性。


6. Probability and Combined Events | 概率与组合事件

The cafe offers three speciality drinks: Latte, Cappuccino and Americano. Among the 32 speciality‑drink customers, 14 chose Latte, 10 Cappuccino and 8 Americano. If we randomly select one speciality customer, the probability they chose Latte is P(L) = 14/32 = 0.4375, and Cappuccino is P(C) = 10/32 = 0.3125. Since a customer can only choose one drink, the events are mutually exclusive. The probability of choosing either Latte or Cappuccino is P(L ∪ C) = P(L) + P(C) = 24/32 = 0.75. These simple probabilities help the owner understand preference shares.

咖啡馆提供三种特色饮品:拿铁、卡布奇诺和美式。32 位购买特色饮品的顾客中,14 人选拿铁,10 人选卡布奇诺,8 人选美式。如果随机选取一位特色饮品顾客,其选择拿铁的概率为 P(L) = 14/32 = 0.4375,卡布奇诺的概率为 P(C) = 10/32 = 0.3125。由于顾客只能选择一种饮品,这些事件互斥。选择拿铁或卡布奇诺的概率为 P(L ∪ C) = P(L) + P(C) = 24/32 = 0.75。这些简单概率能帮助店主了解偏好占比。


7. The Normal Distribution and Standardisation | 正态分布与标准化

Assuming the pre‑refurbishment spend is normally distributed with μ = 4.50 and σ = 1.20, we can calculate the probability that a randomly chosen old‑style customer spent more than ₤5.50. First, compute the z‑score: z = (X − μ)/σ = (5.50 − 4.50)/1.20 = 0.8333. Using the standard normal table, P(Z > 0.833) = 1 − 0.7977 = 0.2023. So about 20% of customers spent over ₤5.50 before the change.

假设装修前消费额服从正态分布, μ = 4.50,σ = 1.20。我们可以计算随机选取一位原有顾客其消费超过 ₤5.50 的概率。首先计算 z 分数:z = (X − μ)/σ = (5.50 − 4.50)/1.20 = 0.8333。查标准正态表得 P(Z > 0.833) = 1 − 0.7977 = 0.2023。因此装修前约有 20% 的顾客花费超过 ₤5.50。


8. Confidence Intervals for the Mean | 均值的置信区间

We construct a 95% confidence interval for the true mean spend after refurbishment. Since n = 50 is large, we can use the z‑interval even if σ is estimated by s (₤1.05). The formula is: x̄ ± z × (σ / √n). Using the known pre‑refurbishment σ = 1.20 as a conservative estimate gives a slightly wider interval: 4.80 ± 1.96 × (1.20 / √50) = 4.80 ± 0.332, i.e. (₤4.47, ₤5.13). Using s = 1.05 yields 4.80 ± 1.96 × (1.05 / √50) = 4.80 ± 0.291, i.e. (₤4.51, ₤5.09). Both intervals lie partly above ₤4.50, hinting at a possible increase.

我们为装修后的真实平均消费构建 95% 置信区间。由于 n = 50 足够大,即使用 s = ₤1.05 估计 σ,也可以使用 z 区间。公式为:x̄ ± z × (σ / √n)。采用保守的装修前 σ = 1.20 会得到略宽的区间:4.80 ± 1.96 × (1.20 / √50) = 4.80 ± 0.332,即 (₤4.47, ₤5.13)。使用 s = 1.05 则得 4.80 ± 1.96 × (1.05 / √50) = 4.80 ± 0.291,即 (₤4.51, ₤5.09)。两个区间都部分高于 ₤4.50,暗示可能有所增加。


9. Hypothesis Testing for a Mean | 均值的假设检验

We now formally test whether the population mean has changed. Let μ be the mean spend after refurbishment. Hypotheses: H₀: μ = 4.50, H₁: μ ≠ 4.50 (two‑tailed). Significance level α = 0.05, so the critical z‑values are ±1.96. Using the known σ = 1.20, the test statistic is z = (x̄ − μ₀) / (σ/√n) = (4.80 − 4.50) / (1.20/√50) = 0.30 / 0.1697 ≈ 1.768. Since |1.768| < 1.96, we do not reject H₀. There is insufficient evidence at the 5% level to conclude that the refurbishment has changed the mean spend.

现在我们正式检验总体均值是否发生变化。设 μ 为装修后平均消费。假设为 H₀: μ = 4.50,H₁: μ ≠ 4.50(双尾检验)。显著性水平 α = 0.05,临界 z 值为 ±1.96。使用已知 σ = 1.20,检验统计量为 z = (x̄ − μ₀) / (σ/√n) = (4.80 − 4.50) / (1.20/√50) = 0.30 / 0.1697 ≈ 1.768。由于 |1.768| < 1.96,我们不拒绝 H₀。在 5% 显著性水平下,没有足够证据表明装修改变了平均消费额。


10. Bivariate Data: Scatter Diagrams, Correlation and Regression | 双变量数据:散点图、相关与回归

The cafe owner also suspects that older customers tend to spend more. We plot age (x) against spend (y) for the 50 customers. The scatter diagram shows a weak positive linear pattern. Pearson’s product‑moment correlation coefficient is calculated as r = 0.38. The least squares regression line is y = 3.20 + 0.045x, meaning each additional year of age is associated with an estimated extra ₤0.045 in spend. The coefficient of determination r² = 0.144 tells us that only about 14.4% of the variation in spend is explained by age.

店主还推测年长顾客往往花费更多。我们绘制 50 位顾客年龄 (x) 对消费额 (y) 的散点图。图形显示微弱的正向线性趋势。积差相关系数经计算为 r = 0.38。最小二乘回归直线为 y = 3.20 + 0.045x,即年龄每增加一岁,消费额估计增加 ₤0.045。决定系数 r² = 0.144 表明消费额变异中仅有约 14.4% 可由年龄解释。


11. Chi‑Squared Test for Association | 卡方独立性检验

We now investigate whether gender and speciality‑drink choice are associated. Observed frequencies:

Latte Cappuccino Americano Total
Male 6 7 5 18
Female 8 3 3 14
Total 14 10 8 32

Expected frequencies are calculated as (row total × column total)/grand total. The χ² statistic is found to be 2.07 with (2−1)×(3−1) = 2 degrees of freedom. The critical value at 5% level is 5.991. Since 2.07 < 5.991, we do not reject H₀: there is no significant association between gender and drink choice.

现在我们考察性别与特色饮品选择是否有关联。观测频数表:

拿铁 卡布奇诺 更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading