📚 Case Study Walkthrough: Applying Year 12 Statistics | AQA 统计案例分析实战演练
Welcome to a practical deep-dive into Year 12 AQA Statistics. In this walkthrough, we will follow a single case study from start to finish, applying the core techniques you need to master for your AS‑Level exams. By seeing how concepts such as descriptive statistics, probability distributions, and estimation work together in a real‑world scenario, you will build confidence in tackling your own exam papers.
欢迎来到 Year 12 AQA 统计的实战演练。本文将通过一个完整的案例分析,引导你从头到尾应用 AS 阶段必须掌握的核心统计方法。当你看到描述性统计、概率分布和区间估计等概念如何在一个真实场景中协同工作时,你应对考试的自信心将得到极大的提升。
1. Setting the Scene and Defining the Question | 背景与问题定义
Imagine a popular coffee shop chain, “Coffee Time”, is reviewing its sales performance. The manager wants to answer two main questions: first, what is the typical amount a customer spends per visit, and how much does it vary? Second, what is the probability that a customer purchasing a coffee also buys a slice of cake? The company’s marketing team believes that 40 % of coffee buyers add a cake, but the manager suspects the true proportion is lower. We have been asked to analyse a random sample of 50 customers observed on a single morning.
设想一家热门连锁咖啡店 “Coffee Time” 正在审视其销售表现。店经理想解答两个主要问题:第一,顾客每次光顾的典型消费额是多少,差异有多大?第二,购买咖啡的顾客同时购买蛋糕的概率是多少?市场部认为 40 % 的咖啡购买者会加购蛋糕,但经理怀疑实际比例更低。我们需要分析在一个上午随机观察到的 50 名顾客样本。
For each customer, two variables were recorded: total spend in GBP (to the nearest pound) and whether a cake was purchased (Yes/No). This gives us a mixture of continuous and categorical data to work with. The first few rows of our dataset look like this:
我们记录了每位顾客的两个变量:总消费额(英镑,精确到整数)以及是否购买了蛋糕(是/否)。这样我们就有了连续数据和分类数据的混合数据集。数据表的前几行如下:
| Customer | Spend (£) | Cake (Y/N) |
|---|---|---|
| 1 | 5 | Y |
| 2 | 4 | N |
| 3 | 7 | Y |
| … | … | … |
2. Data Collection Methods and Sampling | 数据收集方法与抽样
The manager collected data on a Monday morning between 8 am and 10 am using systematic sampling: every 3rd customer entering the coffee shop was approached. This is a time‑efficient method but it may introduce bias if, for example, customers who come at busy times behave differently. A truly random sample would be preferable to make the results more reliable. In AQA Statistics, you are expected to discuss the strengths and weaknesses of sampling methods and suggest improvements.
经理在周一早晨 8 点到 10 点之间采用系统抽样收集了数据:每进店的第 3 名顾客被邀请参与。这种方法省时,但如果高峰期的顾客行为不同,则可能引入偏差。若要使结果更可靠,简单随机抽样会更理想。在 AQA 统计中,你需要能够讨论抽样方法的优缺点并提出改进建议。
Because we are observing a sample of size n = 50, the conclusions we draw about the whole population of Coffee Time customers will involve uncertainty. That is why we need inferential statistics – to quantify how confident we can be in our estimates.
由于我们观察的是样本量 n = 50 的样本,关于 Coffee Time 全体顾客的结论将涉及不确定性。因此我们需要推断性统计——以量化我们对估计值的把握程度。
3. Descriptive Statistics: Central Tendency | 描述性统计:集中趋势
Let’s first summarise the spend variable. The raw spend values (in £) for the 50 customers were: 5, 4, 7, 5, 3, 6, 8, 4, 5, 9, … (we will assume the full list yields a mean of £5.84 and a median of £5.50). The mean is calculated as ∑x/n. With n = 50, the sum came to 292, so x̄ = 292 ÷ 50 = 5.84. The median is the middle value when the data are ordered; here the 25th and 26th ordered values were 5 and 6, so the median is (5 + 6)/2 = 5.50.
我们先汇总消费金额变量。50 位顾客的原始消费数据(英镑)为:5, 4, 7, 5, 3, 6, 8, 4, 5, 9, …(假设完整数据算出均值为 £5.84,中位数为 £5.50)。均值由 ∑x/n 计算。n = 50 时总和为 292,因此 x̄ = 292 ÷ 50 = 5.84。中位数是排序后位于中间的数;此处第 25 和第 26 个有序值是 5 和 6,故中位数为 (5 + 6)/2 = 5.50。
The mean is slightly higher than the median, suggesting a mild positive skew – there may be a few customers who spent considerably more than the typical amount. In exam questions, you need to choose the most appropriate measure: the median is resistant to outliers, while the mean uses all data and is more suitable for symmetric distributions.
均值略高于中位数,表明分布有轻微的正偏态——可能存在少数消费远高于典型值的顾客。在考试中,你需要选择最合适的度量:中位数对异常值稳健,而均值利用了所有数据,更适合对称分布。
4. Descriptive Statistics: Measures of Spread | 描述性统计:离散程度
To describe the variability in spend, we compute the standard deviation and the interquartile range (IQR). Using the formula s = √[∑(x − x̄)²/(n−1)], we obtained s = 1.78 (to 2 d.p.). The quartiles are Q₁ = 4, Q₂ = 5.50, Q₃ = 7, so the IQR = 7 − 4 = 3. The range was 9 − 3 = 6, but the IQR is less affected by extreme values.
为了描述消费的变异性,我们计算标准差和四分位距 (IQR)。使用公式 s = √[∑(x − x̄)²/(n−1)],得到 s = 1.78(保留两位小数)。四分位数为 Q₁ = 4,Q₂ = 5.50,Q₃ = 7,故 IQR = 7 − 4 = 3。全距为 9 − 3 = 6,但 IQR 受极端值的影响较小。
When writing a statistical commentary, always pair a measure of spread with a measure of centre. For example, “The median spend was £5.50, with an IQR of £3.00, indicating that the middle 50 % of customers spent between £4 and £7.”
在撰写统计评述时,务必将离散度量与集中趋势度量配对使用。例如,“消费中位数为 £5.50,IQR 为 £3.00,表明中间 50% 的顾客消费在 £4 至 £7 之间。”
5. Probability Foundations and Tree Diagrams | 概率基础与树状图
Now we switch our attention to the cake‐purchasing variable. Out of 50 customers, 17 bought a cake, so the sample proportion is p̂ = 17/50 = 0.34. We can model the event “a randomly selected customer buys a cake” using basic probability rules. Tree diagrams are particularly helpful if we want to consider two customers in sequence. Suppose we select two customers independently. The probability that both buy a cake is 0.34 × 0.34 = 0.1156, while the probability that at least one buys a cake is 1 − (0.66 × 0.66) = 0.5644.
现在我们将注意力转向蛋糕购买变量。在 50 位顾客中,17 人购买了蛋糕,因此样本比例为 p̂ = 17/50 = 0.34。我们可以用基本概率规则对事件“随机抽取一位顾客购买了蛋糕”进行建模。如果想要依次考虑两位顾客,树状图尤其有用。假设我们独立地选取两位顾客。两人都购买蛋糕的概率为 0.34 × 0.34 = 0.1156,而至少一人购买蛋糕的概率为 1 − (0.66 × 0.66) = 0.5644。
AQA Statistics questions often ask you to draw and label a tree diagram, showing clear probabilities on each branch. You must be careful to distinguish between P(A|B) and P(A ∩ B). In our case, since customers are independent, the probability a second customer buys a cake given the first bought one is still 0.34.
AQA 统计试题常要求你绘制并标注树状图,在每条分支上清楚地标明概率。你必须注意区分 P(A|B) 与 P(A ∩ B)。在我们的案例中,由于顾客相互独立,即使已知第一位顾客购买了蛋糕,第二位顾客购买蛋糕的概率依然是 0.34。
6. Discrete Random Variables and Expected Value | 离散随机变量与期望值
Moving further, we can define a discrete random variable X as the number of cakes sold among a small number of customers. For instance, let X be the number of cake buyers in a group of 2 customers. The probability distribution of X is: P(X = 0) = 0.66² = 0.4356, P(X = 1) = 2 × 0.34 × 0.66 = 0.4488, P(X = 2) = 0.34² = 0.1156. The expected number of cakes sold per two customers is E(X) = 0 × 0.4356 + 1 × 0.4488 + 2 × 0.1156 = 0.68, which is simply 2p, as expected.
进一步地,我们可以定义一个离散随机变量 X,表示少数几位顾客中购买蛋糕的人数。例如,令 X 为 2 位顾客中购买蛋糕的人数。X 的概率分布为:P(X = 0) = 0.66² = 0.4356,P(X = 1) = 2 × 0.34 × 0.66 = 0.4488,P(X = 2) = 0.34² = 0.1156。每两位顾客中期望售出的蛋糕数为 E(X) = 0 × 0.4356 + 1 × 0.4488 + 2 × 0.1156 = 0.68,这正是 2p,与预期一致。
In AQA Statistics, you need to be able to compute E(X) and Var(X) from a probability distribution table, and understand that for a general binomial distribution these are np and np(1−p) respectively. This leads us naturally to the binomial model.
在 AQA 统计中,你需要能够根据概率分布表计算 E(X) 和 Var(X),并理解对于一般的二项分布,它们分别是 np 和 np(1−p)。这自然地将我们引向二项分布模型。
7. Modelling with the Binomial Distribution | 二项分布建模
Suppose we consider a working day where the coffee shop serves 10 independent customers, and each has a constant probability p = 0.34 of buying a cake. Let Y ~ B(10, 0.34). We can calculate probabilities such as P(Y = 4) using the formula: P(Y = r) = C(10,r) × (0.34)ʳ × (0.66)¹⁰⁻ʳ. Using a calculator, P(Y = 4) ≈ 0.224. The probability that at least 5 customers buy a cake is P(Y ≥ 5) = 1 − P(Y ≤ 4) ≈ 0.186.
假设我们考虑一个工作日,咖啡店为 10 位相互独立的顾客提供服务,且每位顾客购买蛋糕的恒定概率为 p = 0.34。设 Y ~ B(10, 0.34)。我们可以使用公式 P(Y = r) = C(10,r) × (0.34)ʳ × (0.66)¹⁰⁻ʳ 计算概率。借助计算器,P(Y = 4) ≈ 0.224。至少 5 位顾客购买蛋糕的概率为 P(Y ≥ 5) = 1 − P(Y ≤ 4) ≈ 0.186。
When solving binomial problems in AQA exams, clearly state the distribution, define the random variable, and mention the assumptions: fixed number of trials, two outcomes, constant probability, independence. Here, independence may be violated if a group of friends influences each other’s choices, but for simplicity we assume it holds.
在 AQA 考试中解答二项分布问题时,要清晰地陈述分布、定义随机变量,并说明假设:固定试验次数、两种结果、恒定的概率、独立性。在我们的场景中,如果一群朋友相互影响彼此的选择,独立性可能会被违背,但为简化起见我们假设成立。
8. Modelling with the Normal Distribution | 正态分布建模
Our spend data are continuous and roughly symmetric, making the normal distribution a plausible model. Using the sample mean 5.84 and standard deviation 1.78 as estimates of μ and σ, we may assume Spend ~ N(5.84, 1.78²). To check if a spend of £10 is unusually high, we calculate the z‑score: z = (10 − 5.84) / 1.78 ≈ 2.34. From normal tables, P(Z > 2.34) ≈ 0.0096, so only about 0.96 % of customers would be expected to spend £10 or more – it is indeed unusual.
我们的消费数据是连续的且大致对称,因此正态分布是一个合理的模型。将样本均值 5.84 和标准差 1.78 作为 μ 和 σ 的估计值,我们可以假设 Spend ~ N(5.84, 1.78²)。要检验消费 £10 是否异常高,我们计算 z 分数:z = (10 − 5.84) / 1.78 ≈ 2.34。查正态分布表,P(Z > 2.34) ≈ 0.0096,因此预计只有约 0.96% 的顾客消费达到或超过 £10——这确实不寻常。
It is important to remember that the normal distribution is a continuous model, so we may need to apply a continuity correction when modelling discrete quantities such as spend rounded to the nearest pound. However, at this stage, the basic standardisation technique is what the exam mainly tests.
要记住正态分布是连续模型,因此在对四舍五入到整数的离散量(如消费金额)建模时,可能需要应用连续性校正。不过现阶段,考试主要考查的是基本的标准化方法。
9. Sampling Distributions and the Central Limit Theorem | 抽样分布与中心极限定理
If we repeatedly took samples of size n = 50 from the population of all Coffee Time customers, the sample mean spend would follow its own distribution. The Central Limit Theorem tells us that, regardless of the population shape, the sampling distribution of the mean will be approximately normal with mean μ and standard error σ/√n, provided n is large enough. Here, using s = 1.78 as an estimate of σ, the standard error is 1.78 / √50 ≈ 0.252.
如果我们从 Coffee Time 的全体顾客中反复抽取容量为 n = 50 的样本,样本平均消费额将有其自身的分布。中心极限定理告诉我们,无论总体分布形状如何,只要 n 足够大,样本均值的抽样分布将近似服从均值为 μ、标准误为 σ/√n 的正态分布。此处,以 s = 1.78 作为 σ 的估计值,标准误为 1.78 / √50 ≈ 0.252。
This concept underpins confidence intervals and hypothesis tests. Understanding that different samples yield different means, and that these means cluster around the true population mean with a spread measured by the standard error, is fundamental to statistical inference.
这一概念是置信区间和假设检验的基础。理解不同的样本会产生不同的均值,并且这些均值围绕真总体均值分布、其分散程度由标准误衡量,是统计推断的基石。
10. Constructing a Confidence Interval for the Mean | 构建均值的置信区间
Using our knowledge of the sampling distribution, we can construct a 95 % confidence interval for the population mean spend. The formula is x̄ ± z* × (σ/√n). With x̄ = 5.84, s = 1.78, n = 50, and z* = 1.96 for a 95 % level, we obtain: 5.84 ± 1.96 × 0.252 → 5.84 ± 0.494, giving an interval of (£5.35, £6.33). We are 95 % confident that the average spend of all customers lies within this range.
利用对抽样分布的理解,我们可以构建总体均值消费额的 95% 置信区间。公式为 x̄ ± z* × (σ/√n)。代入 x̄ = 5.84,s = 1.78,n = 50,95% 水平下 z* = 1.96,得到:5.84 ± 1.96 × 0.252 → 5.84 ± 0.494,即区间 (£5.35, £6.33)。我们有 95% 的把握认为所有顾客的平均消费额落在此区间内。
Exam questions often require you to interpret the confidence level correctly: it means that if we took many such samples and constructed intervals, 95 % of them would capture the true μ. It is not a probability statement about μ being in this specific interval.
考试中常要求你正确解释置信水平:它意味着如果我们抽取许多类似样本并构建区间,其中 95% 会包含真实的 μ。这并不是一个关于 μ 位于此特定区间内的概率陈述。
11. Introduction to Hypothesis Testing for a Proportion | 比例假设检验初步
The manager suspects the true proportion p of cake buyers is less than 0.40. We can perform a one‑tailed hypothesis test using the binomial distribution. Let p be the population proportion. Set H₀: p = 0.40. H₁: p < 0.40. From our sample of 50, we observed 17 successes. Under H₀, X ~ B(50, 0.40). We need the p‑value = P(X ≤ 17 | p = 0.40). Using a calculator or tables, this is approximately 0.037. With a significance level of 5 %, since 0.037 < 0.05, we reject H₀ and conclude there is sufficient evidence that the proportion is indeed less than 0.40.
店经理怀疑蛋糕购买者的真实比例 p 低于 0.40。我们可以利用二项分布进行单侧假设检验。设 p 为总体比例。建立 H₀: p = 0.40,H₁: p < 0.40。在我们的 50 人样本中,我们观察到 17 次成功。在 H₀ 下,X ~ B(50, 0.40)。我们需要计算 p 值 = P(X ≤ 17 | p = 0.40)。使用计算器或查表,该概率约为 0.037。在 5% 显著性水平下,由于 0.037 < 0.05,我们拒绝 H₀,并得出结论:有充分证据表明真实比例确实低于 0.40。
This test relies on the same binomial assumptions we discussed earlier. Always state your conclusion in context, and avoid saying “prove”. The p‑value tells us how likely we are to see 17 or fewer cake buyers if the true proportion were 0.40; a small p‑value casts doubt on H₀.
该检验依赖于我们之前讨论过的二项分布假设。始终要在上下文中陈述结论,并避免使用“证明”。p 值告诉我们,如果真实比例是 0.40,我们观察到 17 个或更少蛋糕购买者的可能性有多大;p 值越小,对 H₀ 的质疑就越大。
12. Bringing It All Together and Exam Tips | 总结与考试技巧
This case study has shown how the different topics in Year 12 AQA Statistics fit together into a coherent analytical pipeline. From descriptive statistics and graphical displays to probabilistic modelling, and finally to formal inference, each step builds on the last. When you face real exam questions, always read the scenario carefully, identify the variables, note any assumptions, and present your reasoning step‑by‑step.
本案例分析展示了 Year 12 AQA 统计课程的不同主题如何组合成一条连贯的分析管线。从描述统计和图形展示,到概率建模,再到正式的统计推断,每一步都建立在前一步的基础上。当你面对真实考题时,务必仔细阅读情景描述,识别变量,注意任何假设条件,并一步步展示你的推理过程。
Here are a few golden rules: use precise statistical language, show your working even when using a calculator, state your conclusions in plain English that a non‑statistician could understand, and double‑check that your answer makes sense in the context (e.g., a negative probability or a confidence interval for spend that includes negative numbers is obviously wrong).
以下是几条黄金法则:使用精确的统计语言,即使使用计算器也要展示计算过程,用通俗易懂的普通英语陈述结论,使得非统计专业人士也能理解,并再次核对答案在背景下是否合理(例如,负的概率或包含负数的消费置信区间显然是错误的)。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导