📚 IGCSE Cambridge Statistics: Case Study Practical Exercises | IGCSE Cambridge 统计:案例分析实战演练
Statistics is not just a collection of formulas; it is a way of making sense of the world through data. In the Cambridge IGCSE curriculum, case studies bridge the gap between theory and real-world application. This article presents a series of practical case study exercises designed to sharpen your analytical skills, covering data collection, presentation, probability, regression, and hypothesis testing. Each scenario is followed by detailed English and Chinese explanations to reinforce understanding and prepare you for examination-style contextual questions.
统计学不只是一堆公式,更是一种通过数据理解世界的方式。在剑桥 IGCSE 课程中,案例分析架起了理论与实际应用之间的桥梁。本文提供一系列实战演练,旨在提升你的分析能力,涵盖数据收集、数据展示、概率、回归和假设检验。每个情景后都配有详细的中英文解释,帮助你深化理解,并应对考试中基于情境的题目。
1. Data Collection Design: A School Canteen Survey | 数据收集设计:学校食堂调查
A student council wants to investigate how satisfied students are with the canteen food. They decide to conduct a survey. One member suggests standing at the canteen entrance and asking every third student who enters. Another suggests handing out questionnaires to all students during form time. Identify the sampling methods used and discuss potential biases.
学生会想调查学生对食堂食物的满意度,决定开展调查。一位成员建议站在食堂入口,每第三个进入的学生进行询问;另一位建议在班会时间向所有学生发放问卷。指出所用的抽样方法并讨论潜在偏倚。
The first method is systematic sampling. By selecting every third student, the sample may still be random if the starting point is random, but there is a risk of periodicity—for example, if students arrive in groups of friends with similar tastes. The second method is a census (or attempt at one), because all students are given the questionnaire. However, non-response bias could be significant if many students do not return the forms, and those who do might feel more strongly about the canteen. In examinations, you should also note that a census is often impractical for large populations and may suffer from coverage error if some students are absent.
第一种方法是系统抽样。通过选取每第三个学生,若起始点随机,样本可能仍具随机性,但存在周期性问题——例如如果学生成团结伴且口味相近。第二种方法试图做普查,因为问卷发给所有学生。然而,如果许多学生不交回问卷,无回答偏倚可能很显著,并且交卷的学生可能对食堂有更强烈的意见。考试中还应注意,普查对大规模总体往往不切实际,若部分学生缺席则会产生覆盖误差。
2. Presenting Data: Choosing the Right Chart | 数据展示:选择正确图表
A geography teacher collected the annual rainfall (in mm) for a town over 20 years and wants to show the trend over time. A student also gathered data on the favorite sports of 80 classmates. For each dataset, recommend a suitable chart and justify your choice.
一位地理老师收集了某镇20年的年降雨量(毫米),想展示时间趋势。一名学生也收集了80位同学最喜好运动的数据。针对每组数据推荐合适的图表并说明理由。
For the rainfall over time, a line graph or time series plot is most suitable because it clearly shows changes and trends over consecutive years. The years are plotted on the x-axis (time) and rainfall on the y-axis. For the favorite sports, which is categorical data, a bar chart or a pie chart is appropriate. A bar chart allows easy comparison of frequencies across categories, while a pie chart shows the proportion of the whole for each sport. In IGCSE, you must label axes, provide a title, and for pie charts, calculate the angle for each sector using (frequency ÷ total) × 360°.
对于多年降雨量,折线图或时间序列图最为合适,因为它能清晰显示连续年份的变化与趋势。年份绘于 x 轴(时间),降雨量绘于 y 轴。对于最喜爱运动这类分类数据,条形图或饼图合适。条形图易于比较各类别的频数,而饼图显示每项运动占整体的比例。在 IGCSE 中,必须标记坐标轴、给出标题,对于饼图要用(频数 ÷ 总数)× 360° 计算每个扇形的角度。
3. Measures of Central Tendency and Spread: Exam Scores Analysis | 集中趋势与离散度量:考试成绩分析
Two classes, A and B, took the same mathematics test. Class A: mean 68, median 70, standard deviation 8. Class B: mean 68, median 60, standard deviation 15. Interpret these statistics and compare the performance of the two classes.
A、B 两班参加同一数学测验。A 班:平均数68,中位数70,标准差8。B 班:平均数68,中位数60,标准差15。解读这些统计量并比较两班表现。
Both classes have the same mean, suggesting average performance is equal. However, the median for Class A is higher than the mean (70 > 68), indicating a slight negative skew; more students scored above the average. In Class B, the median is 60, much lower than the mean, pointing to a positive skew—a few high scores are pulling the mean up. The standard deviation is much larger for Class B (15 vs 8), meaning scores in Class B are more spread out and less consistent. Therefore, Class A performed more consistently and the typical student in Class A did better than the typical student in Class B, even though means are identical. Always look beyond the mean.
两班平均数相同,表明平均表现持平。但 A 班中位数高于平均数(70 > 68),指示轻微负偏态;更多学生得分高于平均。B 班中位数为60,远低于平均数,表明正偏态——少数高分拉高了均值。B 班标准差远大于 A 班(15 对 8),意味着 B 班成绩更分散、一致性更差。因此,尽管平均数相同,A班表现更一致,典型学生得分高于B班典型学生。切勿只看均值。
4. Probability and Tree Diagrams: Defective Light Bulbs | 概率与树状图:不合格灯泡
A factory produces light bulbs on two machines. Machine X makes 60% of the bulbs with a 2% defect rate; Machine Y makes the rest with a 5% defect rate. A bulb is selected at random. (a) Draw a tree diagram. (b) Find the probability that the bulb is defective. (c) Given the bulb is defective, find the probability it came from Machine Y.
某工厂用两台机器生产灯泡。机器 X 生产60%的灯泡,不良率2%;机器 Y 生产其余灯泡,不良率5%。随机抽取一个灯泡。(a) 画树状图。(b) 求该灯泡不合格的概率。(c) 已知灯泡不合格,求它产自机器 Y 的概率。
The tree diagram starts with two branches: X (0.6) and Y (0.4). From X, branches: Defective (0.02) and OK (0.98). From Y: Defective (0.05) and OK (0.95). Overall defective probability P(D) = P(X and D) + P(Y and D) = 0.6×0.02 + 0.4×0.05 = 0.012 + 0.020 = 0.032. For part (c), we use Bayes’ theorem or conditional probability: P(Y|D) = P(Y and D) / P(D) = 0.020 / 0.032 = 20/32 = 5/8 = 0.625. So there is a 62.5% chance that a defective bulb came from Y, even though Y produces only 40% of the bulbs—reflecting its higher defect rate. Always label branches with probabilities and outcomes clearly.
树状图从两个主枝开始:X (0.6) 和 Y (0.4)。从 X 出发:不良 (0.02) 和合格 (0.98)。从 Y 出发:不良 (0.05) 和合格 (0.95)。整体不良概率 P(D) = P(X 且 D) + P(Y 且 D) = 0.6×0.02 + 0.4×0.05 = 0.012 + 0.020 = 0.032。第(c)部分用贝叶斯定理或条件概率:P(Y|D) = P(Y 且 D) / P(D) = 0.020 / 0.032 = 20/32 = 5/8 = 0.625。因此,即便Y仅生产40%灯泡,一个不良品来自Y的概率达62.5%,反映出其更高不良率。务必清晰标注各枝干概率与结局。
5. Scatter Diagrams and Correlation: Study Hours vs Test Scores | 散点图与相关性:学习时长与测验分数
A researcher recorded study hours (x) and test scores (y) for 10 students. The scatter plot shows a positive trend. The calculated Pearson’s correlation coefficient is r = 0.87. Comment on the relationship and caution against assuming causation.
一位研究者记录10名学生的学习时长 (x) 和测验分数 (y)。散点图呈现正向趋势。计算得皮尔逊相关系数 r = 0.87。评论该关系并提醒勿假设因果关系。
An r value of 0.87 indicates a strong positive linear correlation: generally, students who study more hours tend to score higher. However, correlation does not imply causation. A high r could be due to a third variable, such as natural ability or parental support, which influences both study time and test scores. Also, the relationship may not be perfectly linear; always inspect the scatter diagram for outliers or non-linear patterns. In IGCSE, you should mention that even strong correlation does not prove that changing x will cause a change in y. Suggest possible lurking variables: motivation, prior knowledge, or quality of study.
r 值 0.87 表示强正线性相关:通常学习时间越长的学生分数越高。然而,相关不代表因果。高 r 可能源于第三变量,例如天赋或家长支持,同时影响学习时长和分数。此外,关系可能并非完全线性;务必检查散点图有无异常值或非线性模式。在 IGCSE 中,须指出即使强相关也不能证明改变 x 会导致 y 变化。提出可能的潜伏变量:动机、先前知识或学习质量。
6. Regression Line: Predicting Sales from Advertising Spend | 回归直线:由广告支出预测销售额
A company’s monthly advertising spend (£’000) and sales (£’000) over 12 months gave the equation: Sales = 25 + 3.6 × Advertising. (a) Interpret the slope and intercept. (b) Predict sales when advertising spend is £10,000. (c) Explain why prediction for £50,000 spend might be unreliable.
某公司12个月的月广告支出(千镑)与销售额(千镑)得出方程:销售额 = 25 + 3.6 × 广告支出。(a) 解释斜率和截距。(b) 预测广告支出为£10,000时的销售额。(c) 解释为何预测£50,000支出可能不可靠。
The intercept 25 means that if advertising spend is zero, the model predicts sales of £25,000. This may not be meaningful in reality but provides the line’s starting point. The slope 3.6 means that for every additional £1,000 spent on advertising, sales are predicted to increase by £3,600. For a spend of £10,000 (x=10), predicted sales = 25 + 3.6×10 = 25 + 36 = £61,000. However, using the line to predict for x=50 is extrapolation, as 50 lies beyond the range of observed advertising data. The relationship may not remain linear at very high spend levels due to market saturation or diminishing returns, so the prediction may be inaccurate.
截距25表示若广告支出为零,模型预测销售额为£25,000。现实中可能并无意义,但它给出直线的起点。斜率3.6表示广告支出每增加£1,000,销售额预计增加£3,600。支出£10,000 (x=10) 时,预测销售额 = 25 + 3.6×10 = 25 + 36 = £61,000。然而,用该直线预测 x=50 属于外推,因为50超出了观察到的广告数据范围。在极高支出水平下,由于市场饱和或回报递减,关系可能不再保持线性,因此预测可能不准。
7. Probability Distributions: Binomial in Quality Control | 概率分布:二项分布用于质量控制
A factory produces components, and the probability that a component is defective is 0.1. A quality inspector randomly selects 15 components. Find the probability that exactly 2 are defective, and the probability that at most 1 is defective. State the assumptions for a binomial model.
某工厂生产零件,次品概率为0.1。质检员随机抽取15个零件。求恰好有2个次品的概率,以及至多1个次品的概率。列出二项模型的假设。
Let X ~ B(15, 0.1). P(X = 2) = ₁₅C₂ × (0.1)² × (0.9)¹³. Using a calculator or formula, ₁₅C₂ = 105, so P = 105 × 0.01 × 0.2542 ≈ 0.2669 (approx). For P(X ≤ 1) = P(X=0) + P(X=1). P(X=0) = (0.9)¹⁵ ≈ 0.2059; P(X=1) = 15 × 0.1 × (0.9)¹⁴ ≈ 0.3432; sum ≈ 0.5491. The binomial model assumes: fixed number of trials (n=15), each trial is independent, only two outcomes (defective or not), and constant probability of success (p=0.1). In quality control, independence could be violated if components are produced in batches with a common fault.
设 X ~ B(15, 0.1)。P(X = 2) = ₁₅C₂ × (0.1)² × (0.9)¹³。使用计算器或公式,₁₅C₂ = 105,故概率约为 105 × 0.01 × 0.2542 ≈ 0.2669。P(X ≤ 1) = P(X=0) + P(X=1)。P(X=0) = (0.9)¹⁵ ≈ 0.2059;P(X=1) = 15 × 0.1 × (0.9)¹⁴ ≈ 0.3432;总和 ≈ 0.5491。二项模型假设:试验次数固定 (n=15),每次试验独立,只有两种结果,成功概率恒定 (p=0.1)。在质量控制中,若零件按批次生产且存在共同缺陷,独立性可能被破坏。
8. Normal Distribution: Heights of Students | 正态分布:学生身高
The heights of 16-year-old students are normally distributed with mean 165 cm and standard deviation 8 cm. (a) What proportion of students are taller than 175 cm? (b) Find the height that 90% of students exceed. (c) If a sample of 25 students is taken, describe the sampling distribution of the sample mean.
16岁学生身高服从正态分布,均值165 cm,标准差8 cm。(a) 身高高于175 cm的学生比例是多少?(b) 找出90%学生超过的身高。(c) 若抽取一个25名学生的样本,描述样本均值的抽样分布。
(a) Standardize: z = (175 – 165) / 8 = 10/8 = 1.25. Using normal table, P(Z > 1.25) = 1 – 0.8944 = 0.1056, so about 10.6% are taller than 175 cm. (b) We need the 10th percentile (since 90% exceed it). The z-value with 10% in the left tail is approximately -1.2816. So height = 165 + (-1.2816)×8 = 165 – 10.25 ≈ 154.75 cm. (c) For a sample of size n=25, the sampling distribution of the sample mean is also normal (because the population is normal) with mean μ = 165 cm and standard error σ/√n = 8/√25 = 8/5 = 1.6 cm. The distribution is X̄ ~ N(165, 1.6²). Even if the population were not normal, the Central Limit Theorem ensures approximate normality for large n, but here n=25 is moderate; normality of population justifies exact normality.
(a) 标准化:z = (175 – 165) / 8 = 10/8 = 1.25。查正态表,P(Z > 1.25) = 1 – 0.8944 = 0.1056,约有10.6%的学生身高超过175 cm。(b) 需要第10百分位数(因为90%超过该值)。左尾面积0.10对应的 z 值约为 -1.2816。身高 = 165 + (-1.2816)×8 = 165 – 10.25 ≈ 154.75 cm。(c) 对于 n=25 的样本,样本均值的抽样分布也是正态的(因为总体正态),均值为 μ = 165 cm,标准误 σ/√n = 8/√25 = 8/5 = 1.6 cm。分布为 X̄ ~ N(165, 1.6²)。即使总体非正态,中心极限定理保证大样本时近似正态,但此处 n=25 中等;总体正态性确保了精确正态。
9. Confidence Intervals: Estimating Mean Daily Screen Time | 置信区间:估计日均屏幕时间
A random sample of 36 teenagers had a mean daily screen time of 5.2 hours with a standard deviation of 1.8 hours. Construct a 95% confidence interval for the population mean and interpret it.
一个随机样本包含36名青少年,其日均屏幕时间为5.2小时,标准差1.8小时。构建总体均值的95%置信区间并解读。
Since n=36 is large, we can use the z-distribution (or t with 35 df, but z is acceptable at IGCSE for large n). The 95% confidence z-value is 1.96. Standard error = s/√n = 1.8/6 = 0.3 hours. Margin of error = 1.96 × 0.3 = 0.588 hours. Confidence interval = 5.2 ± 0.588, i.e., (4.612, 5.788) hours. Interpretation: we are 95% confident that the true mean daily screen time for all teenagers in the population lies between 4.61 and 5.79 hours. This does not mean there is a 95% probability that the true mean is in this interval; rather, if we repeated the sampling many times, 95% of such intervals would contain the population mean.
由于 n=36 较大,可使用 z 分布(IGCSE 阶段对大样本用 z 值 1.96 可接受)。标准误 = s/√n = 1.8/6 = 0.3小时。误差边际 = 1.96 × 0.3 = 0.588小时。置信区间 = 5.2 ± 0.588,即 (4.612, 5.788) 小时。解读:我们有95%的把握总体所有青少年的真实日均屏幕时间在4.61至5.79小时之间。这并非指真实均值落在该区间的概率是95%,而是指若反复多次抽样,如此构建的区间中有95%会包含总体均值。
10. Hypothesis Testing: Is a Coin Fair? | 假设检验:硬币是否公平?
A student suspects a coin is biased towards heads. In 50 flips, 32 heads are observed. Test at the 5% significance level whether there is evidence that the coin is biased. Clearly state hypotheses, test statistic, critical region, and conclusion.
某学生怀疑一枚硬币偏向正面。投掷50次,观察到32次正面。在5%显著性水平下,检验是否有证据表明硬币有偏。明确写出假设、检验统计量、拒绝域和结论。
Let p be the probability of heads. H₀: p = 0.5 (coin is fair); H₁: p > 0.5 (biased towards heads). One-tailed test at α = 0.05. Under H₀, X ~ B(50, 0.5). We need the critical region: the smallest k such that P(X ≥ k) ≤ 0.05. Using binomial tables or normal approximation: mean = 25, sd = √(50×0.5×0.5) = √12.5 ≈ 3.5355. For continuity correction, we find z = (31.5 – 25)/3.5355 ≈ 1.838. P(Z > 1.838) ≈ 0.033, so the critical region is X ≥ 32. Since 32 falls in the critical region, we reject H₀. There is sufficient evidence at the 5% level to conclude the coin is biased towards heads. Always state conclusion in context and mention significance level.
令 p 为出现正面的概率。H₀: p = 0.5(硬币公平);H₁: p > 0.5(偏向正面)。单尾检验,α = 0.05。在 H₀ 下,X ~ B(50, 0.5)。需要拒绝域:最小的 k 使得 P(X ≥ k) ≤ 0.05。使用二项表或正态近似:均值=25,标准差=√(50×0.5×0.5)=√12.5≈3.5355。连续性校正后,z = (31.5 – 25)/3.5355 ≈ 1.838。P(Z > 1.838) ≈ 0.033,故拒绝域为 X ≥ 32。由于32落在拒绝域,拒绝 H₀。在5%显著性水平下,有足够的证据表明硬币偏向正面。结论要结合情境并提及显著性水平。
11. Chi-Squared Test: Association Between Gender and Favorite Subject | 卡方检验:性别与最喜爱科目间的关联
A survey asked 200 students about their favorite subject: Math, English, or Science. The observed frequencies are: Male: Math 40, English 20, Science 30; Female: Math 30, English 50, Science 30. Conduct a chi-squared test for independence at the 5% level.
一项调查询问200名学生最喜爱的科目:数学、英语或科学。观察频数为:男生:数学40,英语20,科学30;女生:数学30,英语50,科学30。在5%水平下进行独立性卡方检验。
H₀: Gender and favorite subject are independent. H₁: They are not independent. Compute expected frequencies from row and column totals. Row totals: Male 90, Female 110. Column totals: Math 70, English 70, Science 60. Expected for Male-Math = (90×70)/200 = 6300/200 = 31.5. Similarly: Male-English = (90×70)/200 = 31.5; Male-Science = (90×60)/200 = 27; Female-Math = (110×70)/200 = 38.5; Female-English = 38.5; Female-Science = (110×60)/200 = 33. Calculate χ² = Σ (O-E)²/E. Male-Math: (40-31.5)²/31.5 = 2.2857; Male-English: (20-31.5)²/31.5 = 4.2143; Male-Science: (30-27)²/27 = 0.3333; Female-Math: (30-38.5)²/38.5 = 1.8831; Female-English: (50-38.5)²/38.5 = 3.4091; Female-Science: (30-33)²/33 = 0.2727. Sum = 12.3982. Degrees of freedom = (2-1)×(3-1) = 2. Critical value at 5% for 2 df is 5.991. Since 12.40 > 5.991, reject H₀. There is evidence of association between gender and favorite subject.
H₀:性别与最喜爱科目独立。H₁:不独立。由行合计与列合计计算期望频数。行合计:男生90,女生110。列合计:数学70,英语70,科学60。男生-数学期望 = (90×70)/200 = 31.5。类似:男生-英语 = 31.5;男生-科学 = 27;女生-数学 = 38.5;女生-英语 = 38.5;女生-科学 = (110×60)/200 = 33。计算 χ² = Σ (O-E)²/E。男生-数学:(40-31.5)²/31.5 = 2.2857;男生-英语:4.2143;男生-科学:0.3333;女生-数学:1.8831;女生-英语:3.4091;女生-科学:0.2727。总和 = 12.3982。自由度 = (2-1)×(3-1) = 2。5%水平下 df=2 的临界值 5.991。因 12.40 > 5.991,拒绝 H₀。有证据表明性别与最喜爱科目之间存在关联。
12. Real-World Pitfalls: Misleading Statistics in the Media | 现实误区:媒体中的误导性统计
A headline claims: “Eating chocolate boosts exam scores by 50%!” The study behind it asked students whether they ate chocolate before an exam and recorded their scores. Identify three statistical issues that could make this claim misleading.
某标题宣称:“吃巧克力使考试成绩提高50%!”其背后的研究询问学生考前是否吃巧克力并记录分数。指出可能使这一说法产生误导的三个统计问题。
First, the study is observational, not experimental, so causation cannot be established. Students who eat chocolate may differ in other ways (e.g., more relaxed, better socio-economic background). Second, “boosts by 50%” might refer to a relative increase from a very small baseline—e.g., from 2% to 3% is a 50% increase but practically meaningless. Always check absolute differences. Third, self-reported data on chocolate consumption may be unreliable; students may not remember accurately or may exaggerate. Also, the sample may be biased if only volunteers participated. In IGCSE, you must critically evaluate claims by considering study design, confounding variables, and measurement validity.
首先,研究是观察性的而非实验性的,因此不能确立因果关系。吃巧克力的学生可能在其他方面不同(如更放松、社会经济背景更好)。其次,“提高50%”可能指从很小基数的相对增长——如从2%到3%是50%增长但实际无意义。务必检查绝对差值。第三,自我报告的巧克力摄入数据可能不可靠;学生可能记忆不准或夸大。此外,若仅志愿者参与,样本可能有偏倚。在 IGCSE 中,必须通过考虑研究设计、混杂变量和测量效度来批判性评价论断。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply