IGCSE CAIE Statistics: Case Study Practice | IGCSE CAIE 统计:案例分析实战演练

📚 IGCSE CAIE Statistics: Case Study Practice | IGCSE CAIE 统计:案例分析实战演练

Welcome to this case study walkthrough designed for IGCSE CAIE Statistics. We will explore how statistical concepts are applied in a real-world scenario, helping you master the skills needed for your exam. The case involves analysing data from 30 students, looking at their weekly social media usage hours and their mathematics test scores out of 100. By the end, you will see how descriptive statistics, probability, distributions, and hypothesis testing come together in a practical investigation.

欢迎来到这篇专为 IGCSE CAIE 统计学设计的案例分析演练。我们将探讨统计学概念在真实情境中的应用,帮助你掌握考试所需的技能。该案例分析了 30 名学生每周社交媒体使用时长(小时)和数学测验分数(满分 100)的数据。通过练习,你将看到描述性统计、概率、分布和假设检验如何在实践中结合。

1. Introducing the Data Scenario | 数据情境介绍

Our case study concerns a group of 30 Year 11 students at an international school. Each student recorded the average number of hours they spent on social media per week (variable X) and their most recent mathematics test score out of 100 (variable Y). The teacher wants to know whether social media use is associated with test performance and to describe the overall achievement of the class. We will treat this dataset as a sample to practice essential IGCSE Statistics techniques.

本案例涉及一所国际学校 30 名 11 年级学生。每名学生记录了每周平均社交媒使用时长(变量 X)和最近一次数学测验分数(满分 100,变量 Y)。老师希望了解社交媒体使用是否与测验成绩相关,并描述班级的整体学业水平。我们将该数据集视为一个样本,练习关键的 IGCSE 统计学方法。


2. Data Collection Considerations | 数据收集注意点

Before crunching numbers, it is important to reflect on data collection. The students self-reported their social media hours, which could introduce response bias. The sample size is 30, which is moderate but may not be representative of all Year 11 students. The type of data is bivariate: X is continuous, Y is discrete (though we treat scores as continuous). This is primary data collected directly from the individuals, and it is a cross-sectional snapshot.

在计算之前,反思数据收集很重要。学生自报社交媒体使用时长,可能引入回答偏倚。样本容量为 30,属于中等规模,但未必能代表所有 11 年级学生。数据类型是双变量:X 为连续变量,Y 为离散变量(但我们可将分数视为连续)。这是直接从个体收集的一手数据,且为横截面快照。


3. Organising Data: Frequency Tables | 数据整理:频数表

We organise the 30 social media hours into a grouped frequency table. Choosing class intervals of width 5 hours: 0–4, 5–9, 10–14, 15–19, 20–24. The frequencies are: 4, 10, 9, 5, 2 respectively. For test scores, we group them: 40–49 (3), 50–59 (5), 60–69 (8), 70–79 (7), 80–89 (5), 90–100 (2). This summarisation allows us to see the distribution shape and calculate statistics later.

我们将 30 个社交媒体时长数据整理成分组频数表。选择组距为 5 小时:0–4, 5–9, 10–14, 15–19, 20–24。频数分别为:4, 10, 9, 5, 2。对于测验分数,分组如下:40–49(3人),50–59(5人),60–69(8人),70–79(7人),80–89(5人),90–100(2人)。这种汇总有助于我们观察分布形态并进行后续计算。


4. Graphical Displays: Histograms and Scatter Plots | 图形展示:直方图与散点图

For the grouped data, a histogram is appropriate. The horizontal axis represents hours, and the vertical axis shows frequency density (frequency / class width). Since all class widths are equal, the height corresponds to frequency. We would draw bars with continuous boundaries. To explore the relationship between X and Y, a scatter plot is drawn with social media hours on the x-axis and test scores on the y-axis. The plot shows a general negative trend: higher social media use tends to associate with lower scores.

对于分组数据,直方图很合适。横轴表示小时,纵轴表示频数密度(频数 / 组距)。由于所有组距相等,柱高直接对应频数。我们将绘制连续边界的长条。为探究 X 与 Y 关系,绘制散点图,横轴为社交媒体时长,纵轴为测验分数。散点图呈现总体负相关趋势:社交媒体使用越多,分数往往越低。


5. Measures of Central Tendency | 集中趋势度量

Using the raw data (not shown in full here), the sample mean social media hours x̄ is 10.2 hours, and the median is 10 hours. The mode (most frequent interval midpoint) is 7.5 hours (the 5–9 group). For test scores, the mean ȳ is 68.3, median is 69, and the modal class is 60–69. Because the mean and median are close, the distributions are roughly symmetric, though social media hours show a slight right skew (mean > median).

利用原始数据(此处未全显示),社交媒体时长的样本均值 x̄ 为 10.2 小时,中位数为 10 小时。众数(频数最高区间的中点)为 7.5 小时(5–9 组)。测验分数的均值 ȳ 为 68.3,中位数为 69,众数所在组为 60–69。由于均值与中位数接近,分布大致对称,但社交媒体时长略呈右偏(均值 > 中位数)。


6. Measures of Spread and Box Plots | 离散程度度量与箱线图

Spread is quantified by range, interquartile range (IQR) and standard deviation. For social media hours: range = 22 hours, IQR = Q3 – Q1 = 14 – 6.5 = 7.5 hours, and sample standard deviation s₁ = 6.2 hours. For scores: range = 55, IQR = 76 – 60 = 16, s₂ = 12.4. Box plots can be drawn to compare distributions. The score box plot shows a larger spread and a slightly lower median, while social media hours appear more clustered except for two outliers (high users).

离散程度通过极差、四分位距(IQR)和标准差量化。社交媒体时长:极差 = 22 小时,IQR = Q3 – Q1 = 14 – 6.5 = 7.5 小时,样本标准差 s₁ = 6.2 小时。分数:极差 = 55,IQR = 76 – 60 = 16,s₂ = 12.4。可以绘制箱线图比较分布。分数箱线图离散程度更大,中位数稍低;社交媒体时长除两个离群值(高频用户)外,数据更为集中。


7. Elementary Probability Concepts | 基础概率概念

From the frequency table, we estimate probabilities. If a student is selected at random, the probability that their social media use is 15 hours or more is (5+2)/30 = 7/30 ≈ 0.233. The probability that a student scores above 80 is (5+2)/30 = 7/30 ≈ 0.233 as well. The conditional probability that a student scores above 80 given they use social media 15+ hours is (2)/7 ≈ 0.286. These empirical probabilities set the stage for more formal inference.

根据频数表,我们可以估计概率。若随机选取一名学生,其社交媒体使用时长在 15 小时及以上的概率为 (5+2)/30 = 7/30 ≈ 0.233。测验分数高于 80 的概率同样为 (5+2)/30 = 7/30 ≈ 0.233。在已知使用社交媒体 15+ 小时的条件下,分数高于 80 的条件概率为 2/7 ≈ 0.286。这些经验概率为更正式的推断奠定基础。


8. Binomial and Normal Distributions | 二项分布与正态分布

Suppose we define ‘success’ as scoring 80 or above. Then p = 0.233. If we randomly sample 8 students with replacement, the number of high scorers follows a Binomial distribution B(8, 0.233). The probability of exactly 2 high scorers is calculated using the formula. Alternatively, if the test scores are approximately normally distributed with μ = 68.3 and σ = 12.4, we can find the probability a randomly chosen student scores above 80 by standardising:

假设定义“成功”为分数达到 80 或以上,则 p = 0.233。若随机有放回地抽取 8 名学生,高分人数服从二项分布 B(8, 0.233)。恰好有 2 名高分学生的概率可通过公式计算。另外,若测验分数近似服从均值为 68.3、标准差为 12.4 的正态分布,我们可以通过标准化求得随机选出的学生分数高于 80 的概率:

z = (80 – 68.3) / 12.4 ≈ 0.94

Using standard normal tables, P(Z > 0.94) ≈ 0.174. This normal approximation result is close to the empirical relative frequency.

查标准正态表得 P(Z > 0.94) ≈ 0.174。该正态近似结果与经验频率相近。


9. Sampling Distributions and Confidence Intervals | 抽样分布与置信区间

If we consider the mean social media hours for all Year 11 students as an unknown population parameter μ, our sample mean 10.2 is a point estimate. To construct a 95% confidence interval, we use the formula:

若将所有 11 年级学生的平均社交媒体时长视为未知总体参数 μ,样本均值 10.2 为其点估计。要构造 95% 置信区间,使用公式:

x̄ ± z* × (s/√n)

With z* = 1.96, s = 6.2, n = 30, the margin of error is 1.96 × 6.2/√30 ≈ 2.22. Thus the interval is (10.2 – 2.22, 10.2 + 2.22) = (7.98, 12.42) hours. We are 95% confident that the true mean social media hours lies between 8.0 and 12.4.

取 z* = 1.96,s = 6.2,n = 30,误差界限为 1.96 × 6.2/√30 ≈ 2.22。因此区间为 (10.2 – 2.22, 10.2 + 2.22) = (7.98, 12.42) 小时。我们有 95% 信心认为真实平均社交媒体使用时数在 8.0 至 12.4 之间。


10. Hypothesis Testing: One Sample Mean | 假设检验:单样本均值

The teacher claims that students use social media for 9 hours per week on average. We test this claim with a two-tailed hypothesis: H₀: μ = 9 vs H₁: μ ≠ 9. Using the sample statistics (x̄ = 10.2, s = 6.2, n = 30), the test statistic is:

老师声称学生平均每周使用社交媒体 9 小时。我们进行双尾检验:H₀: μ = 9

Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading