📚 Year 11 OCR Statistics: Case Study Practice | 案例分析实战演练
Welcome to this OCR GCSE Statistics case study walkthrough. Instead of isolated textbook questions, we will tackle three linked datasets from a fictional school health survey. You will practise data collection, sampling, presentation, numerical analysis, probability, correlation, and time series – exactly as they appear in your exam. Work through each section, check your calculations, and build confidence for the real paper.
欢迎来到这个 OCR GCSE 统计学案例实战演练。我们不再做孤立的课本习题,而是处理一项虚构学校健康调查中三个相互关联的数据集。你将练习数据收集、抽样、图表展示、数值分析、概率、相关性以及时间序列——这些正是考试中会出现的全部内容。逐节练习,核对自己的计算,为真正的考试树立信心。
1. The School Health Survey Case | 学校健康调查案例
A fictional secondary school, Oakwood Academy, wants to understand the health profiles of its Year 11 students. The wellness team collects data on body measurements, breakfast habits, and weekly weight changes. The overarching question is whether lifestyle factors are linked to academic outcomes. We will analyse three separate datasets: Dataset A – heights and weights of 30 randomly selected Year 11 students; Dataset B – breakfast habits and maths results for 100 students; Dataset C – weekly weight recordings of a single volunteer over 10 weeks. By working through these, you will simulate the full statistical enquiry cycle: pose a question, collect data, analyse, and interpret.
一所虚构的橡木学院希望了解其11年级学生的健康概况。健康小组收集了身体测量数据、早餐习惯和每周体重变化。核心问题是生活方式因素是否与学业表现存在联系。我们将分析三个独立的数据集:数据集A – 30名随机选择的11年级学生的身高和体重;数据集B – 100名学生的早餐习惯与数学成绩;数据集C – 一名志愿者在10周内的每周体重记录。通过逐步处理这些数据,你将模拟完整的统计调查循环:提出问题、收集数据、分析并解读。
2. Data Collection and Sampling Methods | 数据收集与抽样方法
For Dataset A, the school uses simple random sampling. Every Year 11 student is assigned a number, and a computer generates 30 unique numbers to select the sample. This method gives each pupil an equal chance of being chosen, but it may not represent smaller subgroups perfectly. In contrast, Dataset B is gathered using stratified sampling by gender. Oakwood Academy has 52% male and 48% female students in Year 11, so the sample of 100 is drawn with exactly 52 boys and 48 girls to maintain the population proportion. Stratified sampling improves representativeness when a characteristic (here gender) is expected to affect the variable of interest.
对于数据集A,学校采用简单随机抽样。每位11年级学生被分配一个编号,计算机生成30个不重复的号码来选择样本。这种方法让每个学生都有同等被选中的机会,但它可能无法完美体现较小的子群体。相比之下,数据集B采用按性别分层抽样收集。橡木学院的11年级学生中有52%为男生,48%为女生,因此抽取的100个样本中恰好包含52名男生和48名女生,以保持总体比例。当某个特征(此处为性别)预计会影响研究变量时,分层抽样能提高代表性。
Dataset C uses opportunity sampling. One student volunteers to record their weight every Monday morning for 10 weeks. This is quick and practical but may introduce bias because volunteers might be more health‑conscious. The time series data allow us to detect short‑term patterns and smooth fluctuations.
数据集C采用机会抽样。一名学生自愿在每周一早上记录体重,持续10周。这种方法快速且实际,但可能引入偏差,因为志愿者可能更注重健康。时间序列数据则帮助我们检测短期形态并平滑波动。
3. Organising Data with Frequency Tables | 用频率表整理数据
Let us begin with Dataset A. The 30 heights (in cm) are listed below. Having raw data is messy, so we group them.
我们先从数据集A开始。以下是30个身高数据(单位:厘米)。原始数据很凌乱,因此我们将其分组。
Raw heights: 152, 155, 157, 158, 159, 160, 162, 162, 163, 164, 165, 165, 166, 167, 168, 169, 170, 170, 171, 172, 172, 173, 174, 175, 175, 176, 178, 179, 180, 182.
原始身高数据: 152, 155, 157, 158, 159, 160, 162, 162, 163, 164, 165, 165, 166, 167, 168, 169, 170, 170, 171, 172, 172, 173, 174, 175, 175, 176, 178, 179, 180, 182。
We choose a class interval of 5 cm, starting at 150 and ending at 185. A grouped frequency table is shown below. Notice the class boundaries (e.g. 150–154.999) are represented by 150 ≤ h < 155 to avoid gaps.
我们选择组距为5厘米,从150开始,到185结束。下面是一个分组频率表。请注意,组边界(例如150–154.999)写作150 ≤ h < 155,以避免间距。
| Height (cm) | Tally | Frequency |
|---|---|---|
| 150 ≤ h < 155 | | | | 2 |
| 155 ≤ h < 160 | | | | | 3 |
| 160 ≤ h < 165 | | | | | | 4 |
| 165 ≤ h < 170 | | | | | | | 5 |
| 170 ≤ h < 175 | | | | | | | | 6 |
| 175 ≤ h < 180 | | | | | | 4 |
| 180 ≤ h < 185 | | | | | 3 |
This grouped table makes the distribution visible at a glance. Most students are clustered between 165 and 175 cm, with fewer at the extremes. The total frequency checks out: 2+3+4+5+6+4+3 = 30.
这张分组表让人一眼就能看到分布特征。大多数学生集中在165–175厘米,两端人数较少。总频数核对无误:2+3+4+5+6+4+3 = 30。
4. Visualising Data: Histograms and Box Plots | 数据可视化:直方图和箱线图
Because all class intervals are equal width (5 cm), a histogram of frequency is straightforward. The heights of the bars correspond directly to the frequencies: the 170–175 group would have the tallest bar (6). A frequency density calculation is not required here, but if you were given unequal intervals, you would use frequency ÷ class width.
由于所有组距相等(5厘米),绘制频率直方图就很简单。直条的高度直接对应频数:170–175组将拥有最高的柱形(6)。此处不需要计算频率密度,但如果遇到不等组距,你就需要使用频率 ÷ 组宽。
For a box plot, we need the five‑number summary. Sorted data (already given) allow us to find: Minimum = 152; Lower quartile Q₁ = position (30+1)/4 = 7.75th value, so Q₁ = 162; Median Q₂ = 15.5th value, average of 15th and 16th = (169+170)/2 = 169.5; Upper quartile Q₃ = 23.25th value, Q₃ = 175; Maximum = 182. The interquartile range (IQR) = 175 − 162 = 13.
绘制箱线图需要五数概括。排序后的数据(已给出)让我们找到:最小值 = 152;下四分位数 Q₁ = 位置 (30+1)/4 = 第7.75项,因此 Q₁ = 162;中位数 Q₂ = 第15.5项,取第15和第16项的平均值 = (169+170)/2 = 169.5;上四分位数 Q₃ = 第23.25项,Q₃ = 175;最大值 = 182。四分位距 (IQR) = 175 − 162 = 13。
We can now sketch a box plot with a whisker from 152 to 182, a box from Q₁ to Q₃ (162–175), and a median line inside at 169.5. The distribution is slightly skewed to the left because the left whisker is longer (162−152=10) than the right (182−175=7). Box plots are excellent for comparing multiple groups, for example if we split by gender.
现在我们可以绘制箱线图:触须从152到182,箱子从Q₁到Q₃(162–175),内部中位线在169.5处。该分布略向左偏,因为左触须(162−152=10)比右触须(182−175=7)更长。箱线图非常适合比较多个组别,例如若按性别分组。
5. Measures of Central Tendency | 集中趋势测度
The three common averages are all useful here. Using the raw data, the mean height is the sum of all 30 values divided by 30. Sum = 152+155+…+182 = 5083 (please verify). Mean x̄ = 5083 ÷ 30 ≈ 169.4 cm. The median (169.5 cm) is virtually identical, confirming symmetry. The mode from the raw list is a bit less helpful — values 165, 170, 172, 175 each appear twice, so the data are bimodal or multimodal. In grouped data, the modal class is 170–175 with frequency 6.
三种常见平均数在这里都能派上用场。使用原始数据,平均身高 = 所有30个值之和 ÷ 30。总和 = 152+155+…+182 = 5083(请自行验证)。平均值 x̄ = 5083 ÷ 30 ≈ 169.4 厘米。中位数(169.5厘米)几乎完全相等,印证了分布的对称性。从原始列表看众数的帮助不大——165、170、172、175都分别出现了两次,数据呈双峰或多峰。在分组数据中,众数组是170–175,频数为6。
x̄ = Σx ÷ n = 5083 ÷ 30 ≈ 169.4 cm
平均值 x̄ = Σx ÷ n = 5083 ÷ 30 ≈ 169.4 厘米
For skewed data, the median is usually a better measure as it is resistant to outliers. Here the similarity between mean and median suggests the absence of extreme values. Always justify your choice of average in exam contexts.
对于偏斜的数据,中位数通常是更佳的度量,因为它不易受异常值影响。这里平均值和中位数十分接近,表明不存在极端值。在考试中,务必说明你选择该平均数的理由。
6. Measures of Spread | 离散程度测度
Spread helps us understand variability. The simplest measure is the range = 182 − 152 = 30 cm. The interquartile range (IQR) we already computed as 13 cm, which captures the middle 50% of the data. A smaller IQR indicates that the central half of students have fairly similar heights.
离散程度帮助我们理解变异程度。最简单的度量是极差 = 182 − 152 = 30 厘米。我们此前已算出四分位距 (IQR) 为13厘米,它涵盖了中间50%的数据。较小的IQR表明,中间一半学生的身高相当接近。
The standard deviation takes every value into account. For a sample we use the formula with n−1. Let us compute it stepwise. First subtract the mean from each height, square the difference, sum those squares, divide by 29, then take the square root. Manual calculation gives Σ(x − x̄)² ≈ 996.7 (rounded). Thus sample standard deviation s = √(996.7 ÷ 29) ≈ √34.37 ≈ 5.86 cm. This tells us that a typical height deviates from the mean by about 5.9 cm. Combining the mean and standard deviation allows us to make statements about the data’s distribution. Always show your working meticulously.
标准差则考虑了每一个值。对于样本,我们使用包含n−1的公式。让我们逐步计算。首先,用每个身高减去平均值,再将差值平方,求出这些平方和,除以29,最后取平方根。手动计算可得 Σ(x − x̄)² ≈ 996.7(已四舍五入)。因此,样本标准差 s = √(996.7 ÷ 29) ≈ √34.37 ≈ 5.86 厘米。这告诉我们,典型的身高偏离平均值约5.9厘米。结合平均值和标准差,我们就能够对数据分布做出描述。务必详细展现计算过程。
s = √(Σ(x − x̄)² ÷ (n−1)) ≈ √(996.7 ÷ 29) ≈ 5.9 cm
样本标准差 s = √(Σ(x − x̄)² ÷ (n−1)) ≈ √(996.7 ÷ 29) ≈ 5.9 厘米
A low standard deviation relative to the mean signals that most students’ heights are tightly packed around 169 cm. If we compare these with weight data later, we may find larger relative variation.
相对于平均值而言,较低的标准差表明大多数学生的身高紧密聚集在169厘米附近。稍后与体重数据对比时,我们可能会发现更大的相对变异。
7. Probability Analysis: Breakfast and Results | 概率分析:早餐与成绩
We now turn to Dataset B, a two‑way table of 100 students classified by whether they eat breakfast regularly and whether they passed their latest maths test. This is perfect for conditional probability and tree diagrams.
我们转而分析数据集B,这是一个100名学生的双向表,按照是否规律吃早餐以及最近一次数学考试是否及格进行分类。这些数据非常适合处理条件概率和树形图。
| Pass | Fail | Total | |
|---|---|---|---|
| Eats breakfast | 40 | 10 | 50 |
| Does not eat breakfast | 20 | 30 | 50 |
| Total | 60 | 40 | 100 |
From the table, P(pass) = 60/100 = 0.6. But if we condition on eating breakfast, P(pass|breakfast) = 40/50 = 0.8. For those skipping breakfast, P(pass|no breakfast) = 20/50 = 0.4. The probability of passing is twice as high for breakfast eaters, which suggests an association. You could construct a probability tree with first branch “Breakfast (0.5)” and “No breakfast (0.5)”, then second branches for pass/fail. Always test independence: if P(pass|breakfast) = P(pass), they would be independent. Here 0.8 ≠ 0.6, so the events are not independent.
根据表格,P(及格) = 60/100 = 0.6。但如果以吃早餐为条件,P(及格|吃早餐)
Published by TutorHao | Year 11 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导