Year 10 OCR Statistics: Case Study Practical Exercises | 十年级OCR统计:案例分析实战演练

📚 Year 10 OCR Statistics: Case Study Practical Exercises | 十年级OCR统计:案例分析实战演练

In this article, we will work through a series of real-world case studies that bring together the key statistical skills required for the OCR Year 10 Statistics syllabus. Each section focuses on a specific concept – from data collection and presentation to probability and scatter graphs – and shows how to apply it using realistic data sets. By working step by step through the exercises, you will strengthen your ability to choose appropriate methods, perform accurate calculations, and interpret results in context.

在本文中,我们将通过一系列真实世界的案例研究,来综合运用OCR十年级统计教学大纲所要求的关键统计技能。每个部分聚焦一个特定概念——从数据收集与呈现到概率与散点图——并展示如何使用现实数据集来应用这些概念。通过逐步完成这些练习,你将增强选择合适方法、进行准确计算以及结合实际情况解释结果的能力。

1. Case Study 1: Screen Time Survey | 案例一:屏幕时间调查

A group of Year 10 students wants to investigate how many hours per day their peers spend on screens (phones, tablets, computers) outside of schoolwork. The hypothesis is that the average daily screen time exceeds the recommended 2 hours. To investigate this, they need to design a survey, collect data, and analyse it using statistical measures.

一群十年级学生想调查同龄人每天在课外花在屏幕(手机、平板、电脑)上的小时数。假设每日平均屏幕时间超过了建议的2小时。为了验证这一假设,他们需要设计问卷、收集数据,并使用统计指标进行分析。

The first step is to formulate a clear question: ‘How many hours per day do you spend on recreational screen time?’ and to decide on a target population, which might be all Year 10 students in the school. A sample of at least 20 students would be needed for reliable preliminary analysis.

第一步是制定一个明确的问题:“你每天花在娱乐屏幕上的时间是多少小时?”然后确定目标总体,比如学校所有十年级学生。为了进行可靠的初步分析,至少需要20名学生的样本。

2. Data Collection and Sampling Methods | 数据收集与抽样方法

Good data collection starts with a well-designed questionnaire. Leading questions such as ‘Don’t you think you spend too much time on screens?’ must be avoided because they introduce bias. Instead, a neutral wording and a clear response box (e.g. ‘Hours: ____’) should be used. The students could also decide to collect data through a short online form to ensure anonymity and encourage honest answers.

良好的数据收集始于设计良好的问卷。诸如“你不觉得你在屏幕上花了太多时间吗?”这样的引导性问题必须避免,因为它们会引入偏差。相反,应使用中性的措辞和清晰的回答框(例如“小时数:____”)。学生们也可以决定通过简短的在线表单收集数据,以确保匿名性并鼓励诚实回答。

For sampling, a simple random sample can be obtained by assigning a number to every Year 10 student and using a random number generator to select 20 participants. Alternatively, a stratified sample might be used to ensure proportional representation from different tutor groups or gender categories, which can make the sample more reflective of the population. It is important to recognise that convenience sampling (e.g. only asking friends) can lead to unrepresentative results.

对于抽样,可以通过给每位十年级学生分配编号,然后使用随机数生成器选择20名参与者,来获得简单随机样本。另外,也可以使用分层抽样,以确保不同导师小组或性别类别的比例代表,这可以使样本更能反映总体。重要的是要认识到,便利抽样(例如只询问朋友)可能导致不具代表性的结果。

3. Organising Data: Frequency Tables | 整理数据:频率表

Suppose the following data (hours) were collected from 20 students: 2, 1, 3, 4, 2, 5, 3, 2, 4, 3, 3, 4, 5, 2, 3, 4, 4, 3, 2, 5. The first step in analysis is to sort the data in ascending order: 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4, 5, 5, 5. This makes it easier to construct a frequency table and spot patterns.

假设从20名学生收集到以下数据(小时):2, 1, 3, 4, 2, 5, 3, 2, 4, 3, 3, 4, 5, 2, 3, 4, 4, 3, 2, 5。分析的第一步是将数据按升序排列:1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4, 5, 5, 5。这样能更容易地构建频率表并发现规律。

A frequency table groups the data values and shows how often each value occurs. For this data set:

频率表将数据值分组,并显示每个值出现的频率。对于该数据集:

Screen time (hours) Tally Frequency
1 I 1
2 IIII 5
3 IIII I 6
4 IIII 5
5 III 3
Total 20

The frequency table reveals that the most common value (mode) is 3 hours, and very few students reported 1 or 5 hours. It also serves as the basis for drawing charts and calculating summary statistics.

该频率表显示最常见的值(众数)是3小时,而很少有学生报告1小时或5小时。它也为绘制图表和计算汇总统计量提供了基础。

4. Measures of Central Tendency: Mean, Median, Mode | 集中趋势指标:平均数、中位数、众数

The mean is calculated by summing all values and dividing by the number of data points. Using the sorted list, the sum is 1 + 2×5 + 3×6 + 4×5 + 5×3 = 1 + 10 + 18 + 20 + 15 = 64. Therefore:

平均数通过将所有数值求和后除以数据点个数来计算。使用排序列表,总和为 1 + 2×5 + 3×6 + 4×5 + 5×3 = 1 + 10 + 18 + 20 + 15 = 64。因此:

mean = Σx ÷ n = 64 ÷ 20 = 3.2 hours

The median is the middle value when data are ordered. With 20 values, the median lies between the 10th and 11th values. Both the 10th and 11th sorted values are 3, so the median is 3 hours. The mode is also 3 hours, which is the most frequent observation. In this symmetric-looking distribution, the mean, median and mode are all close to 3, suggesting a typical screen time of around 3 hours.

中位数是数据排序后位于中间的值。对于20个数值,中位数位于第10和第11个值之间。排序后第10和第11个值都是3,因此中位数为3小时。众数也是3小时,即出现频率最高的观测值。在这个看似对称的分布中,平均数、中位数和众数都接近3,表明典型的屏幕时间约为3小时。

While the mode is quick to identify, the mean takes every value into account and is useful for further calculations. The median is less affected by extreme values and gives a good central estimate when data are skewed.

虽然众数最容易识别,但平均数考虑了每一个值,对于进一步的计算很有用。中位数受极端值的影响较小,在数据偏斜时能提供一个良好的集中估计。

5. Measures of Spread: Range and Interquartile Range | 离散程度:极差与四分位距

Central tendency alone does not describe how spread out the data are. The simplest measure of spread is the range: maximum minus minimum. Here, the range is 5 − 1 = 4 hours, indicating a wide spread across students.

仅靠集中趋势并不能描述数据的离散程度。最简单的离散度量是极差:最大值减去最小值。这里极差为 5 − 1 = 4小时,表明学生之间差异较大。

A more robust measure is the interquartile range (IQR), which focuses on the middle 50% of data. First, find the lower quartile (Q₁) and upper quartile (Q₃). With 20 values, Q₁ is the median of the lower half (first 10 values): the median of 1, 2, 2, 2, 2, 2, 3, 3, 3, 3 lies between the 5th and 6th values, both 2, so Q₁ = 2. Q₃ is the median of the upper half (last 10 values): 3, 3, 3, 4, 4, 4, 4, 4, 5, 5; the 5th and 6th are 4 and 4, so Q₃ = 4. Thus, IQR = Q₃ − Q₁ = 4 − 2 = 2 hours. This tells us that the middle half of students have screen times within a 2-hour window, suggesting that while some extreme values exist, the bulk of the data is fairly concentrated.

更稳健的度量是四分位距(IQR),它关注中间50%的数据。首先,找到下四分位数(Q₁)和上四分位数(Q₃)。对于20个值,Q₁是下半部分(前10个值)的中位数:1, 2, 2, 2, 2, 2, 3, 3, 3, 3 的中位数位于第5和第6个值之间,两者都是2,因此 Q₁ = 2。Q₃是上半部分(后10个值)的中位数:3, 3, 3, 4, 4, 4, 4, 4, 5, 5;第5和第6个值是4和4,因此 Q₃ = 4。因此,IQR = Q₃ − Q₁ = 4 − 2 = 2小时。这告诉我们,中间一半学生的屏幕时间在2小时窗口内,表明虽然存在一些极端值,但数据主体相当集中。

6. Data Visualisation: Bar Charts and Pie Charts | 数据可视化:柱状图与饼图

A bar chart can be drawn with screen time categories on the horizontal axis and frequency on the vertical axis. The height of each bar corresponds to the frequency: bar for 1 hour height 1, for 2 hours height 5, for 3 hours height 6, for 4 hours height 5, and for 5 hours height 3. The bars should be separated by equal gaps, as screen time hours are discrete numerical data. A bar chart makes it visually clear that the distribution peaks at 3 hours.

可以绘制柱状图,横轴为屏幕时间类别,纵轴为频率。每个柱子的高度对应频率:1小时的柱子高度为1,2小时为5,3小时为6,4小时为5,5小时为3。柱子之间应有相等的间隙,因为屏幕时间小时数是离散数值数据。柱状图清楚地显示出分布峰值在3小时处。

Alternatively, a pie chart can be used to show proportions. The total frequency is 20. The angle for each category is (frequency ÷ 20) × 360°. For example, the 3-hour category accounts for 6 ÷ 20 × 360° = 108°. Using a protractor, you would draw sectors of 18° (1 hour), 90° (2 hours), 108° (3 hours), 90° (4 hours), and 54° (5 hours). A pie chart emphasises that 3 hours is the largest slice, but it does not easily show the order of values.

或者,也可以使用饼图来展示比例。总频率为20。每个类别的角度为 (频率 ÷ 20) × 360°。例如,3小时类别占据 6 ÷ 20 × 360° = 108°。使用量角器,你将画出18°(1小时)、90°(2小时)、108°(3小时)、90°(4小时)和54°(5小时)的扇形。饼图强调了3小时是最大的一块,但它不易显示数值的顺序。

7. Box Plots for Comparison | 箱线图用于比较

Now suppose the students also surveyed a sample of Year 11 pupils and obtained the following screen time data (hours): 3, 4, 4, 5, 5, 5, 6, 6, 6, 6, 7, 7, 7, 8, 8, 8, 9, 9, 10, 11. A box plot is an excellent tool for comparing the two year groups.

现在假设学生们还调查了一组十一年级学生,获得了以下屏幕时间数据(小时):3, 4, 4, 5, 5, 5, 6, 6, 6, 6, 7, 7, 7, 8, 8, 8, 9, 9, 10, 11。箱线图是比较两个年级的绝佳工具。

We first calculate the five-number summary for Year 11: minimum = 3, Q₁ = median of lower half (3,4,4,5,5,5,6,6,6,6) = 5, median (Q₂) = (6+6)/2 = 6, Q₃ = median of upper half (6,7,7,7,8,8,8,9,9,10,11?) careful: 20 values, upper half is 6,7,7,7,8,8,8,9,9,10,11? Actually 20 values sorted, lower half first 10: 3,4,4,5,5,5,6,6,6,6 -> Q₁ = (5+5)/2=5. Upper half next 10: 6,7,7,7,8,8,8,9,9,10,11? Wait there are 11 values here, must take last 10: 7,7,7,8,8,8,9,9,10,11? Need to be precise. Sorted Year 11: 3,4,4,5,5,5,6,6,6,6,7,7,7,8,8,8,9,9,10,11. So first 10: 3,4,4,5,5,5,6,6,6,6 -> Q₁ = average of 5th and 6th = (5+5)/2=5. Last 10: 7,7,7,8,8,8,9,9,10,11 -> Q₃ = average of 5th and 6th in this subset (8 and 8) = 8. Maximum = 11. So Year 11 five-number summary: 3, 5, 6, 8, 11.

我们先计算十一年级的五数概要:最小值 = 3,Q₁ = 下半部分中位数 = 5,中位数 = 6,Q₃ = 8,最大值 = 11。

For Year 10, using sorted data 1,2,2,2,2,2,3,3,3,3,3,3,4,4,4,4,4,5,5,5, the five-number summary is: min=1, Q₁=2, median=3, Q₃=4, max=5.

对于十年级,排序数据 1,2,2,2,2,2,3,3,3,3,3,3,4,4,4,4,4,5,5,5,五数概要为:min=1, Q₁=2, median=3, Q₃=4, max=5。

By drawing parallel box plots on the same scale, we can directly compare centres and spreads. Year 11 has a higher median (6 vs 3), a larger IQR (3 vs 2), and a much higher maximum. The box plot shows that Year 11 students generally have longer screen time, with greater variability. This visual comparison supports a conclusion that there is a difference in screen habits between the two year groups, which could lead to further investigation.

通过在同一尺度上绘制平行箱线图,我们可以直接比较中心和离散度。十一年级的箱线图中位数更高(6 对比 3),IQR 更大(3 对比 2),最大值也高得多。箱线图显示十一年级学生通常屏幕时间更长,且差异更大。这种视觉比较支持一个结论:两个年级的屏幕使用习惯存在差异,这可以引发进一步的调查。

8. Probability Case Study: Rolling Two Dice | 概率案例:掷两个骰子

Probability often features in case studies involving games of chance. Consider the experiment of rolling two fair six-sided dice and recording the sum of the two numbers. One common question is: what is the probability that the sum is 7?

概率经常出现在涉及机会游戏的案例中。考虑一个实验:掷两个公平的六面骰子,并记录两个点数之和。一个常见问题是:点数之和为7的概率是多少?

To solve this, we can list the sample space in a table. Let the outcomes of the first die be rows (1 to 6) and the second die be columns. Each of the 36 possible pairs is equally likely.

为了解决这个问题,我们可以用表格列出样本空间。让第一个骰子的结果作为行(1到6),第二个骰子作为列。36种可能配对都是等可能的。

Die 1 \ Die 2 1 2 3 4 5 6
1 2 3 4 5 6 7
2 3 4 5 6 7 8
3 4 5 6 7 8 9
4 5 6 更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading