Practical Case Study in Statistics | 案例分析实战演练

📚 Practical Case Study in Statistics | 案例分析实战演练

Statistics is not just a collection of numbers and formulas – it is a powerful tool for making sense of the world. In this case study, we will walk through a realistic data investigation step by step, from posing a question to drawing conclusions. You will see how the Year 9 CAIE Statistics syllabus comes alive when we tackle a genuine problem: understanding the weekly exercise habits of students.

统计学不仅仅是数字和公式的集合,它是理解世界的有力工具。在这个案例分析中,我们将一步一步走过真实的数据调查过程,从提出问题到得出结论。你会看到,当我们解决一个真实的问题——了解学生们每周锻炼的习惯时,Year 9 CAIE 统计学的知识点是如何活起来的。

1. Defining the Problem and Data Needs | 定义问题与数据需求

Our school wants to know how many hours Year 9 students spend on physical exercise each week. We turn this curiosity into a clear statistical question: ‘What is the typical amount of weekly exercise for a Year 9 student in our school, and how much do individual students vary?’ To answer this, we need to collect data on hours of exercise per student. The data should be numerical (discrete, measured to the nearest whole hour) and collected from a representative group.

我们学校想知道九年级学生每周花多少小时进行体育锻炼。我们把这种好奇心转变为一个明确的统计问题:“我们学校九年级学生每周典型的锻炼量是多少?不同学生之间有多大差异?”为了回答这个问题,我们需要收集每位学生每周锻炼小时数的数据。这些数据应该是数值型的(离散数据,精确到最接近的整小时),并且从一个有代表性的群体中收集。


2. Designing a Survey Questionnaire | 设计调查问卷

We design a short questionnaire with one main question: ‘How many hours did you spend on physical exercise (e.g., sports, PE, dancing, gym) in the last full 7‑day week?’ We also ask for their form group but keep responses anonymous. We include clear instructions to round to the nearest hour and avoid double-counting multi‑sport sessions. The questionnaire uses simple, unambiguous language and avoids leading questions.

我们设计了一份简短的问卷,其中有一个主要问题:“在刚刚过去的一个完整星期(7天)里,你花了多少小时进行体育锻炼(例如运动、体育课、舞蹈、健身)?”我们还问了他们所在的年级组,但保持匿名。问卷中包含清晰的说明,要求四舍五入到最接近的整小时,并避免重复计算多项运动的时间。问卷语言简洁、无歧义,并避免了诱导性问题。


3. Choosing a Sampling Method | 选择抽样方法

We cannot survey all Year 9 students in the country, so we take a sample of 30 students from our school. We use a stratified sampling method: Year 9 has 180 students divided equally into three form groups (A, B, C). We randomly select 10 students from each form group. This ensures each group is fairly represented and reduces bias, making our sample more representative than a simple convenience sample.

我们不可能调查全国所有九年级学生,因此我们从学校里抽取了30名学生作为样本。我们采用了分层抽样方法:九年级共有180名学生,平均分成三个年级组(A、B、C)。我们从每个组随机抽取10名学生。这保证了每个组都能得到公平的代表,并减少了偏差,使得我们的样本比简单的方便样本更具代表性。


4. Collecting the Raw Data | 收集原始数据

The 30 students returned their questionnaires and we recorded the hours per week. The raw data set is shown below (each value is hours of exercise in one week):

30名学生交回了问卷,我们记录了每周锻炼的小时数。原始数据集如下(每个数值代表一周的锻炼小时数):

2, 5, 3, 0, 7, 4, 1, 6, 3, 2, 8, 4, 3, 5, 1, 0, 6, 2, 4, 7, 3, 5, 1, 9, 4, 2, 6, 3, 0, 5

We double-check for any obvious errors (e.g., a value of 50 would be impossible) and find none. The data is now ready for organisation.

我们检查了是否存在明显的错误(例如,50小时的数值是不可能的),未发现异常。现在数据已准备好进行整理。


5. Organising Data in a Frequency Table | 用频数表整理数据

To see patterns, we group the raw hours into a frequency distribution table. Since the values range from 0 to 9, we can list each discrete value (0 hours, 1 hour, …, 9 hours) and count how many students reported that number.

为了看出规律,我们将原始的小时数整理成一个频数分布表。由于数值范围是0到9,我们可以列出每一个离散值(0小时、1小时、…、9小时),并数出有多少学生报告了该数值。

Hours (x) Tally Frequency (f)
0 III 3
1 III 3
2 IIII 4
3 IIII I 5
4 IIII 4
5 IIII 4
6 III 3
7 II 2
8 I 1
9 I 1
Total 30

This table makes it easy to see that 3 to 5 hours is the most common interval. The distribution appears roughly symmetric with a slight right tail.

这张表格让我们很容易看出,3至5小时是最常见的区间。分布大致对称,略带右尾。


6. Visualising Data with Charts | 用统计图展示数据

We first draw a bar chart for the discrete data. Each hour value gets a bar whose height equals its frequency. The bar chart clearly shows the mode at 3 hours (frequency 5) and the overall shape. We also construct a pie chart after calculating proportions: for hour 3, the angle = (5/30) × 360° = 60°. The pie chart helps compare each category’s share of the whole.

我们首先为离散数据绘制条形图。每个小时值都对应一个柱,其高度等于频数。条形图清晰地显示出众数是3小时(频数为5)以及整体形状。我们还计算了比例后绘制了饼图:对于3小时,角度 = (5/30) × 360° = 60°。饼图有助于比较每个类别在整体中的份额。

These visualisations help communicate findings instantly. For example, nearly half of the students (14 out of 30) exercise between 2 and 4 hours.

这些可视化能帮助人们立即了解调查结果。例如,将近一半的学生(30人中的14人)锻炼时间在2到4小时之间。


7. Measure of Central Tendency: Mean, Median, Mode | 集中趋势度量:平均数、中位数、众数

The mode is the most frequent value. From the frequency table, 3 hours occurs 5 times, so mode = 3 hours. The median is the middle value when data are ordered. With 30 values, the median lies between the 15th and 16th ordered values. Counting up the frequencies: cumulatively, 0–2 hours: 3+3+4=10, add 3‑hours: total 15. The 15th value is 3 and the 16th is 4 (since the next value is 4). Median = (3+4)/2 = 3.5 hours.

众数是出现频率最高的数值。从频数表看,3小时出现了5次,因此众数 = 3小时。中位数是将数据排序后处于中间位置的值。有30个数据,中位数位于第15个和第16个排序值之间。累计频数:0–2小时:3+3+4=10,加上3小时共15。第15个值是3,第16个值是4(因为下一个值是4)。中位数 = (3+4)/2 = 3.5小时。

To find the mean, we multiply each value by its frequency, sum, and divide by 30:

计算平均数,我们将每个值乘以它的频数,求和,再除以30:

∑(x·f) = (0×3)+(1×3)+(2×4)+(3×5)+(4×4)+(5×4)+(6×3)+(7×2)+(8×1)+(9×1) = 0+3+8+15+16+20+18+14+8+9 = 111

Mean = 111 ÷ 30 = 3.7 hours

The mean (3.7 h) is slightly higher than the median (3.5 h), suggesting a small positive skew. All three measures indicate a typical student exercises around 3–4 hours per week.

平均数(3.7小时)略高于中位数(3.5小时),表明分布有轻微的正偏态。三个指标都表明,典型学生的每周锻炼时间在3至4小时左右。


8. Measure of Spread: Range | 离散程度度量:极差

Spread tells us how consistent the exercise habits are. The simplest measure is the range: maximum – minimum. From the raw data, max = 9 hours, min = 0 hours, so the range = 9 – 0 = 9 hours. A range of 9 hours is quite large compared to the typical value, showing that students’ exercise habits vary widely.

离散程度告诉我们锻炼习惯的一致性如何。最简单的度量是极差:最大值 − 最小值。从原始数据看,最大值 = 9小时,最小值 = 0小时,因此极差 = 9 − 0 = 9小时。与典型值相比,9小时的极差相当大,表明学生的锻炼习惯差异很大。


9. Introduction to Probability from Data | 根据数据引入概率

We can use our sample to estimate probabilities. For example, if we pick one student at random from our 30, what is the probability that they exercise more than 4 hours per week? Count students with x > 4: hours 5,6,7,8,9 give frequencies 4+3+2+1+1 = 11. So P(more than 4 hours) = 11/30 ≈ 0.367.

我们可以用样本估计概率。例如,如果我们从这30人中随机选一名学生,他每周锻炼超过4小时的概率是多少?统计x > 4的学生:5,6,7,8,9小时的频数为4+3+2+1+1 = 11。因此P(超过4小时) = 11/30 ≈ 0.367。

Similarly, the probability that a student does no exercise at all is P(0 hours) = 3/30 = 0.1. These probabilities are empirical and only describe our sample, but they give a reasonable estimate for the whole year group if the sample is representative.

类似地,一名学生完全不锻炼的概率为P(0小时) = 3/30 = 0.1。这些概率是经验性的,仅描述我们的样本,但如果样本具有代表性,它们可以为整个年级提供合理的估计。


10. Drawing Conclusions and Writing a Report | 得出结论并撰写报告

Our investigation found that Year 9 students at this school exercise for an average of 3.7 hours per week, but the wide range (0–9 hours) shows inequality in activity levels. The mode is 3 hours, and half the students exercise 3.5 hours or less. About one‑third exceed the 4‑hour mark. We recommend the school promote more opportunities for physical activity, especially targeting students who currently do very little. A full statistical report would include our methodology, tables, charts, calculations, and limitations (e.g., self‑reported data may be inaccurate).

我们的调查发现,该校九年级学生平均每周锻炼3.7小时,但极差很大(0–9小时)反映出活动水平不均衡。众数为3小时,半数学生锻炼3.5小时或更少。约三分之一的学生超过4小时。我们建议学校推广更多体育锻炼机会,特别是针对目前活动量很少的学生。一份完整的统计报告应包括我们的方法、表格、图表、计算以及局限性(例如,自我报告的数据可能不准确)。


11. Reflecting on the Statistical Process | 反思统计过程

This case study illustrates the full statistical enquiry cycle: question, plan, data collection, processing, analysis, and conclusion. Each stage relies on careful thinking. For instance, our choice of stratified sampling reduced bias, and organising data into a frequency table revealed the shape of the distribution. Without such structure, raw numbers would be meaningless.

这个案例分析展示了完整的统计探究循环:提出问题、制定计划、收集数据、处理数据、分析、得出结论。每个阶段都需要谨慎思考。例如,我们选择分层抽样减少了偏差,将数据整理成频数表揭示出分布的形状。没有这样的结构,原始数字将毫无意义。

We also see how statistics helps answer real questions and can guide decision‑making, such as whether the school needs to introduce new sports clubs.

我们也能看到统计学是如何帮助回答真实问题的,并能为决策提供依据,例如学校是否需要推出新的体育俱乐部。


12. Practising with Your Own Data | 用你自己的数据进行练习

You can replicate this exercise by choosing a topic that interests you – screen time, hours of sleep, number of books read in a month. Design a short questionnaire, collect data ethically from friends or family (with permission), and work through all the steps: frequency table, charts, mean, median, mode, range, and simple probability statements. This hands‑on practice builds confidence and deepens your understanding far beyond textbook exercises.

你可以选择一个自己感兴趣的主题来重复这个练习——屏幕使用时间、睡眠时长、一个月读了多少本书。设计一份简短的问卷,在获得许可的前提下从朋友或家人那里收集数据,然后完成所有步骤:频数表、统计图、平均数、中位数、众数、极差和简单的概率表述。这种动手练习能建立信心,并远比课本练习更深入地加深你的理解。

Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version