📚 Case Study in Statistics: Practical Drills | 统计案例分析实战演练
Statistics comes alive when we apply it to real situations. In this article, we will work through a complete case study – investigating the weekly mobile phone usage of Year 10 students at a school. You will see how to design a study, collect data, summarise it with tables and graphs, calculate key measures, and draw meaningful conclusions. Each step is presented as a practical drill to help you master the skills required for the Cambridge IGCSE Statistics examination. Work through the examples and try the mini‑exercises yourself to build confidence.
统计学在应用于真实情境时才真正生动起来。本文我们将完整演练一个案例——调查一所学校10年级学生每周的手机使用时间。你将看到如何设计研究、收集数据、用表格和图表汇总数据、计算关键度量,并得出有意义的结论。每一步都作为实战练习呈现,帮助你掌握剑桥IGCSE统计学考试所需的技能。逐一学习案例并亲自尝试小练习,能有效增强你的信心。
1. Understanding the Problem | 理解问题
The first step in any statistical investigation is to define exactly what we want to know. In our case, the school wishes to understand how many hours per week its Year 10 pupils spend on their mobile phones. The underlying question might be: does excessive phone use affect academic performance? For now, we shall focus on quantifying phone use. A clear, measurable research question is essential: ‘What is the distribution of weekly mobile phone hours among Year 10 students?’
任何统计调查的第一步都是准确定义我们想要知道什么。在我们的案例中,学校希望了解其10年级学生每周在手机上花费多少小时。潜在的问题可能是:过度使用手机是否影响学业?目前我们将专注于量化手机使用。一个清晰、可测量的研究问题是必不可少的:“10年级学生每周手机使用小时数的分布是怎样的?”
When the goal is vague, the data collected will be useless. A well‑defined problem leads to an appropriate data collection method, suitable variables, and meaningful analysis. Always ask yourself: What is the population? What variable am I measuring? Is it discrete or continuous? In this study, the population is all Year 10 pupils in the school, and the variable is hours of phone use per week – a continuous numerical variable.
如果目标模糊,收集到的数据将毫无用处。一个定义明确的问题能引导出合适的数据收集方法、恰当的变量以及有意义的分析。始终要问自己:总体是什么?我测量的是什么变量?它是离散的还是连续的?在本研究中,总体是学校所有10年级学生,变量是每周手机使用的小时数——一个连续型数值变量。
2. Designing a Questionnaire | 设计问卷
To gather data, we need a simple, unbiased questionnaire. A poorly designed question can skew results. For our case, we could ask: ‘How many hours did you spend on your mobile phone last week? (to the nearest hour)’. This is a clear, direct question. We must avoid leading phrases like ‘Do you waste too much time on your phone?’, which implies a judgment. Closed‑format questions are preferable for numerical answers because they are easier to process.
为了收集数据,我们需要一份简单、无偏的问卷。设计不当的问题会扭曲结果。对于我们的案例,可以这样问:“上周你在手机上花费了多少小时?(四舍五入到整数小时)”。这是一个清晰、直接的问题。我们必须避免引导性用语,比如“你是否在手机上浪费了太多时间?”,这暗含了判断。封闭式问题更适合数值型答案,因为它们更容易处理。
-
Provide clear instructions: ‘Please think about all activities on your phone, including social media, games, and calls.’
提供清晰说明:“请回想手机上的一切活动,包括社交媒体、游戏和通话。”
-
Ensure anonymity so that students answer honestly. A statement like ‘Your answers are confidential and will only be used for statistical purposes’ helps.
确保匿名,让学生诚实回答。一句“你的回答将保密,且仅用于统计目的”会有所帮助。
-
Pilot the questionnaire on a small group first to spot confusing wording.
先在小组中试用问卷,以发现令人困惑的措辞。
3. Sampling Methods | 抽样方法
We cannot survey every Year 10 student in the country, so we take a sample from our school. A good sample represents the population fairly. Two key methods for this study are simple random sampling and stratified sampling. In simple random sampling, every student has an equal chance of being chosen. We could assign each pupil a number and use a random number generator to pick 30 students.
我们无法调查全国每一位10年级学生,因此从我们学校抽取一个样本。一个良好的样本能公平地代表总体。本研究的两种关键方法是简单随机抽样和分层抽样。在简单随机抽样中,每个学生被选中的机会均等。我们可以给每个学生分配一个号码,然后用随机数生成器选出30名学生。
However, if we suspect phone use differs by gender or by class group, stratified sampling is better. We divide the population into strata (e.g., Year 10 boys and girls) and randomly select from each stratum in proportion to its size. Suppose the school has 100 Year 10 boys and 100 Year 10 girls. To select a sample of 30, we would take 15 boys and 15 girls randomly. This guarantees representation of both genders.
然而,如果我们怀疑手机使用因性别或班级而异,分层抽样更好。我们把总体分成层(例如,10年级男生和女生),并按各层大小的比例从每层中随机抽取。假设该校有100名10年级男生和100名10年级女生。要选出30个样本,我们会随机抽取15名男生和15名女生。这保证了两种性别的代表性。
Avoid convenience sampling, such as asking only your friends, because it is likely to be biased. In our case study, we shall assume a simple random sample of 30 Year 10 students was successfully obtained.
避免便利抽样,比如只询问朋友,因为这样很可能有偏差。在我们的案例中,我们假定成功获得了一个30名10年级学生的简单随机样本。
4. Organising Raw Data into Frequency Tables | 原始数据整理成频数表
The thirty weekly mobile‑phone hours collected are: 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32. At first glance, this list is hard to interpret. We can group the data into a frequency table with suitable class intervals. Since the values range from 3 to 32, intervals of width 10 are convenient: 0‑9, 10‑19, 20‑29, 30‑39.
收集到的30个每周手机小时数如下:3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32。乍一看,这一串数字难以解读。我们可以将数据按合适的组距整理成频数表。由于值域从3到32,宽度为10的区间很方便:0‑9, 10‑19, 20‑29, 30‑39。
|
Hours (h) 小时数 |
Tally 划记 |
Frequency (f) 频数 (f) |
|
0 – 9 |
|||| || |
7 |
|
10 – 19 |
|||| |||| |
10 |
|
20 – 29 |
|||| |||| |
10 |
|
30 – 39 |
||| |
3 |
Notice that the intervals are continuous and non‑overlapping. For example, 0‑9 includes values from 0 up to but not including 10. The midpoint of the first class is 4.5, which we can use in later calculations. Always check that the total frequency equals the sample size – here 7+10+10+3 = 30, which matches.
请注意这些区间是连续且不重叠的。例如,0‑9 包括从0到10(不含10)的值。第一组的组中值是4.5,我们可以在后续计算中使用。务必检查总频数是否等于样本量——这里 7+10+10+3 = 30,吻合。
5. Choosing the Right Diagram | 选择合适的统计图
With grouped continuous data, a histogram is the most suitable diagram. Unlike a bar chart, a histogram has bars that touch, and the area of each bar represents frequency. When class widths are equal, the height of the bar is proportional to the frequency. For our data, all class widths are 10, so we can simply plot frequency on the vertical axis.
对于分组的连续数据,直方图是最合适的图表。与条形图不同,直方图的条形彼此接触,且每个条形的面积表示频数。当组距相等时,条形的高度与频数成正比。我们的数据所有组距均为10,因此可以直接在纵轴上标出频数。
Alternatively, a frequency polygon could be plotted by joining the midpoints of the histogram bars. This helps in comparing distributions. A pie chart is not appropriate here because the data are continuous and the number of slices would be too few to show shape. A box‑and‑whisker plot (section 10) can also summarise the data effectively.
或者,可以通过连接直方图各条形顶端中点绘制出频数多边形,这有助于比较分布。饼图在这里不合适,因为数据是连续的,而且扇区太少难以展现形状。箱线图(第10节)也能有效地概括数据。
When you are in the exam, think about the type of data: for categorical data, use bar charts or pie charts; for continuous data, use histograms or line graphs. Always label axes clearly and give your diagram a title.
在考试中,要考虑数据的类型:对于分类数据,使用条形图或饼图;对于连续数据,使用直方图或折线图。始终清晰地标注坐标轴,并为你的图表加上标题。
6. Measures of Central Tendency | 集中趋势的度量
Once the data are organised, we calculate summary statistics. The mean (x̄) gives us the average. With all thirty individual values, we simply sum them and divide by 30. The sum is 3+4+…+32 = 525, so x̄ = 525 ÷ 30 = 17.5 hours.
数据整理好之后,我们计算概括统计量。均值(x̄)给出平均数。利用全部三十个原始数据,只需把它们加起来再除以30。总和是 3+4+…+32 = 525,所以 x̄ = 525 ÷ 30 = 17.5 小时。
The median is the middle value when the data are ordered. With 30 data points, the median lies between the 15th and 16th values. The 15th value is 17 and the 16th is 18, so the median is (17+18)/2 = 17.5 hours. The mode is the value that occurs most frequently – in this evenly distributed set, every value from 3 to 32 appears exactly once, so there is no unique mode. However, using the grouped table, the modal class is 10‑19 hours, with the highest frequency of 10.
中位数是数据排序后居中的数值。有30个数据点,中位数位于第15和第16个值之间。第15个值是17,第16个是18,因此中位数为 (17+18)/2 = 17.5 小时。众数是出现频率最高的值——在这个均匀分布的数据集中,3到32的每个值都恰好出现一次,因此没有唯一的众数。但是,使用分组表,众数组是10‑19小时,频数最高为10。
When the mean and median are almost equal, as here, the distribution is roughly symmetrical. If the mean is much larger than the median, the data are positively skewed (tail to the right). This initial check tells us a lot about the shape.
当均值和中位数几乎相等时,如本例所示,分布大致对称。如果均值远大于中位数,数据是正偏态的(尾部向右)。这一初步检查能告诉我们很多关于分布形状的信息。
7. Measures of Spread | 离散程度的度量
Knowing the average is not enough – we must describe how spread out the data are. The range is the simplest measure: maximum minus minimum = 32 – 3 = 29 hours. However, the range is sensitive to extreme values. A more robust measure is the interquartile range (IQR).
只知道平均值是不够的——我们必须描述数据的离散程度。极差是最简单的度量:最大值减最小值 = 32 – 3 = 29 小时。然而,极差对极端值敏感。一个更稳健的度量是四分位距(IQR)。
To find the quartiles, we split the ordered data into two halves. The first half (first 15 values) gives the lower quartile Q₁ = 9 (the 8th value), and the upper half gives Q₃ = 24 (the 23rd value). Thus IQR = Q₃ – Q₁ = 24 – 9 = 15 hours. The middle 50% of students have phone usage spread over 15 hours.
为找出四分位数,我们把排序后的数据分成两半。前半部分(前15个值)给出下四分位数 Q₁ = 9(第8个值),后半部分给出上四分位数 Q₃ = 24(第23个值)。因此 IQR = Q₃ – Q₁ = 24 – 9 = 15 小时。中间50%的学生手机使用时间分布在15个小时的范围内。
For a more precise spread, we calculate the standard deviation. The formula for sample standard deviation s is:
要得到更精确的离散度,我们计算标准差。样本标准差 s 的公式为:
s = √( ∑(xᵢ – x̄)² / (n – 1) )
Using our data and x̄ = 17.5, squaring each deviation and summing gives ∑(x – x̄)² = 2247.5 (you can verify with a calculator). Then s = √(2247.5 ÷ 29) ≈ √77.5 ≈ 8.80 hours. This tells us that, on average, the individual phone usage deviates from the mean by about 8.8 hours.
使用我们的数据和 x̄ = 17.5,计算每个偏差的平方并求和,得到 ∑(x – x̄)² = 2247.5(可用计算器验证)。然后 s = √(2247.5 ÷ 29) ≈ √77.5 ≈ 8.80 小时。这告诉我们,每个手机使用时间平均偏离均值约8.8小时。
8. Probability from Experiments | 实验中的概率
We can use the sample data to estimate probabilities. For example, what is the probability that a randomly chosen Year 10 student uses a mobile phone for more than 20 hours per week? In our sample of 30, there are 13 values above 20 (21 through 32, plus 20? Actually 20 is not above 20; values 21‑32 are 12, but plus 20? Count: 21,22,23,24,25,26,27,28,29,30,31,32 – that’s 12. But we must check: values from 21 to 32 inclusive: 12 values. So the relative frequency is 12/30 = 0.4. Thus, the estimated probability is 0.4 or 40%.
我们可以使用样本数据来估计概率。例如,随机选出一名10年级学生每周使用手机超过20小时的概率是多少?在我们的30人样本中,大于20的有13个值吗?数一数:21到32共12个数,所以12/30 = 0.4。因此,估计概率为0.4或40%。
Probability is fundamental to statistics. The relative frequency idea forms the basis of many simulations and predictions. The more data we collect, the more reliable our probability estimates become – this is the law of large numbers in action. You could also ask: what is the probability that a student uses the phone between 10 and 19 hours? That’s 10/30 ≈ 0.333.
概率是统计学的基础。相对频率观念是许多模拟和预测的基础。收集的数据越多,我们的概率估计就越可靠——这就是大数定律在起作用。你也可以问:学生手机使用时间在10到19小时之间的概率是多少?那是10/30 ≈ 0.333。
9. Bivariate Data and Scatter Plots | 双变量数据与散点图
Often we want to explore relationships between two variables. Suppose the school also recorded each student’s average hours of sleep per night. We could pair each phone‑usage value with a sleep value. A scatter plot of sleep hours (vertical) against phone hours (horizontal) would allow us to see if there is a correlation. For the sake of illustration, imagine the following five data pairs (phone, sleep): (5, 9), (15, 8), (20, 7), (25, 6.5), (30, 6). These suggest a negative correlation: as phone use increases, sleep tends to decrease.
我们常常希望探究两个变量之间的关系。假设学校还记录了每位学生每晚平均睡眠小时数。我们可以把每个手机使用数据与一个睡眠数据配对。以睡眠时间(纵轴)对手机使用时间(横轴)绘制散点图,可以让我们看到是否存在相关性。为便于说明,设想以下五对数据(手机小时数, 睡眠小时数):(5, 9), (15, 8), (20, 7), (25, 6.5), (30, 6)。这暗示着负相关:随着手机使用增加,睡眠时间趋于减少。
We can describe the correlation by its direction (negative), its strength (fairly strong), and its form (linear or non‑linear). In the IGCSE examination, you will not need to calculate the line of best fit by the method of least squares, but you should be able to draw a line of best fit by eye, use it to estimate one value from the other (interpolation), and comment on reliability. Points far from the trend are outliers and should be mentioned.
我们可以从方向(负)、强度(相当强)和形式(线性或非线性)来描述相关性。在IGCSE考试中,你不需要用最小二乘法计算出最佳拟合线,但你应该能够凭目测画出最佳拟合线,用它从一个变量估计另一个变量(内插),并评论可靠性。远离趋势的点是异常值,应当提及。
10. Cumulative Frequency and Box Plots | 累积频率与箱线图
Cumulative frequency diagrams help us find the median and quartiles without sorting the raw data repeatedly. For our grouped data, we first construct a cumulative frequency table. Using the class intervals and frequencies, we add the frequencies successively. For the interval 30‑39, the cumulative frequency is 30, which is the total.
累积频率图帮助我们无需重复排序原始数据便可找到中位数和四分位数。对于分组数据,我们首先构建累积频率表。使用组区间和频数,我们逐次累加频数。对于区间30‑39,累积频率是30,即总数。
|
Hours (upper boundary) 小时数(上边界) |
Frequency 频数 |
Cumulative Frequency 累积频率 |
|
≤ 9.5 |
7 |
7 |
|
≤ 19.5 |
10 |
17 |
更多咨询请联系16621398022(同微信)
CommentsMore posts |
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导