📚 Case Study Practical Exercise | 案例分析实战演练
Statistics is not only about numbers—it is about telling stories with data. In this practical case study, we follow a Year 9 student who investigates whether physical exercise affects maths test scores. You will learn how to design a study, collect data, present it clearly, calculate key statistics, and draw meaningful conclusions.
统计不仅仅是关于数字——它是用数据讲述故事。在这个实践案例分析中,我们将跟随一名九年级学生,他调查体育锻炼是否影响数学考试成绩。你将学习如何设计研究、收集数据、清晰地呈现数据、计算关键统计量,并得出有意义的结论。
1. Setting the Scene | 设定场景
Imagine you are a Year 9 student curious about the link between exercise and academic performance. You decide to conduct a small survey in your school. You ask 10 classmates to record their weekly hours of physical activity and their most recent maths test score out of 100.
假设你是一名九年级学生,对运动与学业表现之间的关系感到好奇。你决定在学校开展一项小调查。你请10位同学记录他们每周体育锻炼的小时数,以及他们最近一次百分制数学考试的成绩。
2. Data Collection | 数据收集
Careful data collection is essential. You give each participant a simple form to fill in over one week. The two variables are: ‘Hours of exercise per week’ and ‘Maths test score’. You ensure the data are recorded accurately and anonymously.
仔细的数据收集至关重要。你给每位参与者一份简单的表格,让他们在一周内填写。两个变量是:“每周运动小时数”和“数学考试成绩”。你确保数据记录准确且匿名。
Here are the raw data collected from the 10 students:
以下是从这10名学生收集到的原始数据:
| Student | Exercise (hours) | Score (out of 100) |
|---|---|---|
| A | 2 | 65 |
| B | 5 | 78 |
| C | 1 | 55 |
| D | 8 | 92 |
| E | 3 | 70 |
| F | 6 | 85 |
| G | 0 | 48 |
| H | 4 | 72 |
| I | 7 | 88 |
| J | 2.5 | 60 |
For example, Student D exercised for 8 hours and scored 92, while Student G did no exercise at all and scored 48. We will use these data to explore patterns.
例如,学生D运动了8小时得分92,而学生G完全没运动得分48。我们将使用这些数据探索模式。
3. Organising Data into Tables | 用表格整理数据
Before drawing graphs, it is helpful to sort the data. Let us arrange the exercise hours in ascending order along with the corresponding scores. This makes it easier to spot trends and calculate statistics.
在绘制图表之前,整理数据很有帮助。我们把运动时间按升序排列,并对应写出成绩。这样更容易发现趋势并计算统计量。
| Exercise (hours) | Score |
|---|---|
| 0 | 48 |
| 1 | 55 |
| 2 | 65 |
| 2.5 | 60 |
| 3 | 70 |
| 4 | 72 |
| 5 | 78 |
| 6 | 85 |
| 7 | 88 |
| 8 | 92 |
Now the data are neatly ordered. We can see that generally, higher exercise hours seem to pair with higher scores, but there are some exceptions, such as 2.5 hours with 60 compared to 2 hours with 65.
现在数据排列整齐。我们可以看出,一般来说运动时间越长似乎分数越高,但也有例外,比如运动2.5小时得60分,而2小时却得65分。
4. Visualising Data: Scatter Graphs | 数据可视化:散点图
A scatter graph is the best choice when we have two numerical variables. We plot ‘Exercise hours’ on the x-axis and ‘Maths score’ on the y-axis. Each student becomes a single point on the graph.
当我们有两个数值变量时,散点图是最佳选择。我们把“运动小时数”放在x轴,“数学成绩”放在y轴。每名学生对应图上的一个点。
Imagine plotting these points: (0,48), (1,55), (2,65), (2.5,60), (3,70), (4,72), (5,78), (6,85), (7,88), (8,92). As you move from left to right, the points tend to rise. This suggests a positive relationship.
想象绘制这些点:(0,48)、(1,55)、(2,65)、(2.5,60)、(3,70)、(4,72)、(5,78)、(6,85)、(7,88)、(8,92)。从左向右移动时,点总体呈上升趋势。这表明可能存在正相关关系。
Scatter graphs allow us to see the overall pattern at a glance, even before we calculate any numbers. They also help us identify any outliers—data points that do not fit the general pattern.
散点图让我们在计算任何数值之前就能一眼看到总体模式。它们还能帮助我们识别离群值——即不符合总体模式的数据点。
5. Calculating the Mean | 计算平均数
The mean is a measure of central tendency, often called the average. For exercise hours, we add all values and divide by 10.
平均数是一种集中趋势的度量,通常被称为平均值。对于运动小时数,我们将所有数值相加再除以10。
Mean exercise hours = (2 + 5 + 1 + 8 + 3 + 6 + 0 + 4 + 7 + 2.5) ÷ 10 = 38.5 ÷ 10 = 3.85
For test scores, the sum is 65 + 78 + 55 + 92 + 70 + 85 + 48 + 72 + 88 + 60 = 713. So the mean score is 71.3.
对于考试成绩,总和为65 + 78 + 55 + 92 + 70 + 85 + 48 + 72 + 88 + 60 = 713。因此平均分是71.3。
Mean score = 713 ÷ 10 = 71.3
The mean gives us a typical value, but it can be affected by very high or very low numbers. For example, Student D’s score of 92 pulls the mean upwards slightly.
平均数给出了一个典型值,但它可能受到极高或极低数值的影响。例如,学生D的92分就把平均分略微拉高了。
6. Finding the Median and Mode | 查找中位数和众数
The median is the middle value when data are ordered. For the 10 exercise hours, the fifth and sixth values in order are 3 and 4. So the median is (3 + 4) ÷ 2 = 3.5 hours.
中位数是排序后位于中间的值。对于10个运动小时数,排序后第五和第六个值是3和4。因此中位数为(3 + 4) ÷ 2 = 3.5小时。
For test scores in order: 48, 55, 60, 65, 70, 72, 78, 85, 88, 92. The median lies between 70 and 72, giving 71.
对于排序后的考试成绩:48, 55, 60, 65, 70, 72, 78, 85, 88, 92。中位数在70和72之间,为71。
The mode is the most frequent value. In our data, no exercise hour repeats exactly, and the scores are all unique. Therefore, there is no mode in this small dataset. However, if we had many students, the mode could show the most common score or exercise habit.
众数是出现频率最高的值。在我们的数据中,没有完全重复的运动小时数,所有成绩也都是唯一的。因此,这个小数据集没有众数。但如果我们有很多学生,众数就能显示最常见的成绩或运动习惯。
7. Measuring Spread with Range | 用极差衡量离散程度
The range tells us how spread out the data are. For exercise hours, the highest is 8 and the lowest is 0, so the range is 8 – 0 = 8 hours.
极差告诉我们数据的分散程度。对于运动小时数,最高为8,最低为0,因此极差是8 – 0 = 8小时。
For test scores, the highest is 92 and the lowest is 48, giving a range of 92 – 48 = 44 marks. This shows that the scores vary quite a bit.
对于考试成绩,最高为92,最低为48,极差为92 – 48 = 44分。这表明分数差异相当大。
Range is simple to calculate but only uses two values. Later, you will learn about interquartile range and standard deviation, which give a more complete picture of spread.
极差计算简单,但只用了两个值。往后你们将学习四分位距和标准差,它们能更全面地反映数据的分散情况。
8. Identifying Correlation | 识别相关性
Correlation measures the strength and direction of a linear relationship between two variables. Looking at our scatter graph, as exercise hours increase, the maths scores also tend to increase. This is called a positive correlation.
相关性衡量两个变量之间线性关系的强度和方向。看我们的散点图,随着运动小时数增加,数学成绩也倾向于提高。这称为正相关。
To judge the strength, we check how closely the points follow a straight line. Here, the points are fairly close to an imaginary line sloping upwards, so we can say there is a moderately strong positive correlation.
要判断强度,我们检查这些点沿一条直线排列的紧密程度。这里,各个点相当接近一条向上倾斜的假想直线,因此我们可以说存在中等强度的正相关。
Remember that correlation does not imply causation. Just because two things are linked does not mean one directly causes the other. There could be other factors, such as students who exercise more also having better time management.
请记住,相关性不代表因果关系。两件事有关联并不意味着一个直接导致另一个。可能还有其他因素,比如运动较多的学生时间管理能力也更好。
9. Drawing Conclusions | 得出结论
From our analysis, we can conclude that in this group of 10 students, there is a positive association between weekly exercise hours and maths test performance. The mean score for those who exercise more than the median (3.5 hours) is noticeably higher.
根据我们的分析,可以得出结论:在这10名学生中,每周运动小时数与数学考试成绩之间存在正相关。运动时间高于中位数(3.5小时)的学生,平均分明显更高。
However, we must be cautious. This was a very small sample from one school. We cannot generalise these findings to all Year 9 students. Also, the data do not prove that exercising more will definitely improve your maths grade.
然而,我们必须谨慎。这只是一个来自一所学校的极小样本。我们不能将这些发现推广到所有九年级学生。而且,这些数据并不能证明多运动一定能提高数学成绩。
A good conclusion always states what the data show and acknowledges the limits of the study. In a real investigation, you would also suggest how to improve the research next time.
一个好的结论总是陈述数据所展示的内容,并承认研究的局限性。在实际调查中,你还需要建议下次如何改进研究。
10. Evaluating the Study | 评估研究
Every statistical study has strengths and weaknesses. This study was quick and easy to carry out, and the variables were clearly defined. The scatter graph and summary statistics gave a clear picture of the relationship.
每项统计研究都有优点和缺点。本研究进行起来快速简便,变量定义清晰。散点图和汇总统计清晰地展现了关系。
Weaknesses include the small sample size, which makes it hard to see reliable patterns. Only one maths test was used, and a single test may not reflect true ability. Other factors like sleep, diet, and study time were not recorded, so we cannot isolate the effect of exercise.
缺点包括样本量小,难以看到可靠模式。只使用了一次数学考试,而单次考试可能无法反映真实能力。其他因素如睡眠、饮食和学习时间没有记录,因此我们无法单独评估运动的影响。
To improve, we could survey more students, collect data over several weeks, and measure exercise with an activity tracker. We might also use a control group and record potential confounding variables.
为了改进,我们可以调查更多学生,收集数周数据,并用活动追踪器测量运动量。我们还可以设立对照组,并记录可能的混杂变量。
11. Extension: Simple Probability Application | 拓展:简单概率应用
We can use the data to calculate simple probabilities. If a student is chosen at random from this group, what is the probability that they exercise more than 5 hours per week?
我们可以用这些数据计算简单的概率。如果从这个小组中随机选一名学生,他每周运动超过5小时的概率是多少?
Three students (D, F, and I) exercised more than 5 hours. So the probability is 3 out of 10, or 0.3. We write P(exercise > 5) = 3/10 = 0.3.
有三名学生(D、F和I)运动超过5小时。因此概率是十分之三,即0.3。我们写作P(运动 > 5) = 3/10 = 0.3。
Similarly, the probability that a randomly selected student scored at least 80 is the number of students with scores ≥80 divided by 10. Students D, F, and I scored 92, 85, and 88, so again 3/10 = 0.3.
类似地,随机选择一名学生,其成绩至少为80的概率是得分≥80的学生人数除以10。学生D、F和I分别得了92、85和88,因此也是3/10 = 0.3。
This simple probability exercise shows how statistics and probability work together to help us make predictions based on data.
这个简单的概率练习展示了统计与概率如何协同工作,帮助我们基于数据做出预测。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply