📚 School Canteen Survey: A Statistical Case Study | 校园食堂调查:统计案例分析实战演练
In this article, we will walk you through a complete statistical case study designed for Year 8 AQA Statistics. You will learn how to design a survey, collect and organise data, create graphs, calculate summary statistics, and draw meaningful conclusions. This practical approach will help you consolidate all the key skills needed for your statistics assessment.
在这篇文章中,我们将带你完整经历一个为 Year 8 AQA 统计课程设计的统计案例研究。你将学习如何设计调查问卷、收集和整理数据、绘制图表、计算汇总统计量,并得出有意义的结论。这种实战式学习能帮助你巩固统计评估所需的所有关键技能。
1. Understanding the Case Study Context | 理解案例背景
Greenwood School has noticed that many students buy snacks from the school canteen, but some staff are concerned about how healthy these choices are. The school council has decided to carry out a statistical investigation into students’ snack preferences. The goal is to collect data, analyse it, and propose improvements to the canteen menu. As a Year 8 statistician, you will guide this project from start to finish.
格林伍德学校发现许多学生从食堂购买零食,但部分教职工担心这些选择的健康程度。学校学生会决定对学生的零食偏好展开一项统计调查。目标是收集数据、进行分析,并向食堂菜单提出改进建议。作为一名 Year 8 统计员,你将从头到尾主导这个项目。
The key statistical question is: ‘What are the snacking habits of students at Greenwood School, and how can the canteen offer healthier choices?’ To answer this, we need to gather both categorical data (favourite snack type) and numerical data (number of snacks bought per week, pocket money). We will then summarise this data using graphs and numbers, before forming a conclusion.
核心统计问题是:“格林伍德学校学生的零食习惯是什么?食堂如何提供更健康的选择?”为了回答这个问题,我们需要同时收集分类数据(最喜欢的零食类型)和数值数据(每周购买零食的数量、零花钱)。然后我们将使用图表和数字来概括这些数据,最终形成结论。
2. Designing Survey Questions and Identifying Data Types | 设计调查问题与识别数据类型
A well-designed survey starts with clear, unbiased questions. We want to collect two types of data: categorical (qualitative) and numerical (quantitative). The planned questionnaire will contain three questions. Question 1: ‘What is your favourite type of snack from the canteen?’ with options Fruit, Crisps, Chocolate, Biscuits, Other. This gives categorical data. Question 2: ‘How many snacks do you buy from the canteen each week?’ This provides discrete numerical data, as it is a count. Question 3: ‘How much pocket money do you receive per week (in £)?’ This yields continuous numerical data, though we might round to the nearest pound.
一个精心设计的调查始于清晰、无偏的问题。我们想收集两类数据:分类(定性)数据和数值(定量)数据。计划中的问卷将包含三个问题。问题 1:“你最喜欢的食堂零食类型是什么?”选项为水果、薯片、巧克力、饼干、其他。这给出分类数据。问题 2:“你每周从食堂购买多少个零食?”这提供离散数值数据,因为它是一个计数。问题 3:“你每周收到多少零花钱(单位:£)?”这产生连续数值数据,尽管我们可能会四舍五入到最接近的整数英镑。
It is important that questions are easy to understand and do not lead respondents towards a particular answer. For instance, asking ‘Do you buy too many unhealthy snacks?’ would be a leading question and introduce bias. Our questions are neutral and allow students to answer honestly. Also, we must decide how many people to survey – this is called the sample size.
问题的措辞必须容易理解,且不能引导受访者选择某个特定答案。例如,问“你是不是买了太多不健康的零食?”就是一个诱导性问题,会引入偏差。我们的问题是中立的,让学生可以诚实回答。此外,我们必须决定调查多少人——这称为样本量。
3. Sampling Methods and Data Collection | 抽样方法与数据收集
We cannot ask every student in the school, so we select a sample. For a quick survey, some may suggest asking students in the Year 8 common room during break – this is a convenience sample. However, a convenience sample could be biased because it only includes Year 8 students who happen to be in that room, not representing the whole school. A better method is a simple random sample: put all students’ names in a hat and draw 30 names. This gives each student an equal chance of being selected.
我们无法询问学校里的每一个学生,因此选择了一个样本。为了快速完成调查,有人可能建议在课间到 Year 8 休息室询问学生——这是一个便利样本。然而,便利样本可能产生偏见,因为它只包含恰好在那个房间的 Year 8 学生,无法代表全校。一个更好的方法是简单随机抽样:把所有学生的名字放进帽子,抽取 30 个名字。这给了每个学生平等的被选中的机会。
In our case study, we used a stratified sampling approach to ensure genders are fairly represented. The school population is roughly 50% boys and 50% girls, so we randomly selected 15 boys and 15 girls from Year 8. This makes the sample more representative. We then handed out the questionnaire and received full responses from all 30 selected students. The collected data forms the backbone of our analysis.
在我们的案例研究中,我们使用了分层抽样方法来确保性别得到公平代表。全校大约 50% 是男生,50% 是女生,因此我们从 Year 8 随机抽取了 15 名男生和 15 名女生。这让样本更具代表性。随后我们发放了问卷,并收到全部 30 名被选学生的完整回答。收集到的数据构成了我们分析的基石。
4. Organising Data with Frequency Tables | 用频数表整理数据
Once the questionnaires are returned, the raw data needs to be organised. For the favourite snack type question, we tally the responses in a frequency table. Here are the results from our 30 students:
一旦问卷回收完毕,原始数据就需要被整理。针对最喜欢的零食类型这一问题,我们用频数表对答案进行计数。以下是来自 30 名学生的结果:
| Favourite Snack Type | Tally | Frequency |
|---|---|---|
| Fruit | |||| | 4 |
| Crisps | |||| |||| | 10 |
| Chocolate | |||| ||| | 8 |
| Biscuits | |||| | 5 |
| Other | ||| | 3 |
This frequency table gives a clear summary: crisps are the most popular, chosen by 10 out of 30 students. Fruit is chosen by only 4 students. The ‘other’ category might include items like yoghurt or cereal bars. The total frequency is 30, as expected. We can also add a relative frequency column to show percentages.
这个频数表给出了清晰的概括:薯片最受欢迎,30 名学生中有 10 人选它。水果仅有 4 名学生选择。“其他”类别可能包括酸奶或谷物棒等。总频数正如预期的 30。我们还可以添加一列相对频数,用来展示百分比。
For the numerical data ‘number of snacks bought per week’, we have 30 values: 2, 5, 3, 0, 4, 6, 2, 1, 3, 4, 5, 2, 3, 7, 1, 4, 3, 0, 2, 5, 3, 4, 1, 6, 2, 3, 4, 2, 5, 3. To handle this data more effectively, we create a grouped frequency table using intervals of 0-1, 2-3, 4-5, 6-7 snacks.
对于数值数据“每周购买零食数量”,我们有 30 个值:2, 5, 3, 0, 4, 6, 2, 1, 3, 4, 5, 2, 3, 7, 1, 4, 3, 0, 2, 5, 3, 4, 1, 6, 2, 3, 4, 2, 5, 3。为了更有效地处理这些数据,我们使用 0-1、2-3、4-5、6-7 个零食的区间创建了一个分组频数表。
| Snacks per week | Tally | Frequency |
|---|---|---|
| 0 – 1 | |||| | 5 |
| 2 – 3 | |||| |||| || | 13 |
| 4 – 5 | |||| || | 8 |
| 6 – 7 | ||| | 4 |
Grouped frequency tables help us see patterns quickly. Most students buy between 2 and 5 snacks weekly, while very few purchase 6 or more. Only five students buy 0–1 snacks. This organisation prepares the data for drawing graphs and calculating statistics.
分组频数表有助于我们快速发现模式。大多数学生每周购买 2 到 5 个零食,极少数购买 6 个或更多。只有五名学生购买 0–1 个零食。这种整理方式为绘制图形和计算统计量做好了准备。
5. Visualising Data: Bar Charts and Pie Charts | 数据可视化:条形图和饼图
Graphs turn numbers into pictures that are easy to compare. For categorical data like ‘favourite snack type’, a bar chart is ideal. We draw bars with equal width spaced apart. The height represents the frequency. From our table, the bar for Crisps would reach 10, Chocolate 8, Biscuits 5, Fruit 4, and Other 3. The bar chart clearly shows crisps as the favourite snack. Pie charts are also useful to show proportions. To draw a pie chart, we calculate the angle for each sector using the formula:
图表把数字变成易于比较的图画。对于“最喜欢的零食类型”这样的分类数据,条形图非常理想。我们画等宽且彼此间隔的条形。高度表示频数。根据我们的表格,薯片的条形高度为 10,巧克力为 8,饼干为 5,水果为 4,其他为 3。条形图清晰地显示薯片是最受欢迎的零食。饼图也有助于显示比例。要绘制饼图,我们使用以下公式计算每个扇形的角度:
angle = (frequency ÷ total frequency) × 360°
For crisps: (10 ÷ 30) × 360° = 120°. For chocolate: (8 ÷ 30) × 360° = 96°. For biscuits: 60°, fruit: 48°, other: 36°. Summing these gives 360°. The pie chart visually shows crisps occupies one-third of the circle. Both charts tell the same story but in different forms—crisps dominate snack choices.
薯片:(10 ÷ 30) × 360° = 120°。巧克力:(8 ÷ 30) × 360° = 96°。饼干:60°,水果:48°,其他:36°。这些角度之和为 360°。饼图直观地展示出薯片占据了圆的三分之一。两种图表以不同形式讲述同一个故事——薯片主导了零食选择。
For the numerical grouped data, a histogram (with bars touching) would be appropriate for the continuous-like intervals, but for discrete counts we can use a simple bar chart with the groups on the horizontal axis. The tallest bar is for 2–3 snacks (frequency 13), confirming the peak shopping behaviour. Graphs must always have clear titles, labelled axes, and a key if necessary.
对于数值分组数据,直方图(条形相连)适用于类似连续型区间,但对于离散计数,我们可以使用简单的条形图,以组别为横轴。最高的条形是 2–3 个零食(频数 13),确认了购买行为的高峰。图表必须始终包含清晰的标题、带标签的坐标轴,并在必要时使用图例。
6. Measures of Central Tendency: Mean, Median and Mode | 集中趋势度量:平均数、中位数和众数
Graphs give a visual summary, but we also need numerical summaries to describe the typical number of snacks bought. We will calculate the mean, median and mode from the raw data list. The mode is the most frequent value. In our unordered list, the number 3 appears 6 times, 2 appears 5 times, 4 appears 5 times. The mode is 3 snacks, but note the data has several peaks – it is almost bimodal. The mode for favourite snack type is crisps (categorical mode).
图表给出了视觉概括,但我们还需要数值概括来描述典型的零食购买数量。我们将从原始数据列表计算平均数、中位数和众数。众数是出现最频繁的值。在我们未排序的列表中,数字 3 出现 6 次,2 出现 5 次,4 出现 5 次。众数是 3 个零食,但注意数据有几个峰值——它几乎是双峰的。而对于最喜欢的零食类型,众数是薯片(分类众数)。
To find the median, we must first order the 30 values from smallest to largest: 0, 0, 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4, 5, 5, 5, 5, 5, 6, 6, 7. Since n = 30 is even, the median is the average of the 15th and 16th values. The 15th value is 3, and the 16th value is 3, so median = 3 snacks. Exactly half of students buy 3 or fewer snacks, and half buy 3 or more.
要找到中位数,我们必须先把 30 个数值从小到大排序:0, 0, 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4, 5, 5, 5, 5, 5, 6, 6, 7。由于 n = 30 是偶数,中位数是第 15 和第 16 个值的平均数。第 15 个值是 3,第 16 个值也是 3,因此中位数 = 3 个零食。恰好一半学生购买的零食数为 3 个或更少,一半为 3 个或更多。
The mean is calculated by summing all values and dividing by the total number. Sum = 0+0+1+1+1+2+2+2+2+2+3+3+3+3+3+3+3+4+4+4+4+4+5+5+5+5+5+6+6+7 = 105. Then mean = 105 ÷ 30 = 3.5 snacks. We can present this as:
平均数是通过将所有值相加再除以总数计算得出。总和 = 0+0+1+1+1+2+2+2+2+2+3+3+3+3+3+3+3+4+4+4+4+4+5+5+5+5+5+6+6+7 = 105。然后平均数 = 105 ÷ 30 = 3.5 个零食。我们可以这样呈现:
mean = (sum of values) ÷ number of values = 105 ÷ 30 = 3.5
The mean of 3.5 is slightly higher than the median of 3, suggesting a small right skew caused by a few students buying many snacks (the 6s and 7). All three measures tell us that a typical student buys about 3 to 3.5 snacks per week.
平均数 3.5 略高于中位数 3,表明由于少数学生购买较多零食(6 和 7),数据有轻微右偏。这三个度量值都告诉我们,一名典型学生每周大约购买 3 到 3.5 个零食。
7. Measures of Spread: Range and Identifying Outliers | 离散度量:极差与识别离群值
Averages alone do not tell the whole story; we need to know how spread out the data are. The simplest measure of spread is the range: maximum value minus minimum value. From our sorted list, the minimum is 0 and the maximum is 7, so the range = 7 – 0 = 7 snacks. This tells us that snack-buying behaviour varies by up to 7 snacks per week.
仅凭平均数不能反映全貌;我们还需要知道数据有多么分散。最简单的离散量数是极差:最大值减去最小值。从我们排序后的列表来看,最小值是 0,最大值是 7,因此极差 = 7 – 0 = 7 个零食。这告诉我们,学生购买零食的行为每周最多相差 7 个。
While the range is easy to compute, it is sensitive to extreme values. The value 7 might be an outlier, as it lies far from the bulk of the data. A very rough check for outliers uses the rule: an outlier is any value more than 1.5 times the interquartile range (IQR) away from the quartiles. For simplicity, you can calculate the quartiles from the ordered list. The lower quartile Q₁ is the median of the first half (first 15 values): the 8th value is 2. The upper quartile Q₃ is the median of the second half: the 23rd value is 5. IQR = Q₃ – Q₁ = 5 – 2 = 3. Then calculate fences: lower fence = Q₁ – 1.5 × IQR = 2 – 4.5 = –2.5 (no low outliers). Upper fence = Q₃ + 1.5 × IQR = 5 + 4.5 = 9.5. Since 7 is less than 9.5, it is not considered an outlier, though the student buying 7 snacks is unusual. The range remains a useful quick indicator of spread.
虽然极差易于计算,但它对极端值敏感。数值 7 可能是一个离群值,因为它远离数据主体。一个非常粗略的离群值检查法则是:任何超出四分位距(IQR)1.5 倍距离的值均为离群值。简单起见,你可以从有序列表中计算四分位数。下四分位数 Q₁ 是前一半(前 15 个值)的中位数:第 8 个值是 2。上四分位数 Q₃ 是后一半的中位数:第 23 个值是 5。IQR = Q₃ – Q₁ = 5 – 2 = 3。然后计算界限:下界限 = Q₁ – 1.5 × IQR = 2 – 4.5 = –2.5(无低离群值)。上界限 = Q₃ + 1.5 × IQR = 5 + 4.5 = 9.5。因为 7 小于 9.5,它不算离群值,但那位购买 7 个零食的学生仍然不常见。极差依然是衡量离散度的快速指标。
A larger
Published by TutorHao | Year 8 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply