Year 9 SQA Statistics: Case Study Practical Exercises | Year 9 SQA 统计:案例分析实战演练

📚 Year 9 SQA Statistics: Case Study Practical Exercises | Year 9 SQA 统计:案例分析实战演练

This article takes you through real-world statistical investigations, exactly as you will meet in SQA Year 9 assessments. We will calculate averages, draw box plots, interpret scatter graphs, and explore probability using genuine data sets. Working through these case studies will sharpen your ability to handle data, choose the right statistical tools, and present clear conclusions.

本文将通过真实的统计调查案例,带你体验 SQA 九年级考试中常见的题型。我们会计算平均数、绘制箱线图、解读散点图,并利用实际数据集探索概率。通过这些案例演练,你将提高处理数据的能力,学会选择合适的统计工具,并能清晰地呈现结论。

1. Case Study 1: Test Scores Dataset | 案例一:考试成绩数据集

Imagine a class of 15 students who sat a maths test. Their scores out of 100 are: 68, 72, 74, 75, 78, 80, 80, 82, 85, 88, 90, 92, 95, 98, 100. We will use this dataset to practise descriptive statistics and graphical displays.

假设一个班有 15 名学生参加了一次数学测验,满分 100 分,他们的成绩为:68, 72, 74, 75, 78, 80, 80, 82, 85, 88, 90, 92, 95, 98, 100。我们将使用这个数据集练习描述性统计和图形展示。

First, let’s find the mean. Add all the scores: 68+72+74+75+78+80+80+82+85+88+90+92+95+98+100 = 1257. Divide the total by the number of scores (15): mean = 1257 ÷ 15 = 83.8. The mean, usually written as x̄, tells us the average performance of the class.

首先,我们来计算平均值。将所有分数相加:68+72+74+75+78+80+80+82+85+88+90+92+95+98+100 = 1257。用总和除以数据个数(15):平均值 = 1257 ÷ 15 = 83.8。平均值通常记作 x̄,它告诉我们这个班的平均表现。

The median is the middle value when all scores are arranged in order. With 15 scores, the 8th value is the median. The sorted list shows the 8th score is 82, so the median is 82. This means half the students scored below 82 and half above.

中位数是将所有分数按顺序排列后处于中间位置的数值。15 个数据中,第 8 个值就是中位数。排序后第 8 个成绩是 82,因此中位数为 82。这意味着有一半学生的分数低于 82,另一半高于 82。

The mode is the most frequent score. In our list, 80 appears twice while all other values appear once, so the mode is 80. The range, the simplest measure of spread, is the difference between the highest and the lowest score: 100 − 68 = 32.

众数是出现次数最多的分数。在我们的数据中,80 出现了两次,其他值只出现一次,因此众数为 80。极差是最简单的离散程度指标,等于最大值与最小值之差:100 − 68 = 32。


2. Quartiles and Interquartile Range (IQR) | 四分位数与四分位距

Box plots require quartiles. For 15 ordered scores, the lower quartile Q₁ is the median of the lower half (first 7 scores). The lower half is: 68, 72, 74, 75, 78, 80, 80; its median is 75, so Q₁ = 75. The upper quartile Q₃ is the median of the upper half (last 7 scores): 85, 88, 90, 92, 95, 98, 100; its median is 92, so Q₃ = 92.

绘制箱线图需要四分位数。对于 15 个已排序的数据,下四分位数 Q₁ 是下半部分(前 7 个数据)的中位数。下半部分:68, 72, 74, 75, 78, 80, 80;其中位数为 75,因此 Q₁ = 75。上四分位数 Q₃ 是上半部分(后 7 个数据)的中位数:85, 88, 90, 92, 95, 98, 100;其中位数为 92,因此 Q₃ = 92。

The interquartile range is IQR = Q₃ − Q₁ = 92 − 75 = 17. The IQR describes the spread of the middle 50% of scores and is not affected by extreme values, unlike the range.

四分位距 IQR = Q₃ − Q₁ = 92 − 75 = 17。四分位距描述了中间 50% 数据的分散程度,且不像极差那样受极端值的影响。

Q₁ = 75, Median = 82, Q₃ = 92, IQR = 17


3. Constructing a Box Plot | 构建箱线图

A box plot (or box-and-whisker diagram) summarizes the five-number summary: minimum = 68, Q₁ = 75, median = 82, Q₃ = 92, maximum = 100. Draw a horizontal scale from about 65 to 105. Above it, draw a box from Q₁ to Q₃ with a vertical line at the median. Extend ‘whiskers’ from the box to the minimum and maximum. This visual instantly shows the central tendency and spread.

箱线图(或盒须图)概括了五数概括法:最小值 = 68, Q₁ = 75, 中位数 = 82, Q₃ = 92, 最大值 = 100。先画一条从大约 65 到 105 的水平刻度轴,然后在轴上方从 Q₁ 到 Q₃ 画一个矩形盒子,并在中位数处画一条竖线。再从盒子两端向最小值和最大值画出“须”。这一图形能立即显示出数据的集中趋势和分散情况。

When interpreting the box plot, note that the right whisker is slightly longer, suggesting a small positive skew, but the median line sits fairly central within the box, indicating a reasonably symmetric interquartile spread.

在解读该箱线图时,我们可以看到右侧的须稍长,表明数据有轻微的正偏态,但中位数线位于盒内相对居中的位置,说明四分位距内的分布大致对称。


4. Stem-and-Leaf Plot | 茎叶图

A stem-and-leaf plot keeps the original data while showing the shape. For our scores, let the tens digit be the ‘stem’ and the units digit be the ‘leaf’. We list stems 6 to 10. The key: 6 | 8 means 68.

茎叶图既能展示数据分布形态,又能保留原始数据。对于我们的成绩,将十位数字作为“茎”,个位数字作为“叶”。列出茎 6 到 10,并注明图例:6 | 8 表示 68。

6 | 8
7 | 2 4 5 8
8 | 0 0 2 5 8
9 | 0 2 5 8
10 | 0
Key: 6|8 = 68

The plot quickly reveals that scores cluster mainly in the 80s, with a single very high score of 100. You can also read off the median as the leaf at the centre.

从茎叶图中可以迅速看出分数主要集中在 80 分段,只有一个很高的分数 100。你也能在图中直接找出中位数,即位于中央位置的叶。


5. Frequency Distribution and Histogram | 频率分布与直方图

When data are grouped into equal-width intervals, we can create a frequency table and draw a histogram. Let’s group the 15 scores into classes of width 10: 60–69, 70–79, 80–89, 90–99, 100–109. Note that 100 falls into the last class. We count how many scores fall into each class.

当数据被分为等宽的区间时,我们可以制作频率表并绘制直方图。我们将这 15 个成绩按组距 10 分组:60–69, 70–79, 80–89, 90–99, 100–109。注意 100 分属于最后一组。我们统计每个组内的数据个数。

Score interval Tally Frequency
60 – 69 | 1
70 – 79 |||| 4
80 – 89 |||| 5
90 – 99 |||| 4
100 – 109 | 1

The histogram would have no gaps between bars because the score scale is continuous. The bar heights represent frequencies, and the area of each bar is proportional to the frequency if class widths are equal.

直方图中条形之间没有间隙,因为分数尺度是连续的。条形高度代表频数,在组距相等的情况下,每个条形的面积与频数成正比。


6. Case Study 2: Scatter Plot and Correlation | 案例二:散点图与相关性

A shop owner recorded daily advertising spend (£) and sales (£) over 7 days. The data are: Spend (x): 10, 15, 20, 25, 30, 35, 40; Sales (y): 220, 250, 300, 340, 390, 430, 480. Plotting these points on a scatter graph with spend on the horizontal axis and sales on the vertical axis reveals a strong positive correlation: as advertising spend increases, sales rise in a linear pattern.

一位店主记录了连续 7 天的每日广告支出(英镑)和销售额(英镑)。数据如下:广告支出 (x):10, 15, 20, 25, 30, 35, 40;销售额 (y):220, 250, 300, 340, 390, 430, 480。将这些点画在散点图上(横轴为广告支出,纵轴为销售额),可以观察到很强的正相关关系:随着广告支出的增加,销售额也以近似线性的方式上升。

We describe this relationship as ‘strong positive correlation’ because the points lie close to a straight line with a positive slope. There are no obvious outliers. Be careful: correlation does not mean causation, but in this context higher advertising may help generate more sales.

我们将这种关系描述为“强正相关”,因为这些点紧密分布在一条具有正斜率的直线附近,且没有明显异常值。注意:相关不代表因果,但在这个背景下,提高广告支出确实可能促进销售额的提升。


7. Line of Best Fit and Prediction | 最佳拟合线与预测

We can draw a line of best fit by eye, passing as close as possible to all points. To make a precise prediction, we can calculate the equation of the line using two well-spaced points on our judged line, say (20, 300) and (35, 430). First find the gradient m: (430 − 300) ÷ (35 − 20) = 130 ÷ 15 = 8.67 (approx). Then using the point-slope form, the equation is: y − 300 = 8.67(x − 20), which simplifies to y = 8.67x + 126.7. You may round to reasonable values, e.g. y = 8.7x + 127.

我们可以靠目测画一条最佳拟合直线,使直线尽可能贴近所有点。为了做出更准确的预测,可以利用所画直线上的两个间距较大的点来计算直线方程,比如取点 (20, 300) 和 (35, 430)。首先计算斜率 m:(430 − 300) ÷ (35 − 20) = 130 ÷ 15 ≈ 8.67。然后利用点斜式得到直线方程:y − 300 = 8.67(x − 20),简化后为 y = 8.67x + 126.7。你可以适当取整,例如写作 y = 8.7x + 127。

Now we can use this equation to predict sales when the shop spends £50 on advertising. Substitute x = 50: y = 8.7 × 50 + 127 = 435 + 127 = 562. So the model predicts approximately £562 in sales. This is an extrapolation beyond the data range, so it should be treated with caution.

现在我们可以用这个方程来预测当广告支出为 50 英镑时的销售额。代入 x = 50:y = 8.7 × 50 + 127 = 435 + 127 = 562。因此模型预测销售额大约为 562 英镑。由于这个预测超出了原有数据范围,属于外推,我们应该谨慎对待。

y = 8.7x + 127; for x = 50, predicted y = 562 (£)


8. Probability: Drawing Coloured Balls from a Bag | 概率:袋中摸球实验

A bag contains 5 red balls, 3 blue balls and 2 green balls. The probability of picking a red ball is P(red) = 5/10 = 1/2, P(blue) = 3/10, and P(green) = 2/10 = 1/5. These are theoretical probabilities.

一个袋子里装有 5 个红球、3 个蓝球和 2 个绿球。抽到红球的理论概率为 P(红) = 5/10 = 1/2,P(蓝) = 3/10,P(绿) = 2/10 = 1/5。

We can simulate an experiment: draw a ball, record its colour, replace it, and repeat 30 times. Suppose we get 12 reds, 11 blues, and 7 greens. The experimental probabilities are: P_red_exp = 12/30 = 0.4, P_blue_exp = 11/30 ≈ 0.367, P_green_exp = 7/30 ≈ 0.233. These are close to the theoretical values but not exactly equal – differences arise because of random variation in small samples.

我们可以进行模拟实验:每次从袋中摸一个球,记录颜色后放回,重复 30 次。假设实验结果得到 12 次红球、11 次蓝球和 7 次绿球。此时实验概率为:P(红) 实验 = 12/30 = 0.4,P(蓝) 实验 = 11/30 ≈ 0.367,P(绿) 实验 = 7/30 ≈ 0.233。这些数值与理论值接近但不完全相等,差异源于小样本中的随机波动。

As the number of trials increases, the experimental probability tends to settle close to the theoretical probability – this is called the law of large numbers.

随着试验次数的增加,实验概率会越来越接近理论概率,这就是大数定律。


9. Tree Diagrams for Combined Events | 复合事件的树状图

Now consider drawing two balls from the same bag without replacement. A tree diagram helps find probabilities of combined outcomes. First draw: branches with probabilities 5/10 (red), 3/10 (blue), 2/10 (green). After one ball is removed, probabilities change. For instance, if the first ball is red, the bag now has 4 red, 3 blue, 2 green (total 9), so the second draw probabilities from the ‘red’ node become 4/9 (red), 3/9 (blue), 2/9 (green).

现在我们考虑从同一个袋子中不放回地连续摸两个球。树状图可以帮助我们计算复合结果的概率。第一次抽取的分支概率为:红 5/10,蓝 3/10,绿 2/10。当一个球被抽走后,概率会发生变化。例如,若第一次抽到红球,袋中还剩 4 红、3 蓝、2 绿(共 9 个球),因此从“红”节点出发的第二抽取概率变为:红 4/9,蓝 3/9,绿 2/9。

To find the probability of getting at least one red ball, we can work out all outcomes that contain at least one red and add their probabilities, or use the complementary approach: 1 − P(no red). The no-red path means first ball not red (5/10) and second ball not red (4/9, since one non-red is removed), so P(no red) = (5/10) × (4/9) = 20/90 = 2/9. Then P(at least one red) = 1 − 2/9 = 7/9.

要计算至少摸到一个红球的概率,我们可以列出所有包含至少一个红球的结果并相加,也可以使用补集的方法:1 − P(无红)。无红的路径意味着第一个球不是红球(概率 5/10),第二个球也不是红球(此时已移走一个非红球,概率 4/9),因此 P(无红) = (5/10) × (4/9) = 20/90 = 2/9。于是 P(至少一红) = 1 − 2/9 = 7/9。


10. Comparing Datasets Using Mean and Range | 使用平均数和极差比较数据集

Let’s introduce a second class of 11 students with these test scores: 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105. Their mean = (55+60+65+…+105) ÷ 11 = 880 ÷ 11 = 80. The range = 105 − 55 = 50. Compare this to our original class (mean 83.8, range 32).

我们引入另一个班级,11 名学生的测验成绩为:55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105。他们的平均值 = (55+60+65+…+105) ÷ 11 = 880 ÷ 11 = 80。极差 = 105 − 55 = 50。与原始班级(平均值 83.8,极差 32)进行比较。

The original class has a higher average, suggesting stronger performance, and a smaller range, indicating more consistent scores. The second class has a lower average and a wider spread, meaning student performance varies more widely. Using only one measure can be misleading; using both mean and range gives a better picture.

原始班级平均分更高,表明整体表现更好,而且极差更小,说明成绩更集中、更一致。第二个班级平均分较低,且极差更大,意味着学生之间的表现差异更悬殊。仅用一个指标去比较可能产生误导,同时使用平均值和极差可以更全面地了解情况。


11. Interpreting Statistical Graphs in Context | 结合情境解读统计图

Imagine you

Published by TutorHao | Year 9 统计 Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading