📚 Year 10 AQA Statistics: Unit Test Mock Paper Walkthrough | AQA统计单元测试模拟卷解析
Welcome to this detailed walkthrough of a Year 10 AQA Statistics unit test mock paper. This resource is designed to help you consolidate key concepts from the first year of the GCSE Statistics course, including data types, summary measures, probability, and data representation. Each question is explained step by step, with clear reasoning in both English and Chinese. You can use this to self-assess, fill in any gaps, and build confidence before your real unit test.
欢迎阅读这份 Year 10 AQA 统计单元测试模拟卷的详细解析。本资源旨在帮助你巩固 GCSE 统计课程第一年的核心概念,包括数据类型、概括性度量、概率以及数据呈现。每道题都配有逐步讲解,并提供清晰的中英文解释。你可以利用它进行自评、查漏补缺,并在真正的单元测试前建立信心。
1. Question 1: Types of Data | 第1题:数据类型
The first question asked you to classify the following variables: ‘Type of pet owned’ (e.g., cat, dog, none), ‘Number of siblings’, and ‘Time taken to run 100 m in seconds’. Understanding data types is crucial because it determines which statistical measures and diagrams are appropriate.
第一题要求你对以下变量进行分类:“拥有的宠物类型”(例如猫、狗、无)、“兄弟姐妹的数量”以及“跑100米所需的时间(秒)”。理解数据类型至关重要,因为它决定了哪些统计度量和图表是合适的。
‘Type of pet owned’ is a qualitative (categorical) variable. It describes a quality and cannot be measured numerically. The categories have no inherent order, so it is nominal qualitative data.
“拥有的宠物类型”是定性(分类)变量。它描述一种属性,无法用数值测量。这些类别没有内在顺序,因此它属于名义型定性数据。
‘Number of siblings’ is quantitative discrete data. It takes numerical values that can be counted, and you cannot have half a sibling. It arises from counting, so it is discrete.
“兄弟姐妹的数量”是定量离散数据。它取可计数的数值,你不可能有半个兄弟姐妹。它源于计数,因此是离散的。
‘Time taken to run 100 m’ is quantitative continuous data. Time can take any value within an interval depending on the precision of the stopwatch, so it is measured and not simply counted.
“跑100米所需的时间”是定量连续数据。时间可以在一个区间内取任何值,取决于秒表的精度,因此它是测量得到的,而不仅仅是计数的。
2. Question 2: Interpreting Bar Charts | 第2题:条形图解读
A vertical bar chart showed the frequency of favourite colours among 120 students: Red (40), Blue (35), Green (25), Yellow (20). You were asked to state the mode and to calculate how many more students chose Red than Green.
一个竖直条形图显示了120名学生中最喜欢颜色的频数:红色(40)、蓝色(35)、绿色(25)、黄色(20)。题目要求你指出众数,并计算选择红色的学生比选择绿色的多多少人。
The mode is the category with the highest frequency. Here, Red has the highest bar with a frequency of 40, so the mode is Red.
众数是频数最高的类别。这里,红色的柱子最高,频数为40,因此众数是红色。
To find the difference in student numbers: Red count minus Green count = 40 – 25 = 15. So 15 more students preferred Red.
要找出人数差:红色人数减去绿色人数 = 40 – 25 = 15。因此,选择红色的学生多15人。
You were also reminded that bar charts are used for categorical or discrete data with gaps between bars, unlike histograms which are for continuous data.
题目也提醒你,条形图用于分类或离散数据,柱子之间有空隙,这不同于适用于连续数据的直方图。
3. Question 3: Mean and Median | 第3题:平均数与中位数
A small dataset was given: 12, 15, 18, 22, 22, 25. You had to calculate the mean, median, and mode, and explain why the median might be preferred when outliers are present.
给出一组小数据:12, 15, 18, 22, 22, 25。要求计算平均数、中位数和众数,并解释为什么存在异常值时中位数可能更可取。
First, to find the mean, sum all values and divide by the count.
Mean = (12 + 15 + 18 + 22 + 22 + 25) ÷ 6 = 114 ÷ 6 = 19
首先,求平均数:将所有数值相加,除以个数。
平均数 = (12 + 15 + 18 + 22 + 22 + 25) ÷ 6 = 114 ÷ 6 = 19
The median is the middle value when the data are ordered. With 6 values, the median lies between the 3rd and 4th values: (18 + 22) ÷ 2 = 20.
中位数是将数据排序后中间的值。有6个值,中位数位于第3和第4个值之间:(18 + 22) ÷ 2 = 20。
The mode is the most frequent value, which is 22 (appearing twice).
众数是出现频率最高的值,即22(出现两次)。
The median is often preferred when there are outliers because it is not affected by extreme values, unlike the mean. For example, if the 25 were replaced by 100, the mean would jump to 31.5, but the median would stay at 20, giving a better idea of a typical value.
存在异常值时通常更倾向于使用中位数,因为它不会像平均数那样受极端值影响。例如,如果把25换成100,平均数会跃升至31.5,但中位数仍为20,能更好地反映典型值。
4. Question 4: Range and Interquartile Range | 第4题:极差与四分位数间距
Using the dataset from Question 3 (now with an extra value: 12, 15, 18, 22, 22, 25, 30), you were asked to find the range and the interquartile range (IQR).
使用第3题的数据集(现增加一个值:12, 15, 18, 22, 22, 25, 30),要求找出极差和四分位数间距(IQR)。
The range gives a quick measure of spread but is sensitive to outliers. Range = maximum – minimum = 30 – 12 = 18.
极差能快速衡量分散程度,但对异常值敏感。极差 = 最大值 – 最小值 = 30 – 12 = 18。
The IQR focuses on the spread of the middle 50% of data. For n = 7, the median is the 4th value (22). The lower quartile (Q1) is the median of the lower half: 12, 15, 18 → Q1 = 15. The upper quartile (Q3) is the median of the upper half: 22, 25, 30 → Q3 = 25. Therefore, IQR = Q3 – Q1 = 25 – 15 = 10.
IQR集中反映中间50%数据的分散程度。n = 7时,中位数是第4个值(22)。下四分位数(Q1)是下半部分的中位数:12, 15, 18 → Q1 = 15。上四分位数(Q3)是上半部分的中位数:22, 25, 30 → Q3 = 25。因此,IQR = Q3 – Q1 = 25 – 15 = 10。
This tells us that the middle half of the data lies within a range of 10 units, which is more robust than the full range.
这告诉我们,中间一半的数据落在10个单位的区间内,这比全距更稳健。
5. Question 5: Box Plots | 第5题:箱形图
Using the dataset 12, 15, 18, 22, 22, 25, 30, you were required to construct a box plot and to interpret the shape of the distribution. The five-number summary was: min = 12, Q1 = 15, median = 22, Q3 = 25, max = 30.
利用数据集12, 15, 18, 22, 22, 25, 30,要求绘制箱形图并解释分布的形态。五数综合为:最小值=12, Q1=15, 中位数=22, Q3=25, 最大值=30。
A box plot is drawn using a scaled axis. The box spans from Q1 to Q3, with a line inside at the median. Whiskers extend to the minimum and maximum (assuming no outliers). In this box plot, the median is closer to Q3 than to Q1, indicating that the data might be slightly negatively skewed. However, with a small dataset, we must be cautious. The whiskers also show that the left tail is shorter than the right tail, but the box itself shows the median offset to the right within the box.
箱形图使用带刻度的轴绘制。箱子从Q1到Q3,内部在中位数处画一条线。须线延伸至最小值和最大值(假设无异常值)。在这个箱形图中,中位数更靠近Q3而非Q1,表明数据可能略呈负偏态。不过,由于数据量小,我们得谨慎。须线还显示左尾比右尾短,但箱子本身显示中位数在箱内偏右。
You can describe the skew: when the median is nearer the upper quartile, the data are negatively skewed. This means the bulk of the data is concentrated towards the higher end, with a tail stretching to the lower values.
你可以描述偏态:当中位数更靠近上四分位数时,数据呈负偏态。这意味着大部分数据集中在较高端,尾部延伸至较低值。
6. Question 6: Basic Probability | 第6题:基础概率
A fair six-sided die is rolled and a letter is chosen from the word STATISTICS. You had to find the probability of rolling a number greater than 4 and selecting a vowel.
一个公平的六面骰子被掷出,并从单词 STATISTICS 中随机选取一个字母。要求计算掷出大于4的点数并选中元音字母的概率。
First, identify the outcomes. Die: numbers greater than 4 are 5 and 6. Probability of this event = 2/6 = 1/3.
首先,确定结果。骰子:大于4的点数是5和6。该事件的概率 = 2/6 = 1/3。
Word STATISTICS has 10 letters. Vowels: A, I, I (A appears once, I appears twice). Total vowels = 3. Probability of picking a vowel = 3/10.
单词 STATISTICS 有10个字母。元音字母:A, I, I(A出现一次,I出现两次)。元音总数 = 3。选中元音的概率 = 3/10。
These events are independent, so we multiply the probabilities: (1/3) × (3/10) = 3/30 = 1/10. This is the joint probability.
这些事件相互独立,因此我们将概率相乘:(1/3) × (3/10) = 3/30 = 1/10。这就是联合概率。
You were also asked to write the probability as a decimal (0.1) and a percentage (10%).
题目也要求将概率写成小数(0.1)和百分数(10%)。
7. Question 7: Probability Tree Diagrams | 第7题:概率树图
A bag contains 4 red counters and 6 blue counters. Two counters are drawn at random without replacement. Construct a tree diagram and find the probability that both counters are blue.
一个袋子装有4个红色计数片和6个蓝色计数片。不放回地随机抽取两个计数片。构建树图并求出两个都是蓝色的概率。
The first draw: P(Blue) = 6/10 = 3/5, P(Red) = 4/10 = 2/5. If a blue is taken first, the second draw probabilities change: now 5 blue and 4 red remain, so P(Blue | first Blue) = 5/9, P(Red | first Blue) = 4/9. If a red is taken first, P(Blue | first Red) = 6/9 and P(Red | first Red) = 3/9.
第一次抽取:P(蓝) = 6/10 = 3/5,P(红) = 4/10 = 2/5。如果第一次抽到蓝色,第二次抽取概率改变:剩下5蓝4红,因此 P(蓝 | 第一次蓝) = 5/9,P(红 | 第一次蓝) = 4/9。如果第一次抽到红色,则P(蓝 | 第一次红) = 6/9,P(红 | 第一次红) = 3/9。
To find the probability both are blue, we follow the branch: P(first blue) × P(second blue | first blue) = (6/10) × (5/9) = 30/90 = 1/3.
要计算两个都是蓝色的概率,我们沿着该分支:P(第一次蓝) × P(第二次蓝|第一次蓝) = (6/10) × (5/9) = 30/90 = 1/3。
You also had to show that the sum of all joint probabilities at the end of the diagram equals 1. This verifies your tree is correct.
你还需要展示树图末尾所有联合概率之和等于1。这能验证你的树图正确无误。
8. Question 8: Expected Value | 第8题:期望值
A game costs £2 to play. You roll a fair die. If you roll a 6, you win £10. If you roll a 5, you win £3. Otherwise, you win nothing. Calculate the expected profit (or loss) per game in the long run.
一个游戏参与费用为2英镑。你掷一个公平骰子。如果掷出6,赢得10英镑;如果掷出5,赢得3英镑;否则无奖金。计算长期来看每局游戏的期望盈利(或亏损)。
First, define the net profit for each outcome. Profit for rolling 6 = £10 – £2 = £8. Profit for rolling 5 = £3 – £2 = £1. Profit for rolling 1,2,3,4 = £0 – £2 = -£2 (a loss).
首先,定义每个结果的净盈利。掷出6的盈利 = 10 – 2 = 8英镑。掷出5的盈利 = 3 – 2 = 1英镑。掷出1,2,3,4的盈利 = 0 – 2 = -2英镑(亏损)。
Now, find the probability of each: P(6) = 1/6, P(5) = 1/6, P(1-4) = 4/6 = 2/3.
现在,找出各自的概率:P(6) = 1/6,P(5) = 1/6,P(1-4) = 4/6 = 2/3。
Expected profit = (8 × 1/6) + (1 × 1/6) + (-2 × 2/3) = (8/6) + (1/6) – (4/3) = (9/6) – (4/3) = (3/2) – (4/3) = (9/6) – (8/6) = 1/6 ≈ £0.1667. So the expected profit is about 16.7p per game.
期望盈利 = (8 × 1/6) + (1 × 1/6) + (-2 × 2/3) = (8/6) + (1/6) – (4/3) = (9/6) – (4/3) = (3/2) – (4/3) = (9/6) – (8/6) = 1/6 ≈ 0.1667英镑。因此每局游戏期望盈利约16.7便士。
In the long run, you would expect to gain a small amount each time, so the game favours the player. If the expected value were negative, the game would favour the organiser.
长期来看,你每次预计会获得少量盈利,所以游戏对玩家有利。如果期望值为负,则游戏对组织者有利。
9. Question 9: Sampling Methods | 第9题:抽样方法
This question presented a scenario: a school head wants to investigate the amount of homework done by students. Discuss whether a census or a sample is more appropriate, and describe a simple random sampling method.
本题情景:一位校长想调查学生完成作业量。讨论普查和抽样哪种更合适,并描述一种简单随机抽样方法。
A census would involve surveying every student in the school. This eliminates sampling error and gives a true picture, but it can be time-consuming, costly, and may not be practical if the school is large. A sample is quicker and cheaper, but it may not be fully representative if the sample is biased or too small.
普查需调查学校里的每一个学生。这能消除抽样误差并给出真实情况,但可能耗时、昂贵,若学校规模大则不太实际。抽样更快、更省钱,但如果样本有偏或过小,可能不完全具有代表性。
For a large school, a sample is usually preferred. A simple random sample can be obtained by assigning each student a number and using a random number generator to select, say, 100 students. This gives every student an equal chance of being chosen, reducing selection bias.
对于大学校来说,通常更倾向于抽样。简单随机抽样可以通过给每名学生分配一个编号,然后使用随机数生成器选取,比如100名学生。这使得每名学生有相同的被选中机会,减少了选择偏差。
You should also mention that the sampling frame (an accurate list of all students) is essential. Without it, a simple random sample cannot be properly carried out.
你也应当提到,抽样框(一份所有学生的准确名单)至关重要。没有它,就无法正确实施简单随机抽样。
10. Question 10: Scatter Graphs and Correlation | 第10题:散点图与相关性
The final problem gave a table of data: hours spent revising (x) and test score (y): (1, 35), (2, 45), (3, 50), (4, 60), (5, 70), (6, 85). You were asked to plot a scatter graph, describe the correlation, and draw a line of best fit to estimate a score for 3.5 hours of revision.
最后一题给出一张数据表:复习小时数(x)和测试分数(y):(1, 35), (2, 45), (3, 50), (4, 60), (5, 70), (6, 85)。要求绘制散点图,描述相关性,并画出最佳拟合线,估算复习3.5小时对应的分数。
When plotting, the x-axis represents hours and the y-axis represents scores. The points form an upward trend, indicating a positive correlation: as revision hours increase, test scores tend to increase.
绘图时,x轴表示小时,y轴表示分数。这些点呈现上升趋势,表明正相关:随着复习时间增加,测试分数往往也会提高。
The correlation is strong because the points lie close to a straight line. You can draw a line of best fit by eye, ensuring roughly equal numbers of points above and below the line. It should pass through the mean point (approximately x̄ = 3.5, ȳ = 57.5).
相关性很强,因为这些点紧贴一条直线。你可以凭目测画出最佳拟合线,确保线上下的点数量大致相等。该线应该经过均值点(大约 x̄ = 3.5, ȳ = 57.5)。
To estimate the score for 3.5 hours, read up from 3.5 on the x-axis to the line, then across to the y-axis. The estimate should be around 58 to 60. This is an interpolation because 3.5 lies within the data range, making the prediction more reliable.
要估算3.5小时的分数,从x轴的3.5处向上作垂线与直线相交,再水平读取y轴值。估算结果应在58至60左右。这属于内插,因为3.5落在数据范围内,使得预测更可靠。
Extrapolation beyond the data range (e.g., 10 hours) would be unreliable as the linear trend may not continue indefinitely.
超出数据范围的外推(例如10小时)是不可靠的,因为线性趋势可能不会无限延续。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导