Year 10 Eduqas Statistics Unit Test Mock Analysis | Eduqas 统计单元测试模拟卷解析

📚 Year 10 Eduqas Statistics Unit Test Mock Analysis | Eduqas 统计单元测试模拟卷解析

Welcome to this in-depth walkthrough of a Year 10 Eduqas Statistics Unit Test mock paper. The following analysis breaks down ten typical questions, each targeting a core skill from the Eduqas GCSE Statistics specification. We will explore not only the correct solutions but also the reasoning behind each step, common pitfalls, and examiner tips. Whether you are reviewing your performance or preparing for the real assessment, this article will help you consolidate key concepts and boost your confidence.

欢迎来到这篇针对 Year 10 Eduqas 统计单元测试模拟卷的深度解析。下文将拆解十道典型试题,每道题都对应 Eduqas GCSE 统计大纲中的核心技能。我们不仅会给出正确答案,还会讲解每一步的推理过程、常见错误以及考官提示。不论你是在回顾自己的表现,还是在为正式测评做准备,本文都将帮助你巩固关键概念并增强自信。

1. Identifying Data Types | 识别数据类型

A local council records the number of library books borrowed by each resident in one month. Question: State whether the resulting data is discrete or continuous, and give a reason for your answer.

某地方议会记录了每位居民一个月内从图书馆借书的数量。问题:判断该数据是离散的还是连续的,并说明理由。

The number of books borrowed can only take whole-number values (0, 1, 2, 3, …). It is impossible to borrow a fractional part of a book. Therefore, the data is discrete because it comes from counting.

借书的数量只能取整数值(0, 1, 2, 3, …)。不可能借出小数本书。因此,该数据是离散的,因为它来源于计数。

A common mistake is confusing ‘discrete’ with ‘qualitative’ or assuming that all numerical data is continuous. Remember: if the data is obtained by measuring, it is continuous; if obtained by counting, it is discrete.

常见的错误是将“离散”与“定性”混淆,或者认为所有数值数据都是连续的。请记住:如果数据是通过测量获得的,就是连续的;如果是通过计数获得的,就是离散的。

Examiners often award one mark for the correct type and one for an accurate justification. Using the exact wording ‘it is counted, not measured’ can secure both marks.

考官通常会给正确的数据类型 1 分,给准确的解释 1 分。使用“这是计数而非测量”这样的精确表述可以稳稳拿下两分。


2. Sampling Methods | 抽样方法

A headteacher wants to survey 50 students about homework habits from a school of 800 students. Describe how a stratified sample could be taken using year groups, and explain one advantage of this method.

一位校长希望对一所 800 名学生的学校中的 50 名学生进行关于作业习惯的调查。描述如何按年级进行分层抽样,并说明该方法的一个优点。

First, the population is divided into distinct strata – here, Year 7, Year 8, Year 9, Year 10 and Year 11. The number of students selected from each year group is proportional to its size. For example, if Year 10 contains 200 out of 800 pupils, then the sample from Year 10 is (200/800) × 50 = 12.5, rounded to 13 students.

首先,将总体划分为不同的层——此处为 7 年级、8 年级、9 年级、10 年级和 11 年级。从每个年级选取的学生人数与该年级的规模成比例。例如,如果 10 年级有 200 人(总数 800),则从 10 年级抽取的样本量为 (200/800) × 50 = 12.5,四舍五入为 13 人。

Within each year group, a simple random sample (e.g., using a random number generator on the register) is then taken to avoid bias. The advantage of stratified sampling is that it guarantees proportional representation of each year group, so the sample reflects the structure of the whole school and reduces sampling error.

在每个年级内部,再进行简单随机抽样(例如,利用随机数生成器在名单上抽取),以避免偏差。分层抽样的优点是它能确保每个年级按比例被代表,因此样本能反映全校的结构,并减小抽样误差。

Many students lose marks by forgetting to mention the within-stratum random selection. A full description must include both the proportional allocation and the use of random sampling inside each group.

许多学生因忘记提及层内随机选择而丢分。完整的描述必须同时包括按比例分配和各组内部的随机抽样。


3. Frequency Tables and Histograms | 频数表与直方图

The grouped frequency table shows the lengths of phone calls (in minutes) made by a receptionist: 0 < t ≤ 2 (frequency 8), 2 < t ≤ 5 (15), 5 < t ≤ 10 (10), 10 < t ≤ 20 (7). Draw a histogram to represent these data and comment on the distribution.

下面的分组频数表显示了某接待员通话时长(分钟):0 < t ≤ 2(频数 8),2 < t ≤ 5(15),5 < t ≤ 10(10),10 < t ≤ 20(7)。绘制直方图表示该数据并评论其分布。

Since the class widths are unequal, we must first calculate frequency density: frequency ÷ class width. The intervals give: 0–2 (width 2, density 4.0), 2–5 (width 3, density 5.0), 5–10 (width 5, density 2.0), 10–20 (width 10, density 0.7). The histogram bars are drawn with these densities as heights, and the area of each bar is proportional to frequency.

鉴于组距不相等,我们必须先计算频率密度:频数 ÷ 组距。各区间为:0–2(宽度 2,密度 4.0),2–5(宽度 3,密度 5.0),5–10(宽度 5,密度 2.0),10–20(宽度 10,密度 0.7)。直方图的条形以这些密度为高度绘制,每个条形的面积与频数成正比。

The vertical axis must be labelled ‘Frequency density’ and the horizontal axis ‘Time, t (minutes)’. The bars should be drawn without gaps because the time variable is continuous. Comment: The distribution is positively skewed, with the modal class 2 < t ≤ 5. Most calls last between 2 and 5 minutes, and a few very long calls create a long tail towards the right.

垂直轴必须标注“频率密度”,水平轴标注“时间 t(分钟)”。条形之间不应留空隙,因为时间变量是连续的。评论:分布呈正偏态,众数组为 2 < t ≤ 5。大多数通话时长为 2–5 分钟,少数极长的通话导致右侧出现长尾。


4. Averages and Range | 平均数与极差

The weekly pocket money (£) for ten Year 10 students is: 5, 8, 10, 10, 12, 15, 15, 18, 20, 25. Calculate the mean, median, mode and range. Which average best represents the data? Justify your choice.

十名 10 年级学生每周的零花钱(英镑)如下:5, 8, 10, 10, 12, 15, 15, 18, 20, 25。计算均值、中位数、众数和极差。哪一个平均数最能代表该数据?论证你的选择。

Mean = (5+8+10+10+12+15+15+18+20+25) ÷ 10 = 138 ÷ 10 = £13.80. Median: with ten values, take the average of the 5th and 6th in order – (12+15)/2 = £13.50. Mode: the values 10 and 15 each appear twice, so the data is bimodal (modes £10 and £15). Range = 25 − 5 = £20.

均值 = (5+8+10+10+12+15+15+18+20+25) ÷ 10 = 138 ÷ 10 = £13.80。中位数:共十个数值,取排序后第 5 和第 6 个数的平均值——(12+15)/2 = £13.50。众数:10 和 15 各出现两次,因此数据是双峰的(众数为 £10 和 £15)。极差 = 25 − 5 = £20。

The median is the most representative average here because the highest value (£25) slightly skews the data, pulling the mean above the majority of the values. The median is not affected by this extreme value and lies close to the centre of the dataset.

这里中位数是最具代表性的平均数,因为最大值(£25)轻微地偏斜了数据,将均值拉到了大多数数值之上。中位数不受该极端值的影响,并且接近数据集的中心。

A full-mark response always explains why one average is preferred by referring to the shape of the distribution (positive skew) and the influence of outliers. Simply stating ‘the median is better’ without justification earns only partial credit.

满分答案总是通过提及分布形状(正偏态)和异常值的影响,来解释为何某一种平均数更受青睐。仅说明“中位数更好”而不加论证,只能得到部分分数。


5. Cumulative Frequency and Box Plots | 累积频率与箱线图

Using the same grouped frequency table for call lengths as in Question 3, construct a cumulative frequency table and draw a cumulative frequency curve. Then estimate the median and interquartile range (IQR), and use them to sketch a box plot.

利用第 3 题中通话时长的分组频数表,构建累积频数表并绘制累积频率曲线。然后估计中位数和四分位距(IQR),并利用它们绘制箱线图。

Cumulative frequencies: ≤2 (8), ≤5 (23), ≤10 (33), ≤20 (40). Plot points at upper class boundaries (2,8), (5,23), (10,33), (20,40) and also at (0,0). Join the points with a smooth curve. Median position: 40/2 = 20th value; from the graph, median ≈ 4.2 minutes. Lower quartile: 40/4 = 10th value, Q₁ ≈ 2.5 minutes. Upper quartile: 30th value, Q₃ ≈ 6.8 minutes. IQR = Q₃ − Q₁ ≈ 4.3 minutes.

累积频数:≤2 (8),≤5 (23),≤10 (33),≤20 (40)。在组距上限 (2,8)、(5,23)、(10,33)、(20,40) 以及 (0,0) 处描点。用光滑曲线连接各点。中位数位置:40/2 = 第 20 个值;从图中读取,中位数 ≈ 4.2 分钟。下四分位数:40/4 = 第 10 个值,Q₁ ≈ 2.5 分钟。上四分位数:第 30 个值,Q₃ ≈ 6.8 分钟。IQR = Q₃ − Q₁ ≈ 4.3 分钟。

Box plot: draw a horizontal scale from 0 to 20. Mark the minimum (just above 0), Q₁ (2.5), median (4.2), Q₃ (6.8) and maximum (20, the upper bound of the last class). The box spans from Q₁ to Q₃ with a vertical line at the median; whiskers extend to the minimum and maximum within 1.5×IQR. No outliers in this range.

箱线图:绘制从 0 到 20 的水平标尺。标出最小值(略高于 0)、Q₁ (2.5)、中位数 (4.2)、Q₃ (6.8) 和最大值(20,即最后一组的上限)。矩形框从 Q₁ 到 Q₃,框内中位数处画竖线;触须延伸到最小值和位于 1.5×IQR 范围内的最大值。此范围内无异常值。


6. Scatter Graphs and Correlation | 散点图与相关性

A student records the number of hours spent on social media (x) and the score on a maths revision test (y) for eight friends. The data are: (10, 42), (8, 55), (12, 38), (6, 65), (14, 30), (5, 72), (11, 45), (7, 60). Plot a scatter graph, describe the correlation, and comment on what this suggests about the relationship.

一名学生记录了八位朋友花在社交媒体上的小时数(x)和数学复习测试的分数(y)。数据为:(10, 42), (8, 55), (12, 38), (6, 65), (14, 30), (5, 72), (11, 45), (7, 60)。绘制散点图,描述相关性,并评论这表明了怎样的关系。

Plotting the points reveals a downward trend: as hours on social media increase, test scores tend to decrease. The points lie reasonably close to an imaginary straight line, though not perfectly. We describe the correlation as negative and moderately strong.

描点后呈现出下降趋势:随着社交媒体使用时间的增加,测试分数趋于下降。数据点相当接近一条假想直线,但并非完美。我们可以将这种相关性描述为负相关且中等偏强。

Correlation does not imply causation. The graph suggests an association, but there may be confounding variables, such as total study time or natural ability, that affect both social media use and test performance. A line of best fit can be drawn to make predictions, but extrapolation beyond the data range would be unreliable.

相关性并不意味着因果关系。该图显示了一种关联,但可能存在混杂变量,比如总学习时间或天赋,它们既影响社交媒体使用也影响考试表现。可以绘制最佳拟合线来进行预测,但超出数据范围的推断并不可靠。

Examiners expect the description to include both the direction (negative) and the strength (moderate). Avoid using vague terms like ‘quite good’ – use ‘strong’, ‘moderate’ or ‘weak’ alongside the direction.

考官期望描述中能同时包含方向(负相关)和强度(中等)。避免使用“相当好”这类模糊的说法——应在方向之外使用“强”、“中等”或“弱”等词。


7. Probability from Experimental Data | 基于实验数据的概率

A dice is suspected of being biased. It is thrown 300 times, and the results are: 1 (45), 2 (50), 3 (48), 4 (52), 5 (40), 6 (65). Calculate the experimental probability of rolling a 6. If the dice were fair, how many times would you expect a 6 to appear? Use the data to argue whether the dice is likely biased.

有人怀疑一颗骰子有偏差。投掷 300 次的结果为:1 (45), 2 (50), 3 (48), 4 (52), 5 (40), 6 (65)。计算掷出 6 的实验概率。如果骰子是公平的,你预期 6 会出现多少次?利用数据论证该骰子是否很可能有偏差。

Experimental probability of a 6 = 65/300 = 13/60 ≈ 0.217. For a fair dice, the expected frequency of each face = 300 × (1/6) = 50. The observed frequency for 6 is 65, which is 15 more than expected. The variation may look large, but a formal hypothesis test or a check using 2×expected range could be used.

掷出 6 的实验概率 = 65/300 = 13/60 ≈ 0.217。对于一颗公平的骰子,每个面的期望频数为 300 × (1/6) = 50。6 的观察频数为 65,比期望值多了 15。这一变异看起来可能很大,但可以使用正式的假设检验或 2×期望值范围来进行判断。

A simple rule of thumb: if an observed frequency differs from the expected frequency by more than about twice the square root of the expected frequency (here √50 ≈ 7.07, so 2×7.07 ≈ 14), we suspect bias. Since 65−50 = 15 > 14, there is some evidence of bias towards rolling a 6. However, a larger number of trials would give a more reliable conclusion.

一个简单的经验法则:如果观察频数与期望频数的差值超过约 2 倍期望频数的平方根(这里 √50 ≈ 7.07,所以 2×7.07 ≈ 14),我们就怀疑存在偏差。由于 65−50 = 15 > 14,存在一些证据表明骰子倾向于掷出 6。然而,更多的试验次数会给出更可靠的结论。


8. Time Series and Moving Averages | 时间序列与移动平均

The quarterly sales (£1000s) of a shop are: Q1 15, Q2 22, Q3 28, Q4 11, Q1 17, Q2 25, Q3 31, Q4 14. Calculate four-point moving averages, plot them on a time series graph, and comment on the trend and any seasonal pattern.

某商店的季度销售额(千英镑)为:Q1 15, Q2 22, Q3 28, Q4 11, Q1 17, Q2 25, Q3 31, Q4 14。计算四点移动平均,并在时间序列图上绘制它们,并评论趋势和任何季节性模式。

Four-point moving averages are calculated by averaging successive groups of four quarters, then centering. For the first set (15+22+28+11)/4 = 19.0; second set (22+28+11+17)/4 = 19.5; third (28+11+17+25)/4 = 20.25; fourth (11+17+25+31)/4 = 21.0; fifth (17+25+31+14)/4 = 21.75. These are plotted at the midpoints of each group (between Q2 and Q3, Q3 and Q4, etc), giving a smooth trend line.

四点移动平均的计算方法是先对连续的四个季度求平均,然后居中。第一组 (15+22+28+11)/4 = 19.0;第二组 (22+28+11+17)/4 = 19.5;第三组 (28+11+17+25)/4 = 20.25;第四组 (11+17+25+31)/4 = 21.0;第五组 (17+25+31+14)/4 = 21.75。这些值被绘制在每组的中间位置(比如 Q2 与 Q3 之间、Q3 与 Q4 之间,依此类推),形成一条平滑的趋势线。

The moving averages show a gentle upward trend in sales over the two years. The raw data reveals a clear seasonal pattern: sales peak in Q3 and dip dramatically in Q4 each year, likely reflecting typical buying habits. This seasonal variation can be quantified by calculating seasonal effects once the trend is removed.

移动平均线显示销售额在两年间呈温和的上升趋势。原始数据展现出明显的季节性模式:每年 Q3 销售额达到顶峰,而在 Q4 急剧下降,这很可能反映了典型的消费习惯。在消除趋势项后,可以通过计算季节效应来量化这一季节性波动。


9. Comparing Distributions Using Summary Statistics | 利用汇总统计量比较分布

Two different revision apps are trialled. Group A (n=25) used App1 and scored a mean of 72 marks with a standard deviation of 8. Group B (n=25) used App2 and scored a mean of 68 marks with a standard deviation of 5. Compare the performance of the two groups, making specific reference to both measures.

两种不同的复习应用软件被试用。A 组(n=25)使用 App1,平均分 72,标准差 8。B 组(n=25)使用 App2,平均分 68,标准差 5。比较两组的成绩,特别提及这两个指标。

On average, Group A achieved higher marks (mean 72 vs 68), suggesting App1 may have been more effective in boosting test performance by 4 marks. However, we must also consider the variability. Group A’s standard deviation of 8 indicates greater spread – scores varied more widely around the mean. Group B, with a smaller standard deviation of 5, was more consistent, meaning most students scored closer to the average of 68.

平均而言,A 组取得了更高的分数(均值 72 对 68),这表明 App1 可能在提高测试表现方面更为有效,平均高出 4 分。然而,我们还必须考虑离散程度。A 组的标准差为 8,表明其成绩分布更广——分数围绕均值的波动更大。B 组的标准差较小,为 5,其成绩更为一致,也就是说大多数学生的分数更接近 68 的平均分。

If consistency is valued (e.g., all students reaching a pass mark), App2 may be preferable despite its lower mean. A statistical test like a two-sample t-test would be needed to determine if the difference in means is significant, but this simple comparison already highlights the trade-off between central tendency and spread.

如果更看重一致性(例如所有学生都能达到及格线),那么尽管 App2 均值较低,却可能更为可取。需要进行像双样本 t 检验这样的统计检验,来判断均值差异是否显著,但这一简单比较已经凸显了集中趋势与离散度之间的权衡。


10. Designing a Statistical Investigation | 设计统计调查

A student wants to investigate whether Year 10 pupils who eat breakfast perform better in morning lessons. Write a short plan outlining the data to collect, how to ensure reliability, and how to present and analyse the findings.

一名学生想调查吃早餐的 10 年级学生是否在上午的课堂上表现更好。撰写一份简短的计划,概述要收集的数据、如何确保可靠性,以及如何呈现和分析结果。

Data to collect: (1) a ‘breakfast frequency’ questionnaire – how many days per week the pupil eats breakfast before school; (2) performance measure – average test scores in morning subjects or teacher ratings on a scale of 1–5 for engagement and focus. The target population is Year 10 students, and a stratified sample by tutor group could be used to ensure representativeness.

要收集的数据:(1) “早餐频率”问卷——学生每周有几天在上学前吃早餐;(2) 表现指标——上午各科的平均测验分数,或者教师对其课堂参与度和专注度的 1–5 评分。目标总体为 10 年级学生,可以按导师小组进行分层抽样以确保代表性。

Reliability: pilot the questionnaire to check that questions are clear; collect performance data from multiple subjects or teachers to avoid bias from a single source; use a large enough sample size (at least 30) to increase reliability; control confounding variables by recording hours of sleep (as it could also affect performance).

可靠性:对问卷进行试测,确保问题清晰;从多门学科或多名教师那里收集表现数据,以避免单一来源的偏差;使用足够大的样本量(至少 30 人)以提高可靠性;通过记录睡眠时长来控制混杂变量(因为睡眠也会影响表现)。

Presentation: a scatter graph of breakfast frequency vs performance, with a line of best fit; summary statistics (mean performance for those eating breakfast 5 days vs 0–2 days). Analysis: calculate Spearman’s rank correlation coefficient to measure the strength of association; discuss any causal limitation and suggest further investigations.

展示方式:散点图(早餐频率 vs. 表现),附上最佳拟合线;汇总统计量(每周吃 5 天早餐的学生与 0–2 天学生的平均表现)。分析:计算斯皮尔曼等级相关系数来衡量关联强度;讨论因果局限并提出进一步调查的建议。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading