📚 Year 10 CIE Statistics Unit Test Mock Paper Analysis | Year 10 CIE 统计单元测试模拟卷解析
This comprehensive walkthrough analyses a typical CIE IGCSE Statistics unit test mock paper, designed for Year 10 students. Each question is broken down step by step, covering data types, averages, spread, probability and correlation. The aim is to deepen your understanding of key statistical concepts and exam technique.
本文对一份典型的 CIE IGCSE 统计单元测试模拟卷进行深度解析,适合 Year 10 学生。我们逐题分解步骤,涵盖数据类型、平均数、离散程度、概率和相关关系,旨在加深你对统计核心概念和考试技巧的理解。
1. Overview of the Mock Paper | 模拟卷概览
The mock paper consists of 8 questions covering the core Unit 1 topics: data classification, measures of central tendency (mean, median, mode), measures of dispersion (range, interquartile range, variance, standard deviation), basic probability with tree diagrams, cumulative frequency, scatter graphs and comparative analysis. It mirrors the style of CIE paper questions, mixing straightforward calculation with interpretation.
模拟卷包含 8 道题目,覆盖单元一的主题:数据分类、集中趋势量数(平均数、中位数、众数)、离散量数(极差、四分位距、方差、标准差)、概率基础与树状图、累积频数、散点图以及比较分析。卷子模仿 CIE 真题风格,将直接计算与解释说明结合在一起。
We will now examine each question in detail, explaining the correct method, common pitfalls and how marks are allocated. Always read the question carefully and show your working clearly.
现在我们将详细分析每一道题目,解释正确方法、常见错误以及如何分配分数。请始终仔细读题,并清晰展示解题步骤。
2. Question 1: Identifying Data Types | 问题一:识别数据类型
Question: Classify each of the following as qualitative, discrete quantitative or continuous quantitative. (a) Heights of students in cm. (b) Blood types A, B, AB, O. (c) Number of books read in a month.
题目:将下列各项分为定性数据、离散定量数据或连续定量数据。(a) 学生身高(厘米)。(b) 血型 A、B、AB、O。(c) 每月阅读书籍的数量。
Understanding data types is fundamental. Quantitative data are numerical and can be discrete (countable, finite steps) or continuous (measured, can take any value in an interval). Qualitative data are non-numerical categories or attributes.
理解数据类型是基础。定量数据是数值型的,可分为离散(可数,步骤有限)或连续(测量值,可在区间内取任意实数值)。定性数据是非数值的类别或属性。
(a) Height is a measurement that can take any value within a range, making it continuous quantitative. (b) Blood types are labeled categories with no intrinsic numerical order, so qualitative. (c) The number of books is obtained by counting, so it is discrete quantitative.
(a) 身高是可在一段范围内取任意值的度量,因此是连续定量数据。(b) 血型是带标签的类别,无内在数值次序,因此是定性数据。(c) 阅读数量通过计数得到,因此是离散定量数据。
A common mistake is to confuse discrete and continuous data when the variable looks numerical. Ask yourself: could the values take on fractions or decimals meaningfully? For counts, fractions are usually impossible (e.g. 2.3 books does not occur naturally), so it is discrete.
常见的错误是当变量看起来是数值时混淆离散与连续。请自问:该数值能有意义地取分数或小数吗?对计数而言,分数通常不可能发生(例如 2.3 本书不会自然出现),所以是离散的。
3. Question 2: Mean, Median, Mode and Range | 问题二:平均数、中位数、众数与极差
Data set: 12, 15, 13, 10, 12, 14, 12, 16. Find the mean, median, mode and range.
数据集:12, 15, 13, 10, 12, 14, 12, 16。求平均数、中位数、众数和极差。
First, compute the mean: sum all values and divide by the count n = 8. Sum = 12+15+13+10+12+14+12+16 = 104. Mean = 104 ÷ 8 = 13.0.
首先计算平均数:将所有值求和并除以个数 n=8。总和 = 12+15+13+10+12+14+12+16 = 104。平均数 = 104 ÷ 8 = 13.0。
To find the median, arrange the data in ascending order: 10, 12, 12, 12, 13, 14, 15, 16. With an even number of values, the median is the average of the 4th and 5th values: (12 + 13)/2 = 12.5.
求中位数需将数据按升序排列:10, 12, 12, 12, 13, 14, 15, 16。当数据个数为偶数时,中位数是第 4 和第 5 个值的平均:(12+13)/2 = 12.5。
The mode is the most frequent value: 12 appears three times, so mode = 12. The range is the difference between the largest and smallest values: 16 – 10 = 6.
众数是出现频率最高的值:12 出现了三次,因此众数 = 12。极差是最大值与最小值之差:16 – 10 = 6。
Always show your ordering step when finding the median; marks are often granted for the method. An incorrect ordering leads to a wrong median even if the mean is correct.
求中位数时务必展示排序步骤;通常评分会给出方法分。即使平均数算对了,排序错误也会导致中位数出错。
4. Question 3: Frequency Table and Estimated Mean | 问题三:频数表与估计均值
Grouped data: Scores: 0–10, 10–20, 20–30, 30–40 with frequencies 4, 10, 12, 4. Calculate an estimate of the mean score.
分组数据:得分区间:0–10、10–20、20–30、30–40,对应频数分别为 4, 10, 12, 4。计算得分均值的估计值。
For grouped data we use the midpoints of each class as representative values. Midpoints: 5, 15, 25, 35. Multiply each midpoint by its frequency, sum these products, and divide by total frequency.
对分组数据,我们用各组的组中值作为代表值。组中值:5, 15, 25, 35。将每组中值乘以频数,求总和,再除以总频数。
Σf = 4+10+12+4 = 30. Σ(f × midpoint) = 4×5 + 10×15 + 12×25 + 4×35 = 20 + 150 + 300 + 140 = 610. Estimated mean = 610 / 30 ≈ 20.33 (or 20.3).
Σf = 4+10+12+4 = 30。Σ(f × 组中值) = 4×5 + 10×15 + 12×25 + 4×35 = 20 + 150 + 300 + 140 = 610。估计均值 = 610 / 30 ≈ 20.33(或 20.3)。
When drawing a histogram for these data, remember that the class widths are equal (10 units each). The vertical axis should represent frequency density, but with equal widths the heights are proportional to frequencies. Label axes clearly.
绘制这些数据的直方图时,注意组距相等(均为 10 单位)。纵轴应表示频数密度,不过由于组距相等,矩形高度与频数成正比。务必清晰标注坐标轴。
Using midpoints provides an estimate because we assume data are evenly distributed within each interval. The true mean could differ slightly if the raw data are skewed.
使用组中值计算得到的是一个估计值,因为我们假定数据在各区间内均匀分布。若原始数据有偏态,实际均值可能略有不同。
5. Question 4: Cumulative Frequency and Quartiles | 问题四:累积频数与四分位数
Using the grouped frequency from Question 3, construct a cumulative frequency table and use the curve to find the median and interquartile range (IQR).
利用问题三的分组频数,构建累积频数表,并通过累积频数曲线求中位数和四分位距(IQR)。
Cumulative frequencies: up to 10 → 4; up to 20 → 4+10 = 14; up to 30 → 14+12 = 26; up to 40 → 26+4 = 30. Plot upper class boundaries (10,20,30,40) against cumulative frequencies, join with a smooth curve.
累积频数:到 10 为止 → 4;到 20 为止 → 14;到 30 为止 → 26;到 40 为止 → 30。将上组界(10,20,30,40)与累积频数对应描点,用平滑曲线连接。
The median corresponds to the 15th value (total frequency/2). From the curve, the median ≈ 23. The lower quartile (Q1) is at 7.5th value ≈ 15; the upper quartile (Q3) is at 22.5th value ≈ 28. Therefore, IQR = Q3 – Q1 ≈ 28 – 15 = 13.
中位数对应第 15 个值(总频数/2)。从曲线上读取,中位数 ≈ 23。下四分位数(Q1)在第 7.5 个值 ≈ 15;上四分位数(Q3)在第 22.5 个值 ≈ 28。因此,IQR = 28 – 15 = 13。
Your graph must have a title and labelled axes. Reading from a cumulative frequency graph involves locating the target cumulative frequency on the y-axis, drawing a horizontal line to the curve, then dropping vertically to the x-axis.
图表必须有标题和标好轴的标签。从累积频数图读取数据的方法是:在纵轴上定位目标累积频数,画水平线交于曲线,再由交点垂直下落到横轴读数。
IQR is a resistant measure of spread that ignores extreme values, making it useful for skewed distributions.
四分位距是一种不受极端值影响的离散度量,对偏态分布尤为有用。
6. Question 5: Variance and Standard Deviation | 问题五:方差与标准差
Using the dataset from Question 2, calculate the variance and standard deviation. Provide an interpretation.
利用问题二的数据集,计算方差和标准差,并给出解释。
We treat the data as a sample, so we use the formula s² = Σ(x – x̄)² / (n – 1). Mean x̄ = 13. Calculate each deviation, square it, and sum.
我们将数据视为样本,因此使用公式 s² = Σ(x – x̄)² / (n – 1)。平均数 x̄ = 13。计算各偏差、平方,并求和。
Deviations: 12–13 = -1, squared = 1; 15–13 = 2, squared = 4; 13–13 = 0, squared = 0; 10–13 = -3, squared = 9; 12–13 = -1 → 1; 14–13 = 1 → 1; 12–13 = -1 → 1; 16–13 = 3 → 9. Sum of squares = 1+4+0+9+1+1+1+9 = 26.
偏差:12-13 = -1,平方 = 1;15-13 = 2,平方 = 4;13-13 = 0,平方 = 0;10-13 = -3,平方 = 9;12-13 = -1 → 1;14-13 = 1 → 1;12-13 = -1 → 1;16-13 = 3 → 9。平方和 = 26。
Variance s² = 26 / (8–1) = 26 / 7 ≈ 3.714. Standard deviation s = √3.714 ≈ 1.927. The standard deviation measures typical deviation from the mean; here the values are on average about 1.93 units away from the mean 13.
方差 s² = 26 / 7 ≈ 3.714。标准差 s = √3.714 ≈ 1.927。标准差衡量与均值的典型偏离程度;此处数据平均约偏离均值 13 约 1.93 单位。
In CIE exams always state which formula you are using. If the question says “population”, divide by n instead of n–1. Using a clear table saves time and reduces errors.
在 CIE 考试中,总需说明所用公式。如果题目指明“总体”,则除以 n 而非 n-1。使用清晰的表格可节省时间并减少错误。
A small standard deviation indicates data points are clustered tightly around the mean, while a larger one suggests greater variability.
标准差小表明数据点紧密聚集在均值周围,标准差大则表示变异性更强。
7. Question 6: Probability with Tree Diagrams | 问题六:概率与树状图
Scenario: A bag contains 5 red, 3 blue and 2 green balls. A ball is drawn, its colour recorded, and replaced. Then a second ball is drawn. Find the probability that (i) both are red, (ii) at least one is blue.
情景:一个袋子里有 5 个红球、3 个蓝球和 2 个绿球。抽出一个球记录颜色后放回,再抽第二个球。求:(i) 两次都是红球的概率;(ii) 至少一次抽到蓝球的概率。
Total balls = 10. P(Red) = 5/10 = 0.5, P(Blue) = 0.3, P(Green) = 0.2. Since the draw is with replacement, the probabilities remain constant for both trials.
总球数 = 10。P(红) = 0.5、P(蓝) = 0.3、P(绿) = 0.2。由于抽取后放回,两次试验的概率保持不变。
(i) Both red: multiply along the branches: 0.5 × 0.5 = 0.25. (ii) At least one blue: 1 – P(no blue) is easiest. P(no blue) = P(not blue, then not blue) = 0.7 × 0.7 = 0.49, so answer = 1 – 0.49 = 0.51. Alternatively, add P(B, not B) + P(not B, B) + P(B, B) = (0.3×0.7)+(0.7×0.3)+(0.3×0.3)=0.51.
(i) 两次红球:沿分支相乘:0.5 × 0.5 = 0.25。(ii) 至少一次蓝球:最简法是 1 – P(无蓝)。无蓝概率 = P(非蓝且非蓝) = 0.7 × 0.7 = 0.49,故答案为 1 – 0.49 = 0.51。也可用概率加法:(0.3×0.7)+(0.7×0.3)+(0.3×0.3)=0.51。
Drawing a tree diagram with labelled branches and writing probabilities along each branch helps to visualise all outcomes and secure method marks.
绘制树状图,标注分支并沿分支写概率,有助于可视化所有结果,从而稳拿方法分。
8. Question 7: Scatter Graphs and Correlation | 问题七:散点图与相关性
Data: Study hours x: 2,3,5,6,8,9,10,12,14,15; Test scores y: 50,55,65,70,78,80,85,88,92,95. Plot the scatter graph and describe the correlation.
数据:学习时间 x:2,3,5,6,8,9,10,12,14,15;测试成绩 y:50,55,65,70,78,80,85,88,92,95。绘制散点图并描述相关性。
Plot each pair (x, y) on a grid, with hours on the horizontal axis and scores on the vertical axis. You should see an upward trend: as study hours increase, test scores tend to increase. This indicates a positive correlation.
在网格图上绘制每一对 (x, y),横轴为学习时间,纵轴为成绩。你将看到上升趋势:学习时间增加,测试成绩趋于提高,表明存在正相关。
Correlation can be described as strong, moderate or weak, and positive or negative. Here points lie close to an imaginary straight line, suggesting a strong positive correlation.
相关性可描述为强、中等或弱,以及正或负。这里各点紧密靠近一条想象中的直线,说明是强正相关。
Correlation does not imply causation. While longer study hours are associated with higher scores, other factors such as prior knowledge may also be at play.
相关性不意味因果关系。虽然较长的学习时间与较高分数相关,但先前知识等其他因素也可能起作用。
You may be asked to draw a line of best fit by eye and estimate a value. Ensure the line passes through the “heart” of the points with roughly equal numbers above and below.
你可能需要凭目测画出最佳拟合线并估计某值。确保直线穿过点群的“心脏”区域,使直线上方和下方的点数大致相等。
9. Question 8: Comparing Two Data Sets | 问题八:两组数据的比较
Scenario: Boys’ test scores: mean = 70, standard deviation = 8. Girls’ scores: mean = 75, standard deviation = 10. Use appropriate statistics to compare the two groups.
情景:男生测试成绩:平均数 = 70,标准差 = 8。女生测试成绩:平均数 = 75,标准差 = 10。使用合适的统计量比较两组数据。
Girls have a higher average score (75 vs 70). However, the standard deviation for girls is larger, indicating greater variability in girls’ performance. A consistent dataset has a smaller standard deviation relative to the mean.
女生的平均分更高(75 vs 70)。但是女生的标准差更大,表明女生成绩的变异性更强。相对于均值而言,标准差较小表示数据更一致。
To make a fair comparison of consistency, we can use the coefficient of variation (CV) = (standard deviation / mean) × 100%. For boys: CV = (8/70)×100% ≈ 11.43%. For girls: CV = (10/75)×100% ≈ 13.33%. The lower CV shows the boys’ scores are more consistent relative to their mean.
为公平比较一致性,可使用变异系数 (CV) =(标准差 / 平均数)× 100%。男生:CV = (8/70)×100% ≈ 11.43%。女生:CV = (10/75)×100% ≈ 13.33%。CV 更低表明男生成绩相对于其均值更一致。
When writing a comparison, always comment on both a measure of central tendency and a measure of spread. State the values explicitly and then interpret them.
写比较时,总要同时讨论集中趋势量数和离散量数。明确写出数值,然后加以解读。
In real exam answers, you might also mention possible reasons for the difference, but statistical comparison alone suffices.
在实际考试答案中,你或许也可以提及产生差异的可能原因,但仅作统计比较已足够。
10. Exam Tips and Review Summary | 考试技巧与复习总结
Mastering a statistics mock paper requires not only correct calculation but also clear communication. Keep these strategies in mind: read word problems carefully, identify the data type, draw diagrams where helpful (tree diagrams, cumulative frequency curves, scatter plots), and always check if you need sample or population formulas.
掌握统计模拟卷不仅需要正确计算,还需要清晰的表达。记住这些策略:仔细阅读文字题,识别数据类型,在必要时画图(树状图、累积频数曲线、散点图),并始终检查需要样本公式还是总体公式。
Time management is crucial. Spend no more than 1 minute per mark on average, and leave time to review your answers for rounding errors or misread scales. Label your graphs fully — titles, axes, and units earn method marks.
时间管理至关重要。平均每题每分不超过一分钟,并留出时间检查答案是否有舍入错误或刻度误读。充分标注图表——标题、坐标
Published by TutorHao | Year 10 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导