📚 Year 11 Edexcel Statistics: Unit Test Mock Exam Analysis | 模拟试卷解析
This article provides a detailed walkthrough of a mock unit test for the Edexcel GCSE Statistics course, covering key topics such as data representation, measures of central tendency and dispersion, probability, and statistical inference. Each question is broken down with step-by-step solutions and examiner insights to help Year 11 students consolidate their understanding and boost exam confidence.
本文详细解析了一份 Edexcel GCSE 统计课程单元测试模拟卷,涵盖数据表示、集中趋势与分散度量、概率及统计推断等关键主题。每个问题均配有分步解答和考官视角提示,帮助 Year 11 学生巩固理解并提升应试信心。
1. Stem-and-Leaf Diagram & Quartiles | 茎叶图与四分位数
Question: The ordered stem-and-leaf diagram shows the test scores of 15 students. Key: 1|2 means 12. Find the median, lower quartile, upper quartile and interquartile range.
0 | 8 91 | 2 4 5 72 | 0 1 1 3 6 83 | 1 44 | 0
题目:有序茎叶图显示了15名学生的测验成绩。键:1|2 表示 12。求中位数、下四分位数、上四分位数和四分位距。数据:0|8 9;1|2 4 5 7;2|0 1 1 3 6 8;3|1 4;4|0。
First, convert the stem-and-leaf plot into an ordered list. The raw scores are: 8, 9, 12, 14, 15, 17, 20, 21, 21, 23, 26, 28, 31, 34, 40. The total number of values n = 15, so the median position is (n + 1)/2 = 8th value. The 8th value is 21, thus median = 21.
首先将茎叶图转换为有序列表。原始成绩为:8, 9, 12, 14, 15, 17, 20, 21, 21, 23, 26, 28, 31, 34, 40。数据总个数 n = 15,中位数位置为 (n + 1)/2 = 第8个值。第8个值为21,故中位数 = 21。
To find the quartiles, we split the data into two halves, omitting the median. The lower half contains the first 7 values: 8, 9, 12, 14, 15, 17, 20. The median of this lower half (the 4th value) gives Q₁ = 14. The upper half contains the last 7 values: 21, 23, 26, 28, 31, 34, 40. Its median (the 4th value) gives Q₃ = 28.
求四分位数时,将数据分为不含中位数的两半。下半部分包含前7个值:8, 9, 12, 14, 15, 17, 20,其中位数(第4个值)为下四分位数 Q₁ = 14。上半部分包含后7个值:21, 23, 26, 28, 31, 34, 40,其中位数为上四分位数 Q₃ = 28。
The interquartile range IQR = Q₃ – Q₁ = 28 – 14 = 14. This measure describes the spread of the middle 50% of the scores.
四分位距 IQR = Q₃ – Q₁ = 28 – 14 = 14。此度量描述了中部50%成绩的分散程度。
2. Box Plots and Outlier Detection | 箱线图与异常值检测
Question: A box plot is to be drawn from the five-number summary: Minimum = 9, Q₁ = 15, Median = 22, Q₃ = 30, Maximum = 55. Identify any outliers and comment on the skewness.
题目:根据五数概括画箱线图:最小值 = 9,Q₁ = 15,中位数 = 22,Q₃ = 30,最大值 = 55。识别任何异常值并评论分布的偏态。
Outlier boundaries are determined using the IQR. IQR = Q₃ – Q₁ = 30 – 15 = 15. The lower fence = Q₁ – 1.5 × IQR = 15 – 22.5 = –7.5. The upper fence = Q₃ + 1.5 × IQR = 30 + 22.5 = 52.5.
异常值界限通过 IQR 确定。IQR = Q₃ – Q₁ = 30 – 15 = 15。下限 = Q₁ – 1.5 × IQR = 15 – 22.5 = –7.5。上限 = Q₃ + 1.5 × IQR = 30 + 22.5 = 52.5。
Any data point below –7.5 or above 52.5 is an outlier. The minimum 9 is greater than –7.5, so no low outlier. However, the maximum 55 exceeds 52.5, so 55 is a high outlier. It should be plotted as a separate point on the box plot.
任何低于 –7.5 或高于 52.5 的数据点均为异常值。最小值9大于 –7.5,故无低异常值。然而最大值55超过52.5,因此55是一个高异常值,应在箱线图上单独绘制点。
To assess skewness, compare the distances from the median to the quartiles and the whiskers. Here Q₃ – median = 8, median – Q₁ = 7; the upper whisker is stretched by the outlier, while the bulk of the data is fairly symmetric. The presence of a high outlier indicates the distribution is positively skewed (right-skewed).
为评估偏态,比较中位数到四分位数及须线的距离。这里 Q₃ – 中位数 = 8,中位数 – Q₁ = 7;上部须线因异常值而拉长,而数据主体较对称。高异常值的存在表明分布为正偏态(右偏)。
3. Cumulative Frequency Graphs and Percentiles | 累积频率图与百分位数
Question: The frequency table shows the times (minutes) taken by 40 students to complete a puzzle. Draw a cumulative frequency graph and estimate the median and the 90th percentile.
Time (t min): 0 ≤ t < 10 (freq 4); 10 ≤ t < 20 (10); 20 ≤ t < 30 (12); 30 ≤ t < 40 (8); 40 ≤ t < 50 (6).
题目:频率表显示40名学生完成拼图的时间(分钟)。绘制累积频率图,并估算中位数和第90百分位数。时间分组及频率:0≤t<10 (4); 10≤t<20 (10); 20≤t<30 (12); 30≤t<40 (8); 40≤t<50 (6)。
Construct a cumulative frequency column by adding frequencies sequentially: 4, 14, 26, 34, 40. Plot the cumulative frequency against the upper class boundary of each interval (10, 20, 30, 40, 50). Join the points with a smooth curve.
构建累积频率列,依次累加:4, 14, 26, 34, 40。将累积频率对照每个区间的上界(10, 20, 30, 40, 50)描点,并用平滑曲线连接。
The median is the value at the 50th percentile, i.e., cumulative frequency = 20. Drawing a horizontal line from 20 to the curve and down to the time axis gives approximately 22.5 minutes. The 90th percentile corresponds to cumulative frequency = 36 (90% of 40). From the graph, this reads about 43 minutes.
中位数为第50百分位数对应的值,即累积频率 = 20。从纵轴20画水平线至曲线,再垂直下至时间轴,约读得22.5分钟。第90百分位数对应累积频率 = 36(40的90%),从图上读得约为43分钟。
Always check that the curve starts at (0,0) and endpoints match the total frequency. These estimates depend on the smoothness of the curve; examiners allow a small tolerance.
务必检查曲线始于 (0,0) 且终点匹配总频数。这些估算值依赖于曲线的平滑度;考官允许小幅容差。
4. Histograms and Frequency Density | 直方图与频率密度
Question: The table gives the lengths of phone calls. Draw a histogram and identify the modal class.
Length (mins): 0–4 (freq 10), 5–9 (16), 10–19 (20), 20–29 (12). (Note: boundaries are 0–4.5, 4.5–9.5, 9.5–19.5, 19.5–29.5 after correcting for gaps.)
题目:表格给出电话通话时长。绘制直方图并确定众数组。时长(分钟): 0–4 (freq 10), 5–9 (16), 10–19 (20), 20–29 (12)。(注意:修正间隔后边界为 0–4.5, 4.5–9.5, 9.5–19.5, 19.5–29.5)
Since class widths are unequal, we must calculate frequency density = frequency ÷ class width. Class widths: 4.5, 5, 10, 10. Frequency densities: 10/4.5 ≈ 2.22, 16/5 = 3.2, 20/10 = 2.0, 12/10 = 1.2.
由于组距不等,须计算频率密度 = 频率 ÷ 组距。组距分别为:4.5, 5, 10, 10。频率密度依次为:10/4.5 ≈ 2.22, 16/5 = 3.2, 20/10 = 2.0, 12/10 = 1.2。
The histogram is drawn with frequency density on the vertical axis and the continuous time scale on the horizontal axis. The modal class is the class with the highest frequency density, which is 5–9 minutes (fd = 3.2). Note that the modal class is not necessarily the class with the highest frequency when widths differ.
直方图以频率密度为纵轴,连续时间尺度为横轴。众数组是频率密度最高的组,即 5–9 分钟 (fd = 3.2)。注意当组距不等时,众数组未必是频率最高的组。
Always label axes with units and give the histogram an appropriate title. The area of each bar represents the frequency of that class.
轴须标注单位,并为直方图加上合适标题。每个条形的面积代表该组的频率。
5. Mean and Standard Deviation | 平均数与标准差
Question: Calculate the mean and the population standard deviation for the dataset: 12, 15, 16, 18, 20, 21.
题目:计算数据集 12, 15, 16, 18, 20, 21 的平均数和总体标准差。
First, find the mean (μ). Sum = 12 + 15 + 16 + 18 + 20 + 21 = 102. Number of values N = 6, so μ = 102/6 = 17.
首先求平均数 (μ)。总和 = 12 + 15 + 16 + 18 + 20 + 21 = 102。数据个数 N = 6,故 μ = 102/6 = 17。
Next, calculate the squared deviations from the mean: (12 – 17)² = 25, (15 – 17)² = 4, (16 – 17)² = 1, (18 – 17)² = 1, (20 – 17)² = 9, (21 – 17)² = 16. Sum of squares = 25 + 4 + 1 + 1 + 9 + 16 = 56.
其次,计算与平均数的离差平方:(12 – 17)² = 25, (15 – 17)² = 4, (16 – 17)² = 1, (18 – 17)² = 1, (20 – 17)² = 9, (21 – 17)² = 16。平方和 = 25 + 4 + 1 + 1 + 9 + 16 = 56。
Population standard deviation σ = √(Σ(x – μ)² / N) = √(56 / 6) = √9.333… ≈ 3.055 (to 3 d.p.). Thus, the mean is 17 and the standard deviation is 3.055.
总体标准差 σ = √(Σ(x – μ)² / N) = √(56 / 6) = √9.333… ≈ 3.055(保留三位小数)。因此平均数为17,标准差为3.055。
6. Scatter Graphs and Regression Lines | 散点图与回归线
Question: The table shows study hours (x) and test scores (y). Plot a scatter graph, describe the correlation, and use the regression line y = 6.17x + 38.2 to predict the score for 10 hours of study.
x: 2, 3, 5, 6, 8, 9; y: 50, 55, 65, 70, 80, 85.
题目:表格显示学习小时数 (x) 与测验分数 (y)。绘制散点图,描述相关性,并使用回归线 y = 6.17x + 38.2 预测学习10小时的分数。x: 2, 3, 5, 6, 8, 9; y: 50, 55, 65, 70, 80, 85。
Plot the points (2,50), (3,55), (5,65), (6,70), (8,80), (9,85). The pattern shows a strong positive linear correlation because as study hours increase, the test scores consistently rise in a linear
Published by TutorHao | Year 11 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导