📚 Year 10 SQA Statistics: High-Frequency Topics and Common Mistake Analysis | Year 10 SQA 统计:高频考点与易错题分析
Preparing for the SQA National 5 Statistics assessment requires a clear understanding of which topics appear most frequently and where students typically lose marks. This article pairs each high-frequency area with an analysis of common mistakes, helping you revise efficiently and avoid costly errors in the exam.
备考 SQA National 5 统计考试,需要清楚哪些主题出现频率最高,以及考生通常在哪里丢分。本文把每一个高频考点与常见错误分析配对讲解,帮助你高效复习,避免在考试中犯代价高昂的错误。
1. Mean, Median, Mode and Range | 平均数、中位数、众数与极差
These measures of central tendency and spread form the foundation of the statistics course. Frequency tables, grouped data and unusual values all feature in exam questions.
这些集中趋势和离散程度的度量构成了统计课程的基础。频数表、分组数据和异常值都会出现在考题中。
A high-frequency exam task asks you to calculate the mean from a frequency table. Remember to multiply each data value by its frequency, sum those products, then divide by the total frequency. For grouped data, use the midpoint of each class interval.
高频考题是要求你从频数表计算平均数。记住用每个数据值乘以其频数,把这些乘积相加,再除以总频数。对于分组数据,要使用每个组区间的中点值。
A common mistake is confusing the position of the median in a list. When data are ordered, the median is at the (n+1)/2 th position. In a frequency table, find the cumulative frequency that first reaches or exceeds (n+1)/2. Many candidates use n/2 incorrectly.
一个常见错误是混淆中位数在列表中的位置。当数据有序时,中位数位于第(n+1)/2个位置。在频数表中,找到累加频数首次达到或超过(n+1)/2的位置。许多考生错误地使用了n/2。
The range, calculated as highest – lowest, is easy to compute but highly sensitive to outliers. Examiners often set questions where an outlier inflates the range, making the interquartile range a more reliable measure of spread.
极差的计算是最大值减最小值,容易计算但对异常值高度敏感。考官常出题让异常值夸大极差,此时四分位距是更可靠的离散度量。
2. Quartiles and Interquartile Range | 四分位数与四分位距
Quartiles split an ordered data set into four equal parts. The interquartile range (IQR = Q3 – Q1) measures the spread of the middle 50% of data and is unaffected by extreme values.
四分位数将有序数据集分成四个相等部分。四分位距(IQR = Q3 – Q1)测量中间50%数据的离散程度,不受极端值影响。
In the SQA National 5 course, the method for finding quartiles from a list of numbers is standard: after finding the median, take the lower half of data (including the median if n is odd) to find Q1, and the upper half for Q3. Many students forget to include the median when the dataset has an odd number of values, leading to a slightly different Q1 or Q3.
在 SQA National 5 课程中,从数字列表中寻找四分位数的方法是标准做法:找到中位数后,取数据的下半部分(如果n为奇数则包含中位数)找Q1,上半部分找Q3。许多学生忘记在数据个数为奇数时包含中位数,从而导致略有不准确的Q1或Q3。
A typical error is calculating semi‑interquartile range (Q3 – Q1)/2 when the question asks for the IQR. Always read the question carefully: IQR is simply Q3 – Q1.
一个典型错误是当题目要求IQR时计算了半四分位距 (Q3 – Q1)/2。一定要仔细审题:IQR 就是 Q3 – Q1。
From a cumulative frequency diagram, Q1, median and Q3 are read off at 25%, 50% and 75% of the total frequency. Misreading the scale on the graph is a frequent source of lost marks.
从累积频率图中,Q1、中位数和Q3分别在总频数的25%、50%和75%处读取。看错图形比例尺是常见的失分原因。
3. Standard Deviation | 标准差
Standard deviation quantifies how much the data values deviate from the mean. A smaller standard deviation indicates data are tightly clustered around the mean; a larger one shows wider spread.
标准差量化了数据值偏离平均数的程度。标准差越小表明数据越紧密地聚集在均值周围;越大表明离散程度越高。
The formula used in National 5 questions is the population standard deviation:
s = √[ Σ(x – x̄)² / n ]
National 5 题目中使用的公式是总体标准差:
s = √[ Σ(x – x̄)² / n ]
The most damaging mistake happens in the intermediate calculations. Forgetting to square the deviations before summing, or accidentally dividing by n – 1 instead of n, are both common. Always set your work out in columns: x, (x – x̄), (x – x̄)², then total the last column.
最具破坏性的错误发生在中间计算步骤。求和之前忘记把偏差平方,或者错误地除以 n – 1 而不是 n,都很常见。一定要把计算排成列:x、(x – x̄)、(x – x̄)²,然后对最后一列求和。
Another pitfall is using the rounded mean in the formula instead of the exact value. Rounding too early can make the standard deviation inaccurate, especially when the data have small variance. Always keep the mean to at least one more decimal place than the original data.
另一个陷阱是在公式中使用四舍五入后的均值,而不是精确值。过早舍入会使标准差失准,尤其是当数据方差很小时。始终让均值比原始数据多保留至少一位小数。
4. Box Plots | 箱线图
Box plots (or box‑and‑whisker diagrams) summarise a dataset using the five‑number summary: minimum, Q1, median, Q3, maximum. They are excellent for comparing two distributions visually.
箱线图(或箱须图)使用五数总结概括数据集:最小值、Q1、中位数、Q3、最大值。它们在目视比较两个分布时非常出色。
A high‑frequency exam question provides two box plots and asks you to compare median and spread. Always comment on both a measure of central tendency (median) and a measure of spread (IQR or range), and link your comparison to the context, e.g. ‘on average, girls scored higher and showed less variation’.
高频考题给你两个箱线图,要求比较中位数和离散程度。一定要同时评论集中趋势度量(中位数)和离散度量(IQR 或极差),并把比较联系到语境中,比如“平均来说,女生得分更高且变异更小”。
A subtle but common mistake is misidentifying the whisker endpoints. Sometimes a box plot uses 1.5 × IQR beyond the quartiles to identify outliers, and the whiskers extend to the most extreme data point that is not an outlier. If outliers are present, they are plotted as individual points. Candidates often forget to adjust the whiskers in such cases.
一个细微但常见的错误是错误识别须的端点。有时箱线图使用1.5 × IQR 超出四分位数来识别离群值,此时须延伸到不是离群值的最极端数据点。如果有离群值,它们会作为单独点绘出。考生常常忘记在这种情况下调整须的长度。
5. Cumulative Frequency Diagrams | 累积频率图
Cumulative frequency diagrams are used to estimate medians, quartiles and percentiles, and to find how many data values lie above or below a given threshold.
累积频率图用来估算中位数、四分位数和百分位数,并找出有多少数据值低于或高于某个给定阈值。
When plotting the curve, join points with a smooth curve, not straight lines. The points should be plotted at the upper boundary of each class interval. A common exam error is plotting at the midpoint or lower boundary, which shifts the entire curve and produces incorrect readings.
绘制曲线时,要用平滑曲线连接各点,不要用直线。点应标在每个组区间的上边界。常见的考试错误是把点标在中点或下边界,这会移动整条曲线,导致读数错误。
Reading values from the graph requires careful use of a ruler. To find the number of data values less than a certain value, go up from the x‑axis value to the curve, then across to the y‑axis. To find how many are greater than that value, subtract the reading from the total frequency. Many candidates forget the subtraction step and lose marks.
从图中读取数值需要仔细使用直尺。要找小于某一数值的数据个数,从 x 轴上的值向上到曲线,再平移到 y 轴。要找大于该值的个数,用总数减去读数。许多考生忘记减法步骤而失分。
6. Scatter Graphs and Correlation | 散点图与相关性
Scatter graphs display the relationship between two variables. The exam expects you to describe correlation as positive, negative or none, and to comment on its strength (strong, moderate, weak).
散点图展示两个变量之间的关系。考试要求你描述相关性为正、负或无,并评论其强度(强、中等、弱)。
Drawing a line of best fit by eye should balance the points, with roughly equal numbers above and below the line. A rushed, uneven line leads to unreliable predictions. When using the line to estimate one variable from another, interpolation (within the data range) is reliable; extrapolation (outside the range) is not. The exam frequently asks you to explain why an extrapolated estimate may be unreliable.
凭目测画一条最佳拟合线应平衡各点,线上方和线下方的点数大致相等。一条匆忙或不平坦的线会导致预测不可靠。使用该线从一个变量估计另一个变量时,内插(在数据范围内)是可靠的;外推(超出范围)则不然。考试常常要求你解释为什么外推估计可能不可靠。
A classic misconception is interpreting correlation as causation. Even a strong correlation does not prove that one variable causes the other to change. Exam commentary questions often test your ability to recognise this and suggest a lurking variable.
一个经典误区是把相关解释为因果。即使强相关也不能证明一个变量会导致另一个变量变化。考试中的评论题常考查你是否能认识到这一点,并提出一个潜在变量。
7. Probability: Tree Diagrams | 概率:树状图
Tree diagrams are essential for structured probability calculations, especially for combined events. The exam features independent events and conditional probabilities where the outcome of one event affects the probability of the next.
树状图对于结构化的概率计算至关重要,特别是对于复合事件。考试会涉及独立事件以及一个事件的结果影响下一个事件概率的条件概率。
The probabilities on the branches from a single point must sum to 1. A very common slip is writing a branch probability incorrectly after a ‘not’ event – for example, if P(success) = 0.3, P(not success) = 0.7, but candidates may misplace the decimal. Always check the sums at each stage.
从同一个点发出的各分支概率之和必须等于1。一个非常常见的失误是在“非”事件后写错分支概率——例如,若 P(成功)=0.3,P(不成功)=0.7,但考生可能搞错小数点的位置。务必在每一步检查概率之和。
When calculating the probability of two events along a path, multiply the branch probabilities. To find the total probability of an event that occurs along several paths, add those path probabilities. Mixing up multiplication and addition is a persistent error.
当计算一条路径上两个事件的概率时,把各分支概率相乘。要找出沿多条路径发生的某事件的总概率,就把那些路径概率相加。将乘法和加法混淆是一个顽固的错误。
For conditional probabilities, carefully adjust the probabilities on the second set of branches. If an item is not replaced, the denominator changes. Many students forget to update the denominator and treat events as independent when they are not.
对于条件概率,要仔细调整第二组分枝上的概率。如果不放回,分母就会改变。许多学生忘记更新分母,而把不独立的事件当作独立事件处理。
8. Normal Distribution | 正态分布
The normal distribution is a continuous, bell‑shaped curve symmetrical about the mean. In SQA National 5, you are given a table of standard normal probabilities and must calculate probabilities for real‑world contexts.
正态分布是一条连续的钟形曲线,关于均值对称。在 SQA National 5 中,会给你标准正态概率表,你必须计算现实世界语境下的概率。
The key skill is standardising a raw score x to a z‑value: z = (x – μ) / σ. After rounding z to two decimal places, use the table to find P(Z < z). The most common error is using the table for a 'greater than' probability without applying the symmetry rule: P(Z > z) = 1 – P(Z < z). Always sketch a quick diagram.
关键技能是把原始分数 x 标准化为 z 值:z = (x – μ) / σ。把 z 四舍五入到两位小数后,用表格找出 P(Z < z)。最常见的错误是对于“大于”的概率没有运用对称规则:P(Z > z) = 1 – P(Z < z)。一定要快速画个草图。
Another mistake arises when the question involves an interval. For P(a < X < b), compute P(Z < z_b) – P(Z < z_a). Candidates sometimes subtract the z‑scores instead of the probabilities, which gives a meaningless result.
另一个错误出现在涉及区间的问题中。对于 P(a < X < b),应计算 P(Z < z_b) – P(Z < z_a)。考生有时会减去 z 分数而不是概率,这会得到无意义的结果。
9. Sampling Methods and Bias | 抽样方法与偏差
Understanding how data are collected is vital. The course covers random, systematic, stratified, quota and convenience sampling, each with advantages and biases.
理解数据如何收集至关重要。课程涵盖随机、系统、分层、配额和便利抽样,各有优缺点和偏差。
In stratified sampling, the sample size from each stratum should be proportional to the stratum size in the population. The formula (stratum size ÷ population size) × total sample size is assessed regularly. A frequent slip is using the wrong total population or forgetting to round the calculated sub‑sample sizes.
在分层抽样中,每层抽取的样本量应与该层在总体中的大小成比例。公式(层大小 ÷ 总体大小)× 总样本量 经常被考查。常见的失误是用错总体大小,或忘记对计算出的子样本量进行舍入。
Exam questions often describe a sampling scenario and ask you to identify the source of bias. For example, a convenience sample of friends may not represent the wider population. Returning to the definition – that a sampling method is biased if it systematically favours certain outcomes – helps structure your written answers.
考题常常描述一个抽样情境,要求你识别偏差来源。例如,一个仅由朋友组成的便利样本可能不代表更广泛的总体。回到定义——如果一个抽样方法系统地偏向某些结果,它就是有偏的——这有助于组织你的书面回答。
10. Comparing Data Sets: Interpretations and Misconceptions | 数据集比较:解释与常见误读
One of the most valuable statistical skills is making valid comparisons between two or more distributions. The exam expects you to link numbers to context, not just quote statistics.
最有价值的统计技能之一是在两个或多个分布之间进行有效的比较。考试期望你把数字与语境联系起来,而不只是引用统计量。
When comparing, always give a measure of average (mean or median) and a measure of spread (range, IQR or standard deviation). For instance, ‘The median waiting time at clinic A is 8 minutes, which is lower than clinic B’s 12 minutes, and the IQR of 3 minutes at clinic A is smaller, indicating more consistent service.’ Avoid saying ‘clinic A is better’ without evidence.
比较时,一定要给出平均数的度量(均值或中位数)和离散程度的度量(极差、IQR 或标准差)。例如,“诊所A的等待时间中位数是8分钟,低于诊所B的12分钟,而且诊所A的IQR为3分钟更小,表明服务更一致。” 避免在没有证据的情况下说“诊所A更好”。
A common error is using the range when outliers are present. If a dataset contains an extreme value, the range becomes misleading. The IQR is then the appropriate spread to compare. Also, do not use standard deviation to compare distributions unless both are roughly symmetric; otherwise, the IQR is safer.
一个常见错误是在存在离群值时使用极差。如果数据集包含极端值,极差就会产生误导。这时 IQR 是比较离散程度的合适度量。此外,除非两个分布都大致对称,否则不要用标准差来比较分布;不然 IQR 更为安全。
11. Common Mistakes in Calculation | 计算中的常见错误
Even when a concept is understood, arithmetic slips can cost many marks. The most frequent computational errors centre on order of operations, sign errors and misreading tables.
即使理解了概念,算术失误也可能丢掉很多分数。最常见的计算错误集中在运算顺序、符号错误以及看错表格上。
For standard deviation, a classic mistake is to subtract first and then square, but do it inconsistently. Write each step carefully, and double‑check the sum of squared deviations. It is easy to miss a negative sign when subtracting the mean from a value smaller than the mean, resulting in a squared deviation that is too small.
对于标准差,一个经典错误是先减后平方,但过程不一致。仔细写下每一步,并复核偏差平方和。从一个小于均值的值减去均值时,很容易遗漏负号,导致平方偏差过小。
When using a cumulative frequency table, ensure you correctly identify the class boundaries. For a class 10‑19, the upper boundary is 19.5 if data are continuous. Using 19 instead will misplace points on the graph. Similarly, misreading the normal distribution table by one decimal place can turn a correct answer into a wrong one; always trace with a finger or ruler.
使用累积频率表时,确保正确识别组边界。对于组 10‑19,如果数据连续,上边界是 19.5。使用 19 会使图上的点错位。同样,看错正态分布表一位小数,就可能把正确答案变成错误答案;始终用手指或直尺进行追踪。
12. Exam Tips: Checking Your Answers | 考试技巧:检查答案
Leaving time to review your work can catch many of the mistakes described above. Develop a systematic checking routine for the statistics paper.
留出时间检查作业可以抓到上述许多错误。为统计试卷建立一套系统的检查流程。
First, verify that your answers make sense in the context. If you calculate a mean height of 250 cm for a group of teenagers, you probably made a data entry error. Probability answers must lie between 0 and 1; a probability of 1.2 is impossible.
首先,验证你的答案在语境中是否合理。如果你为一组青少年算出身高均值为 250 cm,很可能数据录入有误。概率答案必须在 0 和 1 之间;概率 1.2 是不可能的。
Second, recalculate one or two key values using a different method if possible. For the mean from a frequency table, check the total fx and the total frequency separately. For normal distribution, see if the symmetry property gives a quick check: P(Z < -z) = 1 – P(Z < z) should hold.
其次,如果可能,用另一种方法重新计算一两个关键值。对于从频数表得出的均值,单独检查总 fx 和总频数。对于正态分布,看对称性能否提供快速检查:P(Z < -z) = 1 – P(Z < z) 应当成立。
Finally, ensure all graph axes are labelled and scaled properly. A box plot without a scale or a cumulative frequency graph drawn with a ruler and no smooth curve often lose presentation marks. Even a quick visual check can save 2 or 3 marks across the paper.
最后,确保所有图形坐标轴都有标注并正确标尺。没有标尺的箱线图,或用直尺画成折线而非平滑曲线的累积频率图,通常会损失呈现分。即便快速目视检查,也能在全卷中挽回 2 到 3 分。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导