📚 Year 10 CAIE Statistics: High-Frequency Exam Topics and Common Mistakes | Year 10 CAIE 统计:高频考点与易错题分析
Statistics in the Year 10 CAIE syllabus tests your ability to collect, present, and interpret data. While many concepts appear simple, candidates consistently lose marks on specific areas like histogram scaling, standard deviation calculations, and conditional probability. This article walks you through the most frequently examined topics, pinpoints common errors, and provides clear bilingual explanations to help you avoid these costly mistakes and approach your exam with confidence.
CAIE Year 10 统计课程考察你收集、展示和解读数据的能力。虽然许多概念看似简单,但考生总在直方图刻度、标准差计算和条件概率等特定领域失分。本文带你梳理最高频的考点,指出常见错误,并提供清晰的双语解释,帮你避开这些代价高昂的失分点,自信迎考。
1. Histograms and Frequency Density | 直方图与频率密度
Histograms are examined almost every series. The key point is that when class widths are unequal, the vertical axis must show frequency density, not frequency. Many candidates simply plot frequency, producing a completely misleading shape.
直方图几乎每套试卷都会考查。关键点在于:当组距不相等时,纵轴必须表示频率密度,而非频率。许多考生直接绘制频率,导致图形形状完全错误。
Frequency density is calculated as class frequency divided by class width. The area of each bar is proportional to the frequency. Always label the vertical axis ‘Frequency density’ and double‑check that you have used the correct widths before drawing rectangles.
频率密度的计算方法是组频率除以组距。每个条形的面积与频率成正比。务必在纵轴上标注“频率密度”,并在绘制矩形前再次确认使用了正确的组距。
Common mistake: assuming the tallest bar represents the highest frequency. With unequal widths, a narrow bar may look taller but actually contain fewer observations. Interpret the histogram by area, not height.
常见错误:认为最高的条形代表最多的频数。当组距不等时,一个窄条形可能看起来更高,但实际包含的观测值更少。解读直方图要看面积,而不是高度。
2. Cumulative Frequency Curves and Percentiles | 累积频率曲线与百分位数
Cumulative frequency graphs are used to estimate medians, quartiles, and percentiles. The curve is plotted using the upper class boundaries, and the cumulative frequency is always on the vertical axis. A smooth curve is then drawn, not a dot‑to‑dot line.
累积频率图用于估计中位数、四分位数和百分位数。曲线应使用组上限作为横坐标,纵轴总是累积频率。绘制时应连成光滑曲线,而非点对点的折线。
The position for the median is found at n/2 on the cumulative frequency axis, the lower quartile at n/4, and the upper quartile at 3n/4. Some students mistakenly use (n+1)/2 and (n+1)/4, which shifts the position and leads to incorrect values in interpolation.
中位数的位置是在累积频率轴上取 n/2,下四分位数取 n/4,上四分位数取 3n/4。有些学生错误地使用 (n+1)/2 和 (n+1)/4,这会改变位置,导致线性插值后得到的值不准确。
When reading off a value, draw vertical and horizontal construction lines clearly. Any reading without clear guidelines may be penalised. Always state the interpolated value with units and appropriate precision.
读取数值时要清晰地画出竖直和水平辅助线。没有清晰辅助线的读数可能被扣分。务必在插值结果中给出单位和合适的精度。
3. Quartiles, Range and Interquartile Range | 四分位数、极差与四分位距
In a box‑and‑whisker plot, the box spans from Q₁ to Q₃, with the median marked inside. The whiskers extend to the minimum and maximum values, unless the question specifies outliers. The interquartile range (IQR = Q₃ − Q₁) is a measure of spread that ignores extreme values.
在箱线图中,箱体从 Q₁ 延伸到 Q₃,中位数标在箱内。须线延伸到最小值和最大值,除非题目明确规定了异常值。四分位距(IQR = Q₃ − Q₁)是一种忽略极值的离散度量。
A common misinterpretation: the range is not the same as the IQR. The range uses the full extent of the data, while the IQR captures the middle 50%. In skewed distributions, the IQR is a much better summary of spread.
常见误解:极差不同于四分位距。极差用到数据的整个范围,而 IQR 反映中间 50% 的跨度。在偏态分布中,IQR 能更好地概括离散程度。
When comparing two data sets, candidates often mention only the median. A full comparison should comment on both a measure of central tendency and a measure of spread, such as median and IQR, and refer to the context of the data.
比较两组数据时,考生常只提及中位数。一个完整的比较应同时评论集中趋势度量和离散度量,例如中位数和 IQR,并结合数据背景说明。
4. Choosing and Comparing Averages | 选择与比较平均数
The three averages – mean, median, and mode – each have advantages and disadvantages. The mean uses all data but is distorted by outliers. The median is robust to outliers and is preferred for skewed data. The mode is the only average suitable for non‑numerical categorical data.
三种平均数——均值、中位数和众数——各有优缺点。均值用到了所有数据,但易受异常值影响。中位数对异常值稳健,偏态数据下优先选用。众数是唯一适合非数值类别数据的平均数。
Examiners often ask ‘Which average is most appropriate?’. If the data contain outliers, the median is usually better. If the data are symmetric, the mean is often used. For finding the most common category, only the mode works.
考官常问“哪个平均数最合适?”。如果数据包含异常值,中位数通常更好。如果数据对称,常用均值。若要找出最常见的类别,只能使用众数。
Many students lose marks by giving a purely numerical comparison without context. For example, ‘Class A has a higher mean mark, so students performed better on average’ shows understanding of what the mean represents in context.
许多学生因只给出纯数字比较而缺乏背景而失分。例如,“A 班均分更高,因此学生平均表现更好”表明理解了均值在背景中的含义。
5. Standard Deviation: Key Formulas and Pitfalls | 标准差:关键公式与常见陷阱
Standard deviation measures the spread of data around the mean. The sample standard deviation s uses n − 1 in the denominator, while the population standard deviation σ uses n. Always check the question’s wording: ‘sample’ implies n − 1, ‘population’ implies n. If unsure, CAIE often expects the sample formula unless stated otherwise.
标准差衡量数据围绕均值的分散程度。样本标准差 s 的分母用 n−1,而总体标准差 σ 用 n。务必检查题目措辞:“样本”暗示 n−1,“总体”则用 n。如不确定,CAIE 若无特别说明通常期望样本公式。
s = √( Σ(x − x̄)² / (n − 1) )
The most frequent arithmetic mistakes include forgetting to square the deviations, dividing by n instead of n−1, and entering the data incorrectly into a calculator. Always jot down intermediate steps: find x̄, subtract from each value, square, sum, then divide before taking the square root.
最常见的运算错误包括忘记将偏差平方、错误地除以 n 而非 n−1,以及计算器输入数据有误。务必写下中间步骤:求 x̄、每个值减去均值、平方、求和、相除,最后开方。
When using the alternative formula s = √( (Σx² − (Σx)²/n) / (n−1) ), ensure you do not forget to square the total Σx before dividing by n. Many candidates obtain negative values under the square root due to this slip.
在使用替代公式 s = √( (Σx² − (Σx)²/n) / (n−1) ) 时,不要忘记先将 Σx 平方再除以 n。许多考生因此失误,导致根号下出现负值。
6. Probability Trees and ‘At Least One’ Questions | 概率树与“至少一个”问题
Tree diagrams provide a structured way to handle sequential events. Each branch must show a probability, and the sum of probabilities from each node equals 1. Multiply along branches and add the outcomes of different paths to find the probability of an event.
树状图为处理序贯事件提供了结构化方法。每条分支必须标出概率,每个节点射出的概率之和为 1。沿分支相乘,再对不同路径的结果相加,即得事件的概率。
A classic pitfall occurs with ‘at least one’ problems. Direct calculation can be tedious; the complement rule is far more efficient: P(at least one) = 1 − P(none). Yet many candidates attempt to add paths, miscounting duplicate scenarios.
“至少一个”问题中常见陷阱。直接计算可能很繁琐;利用补集规则高效得多:P(至少一个) = 1 − P(零个)。然而许多考生尝试直接相加路径,因重复计数而出错。
For questions involving ‘without replacement’, remember that the probabilities on the second set of branches change. For example, after removing one red ball, the fraction of red balls remaining decreases. Forgetting to adjust these probabilities is a very common error.
对于“不放回”的问题,记住第二组分支上的概率会改变。例如,拿走一个红球后,剩余红球的比例减小。忘记调整这些概率是非常常见的错误。
7. Scatter Diagrams and Correlation | 散点图与相关性
Scatter graphs illustrate the relationship between two variables. The line of best fit (also called a trend line) should pass through the mean point (x̄, ȳ) and have roughly half the points on each side. Do not force the line through the origin unless there is a valid reason.
散点图展示两个变量间的关系。最佳拟合线(也称趋势线)应经过均值点 (x̄, ȳ),并使两侧点数大致相等。除非有正当理由,不要强制直线通过原点。
Describing correlation requires a comment on strength (strong, moderate, weak) and direction (positive or negative). ‘As the temperature increases, the number of ice creams sold increases, showing a strong positive correlation’ links the variables to the context.
描述相关性需要评论强度(强、中等、弱)和方向(正或负)。“随着气温升高,冰淇淋销量增加,呈现强正相关”将变量与上下文联系起来。
Two critical misunderstandings: correlation does not imply causation, and extrapolating beyond the data range is unreliable. A line of best fit can be used to estimate values within the data range (interpolation), but predictions far outside are not trustworthy.
两个关键误区:相关不代表因果,以及超出数据范围的外推不可靠。最佳拟合线可用于估计数据范围内的值(内插),但远距离的预测不可信。
8. Stem‑and‑Leaf Plots (Back‑to‑Back) | 茎叶图(背靠背)
Stem‑and‑leaf diagrams retain the original data while showing shape. In a back‑to‑back plot, leaves for one data set extend left and the other right, sharing a central stem. Always provide a key indicating what the stem and leaves represent (e.g., 3|4 means 34).
茎叶图在展示分布的同时保留原始数据。背靠背图中,一组数据的叶向左延伸,另一组向右,共用居中的茎。务必提供图例说明茎和叶的含义(如 3|4 表示 34)。
Leaves must be ordered. A common exam error is presenting leaves unsorted, which makes it impossible to read the median correctly. After listing, sort each row of leaves from smallest to largest. Also check that no data values are omitted.
叶必须排序。常见考试错误是叶未排序,导致无法正确读取中位数。列出后,将每行的叶从小到大排序。同时检查无数据值遗漏。
From the plot, you can quickly identify the mode (most frequent leaf), median (middle value), and range. Comparing back‑to‑back plots often requires commenting on the shape (skewness) and the typical value, not just stating the medians.
从图中可快速找出众数(最常见叶)、中位数(中间值)和极差。比较背靠背图常需评论分布形状(偏度)和典型值,而不仅仅是给出中位数。
9. Conditional Probability: Misunderstood Concepts | 条件概率:被误解的概念
Conditional probability P(A|B) is the probability of A occurring given that B has happened. The formula is P(A|B) = P(A ∩ B) / P(B). A common mistake is confusing P(A|B) with P(B|A); they are only equal in special symmetric cases.
条件概率 P(A|B) 是在 B 已发生的条件下 A 发生的概率。公式为 P(A|B) = P(A ∩ B) / P(B)。常见
Published by TutorHao | Year 10 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导