📚 Common Misconceptions in Statistics and How to Correct Them | 统计学的常见误区与纠正方法
Statistics is full of subtle traps that can lead even keen students astray. Many learners in Year 10 confuse mean with median, apply the wrong probability rules, or trust a scatter graph to prove causation. This article picks out the ten most frequent pitfalls in the CIE Statistics syllabus and shows exactly how to avoid them, with clear explanations and real examples so you can answer exam questions with confidence.
统计学里藏着许多微妙的陷阱,即使是用功的学生也容易踩坑。很多 Year 10 的学生会把均值和 median 弄混,用错概率法则,或者轻信散点图能证明因果关系。这篇文章梳理了 CIE 统计大纲里最容易翻车的十个误区,用清楚的解释和真实的例子告诉你如何避开它们,让你在考场上自信作答。
1. Confusing Mean, Median and Mode | 混淆均值、中位数与众数
A very common mistake is reaching for the mean as the default measure of central tendency, even when the data set is skewed or contains outliers. For example, if a class test score list includes one extremely low mark, the mean will be pulled downwards, no longer reflecting the typical performance of the class. Students also often report the mode for numerical data without realising that the mode is much more useful for categorical data, such as favourite colour.
一个非常常见的错误是把均值当作集中趋势的默认度量,哪怕数据是偏斜的或者含有异常值。比如,一个班级的考试成绩如果包含一个极低的分数,均值就会被拉低,不再反映班级的典型表现。学生也经常对数值型数据报告众数,却没有意识到众数对分类数据(比如最喜欢的颜色)才更有用。
The correction is to pause and look at the shape of the distribution. For symmetric data with no outliers, the mean is the best summary; for skewed data, the median is more robust because it sits in the middle and ignores extreme values. The mode should be reserved for categorical data or to identify the most frequent value in a discrete set. Always ask yourself: ‘Is my data symmetric? Are there any unusual values?’ before picking a measure of centre.
纠正的方法是先停下来看看数据的分布形状。对于对称且没有异常值的数据,均值是最好的概括;对于偏斜的数据,median(中位数)更稳健,因为它处在正中间,不受极端值干扰。众数则应该留给分类数据,或者用来识别离散数据中出现最频繁的值。在选定中心度量之前,一定要问自己:“我的数据是对称的吗?有没有异常值?”
2. Misunderstanding Sampling Methods | 误解抽样方法
Students regularly assume that any ‘random’ sample is automatically unbiased, but a random sample can still be biased if it is drawn from an incomplete sampling frame. Another frequent error is confusing stratified sampling with quota sampling, and thinking that systematic sampling involves picking people who ‘look right’ to the researcher.
学生通常认为只要是“随机”样本就一定没有偏倚,但如果抽样框不完整,随机样本仍然可能带有偏倚。另一个常见错误是混淆分层抽样与配额抽样,并且以为系统抽样就是调研员挑选“看起来合适”的人。
To correct this, remember the precise definitions. Simple random sampling gives every member of the population an equal chance of being chosen and requires a full list of the population. Stratified sampling divides the population into distinct groups (strata) and takes a random sample from each in proportion to its size; this guarantees representation. Systematic sampling selects every k-th individual from a list, but can introduce bias if the list has a hidden pattern. A convenience sample, like asking friends, is fast but almost always biased. Describe the method clearly in your answers and explain why it is suitable, or unsuitable, for the given context.
纠正方法是记住准确的定义。简单随机抽样让总体中每个成员被抽到的机会相等,需要一个完整的总体名单。分层抽样把总体划成不同的组(层),然后从每一层按比例随机抽取样本,这能保证代表性。系统抽样是从名单中每隔 k 个选一个人,但如果名单本身有隐藏的规律就可能引入偏倚。便利抽样(比如问朋友)很快,但几乎总是有偏的。在答案里要清楚地描述抽样方法,并解释它为什么适合或不适合所给的情境。
3. Misreading Bar Charts, Histograms and Graph Labels | 误读条形图、直方图与坐标轴标签
A classic trap is treating a histogram like a bar chart and reading the height of each bar as the frequency. In a histogram, it is the area of the bar that represents frequency, so when class intervals have unequal widths, the height alone is misleading. Another common oversight is failing to check the scale on the vertical axis: a chart that does not start at zero can exaggerate differences and lead to incorrect conclusions.
一个经典的陷阱就是把直方图当成条形图,把每个柱子的高度当作频数。在直方图中,表示频数的是面积,所以当组距宽度不相等时,只看高度就会产生误导。另一个常见疏忽是没有检查纵轴的刻度:不从零开始的图表会放大差异,导致错误的结论。
To avoid these errors, first identify the type of graph. Bar charts are for categorical data, with gaps between bars and height showing frequency or value. Histograms are for continuous data, with no gaps, and frequency is proportional to area. Calculate frequency density when needed. Always read the axis labels, check the starting point of the scale, and note the units. If a graph does not have a title or axis labels, point out that it is incomplete.
要避免这些错误,首先弄清楚图表类型。条形图用于分类数据,柱子之间有间隔,高度表示频数或数值。直方图用于连续数据,柱子之间没有间隔,频数与面积成比例。需要时请计算频数密度。一定要阅读坐标轴标签,检查刻度的起点,并留意单位。如果图表没有标题或坐标轴标签,要指出它是不完整的。
4. Cumulative Frequency Graph Errors | 累积频率图的错误
Many candidates confuse cumulative frequency with ordinary frequency, and simply plot the frequencies from the table without adding them up. Others use the midpoint of the class interval as the plotting point instead of the upper boundary, or connect the points with straight line segments when a smooth curve is expected. When reading off the median and quartiles, they often misread the scale or forget to state the units.
很多考生把累积频率和普通频率搞混,没有相加就直接把表格里的频数描点。有些人用的是区间的中点,而不是上界来描点,或者在应该画光滑曲线的地方用直线段连接点。在读取中位数和四分位数时,他们经常读错刻度,或者忘记标出单位。
The correction is straightforward: plot cumulative frequency against the upper class boundary of each interval. Accumulate the frequencies step by step, checking that the final cumulative total equals the total number of data values. Draw a smooth curve that passes as close as possible to all the plotted points. To find the median, go to half the total frequency on the vertical axis, travel across to the curve, then down to the horizontal axis and read the value carefully, including units.
纠正方法很清楚:用每个区间的上界来画累积频率。一步一步地累加频数,检查最终的累积总数是否等于数据总数。画一条光滑曲线,让它尽量靠近所有已描好的点。找中位数时,从纵轴上总频数的一半出发,水平移动到曲线,再垂直到横轴上,仔细读出数值,记得带单位。
5. Independent vs. Mutually Exclusive Events | 独立事件与互斥事件混淆
This is one of the most persistent misunderstandings in probability. Students often think that ‘independent’ and ‘mutually exclusive’ mean the same thing. They then misapply the formulas, writing P(A and B) = 0 for independent events or adding probabilities where they should be multiplied. In truth, mutually exclusive events cannot happen at the same time, while independent events do not affect each other’s probabilities.
这是概率里最常见的一个误解。学生们常以为“独立”和“互斥”是一回事。接着就用错公式,对独立事件写出 P(A and B) = 0,或者在应该相乘的地方却用了加法。事实上,互斥事件不可能同时发生,而独立事件则不会影响彼此发生的概率。
Remember the key rules: for mutually exclusive events, P(A or B) = P(A) + P(B). For independent events, P(A and B) = P(A) × P(B). Two events can be both mutually exclusive and independent only in trivial cases where at least one probability is zero. Always test the conditions: if A happens, can B still happen? Does the probability of B change when A occurs? These two questions will guide you to the correct formula.
记住关键的法则:对于互斥事件,P(A or B) = P(A) + P(B);对于独立事件,P(A and B) = P(A) × P(B)。只有在至少一个概率为零的平凡情况下,两个事件才可能既互斥又独立。一定要检验条件:如果 A 发生了,B 还能发生吗?A 的发生会不会改变 B 的概率?这两个问题就能指引你套用正确的公式。
6. Probability Tree Diagram Mistakes | 概率树形图的错误
Even when students draw tree diagrams, they frequently make small mistakes that cost marks. They forget to label branches with probabilities, omit the second set of branches, or fail to check that the probabilities on branches from the same point sum to 1. In ‘without replacement’ scenarios, they often copy the first-stage probabilities onto the second stage without adjusting for the changed sample space.
即使学生画出了树形图,也常常犯一些小错误而丢分。他们忘记在分支上标注概率,漏掉第二组分支,或者没有检查从同一点分出的分支概率和是否为 1。在“不放回”的题目里,他们经常把第一阶段的概率照搬到第二阶段,却没有根据样本空间的变化进行调整。
Build the tree methodically. Start by writing the probabilities on the first set of branches and check they add to 1. Then, for each second-stage branch, write the probability given the first outcome; these should also sum to 1 for each group. For ‘without replacement’, reduce the denominator and, if necessary, the numerator. To find the probability of a combined outcome, multiply along the branches. To find the probability of an event that can happen along multiple paths, add the probabilities of those final outcomes. Always label the outcomes at the end of the branches clearly.
系统地构建树形图。从第一组分支开始,写上概率并检查它们相加是否为 1。然后,对于每个第二阶段分支,根据第一阶段的结果写出概率;每一组分支的概率和也应该为 1。对于“不放回”的情况,要减少分母,有时也要调整分子。求组合结果的概率时,沿着分支相乘。求能通过多条路径发生的事件的概率时,把那些最终结果的概率加起来。始终在分支末端清楚地写出结果。
7. Misinterpreting Scatter Graphs and Correlation | 错误解读散点图与相关性
A huge number of students jump from ‘there is a correlation’ to ‘one variable causes the other’. This leap ignores possible lurking variables and common causes. Another frequent error is drawing a line of best fit by simply joining the first and last points, rather than balancing the points on either side. Students also sometimes misjudge the strength of correlation, calling a very weak trend ‘strong’ because it looks roughly diagonal.
大量学生从“存在相关性”直接跳到“一个变量导致了另一个变量”,这种跳跃忽略了可能存在的潜在变量和共同原因。另一个常见错误是画最佳拟合线时只是连接第一个和最后一个点,而不是让线两侧的点大致平衡。学生们有时也会错误判断相关的强弱,把一个很弱的趋势说成“强相关”,只因为它看上去大致沿对角线方向。
Correlation does not imply causation. When you describe a scatter graph, use phrases like ‘there appears to be a positive association between …’, and always add ‘but this does not prove that one causes the other’. To draw the line of best fit, position a ruler so that roughly half the points lie above and half below the line, and try to make the line pass through the mean point (mean of x, mean of y). Avoid forcing the line through the origin unless there is a good reason. Comment on the strength using terms like weak, moderate or strong, and refer to the spread of the points.
相关性并不意味着因果关系。描述散点图时,要使用像“在……之间似乎存在正向关联”这样的说法,并总是补充“但这并不能证明一个是另一个的原因”。画最佳拟合线时,让直尺处于大致一半点在线上方、一半在线下方的位置,尽量让线通过均值点(x 的均值, y 的均值)。除非有充分理由,否则不要强迫直线通过原点。用弱、中等或强等词来评论相关的强度,并结合点的分散程度来说明。
8. Inappropriate Measures of Spread | 使用不恰当的离散程度度量
It is tempting to report only an average and assume this gives a full picture of the data, but two data sets can have the same mean yet very different spreads. A common error is to use the range as the default measure of spread without recognising how sensitive it is to extreme values. Another is to mix up the interquartile range (IQR) with the range, or to calculate the IQR without first ordering the data.
很多人倾向于只报告一个平均数,以为这就给出了数据的全貌,但两个数据集的均值可以完全相同,离散程度却差别很大。一个常见的错误是把极差当作默认的离散度量,却没有意识到它对极端值有多么敏感。另一个是混淆四分位距(IQR)和极差,或者没有先把数据排序就直接计算 IQR。
Choose your measure of spread to match your measure of centre. When you are using the median, report the IQR, because both are resistant to outliers. When using the mean, the standard deviation is the natural partner, but the range can also be given as a quick summary, as long as you acknowledge its limitations. To find the IQR, order the data, locate the lower quartile (Q₁) and upper quartile (Q₃), then calculate IQR = Q₃ − Q₁. The semi-interquartile range is half of this. Always interpret the spread in context: a small IQR means the middle 50% of the data are tightly packed.
选择的离散度量要与中心度量相匹配。使用中位数时,用 IQR 来报告,因为两者都抗异常值。使用均值时,标准差是天然的搭档,但极差也可以作为一种快速总结,只要承认它的局限。求 IQR 时,先把数据排序,找出下四分位数(Q₁)和上四分位数(Q₃),然后计算 IQR = Q₃ − Q₁。半四分位距是它的一半。始终结合背景解释离散程度:小的 IQR 意味着中间 50% 的数据非常集中。
9. Unjustified Outlier Removal | 无理由地剔除异常值
Seeing an extreme value and immediately deleting it is a serious error, because that outlier could be a genuine piece of data that reveals something important. Some students use the 1.5 × IQR rule to identify outliers but then remove them without comment, or they remove points that lie outside an arbitrary range they have invented themselves.
看到极端值就立刻删除是一个严重错误,因为那个异常值可能是揭示重要信息的真实数据。有些学生用 1.5 × IQR 规则来识别异常值,但随后不加说明就将其移除,或者移除那些落在自己凭空捏造的“范围”之外的点。
The correct scientific practice is to first check for data entry errors. If the value is clearly impossible (e.g. a height of 300 cm for a Year 10 student), it may be a recording mistake and can be corrected or removed with a justification stated. If the value is plausible but unusual, keep it and note its influence. You can apply the outlier fences: lower fence = Q₁ − 1.5 × IQR, upper fence = Q₃ + 1.5 × IQR. Values outside these fences are flagged as potential outliers. Report any decision you make, and consider carrying out the analysis both with and without the outlier to see the difference.
正确的科学做法是首先检查数据录入错误。如果数值明显不可能(比如 Year 10 学生身高 300 cm),那可能是记录错误,可以在说明理由后纠正或移除。如果数值合理但不寻常,保留它并注明其影响。你可以应用异常值界限:下界限 = Q₁ − 1.5 × IQR,上界限 = Q₃ + 1.5 × IQR。落在这些界限之外的值被标记为潜在异常值。无论做了什么决定都要报告出来,并考虑分别在包含和剔除异常值的情况下进行分析,看看有什么差异。
10. Questionnaire Design Pitfalls and Bias | 问卷设计的陷阱与偏倚
Writing a good questionnaire is a skill, and Year 10 projects often fall into traps such as using leading questions (‘Don’t you agree that homework is boring?’), overlapping response options, or forgetting to include a time frame. Another common flaw is sampling only friends, which produces a biased sample that cannot be generalised to the whole population.
设计一份好的问卷是一种技能,Year 10 的项目经常会落入一些陷阱,比如使用引导性问题(“你难道不觉得家庭作业很无聊吗?”)、设定有重叠的选项,或者忘记提供时间范围。另一个常见缺陷是只抽取朋友作为样本,这会产出一个有偏的样本,无法推广到整个总体。
To write a fair questionnaire, keep questions neutral and specific. Use phrases like ‘How often do you …’ and give mutually exclusive response boxes with an ‘Other’ option where needed. Include clear units and time periods. Pilot the questionnaire on a small group to spot ambiguous wording. When collecting data, describe the sampling method, state the sample size, and discuss any limitations honestly. If your sample is biased, acknowledge this and suggest how a more representative sample could be obtained in the future.
要设计一份公平的问卷,问题要保持中立、具体。使用像“你多久一次……”这样的措辞,并提供互斥的回答框,必要时加入“其他”选项。要写清楚单位与时间段。先在小范围内试用问卷,以发现含混的措辞。在收集数据时,说明抽样方法、样本量,并诚实地讨论任何局限性。如果样本有偏,就承认这一点,并建议未来如何获得更具代表性的样本。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导