📚 Common Statistical Mistakes and How to Fix Them | 常见统计误区与纠正方法
Statistics helps us understand the world through data, but even simple ideas can trip us up. Year 7 students often make similar mistakes when collecting, displaying, or interpreting data. Recognising these common pitfalls is the first step to becoming a confident statistician.
统计学帮助我们通过数据理解世界,但即使简单的概念也可能让我们犯错。7年级学生在收集、展示或解读数据时,常常会犯相似的错误。识别这些常见陷阱是成为自信统计学家的第一步。
1. Misunderstanding Mean, Median, and Mode | 误解平均数、中位数与众数
Many students believe the ‘average’ is always the mean and that it represents a typical value. They may ignore extreme values (outliers), leading to a distorted picture of the data.
许多学生认为“平均值”总是指平均数,并且它代表典型值。他们可能忽略极端值(异常值),导致对数据的描述失真。
A common mistake is computing the median before sorting the numbers. For example, from the list 8, 3, 5, 10, a student might pick 5 as the median without ordering. The correct approach is to order: 3, 5, 8, 10, then find the middle (5.5).
一个常见错误是在排序之前计算中位数。例如,从列表8、3、5、10中,学生可能直接取5作为中位数而没有排序。正确的方法是先排序:3、5、8、10,然后找到中间值(5.5)。
Mode confusion also occurs when there is no repeating number; students may say the mode is 0 or none instead of stating there is no mode. When a dataset has two modes, calling it ‘no mode’ is another error – it is in fact bimodal.
众数混淆也发生在没有重复数字时;学生可能会说众数是0或没有,而不是说明没有众数。当数据集有两个众数时,将其称为“没有众数”是另一个错误——它实际上是双众数的。
2. Bar Chart and Pictogram Errors | 条形图与象形图错误
When reading bar charts, students often compare the heights of bars without checking the scale on the vertical axis. If the scale jumps by 5s but a bar stops between two gridlines, they might misread the value.
阅读条形图时,学生经常只比较条形的高度而不检查纵轴的刻度。如果刻度间隔是5,而条形停在两条网格线之间,他们可能会误读数值。
With pictograms, a key common mistake is ignoring the symbol’s value. For instance, if one smiley face represents 4 students, a half-face represents 2, but many treat it as a full face and count it as 4.
在象形图中,一个常见错误是忽略符号代表的值。例如,如果一个笑脸代表4名学生,那么半个笑脸代表2名,但许多学生将其当作整个笑脸,并计为4。
Another error is drawing bar charts with unequal bar widths, which can mislead the viewer into thinking area represents frequency. In a bar chart, only the height matters because bars are separate and represent categories.
另一个错误是绘制条形图时条形宽度不一致,这会误导读者认为面积代表频数。在条形图中,只有高度重要,因为条形是分离的并代表类别。
3. Ignoring Scale Manipulation on Graphs | 忽视图表上的刻度操控
A graph with a vertical axis that does not start at zero can exaggerate small differences. Students often fail to notice this and overstate the change, saying a value ‘doubled’ when it increased only slightly.
纵轴不从零开始的图表可能夸大小差异。学生常常未能注意到这一点,从而夸大变化,声称某个值“翻倍”了,而实际上它只增加了一点点。
Similarly, irregular intervals (e.g., 0, 10, 50, 100) can distort patterns. Always check the labels and spacing before making comparisons.
同样,不规则的间隔(如0、10、50、100)会扭曲模式。在进行比较之前,务必检查标签和间距。
4. Misinterpreting Pie Charts | 误读饼图
Students often assume a larger slice always means a bigger number, forgetting to relate it to the total. A slice of 25% of 200 is 50, but 40% of 100 is 40 – the percentages alone are not comparable without knowing the whole.
学生常常认为更大的扇区总是代表更大的数量,却忘记将其与总数关联。200的25%是50,而100的40%是40。在不了解整体的情况下,单独比较百分比没有意义。
When two pie charts have different totals, it is a mistake to directly compare their sector sizes. Always look for the actual counts or the total given alongside the chart.
当两个饼图的总数不同时,直接比较扇区大小是错误的。务必查看实际计数或图表旁给出的总数。
5. Probability Misconceptions | 概率误解
The gambler’s fallacy is common: after flipping a coin and getting heads five times, a student thinks tails is ‘due’ on the next flip. In reality, each flip is independent, and the probability remains 1/2.
赌徒谬误很常见:抛硬币五次得到正面后,学生认为下一次“该”反面了。实际上,每次抛掷是独立的,概率始终是1/2。
Some students think a probability of 0.3 means the event will happen exactly 3 times out of 10, not realising it is a long-term average. Short sequences can vary greatly.
有些学生认为概率0.3意味着事件在10次中恰好发生3次,没有意识到这是长期平均值。短序列可能会有很大变化。
They also confuse ‘even chance’ with ‘certain’. Even chance is 0.5 or 50%, not 1 or 100%.
他们还混淆“等可能性”(一半机会)与“必然发生”。一半机会是0.5或50%,而不是1或100%。
6. Sampling Bias | 抽样偏差
When collecting data, Year 7 pupils often ask only their friends, resulting in a convenience sample that does not represent the whole population. This can lead to wrong conclusions.
收集数据时,7年级学生常常只询问自己的朋友,导致便利抽样,不能代表整个总体。这可能导致错误的结论。
To avoid bias, they should use random sampling: give every person an equal chance to be chosen. A simple method is to draw names from a hat or use a random number generator.
为避免偏差,他们应该使用随机抽样:给每个人均等的机会被选中。一个简单的方法是从帽子中抽名字或使用随机数生成器。
7. Confusing Correlation with Causation | 混淆相关与因果
A classic error is seeing two trends together and assuming one causes the other. For example, ice cream sales and drowning incidents both rise in summer, but hot weather causes both – ice cream does not cause drowning.
一个经典错误是看到两个趋势同时出现,就假定一个导致另一个。例如,冰淇淋销量和溺水事件都在夏季增加,但炎热的天气同时导致这两者——冰淇淋并不会导致溺水。
Always ask: ‘Could there be a third factor?’ before claiming a cause-and-effect relationship. Just because two things happen together does not mean one makes the other happen.
在声称因果关系之前,始终要问:“是否存在第三个因素?”仅仅因为两件事同时发生,并不意味着一个导致了另一个。
8. Calculating the Range Incorrectly | 错误计算极差
Students sometimes subtract the first number from the last instead of the smallest from the largest. For the set 12, 7, 15, 9, the range is 15 – 7 = 8, not 12 – 9 = 3.
学生有时用第一个数减去最后一个数,而不是用最大值减去最小值。对于数据集12、7、15、9,极差是15 – 7 = 8,而不是12 – 9 = 3。
They may also forget to include units or write the range as a single number without context, e.g., ‘8’ instead of ‘8 cm’. The range is not the ‘average’ and should not be used alone to describe spread.
他们还可能忘记包含单位,或写出的极差没有上下文,比如只写“8”而不是“8 cm”。极差不是“平均值”,不应单独用来描述分散程度。
9. Ignoring Outliers | 忽略异常值
An outlier is a value that is much higher or lower than the rest. Students often calculate the mean including the outlier, which skews the average. For example, test scores: 60, 65, 70, 100 – the mean is 73.75, but most scores are around 65.
异常值是一个远高于或低于其他数据的值。学生常在计算平均数时包含异常值,从而导致平均值偏斜。例如,测试分数:60、65、70、100——平均数是73.75,但大多数分数在65左右。
It is wise to identify outliers and consider using the median instead, or to note the effect of the outlier on the mean. Spotting an outlier helps you understand whether the mean is truly representative.
明智的做法是识别异常值并考虑使用中位数,或者说明异常值对平均数的影响。发现异常值有助于你理解平均数是否真正具有代表性。
10. Using the Wrong Type of Average | 使用不当的平均值类型
For categorical data (words, not numbers), the mean is meaningless. If students are asked about favourite colours, the mode (most frequent) is the appropriate average, not the mean.
对于分类数据(词语而非数字),平均数没有意义。如果学生被问及最喜欢的颜色,众数(出现最频繁的)是合适的平均指标,而不是平均数。
Using the mean for a skewed distribution can be misleading. The median often better represents a typical value when data is not symmetric, such as house prices or incomes.
对于偏斜分布使用平均数可能产生误导。当数据不对称时,中位数通常更能代表典型值,例如房价或收入数据。
11. Data Collection Errors: Leading Questions | 数据收集错误:诱导性问题
A survey question like ‘Don’t you think homework is a waste of time?’ pushes the respondent towards a particular answer. This is a leading question and produces biased data.
像“你不觉得家庭作业是浪费时间吗?”这样的调查问题会将回答者推向特定答案。这就是诱导性问题,会产生有偏差的数据。
To get honest responses, phrase questions neutrally: ‘What is your opinion on the amount of homework?’ Also offer balanced options rather than only extreme ones.
为了获得诚实的回答,要措辞中立:“你对家庭作业量有什么看法?”还要提供平衡的选项,而不是只给极端选项。
12. Misreading Two-Way Tables | 误读双向表
Two-way tables show counts for two categories. A common mistake is reading the wrong row or column total. For example, finding how many boys like football means looking at the intersection of ‘Boys’ row and ‘Football’ column, not the total of the row.
双向表显示两个类别的计数。常见错误是看错行或列的合计。例如,找出有多少男孩喜欢足球,需要查看“男孩”行与“足球”列的交叉点,而不是该行的合计。
When calculating percentages, students often use the wrong denominator, such as using the overall total instead of the row or column total for a conditional percentage. Always ask: ‘Percentage of what?’ before dividing.
计算百分比时,学生经常用错分母,比如计算条件百分比时用总计数而不是行或列的合计。在相除之前,始终要问:“要计算什么的百分比?”
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply