📚 Common Statistical Mistakes and How to Correct Them | 常见统计误区与纠正方法
Statistics helps us make sense of data, but even small misunderstandings can lead to big mistakes. In Year 8 WJEC Statistics, many students lose marks not because they cannot calculate, but because they apply the wrong method or misinterpret results. This article highlights ten common pitfalls and shows you how to avoid them, giving you clearer thinking and better exam performance.
统计学帮助我们理解数据,但即使很小的误解也可能导致严重的错误。在 WJEC 八年级统计课程中,许多学生丢分并非因为不会计算,而是因为用错方法或错误地解释了结果。本文列出十个常见误区,并告诉你如何避免它们,让你的思路更清晰,考试成绩更出色。
1. Mixing Up Mean, Median and Mode | 混淆平均数、中位数和众数
A classic mistake is using the word ‘average’ to mean any measure of central tendency, then calculating the wrong one. The mean is the sum of all values divided by the count; the median is the middle value when data are ordered; the mode is the most frequent value. If a question asks for the ‘average’ in a dataset with an extreme outlier, the mean will be pulled towards that outlier and may not represent a typical value. Many students blindly calculate the mean without checking which measure is appropriate.
一个经典错误是把“平均数”当作任何集中趋势度量,然后计算出错误的指标。平均数(均值)是所有数值之和除以个数;中位数是排序后位于中间的值;众数是出现频率最高的值。如果题目问的是含极端异常值的数据集的“典型值”,均值会被拉向异常值而无法代表典型水平。许多学生不假思索地计算均值,却没有检查哪种度量更合适。
To correct this, always read the question carefully. Use a table to keep definitions straight:
要纠正这一点,务必仔细审题。用一张表格来理清定义:
| Measure | How to find it | Best used when… |
|---|---|---|
| Mean | Sum of values ÷ number of values | Data are fairly symmetrical with no extreme outliers |
| Median | Middle value in an ordered list | Data are skewed or contain outliers |
| Mode | Value that appears most often | Data are categorical or you need the most popular item |
For example, if nine people earn £20,000 and one earns £150,000, the mean is £33,000, but the median is £20,000 – a much better summary of a typical salary. Always justify your choice in an exam.
例如,若九人收入为 20,000 英镑,一人收入为 150,000 英镑,均值是 33,000 英镑,但中位数是 20,000 英镑——后者更能代表典型薪资。考试中一定要说明你选择的理由。
2. Confusing Bar Charts with Histograms | 混淆条形图与直方图
Many Year 8 students treat bar charts and histograms as the same thing. A bar chart is used for categorical data (e.g. favourite fruit, hair colour) and the bars have gaps between them. A histogram is for continuous data grouped into intervals (e.g. heights, times) and the bars touch, with the area of each bar proportional to frequency. A common error is drawing a histogram with gaps or using a bar chart to display grouped continuous data, which misrepresents the distribution.
很多八年级学生把条形图和直方图混为一谈。条形图用于分类数据(如最爱的水果、头发颜色),条形之间有间隔。直方图则用于分组连续数据(如身高、时间),条形紧挨着,每个条形的面积与频数成比例。常见错误是画直方图时留出间隔,或者用条形图表示分组的连续数据,这都会歪曲数据分布。
Another pitfall is ignoring unequal interval widths in a histogram. If the class widths are different, you must use frequency density (frequency ÷ class width) for the height of the bars. Drawing bars simply by frequency will make wider intervals look artificially taller. Always check the data type before you start drawing.
另一个陷阱是忽略直方图中组距不等的区间。如果组距不同,必须用频率密度(频数 ÷ 组距)作为条形高度。仅按频数画条会使较宽的区间显得过高。始终在绘图前检查数据类型。
3. Misusing Probability Language: ‘Certain’, ‘Likely’, ‘Even Chance’ | 概率用语错误:“一定发生”与“可能发生”
Probability is measured on a scale from 0 to 1, but everyday words can cause confusion. Students sometimes say an event is ‘certain’ when it is merely ‘very likely’. For example, if the probability of rain tomorrow is 0.9, it is not certain – there is still a 10% chance it will not rain. In WJEC exams, precision matters: write probabilities as fractions, decimals or percentages, and label the words correctly.
概率的度量范围是 0 至 1,但日常用语可能导致混淆。学生偶尔会把“极有可能”说成“一定发生”。比如,明天下雨的概率是 0.9,并不意味着必定下雨——仍然有 10% 的可能性不下雨。在 WJEC 考试中,表述必须准确:将概率写成分数、小数或百分数,并正确标注词汇。
A helpful correction exercise is to match phrases to numerical ranges:
一个有效的纠正练习是将词语与数值范围配对:
- Impossible (0): cannot happen, e.g. rolling a 7 on a standard die. / 不可能 (0):不会发生,如掷标准骰子得到 7。
- Unlikely (close to 0): a small chance, e.g. picking a black sock from a drawer of 9 white socks and 1 black sock. / 不太可能(接近0):概率很小,如从 9 只白袜和 1 只黑袜中抽到黑袜。
- Even chance (0.5): equally likely to happen or not, e.g. getting heads on a fair coin. / 对等机会 (0.5):发生与不发生的可能性相同,如公平硬币抛出正面。
- Likely (close to 1): a strong chance but not certain. / 很可能(接近1):机会很大,但并非一定。
- Certain (1): will definitely happen, e.g. the Sun rising tomorrow. / 一定发生 (1):必定发生,如太阳明天升起。
Avoid treating ‘likely’ as a synonym for ‘certain’ – exam questions often test this distinction.
不要把“很可能”当作“一定”的同义词——考题常会检验这一区别。
4. Falling for the Gambler’s Fallacy with Independent Events | 独立事件中的赌徒谬误
An extremely common mistake is believing that past outcomes of independent events affect future probabilities. If a fair coin lands heads five times in a row, the chance of tails on the next toss is still ½. Many students think tails is ‘due’ – this is the gambler’s fallacy. In WJEC statistics, you must state clearly that the events are independent and the probability stays the same.
一个极其常见的错误是认为独立事件的过往结果会影响未来的概率。若一枚公平硬币连续五次正面朝上,下一次掷出反面的概率仍然是 ½。许多学生认为反面该“出现了”——这就是赌徒谬误。在 WJEC 统计中,你必须明确指出事件是独立的,概率保持不变。
To correct this, always ask: ‘Does one trial affect the next?’ If not, treat each probability calculation separately. Use tree diagrams to show that the probabilities on each branch remain unchanged. In questions about spinners, dice or cards with replacement, remind yourself to write ‘independent events – probabilities do not change’.
要纠正这一点,永远问自己:“上一次试验是否影响下一次?”若不,则单独计算每次的概率。使用树状图展示每条分支的概率不变。在涉及转盘、骰子或有放回抽牌的题目中,提醒自己注明“独立事件——概率不改变”。
5. Drawing Conclusions from a Biased Sample | 从有偏样本中得出结论
It is tempting to ask your friends for their opinions and treat the results as true for the whole school. But a sample must be representative. If you only survey your classmates who enjoy football, you cannot conclude that 90% of all Year 8 students love football. Sampling bias occurs when some groups are over- or under‑represented. This leads to unreliable conclusions.
只询问朋友的意见,然后把结果当成全校的普遍情况,这很有诱惑力。但样本必须具有代表性。如果你只调查喜欢足球的同班同学,就不能得出“90% 的八年级学生热爱足球”的结论。当某些群体被过度代表或代表不足时,就会出现抽样偏差,这会导致不可靠的结论。
To avoid this, design a sample that gives every member of the population an equal chance of being chosen – a random sample. In an exam, you may be asked to critique a sampling method. Check for: size (too small?), selection method (only friends? only one class?), and whether it covers all relevant subgroups. Suggest improvements like using a random number generator or stratified sampling by form group.
要避免这一点,需设计一种让总体中每个成员都有同等机会被选中的样本——即随机样本。考试中可能要求你评论一种抽样方法。检查:样本大小(是否太小?)、选择方式(只选朋友?只选一个班?),以及是否涵盖所有相关子群。提出改进建议,如使用随机数生成器或按年级分层抽样。
6. Being Tricked by Charts with a Non-Zero Vertical Axis Start | 被纵轴不从零开始的图表误导
A sensational-looking bar chart sometimes hides a simple truth: the vertical axis does not start at zero. When the axis is cut (e.g. starting at 50 instead of 0), small differences appear huge. A common mistake is to read the height of bars at face value without checking the scale. In WJEC assessments, you may see a chart showing ‘sales rose dramatically’ when in reality they only increased from 83 to 87 – a tiny relative change.
一张看起来很夸张的条形图有时隐藏着一个简单的事实:纵轴并非从零开始。当轴线被截断(例如从 50 而非 0 开始)时,微小的差异看起来会非常巨大。常见的错误是不检查刻度就直接相信条形的高度。在 WJEC 测试中,你可能会看到一张声称“销售额剧增”的图表,而实际上只是从 83 增加到 87——相对变化微不足道。
To correct this, train yourself to look at the axis labels first. Ask: ‘What is the starting value?’ and ‘What would the bars look like if the axis started at zero?’ When describing trends, avoid emotive words like ‘soared’ or ‘collapsed’ unless the data truly support them. Always interpret the graph by quoting actual numbers and the size of the change, not just the visual impression.
要纠正这一点,训练自己首先查看轴标签。问问:“起点是多少?”以及“如果坐标轴从零开始,条形会是什么样子?”在描述趋势时,避免使用“飙升”或“崩盘”等情绪化词汇,除非数据确实支持。解读图表时始终引用实际数字和变化幅度,而不仅仅依赖视觉印象。
7. Mishandling Zeros and Missing Values When Calculating the Mean | 计算平均数时误用零和缺失值
A frequent error is treating a missing data value as zero. If you do not know a student’s test score, you cannot assume they scored zero. Including a false zero drags the mean down artificially. Similarly, some students include genuine zero values incorrectly – for example, if a temperature reading is 0 °C, that 0 is a valid measurement and must be included in the sum and the count. Confusing ‘no data’ with ‘zero measurement’ leads to incorrect means.
一个常见错误是将缺失的数据值当作零处理。如果你不知道某位学生的测试分数,不能假设他得了零分。加入虚假的零分会人为地拉低平均值。同样地,有些学生错误地处理了真正的零值——例如,温度读数为 0 °C 时,这个 0 是有效测量值,必须先加入总和并计入个数。混淆“无数据”与“测量值为零”会导致均值错误。
As a rule: if a value is missing or unknown, omit that entry from both the sum and the count (or use a statistical method to estimate it – but at Year 8 level, omit it). If the value is a genuine zero, include it. Always check the context: a zero score in a quiz means the pupil attempted but got none right; a blank means they were absent. Mark clearly on your working which values you have used.
一条规则:如果数值缺失或未知,就不要将其纳入总和与计数(或使用统计方法估算——但八年级阶段直接略过)。如果数值是真实的零,就要计入。务必检查数据背景:测验得零分表示学生参加了但全错;空白表示缺考。在计算过程中要标出你使用了哪些数值。
8. Percentage Increase vs Percentage Points | 百分比增加与百分点混淆
When a quantity changes from 20% to 30%, students often say it has increased by 10%. In everyday language this may be accepted, but in statistics it is ambiguous. The increase is 10 percentage points, but the percentage increase is actually (30 − 20) ÷ 20 × 100% = 50%. Using the wrong measure can massively exaggerate or understate a change.
当某个量从 20% 变为 30% 时,学生常说它“增加了 10%”。日常语言或许可以接受,但在统计学中这是含糊的。实际上增加了 10 个百分点,但百分比增幅是 (30 − 20) ÷ 20 × 100% = 50%。用错度量会严重夸大或低估变化。
To keep them apart, always clarify whether you are talking about an absolute change in percentage terms (percentage points) or a relative change (percentage increase). A helpful table:
为了区分清楚,始终明确你所指的是百分数的绝对变化(百分点)还是相对变化(百分比增幅)。一张助记表:
| Situation | Change in percentage points | Percentage increase |
|---|---|---|
| Pass rate from 40% to 50% | +10 percentage points | +25% |
| Interest rate from 2% to 3% | +1 percentage point | +50% |
| Unemployment from 8% to 6% | -2 percentage points | -25% |
In your written answers, use the exact phrase ‘percentage points’ when referring to the difference between two percentages. This small habit will protect you from a common mark-losing slip.
书面回答中,提及两个百分数之差时务必使用“百分点”一词。这个小习惯能让你避免一个常见的失分疏忽。
9. Confusing Correlation with Causation | 混淆相关性与因果关系
Two variables may rise or fall together, but that does not prove one causes the other. For example, ice cream sales and sunburn cases both increase in summer, but buying ice cream does not cause sunburn. The common cause is hot, sunny weather. Students often spot a pattern in a scatter graph and immediately claim ‘X causes Y’. In statistics, we say there is a correlation, but causation requires stronger evidence, such as a controlled experiment.
两个变量可能同时上升或下降,但这并不能证明一个导致了另一个。例如,冰淇淋销量和晒伤病例都在夏季增加,但购买冰淇淋并不会导致晒伤。共同的原因是炎热晴朗的天气。学生常在散点图中发现规律,就立刻宣称“X 导致 Y”。在统计学中,我们只能说存在相关,而因果关系需要更有力的证据,比如对照实验。
To avoid this trap, always look for a third, ‘lurking’ variable that might explain the link. When interpreting a scatter diagram, write ‘there is a positive/negative correlation’ rather than ‘A causes B’. If a question asks whether a relationship is causal, think: ‘Could there be another factor?’ and ‘Did anyone deliberately change one variable and measure the other?’. Unless the study is an experiment, stick to describing association, not causation.
要避开这个陷阱,始终寻找可能解释联系的第三个“隐藏”变量。解释散点图时,写下“存在正/负相关”,而非“A 导致 B”。如果题目问关系是否为因果,想想:“会否有其他因素?”以及“是否有人为控制一个变量并测量另一个?”。除非研究是实验,否则只描述关联,不声称因果。
10. Treating Short-Run Frequency as Probability | 将短期频率当作概率
If you spin a spinner 10 times and get red 7 times, you cannot claim the probability of red is 0.7. That is the experimental frequency – an estimate based on a small sample. Probability is the long-run relative frequency, expected to settle around the true value after many, many trials. A common mistake is to write down the experimental frequency as if it were the theoretical probability, especially in questions where students are asked to compare.
如果你转动转盘 10 次,得到 7 次红色,你不能因此声称红色的概率是 0.7。那是实验频率——基于小样本的估计值。概率是长期相对频率,经过非常多次试验后会稳定在真实值附近。常见错误是把实验频率当成理论概率写下来,尤其是在要求学生进行比较的题目中。
Correct this by distinguishing clearly between experimental probability (relative frequency from an experiment) and theoretical probability (based on equally likely outcomes). In your answer, use phrases like ‘the experimental probability was…, but with more trials we would expect it to get closer to the theoretical value of…’ Always acknowledge that a small number of trials gives an unreliable estimate.
要纠正这一点,需明确区分实验概率(来自实验的相对频率)和理论概率(基于等可能结果)。在答案中使用这样的表述:“实验概率为……,但随着试验次数增加,我们预计它会趋近于理论值……”。始终承认少量试验会给出不可靠的估计。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导