Common Mistakes in OxfordAQA MA03 Statistics June 2023 Exam | OxfordAQA MA03 统计学2023年6月考试易错点解析

📚 Common Mistakes in OxfordAQA MA03 Statistics June 2023 Exam | OxfordAQA MA03 统计学2023年6月考试易错点解析

Analysing the June 2023 OxfordAQA MA03 Statistics mark scheme reveals several recurring errors that cost students valuable marks. This article highlights the most frequent pitfalls, explains why they occur, and shows how to avoid them in future exams.

分析2023年6月 OxfordAQA MA03 统计学评分方案可以看出,许多反复出现的错误让考生白白丢分。本文梳理了最常见的失分陷阱,解释错误原因,并指导如何在今后的考试中避开这些雷区。

1. Misinterpreting Probability Notation | 误解概率符号

Many candidates confused P(A|B) with P(A ∩ B). When a question asked for a conditional probability, answers often used the intersection probability instead. Remember, P(A|B) = P(A ∩ B) / P(B), provided P(B) > 0.

很多考生混淆了 P(A|B) 和 P(A ∩ B)。当题目要求计算条件概率时,答案却经常取交集概率。请记住,只要 P(B) > 0,就有 P(A|B) = P(A ∩ B) / P(B)。

Another common notation error involved writing P(A ∪ B) as a product of individual probabilities without checking independence. In the exam, some assumed independence without justification, causing incorrect values for union probabilities.

另一个常见符号错误是在未检验独立性的前提下,就将 P(A ∪ B) 表达成各自概率的乘积。考试中有人不经论证就假定事件独立,导致并集概率计算错误。


2. Failures in Checking Binomial Conditions | 忽略二项分布条件

When modelling with the binomial distribution B(n, p), candidates frequently omitted the necessary conditions: fixed number of trials, two possible outcomes, constant probability of success, and independent trials. The mark scheme consistently penalised missing or incomplete condition statements.

在使用二项分布 B(n, p) 建模时,考生经常漏掉必要的前提条件:试验次数固定、每次试验只有两种可能结果、每次成功的概率恒定、各次试验相互独立。评分方案对缺失或不完整的条件说明一贯扣分。

Specifically, some wrote ‘random sample’ but did not link it to independence. Others used a binomial model when the probability changed due to ‘without replacement’ sampling, failing to check that the population size is large enough for a binomial approximation.

具体来说,有人只写了“随机样本”却没有说明其独立性。还有人在“不放回”抽样情形下仍使用二项模型,而没有验证总体容量是否足够大从而满足二项近似条件。


3. Normal Distribution: Errors in Standardisation | 正态分布:标准化错误

A classic mistake was mishandling the standardisation formula z = (X – μ) / σ. Some candidates subtracted μ in the wrong order or used σ² instead of σ. For example, for X ~ N(100, 15²), candidates wrongly computed z = (X – 100) / 15² instead of dividing by 15.

一个典型错误是标准化公式 z = (X – μ) / σ 的处理不当。有人把 μ 的减法顺序搞反,或者用方差 σ² 代替标准差 σ。例如,对于 X ~ N(100, 15²),有人错误地计算 z = (X – 100) / 15² 而不是除以 15。

Furthermore, when finding an unknown μ or σ, many students solved the backwards normal problem incorrectly by not using the correct z-value from the tables. The MS showed that using the wrong tail probability (e.g., aiming for 0.95 when the lower tail was required) was a frequent error.

此外,在求未知的 μ 或 σ 时,很多学生不会正确地逆向解正态分布问题,未能从统计用表里查出正确的 z 值。评分方案显示出,误用尾部概率(例如需要下尾时却套用 0.95)是高频错误。


4. Discrete Random Variables: Probability Sum Not Equal to 1 | 离散随机变量:概率和不等于1

A very basic mistake was writing a probability distribution where ∑P(X = x) ≠ 1. In exam questions asking to find an unknown constant, candidates often set up the equation correctly but made arithmetic errors in solving it. Always double-check that the sum of your final probabilities is exactly 1.

一个非常基础的错误是列出的概率分布表中 ∑P(X = x) ≠ 1。在题目要求计算未知常数时,考生往往能正确列出方程,却在求解过程中犯了算术错误。务必二次检验最终概率之和是否正好为 1。

Also, some candidates gave probabilities greater than 1 or negative, which is impossible. The mark scheme instructed examiners not to award full marks for answers that included invalid probabilities, even if the formula for the constant was correct.

此外,有人竟然给出了大于 1 或负的概率值,这根本不可能。评分方案指示阅卷人,只要答案包含无效的概率,即便常数求解公式正确,也不给满分。


5. Confusion Between E(X) and E(X²) in Variance Calculation | 方差计算中混淆E(X)和E(X²)

The variance formula Var(X) = E(X²) – [E(X)]² was often misapplied. Several candidates calculated E(X) correctly, then squared it and subtracted from E(X²) without remembering the order. Others computed E(X²) wrongly by squaring the x-values before multiplying by the probabilities, or by mishandling fractions.

方差公式 Var(X) = E(X²) – [E(X)]² 经常被错误使用。一些考生正确求出了 E(X),接着未记清顺序,直接用它减去 E(X²) 的平方。另一些人在计算 E(X²) 时,要么没有在乘以概率之前对 x 值取平方,要么在分数处理上出了错。

A specific issue in the June 2023 paper involved a discrete random variable with a large constant. Candidates mislaid the constant term when substituting into the E(X²) expression, leading to a completely wrong variance. Pay careful attention to every term and bracket.

2023年6月试卷中有一个涉及大常数项的离散随机变量问题。考生在代入 E(X²) 表达式时弄丢了常数项,导致方差完全算错。对每一项和括号都要格外小心。


6. Linear Regression: Incorrect Identification of Variables | 线性回归:变量识别错误

In regression questions, the explanatory (independent) variable is the one used to predict the response (dependent) variable. A common blunder was to swap these roles, obtaining the regression line of x on y instead of y on x. This mistake invalidated all subsequent predictions and interpretations.

在回归题中,解释变量(自变量)是用来预测响应变量(因变量)的。一个常见的大错是将两者的角色互换,求得的是 x 对 y 的回归直线而非 y 对 x 的回归。这个错误致使后续所有的预测和解读无效。

The mark scheme noted that some candidates even used the correct formula for the gradient b = Sxy / Sxx but then labelled the equation incorrectly, or they misinterpreted the context. Always read the question stem carefully to understand which variable is being predicted.

评分方案指出,有些考生甚至用了正确的梯度公式 b = Sxy / Sxx,却把方程标注错误,或者误解了背景。一定要仔细阅读题干,弄清要预测的是哪个变量。


7. Correlation vs. Causation | 相关性与因果关系混淆

A context-based question required interpreting a high correlation coefficient. Many candidates asserted a causal relationship without any evidence, losing marks. The correct interpretation is ‘there is a strong positive linear association’, not ‘x causes y’.

有一道基于情境的题目要求解读较高的相关系数。很多考生在没有任何依据的情况下断言存在因果关系,因而失分。正确的解读是“存在很强的正线性关联”,而不是“x 导致 y”。

Additionally, some did not comment on the strength or direction in context, simply saying ‘the correlation is 0.9’. The mark scheme expected a description relating to the real-world variables, such as ‘as the number of hours of revision increases, the exam score tends to increase’.

此外,有人没有结合情境说明强度和方向,只是说“相关系数是 0.9”。评分方案期望的是与真实变量相关的描述,比如“随着复习时数的增加,考试成绩呈上升趋势”。


8. Histograms: Forgetting Frequency Density | 直方图:忘记频率密度

Drawing a histogram with unequal class widths, students often plotted frequency instead of frequency density. The mark scheme was explicit: bars must have heights proportional to frequency density = frequency / class width. Using raw frequency directly on the vertical axis resulted in zero for the graph part.

在绘制组距不等的直方图时,学生经常直接用频数作图,而非频率密度。评分方案写得很明确:柱状的高度必须与频率密度 = 频数 / 组距成正比。在纵轴上直接使用原始频数的,绘图部分为零分。

Another subtle point was careless labelling of axes. Several papers omitted the units or labelled the vertical axis simply ‘Frequency’ when it was actually ‘Frequency density’. The MS deducted marks unless both axes were correctly labelled with appropriate scales.

另一个细微之处是坐标轴的粗心标注。多份答卷遗漏了单位,或在纵轴实际上应为“频率密度”时仍标注为“频数”。除非两轴都正确标注且刻度合理,否则评分方案就会扣分。


9. Box Plots and Outlier Detection | 箱线图与异常值检测

Outlier fences are defined by Q1 – 1.5 × IQR and Q3 + 1.5 × IQR. Many candidates forgot to multiply the IQR by 1.5, simply using Q1 – IQR as the lower fence. This error caused them to misclassify outliers and draw inaccurate box plots.

异常值界限的定义是 Q1 – 1.5 × IQR 和 Q3 + 1.5 × IQR。很多考生忘了将 IQR 乘以 1.5,直接用 Q1 – IQR 作为下界。这个错误导致他们误判异常值,并绘制出不准确的箱线图。

Additionally, some students confused outliers with extreme values that should be shown as separate points. In the June 2023 paper, a data set included an outlier which should have been plotted as an isolated cross beyond the whisker. Candidates who adjusted whiskers to that point lost marks.

另外,有些学生分不清异常值和该用单独点标出的极端值。2023年6月试卷的数据集中包含一个异常值,本应将它在须线之外画成独立的叉号点。那些把须线直接拉到该点的考生因而失分。


10. Sampling Methods: Strengths and Weaknesses | 抽样方法优缺点混淆

Questions on simple random sampling, stratified sampling, and quota sampling required evaluating advantages and disadvantages. A frequent error was attributing a characteristic of one method to another. For instance, claiming that a quota sample is unbiased (it is not, due to non-random selection) or that a simple random sample always perfectly represents the population (it may suffer from sampling error).

关于简单随机抽样、分层抽样和配额抽样的题目要求评价优缺点。常见错误是把一种方法的特点张冠李戴到另一种方法上。例如,声称配额样本是无偏的(由于非随机选择,其实有偏),或者说简单随机样本始终完美代表总体(它可能受到抽样误差的影响)。

In one mark scheme, a candidate stated that stratified sampling is ‘easier and cheaper’ than simple random sampling. This contradicted accepted knowledge that stratification requires a complete sampling frame and can be more costly. Know the key features of each technique precisely.

在一份评分方案中,有考生写道分层抽样比简单随机抽样“更容易、更便宜”。这与公认的知识相悖,因为分层需要完整的抽样框,成本可能更高。准确掌握每种技术的关键特征。


11. Continuous Random Variable: Using c.d.f. for Probabilities | 连续随机变量:使用累积分布函数求概率时的错误

For a continuous random variable, probabilities are found by evaluating the cumulative distribution function F(x) = P(X ≤ x). A significant error was to compute P(X < a) as F(a) – P(X = a). Since continuous distributions have zero probability at a single point, P(X < a) = F(a) without subtraction.

对于连续随机变量,概率通过计算累积分布函数 F(x) = P(X ≤ x) 来求。一个严重的错误是将 P(X < a) 算作 F(a) – P(X = a)。因为连续分布在单点上的概率为零,P(X < a) 直接等于 F(a),无需做减法。

When the c.d.f. was given piecewise, some candidates used the wrong branch for the given value, leading to completely incorrect probabilities. Cross-check the interval definitions carefully before substituting.

当累积分布函数以分段形式给出时,一些考生对给定值用了错误的分支,导致概率完全算错。代入前一定要仔细对照区间定义。


12. Conditional Probability Tree Diagrams | 条件概率树状图错误

Tree diagrams often carry heavy marks. A typical slip was labelling second-branch probabilities as unconditional rather than conditional. For instance, after taking a red ball, the probability of taking another red ball should be updated according to the reduced sample space. Candidates who used the original probability lost all subsequent method marks.

树状图往往占分很重。典型疏忽是将第二层分支上的概率标成无条件概率,而非条件概率。例如,取出一个红球后,再取红球的概率应根据缩小的样本空间更新。用原始概率的考生会丢掉之后所有的方法分。

In the June 2023 MS, some candidates also drew incomplete trees or failed to label probabilities at the ends. Even when the probabilities were numerically correct, missing labels resulted in deduction. Ensure every branch carries the appropriate probability and that outcomes are clearly denoted.

2023年6月评分方案中,部分考生画的树状图不完整,或未在末梢标注概率。即使数值正确,标签缺失也会扣分。务必让每个分支都有对应的概率,且结果端清晰标示。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading