IGCSE CAIE Statistics: High-Frequency Topics and Common Mistake Analysis | IGCSE CAIE 统计:高频考点与易错题分析

📚 IGCSE CAIE Statistics: High-Frequency Topics and Common Mistake Analysis | IGCSE CAIE 统计:高频考点与易错题分析

This revision guide covers the most frequently examined topics in the CAIE IGCSE Statistics syllabus and highlights the mistakes that candidates often make. Each section pairs English and Chinese explanations so that you can master key methods and avoid losing marks on common pitfalls.

本复习指南覆盖 CAIE IGCSE 统计大纲中最常考的知识点,并重点分析考生常犯的错误。每部分均配有中英文对照讲解,帮助你掌握核心方法,避免常见失分点。

1. Types of Data and Sampling Methods | 数据类型与抽样方法

In CAIE IGCSE Statistics, data can be qualitative, such as categories like colour or mode of transport, or quantitative, meaning numerical. Quantitative data can be discrete when values are counted and take specific numbers, or continuous when values are measured on a scale.

在 CAIE IGCSE 统计中,数据可分为定性数据,例如颜色、交通方式等类别,以及定量数据,即数值型数据。定量数据又可分为离散型——通常是计数数据,只能取特定数值,以及连续型——通过测量得到的数据。

Common sampling methods include simple random sampling, stratified sampling, systematic sampling and quota sampling. A very common exam mistake is to name a method without describing how the sample is actually selected or without explaining how the method reduces bias.

常见抽样方法包括简单随机抽样、分层抽样、系统抽样和配额抽样。考试中非常常见的错误是只写出方法名称,却没有说明样本是如何实际抽取的,也没有解释该方法如何减少偏差。

A random sample gives every member of the population an equal chance of being chosen. Stratified sampling keeps the same proportion from each group, which is useful when groups differ in size or characteristics.

随机抽样使总体中每个成员都有相等的机会被选中。分层抽样则保持每个组的比例相同,当各组大小或特征差异较大时非常有用。


2. Frequency Tables, Histograms and Frequency Density | 频数表、直方图与频率密度

For a histogram, the vertical axis must show frequency density, not raw frequency. The formula is: frequency density = frequency ÷ class width. The area of each bar then represents the frequency in that class.

在直方图中,纵轴必须表示频率密度,而不是原始频数。公式为:频率密度 = 频数 ÷ 组距。这样每一条柱形的面积就代表该组的频数。

If class widths are unequal, many candidates mistakenly plot raw frequency on the vertical axis, which makes the histogram misleading because wider classes appear too tall. When class widths differ, the height of each bar must be adjusted using frequency density.

当组距不相等时,许多考生错误地用原始频数作为纵轴,这样会使直方图产生误导,因为较宽的组会显得过高。当组距不同时,必须使用频率密度来调整每条柱形的高度。

Another common error is using the wrong class boundaries for continuous data. For example, the interval 10-19 measured to the nearest whole number should be taken as 9.5 ≤ x < 19.5, not 10 ≤ x < 20.

另一个常见错误是连续数据使用错误的组界。例如,精确到整数的区间 10-19 应取作 9.5 ≤ x < 19.5,而不是 10 ≤ x < 20。


3. Averages and Measures of Spread | 平均数与离散程度指标

For ungrouped data, the mean is calculated as Σx ÷ n. For grouped data, the midpoint of each class represents the class, so the mean is approximated by Σfx ÷ Σf, where f is the frequency and x is the class midpoint.

对于未分组数据,平均数计算为 Σx ÷ n。对于分组数据,每组的组中值代表该组,因此平均数近似为 Σfx ÷ Σf,其中 f 为频数,x 为组中值。

The range is the difference between the largest and smallest values. The interquartile range is upper quartile − lower quartile, often written as IQR = Q₃ − Q₁. Variance is the average squared deviation from the mean: variance = Σx² ÷ n − (Σx ÷ n)², and standard deviation is the square root of variance.

极差是最大值与最小值之差。四分位距为上四分位数减下四分位数,通常写作 IQR = Q₃ − Q₁。方差是各数据与平均数之差的平方的平均值:方差 = Σx² ÷ n − (Σx ÷ n)²,标准差是方差的平方根。

Common mistakes include dividing by n − 1 when the syllabus expects population variance, forgetting to square midpoints in grouped data, or using frequency instead of fx in grouped calculations.

常见错误包括:在大纲要求总体方差时误用 n − 1 作分母;在分组数据中忘记对组中值平方;或者在分组计算中使用频数 f 而不是 fx。

When combining two data sets, the overall mean is (n₁x̄₁ + n₂x̄₂) ÷ (n₁ + n₂). Many candidates incorrectly add the two means and divide by 2, which only works when the two groups have equal size.

当合并两组数据时,总平均数为 (n₁x̄₁ + n₂x̄₂) ÷ (n₁ + n₂)。许多考生错误地将两个平均数相加后除以 2,这仅在两组样本量相等时才正确。


4. Cumulative Frequency and Box Plots | 累积频数与箱线图

Cumulative frequency is plotted against the upper class boundary. The cumulative frequency curve is used to estimate the median, quartiles and percentiles. The median is read at 50% of the total frequency, the lower quartile at 25%, and the upper quartile at 75%.

累积频数应相对于各组上组界绘制。累积频数曲线用于估计中位数、四分位数和百分位数。总频数的 50% 处读取中位数,25% 处读取下四分位数,75% 处读取上四分位数。

A box plot is drawn using the minimum value, lower quartile, median, upper quartile and maximum value. The length of the box is the IQR. A common mistake is to read the value directly from the cumulative frequency number instead of from the horizontal axis when constructing the box plot.

箱线图使用最小值、下四分位数、中位数、上四分位数和最大值绘制。箱体长度即为四分位距。常见错误是在绘制箱线图时直接使用累积频数数字,而不是从横轴上读取对应的数据值。

Another typical error is forgetting to add the whiskers or using the class midpoints rather than the upper class boundaries when plotting a cumulative frequency graph.

另一个典型错误是忘记绘制箱线图两端的须线,或者在绘制累积频数图时使用组中值而不是组上界。


5. Scatter Diagrams and Correlation | 散点图与相关性

Scatter diagrams show the relationship between two variables. Correlation can be positive, negative or zero. If points tend to rise as the x variable increases, correlation is positive; if one variable increases while the other decreases, correlation is negative.

散点图用于展示两个变量之间的关系。相关性可分为正相关、负相关或不相关。如果随着 x 变量增大,点整体呈上升趋势,则为正相关;如果一个变量增大而另一个减小,则为负相关。

A line of best fit should pass close to the mean point (x̄, ȳ) and balance the plotted points. It can be used for interpolation within the data range, but extrapolation outside the range is unreliable because the pattern may not continue.

最佳拟合线应经过均值点 (x̄, ȳ),并使各点大致均匀分布在直线两侧。该线可用于数据范围内的插值估算,但超出范围的外推不可靠,因为变化趋势可能不会延续。

Correlation does not imply causation. A strong correlation alone does not prove that one variable causes the other to change; there may be a third variable or a coincidence.

相关关系不等于因果关系。仅凭高度相关并不能证明一个变量导致另一个变量变化,可能存在第三个变量,也可能只是巧合。


6. Probability, Venn Diagrams and Tree Diagrams | 概率、维恩图与树形图

The probability of an event is the number of favourable outcomes divided by the total number of outcomes. For any event A, the probability of not A is P(A’) = 1 − P(A). The intersection P(A ∩ B) means the probability that both A and B occur.

事件发生的概率等于有利结果数除以所有可能结果数。对于任何事件 A,不发生 A 的概率为 P(A’) = 1 − P(A)。交集 P(A ∩ B) 表示事件 A 和 B 同时发生的概率。

If events are mutually exclusive, then P(A ∪ B) = P(A) + P(B). If they are not mutually exclusive, the addition rule is P(A ∪ B) = P(A) + P(B) − P(A ∩ B). Candidates often forget to subtract the intersection and therefore double count outcomes.

如果事件互斥,则 P(A ∪ B) = P(A) + P(B)。如果事件不互斥,则应使用加法公式 P(A ∪ B) = P(A) + P(B) − P(A ∩ B)。考生常忘记减去交集部分,导致重复计算。

In a tree diagram, you multiply along branches for successive events and add the probabilities of different final outcomes. When items are selected without replacement, the denominator changes for the second stage; with replacement, probabilities remain the same.

在树形图中,连续发生的事件应沿线相乘,不同最终结果的概率应相加。当抽样后不放回时,第二阶段的概率分母会发生变化;如果放回,则各阶段的概率保持不变。


7. Conditional Probability | 条件概率

Conditional probability is the probability of event A given that event B has already occurred. The formula is P(A | B) = P(A ∩ B) ÷ P(B). It should not be confused with the joint probability P(A ∩ B).

条件概率是指在事件 B 已经发生的条件下事件 A 发生的概率。公式为 P(A | B) = P(A ∩ B) ÷ P(B)。它不应与联合概率 P(A ∩ B) 混淆。

A two-way table is often the most reliable way to organise information for conditional probability questions. Many candidates read the wrong row or column total when constructing the table, so always check that row and column totals match the given data.

双向表通常是解答条件概率题时最可靠的信息整理方式。许多考生在构建表格时会读错行或列的总计,因此务必检查行合计与列合计是否与已知数据一致。

Two events A and B are independent if P(A | B) = P(A) or equivalently P(A ∩ B) = P(A) × P(B). If independence is not stated, you should use the conditional formula or the table to find the missing value.

如果 P(A | B) = P(A) 或等价地 P(A ∩ B) = P(A) × P(B),则事件 A 与 B 相互独立。如果题目没有说明独立,应使用条件概率公式或表格来求未知值。


8. Discrete Random Variables and Binomial Distribution | 离散随机变量与二项分布

A discrete random variable X has expectation E(X) = Σx × P(X = x), often written as Σxp. Its variance can be calculated as Var(X) = Σx²p − μ² = E(X²) − [E(X)]².

离散随机变量 X 的期望为 E(X) = Σx × P(X = x),通常写作 Σxp。它的方差可计算为 Var(X) = Σx²p − μ² = E(X²) − [E(X)]²。

For a binomial distribution X ~ B(n, p), the probability of exactly r successes is P(X = r) = nCr × pʳ × (1 − p)ⁿ⁻ʳ. The mean is np and the variance is np(1 − p).

对于二项分布 X ~ B(n, p),恰好有 r 次成功的概率为 P(X = r) = nCr × pʳ × (1 − p)ⁿ⁻ʳ。其平均数为 np,方差为 np(1 − p)。

Binomial distribution applies only when there are a fixed number of independent trials, exactly two possible outcomes, and a constant probability of success. A frequent mistake is using n = total number of items when n should be the number of trials actually selected or tested.

二项分布仅适用于固定次数的独立试验、只有两种可能结果且成功概率不变的情形。常见错误是当 n 应为实际选取或试验的次数时,误把 n 设为所有物品的总数。


9. Normal Distribution | 正态分布

The normal distribution is continuous, symmetric and bell-shaped. Its mean, median and mode are equal. The total area under the curve is 1, and probabilities correspond to areas under the curve.

正态分布是连续、对称、呈钟形的分布。其平均数、中位数和众数相等。曲线下的总面积为 1,概率对应曲线下的面积。

To find probabilities, convert X to a standard score using z = (X − μ) ÷ σ. Then use normal tables or a calculator. For inverse problems, find z from the given probability and convert back using X = μ + zσ.

求概率时,

Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version