📚 PDF资源导航

GCSE CIE Maths: Statistics Key Points | GCSE CIE 数学:统计 考点精讲

📚 GCSE CIE Maths: Statistics Key Points | GCSE CIE 数学:统计 考点精讲

Understanding statistics is essential for CIE GCSE Mathematics. This section covers how to collect, organise, and interpret data, including measures of central tendency, dispersion, graphical representations, and basic probability. Mastering these concepts helps you make data-driven decisions and solve real-world problems. Here we break down the core topics in a simple, structured way.

理解统计学对 CIE GCSE 数学至关重要。这部分涵盖如何收集、整理和解释数据,包括集中趋势度量、离散程度、图形表示和基础概率。掌握这些概念有助于做出数据驱动的决策并解决实际问题。下面我们以简单、结构化的方式分解核心主题。


1. Mean, Median and Mode | 平均数、中位数和众数

The mean is the sum of all values divided by the number of values. For the data set 3, 5, 7, 7, 9, the mean = (3+5+7+7+9) ÷ 5 = 31 ÷ 5 = 6.2.

平均数(均值)是所有数值之和除以数值的个数。对于数据集 3, 5, 7, 7, 9,平均数为(3+5+7+7+9)÷ 5 = 31 ÷ 5 = 6.2。

The median is the middle value when data is ordered. With five numbers in order: 3, 5, 7, 7, 9, the median is the third value, 7. If there is an even number of values, take the mean of the two middle numbers.

中位数是将数据排序后位于中间的值。五个数字按顺序排列:3, 5, 7, 7, 9,中位数是第三个值 7。如果有偶数个数值,则取中间两个数的平均数。

The mode is the value that appears most frequently. In 3, 5, 7, 7, 9, the mode is 7. A set can have more than one mode or no mode at all.

众数是出现次数最多的值。在 3, 5, 7, 7, 9 中,众数是 7。一组数据可以有一个以上的众数,也可以没有众数。

For grouped frequency tables, the mean is estimated using the formula: Mean ≈ Σ(f × midpoint) ÷ Σf. The modal class is the class with the highest frequency; the median class is found using cumulative frequency.

对于分组频数表,平均数用公式估算:平均数 ≈ Σ(频数 × 组中点)÷ 总频数。众数类别是频数最高的组;中位数类别通过累积频数找出。


2. Range and Quartiles | 极差和四分位数

Range = largest value – smallest value. It measures the spread of the data. In 3, 5, 7, 7, 9, the range is 9 – 3 = 6.

极差 = 最大值 – 最小值。它衡量数据的分散程度。在 3, 5, 7, 7, 9 中,极差为 9 – 3 = 6。

Quartiles split ordered data into four equal parts. The lower quartile (Q₁) is the median of the lower half, the upper quartile (Q₃) is the median of the upper half. The median is Q₂.

四分位数将有序数据分成四个相等部分。下四分位数(Q₁)是下半部分的中位数,上四分位数(Q₃)是上半部分的中位数。中位数就是 Q₂。

For the set 2, 4, 5, 7, 8, 9, 10: lower half = 2,4,5 → Q₁ = 4; upper half = 8,9,10 → Q₃ = 9. Interquartile range (IQR) = Q₃ – Q₁ = 5.

对于数据集 2, 4, 5, 7, 8, 9, 10:下半部分 2,4,5 → Q₁ = 4;上半部分 8,9,10 → Q₃ = 9。四分位距(IQR)= Q₃ – Q₁ = 5。

The IQR is a measure of spread that ignores extremes. It is less affected by outliers than the range.

四分位距是一种忽略极端值的离散程度度量。与极差相比,它受异常值的影响较小。


3. Box Plots | 箱线图

A box plot (box-and-whisker diagram) displays the five-number summary: minimum, Q₁, median, Q₃, maximum. It shows the distribution shape and identifies outliers.

箱线图(盒须图)显示五数总结:最小值、下四分位数、中位数、上四分位数、最大值。它显示分布形状并识别异常值。

Box plots are drawn on a scale. The box spans from Q₁ to Q₃, with a line at the median. Whiskers extend to the minimum and maximum values, unless there are outliers.

箱线图在标尺上绘制。箱子从 Q₁ 延伸到 Q₃,中间有一条中位数线。须线延伸到最小值和最大值,除非存在异常值。

An outlier is usually defined as a value less than Q₁ – 1.5×IQR or greater than Q₃ + 1.5×IQR. Such points are marked individually.

异常值通常定义为小于 Q₁ – 1.5×IQR 或大于 Q₃ + 1.5×IQR 的值。这些点应单独标记。

Box plots make it easy to compare two or more distributions side by side, showing central tendency, spread, and skewness.

箱线图便于并排比较两个或多个分布,展示集中趋势、离散程度和偏斜情况。

When interpreting a box plot, a longer box means more variability in the middle half of the data. A median closer to Q₁ suggests positive skew.

在解读箱线图时,较长的箱子表示数据中间一半的变异性较大。中位数靠近下四分位数说明数据呈现正偏态。


4. Frequency Tables and Histograms | 频数表与直方图

A frequency table records how often each value or group occurs. Ungrouped data lists individual values; grouped data uses class intervals.

频数表记录每个值或每个组出现的次数。未分组数据列出单个值;分组数据则使用组距。

Histograms display grouped frequency data. Unlike bar charts, histogram bars touch, and the area of each bar is proportional to frequency. For unequal class widths, use frequency density = frequency ÷ class width.

直方图展示分组频数数据。与条形图不同,直方图的条形彼此相连,且每个条形的面积与频数成正比。组距不等时,需使用频数密度 = 频数 ÷ 组距。

When constructing a histogram, the vertical axis is frequency density. This ensures that the total area equals total frequency.

在构建直方图时,纵轴为频数密度。这样可以确保总面积等于总频数。

For example, if a class 0 ≤ x < 10 has frequency 8, frequency density = 8/10 = 0.8. A class 10 ≤ x < 15 with frequency 6 gives density = 6/5 = 1.2.

例如,若组 0 ≤ x < 10 的频数为 8,则频数密度 = 8/10 = 0.8。组 10 ≤ x < 15 的频数为 6,密度 = 6/5 = 1.2。

Interpret histograms by noting the modal class (tallest bar) and the shape of the distribution (symmetrical, skewed, bimodal).

解读直方图时要注意众数类别(最高的条形)和分布形状(对称、偏斜、双峰)。


5. Cumulative Frequency | 累积频数

Cumulative frequency is the running total of frequencies. It helps find the median, quartiles, and percentiles.

累积频数是频数的逐次累加总和。它有助于找出中位数、四分位数和百分位数。

To draw a cumulative frequency diagram, plot cumulative frequency against the upper class boundary. Points are joined with a smooth curve.

要绘制累积频数图,需要以各组的组上限为横坐标,累积频数为纵坐标描点,然后用光滑曲线连接。

From the graph, the median is found by reading the value at half the total frequency. The lower quartile is at one-quarter of the total frequency; the upper quartile at three-quarters.

从图中找到中位数的方法是:在总频数一半处作水平线交曲线,再垂直向下读取数值。下四分位数在总频数四分之一处,上四分位数在四分之三处。

The interquartile range can be read directly from the graph. Box plots can then be drawn using these estimates.

四分位距可以直接从图中读取。然后可以根据这些估计值绘制箱线图。

Always check that the cumulative frequency curve starts at the first lower boundary with frequency 0 and ends at the total frequency.

务必检查累积频数曲线是否从第一个组下限、频数为 0 开始,并结束于总频数。


6. Scatter Diagrams and Correlation | 散点图与相关性

A scatter diagram plots two variables to see if there is a relationship. Each point represents a pair of values (x, y).

散点图用两个变量绘图,观察它们之间是否存在关系。每个点代表一对数值(x, y)。

Correlation describes the direction and strength of a relationship. Positive correlation: as x increases, y tends to increase. Negative correlation: as x increases, y tends to decrease. No correlation means no clear pattern.

相关性描述关系的方向和强度。正相关:x 增加时,y 往往增加。负相关:x 增加时,y 往往减少。无相关意味着没有明确的模式。

A line of best fit (trend line) can be drawn by eye through the points. It should have roughly equal numbers of points above and below, and pass through the mean point (mean of x, mean of y) if possible.

最佳拟合线(趋势线)可以用肉眼画过各点。应力争使线上下方的点数大致相等,并尽可能通过平均点(x 的平均值,y 的平均值)。

The line can be used to estimate values. Interpolation is estimating within the data range; extrapolation is estimating outside the range and is less reliable.

这条线可用于估计数值。内插是指在数据范围内估计;外推是指在范围外估计,可靠性较低。

Correlation does not imply causation. Always consider possible lurking variables.

相关关系并不意味着因果关系。始终需要考虑潜在的混杂变量。


7. Probability Basics | 概率基础

Probability is a measure of how likely an event is, on a scale from 0 (impossible) to 1 (certain). The probability of an event A is written as P(A) = number of favourable outcomes / total number of possible outcomes, if all outcomes are equally likely.

概率是衡量事件发生可能性的尺度,范围从 0(不可能)到 1(必然)。如果所有结果等可能发生,事件 A 的概率记为 P(A) = 有利结果数 / 可能结果总数。

For a fair six-sided die, P(rolling a 3) = 1/6. P(rolling an even number) = 3/6 = 1/2.

对于一个公平的六面骰子,P(掷出 3) = 1/6。P(掷出偶数) = 3/6 = 1/2。

The sum of probabilities of all mutually exclusive outcomes is 1. The probability of an event not happening is 1 – P(event).

所有互斥结果的概率之和为 1。事件不发生的概率是 1 – P(事件)。

For combined events, we often use Venn diagrams, tree diagrams, or space diagrams.

对于组合事件,我们常使用维恩图、树状图或样本空间图。


8. Tree Diagrams | 树状图

A tree diagram shows all possible outcomes of two or more events, with branches representing probabilities. Probabilities multiply along branches.

树状图显示两个或多个事件的所有可能结果,分枝代表概率。沿分枝进行概率相乘。

To find the probability of a particular sequence, multiply the probabilities on the relevant path. For example, a bag with 3 red and 2 blue balls: P(red then blue with replacement) = (3/5) × (2/5) = 6/25.

要求某一特定序列的概率,将相关路径上的概率相乘。例如,一个袋子里有 3 个红球和 2 个蓝球:有放回地抽取,P(先红后蓝) = (3/5) × (2/5) = 6/25。

Without replacement, the probabilities change after each event. The second event’s probability depends on what happened first (conditional probability).

如果是不放回抽取,每次事件后概率都会改变。第二次事件的概率取决于第一次的结果(条件概率)。

Tree diagrams help avoid errors with “and” and “or” rules. For “or” events on different branches, add the final probabilities of those branches.

树状图有助于避免”且”和”或”规则出错。对于不同分枝上的”或”事件,将这些分枝的最终概率相加。

Always check that the sum of probabilities on branches from a point equals 1.

务必检查从同一点出发的各枝概率之和是否为 1。


9. Combined Events and Conditional Probability | 组合事件与条件概率

The probability that both A and B occur: P(A and B) = P(A) × P(B) if A and B are independent. If they are not independent, use P(A and B) = P(A) × P(B|A).

如果 A 和 B 独立,两者都发生的概率为 P(A 且 B) = P(A) × P(B)。如果不独立,则使用 P(A 且 B) = P(A) × P(B|A)。

The probability that A or B (or both) occurs: P(A or B) = P(A) + P(B) – P(A and B). If mutually exclusive, P(A and B) = 0.

A 或 B(或两者都)发生的概率为 P(A 或 B) = P(A) + P(B) – P(A 且 B)。如果互斥,则 P(A 且 B) = 0。

Conditional probability P(B|A) means the probability of B given that A has occurred. From a two-way table or tree diagram, identify the restricted sample space.

条件概率 P(B|A) 是指在 A 已经发生的条件 B 发生的概率。使用双向表或树状图,确定受限的样本空间。

For example, in a class of 30, 12 are boys, 8 boys have glasses. P(glasses|boy) = 8/12 = 2/3.

例如,在一个 30 人的班级中,12 人是男生,8 名男生戴眼镜。P(戴眼镜|男生) = 8/12 = 2/3。

Venn diagrams can be used to visualise combined events and conditional probabilities. The intersection shows “A and B”.

维恩图可用于可视化组合事件和条件概率。交集表示”A 且 B”。

Always note the difference between P(A|B) and P(B|A). They are not equal in general.

务必注意 P(A|B) 和 P(B|A) 的区别。它们通常不相等。


Published by TutorHao | Maths Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading