Cumulative Frequency: Tables, Curves, and Median Estimation | 累积频率:表格、曲线与中位数估计

📚 Cumulative Frequency: Tables, Curves, and Median Estimation | 累积频率:表格、曲线与中位数估计

Cumulative frequency is a cornerstone of statistical analysis at IGCSE and A-Level. It transforms a simple frequency table into a powerful tool for estimating medians, quartiles, and percentiles, and it provides a visual platform for comparing two or more distributions through the famous ‘ogive’ curve.

累积频率是 IGCSE 和 A-Level 统计学中的基石内容。它将简单的频数表转化为一种强力工具,用于估计中位数、四分位数和百分位数,并通过著名的”卵形线”(ogive curve)为比较两个或多个分布提供了可视化平台。


1. What Is Cumulative Frequency? | 什么是累积频率?

Cumulative frequency is the running total of frequencies as we move through ordered class intervals. For any given class, it tells us how many data points lie at or below the upper boundary of that class. It is always a non-decreasing sequence that ends with the total number of observations, n.

累积频率是当我们按顺序经过各组区间时频数的累加总和。对于任意一组,它告诉我们有多少个数据点位于该组的上组界处或以下。它始终是一个不递减的序列,最终结束于观测总数 n。

Unlike a simple frequency, which counts observations within one interval, a cumulative frequency answers the question: ‘How many values are less than or equal to this point?’ This makes it essential for locating positions such as the median (the 50th percentile) and quartiles.

与仅统计一个区间内观测次数的简单频数不同,累积频率回答的问题是:”有多少个数值小于或等于这一点?”这使得它对于定位中位数(第 50 百分位数)和四分位数等位置至关重要。

  • It is a running total, never a count within a single class.
  • 它是累加总和,绝不是单组内的计数。
  • It never decreases as the value of the variable increases.
  • 随着变量值的增大,它绝不减小。
  • Its final entry always equals the total frequency n.
  • 它的最后一项始终等于总频数 n。

2. Constructing a Cumulative Frequency Table | 构建累积频率表

To build a cumulative frequency table, arrange the class intervals in ascending order, then add the frequencies one by one. Against the upper boundary of each class, write the sum of all frequencies from the first class up to and including that class.

要构建累积频率表,请将组区间按升序排列,然后逐项累加频数。在每组的”上组界”处,写下从第一组到该组(含该组)所有频数之和。

Consider the heights of 40 students recorded in the table below:

考虑下表记录的 40 名学生的身高数据:

Height (cm) Frequency f Cumulative Frequency cf
140 ≤ h < 145 4 4
145 ≤ h < 150 9 13
150 ≤ h < 155 12 25
155 ≤ h < 160 10 35
160 ≤ h < 165 5 40

Notice that 145 is the upper boundary of the first class; at that point, 4 students are shorter than 145 cm. At 150 cm, the cumulative total is 4 + 9 = 13 students, meaning 13 students are shorter than 150 cm, and so on. The final cumulative frequency 40 equals the total number of students.

注意 145 是第一组的上组界;在这一点上,有 4 名学生身高低于 145 cm。在 150 cm 处,累积总数为 4 + 9 = 13 名学生,也就是说有 13 名学生身高低于 150 cm,依此类推。最后的累积频率 40 等于学生总数。

Important: When classes are continuous, cumulative frequency must be plotted against the upper class boundary, not the mid-point or lower boundary. Discontinuous intervals such as ’10–14′ take 14.5 as the upper boundary for plotting purposes.

重要提示:当组为连续数据时,累积频率必须对应上组界进行绘图,而不是组中值或下组界。对于不连续区间(如”10–14″),绘图时应取 14.5 作为上组界。


3. Plotting the Cumulative Frequency Curve (Ogive) | 绘制累积频率曲线(卵形线)

An ogive is a graph of cumulative frequency against the upper boundary of each class. The points are plotted and then joined with a smooth curve. This curve is always increasing (or level), starting from the lower boundary of the first class where the cumulative frequency is zero, and rising to the point (final upper boundary, n).

卵形线是以各组上组界为横坐标、累积频率为纵坐标绘制而成的图形。将各点标出后用平滑曲线连接。这条曲线始终递增(或持平),从第一组下组界处(累积频率为零)开始,一直上升到(最终上组界, n)这一点。

For the heights example, the plotting points are (145, 4), (150, 13), (155, 25), (160, 35), and (165, 40). The curve begins at the lower boundary of the first class, 140, with cumulative frequency 0.

对于身高示例,绘图点为 (145, 4)、(150, 13)、(155, 25)、(160, 35) 和 (165, 40)。曲线从第一组的下组界 140、累积频率为 0 处开始。

  • Use a continuous scale on the x-axis, starting at the first lower boundary.
  • 横轴使用连续刻度,从第一个下组界开始。
  • Plot each point at the upper boundary of its class.
  • 将每个点绘制在其组的上组界处。
  • Join the points with a smooth freehand curve, not straight line segments.
  • 用平滑的手绘曲线连接各点,而不是直线线段。

Cumulative frequency → vertical axis; Upper boundary → horizontal axis

纵轴:累积频率;横轴:上组界


4. Estimating the Median from the Ogive | 从卵形线估计中位数

The median is the value that splits the data into two equal halves. For n observations, the median is located at the \(\frac{n}{2}\)th position (or the \(\frac{n+1}{2}\)th value for raw data; for a grouped data ogive, we use \(\frac{n}{2}\)).

中位数是将数据分为两个相等半部分的值。对于 n 个观测值,中位数位于第 \(\frac{n}{2}\) 个位置(对于原始数据为第 \(\frac{n+1}{2}\) 个值;对于分组数据的卵形线,我们使用 \(\frac{n}{2}\))。

Wait, I need to avoid LaTeX. Let me rephrase without \frac.

The median is the value that splits the data into two equal halves. For n observations, the median position is located at n ÷ 2. On the ogive, locate n ÷ 2 on the vertical (cumulative frequency) axis, draw a horizontal line to the curve, and then drop a vertical line to the horizontal axis. The value where this vertical line meets the x-axis is the estimated median.

中位数是将数据分成两个相等部分的值。对于 n 个观测值,中位数位置位于 n ÷ 2 处。在卵形线上,在纵轴(累积频率)上找到 n ÷ 2 的位置,向曲线画一条水平线,然后向横轴画一条垂直线。这条垂直线与 x 轴相交处的值即为估计的中位数。

Using our height example, n = 40, so the median position is 40 ÷ 2 = 20. On the ogive, find 20 on the y-axis, trace horizontally to the curve, then vertically downward. The x-coordinate of this point is approximately 152 cm. This means roughly half the students are shorter than 152 cm and half are taller.

在我们的身高示例中,n = 40,因此中位数位置为 40 ÷ 2 = 20。在卵形线上,在 y 轴上找到 20,水平追踪到曲线,然后垂直向下。该点的 x 坐标约为 152 cm。这意味着大约一半学生身高低于 152 cm,一半高于 152 cm。

Note that the median from an ogive is an estimate, not an exact value, because we have lost the original raw data and are working from grouped intervals.

请注意,通过卵形线得到的中位数是一个估计值,而非精确值,因为我们已经丢失了原始数据,只能基于分组区间工作。


5. Quartiles and Interquartile Range | 四分位数与四分位距

Quartiles divide the data into four equal parts. The lower quartile (Q₁) is the value below which 25% of the data lies; the upper quartile (Q₃) is the value below which 75% of the data lies. These are found from the ogive in the same way as the median, but using positions n ÷ 4 and 3n ÷ 4 respectively.

四分位数将数据分为四个相等的部分。下四分位数(Q₁)是 25% 数据所在的阈值;上四分位数(Q₃)是 75% 数据所在的阈值。这些值以与中位数相同的方式从卵形线中找到,但分别使用位置 n ÷ 4 和 3n ÷ 4。

The interquartile range (IQR) is defined as Q₃ − Q₁. It measures the spread of the middle 50% of the data and is more robust to outliers than the full range.

四分位距(IQR)定义为 Q₃ − Q₁。它衡量中间 50% 数据的离散程度,并且比全距对异常值更稳健。

IQR = Q₃ − Q₁

For our height example, Q₁ is at position 40 ÷ 4 = 10, and Q₃ is at position 3 × 40 ÷ 4 = 30. Reading from the ogive, Q₁ ≈ 147.5 cm and Q₃ ≈ 157 cm, giving IQR ≈ 157 − 147.5 = 9.5 cm.

对于我们的身高示例,Q₁ 位于位置 40 ÷ 4 = 10 处,Q₃ 位于位置 3 × 40 ÷ 4 = 30 处。从卵形线读取,Q₁ ≈ 147.5 cm,Q₃ ≈ 157 cm,因此 IQR ≈ 157 − 147.5 = 9.5 cm。

  • Q₁ position: n ÷ 4 (25th percentile)
  • Q₁ 位置:n ÷ 4(第 25 百分位数)
  • Median position: n ÷ 2 (50th percentile)
  • 中位数位置:n ÷ 2(第 50 百分位数)
  • Q₃ position: 3n ÷ 4 (75th percentile)
  • Q₃ 位置:3n ÷ 4(第 75 百分位数)

6. Comparing Distributions Using Cumulative Frequency | 利用累积频率比较分布

One of the most powerful uses of the ogive is comparing two data sets. By drawing two cumulative frequency curves on the same axes, we can visually compare their medians and spreads. The curve shifted to the right generally represents data with larger values, while a steeper curve indicates less variability.

卵形线最强大的用途之一是比较多组数据集。在同一坐标系上绘制两条累积频率曲线,我们可以直观地比较它们的中位数和离散程度。向右偏移的曲线通常代表数值较大的数据,而更陡峭的曲线表示变异性较小。

When comparing, always report both a measure of central tendency (the median) and a measure of spread (the IQR). For example, ‘Class A has a higher median than Class B (155 cm vs 150 cm), but Class A also has a larger IQR (12 cm vs 8 cm), meaning its results are more varied.’

比较时,务必同时报告集中趋势的度量(中位数)和离散程度的度量(IQR)。例如:”A 班的中位数高于 B 班(155 cm 对 150 cm),但 A 班的 IQR 也更大(12 cm 对 8 cm),这意味着其成绩差异更大。”

To find which group has greater consistency, compare their IQRs: the smaller IQR indicates more consistent data. Additionally, if one curve lies entirely above and to the left of another, that distribution has smaller values throughout.

要判断哪一组一致性更高,比较它们的 IQR:较小的 IQR 表示数据更一致。此外,如果一条曲线完全位于另一条曲线的左上方,则该分布的所有数值都更小。


7. Common Mistakes and Exam Tips | 常见错误与考试技巧

Many students lose marks on cumulative frequency questions due to avoidable errors. Here are the most frequent pitfalls and how to dodge them:

许多学生在累积频率题目上因为可避免的错误而失分。以下是最常见的陷阱及规避方法:

  • Using mid-points instead of upper boundaries. Always plot cumulative frequency against the upper boundary of each class.
  • 使用组中值而非上组界。始终用各组的上组界来绘制累积频率。
  • Forgetting the starting point. Begin the ogive at the lower boundary of the first class with cumulative frequency 0.
  • 忘记起始点。卵形线应从第一组的下组界、累积频率为 0 处开始。
  • Wrong median position. For grouped data, use n ÷ 2, not (n + 1) ÷ 2. For raw data lists, (n + 1) ÷ 2 is correct.
  • 中位数位置错误。对于分组数据,使用 n ÷ 2,而不是 (n + 1) ÷ 2。对于原始数据列表,(n + 1) ÷ 2 才是正确的。
  • Connecting points with straight lines. Always draw a smooth curve through the points.
  • 用直线连接各点。务必用平滑曲线穿过各点。
  • Poor reading accuracy. When estimating the median, use a ruler and read perpendicularly to the axes.
  • 读取精度不足。估计中位数时,使用直尺并垂直于坐标轴读取。

In exams, always label your axes carefully, include the units, and write down the reading position (e.g., ‘n ÷ 2 = 20’) before drawing the guide lines. This shows the examiner your method clearly.

在考试中,务必备注坐标轴、包含单位,并在绘制参考线前写下读取位置(例如”n ÷ 2 = 20″)。这将向考官清晰地展示你的方法。


8. Worked Example | 完整例题解析

The masses of 60 students were recorded and grouped as follows:

60 名学生的体重记录分组如下:

Mass (kg) 45 ≤ m < 50 50 ≤ m < 55 55 ≤ m < 60 60 ≤ m < 65 65 ≤ m < 70 70 ≤ m < 75
Frequency 6 11 15 14 9 5
Cumulative frequency 6 17 32 46 55 60

Step 1: Plot the ogive. Plot points at upper boundaries: (50, 6), (55, 17), (60, 32), (65, 46), (70, 55), (75, 60). Start the curve at (45, 0).

第一步:绘制卵形线。在上组界处描点:(50, 6)、(55, 17)、(60, 32)、(65, 46)、(70, 55)、(75, 60)。从 (45, 0) 处开始画曲线。

Step 2: Find the median. Median position = 60 ÷ 2 = 30. Trace 30 on the y-axis to the curve and read the x-value: approximately 59.4 kg.

第二步:求中位数。中位数位置 = 60 ÷ 2 = 30。在 y 轴上找 30,追踪到曲线并读取 x 值:约为 59.4 kg。

Step 3: Find quartiles. Q₁ position = 60 ÷ 4 = 15, giving Q₁ ≈ 54.2 kg. Q₃ position = 3 × 60 ÷ 4 = 45, giving Q₃ ≈ 64.7 kg.

第三步:求四分位数。Q₁ 位置 = 60 ÷ 4 = 15,得 Q₁ ≈ 54.2 kg。Q₃ 位置 = 3 × 60 ÷ 4 = 45,得 Q₃ ≈ 64.7 kg。

Step 4: Calculate IQR. IQR = Q₃ − Q₁ = 64.7 − 54.2 = 10.5 kg.

第四步:计算 IQR。IQR = Q₃ − Q₁ = 64.7 − 54.2 = 10.5 kg。

Step 5: Interpret. The median mass is 59.4 kg, meaning half the students weigh less than this. The middle 50% of masses span 10.5 kg, indicating a moderate spread.

第五步:解读。中位体重为 59.4 kg,意味着一半学生的体重低于此值。中间 50% 的体重跨度约为 10.5 kg,表明离散程度适中。


9. Practice Questions | 练习题

Test your understanding with these short questions:

通过以下短题测试你的理解:

Question 1. A survey records the time (in minutes) spent by 80 commuters travelling to work. The cumulative frequency table is: (20, 5), (30, 18), (40, 42), (50, 63), (60, 75), (70, 80). Estimate the median travel time.

第 1 题。一项调查记录了 80 名通勤者的通勤时间(分钟)。累积频率表为:(20, 5)、(30, 18)、(40, 42)、(50, 63)、(60, 75)、(70, 80)。估计通勤时间的中位数。

Solution: Median position = 80 ÷ 2 = 40. The 40th value lies at cumulative frequency 40, which falls between 30 and 40 minutes. From the ogive, the median is approximately 39 minutes.

解答:中位数位置 = 80 ÷ 2 = 40。第 40 个值位于累积频率 40 处,介于 30 和 40 分钟之间。从卵形线读取,中位数约为 39 分钟。

Question 2. For the same data, find Q₁ and Q₃, then compute the interquartile range.

第 2 题。对于同一组数据,求 Q₁ 和 Q₃,并计算四分位距。

Solution: Q₁ position = 80 ÷ 4 = 20, giving Q₁ ≈ 31 minutes. Q₃ position = 3 × 80 ÷ 4 = 60, giving Q₃ ≈ 48.5 minutes. IQR = 48.5 − 31 = 17.5 minutes.

解答:Q₁ 位置 = 80 ÷ 4 = 20,得 Q₁ ≈ 31 分钟。Q₃ 位置 = 3 × 80 ÷ 4 = 60,得 Q₃ ≈ 48.5 分钟。IQR = 48.5 − 31 = 17.5 分钟。

Question 3. Why must the ogive pass through the point (lower boundary of first class, 0)?

第 3 题。为什么卵形线必须经过(第一组下组界, 0)这一点?

Solution: Before the first class, no observations have been counted, so the cumulative frequency is

Published by TutorHao | Mathematics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading