📚 PDF资源导航

IGCSE Edexcel Maths: Statistics – Comprehensive Revision Notes | IGCSE Edexcel 数学:统计考点精讲

📚 IGCSE Edexcel Maths: Statistics – Comprehensive Revision Notes | IGCSE Edexcel 数学:统计考点精讲

Statistics is a fundamental part of the IGCSE Edexcel Mathematics syllabus, equipping students with essential tools to collect, represent, analyse, and interpret data. This comprehensive revision guide covers all key topics, from data types and charts to cumulative frequency, histograms, scatter diagrams, and sampling methods. Each section provides clear explanations paired with step‑by‑step examples, ensuring you can approach exam questions with confidence.

统计是 IGCSE Edexcel 数学大纲的核心组成部分,帮助学生掌握收集、表示、分析和解读数据的基本工具。本篇精讲覆盖所有关键考点,从数据类型和图表到累积频率、直方图、散点图及抽样方法。每个部分均提供清晰的解释和分步示例,确保你能自信地应对考试题目。


1. Types of Data and Data Collection | 数据类型与数据收集

Data is classified into two broad categories: categorical (qualitative) and numerical (quantitative). Categorical data describe qualities, such as hair colour or types of pet. Numerical data represent quantities and are further divided into discrete and continuous. Discrete data can only take specific separate values, often whole numbers (e.g. number of goals scored), while continuous data can take any value within a range (e.g. height, mass, temperature).

数据分为两大类:分类数据(定性)和数值数据(定量)。分类数据描述的是属性,如头发颜色或宠物种类。数值数据表示数量,并进一步分为离散数据和连续数据。离散数据只能取特定的独立值,通常是整数(例如进球数),而连续数据可以在一定范围内取任意值(例如身高、质量、温度)。

Reliable data collection is essential for valid conclusions. Common methods include surveys (using questionnaires or interviews), experiments (controlled investigations), and observational studies. In surveys, questions must be clear, unbiased, and relevant. A pilot survey can help refine the questions before the main data collection.

可靠的数据收集是得出有效结论的基础。常用的方法包括调查(使用问卷或访谈)、实验(对照研究)和观察研究。在调查中,问题必须清晰、无偏见且相关。先进行试点调查有助于在正式收集数据前完善问题。

When designing an experiment or survey, it is important to consider the target population and ensure the sample represents it fairly. Poorly chosen samples lead to biased results.

在设计实验或调查时,考虑目标人群并确保样本能够公平地代表整体是非常重要的。样本选择不当会导致结果有偏差。


2. Frequency Tables and Charts | 频率表与图表

Once data is collected, it is often organised into frequency tables. A frequency table lists each data value (or category) alongside its frequency—the number of times it occurs. For categorical data, a simple tally chart can be used to count occurrences before constructing a bar chart or pie chart.

收集数据后,通常将其整理成频率表。频率表列出每个数据值(或类别)及其频率,即它出现的次数。对于分类数据,可以使用简单的划记表来计数,然后绘制条形图或饼图。

Bar charts represent categorical or discrete data with bars of equal width; the height of each bar corresponds to the frequency. Pie charts display data as sectors of a circle, where each sector angle is calculated as (frequency ÷ total frequency) × 360°. Pictograms use symbols to represent a certain number of units, making data visually appealing.

条形图用等宽的条形表示分类或离散数据,每个条形的高度对应于频率。饼图将数据表示为圆的扇形,每个扇形的角度通过(频率 ÷ 总频率)× 360° 计算得出。象形图使用符号表示一定数量的单位,使数据更具视觉吸引力。

For discrete numerical data, a frequency diagram (vertical line chart) can also be used. In all chart work, remember to label axes clearly and give the chart a title.

对于离散数值数据,也可以使用频率图(垂直线图)。在所有图表作业中,请记住清晰标注坐标轴并给图表加上标题。


3. Measures of Central Tendency | 集中趋势的测量

Measures of central tendency describe the centre of a data set. The three main averages are the mean, median, and mode. Choosing the most appropriate average depends on the data type and the presence of outliers.

集中趋势的测量描述数据集的中心位置。三个主要的平均数是均值、中位数和众数。选择最合适的平均数取决于数据类型及是否存在异常值。

The mode is the value that appears most frequently. A data set may have one mode (unimodal), more than one mode (multimodal), or no mode at all. The mode is the only average that can be used for categorical data.

众数是出现频率最高的值。一个数据集可能有一个众数(单峰)、多个众数(多峰)或根本没有众数。众数是唯一可用于分类数据的平均数。

The median is the middle value when data is arranged in order. For an odd number of values, it is the exact middle; for an even number, it is the mean of the two middle values. The median is unaffected by extreme values, making it useful for skewed distributions.

中位数是将数据排序后位于中间的值。若数据个数为奇数,则取正中间的值;若为偶数,则取中间两个值的平均数。中位数不受极端值的影响,因此在偏态分布中非常有用。

The mean is calculated as the sum of all values divided by the number of values. It is sensitive to every data point and can be pulled in the direction of outliers.

Mean = Σx / n

均值用所有值的总和除以值的个数来计算。它对每一个数据点都敏感,并可能被异常值拉向极端方向。

In the formula, Σx represents the sum of all data values and n is the total frequency. Always check whether a calculated mean is reasonable in the context of the data.

在公式中,Σx 表示所有数据值的总和,n 是总频率。务必检查计算出的均值在数据背景下是否合理。


4. Measures of Spread | 离散程度的测量

Measures of spread indicate how spread out the data values are. The simplest measure is the range, which is the difference between the largest and smallest values. However, the range is heavily affected by outliers.

离散程度的测量显示数据值分散的程度。最简单的度量是极差,即最大值与最小值之差。但极差极易受异常值的影响。

A more robust measure is the interquartile range (IQR). The IQR is the difference between the upper quartile (Q₃) and the lower quartile (Q₁). It represents the spread of the middle 50% of the data and ignores extreme values.

一个更稳健的度量是四分位距(IQR)。四分位距是上四分位数(Q₃)与下四分位数(Q₁)之差,它代表中间 50% 数据的散布范围,并忽略极端值。

To find quartiles, first sort the data. The lower quartile is the median of the lower half, and the upper quartile is the median of the upper half. If the position is not an integer, use the corresponding value or the mean of adjacent values as per your syllabus convention.

要找到四分位数,首先对数据排序。下四分位数是下半部分数据的中位数,上四分位数是上半部分数据的中位数。若位置不是整数,则根据大纲惯例使用相应的值或相邻值的平均数。

Together, the median and IQR give a good picture of the data’s centre and variability, especially when outliers are present.

中位数和四分位距结合在一起,能够很好地展示数据的中心和变异性,尤其是在存在异常值时。


5. Grouped Data and Estimated Mean | 分组数据与估算平均值

When data is grouped into class intervals, the exact raw values are lost. To estimate the mean, we assume that all values in an interval are concentrated at the midpoint (class mark). The estimated mean is then calculated as:

Estimated mean = Σ(f × x) / Σf

当数据被归入组距区间时,确切的原始值就丢失了。为了估算均值,我们假设一个区间内的所有值都集中在组中点(组标)上。估算均值的计算公式为:

估算均值 = Σ(f × x) / Σf

where f is the frequency of each class and x is the class midpoint. The midpoint is found by adding the lower and upper boundaries of the interval and dividing by 2.

其中 f 是各组的频率,x 是组中点。组中点通过将区间的下限和上限相加再除以 2 求得。

For continuous data, class boundaries must be used consistently. For example, the interval 10‑20 may have boundaries 10 and 20, giving midpoint 15, but if data is measured to the nearest unit, the true boundaries may be 9.5‑20.5, still yielding midpoint 15. Always check the context.

对于连续数据,必须统一使用组界。例如,区间 10‑20 的组界可能是 10 和 20,组中点为 15;但如果数据精确到最接近的整数值,真实的组界可能是 9.5‑20.5,组中点仍然是 15。务必根据题意判断。

The modal class is the class interval with the highest frequency. The class interval containing the median can be found by identifying where the cumulative frequency reaches half of the total frequency.

众数所在的组是频率最高的组距区间。中位数所在的组可以通过寻找累积频率达到总频率一半时的位置来确定。


6. Cumulative Frequency | 累积频率

Cumulative frequency is the running total of frequencies. To construct a cumulative frequency table, add each frequency to the sum of all previous frequencies. The final entry should equal the total frequency.

累积频率是频率的累加和。要构建累积频率表,将每个频率加到前面所有频率的和上。最后的累积频率应等于总频率。

A cumulative frequency graph (or ogive) is drawn by plotting each upper class boundary against its cumulative frequency and joining the points with a smooth curve. The graph is always increasing.

绘制累积频率图(或累频曲线)时,以每个组的上限为横坐标、对应的累积频率为纵坐标描点,并用光滑曲线连接各点。曲线总是单调上升的。

From the graph, we can estimate the median and quartiles. The median corresponds to the value at 50% of the total frequency, Q₁ at 25%, and Q₃ at 75%. Draw horizontal lines from the appropriate cumulative frequency value to the curve, then drop vertical lines to read the data values.

从图中我们可以估计中位数和四分位数。中位数对应总频率的 50% 处的值,Q₁ 对应 25%,Q₃ 对应 75%。从相应的累积频率值画水平线与曲线相交,再向下画垂直线即可读出数据值。

Percentiles (e.g. the 90th percentile) can be found in a similar way by using the required percentage of the total frequency. Cumulative frequency graphs are also useful for comparing two distributions on the same axes.

百分位数(例如第 90 百分位数)可以用类似的方法通过总频率的所需百分比来求得。累积频率图也非常适合在同一坐标轴上比较两个分布。


7. Box Plots | 箱线图

A box plot (or box-and-whisker diagram) is a graphical summary of a data set using the five-number summary: minimum value, lower quartile (Q₁), median (Q₂), upper quartile (Q₃), and maximum value. It clearly shows the spread and skewness of the data.

箱线图(或盒须图)是利用五数概括法对数据集进行图形化总结的图表,五数包括:最小值、下四分位数(Q₁)、中位数(Q₂)、上四分位数(Q₃)和最大值。它能清晰地显示数据的散布和偏斜情况。

To draw a box plot, draw a horizontal scale and mark the five values. Draw a box from Q₁ to Q₃ with a vertical line at the median. Then draw whiskers from the box to the minimum and maximum values. The length of the box is the interquartile range.

绘制箱线图时,先画一条水平数值标尺并标出五个数值。从 Q₁ 到 Q₃ 画一个矩形盒,并在中位数处画一条垂直线。然后从盒子两端画须线连接到最小值和最大值。盒子的长度就是四分位距。

Box plots make it easy to compare data sets side by side. Outliers, if defined, are sometimes plotted as individual points beyond the whiskers, but the basic IGCSE syllabus typically uses the minimum and maximum as the whisker ends.

箱线图便于并排比较数据集。若定义了异常值,有时会将异常值作为须线以外的单独点标出,但 IGCSE 基础大纲通常将最小值和最大值作为须线的端点。

When interpreting a box plot, note that a smaller box and shorter whiskers indicate less variability. A median closer to Q₁ suggests a right-skewed distribution.

解读箱线图时,注意较小的盒子和较短的须线表明变异性较小。中位数更靠近 Q₁ 则表明分布呈右偏态。


8. Histograms | 直方图

A histogram is a special type of bar chart used for continuous data grouped into class intervals, where the area of each bar is proportional to the frequency. If the class widths are unequal, frequency alone cannot determine the bar height. Instead, we use frequency density:

Frequency density = Frequency ÷ Class width

直方图是一种特殊的条形图,用于表示已被归入组距区间的连续数据,其中每个条形的面积与频率成正比。如果组距宽度不相等,仅用频率无法确定条形的高度,此时我们需要使用频率密度:

频率密度 = 频率 ÷ 组距

To draw a histogram, plot the class boundaries on the horizontal axis and the frequency density on the vertical axis. Draw rectangles whose widths equal the class widths and whose heights equal the frequency densities. The total area of all rectangles equals the total frequency.

绘制直方图时,以组界为横轴,以频率密度为纵轴,绘制宽度等于组距、高度等于频率密度的矩形。所有矩形的总面积等于总频率。

When calculating frequencies from a histogram, multiply the frequency density by the class width for each bar. Always check that class boundaries are correctly identified—there should be no gaps or overlaps between bars.

当从直方图计算频率时,将每个条形的频率密度乘以组距即可得到该组的频率。务必确保组界识别正确——条形之间不应有间隙或重叠。

Histograms are powerful for showing the shape of a distribution, such as symmetry or skewness, and for identifying peaks in the data.

直方图在展示分布形状(如对称性或偏态)以及识别数据的峰值方面非常有效。


9. Scatter Diagrams and Correlation | 散点图与相关性

A scatter diagram (scatter graph) is used to display the relationship between two numerical variables. Each point represents a pair of values. The pattern of points reveals the type and strength of correlation.

散点图(散点图)用于展示两个数值变量之间的关系。每个点代表一对数值。点的分布形态揭示了相关性的类型和强度。

Positive correlation means that as one variable increases, the other tends to increase. Negative correlation indicates that as one variable increases, the other tends to decrease. No correlation means there is no apparent relationship.

正相关意味着当一个变量增加时,另一个变量也倾向于增加。负相关表明当一个变量增加时,另一个变量倾向于减少。无相关则意味着没有明显的关系。

The strength of correlation can be described as strong, moderate, or weak, depending on how closely the points follow a straight line. Correlation does not imply causation—just because two variables are correlated does not mean that one causes the other.

相关性的强度可根据点靠近一条直线的密切程度描述为强、中等或弱。相关性并不意味着因果关系——两个变量相关并不代表其中一个导致了另一个。

A line of best fit (a straight line drawn through the points as evenly as possible) can be used to make predictions. The line should have roughly equal numbers of points above and below it. Predictions made within the range of the data are called interpolations and are considered reliable; predictions outside the range (extrapolation) are less reliable.

最佳拟合线(一条尽可能均匀地通过各点的直线)可用于进行预测。该线上方和下方的点数应大致相等。在数据范围内做出的预测称为内插,被认为是可靠的;在范围之外做出的预测(外推)则不太可靠。

To find the equation of the line of best fit, select two points on the line (not necessarily data points) and calculate the gradient and y‑intercept. The equation can then be used to estimate one variable for a given value of the other.

要求出最佳拟合线的方程,在线段上选取两个点(不一定是原始数据点),计算斜率和 y 截距。然后可以利用该方程,根据一个变量的给定值估算另一个变量的值。


10. Sampling Methods | 抽样方法

Sampling is the process of selecting a subset of individuals from a population to estimate characteristics of the whole population. A good sample should be representative and free from bias. Common sampling techniques include random, stratified, systematic, and convenience sampling.

抽样是从总体中选取一部分个体以估计整体特征的过程。一个好的样本应具有代表性且无偏差。常用的抽样技术包括随机抽样、分层抽样、系统抽样和便利抽样。

In simple random sampling, every member of the population has an equal chance of being selected. This can be done using random number tables or a random number generator. It reduces selection bias but requires a complete list of the population.

在简单随机抽样中,总体中的每个成员被选中的机会均等。这可以通过随机数表或随机数生成器来实现。该方法能减少选择偏差,但需要总体的完整名单。

Stratified sampling divides the population into distinct groups (strata) based on a characteristic (e.g. age, gender), then a random sample is taken from each stratum in proportion to its size. The formula for the number sampled in a stratum is:

Sample size for stratum = (size of stratum ÷ total population) × total sample size

分层抽样根据某种特征(如年龄、性别)将总体划分为不同的组(层),然后按比例从每一层中随机抽取样本。每层样本量的计算公式为:

每层样本量 = (层的大小 ÷ 总体大小) × 总样本量

Systematic sampling selects every k‑th individual after a random starting point, where k = population size ÷ sample size. It is easy to implement but can introduce bias if there is an underlying pattern.

系统抽样在随机确定起点后,每隔 k 个个体抽取一个,其中 k = 总体大小 ÷ 样本大小。该方法易于实施,但如果存在潜在规律,可能会引入偏差。

Convenience sampling (or opportunity sampling) uses individuals who are easily available. It is quick and inexpensive but is highly likely to produce a biased sample and should be used with caution.

便利抽样(或机会抽样)选取的是容易获取的个体。这种方法快速且成本低,但极有可能产生有偏的样本,应谨慎使用。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading