IGCSE Cambridge Statistics: Core Knowledge Review | IGCSE Cambridge 统计:核心知识点梳理

📚 IGCSE Cambridge Statistics: Core Knowledge Review | IGCSE Cambridge 统计:核心知识点梳理

The IGCSE Cambridge Statistics syllabus equips students with essential skills in collecting, presenting, analysing, and interpreting data. A solid grasp of core concepts is crucial for success in examinations. This article summarises key topics, including data handling, measures of central tendency and dispersion, probability, correlation, and more. By understanding these fundamentals, you can tackle statistical problems with confidence.

IGCSE Cambridge 统计课程培养学生收集、呈现、分析和解读数据的基本能力。扎实掌握核心概念对考试成功至关重要。本文梳理了数据处理、集中趋势与离散程度度量、概率、相关性等关键知识点。理解这些基础,你将能自信应对各类统计问题。

1. Types of Data and Data Collection | 数据类型与数据收集

Data can be classified as qualitative (categorical) or quantitative (numerical). Quantitative data are further split into discrete data, which can only take specific values (e.g. number of children), and continuous data, which can take any value within a range (e.g. weight).

数据可分为定性(分类)数据和定量(数值)数据。定量数据又分为离散数据,只能取特定值(如孩子数量),以及连续数据,可在一定范围内取任意值(如体重)。

A census collects information from every member of a population. A sample studies a subset of the population, saving time and cost. Common sampling methods include simple random sampling, stratified sampling, systematic sampling, and quota sampling. Random sampling avoids bias, whereas non-random methods may be quicker but risk bias.

普查收集总体中每个成员的信息。抽样研究总体的一部分,节省时间和成本。常见的抽样方法包括简单随机抽样、分层抽样、系统抽样和配额抽样。随机抽样可避免偏差,而非随机方法可能更快但有偏差风险。

Data can be primary (collected directly by you) or secondary (already available from other sources). Each source has advantages in reliability and relevance.

数据可以是原始数据(自己直接收集)或二手数据(从其他来源获得)。每种来源在可靠性和相关性上各有优势。


2. Frequency Distributions and Charts | 频数分布与图表

A frequency table organises raw data into groups, showing how often each value or interval occurs. For grouped data, class intervals define the groups. A cumulative frequency column adds frequencies progressively.

频数表将原始数据分组,显示每个取值或区间出现的次数。对于分组数据,组距定义了各小组。累积频率列逐步累加各组的频率。

Bar charts display categorical data with equal-width bars; the height represents frequency or relative frequency. Pie charts show proportions as sectors of a circle, where the angle = (category frequency ÷ total frequency) × 360°.

条形图用等宽的长条展示分类数据,高度表示频数或相对频率。饼图用扇形表示比例,扇形的角度 =(类别频数 ÷ 总频数)× 360°。

Stem-and-leaf diagrams preserve original data and reveal shape. A back-to-back stem-and-leaf plot compares two datasets. Line graphs are used for time series and show trends.

茎叶图保留原始数据并显示分布形状。背靠背茎叶图可比较两组数据。折线图用于时间序列,展示趋势。

Relative frequency = Frequency ÷ Total frequency


3. Measures of Central Tendency | 集中趋势的度量

The mean (x̄) is the arithmetic average, calculated as sum of all values divided by the number of values. For grouped data, use midpoints: Mean = Σ(fx) ÷ Σf, where f is frequency and x is midpoint.

均值(x̄)是算术平均数,等于所有值的总和除以值的个数。对于分组数据,使用组中值:均值 = Σ(fx) ÷ Σf,其中f为频数,x为组中值。

Mean = Σx / n

The median is the middle value when data are ordered. For n values, position = (n + 1)/2. If n is even, take the mean of the two middle values. The mode is the most frequent value. The mean is sensitive to outliers; median is robust.

中位数是数据排序后的中间值。对于n个数据,中位数的位置 = (n + 1)/2。若n为偶数,取中间两个数的平均值。众数是出现次数最多的值。均值易受极端值影响,中位数则较为稳健。

In a symmetrical distribution, mean, median and mode coincide. In a skewed distribution, the mean is pulled towards the tail.

在对称分布中,均值、中位数和众数重合。在偏态分布中,均值会被拉向尾部。


4. Measures of Dispersion | 离散程度的度量

Range = Maximum value – Minimum value. It is simple but affected by outliers.

极差 = 最大值 – 最小值。计算简单,但受极端值影响。

Range = Max – Min

The interquartile range (IQR) = Upper quartile (Q₃) – Lower quartile (Q₁). It measures the spread of the middle 50% and is resistant to outliers.

四分位距(IQR)= 上四分位数(Q₃) – 下四分位数(Q₁)。它衡量中间50%数据的离散程度,不受极端值影响。

Standard deviation (σ for population, s for sample) is the most widely used measure. For raw data, s = √[ Σ(x – x̄)² / (n – 1) ]. A larger standard deviation indicates greater spread.

标准差(总体用σ,样本用 s)是最常用的离散指标。对于原始数据,s = √[ Σ(x – x̄)² / (n – 1) ]。标准差越大,离散程度越高。

s = √[ Σ(x – x̄)² / (n – 1) ]

Variance is the square of standard deviation. An outlier is usually defined as a value more than 1.5 × IQR below Q₁ or above Q₃.

方差是标准差的平方。异常值通常定义为低于 Q₁ – 1.5×IQR 或高于 Q₃ + 1.5×IQR 的数值。


5. Cumulative Frequency and Box Plots | 累积频率与箱线图

A cumulative frequency curve (ogive) is plotted using the upper class boundary and the cumulative frequency. It can be used to estimate the median, quartiles, and percentiles.

累积频率曲线(拱形图)使用组上限和累积频率绘制,可用于估计中位数、四分位数和百分位数。

To find the median, locate the position N/2 on the cumulative frequency axis, draw a horizontal line to the curve, then down to the data axis. Similarly, Q₁ uses N/4 and Q₃ uses 3N/4.

要找出中位数,在累积频率轴上找到 N/2 的位置,画水平线交曲线,再向下交数据轴。同理,Q₁ 用 N/4,Q₃ 用 3N/4。

A box-and-whisker plot displays the minimum, Q₁, median, Q₃, and maximum, and shows skewness and spread at a glance. Outliers, if any, are shown as separate crosses.

箱线图(盒须图)显示了最小值、Q₁、中位数、Q₃ 和最大值,一目了然地呈现偏态和离散程度。若存在异常值,通常用叉号单独标出。


6. Histograms | 直方图

A histogram is used for continuous grouped data. Unlike a bar chart, the area of each bar represents frequency. When class intervals are unequal, frequency density must be calculated.

直方图用于连续型分组数据。与条形图不同,直方图每个矩形的面积代表频数。当组距不等时,必须计算频率密度。

Frequency density = Frequency ÷ Class width

The height of a histogram bar is the frequency density. The total area equals total frequency. Drawing and interpreting histograms require careful reading of axis scales and construction of the frequency density table.

直方图柱的高度为频率密度。总面积等于总频数。绘制和解读直方图需要仔细阅读轴刻度,并构建频率密度表。

To find the median or quartiles from a histogram, you can use linear interpolation within the relevant class, assuming data are evenly spread.

从直方图中求中位数或四分位数,可在相关组内使用线性插值法,假设数据均匀分布。


7. Probability Basics | 概率基础

Probability measures the chance of an event, always between 0 (impossible) and 1 (certain). For equally likely outcomes, P(A) = Number of favourable outcomes ÷ Total number of outcomes.

概率衡量事件发生的可能性,值介于0(不可能)和1(必然)之间。对于等可能结果,P(A) = 有利结果数 ÷ 总结果数。

The complement rule: P(not A) = 1 – P(A). Mutually exclusive events cannot both happen; P(A or B) = P(A) + P(B). For non-mutually exclusive events, subtract the intersection: P(A or B) = P(A) + P(B) – P(A and B).

互补规则:P(非A) = 1 – P(A)。互斥事件不能同时发生,P(A 或 B) = P(A) + P(B)。若事件不互斥,需减去交集:P(A 或 B) = P(A) + P(B) – P(A 且 B)。

Independent events do not affect each other: P(A and B) = P(A) × P(B). The estimated probability from experiments is relative frequency = number of successes ÷ number of trials.

独立事件互不影响:P(A 且 B) = P(A) × P(B)。通过实验估计概率用相对频率 = 成功次数 ÷ 试验次数。


8. Tree Diagrams and Conditional Probability | 树形图与条件概率

Tree diagrams show all possible outcomes of multi-stage trials. Probabilities are written on branches, and final outcomes are obtained by multiplying along paths. The probabilities of all branches from a node sum to 1.

树形图展示多步试验的所有可能结果。概率标在分枝上,沿路径相乘得到最终结果的概率。从同一节点出发的所有分枝概率之和为1。

Conditional probability, P(A|B), is the probability of A occurring given that B has occurred: P(A|B) = P(A and B) ÷ P(B), provided P(B) > 0. Tree diagrams are particularly useful for sequential events, such as drawing without replacement.

条件概率 P(A|B) 是指在事件B发生的条件下A发生的概率:P(A|B) = P(A 且 B) ÷ P(B),其中 P(B) > 0。树形图特别适用于后续事件,例如不放回抽取。

In without-replacement problems, the probabilities on the second set of branches change depending on the first outcome. Always check whether events are independent before multiplying probabilities.

在不放回问题中,第二组分枝上的概率会根据第一次结果而变化。在相乘概率之前,务必检查事件是否独立。


9. Scatter Diagrams and Correlation | 散点图与相关性

A scatter diagram plots bivariate data, each point (x, y). Correlation describes the relationship: positive correlation means as x increases, y tends to increase; negative correlation means as x increases, y tends to decrease; no correlation shows no clear pattern.

散点图用于展示双变量数据,每点坐标为(x, y)。相关性描述变量间的关系:正相关表示x增大时y倾向于增大;负相关表示x增大时y倾向于减小;零相关则无明显趋势。

The strength of correlation can be ‘strong’, ‘moderate’ or ‘weak’. Spearman’s rank correlation coefficient (rₛ) provides a numerical measure between -1 and +1 for monotonic relationships.

相关强度可分为“强”、“中等”或“弱”。斯皮尔曼秩相关系数(rₛ)为单调关系提供介于 -1 与 +1 的数值度量。

rₛ = 1 – [6 Σ d² / (n(n² – 1))]

where d is the difference in ranks. A line of best fit can be drawn by eye, passing through the mean point (x̄, ȳ). For a linear model, the regression equation y = a + bx can be found by calculation or by using given values from the exam. Use the line to estimate values, but extrapolation beyond the data range is unreliable.

其中 d 为秩次之差。通过目测可画出一条最佳拟合直线,它经过均值点 (x̄, ȳ)。对于线性模型,可通过计算或题目给出的值得到回归方程 y = a + bx。利用该直线可进行估计,但超出数据范围的推测不可靠。


10. Time Series and Index Numbers | 时间序列与指数

A time series is a set of data recorded over time. Key components are trend, seasonal variation, and random fluctuations. Moving averages smooth out short-term variations to reveal the underlying trend.

时间序列是按时间顺序记录的数据序列。其主要组成包括趋势、季节变动和随机波动。移动平均法可平滑短期波动,揭示潜在趋势。

To calculate a 3-point moving average, average groups of three consecutive values. For seasonal data, average the moving averages and seasonal components can then be estimated using additive or multiplicative models.

计算三点移动平均时,将连续三个值分组求平均。对于季节性数据,先求移动平均,然后通过加法或乘法模型估计季节分量。

An index number compares the value of a quantity over time relative to a base period. Simple index = (Value in given period ÷ Value in base period) × 100.

指数用于比较某一数量在不同时期相对于基期的变化。简单指数 =(给定时期数值 ÷ 基期数值)× 100。

Price index = (Pₙ / P₀) × 100

Weighted index numbers, such as the Laspeyres index, account for the importance of items. Laspeyres price index = Σ(Pₙ × Q₀) ÷ Σ(P₀ × Q₀) × 100, where P₀ and Q₀ are base prices and quantities. An understanding of how to interpret and calculate such indices is required.

加权指数,如拉斯拜尔指数,考虑了各项的重要性。拉斯拜尔价格指数 = Σ(Pₙ × Q₀) ÷ Σ(P₀ × Q₀) × 100,其中 P₀ 和 Q₀ 为基期价格与数量。考纲要求学生能理解和计算这类指数。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version