📚 Year 10 CAIE Statistics: Quick Reference Handbook of Formulas and Theorems | Year 10 CAIE 统计:公式定理速查手册
This handbook provides a concise summary of all the essential formulas, theorems, and statistical measures required for the Year 10 CAIE IGCSE Statistics course. Use it as a quick revision tool to reinforce your understanding of data handling, probability, correlation, and more. Each section pairs definitions and formula statements in English with their corresponding Chinese translations, ensuring bilingual clarity.
这本手册简明扼要地汇总了 Year 10 CAIE IGCSE 统计学课程所需的核心公式、定理与统计度量。可将其作为速查工具,巩固你在数据处理、概率、相关性等方面的知识。每个小节均以中英语对照的方式给出定义与公式,确保双语理解的准确性。
1. Measures of Central Tendency | 集中趋势度量
The arithmetic mean for ungrouped data is the sum of all values divided by the number of values.
未分组数据的算术平均值等于所有数值之和除以数值的个数。
x̄ = Σx / n
For grouped frequency distributions, the mean uses the midpoints of each class (m) and the frequencies (f).
对于分组频数分布,均值使用各组的组中点 (m) 与频数 (f) 计算。
x̄ = Σ(f × m) / Σf
The median is the middle value when data are ordered. For ungrouped data with n values, its position is the (n+1)/2 th value.
中位数是排序后居中的数值。对于有 n 个值的未分组数据,中位数的位置是第 (n+1)/2 个值。
In a grouped frequency table, the median is estimated by linear interpolation inside the median class.
在分组频数表中,中位数通过中位数所在组的线性插值进行估算。
Median = L + ( (n/2 – cfₚᵣₑ) / f_med ) × w
Here L is the lower boundary of the median class, cfₚᵣₑ is the cumulative frequency before the median class, f_med is the frequency of the median class, and w is the class width.
其中 L 是中位数所在组的下限,cfₚᵣₑ 是该组之前的累积频数,f_med 是该组的频数,w 为组距。
The mode is the most frequently occurring value. A frequency distribution may have no mode, one mode, or several modes. For grouped data, the modal class is the class with the highest frequency.
众数是出现次数最多的数值。频数分布可能没有众数、有一个众数或多个众数。对于分组数据,众数组是频数最高的组。
A weighted mean accounts for different importance of values by assigning weights w.
加权平均数通过赋予权重 w 来体现不同数值的重要程度。
Weighted mean = Σ(w × x) / Σw
2. Measures of Dispersion | 离散程度度量
The range is the simplest measure of spread, defined as the difference between the largest and smallest values.
极差(全距)是最简单的离散度量,定义为最大值与最小值之差。
Range = X_max – X_min
The interquartile range (IQR) measures the spread of the middle 50% of the data and is resistant to outliers.
四分位距 (IQR) 衡量中间 50% 数据的分散程度,且不受极端值影响。
IQR = Q₃ – Q₁
For ungrouped data, Q₁ is the value at position (n+1)/4 and Q₃ at position 3(n+1)/4 after ordering.
对于未分组数据,Q₁ 位于排序后第 (n+1)/4 个位置,Q₃ 位于第 3(n+1)/4 个位置。
Variance quantifies the average squared deviation from the mean. For a set of data the variance is given by
方差衡量各数据与均值差异平方的平均值。对于一组数据,方差公式为
Variance (σ²) = Σ(x – x̄)² / n
Note: In some contexts the denominator (n-1) is used for a sample, but the CAIE IGCSE Statistics syllabus typically employs n.
注意:某些情况下样本方差分母为 (n-1),但 CAIE IGCSE 统计学大纲通常使用 n 作为分母。
The standard deviation is the square root of the variance and is expressed in the same units as the original data.
标准差是方差的平方根,单位与原始数据相同。
Standard deviation (s) = √[ Σ(x – x̄)² / n ]
For grouped data, replace x by the class midpoint m and multiply the squared deviation by the frequency f.
对于分组数据,用组中点 m 替代 x,并将平方偏差乘以频数 f。
Grouped variance = Σ[f × (m – x̄)²] / Σf
3. Cumulative Frequency and Box Plots | 累积频率与箱线图
A cumulative frequency curve (ogive) is plotted by adding frequencies sequentially. It allows you to read off the median, quartiles, and percentiles directly from the graph.
累积频率曲线(折线图)通过依次累加频数绘制而成。可以从图上直接读出中位数、四分位数和百分位数。
The median corresponds to the value at a cumulative frequency of n/2, Q₁ at n/4, and Q₃ at 3n/4 on the y-axis.
中位数对应 y 轴上累积频数为 n/2 时的 x 值,Q₁ 对应 n/4,Q₃ 对应 3n/4。
A box-and-whisker plot (box plot) displays the five-number summary: minimum, Q₁, median, Q₃, and maximum.
箱线图(盒须图)展示五数概括:最小值、Q₁、中位数、Q₃ 和最大值。
The box spans Q₁ to Q₃ with a line at the median. Whiskers extend to the minimum and maximum, or to 1.5 × IQR beyond the quartiles when outliers are considered.
箱子从 Q₁ 延伸至 Q₃,并在中位数处画线。须线延伸到最小值和最大值,或者在考虑离群值时延伸到四分位数 1.5 倍 IQR 范围处。
Outliers are data points that lie more than 1.5 × IQR below Q₁ or above Q₃.
离群值是指低于 Q₁ – 1.5×IQR 或高于 Q₃ + 1.5×IQR 的数据点。
4. Basic Probability | 基础概率
The probability of an event A, written P(A), lies between 0 and 1 and can be expressed as a fraction, decimal, or percentage.
事件 A 的概率记为 P(A),取值在 0 到 1 之间,可用分数、小数或百分数表示。
P(A) = Number of favourable outcomes / Total number of equally likely outcomes
For any two events A and B, the addition rule states:
对于任意两个事件 A 和 B,加法法则为:
P(A ∪ B) = P(A) + P(B) – P(A ∩ B)
If A and B are mutually exclusive (cannot occur together), then P(A ∩ B) = 0, so P(A ∪ B) = P(A) + P(B).
若 A 与 B 互斥(不能同时发生),则 P(A ∩ B) = 0,因此 P(A ∪ B) = P(A) + P(B)。
Conditional probability is the probability of A given that B has occurred.
条件概率是在 B 已发生的条件下 A 发生的概率。
P(A|B) = P(A ∩ B) / P(B), provided P(B) > 0
Two events are independent if the occurrence of one does not affect the probability of the other. In that case, P(A ∩ B) = P(A) × P(B) and P(A|B) = P(A).
两个事件独立意味着一事件的发生不影响另一事件的概率。此时 P(A ∩ B) = P(A) × P(B) 且 P(A|B) = P(A)。
Tree diagrams help to map out combined events and multiply along branches for ‘and’ and add across branches for ‘or’ probabilities.
树形图有助于理清复合事件,沿分支相乘求“且”的概率,跨分支相加求“或”的概率。
5. Permutations and Combinations | 排列与组合
Factorial notation n! means the product of all positive integers from n down to 1.
阶乘符号 n! 表示从 n 到 1 的所有正整数的乘积。
n! = n × (n-1) × (n-2) × … × 2 × 1, and 0! = 1
A permutation is an ordered arrangement of r items chosen from n distinct items.
排列是从 n 个不同物品中选取 r 个进行有序排列。
nPr = n! / (n – r)!
A combination is a selection of r items from n distinct items where the order does not matter.
组合是从 n 个不同物品中选取 r 个而不考虑顺序。
nCr = n! / [r! × (n – r)!]
Key relationships: nCr = nPr / r! and nCr = nC(n – r). Also note nC₀ = nCn = 1.
主要关系:nCr = nPr / r!,nCr = nC(n – r)。此外 nC₀ = nCn = 1。
These counting principles underpin many probability problems, especially those involving selections and arrangements.
这些计数原理是许多概率问题的基础,尤其是涉及选取与排列的题目。
6. Correlation and Scatter Diagrams | 相关与散点图
A scatter diagram plots bivariate data (x, y) to reveal the relationship between two variables. The overall pattern may show positive, negative, or no correlation.
散点图将双变量数据 (x, y) 描点,以揭示两变量之间的关系。整体形态可显示正相关、负相关或无相关。
The strength of a linear relationship is measured by the Pearson product-moment correlation coefficient r.
线性关系的强弱由皮尔逊积矩相关系数 r 衡量。
r = Σ(x – x̄)(y – ȳ) / √[ Σ(x – x̄)² × Σ(y – ȳ)² ]
Using shorthand notation: Sxx = Σ(x – x̄)², Syy = Σ(y – ȳ)², Sxy = Σ(x – x̄)(y – ȳ). Then r = Sxy / √(Sxx × Syy).
使用简写符号:Sxx = Σ(x – x̄)²,Syy = Σ(y – ȳ)²,Sxy = Σ(x – x̄)(y – ȳ),则 r = Sxy / √(Sxx × Syy)。
The value of r always lies between -1 and +1. r = +1 indicates perfect positive correlation, r = -1 perfect negative correlation, and r = 0 no linear correlation.
r 的值始终在 -1 与 +1 之间。r = +1 表示完全正相关,r = -1 表示完全负相关,r = 0 表示无线性相关。
It is important to remember that correlation does not imply causation. A line of best fit drawn by eye can approximate the trend.
必须牢记相关不意味着因果。用目测法画出的最佳拟合线可近似描述趋势。
7. Regression Line | 回归线
The least squares regression line of y on x is the straight line that minimises the sum of the squares of the vertical deviations. Its equation is written as
y 对 x 的最小二乘回归线是使垂直偏差的平方和最小的直线,其方程写作
y = a + b x
The slope b is calculated from the data using
斜率 b 由数据通过下式计算:
b = Sxy / Sxx = Σ(x – x̄)(y – ȳ) / Σ(x – x̄)²
The intercept a ensures the line passes through the mean point (x̄, ȳ).
截距 a 确保回归线通过均值点 (x̄, ȳ)。
a = ȳ – b x̄
Once the line is determined, it can be used to estimate y for a given x within the range of the data (interpolation). Extrapolation beyond the data range is less reliable.
确定回归线后,可用它在数据范围内给定 x 估计 y(内插法)。超出数据范围的外推则不太可靠。
Note that the regression line of x on y has a different slope and is not simply the inverse of the above; the independent and dependent variables must be clearly identified.
注意 x 对 y 的回归线具有不同的斜率,并非上述直线的简单反函数;必须明确区分自变量与因变量。
8. Time Series and Moving Averages | 时间序列与移动平均
A time series records data points at successive time intervals. Its components can include trend, seasonal variation, cyclical fluctuation, and random noise.
时间序列记录连续时间间隔上的数据点。其构成可包括趋势、季节变动、循环波动和随机噪声。
The trend is the long-term movement and can be smoothed using a moving average. For a set of n values, a moving average of order k replaces each point with the mean of itself and k-1 neighbours.
趋势是长期运动,可通过移动平均来平滑。对于 n 个值,k 阶移动平均用该点及其 k-1 个邻近点的均值来替换每个点。
Typically, a 3-point or 4-point moving average is used. If k is even, the resulting averages are centred by taking two-period moving averages of the moving averages.
通常使用 3 点或 4 点移动平均。若 k 为偶数,则需对移动平均再求两期移动平均以实现中心化。
Seasonal variation is the regular pattern that repeats over a fixed period (e.g. quarterly sales). It can be estimated by subtracting the trend (additive model) or by dividing by the trend (multiplicative model).
季节变动是固定周期(如季度销售)内重复的规律性模式。可通过减去趋势(加法模型)或除以趋势(乘法模型)来估算。
Additive: Seasonal effect = Actual value – Trend
Multiplicative: Seasonal factor = Actual value / Trend
Forecasting can then be done by extending the trend line and adding or multiplying the appropriate seasonal component.
预测可通过延伸趋势线并加上或乘以合适的季节成分来完成。
9. Index Numbers | 指数
An index number measures the relative change in a variable, such as price or quantity, compared with a base period. The base period index is usually set at 100.
指数衡量某一变量(如价格或数量)相对于基期的相对变化。基期指数通常设为 100。
The simple price relative for a single item is
单一商品的简单价比为
Price relative = (P₁ / P₀) × 100
where P₁ is the current price and P₀ is the base-period price.
其中 P₁ 为现期价格,P₀ 为基期价格。
An unweighted aggregate index uses the same base for all items:
未加权综合指数对所有商品使用相同的基期:
Simple aggregate index = (ΣP₁ / ΣP₀) × 100
Weighted indices incorporate quantities (q) to reflect relative importance. The Laspeyres index uses base-period quantities q₀, while the Paasche index uses current-period quantities q₁.
加权指数引入数量 (q) 以反映相对重要性。拉氏指数使用基期数量 q₀,而派氏指数使用现期数量 q₁。
Laspeyres price index = ( Σ(P₁ × q₀) / Σ(P₀ × q₀) ) × 100
<
Published by TutorHao | Year 10 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导