IGCSE AQA Statistics: Formula & Theorem Quick Reference Handbook | IGCSE AQA 统计:公式定理速查手册

📚 IGCSE AQA Statistics: Formula & Theorem Quick Reference Handbook | IGCSE AQA 统计:公式定理速查手册

This handbook collates all essential formulae and theorems required for the IGCSE AQA Statistics examination. Each entry is presented as a concise English explanation immediately followed by its Chinese equivalent, enabling bilingual revision. The content is arranged by topic to mirror the specification, covering descriptive measures, probability, discrete random variables, the binomial distribution, the normal distribution, correlation and regression, sampling, data representation, and index numbers. Mastery of these results will equip you to handle both routine calculations and multi-step problem solving with confidence.

本手册汇集了 IGCSE AQA 统计考试所需的所有核心公式与定理。每个条目先给出简明的英文解释,紧接着提供对应的中文,便于双语复习。内容按考纲主题编排,涵盖描述性度量、概率、离散随机变量、二项分布、正态分布、相关与回归、抽样、数据表示以及指数。熟练掌握这些结论将帮助你自信地应对常规计算和多步骤解题。


1. Measures of Central Tendency | 集中趋势的度量

The arithmetic mean for a sample of size n is given by x̄ = Σx / n. For grouped data, x̄ = Σfx / Σf, where x is the class midpoint and f is the frequency.

样本量为 n 的算术平均数公式为 x̄ = Σx / n。对于分组数据,x̄ = Σfx / Σf,其中 x 为组中值,f 为频数。

The median is the middle value when data are ordered. For n values, the position of the median is (n + 1) / 2. From a cumulative frequency graph, read off the value at n/2.

中位数是排序后位于中间的值。对于 n 个数据,中位数的位置为 (n + 1) / 2。在累积频率图中,从 n/2 处读取对应的值。

The mode is the value that occurs most frequently. From a histogram, the modal class is the interval with the highest frequency density.

众数是出现次数最多的值。在直方图中,众数所在的组是频率密度最高的区间。


2. Measures of Dispersion | 离散程度的度量

The range = maximum value – minimum value. It is quick to compute but sensitive to outliers.

极差 = 最大值 – 最小值。计算快捷,但容易受异常值影响。

For ungrouped data, the sample variance s² = Σ(x – x̄)² / (n – 1). Equivalently, use s² = (Σx² – (Σx)²/n) / (n – 1). The standard deviation s = √(variance).

对于未分组数据,样本方差 s² = Σ(x – x̄)² / (n – 1)。等价公式为 s² = (Σx² – (Σx)²/n) / (n – 1)。标准差 s = √(方差)。

For grouped data, x is the midpoint and Σf (x – x̄)² is divided by Σf – 1 for the sample variance. The calculator formula: s² = (Σfx² – (Σfx)²/Σf) / (Σf – 1).

对于分组数据,x 取组中值,样本方差的分母为 Σf – 1。计算器公式:s² = (Σfx² – (Σfx)²/Σf) / (Σf – 1)。

Percentiles and quartiles: the lower quartile Q₁ is at position (n+1)/4; the upper quartile Q₃ is at 3(n+1)/4. The interquartile range IQR = Q₃ – Q₁. In cumulative frequency diagrams, use n/4 and 3n/4.

百分位数与四分位数:下四分位数 Q₁ 位于 (n+1)/4 处;上四分位数 Q₃ 位于 3(n+1)/4 处。四分位距 IQR = Q₃ – Q₁。在累积频率图中,使用 n/4 和 3n/4 读取。


3. Skewness | 偏度

Data can be symmetrical, positively skewed (long tail to the right), or negatively skewed (long tail to the left). In a positively skewed distribution: mode < median < mean. In a negatively skewed distribution: mean < median < mode.

数据可以是对称、正偏(右尾长)或负偏(左尾长)的。在正偏分布中:众数 < 中位数 < 平均数。在负偏分布中:平均数 < 中位数 < 众数。

A formula for skewness is (3(mean – median)) / standard deviation. A positive value indicates positive skew; a negative value indicates negative skew. Pearson’s coefficient of skewness can also be given by (mean – mode)/s.

偏度的一个公式为 (3(平均数 – 中位数)) / 标准差。正值表示正偏;负值表示负偏。皮尔逊偏度系数也可使用 (平均数 – 众数)/s。

For quartile-based skewness: (Q₃ + Q₁ – 2 median) / (Q₃ – Q₁). Positive indicates positive skew.

基于四分位数的偏度:(Q₃ + Q₁ – 2 中位数) / (Q₃ – Q₁)。正值表示正偏。


4. Probability Basics and Set Notation | 概率基础与集合符号

P(A) = number of favourable outcomes / total number of outcomes, provided outcomes are equally likely. The complement rule: P(not A) = 1 – P(A).

在等可能的结果下,P(A) = 有利结果数 / 总结果数。补集规则:P(非 A) = 1 – P(A)。

Addition rule: P(A ∪ B) = P(A) + P(B) – P(A ∩ B). For mutually exclusive events, P(A ∩ B) = 0, so P(A ∪ B) = P(A) + P(B).

加法规则:P(A ∪ B) = P(A) + P(B) – P(A ∩ B)。对于互斥事件,P(A ∩ B) = 0,因此 P(A ∪ B) = P(A) + P(B)。

Multiplication rule for independent events: P(A ∩ B) = P(A) × P(B). Conditional probability: P(A|B) = P(A ∩ B) / P(B). Rearranged: P(A ∩ B) = P(B) × P(A|B).

独立事件的乘法规则:P(A ∩ B) = P(A) × P(B)。条件概率:P(A|B) = P(A ∩ B) / P(B)。变形:P(A ∩ B) = P(B) × P(A|B)。

Tree diagrams multiply along branches and add across final nodes. Probabilities on each set of branches must sum to 1.

树状图沿分枝相乘,在终端节点相加。每组分枝上的概率之和必须为 1。

In Venn diagrams, the rectangle represents the sample space with total probability 1. A and B are drawn as overlapping circles. P(A only) = P(A) – P(A ∩ B).

在文氏图中,矩形代表样本空间,总概率为 1。A 和 B 画为相交的圆。仅有 A 的概率 = P(A) – P(A ∩ B)。


5. Discrete Random Variables | 离散随机变量

A discrete random variable X takes values x with probabilities P(X=x). ΣP(X=x) = 1 over all possible values.

离散随机变量 X 以概率 P(X=x) 取值 x。对所有可能的值,ΣP(X=x) = 1。

Expected value E(X) = μ = Σ [x · P(X=x)]. This is the long-term average.

期望值 E(X) = μ = Σ [x · P(X=x)]。这是长期平均值。

Variance Var(X) = E[(X – μ)²] = Σ [(x – μ)² · P(X=x)]. It is also calculated as Var(X) = E(X²) – [E(X)]², where E(X²) = Σ [x² · P(X=x)].

方差 Var(X) = E[(X – μ)²] = Σ [(x – μ)² · P(X=x)]。也可通过 Var(X) = E(X²) – [E(X)]² 计算,其中 E(X²) = Σ [x² · P(X=x)]。

The standard deviation of X is σ = √Var(X).

X 的标准差 σ = √Var(X)。

For a linear function, E(aX + b) = aE(X) + b and Var(aX + b) = a² Var(X). Constants are additive for expectation but only scale for variance.

对于线性函数,E(aX + b) = aE(X) + b,Var(aX + b) = a² Var(X)。常数在期望中可加,在方差中仅影响缩放。


6. Binomial Distribution | 二项分布

The binomial distribution models the number of successes in n independent trials, each with probability p of success. X ~ B(n, p).

二项分布描述了 n 次独立试验中成功的次数,每次试验成功概率为 p。记作 X ~ B(n, p)。

Probability of exactly r successes: P(X = r) = C(n, r) p^r (1 – p)^(n – r), where C(n, r) = n! / [r! (n – r)!] is the binomial coefficient.

恰好 r 次成功的概率:P(X = r) = C(n, r) p^r (1 – p)^(n – r),其中 C(n, r) = n! / [r! (n – r)!] 为二项式系数。

Mean of a binomial: E(X) = np. Variance: Var(X) = np(1 – p). Standard deviation: σ = √[np(1 – p)].

二项分布的均值:E(X) = np。方差:Var(X) = np(1 – p)。标准差:σ = √[np(1 – p)]。

Cumulative probabilities: P(X ≤ k) or P(X ≥ k) may be found using tables or the formula for sums. Remember P(X ≥ k) = 1 – P(X ≤ k – 1).

累积概率:P(X ≤ k) 或 P(X ≥ k) 可使用表格或求和公式求得。注意 P(X ≥ k) = 1 – P(X ≤ k – 1)。

Assumptions: fixed number of trials, independent trials, constant probability p, each trial has two outcomes (success/failure).

假设条件:试验次数固定,试验独立,概率 p 恒定,每次试验只有两种结果(成功/失败)。


7. The Normal Distribution | 正态分布

The normal distribution is a continuous symmetric bell-shaped curve defined by its mean μ and standard deviation σ. Notation: X ~ N(μ, σ²).

正态分布是一个由均值 μ 和标准差 σ 确定的连续对称钟形曲线。记作 X ~ N(μ, σ²)。

Standardisation: to find probabilities, convert X to the standard normal Z ~ N(0, 1) using Z = (X – μ) / σ.

标准化:为求概率,将 X 转换成标准正态 Z ~ N(0, 1),使用 Z = (X – μ) / σ。

Probability tables give P(Z < z) for positive z. For negative z, use P(Z < –z) = 1 – P(Z < z). For P(Z > z) = 1 – P(Z < z).

概率表提供正 z 值的 P(Z < z)。对于负 z,使用 P(Z < –z) = 1 – P(Z < z)。对于 P(Z > z) = 1 – P(Z < z)。

To find an unknown mean or standard deviation, set up the Z equation and solve using inverse normal methods.

若要求未知的均值或标准差,建立 Z 方程,利用逆正态方法求解。

Approximately 68% of data lie within μ ± σ, 95% within μ ± 2σ, and 99.7% within μ ± 3σ (the empirical rule).

约 68% 的数据落在 μ ± σ 内,95% 落在 μ ± 2σ 内,99.7% 落在 μ ± 3σ 内(经验法则)。


8. Correlation and Regression | 相关与回归

The product moment correlation coefficient (PMCC) r measures linear correlation. Formula: r = Sxy / √(Sxx Syy), where Sxy = Σxy – (Σx Σy)/n, Sxx = Σx² – (Σx)²/n, Syy = Σy² – (Σy)²/n.

积矩相关系数 (PMCC) r 衡量线性相关。公式:r = Sxy / √(Sxx Syy),其中 Sxy = Σxy – (Σx Σy)/n,Sxx = Σx² – (Σx)²/n,Syy = Σy² – (Σy)²/n。

Interpretation: –1 ≤ r ≤ 1. r close to 1 indicates strong positive linear correlation; r close to –1 indicates strong negative; r near 0 suggests weak or no linear correlation.

解释:–1 ≤ r ≤ 1。r 接近 1 表示强正线性相关;接近 –1 表示强负线性相关;接近 0 表示弱或无线性相关。

Spearman’s rank correlation coefficient uses the ranks of the data rather than raw values. When there are tied ranks, assign the average rank. Spearman’s coefficient is often used for non-linear monotonic relationships.

斯皮尔曼等级相关系数使用数据的等级而非原始值。存在并列时,赋予平均等级。斯皮尔曼系数常用于非线性单调关系。

The equation of the least squares regression line of y on x is y = a + bx, where b = Sxy / Sxx and a = ȳ – b x̄. The regression line passes through (x̄, ȳ).

y 对 x 的最小二乘回归线方程为 y = a + bx,其中 b = Sxy / Sxx,a = ȳ – b x̄。回归线通过点 (x̄, ȳ)。

Extrapolation is using the regression line to predict values outside the range of observed x-values; such predictions may be unreliable.

外推是使用回归线预测观测 x 范围之外的值;此类预测可能不可靠。


9. Sampling Methods | 抽样方法

A simple random sample gives each member of the population an equal chance of selection, using random number tables or generators.

简单随机抽样使总体中每个成员有相等被选中的概率,使用随机数表或生成器实现。

In stratified sampling, the population is divided into distinct strata and a random sample is taken from each, often proportional to stratum size. This ensures representation across subgroups.

在分层抽样中,总体被划分为不同的层,从每一层中随机抽样,通常按层的大小比例抽取。这确保了各子群体的代表性。

Systematic sampling selects every kth element from a list after a random start. Quick but can introduce periodicity bias.

系统抽样从一个随机起点开始,从列表中每隔 k 个元素抽取一个。快速,但可能引入周期性偏差。

Cluster sampling divides the population into clusters, then randomly selects a number of clusters and surveys all members within them. Useful when a complete list of individuals is difficult to obtain.

整群抽样将总体分为群,随机选取若干群,并调查群内所有成员。当难以获得完整的个体名单时很有用。

Quota sampling is non-random; the interviewer selects a predetermined number of people in specific categories. Prone to selection bias but practical for market research.

配额抽样是非随机的;访问员按预定类别选取一定数量的人。容易产生选择偏差,但在市场调查中切实可行。


10. Data Presentation and Histograms | 数据展示与直方图

Histograms display grouped continuous data. The vertical axis is frequency density = frequency / class width. The area of each bar is proportional to frequency.

直方图用于展示分组连续数据。纵轴为频率密度 = 频数 / 组距。每个矩形的面积与频数成正比。

For equal class widths, frequency density is proportional to frequency, and the shape mimics a bar chart with touching bars. For unequal widths, frequency density must be used.

组距相等时,频率密度与频数成正比,图形类似于紧贴的条形图。组距不等时,必须使用频率密度。

Cumulative frequency curves (ogives) plot cumulative frequency against the upper class boundary. They are used to estimate medians, quartiles, and percentiles.

累积频率曲线 (ogive) 将累积频率相对于组上界绘制。用于估计中位数、四分位数和百分位数。

Stem-and-leaf diagrams preserve original data values while showing the shape of the distribution. Key must be provided.

茎叶图在展示分布形态的同时保留原始数据。必须提供图例。

Box-and-whisker plots display the five-number summary: minimum, Q₁, median, Q₃, maximum. Outliers are often plotted as separate points.

箱线图显示五数概括:最小值、Q₁、中位数、Q₃、最大值。异常值通常单独画出。


11. Index Numbers | 指数

An index number measures relative change in a variable over time. Base year index is usually set to 100.

指数衡量变量随时间变化的相对变化。基年指数通常设为 100。

Simple price relative = (current price / base price) × 100.

简单价比 = (当前价格 / 基期价格) × 100。

A weighted aggregate index such as Laspeyres uses base period quantities: Index = Σ(pc q0) / Σ(p0 q0) × 100. Paasche uses current period quantities.

加权总合指数,例如拉氏指数使用基期数量:指数 = Σ(pc q0) / Σ(p0 q0) × 100。帕氏指数使用现期数量。

Chain base index links period-to-period changes. The Retail Price Index (RPI) and Consumer Price Index (CPI) are common real-world examples.

链基指数将逐期变化链接起来。零售物价指数 (RPI) 和消费者物价指数 (CPI) 是常见的现实例子。

If an index is 130, prices have risen by 30% since the base year. Deflating a value: real value = (nominal value / index) × 100.

若指数为 130,则自基年以来价格上涨了 30%。价值平减:实际价值 = (名义价值 / 指数) × 100。


12. Time Series and Moving Averages | 时间序列与移动平均

A time series consists of data collected at regular intervals. It can be decomposed into trend, seasonal, cyclical, and random components.

时间序列由定期收集的数据组成。可分解为趋势、季节、周期和随机成分。

Moving averages smooth out short-term fluctuations to reveal the trend. For quarterly data, a 4-point moving average is centred to align with time points.

移动平均用于平滑短期波动以揭示趋势。对于季度数据,使用 4 点移动平均并居中,以对齐时间点。

Additive model: Data = Trend + Seasonal + Residual. Seasonal effect is constant over time.

加法模型:数据 = 趋势 + 季节 + 残差。季节效应随时间保持不变。

Multiplicative model: Data = Trend × Seasonal × Residual. Seasonal effect changes proportionally with the trend.

乘法模型:数据 = 趋势 × 季节 × 残差。季节效应随趋势成比例变化。

To estimate seasonal variations, subtract the centred moving average (trend) from the raw data in the additive model, or divide in the multiplicative model. Average these across corresponding seasons.

为估计季节变动,在加法模型中用原始数据减去中心移动平均(趋势),在乘法模型中用原始数据除以中心移动平均。在相应的季节上取平均值。

Forecasting involves extending the trend and adding back the seasonal component. Caution: extrapolation beyond the data range carries uncertainty.

预测涉及延伸趋势并加回季节成分。注意:超越数据范围的外推存在不确定性。

Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version