Year 10 SQA Statistics: Formula & Theorems Quick Reference Handbook | Year 10 SQA 统计:公式定理速查手册

📚 Year 10 SQA Statistics: Formula & Theorems Quick Reference Handbook | Year 10 SQA 统计:公式定理速查手册

Whether you are preparing for an end‑of‑topic test or a final SQA assessment, having all the key statistical formulas at your fingertips is essential. This quick reference handbook collects the most important definitions, theorems, and formulas you need for Year 10 Statistics. For each item you will find the English explanation immediately followed by its Chinese equivalent, so you can master the content bilingually.

无论你是在准备单元小测验还是最终的 SQA 考试,手边有一份核心统计公式速查表都至关重要。这本速查手册汇集了 Year 10 统计中最重要的定义、定理和公式,每个条目都采用英文解释紧接中文对照的形式,帮助你双语掌握内容。


1. Measures of Central Tendency | 集中趋势的度量

The mean of a set of n numbers is calculated by summing all the values and dividing by n. The formula is: x̄ = Σx / n.

一组 n 个数的平均数等于所有数值之和除以 n,公式为 x̄ = Σx / n。

The median is the middle value when data are ordered. If n is odd, median = the (n+1)/2 th value; if n is even, median = average of the n/2 th and (n/2 + 1) th values.

中位数是排序后位于中间位置的数据值。n 为奇数时,中位数 = 第 (n+1)/2 个值;n 为偶数时,中位数 = 第 n/2 个与第 (n/2+1) 个值的平均数。

The mode is the value that occurs most frequently. A dataset can have no mode, one mode (unimodal), or multiple modes.

众数是出现次数最多的值。一组数据可能没有众数、有一个众数(单峰)或多个众数。

The lower quartile Q₁ is the median of the lower half of the data; the upper quartile Q₃ is the median of the upper half. The interquartile range IQR = Q₃ – Q₁.

下四分位数 Q₁ 是数据下半部分的中位数;上四分位数 Q₃ 是数据上半部分的中位数。四分位距 IQR = Q₃ – Q₁。


2. Measures of Spread | 离散程度的度量

The range is the difference between the maximum and minimum values: Range = xₘₐₓ – xₘᵢₙ.

极差是最大值与最小值之差:极差 = xₘₐₓ – xₘᵢₙ。

The variance (population) is the average of the squared differences from the mean: σ² = Σ(x – μ)² / n. For a sample, use s² = Σ(x – x̄)² / (n – 1).

(总体)方差是各数据与平均数之差的平方的平均值:σ² = Σ(x – μ)² / n。对于样本,使用 s² = Σ(x – x̄)² / (n – 1)。

The standard deviation is the square root of the variance: σ = √[Σ(x – μ)² / n] or s = √[Σ(x – x̄)² / (n – 1)].

标准差是方差的平方根:σ = √[Σ(x – μ)² / n] 或 s = √[Σ(x – x̄)² / (n – 1)]。

For grouped data, use midpoints to estimate the mean and variance: x̄ = Σfx / Σf, and σ² = Σfx² / Σf – x̄².

对于分组数据,用组中值来估算平均数和方差:x̄ = Σfx / Σf,σ² = Σfx² / Σf – x̄²。


3. Probability Basics | 概率基础

The probability of an event A is P(A) = Number of favourable outcomes / Total number of possible outcomes, assuming equally likely outcomes.

事件 A 的概率为 P(A) = 有利结果数 / 所有可能结果的总数,前提是每个结果等可能发生。

The complement rule: P(not A) = 1 – P(A).

互补规则:P(非 A) = 1 – P(A)。

The addition rule for mutually exclusive events: If A and B cannot happen together, then P(A or B) = P(A) + P(B).

互斥事件的加法规则:若 A 和 B 不可能同时发生,则 P(A 或 B) = P(A) + P(B)。

The general addition rule: P(A or B) = P(A) + P(B) – P(A and B).

一般加法规则:P(A 或 B) = P(A) + P(B) – P(A 且 B)。

The multiplication rule for independent events: If A and B are independent, then P(A and B) = P(A) × P(B).

独立事件的乘法规则:若 A 和 B 独立,则 P(A 且 B) = P(A) × P(B)。

Conditional probability: P(A | B) = P(A and B) / P(B), provided P(B) > 0.

条件概率:P(A|B) = P(A 且 B) / P(B),其中 P(B) > 0。


4. Discrete Random Variables | 离散随机变量

A discrete random variable X takes a countable set of values with associated probabilities. The sum of all probabilities must equal 1: ΣP(X = x) = 1.

离散随机变量 X 取有限或可数个值,每个取值对应一个概率,且所有概率之和为 1:ΣP(X = x) = 1。

The expected value (mean) of X is E(X) = Σ x · P(X = x).

X 的期望值(均值)为 E(X) = Σ x · P(X = x)。

The variance of X is Var(X) = E(X²) – [E(X)]², where E(X²) = Σ x² · P(X = x).

X 的方差为 Var(X) = E(X²) – [E(X)]²,其中 E(X²) = Σ x² · P(X = x)。

The standard deviation is σ = √Var(X).

标准差为 σ = √Var(X)。


5. Binomial Distribution | 二项分布

A binomial distribution B(n, p) models the number of successes in n independent trials, each with success probability p. The probability of exactly r successes is given by the binomial formula: P(X = r) = ⁿCᵣ · pʳ · qⁿ⁻ʳ, where q = 1 – p.

二项分布 B(n, p) 描述在 n 次独立试验中成功的次数,每次试验成功概率为 p。恰好成功 r 次的概率由二项公式给出:P(X = r) = ⁿCᵣ · pʳ · qⁿ⁻ʳ,其中 q = 1 – p。

The mean of a binomial distribution is μ = E(X) = n p.

二项分布的均值为 μ = E(X) = n p。

The variance is Var(X) = n p q, and the standard deviation is σ = √(n p q).

方差为 Var(X) = n p q,标准差为 σ = √(n p q)。

The binomial coefficient ⁿCᵣ (also written as n! / [r!(n – r)!]) counts the number of ways to choose r successes from n trials.

二项式系数 ⁿCᵣ(也可写为 n! / [r!(n – r)!])表示从 n 次试验中选出 r 次成功的方法数。

When n is large and p is close to 0.5, the binomial distribution can be approximated by a normal distribution with the same mean and variance, provided np ≥ 5 and nq ≥ 5.

当 n 较大且 p 接近 0.5 时,二项分布可用相同均值和方差的正态分布来近似,前提是 np ≥ 5 且 nq ≥ 5。


6. Normal Distribution | 正态分布

The normal distribution is a continuous probability distribution with a bell‑shaped curve. It is fully described by its mean μ and standard deviation σ. The notation is X ~ N(μ, σ²).

正态分布是一种连续概率分布,呈钟形曲线。它由均值 μ 和标准差 σ 完全确定,记作 X ~ N(μ, σ²)。

The standard normal distribution Z has mean 0 and standard deviation 1: Z ~ N(0, 1). Any normal variable can be standardised using the z‑score formula:

标准正态分布 Z 的均值为 0,标准差为 1:Z ~ N(0, 1)。任意正态变量可通过 z 分数公式进行标准化:

z = (x – μ) / σ

The empirical rule (68–95–99.7 rule) states that approximately 68% of data lie within 1 standard deviation of the mean, 95% within 2 standard deviations, and 99.7% within 3 standard deviations.

经验法则(68–95–99.7 规则)指出,大约 68% 的数据落在均值 ±1 个标准差内,95% 落在 ±2 个标准差内,99.7% 落在 ±3 个标准差内。

To find probabilities for a normal distribution, convert x to a z‑score and use standard normal tables or technology. For example, P(X < a) = P(Z < (a – μ)/σ).

要计算正态分布的概率,先将 x 转换为 z 分数,然后查询标准正态分布表或使用技术工具。例如,P(X < a) = P(Z < (a – μ)/σ)。

The inverse normal function finds the x‑value corresponding to a given cumulative probability.

逆正态函数可求出给定累积概率所对应的 x 值。


7. Scatter Graphs and Correlation | 散点图与相关

A scatter graph displays the relationship between two quantitative variables. The explanatory variable is plotted on the x‑axis and the response variable on the y‑axis.

散点图展示两个定量变量之间的关系。解释变量画在 x 轴上,响应变量画在 y 轴上。

Correlation measures the strength and direction of a linear relationship. The Pearson product‑moment correlation coefficient r is given by:

相关衡量线性关系的强度和方向。皮尔逊积矩相关系数 r 的计算公式为:

r = [nΣxy – (Σx)(Σy)] / √[nΣx² – (Σx)²] · √[nΣy² – (Σy)²]

The value of r lies between –1 and 1. r = 1 indicates perfect positive correlation, r = –1 perfect negative correlation, and r = 0 no linear correlation. The sign shows the direction, and the magnitude shows the strength.

r 的值在 –1 到 1 之间。r = 1 表示完全正相关,r = –1 表示完全负相关,r = 0 表示没有线性相关。符号表示方向,绝对值大小表示强度。

Spearman’s rank correlation coefficient ρ is used when data are ordinal or when the relationship is monotonic but not necessarily linear. It is computed from the differences in ranks d:

斯皮尔曼等级相关系数 ρ 用于有序数据或单调但非线性的关系。它通过对等级差 d 计算得到:

ρ = 1 – (6 Σd²) / [n(n² – 1)]

Always remember that correlation does not imply causation.

始终记住:相关关系不等于因果关系。


8. Linear Regression | 线性回归

The equation of the least‑squares regression line (line of best fit) is y = a + b x, where:

最小二乘回归线(最佳拟合线)的方程为 y = a + b x,其中:

b = [nΣxy – (Σx)(Σy)] / [nΣx² – (Σx)²]

a = ȳ – b x̄

The slope b represents the change in y when x increases by 1 unit. The intercept a is the predicted value of y when x = 0, but it may not always have a meaningful interpretation.

斜率 b 表示当 x 增加一个单位时 y 的变化量。截距 a 是 x = 0 时 y 的预测值,但它并不总具有实际解释意义。

Residuals are the differences between observed and predicted values: residual = yₒ − yₚ. A residual plot helps check the linearity and constant variance assumptions.

残差是观测值与预测值之差:残差 = yₒ − yₚ。残差图有助于检验线性和等方差假设。

The coefficient of determination R² = r² indicates the proportion of the variation in y that is explained by x. An r² value close to 1 indicates a strong model.

决定系数 R² = r² 表示 y 的变异中有多少比例可以被 x 解释。r² 接近 1 表明模型拟合良好。


9. Probability Distributions in Context | 实际情境中的概率分布

When modelling real‑world scenarios, check whether trials are independent and whether the probability of success is constant. If so, a binomial model may be appropriate.

在建立实际情境的模型时,应检查试验是否独立、成功概率是否恒定。若满足,便可使用二项模型。

For continuous measurements that cluster around a mean and follow a symmetric bell shape, a normal distribution is often a good model.

对于围绕均值聚集且呈对称钟形的连续测量值,正态分布通常是一个很好的模型。

The normal approximation to the binomial requires a continuity correction when calculating probabilities for discrete outcomes. For example, to approximate P(X ≤ k), use P(Y < k + 0.5) where Y ~ N(np, npq).

二项分布的正态近似在计算离散结果的概率时需要进行连续性校正。例如,近似 P(X ≤ k) 时,使用 P(Y < k + 0.5),其中 Y ~ N(np, npq)。

Read questions carefully to decide which distribution to use and always state assumptions.

仔细读题,决定使用哪种分布,并始终说明假设条件。


10. Sampling and Data Collection | 抽样与数据收集

A simple random sample gives every member of the population an equal chance of being selected. This eliminates bias.

简单随机抽样使总体中每个个体被选中的机会均等,从而消除偏差。

Stratified sampling divides the population into groups (strata) and samples proportionally from each, ensuring representation.

分层抽样将总体分成若干层,按比例从每层抽样,保证代表性。

Systematic sampling selects every k‑th individual from a list. It is quick but can introduce bias if there is a hidden pattern.

系统抽样从名单中每隔 k 个个体抽取一个,快速简便,但若存在隐藏模式可能引入偏差。

In cluster sampling, the population is divided into clusters, and some clusters are randomly selected entirely. This is practical for geographically spread populations.

整群抽样将总体分成群组,随机选取若干群组进行全面调查,适用于地理分散的总体。

Pilot surveys help test questionnaires and identify potential problems before the main data collection.

先导调查有助于在正式收集数据前测试问卷、发现潜在问题。


11. Presenting Data | 数据展示

Frequency tables, bar charts, and pie charts are used for categorical data. Histograms display grouped quantitative data, with area proportional to frequency.

频数表、条形图和饼图用于分类数据。直方图展示分组定量数据,面积与频数成正比。

For histograms with unequal class widths, frequency density = frequency / class width. Plot frequency density on the vertical axis.

对于组距不等的直方图,使用频率密度 = 频数 / 组距,并将其标在纵轴上。

Cumulative frequency diagrams plot running totals against upper class boundaries, producing an S‑shaped curve. The median and quartiles can be read from it.

累积频数图将累计频数对应上组界绘制,形成 S 形曲线,可从中读取中位数和四分位数。

Box‑and‑whisker plots show the minimum, Q₁, median, Q₃, and maximum, giving a clear visual summary of spread and skewness.

箱形图展示最小值、Q₁、中位数、Q₃ 和最大值,直观反映数据离散程度与偏态。


12. Quick Revision Checklist | 快速复习清单

Before the exam, make sure you can:

考试前请确认你能够:

  • Calculate mean, median, mode, range, quartiles, standard deviation (for raw and grouped data) | 计算均值、中位数、众数、极差、四分位数、标准差(适用于原始数据和分组数据)
  • Use probability rules (addition, multiplication, complement, conditional) correctly | 正确使用概率规则(加法、乘法、互补、条件概率)
  • Recognise and work with binomial and normal distributions, including standardisation and continuity correction | 识别并处理二项分布和正态分布,包括标准化和连续性校正
  • Find correlation coefficient r and Spearman’s rank ρ, and interpret them | 求出相关系数 r 和斯皮尔曼等级相关系数 ρ,并解释其含义
  • Determine the equation of the regression line, make predictions, and analyse residuals | 确定回归线方程,进行预测,分析残差
  • Select appropriate sampling methods and identify potential biases | 选择合适的抽样方法并识别潜在偏差
  • Read and construct histograms, cumulative frequency graphs, and box plots accurately | 准确阅读并绘制直方图、累积频数图和箱形图

Use this handbook alongside your class notes and past paper practice to build confidence and speed.

将本手册与课堂笔记、历年真题练习结合使用,以增强信心并提高解题速度。

Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading