📚 Mastering Pre-U OCR Statistics: In-depth Analysis of Past Papers | 掌握 Pre-U OCR 统计:历年真题深度解析
Past papers are the most powerful revision tool for Pre-U OCR Statistics (9794). By dissecting real exam questions, you develop the ability to recognise underlying structures, apply the correct distributional theory, and avoid subtle traps that separate A* candidates from the rest. This article explores key topics through the lens of genuine past-paper challenges, providing step-by-step reasoning and bilingual commentary.
历年真题是复习 Pre-U OCR 统计(9794)最强大的武器。通过剖析真实的考试题目,你可以培养识别底层结构、应用恰当分布理论的能力,并避开那些区分 A* 考生与常人的陷阱。本文透过一道道具代表性的真题,提供逐步解析和中英双语点评。
1. Discrete Random Variables & Probability Distributions | 离散随机变量与概率分布
A typical Pre-U question gives a small probability mass function with an unknown constant and asks for the full distribution, E(X) and Var(X). Often candidates misread the condition that two probabilities are equal or that a cumulative probability is provided. Always verify that the sum of probabilities equals one before proceeding.
典型的 Pre-U 题目会给出一个含有未知常数的小概率质量函数,并要求写出完整分布、E(X) 和 Var(X)。考生经常误读两个概率相等的条件,或者忽略给出的累积概率。务必在继续之前验证概率之和为 1。
Example: The random variable X takes values 0, 1, 2, 3. It is given that P(X=0) = 0.1, P(X=1) = 2k, P(X=2) = k, and P(X=3) = 0.3.
例题:随机变量 X 取值 0, 1, 2, 3。已知 P(X=0) = 0.1, P(X=1) = 2k, P(X=2) = k, P(X=3) = 0.3。
- Step 1: Solve 0.1 + 2k + k + 0.3 = 1 ⇒ 3k = 0.6 ⇒ k = 0.2.
- 步骤1:解方程 0.1 + 2k + k + 0.3 = 1 ⇒ 3k = 0.6 ⇒ k = 0.2。
- Step 2: Hence P(X=1)=0.4, P(X=2)=0.2. Use these to find E(X) = Σ x·p(x) = 0×0.1 + 1×0.4 + 2×0.2 + 3×0.3 = 1.7.
- 步骤2:因此 P(X=1)=0.4, P(X=2)=0.2。计算 E(X) = Σ x·p(x) = 0×0.1 + 1×0.4 + 2×0.2 + 3×0.3 = 1.7。
- Step 3: Then E(X²) = 0²×0.1 + 1²×0.4 + 2²×0.2 + 3²×0.3 = 0 + 0.4 + 0.8 + 2.7 = 3.9. Hence Var(X) = E(X²) – [E(X)]² = 3.9 – 1.7² = 3.9 – 2.89 = 1.01.
- 步骤3:然后 E(X²) = 0²×0.1 + 1²×0.4 + 2²×0.2 + 3²×0.3 = 0 + 0.4 + 0.8 + 2.7 = 3.9。因此 Var(X) = E(X²) – [E(X)]² = 3.9 – 1.7² = 3.9 – 2.89 = 1.01。
This foundation is crucial because it reappears in PGFs and linear combinations, where errors in basic expectation algebra cascade.
这个基础至关重要,因为它在 PGF 和线性组合中反复出现,基础期望代数中的错误会引发连锁反应。
2. PGFs: Finding Distributions and Moments | 概率生成函数:求分布与矩
Probability generating functions are a favourite in Pre-U. A common question gives G(t) and asks to recover the probability distribution, then find E(X) and Var(X) using G'(1) and G”(1). Remember that the coefficient of tˣ in the power series expansion gives P(X=x).
概率生成函数是 Pre-U 的高频考点。常见题型给出 G(t),要求还原概率分布,然后利用 G'(1) 和 G”(1) 求 E(X) 与 Var(X)。记住,幂级数展开式中 tˣ 的系数就是 P(X=x)。
Example from a past paper: The PGF of X is G(t) = (0.2 + 0.8t)⁴.
真题示例:X 的 PGF 为 G(t) = (0.2 + 0.8t)⁴。
By inspection, this is a binomial expansion. Expanding gives P(X=x) = ⁴Cₓ (0.8)ˣ (0.2)⁴⁻ˣ for x = 0,1,2,3,4. Differentiation is quicker: G'(t) = 4×0.8 (0.2+0.8t)³ = 3.2 (0.2+0.8t)³, so E(X) = G'(1) = 3.2×1³ = 3.2, as expected since np = 4×0.8 = 3.2.
通过观察可知这是二项展开式。展开后得到 P(X=x) = ⁴Cₓ (0.8)ˣ (0.2)⁴⁻ˣ,x=0至4。用求导更快:G'(t) = 4×0.8 (0.2+0.8t)³ = 3.2 (0.2+0.8t)³,所以 E(X) = G'(1) = 3.2×1³ = 3.2,与二项分布的 np = 4×0.8 = 3.2 一致。
Next, G”(t) = 3.2×3×0.8 (0.2+0.8t)² = 7.68 (0.2+0.8t)², so G”(1) = 7.68. Then Var(X) = G”(1) + G'(1) – [G'(1)]² = 7.68 + 3.2 – 3.2² = 10.88 – 10.24 = 0.64, matching npq = 4×0.8×0.2 = 0.64.
接着,G”(t) = 3.2×3×0.8 (0.2+0.8t)² = 7.68 (0.2+0.8t)²,故 G”(1) = 7.68。于是 Var(X) = G”(1) + G'(1) – [G'(1)]² = 7.68 + 3.2 – 3.2² = 10.88 – 10.24 = 0.64,与 npq = 4×0.8×0.2 = 0.64 相符。
3. Continuous Random Variables: PDFs and CDFs | 连续随机变量:PDF 与 CDF
A popular Pre-U question supplies a piecewise pdf f(x) and requires the constant, the cumulative distribution function F(x), and the median. Always check that the integral of the pdf over its domain equals 1. For the median m, solve F(m) = 0.5.
Pre-U 常考分段概率密度函数 f(x),要求求出常数、累积分布函数 F(x) 和中位数。务必验证 pdf 在定义域上的积分等于 1。对于中位数 m,解方程 F(m) = 0.5。
Example: f(x) = { kx² 0 ≤ x ≤ 2,
{ k(4 – x) 2 < x ≤ 4,
{ 0 otherwise. Find k, F(x), and the median.
例题:f(x) = { kx² 0 ≤ x ≤ 2,
{ k(4 – x) 2 < x ≤ 4,
{ 0 其他。求 k, F(x) 和 中位数。
Integrate: ∫₀² kx² dx + ∫₂⁴ k(4 – x) dx = k [x³/3]₀² + k [4x – x²/2]₂⁴ = k(8/3) + k[(16 – 8) – (8 – 2)] = k(8/3 + 8 – 6) = k(8/3 + 2) = (14k/3) = 1 ⇒ k = 3/14.
积分:∫₀² kx² dx + ∫₂⁴ k(4 – x) dx = k [x³/3]₀² + k [4x – x²/2]₂⁴ = k(8/3) + k[(16 – 8) – (8 – 2)] = k(8/3 + 8 – 6) = k(8/3 + 2) = (14k/3) = 1 ⇒ k = 3/14。
For 0 ≤ x ≤ 2: F(x) = ∫₀ˣ (3/14) t² dt = (3/14)(x³/3) = x³/14.
For 2 < x ≤ 4: F(x) = F(2) + ∫₂ˣ (3/14)(4 – t) dt = 8/14 + (3/14)[4t – t²/2]₂ˣ = 8/14 + (3/14)[(4x – x²/2) – (8 – 2)] = 4/7 + (3/14)(4x – x²/2 – 6) = 4/7 + (12x/14 – 3x²/28 – 18/14) = (8/14 – 18/14) + (12x/14) – (3x²/28) = –5/7 + (6x/7) – (3x²/28).
0 ≤ x ≤ 2 时: F(x) = x³/14.
2 < x ≤ 4 时: F(x) = 4/7 + (3/14)(4x – x²/2 – 6),整理得 F(x) = –5/7 + (6x/7) – (3x²/28)。
The median lies where F(m)=0.5. Test F(2)=8/14≈0.571, so median < 2. Solve m³/14 = 0.5 → m³=7 → m = ∛7 ≈ 1.913.
中位数满足 F(m)=0.5。检验 F(2)=8/14≈0.571,因此中位数 < 2。解 m³/14 = 0.5 → m³=7 → m = ∛7 ≈ 1.913。
4. Expectation Algebra & Linear Combinations | 期望代数与线性组合
Pre-U paper 2 frequently asks for the expectation and variance of a linear combination of independent random variables, such as Y = aX₁ + bX₂ + c. Use E(Y) = aE(X₁) + bE(X₂) + c, and Var(Y) = a²Var(X₁) + b²Var(X₂) if independent. A common trap is forgetting to square the coefficients for variance.
Pre-U 卷二经常要求计算独立随机变量线性组合的期望与方差,如 Y = aX₁ + bX₂ + c。使用 E(Y) = aE(X₁) + bE(X₂) + c,以及独立时 Var(Y) = a²Var(X₁) + b²Var(X₂)。常见的陷阱是对方差忘记系数平方。
Example: X₁ ~ N(μ₁, σ₁²) and X₂ ~ N(μ₂, σ₂²) independently. Let T = 3X₁ – 2X₂. Find E(T) and Var(T).
例题:X₁ ~ N(μ₁, σ₁²) 且 X₂ ~ N(μ₂, σ₂²) 相互独立。设 T = 3X₁ – 2X₂。求 E(T) 和 Var(T)。
E(T) = 3μ₁ – 2μ₂. Var(T) = 3²σ₁² + (–2)²σ₂² = 9σ₁² + 4σ₂². This result is used later in confidence intervals for the difference of means or in hypothesis tests about contrasts.
E(T) = 3μ₁ – 2μ₂。Var(T) = 3²σ₁² + (–2)²σ₂² = 9σ₁² + 4σ₂²。这一结果后续会用于均值差值的置信区间或关于对照的假设检验。
5. Sampling & The Central Limit Theorem | 抽样与中心极限定理
Questions on the Central Limit Theorem (CLT) often give a non-normal population with known mean μ and variance σ². For a sample of size n, the sample mean X̄ is approximately N(μ, σ²/n) if n is large. The Pre-U exam expects you to state the CLT clearly and justify the approximation before calculating probabilities.
中心极限定理(CLT)的题目常常给定一个非正态总体,已知均值 μ 和方差 σ²。对于样本量 n 足够大的样本,样本均值 X̄ 近似服从 N(μ, σ²/n)。Pre-U 考试要求你清晰陈述 CLT 并进行近似论证,然后计算概率。
Example: A shop’s daily sales X are skewed with mean £820 and standard deviation £150. Find the probability that the mean sales over 50 randomly chosen days exceed £850.
例题:某店日销售额 X 呈偏态,均值 £820,标准差 £150。求随机50天的日均销售额超过 £850 的概率。
By the CLT, X̄ ≈ N(820, 150²/50 = 450) for n=50. Thus Z = (850 – 820)/√450 ≈ 30 / 21.21 = 1.414. P(X̄ > 850) = P(Z > 1.414) ≈ 0.0786. Always check that n is large enough; the exam may ask why n=50 justifies the approximation.
由 CLT,n=50 时 X̄ ≈ N(820, 150²/50 = 450)。因此 Z = (850 – 820)/√450 ≈ 30 / 21.21 = 1.414。P(X̄ > 850) = P(Z > 1.414) ≈ 0.0786。务必检查 n 是否足够大;试题可能要求解释 n=50 为何可以合理近似。
6. Confidence Intervals for the Mean | 均值的置信区间
Constructing a 95% confidence interval when σ is known uses the formula X̄ ± z₀.₀₂₅ × σ/√n. Past papers love to then ask for the interpretation: “We are 95% confident that the true population mean lies within this interval.” The interval does not make a probability statement about μ once it is computed.
当 σ 已知时,构建 95% 置信区间的公式为 X̄ ± z₀.₀₂₅ × σ/√n。真题常要求解释:“我们有 95% 的信心认为真实总体均值位于该区间内。”区间一旦计算出来,就不再是关于 μ 的概率陈述。
Example: A random sample of 40 items gives X̄ = 52.0, population σ = 5.0. Find the 90% confidence interval.
例题:一个 40 个产品的随机样本给出 X̄ = 52.0,总体 σ = 5.0。求 90% 置信区间。
For 90% confidence, z₀.₀₅ = 1.645. The interval is 52.0 ± 1.645 × (5.0/√40) = 52.0 ± 1.645 × 0.7906 = (50.70, 53.30). If the question asked for the minimum sample size to achieve a width of 2, solve 2 × 1.645 × 5/√n = 2 ⇒ √n = 8.225 ⇒ n ≈ 67.7, so need 68.
90% 置信水平下,z₀.₀₅ = 1.645。区间为 52.0 ± 1.645 × (5.0/√40) = 52.0 ± 1.645 × 0.7906 = (50.70, 53.30)。若题目要求达到宽度为 2 的最小样本量,解 2 × 1.645 × 5/√n = 2 ⇒ √n = 8.225 ⇒ n ≈ 67.7,故需 68 个样本。
7. Hypothesis Testing: Type I/II Errors & Power | 假设检验:两类错误与功效
A well-crafted Pre-U question asks for the probability of a Type I error (α) and the power of a test (1 – β). You need to identify the critical region from the null distribution, then compute the probability under the alternative distribution.
精心设计的 Pre-U 题目会要求计算 I 型错误的概率(α)和检验功效(1 – β)。你需要从原假设分布确定拒绝域,然后在备择假设分布下计算概率。
Example: H₀: μ = 100, σ = 15, n = 25. Test at the 5% significance level against H₁: μ > 100. Determine the power when μ = 105.
例题:H₀: μ = 100, σ = 15, n = 25。在 5% 显著性水平下检验 H₁: μ > 100。求 μ = 105 时检验的功效。
Under H₀, X̄ ~ N(100, 15²/25 = 9). Critical value c = 100 + 1.645×3 = 104.935. Reject if X̄ > 104.935. Type I error = P(reject | H₀ true) = 0.05. Under H₁ with μ = 105, X̄ ~ N(105, 9). Power = P(X̄ > 104.935 | μ=105) = P( Z > (104.935 – 105)/3 ) = P(Z > –0.0217) ≈ 0.5087. So the test has about 51% power to detect a shift of 5 units.
在原假设下 X̄ ~ N(100, 9),临界值 c = 100 + 1.645×3 = 104.935。当 X̄ > 104.935 时拒绝。I 型错误 = 0.05。在备择假设 μ = 105 时,X̄ ~ N(105, 9)。功效 = P(X̄ > 104.935) = P( Z > –0.0217 ) ≈ 0.5087。因此该检验检测 5 个单位偏移的能力约为 51%。
8. Chi-Squared Goodness of Fit | 卡方拟合优度检验
For goodness-of-fit tests, the test statistic is Σ (Oᵢ – Eᵢ)² / Eᵢ, which under the null follows a χ² distribution with (k – 1 – p) degrees of freedom, where p parameters are estimated from the data. A common mistake is using the wrong degrees of freedom or forgetting to combine categories with small expected frequencies.
拟合优度检验的检验统计量为 Σ (Oᵢ – Eᵢ)² / Eᵢ,在原假设下服从自由度为 (k – 1 – p) 的 χ² 分布,其中 p 是从数据中估计的参数个数。常见错误是使用错误的自由度,或忘记合并期望频数较小的类别。
Example: Observed counts for 5 categories: 18, 25, 22, 20, 15. Test the hypothesis that they follow a uniform distribution.
例题:5 个类别的观测频数:18, 25, 22, 20, 15。检验它们是否服从均匀分布。
Total n = 100. Under uniform, each Eᵢ = 20. χ² = (18–20)²/20 + (25–20)²/20 + (22–20)²/20 + (20–20)²/20 + (15–20)²/20 = 4/20 + 25/20 + 4/20 + 0 + 25/20 = 58/20 = 2.9. Degrees of freedom = 5 – 1 = 4. Critical value at 5% is 9.488, so we do not reject H₀. The fit to a uniform distribution is adequate.
总样本量 n = 100。均匀分布下每个 Eᵢ = 20。χ² = (18–20)²/20 + (25–20)²/20 + (22–20)²/20 + (20–20)²/20 + (15–20)²/20 = 2.9。自由度 df = 5 – 1 = 4。5% 临界值为 9.488,不拒绝 H₀,数据与均匀分布拟合尚可。
9. Correlation & Regression Inference | 相关与回归推断
Pre-U often includes a full analysis of bivariate data: compute Pearson’s r, find the least-squares regression line, and perform a t-test for the correlation coefficient or the slope. Remind yourself of the formula Sxy / √(Sxx Syy) for r and b = Sxy / Sxx for the slope.
Pre-U 经常要求对双变量数据进行完整分析:计算 Pearson 相关系数 r,求最小二乘回归直线,并对相关系数或斜率进行 t 检验。请熟记公式:r = Sxy / √(Sxx Syy),斜率 b = Sxy / Sxx。
Example: For n = 12, Σx = 60, Σy = 48, Sxx = 100, Syy = 80, Sxy = 70. Find the regression line y = a + bx and test H₀: β = 0 at 5% level.
例题:n = 12, Σx = 60, Σy = 48, Sxx = 100, Syy = 80, Sxy = 70。求回归直线 y = a + bx,并在 5% 水平检验 H₀: β = 0。
Mean x̄ = 5, ȳ = 4. Slope b = Sxy/Sxx = 70/100 = 0.7. Intercept a = ȳ – b x̄ = 4 – 0.7×5 = 0.5. So line: y = 0.5 + 0.7x.
Published by TutorHao | Pre-U 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导