Tag: 统计

  • Year 12 CIE Statistics: Core Concepts Review | Year 12 CIE 统计:核心知识点梳理

    📚 Year 12 CIE Statistics: Core Concepts Review | Year 12 CIE 统计:核心知识点梳理

    The Cambridge International AS Level Mathematics (9709) Probability & Statistics 1 paper demands a solid grasp of fundamental statistical methods and probability theory. This revision guide breaks down the core concepts you need to command as a Year 12 student, helping you build confidence for the exam.

    剑桥国际AS数学(9709)概率与统计1试卷要求考生扎实掌握基本的统计方法和概率论。这篇复习指南梳理了Year 12学生必须掌握的核心概念,帮助你在考试中建立信心。

    1. Types of Data and Data Representation | 数据类型与数据表示

    Data can be qualitative (categorical), such as favourite colour or type of vehicle, or quantitative (numerical). Quantitative data is further classified as discrete – involving counts like ‘number of students’ – or continuous – involving measurements like time or height. Recognising the data type helps you select an appropriate diagram.

    数据可以是定性(分类)的,例如最喜欢的颜色或车辆类型,也可以是定量(数值)的。定量数据进一步分为离散型——涉及计数如“学生人数”——或连续型——涉及测量值如时间或身高。识别数据类型有助于选择恰当的图表。

    Common representations include stem-and-leaf diagrams, which preserve the original data while displaying order; box-and-whisker plots, which highlight the median, quartiles and potential outliers; and histograms, where the area of each bar is proportional to the frequency. For cumulative data, cumulative frequency curves (ogives) allow you to estimate medians, quartiles and percentiles by reading off the graph.

    常用的表示方法有茎叶图,它在展示顺序的同时保留了原始数据;盒须图突出中位数、四分位数和潜在异常值;直方图中每个条形的面积与频率成比例。对于累积数据,累积频率曲线(拱形图)可以通过读取图形来估计中位数、四分位数和百分位数。


    2. Measures of Central Tendency: Mean, Median, Mode | 集中趋势度量

    The mean (arithmetic average) for ungrouped data is x̄ = (Σx)/n. For grouped data, use midpoints: x̄ ≈ (Σfx)/Σf. The mean is sensitive to extreme values. The median is the middle value when data are ordered; its position is (n+1)/2. In a cumulative frequency graph, the median corresponds to the 50% mark.

    未分组数据的均值(算术平均值)为 x̄ = (Σx)/n。对于分组数据,使用组中点:x̄ ≈ (Σfx)/Σf。均值易受极端值影响。中位数是数据排序后的中间值;位置为 (n+1)/2。在累积频率曲线中,中位数对应50%标记。

    The mode is the most frequent value or the modal class for grouped data. When data are symmetric, mean ≈ median ≈ mode. In skewed distributions the mean is pulled towards the tail. Linear coding y = ax + b changes the mean to ȳ = ax̄ + b, which is very useful when handling large data sets.

    众数是出现频率最高的值,或分组数据的众数组。当数据对称时,均值 ≈ 中位数 ≈ 众数。在偏态分布中,均值会被拖向尾部。线性编码 y = ax + b 将均值变为 ȳ = ax̄ + b,这在处理大数据集时非常有用。


    3. Measures of Dispersion: Range, IQR, Variance, Standard Deviation | 离散程度度量

    Spread can be measured by the range (max – min), the interquartile range (IQR = Q₃ – Q₁) and the variance/standard deviation. The IQR ignores extreme values and is found from a cumulative frequency graph or ordered list. A box plot displays Q₁, Q₂, Q₃ and any outliers.

    离散程度可以用极差(最大值 – 最小值)、四分位距(IQR = Q₃ – Q₁)以及方差/标准差来衡量。四分位距忽略极端值,可从累积频率图或排序列表中得到。箱线图展示了 Q₁、Q₂、Q₃ 以及任何异常值。

    For a set of n values, the variance σ² = Σ(x – x̄)² / n, which can also be written as σ² = (Σx²/n) – x̄². The standard deviation σ = √σ². For grouped data, use Σfx² and Σfx. These measures are unaffected by adding a constant, but multiplying by a factor a scales the variance by a².

    对于一组 n 个数值,方差 σ² = Σ(x – x̄)² / n,也可以写成 σ² = (Σx²/n) – x̄²。标准差 σ = √σ²。对于分组数据,使用 Σfx² 和 Σfx。这些度量在加上常数时不变,但乘以因子 a 时方差缩放为原来的 a² 倍。


    4. Probability Basics and Venn Diagrams | 概率基础与维恩图

    The probability of an event A is P(A) = number of favourable outcomes / total number of outcomes, with 0 ≤ P(A) ≤ 1. The complement rule gives P(A’) = 1 – P(A). For two events A and B, the general addition rule is P(A ∪ B) = P(A) + P(B) – P(A ∩ B). If A and B are mutually exclusive, P(A ∩ B) = 0 and the formula simplifies.

    事件 A 的概率为 P(A) = 有利结果数 / 总结果数,且满足 0 ≤ P(A) ≤ 1。补规则给出 P(A’) = 1 – P(A)。对于两个事件 A 和 B,一般加法法则是 P(A ∪ B) = P(A) + P(B) – P(A ∩ B)。如果 A 与 B 互斥,则 P(A ∩ B) = 0,公式简化。

    Venn diagrams help visualise intersections, unions and complements. They are especially useful for problems involving overlapping categories. When probabilities are calculated from a sample space, always check that the total probability equals 1.

    维恩图有助于直观展示交集、并集和补集。在涉及重叠类别的问题中尤为有用。当从样本空间计算概率时,务必检查总概率是否等于 1。


    5. Conditional Probability and Tree Diagrams | 条件概率与树状图

    Conditional probability is defined as P(A|B) = P(A ∩ B) / P(B), provided P(B) > 0. It represents the probability of A occurring given that B has already happened. Tree diagrams are a powerful tool: multiply along branches for ‘and’ probabilities and add across different branches for ‘or’ probabilities.

    条件概率定义为 P(A|B) = P(A ∩ B) / P(B),前提是 P(B) > 0。它表示在 B 已发生的情况下 A 发生的概率。树状图是一种强大的工具:沿着分支相乘得到“且”的概率,把不同分支相加得到“或”的概率。

    Two events are independent if P(A|B) = P(A) or equivalently P(A ∩ B) = P(A)P(B). This is a key concept to check when drawing successive trials, such as picking coloured balls with or without replacement. With replacement events are independent; without replacement the probabilities change.

    如果 P(A|B) = P(A),或等价地 P(A ∩ B) = P(A)P(B),则两个事件独立。这是检查连续试验(如有放回或无放回地抽取彩色球)的关键概念。有放回时事件独立;无放回时概率会发生变化。


    6. Permutations and Combinations | 排列与组合

    When arranging objects, the number of ways to arrange n distinct items is n! . The number of permutations of r items from n is nPr = n!/(n – r)!, used when order matters. If some items are repeated, divide by factorials of the repetitions: n!/(p!q!…).

    在排列物体时,排列 n 个不同项目的方式数为 n! 。从 n 项中选取 r 项排成一列(顺序有关)的排列数为 nPr = n!/(n – r)!。如果有重复项目,需除以重复次数的阶乘:n!/(p!q!…)。

    Combinations count selections where order does not matter: nCr = n!/(r!(n – r)!). This is the binomial coefficient, essential for the binomial distribution. Combinations are used when choosing a committee or a set of items, while permutations apply to races or passwords.

    组合计算顺序无关的选取方式:nCr = n!/(r!(n – r)!)。这就是二项式系数,对二项分布至关重要。选择委员会或一组物品时用组合,而排列适用于比赛名次或密码。


    7. Discrete Random Variables and Probability Distributions | 离散随机变量与概率分布

    A discrete random variable X takes countable values. A probability distribution lists x and P(X = x) such that Σ P(X = x) = 1. The expected value E(X) is the mean: E(X) = Σ [x · P(X = x)]. The variance Var(X) = Σ (x – μ)² P(X = x) = Σ [x² P(X = x)] – [E(X)]².

    离散随机变量 X 取可数值。概率分布列出了 x 和 P(X = x),且满足 Σ P(X = x) = 1。期望值 E(X) 就是均值:E(X) = Σ [x · P(X = x)]。方差 Var(X) = Σ (x – μ)² P(X = x) = Σ [x² P(X = x)] – [E(X)]²。

    For linear functions, E(aX + b) = aE(X) + b and Var(aX + b) = a²Var(X). Adding or subtracting independent random variables gives E(X ± Y) = E(X) ± E(Y) and Var(X ± Y) = Var(X) + Var(Y). These results allow you to analyse combined experiments.

    对于线性函数,E(aX + b) = aE(X) + b,Var(aX + b) = a²Var(X)。对独立的随机变量进行加减,有 E(X ± Y) = E(X) ± E(Y) 及 Var(X ± Y) = Var(X) + Var(Y)。这些结果可用于分析组合实验。


    8. The Binomial Distribution | 二项分布

    A binomial situation has a fixed number of trials n, each trial independent, only two outcomes (success/failure), and constant probability p of success. We write X ~ B(n, p). The probability of exactly r successes is P(X = r) = ⁿCᵣ pʳ (1 – p)ⁿ⁻ʳ.

    二项分布情形具有固定的试验次数 n,每次试验独立,只有两种结果(成功/失败),且成功的概率 p 保持不变。我们记作 X ~ B(n, p)。恰好 r 次成功的概率为 P(X = r) = ⁿCᵣ pʳ (1 – p)ⁿ⁻ʳ。

    The mean is E(X) = np, and the variance is Var(X) = np(1 – p). CIE provides binomial cumulative probability tables; you can also use the formula for individual probabilities and sum them. Be careful to distinguish between ‘at most’, ‘at least’ and ‘exactly’ phrasing.

    均值为 E(X) = np,方差为 Var(X) = np(1 – p)。CIE 提供二项累积概率表;你也可以使用公式计算单个概率并求和。注意区分“至多”、“至少”和“恰好”等措辞。


    9. The Normal Distribution | 正态分布

    The normal distribution models continuous data with a symmetric bell-shaped curve. If X ~ N(μ, σ²), then the standardised variable Z = (X – μ)/σ follows the standard normal distribution N(0, 1). Tables give P(Z < z) for positive z; use symmetry for negative values.

    正态分布用对称的钟形曲线对连续数据建模。如果 X ~ N(μ, σ²),则标准化变量 Z = (X – μ)/σ 服从标准正态分布 N(0, 1)。正态分布表给出了正 z 值下的 P(Z < z);对负值使用对称性即可。

    To find probabilities, draw a sketch, standardise, then use the table. For inverse problems (finding a value given a probability), use the percentage points table. When n is large and both np and nq are greater than 5, the binomial B(n, p) can be approximated by a normal distribution N(np, npq) with a continuity correction.

    求概率时,画出图形,标准化,然后查表。对于逆向问题(给定概率求取值),使用百分位数表。当 n 很大且 np 与 nq 都大于 5 时,可用正态分布 N(np, npq) 来近似二项分布,并进行连续性校正。


    10. Data Coding and Transformations | 数据编码与变换

    Coding data simplifies calculations. A common transformation is y = (x – a)/b or y = ax + b. The mean transforms as ȳ = ax̄ + b. The variance is only affected by the multiplicative factor: Var(y) = a² Var(x), and the standard deviation scales by |a|.

    数据编码可简化计算。常见的变换是 y = (x – a)/b 或 y = ax + b。均值变换为 ȳ = ax̄ + b。方差只受乘法因子影响:Var(y) = a² Var(x),标准差按 |a| 缩放。

    Using coding you can convert raw data into manageable integers, calculate statistics, and then decode the results. This technique is extremely helpful in exam questions involving large numbers. Always remember to reverse the coding at the end if the question asks for the original mean or variance.

    通过编码,你可以将原始数据转换为易处理的整数,计算统计量,然后将结果解码。在涉及大数的考题中这一技巧极其有用。如果题目要求原始均值或方差,务必在最后将编码还原。

    Published by TutorHao | Statistics Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Applying AQA Statistics to Compare UK University Entry Requirements | 运用AQA统计学比较英国大学入学要求

    📚 Applying AQA Statistics to Compare UK University Entry Requirements | 运用AQA统计学比较英国大学入学要求

    As Year 13 students prepare for university applications, understanding entry requirements across different institutions is crucial. Statistics provides powerful tools to analyse and compare these requirements, helping you make data-driven decisions. This article explores how AQA A-Level statistical techniques can be applied to compare UK university entry requirements effectively.

    作为13年级学生准备大学申请,理解不同院校的入学要求至关重要。统计学提供了强大的工具来分析和比较这些要求,帮助您做出基于数据的决策。本文探讨如何有效运用AQA A-Level统计技术来比较英国大学的入学要求。

    1. Data Collection and UCAS Tariff Points | 数据收集与UCAS分数

    The first step is collecting relevant data, such as typical A-level grade offers from universities. UCAS tariff points convert grades into numerical values, enabling quantitative analysis. For instance, an A* at A-level is worth 56 points, A is 48, B is 40, C is 32, D is 24, and E is 16. You can gather data for a specific course, like Economics, across multiple universities.

    第一步是收集相关数据,如大学典型的A-level成绩录取要求。UCAS分数将成绩转换为数值,便于定量分析。例如,A-level的A*为56分,A为48分,B为40分,C为32分,D为24分,E为16分。您可以收集多所大学同一专业(如经济学)的数据。

    A-level Grade UCAS Tariff Points
    A* 56
    A 48
    B 40
    C 32
    D 24
    E 16

    Ensure your sample includes a range of universities, from Russell Group to modern universities, to reflect the full spectrum of entry standards. A typical offer of AAA corresponds to 144 tariff points, while A*AA gives 152 points.

    确保样本涵盖从罗素集团到现代大学的一系列大学,以反映入学标准的全貌。典型的AAA录取对应144分,而A*AA则为152分。


    2. Descriptive Statistics for Entry Grades | 入学成绩的描述性统计

    Once data is collected, calculate measures of central tendency: mean, median, and mode of UCAS tariff points required. For example, the mean tariff for Economics might be 144 points, corresponding to AAA. The median might also be 144, but if the distribution is skewed, the median provides a more robust typical value.

    收集数据后,计算集中趋势的度量:所需UCAS分数的均值、中位数和众数。例如,经济学专业的平均分数可能为144分,对应A-Level成绩AAA。如果分布偏斜,中位数能提供更稳健的典型值。

    Also compute the range, interquartile range (IQR), and standard deviation to understand the variability in requirements. A small IQR indicates that most universities set similar entry standards, while a large IQR shows diverse expectations across the sector.

    同时计算极差、四分位距(IQR)和标准差,以了解要求的变异程度。较小的IQR表明大多数大学设定相似的入学标准,而较大的IQR则显示出整个行业内的期望差异很大。


    3. Visual Comparison: Box Plots and Cumulative Frequency | 可视化比较:箱线图与累积频率

    Box plots are ideal for comparing distributions of entry requirements across different university groups (e.g., Russell Group vs non-Russell Group). They display the minimum, lower quartile, median, upper quartile, and maximum. You can quickly identify outliers, such as a university demanding exceptionally high tariff points or an institution with a contextual offer.

    箱线图非常适合比较不同大学群体(如罗素集团与非罗素集团)入学要求的分布。它们显示出最小值、下四分位数、中位数、上四分位数和最大值。您可以快速识别异常值,例如一所要求异常高分的大学,或提供背景性录取的院校。

    Cumulative frequency curves can also illustrate the proportion of universities requiring up to a certain tariff. This helps you gauge where you stand relative to the entry landscape. For instance, the 70th percentile might be 160 points (A*AA), meaning 70% of courses require 160 points or fewer.

    累积频率曲线也能说明有多少比例的大学要求达到某个分数。这有助于您评估自己在入学环境中的定位。例如,第70百分位数可能是160分(A*AA),意味着70%的课程要求160分或更低。


    4. Measures of Spread and Consistency | 离散程度与一致性度量

    Calculating standard deviation and variance allows you to quantify consistency. For a highly competitive subject like Medicine, entry requirements might show low variance across top universities, all demanding A*AA or above. In contrast, for Business Studies, variance may be higher due to a wider range of course providers.

    计算标准差和方差可以让您量化一致性。对于像医学这样竞争激烈的专业,顶尖大学之间的入学要求可能显示出很小的方差,全部要求A*AA或以上。相比之下,商业研究专业由于提供者范围更广,方差可能较大。

    The coefficient of variation (CV) is useful when comparing the relative spread of tariff points across different courses. CV = (standard deviation / mean) × 100%. A lower CV suggests more uniform standards, allowing you to anticipate offers with greater certainty.

    变异系数(CV)在比较不同课程分数点的相对离散程度时很有用。CV = (标准差 / 均值) × 100%。CV较低表明标准更统一,让您能够更有把握地预测录取。


    5. Correlation Between Entry Requirements and University Rankings | 入学要求与大学排名之间的相关性

    You may hypothesize that higher-ranked universities have higher entry requirements. You can test this by plotting a scatter graph of tariff points against a ranking metric (e.g., Complete University Guide ranking score) and calculating Spearman’s rank correlation coefficient (ρ) or Pearson’s r.

    您可能假设排名较高的大学入学要求更高。您可以通过绘制分数点与排名指标(例如《完全大学指南》排名分数)的散点图,并计算斯皮尔曼等级相关系数(ρ)或皮尔逊相关系数r来检验。

    A strong positive correlation (e.g., ρ = 0.85) would confirm the link, but note that correlation does not imply causation. Other factors, such as course popularity, location, and widening participation policies, also influence entry standards.

    强正相关(例如 ρ = 0.85)将证实这种联系,但请注意相关性并不意味着因果关系。其他因素,如课程受欢迎程度、地理位置和扩大参与政策,也会影响入学标准。


    6. Hypothesis Testing: Russell Group vs Non-Russell Group | 假设检验:罗素集团与非罗素集团

    A two-sample t-test can determine whether there is a significant difference in mean tariff points between Russell Group and other universities. Let μ₁ be the mean tariff for Russell Group, μ₂ for non-Russell Group. H₀: μ₁ = μ₂, H₁: μ₁ > μ₂ (one-tailed) or μ₁ ≠ μ₂ (two-tailed).

    双样本t检验可以确定罗素集团大学与其他大学的平均分数是否存在显著差异。设 μ₁ 为罗素集团的均值,μ₂ 为非罗素集团的均值。H₀: μ₁ = μ₂,H₁: μ₁ > μ₂(单尾)或 μ₁ ≠ μ₂(双尾)。

    t = (x̄₁ − x̄₂) / (sp √(1/n₁ + 1/n₂))

    where sp is the pooled standard deviation. Compare the calculated t with the critical t-value at a 5% significance level. If the p-value is less than 0.05, reject H₀, concluding a significant difference exists.

    其中 sp 为合并标准差。将计算出的t值与5%显著性水平下的临界t值比较。如果p值小于0.05,则拒绝H₀,得出存在显著差异的结论。


    7. Regression Analysis for Predicted Offers | 预测录取的回归分析

    Using least squares regression, you can model the relationship between a university’s ranking score (independent variable) and its typical tariff offer (dependent variable). The equation of the regression line: y = a + bx, where the slope b = Sxy / Sxx.

    利用最小二乘回归,您可以对大学排名分数(自变量)与其典型分数录取(因变量)之间的关系进行建模。回归线方程:y = a + bx,其中斜率 b = Sxy / Sxx

    This model can predict the likely tariff offer for a university given its ranking. Residual analysis helps assess the model’s accuracy; a large residual suggests that the university’s requirements are either exceptionally high or low relative to its rank, indicating other influential factors.

    该模型可以根据大学的排名预测其可能的分数录取。残差分析有助于评估模型的准确性;较大的残差表明该大学的要求相对于其排名异常高或低,表明有其他影响因素。


    8. Probability of Receiving an Offer | 获得录取的概率

    Understanding conditional probabilities can help you estimate the likelihood of getting an offer. For example, P(Offer | Predictions ≥ AAA) might be estimated from historical data or published offer rates. You could use Bayes’ theorem if you have additional information about subgroups.

    理解条件概率可以帮助您估计获得录取的可能性。例如,P(录取 | 预计成绩 ≥ AAA) 可以根据历史数据或公布的录取率进行估计。如果有关于子群体的额外信息,您可以使用贝叶斯定理。

    Also, if a university has a 20% offer rate for a course, and your predicted grades match the standard offer, you might assume a base probability. However, remember that probability is not certainty—focus on meeting or exceeding the requirements and strengthening your application.

    此外,如果一所大学某专业的录取率为20%,且您的预计成绩符合标准要求,您可以假设一个基础概率。但请注意,概率并非确定性——专注于达到或超越要求并加强您的申请。


    9. Interpreting Conditional Probabilities | 解释条件概率

    Consider multiple conditions: what is the probability of being accepted given you have the required grades AND

    Published by TutorHao | Year 13 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Case Study in Statistical Inference: Real-World Data Analysis | 统计推断案例分析:真实世界数据分析实战

    📚 Case Study in Statistical Inference: Real-World Data Analysis | 统计推断案例分析:真实世界数据分析实战

    Statistical inference transforms raw data into actionable insights, enabling organisations to make evidence-based decisions under uncertainty. In this AQA Year 13 masterclass, we walk through a real-world scenario involving LED bulb manufacturing, applying a full suite of inferential techniques – from confidence intervals and hypothesis tests to regression, chi‑squared analysis and control charts.

    统计推断将原始数据转化为可操作的见解,使组织能够在不确定条件下做出基于证据的决策。在这堂 AQA 13 年级大师课中,我们将走进一个 LED 灯泡制造的真实场景,全面运用置信区间、假设检验、回归分析、卡方检验以及控制图等整套推断方法。


    1. Background and Dataset | 案例背景与数据集

    A leading LED manufacturer claims that its new bulb has an average lifetime of at least 1000 hours. To verify this claim, the quality‑engineering team selects a simple random sample of 30 bulbs and records their lifetimes (in hours) under controlled testing conditions.

    一家领先的 LED 制造商声称其新灯泡的平均寿命不低于 1000 小时。为了验证这一声明,质量工程团队在受控测试条件下抽取了一个 30 只灯泡的简单随机样本,并记录了它们的寿命(小时)。

    The measurement process is carefully standardised, and the data are treated as continuous. Summary statistics are computed as the foundation for all later inference. The sample is large enough to invoke the Central Limit Theorem, ensuring the sampling distribution of the mean is approximately normal.

    测量过程经过仔细标准化,数据被视为连续型。计算汇总统计量作为后续所有推断的基础。样本量足够大,可以调用中心极限定理,从而保证均值的抽样分布近似正态。


    2. Descriptive Statistics and Data Exploration | 描述性统计与数据探索

    From the 30 observations we obtain: sample mean x̄ = 985.2 h, sample standard deviation s = 48.7 h, median = 988 h, lower quartile = 950 h, upper quartile = 1018 h. The range is 138 h and the interquartile range is 68 h. A boxplot would show a roughly symmetric distribution with no extreme outliers.

    根据 30 个观测值,我们得到:样本均值 x̄ = 985.2 小时,样本标准差 s = 48.7 小时,中位数 = 988 小时,下四分位数 = 950 小时,上四分位数 = 1018 小时。极差为 138 小时,四分位距为 68 小时。箱线图将显示大致对称的分布,没有极端离群值。

    The standard deviation suggests that the lifetimes vary by about 49 hours around the mean. Because the sample mean is below 1000, we already suspect that the manufacturer’s claim may not be fulfilled, but we need formal statistical evidence.

    标准差表明,灯泡寿命在均值上下波动约 49 小时。由于样本均值低于 1000 小时,我们已怀疑制造商的声明可能不成立,但还需要正式的统计证据。

    Statistic Value (hours)
    Sample Mean (x̄) 985.2
    Sample Std Dev (s) 48.7
    Sample Size (n) 30
    Minimum 906
    Maximum 1044

    The summary table provides a quick snapshot of the data’s central tendency and spread, guiding the choice of subsequent inferential methods.

    汇总表快速呈现出数据的集中趋势和离散程度,指导着后续推断方法的选择。


    3. Confidence Interval for the Population Mean | 总体均值的置信区间

    Because the population standard deviation is unknown, we construct a 95% confidence interval for the true mean lifetime μ using the t‑distribution. The standard error is SE = s/√n = 48.7/√30 ≈ 8.891 h. With df = 29, the two‑tailed critical value is t* = 2.045. The margin of error is ME = 2.045 × 8.891 ≈ 18.18 h.

    由于总体标准差未知,我们使用 t 分布构建真实平均寿命 μ 的 95% 置信区间。标准误为 SE = s/√n = 48.7/√30 ≈ 8.891 小时。自由度为 29 时,双尾临界值为 t* = 2.045。误差幅度 ME = 2.045 × 8.891 ≈ 18.18 小时。

    95% CI for μ: x̄ ± t* × SE = 985.2 ± 18.18 → (967.02, 1003.38)

    The interval (967.0, 1003.4) suggests that the plausible values for the true mean include 1000 hours, so the claim cannot be discounted solely by interval estimation. However, the interval leans slightly to the left of 1000, motivating a formal hypothesis test.

    该区间 (967.0, 1003.4) 表明真实均值的合理取值包含 1000 小时,因此仅凭区间估计无法否定声称。然而区间略微偏左,这促使我们进行正式的假设检验。


    4. One‑Sample t‑Test: Testing the Manufacturer’s Claim | 单样本 t 检验:检验制造商声明

    We set up a one‑tailed test to see if the true mean is below 1000 hours: H₀: μ = 1000 versus H₁: μ < 1000. The significance level is α = 0.05.

    我们设定单尾检验,以判断真实均值是否低于 1000 小时:H₀: μ = 1000,H₁: μ < 1000。显著性水平 α = 0.05。

    t = (x̄ – μ₀) / (s/√n) = (985.2 – 1000) / 8.891 ≈ –1.665

    The test statistic is t ≈ –1.665 with 29 degrees of freedom. The critical value from t‑tables is –1.699 (one‑tailed). Since –1.665 > –1.699, we fail to reject H₀ at the 5% level. The p‑value is approximately 0.053, just above the threshold.

    检验统计量 t ≈ –1.665,自由度 29。查 t 分布表得单尾临界值为 –1.699。由于 –1.665 > –1.699,在 5% 水平上未能拒绝 H₀。p 值约为 0.053,刚好高于阈值。

    In context, there is insufficient evidence to conclude that the average lifetime is less than 1000 hours. The borderline result suggests that a larger sample size might provide a clearer picture, but the manufacturer’s claim stands for now.

    在实际情形中,没有足够证据表明平均寿命低于 1000 小时。这一临界结果说明,增大样本量可能让结论更清晰,但目前制造商声明得以维持。


    5. Two‑Sample t‑Test: Comparing Production Lines A and B | 双样本 t 检验:比较 A 和 B 生产线

    The factory operates two production lines. A rival hypothesis suggests line B may produce bulbs with a shorter average life. We draw independent random samples: from line A, n₁ = 25, x̄₁ = 1002 h, s₁ = 42 h; from line B, n₂ = 30, x̄₂ = 972 h, s₂ = 55 h. We assume equal population variances after an F‑test (p > 0.10).

    工厂运行两条生产线。有人提出,生产线 B 生产的灯泡平均寿命可能更短。我们抽取独立随机样本:A 线 n₁ = 25,x̄₁ = 1002 h,s₁ = 42 h;B 线 n₂ = 30,x̄₂ = 972 h,s₂ = 55 h。F 检验后(p > 0.10),假定总体方差相等。

    sₚ² = [(24×42² + 29×55²)] / (25+30–2) ≈ 2454.0 → sₚ ≈ 49.54

    t = (1002 – 972) / [49.54 × √(1/25 + 1/30)] ≈ 2.236, df = 53

    We test H₀: μ₁ – μ₂ = 0 against H₁: μ₁ – μ₂ ≠ 0 at α = 0.05. The two‑tailed critical value is around 2.006. Since 2.236 > 2.006, we reject H₀. There is significant evidence of a difference in true mean lifetimes between the two lines. The point estimate suggests line A produces bulbs lasting about 30 hours longer on average.

    我们检验 H₀: μ₁ – μ₂ = 0 对 H₁: μ₁ – μ₂ ≠ 0,α = 0.05。双尾临界值约为 2.006。因为 2.236 > 2.006,拒绝 H₀。有显著证据表明两条生产线的真实平均寿命存在差异。点估计显示 A 线灯泡平均寿命长约 30 小时。


    6. Linear Regression: Modelling Lifetime against Voltage | 线性回归:对寿命与电压建模

    Engineers suspect that the applied voltage influences bulb lifetime. Eight prototype bulbs are tested at different voltages (V): 110, 120, 130, 140, 150, 160, 170, 180, yielding lifetimes (h): 1020, 990, 965, 940, 910, 885, 860, 830. A scatterplot suggests a strong negative linear relationship.

    工程师怀疑施加的电压会影响灯泡寿命。在 110, 120, 130, 140, 150, 160, 170, 180 V 下测试了 8 只样品的寿命:1020, 990, 965, 940, 910, 885, 860, 830 小时。散点图显示强烈的负线性关系。

    Using the least‑squares method we find the fitted model: predicted lifetime = β₀̂ + β₁̂ × Voltage. Calculations give β₁̂ ≈ –2.150 and β₀̂ ≈ 1230. So for every extra volt, the lifetime drops by about 2.15 hours. The R² value is 0.95, indicating the model explains 95% of the variability in lifetime.

    用最小二乘法求得拟合模型:预测寿命 = β₀̂ + β₁̂ × 电压。计算得到 β₁̂ ≈ –2.150,β₀̂ ≈ 1230。因此电压每增加 1 伏特,寿命约减少 2.15 小时。R² = 0.95,表明模型解释了寿命 95% 的变异。

    The residuals show no obvious pattern, confirming that a linear model is appropriate. This simple regression gives engineers a powerful predictive tool for setting voltage limits.

    残差图无明显模式,确认线性模型适用。这一简单回归为工程师设定电压限制提供了有力的预测工具。


    7. Inferences for the Regression Slope | 回归斜率的推断

    To test whether the slope is genuinely non‑zero, we conduct a two‑tailed t‑test on β₁: H₀: β₁ = 0, H₁: β₁ ≠ 0. The standard error of β₁̂ is SE(β₁̂) ≈ 0.180. The test statistic is t = –2.150 / 0.180 ≈ –11.94 with df = 6. The p‑value is smaller than 0.001, so we strongly reject H₀.

    Published by TutorHao | Year 13 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Cross-Disciplinary Exam Practice for AQA Statistics | AQA统计跨学科综合题型训练

    📚 Cross-Disciplinary Exam Practice for AQA Statistics | AQA统计跨学科综合题型训练

    Welcome to this targeted cross-disciplinary revision resource for Year 13 AQA Statistics. In your examination, you will be expected to apply a wide range of statistical techniques to authentic contexts drawn from biology, psychology, economics, physics, geography, and beyond. The worked examples that follow will guide you through hypothesis tests, probability distributions, correlation, regression, and decision-making, always with a focus on the precise reasoning and interpretation demanded by AQA mark schemes.

    欢迎使用这份专为13年级AQA统计设计的跨学科综合题型训练。在考试中,你需要将多种统计方法应用于生物学、心理学、经济学、物理学、地理学等真实情境。下面的例题将带你一步步完成假设检验、概率分布、相关与回归以及决策分析,始终紧扣AQA评分方案所要求的严谨推理与解释。

    1. Genetics: Chi-Squared Goodness-of-Fit Test | 遗传学:卡方拟合优度检验

    Gregor Mendel’s classic pea experiments predicted that the phenotypes of dihybrid crosses should appear in a 9:3:3:1 ratio. A modern replication records the following counts: round-yellow 315, round-green 108, wrinkled-yellow 101, wrinkled-green 32. Use a 5% significance level to test whether the data are consistent with the theoretical ratio.

    孟德尔的经典豌豆实验预测双因子杂交的表型应按9:3:3:1的比例出现。一项现代重复实验记录了以下数据:圆黄315,圆绿108,皱黄101,皱绿32。使用5%的显著性水平检验数据是否符合理论比例。

    Step 1: State the hypotheses. H0: The data follow a 9:3:3:1 ratio; H1: The data do not follow this ratio.

    步骤1:提出假设。H0:数据符合9:3:3:1比例;H1:数据不符合该比例。

    Step 2: Calculate expected frequencies. Total observations n = 556. Expected counts are 9/16 × 556 = 312.75, 3/16 × 556 = 104.25, 3/16 × 556 = 104.25, 1/16 × 556 = 34.75.

    步骤2:计算期望频数。总观测数 n = 556。期望频数为:9/16 × 556 = 312.75,3/16 × 556 = 104.25,3/16 × 556 = 104.25,1/16 × 556 = 34.75。

    Step 3: Compute the test statistic.

    步骤3:计算检验统计量。

    χ2 = Σ (Oi – Ei)2 / Ei

    Phenotype Oi Ei (Oi-Ei)2/Ei
    Round-yellow 315 312.75 0.016
    Round-green 108 104.25 0.135
    Wrinkled-yellow 101 104.25 0.101
    Wrinkled-green 32 34.75 0.218

    χ2 = 0.016 + 0.135 + 0.101 + 0.218 = 0.470

    χ2 = 0.016 + 0.135 + 0.101 + 0.218 = 0.470

    Step 4: Degrees of freedom = number of categories – 1 = 3. Critical value from chi-squared table at 5% (3 d.f.) is 7.815.

    步骤4:自由度 = 类别数 – 1 = 3。查卡方分布表,在5%显著性水平、自由度为3时,临界值为7.815。

    Step 5: Since 0.470 < 7.815, we do not reject H0. There is insufficient evidence to suggest the data deviate from the 9:3:3:1 ratio.

    步骤5:由于0.470 < 7.815,我们不拒绝H0。没有足够证据表明数据偏离9:3:3:1比例。


    2. Psychology: Two-Sample t-Test for Independent Groups | 心理学:独立样本t检验

    A psychologist measures response times (ms) for a cognitive task under two conditions: a control group (n1=10) and a sleep-deprived group (n2=10). The summary statistics are: control mean x¯1 = 245 ms, standard deviation s1 = 18 ms; sleep-deprived mean x¯2 = 267 ms, s2 = 22 ms. Assume equal population variances and test whether sleep deprivation increases response time at the 5% significance level.

    一位心理学家测量了两种条件下的认知任务反应时间(毫秒):控制组(n1=10)和睡眠剥夺组(n2=10)。汇总统计量为:控制组均值 x¯1 = 245 ms,标准差 s1 = 18 ms;睡眠剥夺组均值 x¯2 = 267 ms,s2 = 22 ms。假设总体方差相等,在5%显著性水平下检验睡眠剥夺是否增加反应时间。

    Step 1: Hypotheses. H0: μ1 = μ2; H1: μ2 > μ1 (one-tailed).

    步骤1:假设。H0:μ1 = μ2;H1:μ2 > μ1(单侧检验)。

    Step 2: Pooled variance estimate.

    步骤2:合并方差估计。

    sp2 = [(n1-1)s12 + (n2-1)s22] / (n1+n2-2)

    = [9×324 + 9×484] / 18 = (2916 + 4356)/18 = 404.0, so sp = 20.10 ms.

    = [9×324 + 9×484] / 18 = (2916 + 4356)/18 = 404.0,因此 sp = 20.10 ms。

    Step 3: Test statistic.

    步骤3:检验统计量。

    t = (x¯2 – x¯1) / (sp × √(1/n1 + 1/n2))

    = (267 – 245) / (20.10 × √(0.1+0.1)) = 22 / (20.10 × 0.4472) = 22 / 8.988 = 2.448

    = (267 – 245) / (20.10 × √(0.1+0.1)) = 22 / (20.10 × 0.4472) = 22 / 8.988 = 2.448

    Step 4: Degrees of freedom = 18. Critical value for one-tailed t-test at 5% is 1.734. Since 2.448 > 1.734, we reject H0. There is significant evidence that sleep deprivation increases response time.

    步骤4:自由度 = 18。单侧5%显著性水平的t临界值为1.734。因为2.448 > 1.734,我们拒绝H0。有显著证据表明睡眠剥夺会增加反应时间。


    3. Economics: Linear Regression and Correlation | 经济学:线性回归与相关

    An economist collects annual data on household disposable income x (in £1000s) and consumption expenditure y (in £1000s) for five years: (30, 24), (35, 28), (40, 31), (45, 35), (50, 38). Determine the product-moment correlation coefficient, the regression equation of y on x, and predict expenditure when income is £52 000. Test whether the correlation is significant at the 1% level.

    一位经济学家收集了五年家庭可支配收入x(千英镑)和消费支出y(千英镑)的数据:(30, 24), (35, 28), (40, 31), (45, 35), (50, 38)。计算积矩相关系数、y对x的回归方程,并预测收入为52000英镑时的支出。在1%显著性水平下检验相关是否显著。

    Step 1: Compute summations: n=5, Σx=200, Σy=156, Σx2=8250, Σy2=4942, Σxy=6350.

    步骤1:计算总和:n=5,Σx=200,Σy=156,Σx2=8250,Σy2=4942,Σxy=6350。

    Step 2: Sxy = 6350 – (200×156)/5 = 6350 – 6240 = 110; Sxx = 8250 – 2002/5 = 8250 – 8000 = 250; Syy = 4942 – 1562/5 = 4942 – 4867.2 = 74.8.

    步骤2:Sxy = 6350 – (200×156)/5 = 110;Sxx = 8250 – 2002/5 = 250;Syy = 4942 – 1562/5 = 74.8。

    Step 3: PMCC r = Sxy / √(SxxSyy) = 110 / √(250×74.8) = 110 / √18700 = 110 / 136.75 = 0.804.

    步骤3:积矩相关系数 r = 110 / √(250×74.8) = 110 / 136.75 = 0.804。

    Step 4: For n=5, critical value for two-tailed test at 1% (3 d.f.) is 0.959. Since 0.804 < 0.959, we cannot reject H0: ρ=0. The correlation is not significant at the 1% level.

    步骤4:n=5,双尾1%显著性水平(自由度3)的临界值为0.959。由于0.804 < 0.959,不能拒绝H0: ρ=0。在1%水平下相关不显著。

    Step 5: Regression line: slope b = Sxy/Sxx = 110/250 = 0.44; intercept a = &ymacr; – b&xmacr; = (156/5) – 0.44×(200

    Published by TutorHao | Year 13 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Winter Intensive Revision Plan for OCR Year 12 Statistics | OCR 12年级统计:寒假强化复习计划

    📚 Winter Intensive Revision Plan for OCR Year 12 Statistics | OCR 12年级统计:寒假强化复习计划

    This plan is designed to help you use the winter break effectively to consolidate all the AS-level Statistics topics required for OCR Year 12. By following a structured approach, you can address knowledge gaps, build confidence, and be fully prepared for your upcoming assessments.

    这份复习计划旨在帮助你利用寒假有效巩固 OCR 12 年级统计所需的所有 AS 级别主题。通过结构化复习,你可以弥补知识漏洞,建立信心,并为即将到来的考试做好充分准备。


    1. Diagnostic Self-Assessment | 摸底自测

    Begin by completing a full AS Statistics past paper under timed conditions. This will highlight your strengths and pinpoint topics that need more attention.

    首先在计时条件下完成一套完整的 AS 统计历年真题。这将突出你的优势,并查明需要更多关注的主题。

    Record your scores for each section: data collection, representation, probability, distributions, and hypothesis testing. Create a list of errors and the associated concepts.

    记录每个部分的得分:数据收集、表示、概率、分布和假设检验。列出错误及相关概念。


    2. Mastering Data Collection | 掌握数据收集

    Review all sampling techniques: simple random, stratified, systematic, quota, and convenience sampling. Understand their advantages, disadvantages, and when each is appropriate.

    复习所有抽样技术:简单随机抽样、分层抽样、系统抽样、配额抽样和便利抽样。理解它们的优点、缺点以及适用场合。

    Distinguish between a population and a sample, and between a sampling frame and a census. Know the types of data: qualitative/categorical vs quantitative, discrete vs continuous.

    区分总体与样本,以及抽样框与普查。了解数据类型:定性/分类数据与定量数据,离散数据与连续数据。

    Be able to criticise data collection methods, identifying bias, lack of representativeness, and practical constraints.

    能够批判数据收集方法,识别偏差、缺乏代表性及实际限制。


    3. Data Presentation & Interpretation | 数据表示与解读

    Revisit how to construct and interpret histograms, cumulative frequency diagrams, box plots, and stem-and-leaf diagrams. For grouped data, ensure you can calculate frequency density (frequency ÷ class width) to draw histograms.

    重温如何绘制并解读直方图、累积频率图、箱线图以及茎叶图。对于分组数据,确保能计算频率密度(频率 ÷ 组距)来绘制直方图。

    Learn to identify outliers using the rule: outlier < Q₁ − 1.5 × IQR or > Q₃ + 1.5 × IQR, where IQR = Q₃ − Q₁. Interpret the shape of a distribution (symmetric, positively or negatively skewed) from these diagrams.

    学会使用规则识别离群值:离群值 < Q₁ − 1.5 × IQR 或 > Q₃ + 1.5 × IQR,其中 IQR = Q₃ − Q₁。根据这些图表解读分布形态(对称、正偏态或负偏态)。


    4. Measures of Central Tendency & Dispersion | 集中趋势与离散度量

    Calculate the mean, median, and mode for raw and grouped data. The sample mean is given by x̄ = Σx / n. For grouped data, use midpoints.

    计算原始数据和分组数据的均值、中位数和众数。样本均值公式为 x̄ = Σx / n。对于分组数据,使用组中值。

    Measures of spread include range, interquartile range (IQR), variance, and standard deviation. The sample variance formula is s² = Σ(x − x̄)² / (n − 1), and the population variance is σ² = Σ(x − μ)² / N. The standard deviation is the square root of the variance.

    离差的度量包括极差、四分位距 (IQR)、方差和标准差。样本方差公式为 s² = Σ(x − x̄)² / (n − 1),总体方差为 σ² = Σ(x − μ)² / N。标准差是方差的平方根。

    Understand that adding a constant shifts the mean but does not change the variance; multiplying by a constant scales both the mean and the standard deviation.

    理解加上常数只会平移均值但不改变方差;乘以常数会缩放均值和标准差。

    Be able to select the most appropriate measure of central tendency and spread for a given data set, especially when outliers are present.

    能够为给定数据集选择最合适的集中趋势和离散度量,尤其是存在离群值时。


    5. Probability Fundamentals | 概率基础

    Review the basic probability rules: for any event A, 0 ≤ P(A) ≤ 1; the total probability of all mutually exclusive outcomes is 1. Use Venn diagrams and two-way tables to organize probabilities.

    复习基本概率规则:对于任何事件 A,0 ≤ P(A) ≤ 1;所有互斥结果的总概率为 1。使用维恩图和双向表格来组织概率。

    The addition formula for two events: P(A ∪ B) = P(A) + P(B) − P(A ∩ B). For mutually exclusive events, P(A ∩ B) = 0, so P(A ∪ B) = P(A) + P(B).

    两个事件的加法公式:P(A ∪ B) = P(A) + P(B) − P(A ∩ B)。对于互斥事件,P(A ∩ B) = 0,因此 P(A ∪ B) = P(A) + P(B)。

    Independent events satisfy P(A ∩ B) = P(A) × P(B). Be careful not to confuse independence with mutual exclusivity.

    独立事件满足P(A ∩ B) = P(A) × P(B)。注意不要混淆独立与互斥。


    6. Conditional Probability & Tree Diagrams | 条件概率与树状图

    Conditional probability P(A|B) = P(A ∩ B) / P(B). This measures the probability of A given that B has occurred. Use tree diagrams to handle successive events and calculate combined probabilities.

    条件概率 P(A|B) = P(A ∩ B) / P(B)。这衡量在 B 已经发生的条件下 A 发生的概率。使用树状图处理连续事件并计算组合概率。

    When drawing a tree diagram, label branches with probabilities. For a sequence of two events, multiply along branches to find the probability of the intersection. Add probabilities of branches that satisfy a condition.

    绘制树状图时,用概率标记分支。对于两个事件的序列,沿分支相乘求交集的概率。将满足条件的分支概率相加。

    Practice reverse conditional probability problems where you are given a later probability and asked to find an earlier branch probability, often using the formula or a tree diagram with unknown probabilities.

    练习逆向条件概率问题,即给定后期概率求前期分支概率,通常使用公式或带有未知概率的树状图。


    7. Discrete Random Variables & Binomial Distribution | 离散随机变量与二项分布

    A discrete random variable X takes distinct values with probabilities P(X = x). You can display its probability distribution in a table. The expected value (mean) is E(X) = Σ [x · P(X = x)], and the variance is Var(X) = Σ [x² · P(X = x)] − [E(X)]².

    离散随机变量 X 取不同值,对应概率 P(X = x)。其概率分布可以用表格显示。期望值(均值)为 E(X) = Σ [x · P(X = x)],方差为 Var(X) = Σ [x² · P(X = x)] − [E(X)]²

    The binomial distribution models the number of successes in n independent trials, each with

    Published by TutorHao | Year 12 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Year 12 OCR Statistics: Case Study in Action | OCR Year 12 统计:案例分析实战演练

    📚 Year 12 OCR Statistics: Case Study in Action | OCR Year 12 统计:案例分析实战演练

    Welcome to a full walkthrough of a statistical investigation designed for Year 12 OCR Statistics students. In this case study, we examine a basketball player’s claim about his free-throw shooting accuracy. We will collect data, summarise it with charts, model the number of successes with a binomial distribution, and perform a hypothesis test to see whether the data cast doubt on the player’s stated success rate. Every step mirrors the content you meet in the OCR AS Statistics syllabus, helping you see how classroom theory turns into real-world practice.

    欢迎来到专为 OCR Year 12 统计学生设计的完整统计分析演练。本案例中,我们将考察一名篮球运动员对其罚球命中率的声称。我们将收集数据、用图表加以概括、用二项分布对成功次数建模,并进行假设检验,以判断数据是否对球员声称的成功率产生质疑。每一步都对应 OCR AS 统计课程的内容,帮助你直观理解课堂理论如何转为实际应用。


    1. Case Introduction: A Claim about Free-Throw Accuracy | 案例介绍:有关罚球命中率的声称

    A professional basketball player states that his free-throw success rate is 80%. In other words, he claims that on any given free-throw attempt, the probability of scoring is p = 0.8. We want to investigate whether there is statistical evidence that his true success rate is actually lower than 80%. To do this, we randomly select 20 of his free-throw attempts from recent matches and count how many are successful.

    一位职业篮球运动员称其罚球命中率为 80%。也就是说,他声称在任意一次罚球中,得分的概率为 p = 0.8。我们希望调查是否有统计证据表明其真实命中率实际上低于 80%。为此,我们从其近期比赛中随机选取 20 次罚球,统计其中成功的次数。


    2. Sampling and Data Collection: Ensuring a Fair Test | 抽样与数据收集:确保检验公正

    We decided to use a simple random sample of 20 free-throw attempts. The team video analyst assigned a number to every free-throw taken by the player in the last three months and used a random number generator to pick 20 distinct attempts. This method avoids selection bias; attempts from high-pressure moments and routine situations are equally likely to be chosen. It also approximates the independence condition needed for the binomial model, because the outcomes of widely separated attempts are less likely to influence each other.

    我们决定采用简单随机抽样,从该球员过去三个月内的所有罚球中抽取 20 次。球队视频分析师为每一次罚球编号,利用随机数生成器选出 20 次互不相同的尝试。这一方法避免了选择偏倚;高压时刻的罚球与一般情况下的罚球被选中的概率相同。同时,它也为二项模型所需的独立性条件提供了近似保证,因为间隔较远的罚球结果不太可能相互影响。


    3. Representing the Data: Frequency Table and Bar Chart | 数据表示:频数表与柱状图

    After watching the 20 selected attempts, we recorded the outcome of each as either ‘success’ or ‘failure’. The raw data are summarised in a frequency table and a bar chart.

    观看所选 20 次罚球后,我们将每次结果记录为“成功”或“失败”。原始数据通过频数表和柱状图进行总结。

    Outcome | 结果 Frequency | 频数
    Success | 成功 13
    Failure | 失败 7

    A bar chart (not shown here) would have two bars of height 13 and 7, instantly showing that the sample contains more successes than failures, but with a success proportion of 13/20 = 0.65. This is noticeably below the claimed 0.80.

    一幅柱状图(此处未显示)将有高度为 13 和 7 的两根柱,瞬间展现出样本中成功多于失败,但成功比例为 13/20 = 0.65,明显低于声称的 0.80。


    4. Summary Statistics: Sample Proportion and Variability | 摘要统计量:样本比例与变异性

    Let p̂ = 13/20 = 0.65 denote the sample proportion of successes. This is a point estimate of the player’s true success probability. Although a single number gives a snapshot, we also want a sense of how much it might vary from sample to sample. Under the assumed model with p = 0.8, the standard deviation of the number of successes, X, would be √(np(1-p)) = √(20 × 0.8 × 0.2) = √3.2 ≈ 1.7889, so the standard deviation of the sample proportion is 1.7889/20 = 0.0894.

    p̂ = 13/20 = 0.65 表示样本成功比例。它是球员真实成功概率的点估计。尽管单个样本比例提供了一个快照,我们也希望了解不同样本之间可能的波动程度。在假定 p = 0.8 的模型下,成功次数 X 的标准差为 √(np(1-p)) = √(20 × 0.8 × 0.2) = √3.2 ≈ 1.7889,因此样本比例的标准差为 1.7889/20 = 0.0894。


    5. The Binomial Distribution as a Model | 二项分布模型

    The number of successful free-throws in a fixed number of independent attempts with the same probability of success can be modelled by a binomial distribution. Let X be the random variable ‘number of successes in 20 attempts’. If the player’s claim is true, then X ~ B(20, 0.8). This means the probability that X equals a specific value x is given by:

    在相同成功概率的独立尝试中,固定次数的成功次数可以用二项分布建模。设随机变量 X 为“20 次尝试中的成功次数”。如果球员的声称属实,则 X ~ B(20, 0.8)。这意味着 X 取某个特定值 x 的概率由下式给出:

    P(X = x) = C(20, x) × 0.8ˣ × 0.2²⁰⁻ˣ

    Here C(20, x

    Published by TutorHao | Year 12 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Year 12 OCR Statistics: Unit Test Mock Paper Walkthrough | 单元测试模拟卷解析

    📚 Year 12 OCR Statistics: Unit Test Mock Paper Walkthrough | 单元测试模拟卷解析

    This article provides a detailed walkthrough of a mock unit test for Year 12 OCR Statistics, covering core topics such as probability, discrete random variables, the binomial distribution, and hypothesis testing. Each section breaks down a typical exam‑style question, offering step‑by‑step solutions and key revision points to help you master the techniques required for success.

    本文详细解析一份为 Year 12 OCR 统计课程设计的单元测试模拟卷,涵盖概率、离散随机变量、二项分布和假设检验等核心专题。每一节围绕一个典型试题展开,提供分步解答和关键复习要点,帮助你掌握考试所需的方法。


    1. Basic Probability and Venn Diagrams | 基础概率与维恩图

    A survey of 50 students finds that 28 like football (F), 18 like basketball (B), and 10 like both. One student is chosen at random. We illustrate the information using a Venn diagram and calculate probabilities.

    对50名学生调查显示,28人喜欢足球(F),18人喜欢篮球(B),10人两者都喜欢。我们使用维恩图表示信息并计算概率。

    Draw two overlapping circles. Place 10 in the intersection. The number who like only football is 28 − 10 = 18, only basketball is 18 − 10 = 8, and neither is 50 − (18 + 10 + 8) = 14.

    画出两个相交的圆。交集填10。只喜欢足球的人数为28−10=18,只喜欢篮球为18−10=8,都不喜欢为50−(18+10+8)=14。

    (i) P(F ∪ B) = (18 + 10 + 8) / 50 = 36/50 = 0.72. This is the probability a student likes at least one sport.

    (i) P(F ∪ B) = (18+10+8)/50 = 36/50 = 0.72。 这是学生喜欢至少一种运动的概率。

    (ii) P(F’ ∩ B) = 8/50 = 0.16. These students like only basketball.

    (ii) P(F’ ∩ B) = 8/50 = 0.16。 这些学生只喜欢篮球。

    (iii) P(B | F) = P(B ∩ F) / P(F) = 10/28 ≈ 0.357. Given a student likes football, there is a roughly 35.7% chance they also like basketball.

    (iii) P(B | F) = P(B ∩ F) / P(F) = 10/28 ≈ 0.357。 已知一名学生喜欢足球,他们同时也喜欢篮球的概率约为35.7%。


    2. Discrete Random Variables and Expectation | 离散随机变量与期望

    The probability distribution of a discrete random variable X is given in the table below. Show that it is a valid distribution and find E(X) and Var(X).

    离散随机变量X的概率分布如下表。证明它是合法的分布并求E(X)和Var(X)。

    x 0 1 2 3
    P(X=x) 0.2 0.35 0.3 0.15

    Check Σ P(X=x) = 0.2 + 0.35 + 0.3 + 0.15 = 1.00, and all probabilities are between 0 and 1. The distribution is valid.

    验证 Σ P(X=x) = 0.2+0.35+0.3+0.15 = 1.00,且所有概率介于0和1之间,分布合法。

    E(X) = Σ x·P(X=x) = (0×0.2) + (1×0.35) + (2×0.3) + (3×0.15) = 0 + 0.35 + 0.6 + 0.45 = 1.4.

    E(X²) = (0²×0.2) + (1²×0.35) + (2²×0.3) + (3²×0.15) = 0 + 0.35 + 1.2 + 1.35 = 2.9.

    Var(X) = E(X²) − [E(X)]² = 2.9 − 1.4² = 2.9 − 1.96 = 0.94.

    The expected value is 1.4 and the variance is 0.94.

    期望值为1.4,方差为0.94。


    3. Identifying the Binomial Setting | 识别二项分布的条件

    A factory produces components with a known defect rate of 2%. A quality inspector randomly selects 30 components and counts the number of defectives, X. Explain why X can be modelled by a binomial distribution.

    某工厂生产零件,已知次品率为2%。质检员随机抽取30个零件并记录次品数X。解释为什么X可以用二项分布建模。

    A binomial distribution B(n, p) requires: a fixed number of trials (n = 30); each trial is independent; only two outcomes – defective or not defective; and a constant probability of success (defect) p = 0.02. All conditions are met, so X ~ B(30, 0.02).

    二项分布B(n, p)要求:试验次数固定(n=30);每次试验独立;只有两种结果——次品或非次品;每次成功的概率恒定p=0.02。所有条件均满足,因此X ~ B(30, 0.02)。

    Always check for independence: the outcome of one component must not affect another. Random sampling from a large production run ensures this holds.

    务必检查独立性:一个零件的检测结果不得影响另一个。从大批量生产中随机抽样可保证该条件成立。


    4. Calculating Binomial Probabilities | 计算二项概率

    Using the model X ~ B(30, 0.02), find the probability that exactly two components are defective, P(X = 2).

    利用模型X ~ B(30, 0.02),计算恰好有两个次品的概率P(X = 2)。

    P(X = 2) = ³⁰C₂ × (0.02)² × (0.98)²⁸

    ³⁰C₂ = 435, so P(X = 2) ≈ 435 × 0.0004 × 0.5688 ≈ 0.0988.

    ³⁰C₂ = 435,因此P(X = 2) ≈ 435 × 0.0004 × 0.5688 ≈ 0.0988。

    To find P(X ≤ 2), sum P(X = 0) + P(X = 1) + P(X = 2). Using respective terms:
    P(X = 0) = ⁰.⁹⁸³⁰ ≈ 0.5455, P(X = 1) = ³⁰C₁ × 0.02 × 0.98²⁹ ≈ 0.3340.
    P(X ≤ 2) ≈ 0.5455 + 0.3340 + 0.0988 = 0.9783.

    求P(X ≤ 2),将P(X=0)+P(X=1)+P(X=2)相加。各项分别为:P(X=0)≈0.5455,P(X=1)≈0.3340,P(X≤2)≈0.9783。


    5. Mean and Variance of a Binomial Distribution | 二项分布的均值与方差

    For X ~ B(30, 0.02), calculate the expected number of defectives and the standard deviation.

    对于X ~ B(30, 0.02),求次品的期望个数和标准差。

    E(X) = np = 30 × 0.02 = 0.6.
    Var(X) = np(1−p) = 30 × 0.02 × 0.98 = 0.588.
    Standard deviation = √Var(X) = √0.588 ≈ 0.767.

    E(X) = np = 30×0.02 = 0.6;Var(X) = np(1−p) = 30×0.02×0.98 = 0.588;标准差≈0.767。

    On average we expect 0.6 defective components in a sample of 30, with a spread of about 0.77.

    在30个样本中平均预期有0.6个次品,标准差的波动约0.77。


    6. Introduction to Hypothesis Testing | 假设检验入门

    A restaurant claims that 90% of customers are satisfied. In a random sample of 50 customers, 42 report being satisfied. Test, at the 5% significance level, whether there is evidence that the satisfaction rate is lower than claimed.

    一家餐馆声称顾客满意度为90%。随机抽取50名顾客,其中42人表示满意。在5%显著性水平下检验是否有证据表明满意度低于声称值。

    Let p be the true proportion of satisfied customers. Set up hypotheses:
    H₀: p = 0.9, H₁: p < 0.9 (one‑tailed test).
    Under H₀, the number of satisfied customers in the sample X ~ B(50, 0.9). The observed value is 42.

    令p为真实满意比例。建立假设:H₀: p = 0.9,H₁: p < 0.9(单尾检验)。在H₀下,样本中满意人数X ~ B(50, 0.9)。观测值为42。

    Calculate p‑value = P(X ≤ 42 | p=0.9). Using binomial tables or calculator, P(X ≤ 42) ≈ 0.043. Since 0.043 < 0.05, we reject H₀.

    计算p值 = P(X ≤ 42 | p=0.9)。利用二项分布表或计算器得P(X ≤ 42)≈0.043。由于0.043 < 0.05,拒绝H₀。

    Conclusion: There is sufficient evidence at the 5% level to suggest that the satisfaction rate is lower than 90%.

    结论:在5%显著性水平下,有充分证据表明满意度低于90%。


    7. Setting Up Hypotheses Correctly | 正确建立假设

    Many errors occur when students write hypotheses. The null hypothesis H₀ must include an equality (=). The alternative H₁ reflects the suspicion being tested.

    许多错误出现在建立假设时。原假设H₀必须包含等号(=)。备择假设H₁反映要检验的疑虑。

    • Two‑tailed test: H₁: p ≠ value. Use when the claim is “different from” or “changed”.
    • 双尾检验:H₁: p ≠ 数值。当宣称“不同于”或“改变”时使用。
    • One‑tailed: H₁: p < value (lower tail) or H₁: p > value (upper tail), depending on wording such as “reduced” or “increased”.
    • 单尾:H₁: p < 数值(左侧)或 H₁: p > 数值(右侧),视“降低”或“提高”等措辞而定。

    For example, “test whether the proportion of faulty items is greater than 5%” gives H₁: p > 0.05.

    例如,“检验次品率是否大于5%”给出H₁: p > 0.05。

    Always define p clearly: “p = probability that …” or “p = true proportion of …”.

    始终明确定义p:“p = …的概率”或“p = …的真实比例”。


    8. Finding Critical Regions and p‑values | 求临界区域与p值

    For the coin‑bias example: X ~ B(20, 0.5), H₁: p > 0.5, α = 0.05. Find the critical region.

    对于硬币偏差举例:X ~ B(20, 0.5),H₁: p > 0.5,α=0.05。求临界区域。

    We need the smallest c such that P(X ≥ c) ≤ 0.05.
    P(X ≥ 15) = 0.0207 ≤ 0.05, P(X ≥ 14) = 0.0577 > 0.05. Hence the critical region is X ≥ 15.

    需要最小的c使得P(X ≥ c) ≤ 0.05。P(X ≥ 15)=0.0207 ≤ 0.05,P(X ≥ 14)=0.0577 > 0.05。因此临界区域为X ≥ 15。

    The p‑value for an observed x = 15 is P(X ≥ 15) = 0.0207. If the test statistic falls in the critical region, or p‑value < α, reject H₀.

    当观测值x=15时,p值 = P(X ≥ 15) = 0.0207。若检验统计量落入临界区域,或p值 < α,则拒绝H₀。

    CR (critical region) method is precise for discrete distributions; p‑value method is more common with technology. Both lead to the same conclusion.

    临界区域法在离散分布中很精确;p值法使用技术时更常见。两者结论一致。


    9. Drawing a Conclusion in Context | 在上下文中得出结论

    A conclusion must be written in the context of the problem, not as a generic statistical statement.

    Published by TutorHao | Year 12 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Core Concepts of Year 12 OCR Statistics | Year 12 OCR 统计核心知识点梳理

    📚 Core Concepts of Year 12 OCR Statistics | Year 12 OCR 统计核心知识点梳理

    This comprehensive guide covers the essential statistical topics for Year 12 OCR students, including data collection, presentation, summary statistics, probability, discrete random variables, the binomial distribution, and hypothesis testing. Mastering these concepts will build a strong foundation for both the AS-level examination and further statistical study.

    这份综合指南涵盖了Year 12 OCR学生必修的核心统计主题,包括数据收集、数据展示、汇总统计量、概率、离散随机变量、二项分布以及假设检验。掌握这些概念将为AS考试和进一步的统计学习奠定坚实的基础。

    1. Sampling Methods and Data Collection | 抽样方法与数据收集

    In statistics, a population is the entire set of individuals or items of interest, while a sample is a subset of the population used to draw inferences. A sampling frame is a list of all members of the population from which the sample is drawn.

    在统计学中,总体是研究对象的全体,而样本是总体中的一个子集,用于推断总体的特征。抽样框是包含总体所有成员并用于抽取样本的列表。

    Random sampling methods include simple random sampling, where every member has an equal chance of being selected, usually via random number generators; systematic sampling, where you select every kth member from the sampling frame after a random start; and stratified sampling, where the population is divided into strata and a random sample is taken from each stratum proportional to its size.

    随机抽样方法包括简单随机抽样(每个成员有相等的被选中的机会,通常通过随机数生成器实现)、系统抽样(在随机起始点后每隔k个成员抽取一个)以及分层抽样(将总体分成不同的层,然后按各层规模比例随机抽样)。

    Non-random sampling methods such as quota sampling (selecting individuals to meet predetermined quotas) and opportunity sampling (selecting those readily available) are easier but can introduce bias and do not represent the population fairly.

    非随机抽样方法如配额抽样(按预设配额选择个体)和便利抽样(选取最容易获得的个体)虽然操作简便,但可能引入偏差,无法公平地代表总体。

    Primary data is collected first-hand by the researcher, while secondary data has been collected by someone else. Secondary data can be cheaper and quicker but may not exactly fit the research question.

    一手数据由研究者直接收集;二手数据则由他人收集。二手数据成本较低且获取速度快,但可能与研究问题不完全匹配。


    2. Data Presentation and Interpretation | 数据展示与解读

    A stem-and-leaf diagram displays data by splitting each value into a stem (all but the last digit) and a leaf (the last digit). It retains the original data while showing the shape of the distribution.

    茎叶图通过将每个数据值分为“茎”(除最后一位数字外的部分)和“叶”(最后一位数字)来展示数据。它既能保留原始数据,又能显示分布形态。

    A box plot (or box-and-whisker plot) visually shows the median, lower quartile Q1, upper quartile Q3, and the range. Outliers are identified using the interquartile range (IQR): any data point below Q1 − 1.5 × IQR or above Q3 + 1.5 × IQR is considered an outlier.

    箱线图(或盒须图)直观地显示中位数、下四分位数Q1、上四分位数Q3以及极差。离群点通过四分位距(I

    Published by TutorHao | Year 12 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Year 13 Edexcel Statistics: A Parent’s Guide to Supporting Your Child | Year 13 Edexcel 统计:家长辅导指南

    📚 Year 13 Edexcel Statistics: A Parent’s Guide to Supporting Your Child | Year 13 Edexcel 统计:家长辅导指南

    As your child enters Year 13, the Edexcel A Level Statistics component becomes a major determinant of their final Mathematics grade. This parent’s guide explains what the syllabus covers, highlights common challenges and offers practical ways to support organised, confident revision.

    当您的孩子进入13年级,Edexcel A Level统计部分将成为他们数学最终成绩的重要决定因素。这份家长指南解释了课程大纲涵盖的内容,指出常见难点,并提供实际方法支持有条理、自信的复习。

    1. Understanding the Edexcel A Level Statistics Curriculum | 理解Edexcel A Level统计课程大纲

    Edexcel’s A Level Mathematics (9MA0) integrates Statistics and Mechanics. Statistics accounts for roughly half of the assessment, embedded within three 2‑hour papers. Students must apply statistical models, interpret findings in real‑world contexts and critique the limitations of those models.

    Edexcel A Level数学(9MA0)融合了统计与力学。统计约占评估内容的一半,分布在三份两小时的试卷中。学生必须应用统计模型,在真实情境中解释结果,并评判模型的局限性。

    The specification demands fluency with data summaries, probability distributions, hypothesis testing and correlation/regression. A large set of formulae is provided in the exam booklet, so the emphasis is on selecting and using the correct formula rather than rote memorisation.

    课程大纲要求学生熟练掌握数据汇总、概率分布、假设检验以及相关与回归。考试提供公式手册,因此重点在于选择并使用正确的公式,而非死记硬背。


    2. Key Topics in Year 13 Statistics | Year 13 统计的关键主题

    The core themes include: sampling techniques and bias; measures of centre and spread (mean, median, standard deviation, interquartile range); probability rules and Venn diagrams; discrete distributions (binomial and Poisson); the normal distribution; hypothesis testing for means and proportions; correlation and linear regression.

    核心主题包括:抽样方法与偏差;中心与离散程度(均值、中位数、标准差、四分位距);概率法则与维恩图;离散分布(二项与泊松);正态分布;均值与比例的假设检验;相关与线性回归。

    Within these, Edexcel frequently tests the conditions that justify using a particular distribution—for example, binomial requires fixed number of independent trials with constant probability. Understanding when to apply a normal approximation (np > 5, n(1 – p) > 5) is also essential.

    在这些主题中,Edexcel经常考查适用特定分布的条件——例如,二项分布需要固定次数的独立试验且每次概率恒定。理解何时应用正态近似(np > 5, n(1 – p) > 5)也至关重要。


    3. The Role of Data Representations and Summary Statistics | 数据表示与汇总统计的作用

    Students routinely work with histograms, box plots and cumulative frequency curves. Interpreting skewness, identifying outliers and comparing data sets using summary statistics are fundamental skills. Parents can help by asking their child to explain what a histogram’s area tells you about frequency.

    学生经常处理直方图、箱线图和累积频率曲线。解读偏度、识别异常值以及利用汇总统计比较数据集是基本技能。家长可以帮助,让孩子解释直方图的面积如何反映频率。

    Calculating the sample mean (x̄) and standard deviation (s) underpins most inference. The formula for sample standard deviation is one of the most used in the specification:

    计算样本均值(x̄)和标准差(s)是大多数推断的基础。样本标准差公式是课程大纲中最常用的公式之一:

    s = √( Σ(x – x̄)2 / (n – 1) )

    Encourage your child to practise typing these calculations into their calculator efficiently, checking that they are using the correct divisor (n‑1 for a sample).

    鼓励孩子练习在计算器上高效输入这些计算,并检查是否使用了正确的除数(样本用n‑1)。


    4. Probability Distributions: Binomial, Poisson and Normal | 概率分布:二项、泊松与正态

    The binomial model, B(n, p), counts successes in n independent trials: P(X = k) = ⁿCₖ pᵏ (1 – p)ⁿ⁻ᵏ. The Poisson distribution, Po(λ), models rare events and has probability function P(X = k) = (λᵏ e–λ) / k!.

    二项模型B(n, p)计算n次独立试验中的成功次数:P(X = k) = ⁿCₖ pᵏ (1 – p)ⁿ⁻ᵏ。泊松分布Po(λ)对稀有事件建模,其概率函数为P(X = k) = (λᵏ e–λ) / k!。

    The normal distribution N(μ, σ2) is continuous and symmetric. Standardising with Z = (X – μ) / σ allows the use of standard normal tables. Many students struggle to decide when to apply continuity corrections; regular practice with past papers helps embed these choices.

    正态分布N(μ, σ2)是连续且对称的。通过Z = (X – μ) / σ进行标准化后,便可使用标准正态表。许多学生难以决定何时应用连续性校正;定期练习往年真题有助于巩固这些判断。


    5. Hypothesis Testing and Confidence Intervals | 假设检验与置信区间

    A hypothesis test starts with a null hypothesis H₀ and an alternative H₁. Students calculate a test statistic or p‑value and compare it to the significance level α. The conclusion must be phrased in the context of the problem, never just ‘reject H₀’.

    假设检验从零假设H₀和备择假设H₁开始。学生计算检验统计量或p值,并将其与显著性水平α进行比较。结论必须结合问题背景,绝不仅仅是“拒绝H₀”。

    For the mean of a normal population with known variance, a confidence interval is x̄ ± z × (σ / √n). Edexcel expects learners to interpret confidence intervals correctly—stating, for example, that we are 95% confident the interval captures the true population mean.

    对于方差已知的正态总体均值,置信区间为x̄ ± z × (σ / √n)。Edexcel期望学生能正确解释置信区间——例如,说明我们有95%的信心认为该区间包含了真实的总体均值。

    Parents can check understanding by asking simple questions: ‘What does the p‑value actually measure?’ or ‘Why might we choose a one‑tailed test instead of two‑tailed?’ Such conversations reinforce reasoning.

    家长可以通过简单提问来检查理解程度:“p值实际上衡量什么?”或“为什么我们可能选择单尾检验而不是双尾检验?”此类对话能强化推理能力。


    6. Correlation and Regression Analysis | 相关与回归分析

    Pearson’s product‑moment correlation coefficient r measures linear association. The formula, given in the booklet, uses sums of squares:

    皮尔逊积矩相关系数r衡量线性关联程度。公式手册中给出的公式基于离差平方和:

    r = Sxy / √(Sxx Syy)

    where Sxx = Σ(x – x̄)2, Syy = Σ(y – ȳ)2 and Sxy = Σ(x – x̄)(y – ȳ). A value close to 1 or –1 indicates strong linear correlation, but students must also check for outliers that could distort r.

    其中Sxx = Σ(x – x̄)2, Syy = Σ(y – ȳ)2, Sxy = Σ(x – x̄)(y – ȳ)。接近1或–1的值表明强线性相关,但学生还需检查可能扭曲r的异常值。

    The regression line of y on x is y = a + bx, with b = Sxy / Sxx and a = ȳ – b x̄. Interpreting the gradient and intercept in real‑life terms is a common exam requirement.

    y对x的回归直线为y = a + bx,其中b = Sxy / Sxx,a = ȳ – b x̄。在现实情境中解释斜率和截距是常见的考试要求。


    7. Using Technology:

    Published by TutorHao | Year 13 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Year 13 Edexcel Statistics: Summer Preparation and Bridging Course | Year 13 爱德思统计:暑期预习与衔接课程

    📚 Year 13 Edexcel Statistics: Summer Preparation and Bridging Course | Year 13 爱德思统计:暑期预习与衔接课程

    The step from Year 12 to Year 13 in Edexcel Statistics marks a significant leap in mathematical rigour. While S1 introduces fundamental concepts such as probability, discrete random variables and the normal distribution, Year 13 Statistics (S2) demands a deeper understanding of continuous distributions, advanced hypothesis testing and approximations. This bridging article will help you consolidate your S1 knowledge and build a robust foundation for the topics ahead.

    从 Year 12 进入 Year 13 的爱德思统计课程,标志着数学严谨性的一次重大飞跃。S1 引入了概率、离散随机变量和正态分布等基本概念,而 Year 13 统计(S2)则要求深入理解连续分布、高级假设检验与近似方法。本文旨在帮助你在暑期巩固 S1 知识,为即将到来的新课题打下坚实基础。


    1. The Transition from S1 to S2 | 从 S1 到 S2 的过渡

    In S1, you focused on discrete random variables and the binomial distribution. Moving into Year 13, you will encounter continuous random variables where probabilities are defined over intervals rather than at exact points. A key mindset shift is from summation to integration.

    在 S1 中,你专注于离散随机变量和二项分布。进入 Year 13 后,你会遇到连续随机变量,其概率定义在区间上而非精确点。一个关键思维方式转变是从求和到积分的转变。

    You will also meet new distributions such as Poisson, and apply hypothesis tests to these contexts. Strong algebraic skills and familiarity with the normal distribution tables are essential.

    你还会接触泊松分布等新分布,并在这些情境下进行假设检验。扎实的代数功底和对正态分布表的熟悉至关重要。


    2. Recap: Discrete Random Variables and Distributions | 复习:离散随机变量及其分布

    A discrete random variable X takes a countable number of values. The probability mass function P(X = x) must satisfy ∑ P(X = x) = 1. In S1, you used the binomial distribution for a fixed number of independent trials.

    离散随机变量 X 取可数个值。概率质量函数 P(X = x) 必须满足 ∑ P(X = x) = 1。在 S1 中,你使用二项分布处理固定次数的独立试验。

    Be sure you can calculate E(X) = ∑ x P(X = x) and Var(X) = ∑ (x – μ)² P(X = x) = ∑ x² P(X = x) – μ². These formulas will be extended to continuous data in S2.

    确保你能计算 E(X) = ∑ x P(X = x) 和 Var(X) = ∑ (x – μ)² P(X = x) = ∑ x² P(X = x) – μ²。这些公式在 S2 中会推广到连续数据。


    3. Continuous Random Variables: The New Frontier | 连续型随机变量:新的前沿

    For a continuous random variable, probabilities are represented by the area under a curve, not by individual point probabilities. Consequently, P(X = c) = 0 for any single value c.

    对于连续型随机变量,概率由曲线下的面积表示,而不是个别点的概率。因此,对任意单个值 c,P(X = c) = 0。

    The total area under the probability density function (pdf) must equal 1. Instead of using sums, we integrate: ∫ f(x) dx over the range = 1.

    概率密度函数 (pdf) 曲线下的总面积必须等于 1。我们使用积分而非求和:在整个取值范围内 ∫ f(x) dx = 1。


    4. Probability Density Functions (PDFs) Unpacked | 概率密度函数(PDF)详解

    A pdf, typically denoted f(x), is valid if f(x) ≥ 0 for all x and the total integral over the sample space equals 1. You may be asked to find unknown constants in a pdf by setting the integral to 1.

    概率密度函数通常记作 f(x),若满足 f(x) ≥ 0 且在整个样本空间上的积分为 1,则是有效的。你可能需要通过令积分等于 1 来求出 pdf 中的未知常数。

    To calculate probabilities, compute the definite integral of f(x) over the desired interval: P(a < X < b) = ∫ab f(x) dx.

    计算概率时,求 f(x) 在所需区间上的定积分:P(a < X < b) = ∫ab f(x) dx。


    5. The Cumulative Distribution Function (CDF) | 累积分布函数(CDF)

    The cumulative distribution function F(x) gives the probability that X is less than or equal to x. For continuous distributions, F(x) = P(X ≤ x) = ∫-∞x f(t) dt.

    累积分布函数 F(x) 给出 X 小于等于 x 的概率。对于连续分布,F(x) = P(X ≤ x) = ∫-∞x f(t) dt。

    You can differentiate the CDF to obtain the pdf: f(x) = F'(x). This relationship is extremely useful for finding the median, quartiles and solving for unknown constants.

    你可以对 CDF 求导得到 pdf:f(x) = F'(x)。这种关系在求中位数、四分位数以及解未知常数时非常有用。


    6. Mean and Variance of Continuous Distributions | 连续型分布的期望与方差

    The mean (expected value) of a continuous random variable is E(X) = μ = ∫ x f(x) dx, over the appropriate domain. This directly parallels the discrete formula with integration replacing summation.

    连续型随机变量的均值(期望)为 E(X) = μ = ∫ x f(x) dx,积分区间对应其定义域。这与离散公式直接类似,只不过以积分代替求和。

    Variance is given by Var(X) = E(X²) – μ², where E(X²) = ∫ x² f(x) dx. You will often need to evaluate these integrals to compare distributions or to prepare for further inference.

    方差由 Var(X) = E(X²) – μ² 给出,其中 E(X²) = ∫ x² f(x) dx。你经常需要计算这些积分来比较分布或为后续推断做准备。


    7. The Normal Distribution Revisited | 再探正态分布

    S1 introduced the normal distribution and Z-scores. Year 13 deepens this by requiring you to find unknown means or standard deviations given probabilities. You must be proficient with the standard normal table, Φ(z).

    S1 介绍了正态分布和 Z 分数。Year 13 对此加深,要求你根据给定的概率求出未知的均值或标准差。你必须熟练掌握标准正态分布表 Φ(z)。

    If X ~ N(μ, σ²), then Z = (X – μ) / σ. Often you will set up equations like P(X < a) = 0.95 and solve for μ. Reverse-table reading is a key skill.

    若 X ~ N(μ, σ²),则 Z = (X – μ) / σ。你通常需要建立如 P(X < a) = 0.95 的方程来求解 μ。反向查表是一项关键技能。


    8. The Poisson Distribution | 泊松分布

    The Poisson distribution models the number of events occurring in a fixed interval of time or space, with a known constant mean rate λ. The probability mass function is P(X = x) = e λx / x!.

    泊松分布模拟在固定时间或空间区间内发生的事件次数,具有已知的恒定平均率 λ。其概率质量函数为 P(X = x) = e λx / x!。

    You will learn to use the Poisson distribution in contexts such as call arrivals, defects per metre, or radioactive decay. The mean and variance are both equal to λ.

    你将学习在如电话呼入、每米缺陷数或放射性衰变等情境下应用泊松分布。其期望和方差都等于 λ。


    9. Hypothesis Testing: From Binomial to Poisson | 假设检验:从二项到泊松

    In S1, you performed binomial hypothesis tests for proportions. Year 13 extends this to Poisson tests, testing whether the rate λ has increased or decreased. The logic of null and alternative hypotheses remains identical.

    在 S1 中,你针对比例进行了二项检验。Year 13 将其拓展到泊松检验,检验比率 λ 是否增加或减少。原假设与备择假设的逻辑完全相同。

    You will need to calculate P(X ≥ observed | H0) or P(X ≤ observed | H0) using the Poisson distribution, and compare with the significance level. Critical region finding and p-values are central.

    你需要用泊松分布计算 P(X ≥ 观测值 | H0) 或 P(X ≤ 观测值 | H0),并与显著性水平比较。寻找临界区域以及计算 p 值是核心内容。


    10. Normal Approximations (Binomial and Poisson) | 正态近似(二项与泊松)

    When n is large, the binomial distribution B(n, p) can be approximated by a normal distribution N(np, np(1-p)). Similarly, a Poisson with large λ can be approximated by N(λ, λ). Continuity corrections are essential.

    当 n 很大时,二项分布 B(n, p) 可用正态分布 N(np, np(1-p)) 近似。类似地,λ 较大的泊松分布可用 N(λ, λ) 近似。连续性校正是必不可少的。

    You must apply a correction of ±0.5 to the interval bounds when approximating a discrete distribution with a continuous one. This refines the accuracy of the approximation.

    在用连续分布近似离散分布时,必须对区间边界做 ±0.5 的校正。这能提高近似的精确程度。


    11. Practical Tips for Summer Study | 暑期学习实用建议

    Revisit your S1 notes, especially the normal distribution and hypothesis testing chapters. Practise integration of polynomials and exponential functions, as these frequently appear in pdf and CDF problems.

    重温你的 S1 笔记,尤其是正态分布和假设检验章节。练习多项式和指数函数的积分,因为这些在 pdf 和 CDF 题中频繁出现。

    Work through a few introductory S2 problems on continuous random variables and Poisson distribution without pressure. Aim for conceptual understanding rather than speed. A little each day goes a long way.

    轻松地尝试一些有关连续随机变量和泊松分布的入门 S2 题目。注重概念理解而非速度。每天学一点,积少成多。


    12. Looking Ahead: Exam Success | 展望:考试成功

    Year 13 Statistics rewards clarity of thought and a systematic approach. Label your hypotheses, show your integration steps, and always check conditions before using approximations. Consistent practice with past paper questions will build your confidence.

    Year 13 统计奖励思维清晰、条理有序。写明你的假设,展示积分步骤,并在使用近似前始终检查条件。持续练习历年真题将树立你的信心。

    Embrace the bridging period as an opportunity to strengthen fundamentals without the pressure of deadlines. With a solid summer preparation, you will hit the ground running in September.

    将衔接期视为在没有时间压力下巩固基础的机会。通过扎实

    Published by TutorHao | Year 13 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Year 13 Edexcel Statistics: Quick Vocabulary Mnemonic Guide | 爱德思Year 13统计:词汇术语速记指南

    📚 Year 13 Edexcel Statistics: Quick Vocabulary Mnemonic Guide | 爱德思Year 13统计:词汇术语速记指南

    Mastering the vocabulary of Year 13 Edexcel Statistics is essential for interpreting questions accurately and scoring top marks. This guide provides clear explanations, paired examples, and memorable mnemonics for all key terms you will encounter in the A2 statistics syllabus, including hypothesis testing, distributions, and sampling.

    掌握Year 13爱德思统计学科的术语是准确理解题目并取得高分的关键。本指南将为你清晰讲解A2统计大纲中的所有核心术语,包括假设检验、概率分布和抽样方法,并提供中英配对释义和实用记忆法。

    1. Key Concepts in Hypothesis Testing | 假设检验核心术语

    The null hypothesis (H₀) is a statement of no effect or no difference, assumed true until evidence suggests otherwise. The alternative hypothesis (H₁ or Hₐ) is what you want to prove, often indicating a change, difference, or association.

    原假设 (H₀) 表述为“无效应”或“无差异”,在被证据否定之前我们都假定它为真。备择假设 (H₁ 或 Hₐ) 则是研究者希望证实的命题,通常表明存在变化、差异或关联。

    The significance level (α) is the probability of rejecting H₀ when it is actually true (Type I error). The p-value is the probability of obtaining a test statistic at least as extreme as the observed one, assuming H₀ is true. If p-value ≤ α, we reject H₀.

    显著性水平 (α) 是当原假设为真时拒绝它的概率(第一类错误)。p值是在原假设成立的条件下,得到当前及更极端检验统计量的概率。若 p ≤ α,则拒绝原假设。

    A one-tailed test examines whether a parameter is greater than or less than a specified value, while a two-tailed test checks for any difference (either direction). The critical region is chosen accordingly, with the significance level split for two tails.

    单尾检验考察参数是否大于或小于某一特定值;双尾检验则检验是否存在任何方向上的差异。临界域相应划分,双尾检验将α平分至左右两侧。


    2. Probability Distributions Overview | 概率分布总览

    A probability distribution describes how probabilities are allocated across all possible outcomes of a random variable. For discrete variables, we use a probability mass function (PMF) giving P(X=x). For continuous variables, a probability density function (PDF) f(x) is used, where probabilities are found as areas under the curve.

    概率分布描述随机变量所有可能结果的概率分配方式。离散随机变量使用概率质量函数(PMF)表示 P(X=x);连续随机变量则用概率密度函数(PDF) f(x) 刻画,概率对应曲线下面积。

    The cumulative distribution function (CDF) F(x)=P(X ≤ x) gives the accumulated probability up to x. The expectation E(X) is the long-run average, and variance Var(X) measures spread.

    累积分布函数(CDF) F(x)=P(X ≤ x) 给出到x为止的累计概率;期望 E(X) 是长期平均,方差 Var(X) 衡量离散程度。


    3. Poisson Distribution in Detail | 泊松分布详解

    A Poisson distribution models the number of independent, random events occurring in a fixed interval of time or space, given a known average rate λ (lambda). Conditions: events occur singly, at a constant average rate, and independently of each other.

    泊松分布用于模拟固定时间或空间间隔内独立随机事件的发生次数,已知平均发生率 λ(lambda)。条件:事件单独发生,平均速率恒定,且相互独立。

    The probability mass function is:

    P(X = x) = (e⁻λ λˣ) / x! for x = 0, 1, 2, …

    概率分布为:P(X = x) = (e⁻λ λˣ) / x! ,x取非负整数。

    Important properties: The mean of a Poisson is λ and the variance is also λ. If X~Po(λ₁) and Y~Po(λ₂) are independent, then X+Y~Po(λ₁+λ₂). This additivity is unique to the Poisson distribution.

    重要性质:均值为 λ,方差也为 λ。若独立的 X~Po(λ₁) 和 Y~Po(λ₂),则 X+Y~Po(λ₁+λ₂)。这种可加性是泊松分布的一个特色。


    4. Continuous Distributions: Normal & Others | 连续分布:正态及其他

    The normal distribution N(μ, σ²) is symmetric and bell-shaped, described by its mean μ and standard deviation σ. About 68% of data lie within μ ± σ, 95% within μ ± 2σ, and 99.7% within μ ± 3σ.

    正态分布 N(μ, σ²) 呈对称钟形,由均值 μ 和标准差 σ 决定。约68%的数据落在 μ ± σ 内,95%在 μ ± 2σ 内,99.7%在 μ ± 3σ 内。

    To standardise, we use Z = (X – μ) / σ, giving Z~N(0,1). Probability calculations use the standard normal table or inverse normal functions.

    标准化公式为 Z = (X – μ) / σ,所得 Z~N(0,1)。概率计算借助标准正态分布表或逆正态函数。

    Other continuous distributions that may appear include the continuous uniform distribution where probabilities are proportional to length, and Student’s t-distribution for small samples when σ is unknown.

    其他可能出现的连续分布包括:连续均匀分布(概率与区间长度成正比)和用于小样本且总体标准差未知时的学生 t 分布。


    5. Sampling Methods & Bias | 抽样方法与偏差

    A simple random sample gives every member of the population an equal chance of being selected. Stratified sampling divides the population into distinct

    Published by TutorHao | Year 13 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Edexcel Year 13 Statistics: Unit Test Mock Paper Walkthrough | Edexcel 高三统计单元测试模拟卷解析

    📚 Edexcel Year 13 Statistics: Unit Test Mock Paper Walkthrough | Edexcel 高三统计单元测试模拟卷解析

    This walkthrough breaks down a representative mock paper for the Year 13 Edexcel Statistics component (Paper 3: Statistics & Mechanics). The mock covers the core A2 topics: normal distributions, sampling distributions, normal approximations, hypothesis testing for means, regression, correlation, and hypothesis tests on correlation. By working through these model solutions, you will reinforce the key techniques and examiner expectations, building confidence for the real examination.

    本解析逐题拆解一份具有代表性的 Edexcel 高三统计模拟卷(试卷三:统计与力学)。模拟卷涵盖 A2 核心主题:正态分布、抽样分布、正态近似、均值假设检验、回归、相关以及相关性的假设检验。通过研读这些范例解答,你将巩固关键技巧和考官所期望的答题规范,为实考积累信心。


    1. Mock Paper Overview | 模拟卷概览

    The mock paper is structured like a real Edexcel Statistics paper, blending short-answer questions with multi-step contextual problems. Time management is crucial: allocate about 1 minute per mark. Questions often build from simple probability calculations to hypothesis testing and interpretation, so reading the full scenario before jumping in helps avoid missing details.

    本模拟卷的编排贴近真实的 Edexcel 统计试卷,将简答题与多步骤情境题相结合。时间管理至关重要:建议按每分钟一分的节奏分配答题时间。题目经常从简单的概率计算逐步过渡到假设检验与解读,因此在动笔前通读完整的情景有助于避免遗漏细节。


    2. Working with the Normal Distribution | 正态分布计算

    To find a probability such as P(X > 55) when X ~ N(50, 4²), first standardise: Z = (55 – 50) / 4 = 1.25. The standard normal table gives P(Z < 1.25) = 0.8944, so P(X > 55) = 1 – 0.8944 = 0.1056. Always sketch the normal curve to visualise the tail area.

    若要计算 P(X > 55),其中 X ~ N(50, 4²),首先标准化:Z = (55 – 50) / 4 = 1.25。标准正态表给出 P(Z < 1.25) = 0.8944,因此 P(X > 55) = 1 – 0.8944 = 0.1056。务必画一条正态曲线来直观显示尾部面积。

    Z = (X – μ) / σ

    When a percentile is given, e.g., find x such that P(X < x) = 0.9, use the inverse normal function. The standard normal quantile is z = Φ⁻¹(0.9) ≈ 1.2816. Then x = μ + z σ = 50 + 1.2816 × 4 = 55.13. On a calculator, this is accessed via the Inverse Normal menu with area 0.9, μ = 50, σ = 4.

    当给出分位数,例如求 x 使 P(X < x) = 0.9,使用逆正态函数。标准正态分位数为 z = Φ⁻¹(0.9) ≈ 1.2816,则 x = μ + z σ = 50 + 1.2816 × 4 = 55.13。在计算器上,通过逆正态菜单输入面积 0.9、μ = 50、σ = 4 即可求得。


    3. Inverse Normal for Unknown Parameters | 逆正态求未知参数

    Some exam questions give two probability conditions while both μ and σ are unknown. For instance, P(X < 15) = 0.2 and P(X > 35) = 0.1. Set up standardised equations: (15 – μ)/σ = Φ⁻¹(0.2) ≈ -0.8416 and (35 – μ)/σ = Φ⁻¹(0.9) ≈ 1.2816 (since P(X > 35)=0.1 implies P(X < 35)=0.9). Solving the two simultaneous equations yields μ and σ.

    某些考题会在 μ 和 σ 均未知时给出两个概率条件。例如 P(X < 15) = 0.2 且 P(X > 35) = 0.1。建立标准化方程:(15 – μ)/σ = Φ⁻¹(0.2) ≈ -0.8416,(35 – μ)/σ = Φ⁻¹(0.9) ≈ 1.2816(因为 P(X > 35)=0.1 意味着 P(X < 35)=0.9)。联立这两个方程即可解出 μ 和 σ。

    Subtract the first from the second: (35 – μ) – (15 – μ) = σ(1.2816 – (-0.8416)) ⇒ 20 = σ × 2.1232 ⇒ σ ≈ 9.42. Substitute back into the first equation: (15 – μ) = -0.8416 × 9.42, giving μ ≈ 22.93. Always check that the resulting probabilities are consistent.

    用第二个方程减去第一个:(35 – μ) – (15 – μ) = σ(1.2816 – (-0.8416)) ⇒ 20 = σ × 2.1232 ⇒ σ ≈ 9.42。代回第一个方程,(15 – μ) = -0.8416 × 9.42,得 μ ≈ 22.93。最后务必检查所得概率是否一致。


    4. Sample Means and the Central Limit Theorem | 样本均值与中心极限定理

    If a random sample of size n is taken from a population with mean μ and variance σ², the sample mean X̄ has mean μ and variance σ²/n. For a normal population, X̄ is exactly normal; for non-normal populations with large n (typically n ≥ 30), the Central Limit Theorem ensures X̄ is approximately normal.

    若从均值为 μ、方差为 σ² 的总体中抽取容量为 n 的随机样本,样本均值 X̄ 的均值为 μ,方差为 σ²/n。若总体服从正态分布,X̄ 也精确服从正态分布;对于非正态总体且 n 较大(一般 n ≥ 30),中心极限定理保证 X̄ 近似服从正态分布。

    Example: The weight of apples has μ = 150 g and σ = 20 g. For a random sample of 25 apples, X̄ ~ N(150, 20²/25) i.e. N(150, 16). The standard error is σ/√n = 4 g. To find P(X̄ > 155), compute Z = (155 – 150)/4 = 1.25, giving a probability of about 0.1056.

    例如:苹果重量 μ = 150 g, σ = 20 g,随机抽取 25 个苹果,则 X̄ ~ N(150, 20²/25) 即 N(150, 16)。标准误为 σ/√n = 4 g。计算 P(X̄ > 155),Z = (155 – 150)/4 = 1.25,概率约为 0.1056。


    5. Normal Approximation to the Binomial | 二项分布的正态近似

    When X ~ B(n, p) and both np and nq are > 5 (some texts use 10), the distribution of X can be approximated by Y ~ N(np, npq). The continuity correction adjusts for the discrete nature: P(X ≥ r) ≈ P(Y > r – 0.5), P(X ≤ r) ≈ P(Y < r + 0.5), and P(X = r) ≈ P(r - 0.5 < Y < r + 0.5).

    当 X ~ B(n, p) 且 np 与 nq 均大于 5(有些教材用 10),X 的分布可由 Y ~ N(np, npq) 近似。连续性校正针对离散特性进行调整:P(X ≥ r) ≈ P(Y > r – 0.5),P(X ≤ r) ≈ P(Y < r + 0.5),而 P(X = r) ≈ P(r - 0.5 < Y < r + 0.5)。

    Suppose X ~ B(100, 0.35). Then np = 35, npq = 22.75, σ ≈ 4.77. To find P(X ≥ 40), apply the continuity correction: P(X ≥ 40) ≈ P(Y > 39.5). Z = (39.5 – 35)/4.77 ≈ 0.944, giving a tail probability of about 0.1726. Without the correction, the result would be noticeably less accurate.

    假设 X ~ B(100, 0.35),np = 35, npq = 22.75,σ ≈ 4.77。求 P(X ≥ 40),使用连续性校正:P(X ≥ 40) ≈ P(Y > 39.5)。Z = (39.5 – 35)/4.77 ≈ 0.944,尾部概率约为 0.1726。若不进行校正,结果的准确性会明显下降。


    6. Hypothesis Test for a Population Mean | 总体均值的假设检验

    When the population standard deviation σ is known, the test statistic for the mean is Z = (X

    Published by TutorHao | Year 13 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Year 13 Edexcel Statistics: Oral & Listening Exam Preparation | Edexcel 统计:口语听力备考专项

    📚 Year 13 Edexcel Statistics: Oral & Listening Exam Preparation | Edexcel 统计:口语听力备考专项

    In the Year 13 Edexcel Statistics examination, success depends not only on calculations but also on your ability to explain reasoning clearly and interpret question cues accurately — skills akin to “speaking” and “listening” in a conversation about data. This article guides you through the essential techniques to articulate statistical concepts verbally, decode exam prompts effectively, and present your written answers with the precision of a spoken explanation.

    在 Year 13 Edexcel 统计考试中,成功不仅取决于计算,还取决于你清晰解释推理和准确解读题目线索的能力——这类似于在数据对话中“说”和“听”的技能。本文指导你掌握口头表达统计概念的核心技巧,有效破解考题提示,并以口语解释般的精确度呈现书面答案。

    1. The Speaking Skill: Articulating Statistical Reasoning | 口语技能:清晰表达统计推理

    To “speak” in a statistics exam means translating numerical findings into plain English. When you conclude a hypothesis test, state clearly whether you reject H₀, what the p-value indicates, and the context of the problem.

    在统计考试中“说话”是指将数值发现转化为通俗英语。当你完成假设检验时,要清楚说明是否拒绝 H₀、p 值的含义以及问题的上下文。


    2. Listening to the Question: Decoding Command Words | 倾听题目:解读指令词

    “Listening” in exam terms means carefully reading command words such as ‘interpret’, ‘compare’, or ‘suggest’. For example, “interpret a confidence interval” requires you to state what the interval means, not just calculate it.

    考试中的“倾听”是指仔细阅读指令词,例如“interpret”、“compare”或“suggest”。例如,“interpret a confidence interval”要求你说出区间的含义,而不仅仅是计算。


    3. Describing Distributions Like a Narrator | 像叙述者一样描述分布

    When asked to describe a box plot or histogram, use a clear narrative: mention shape (symmetric/skewed), centre (median/mean), spread (range/IQR), and any outliers. Speak as if you are telling a story about the data.

    当被要求描述箱线图或直方图时,使用清晰的叙述:提及形状(对称/偏斜)、中心(中位数/均值)、离散度(极差/四分位距)以及任何异常值。像讲述数据故事一样去表达。


    4. Hypothesis Testing: The Verbal Summary | 假设检验:口头总结

    After computing the test statistic, articulate the decision: “Since the test statistic z = 1.96 exceeds the critical value 1.645, we reject the null hypothesis at the 5% significance level.” Practice saying this aloud mentally as you write.

    在计算检验统计量后,清晰地陈述决策:“由于检验统计量 z = 1.96 超过临界值 1.645,我们在 5% 显著性水平下拒绝原假设。” 书写时在心里默念这句话,如同口述。


    5. Confidence Intervals in Conversation | 用对话方式说明置信区间

    Interpret a 95% confidence interval for a mean: “We are 95% confident that the true population mean lies between 72.3 and 78.1.” Avoid saying “there is a 95% chance” — use the correct phrasing as if you were explaining to a listener.

    解释均值的 95% 置信区间:“我们 95% 确信真正的总体均值在 72.3 到 78.1 之间。” 避免说“有 95% 的可能性”,要像向听众解释那样使用正确的措辞。


    6. Correlation and Regression: Tell the Relationship | 相关与回归:讲述关系

    When stating the conclusion from a PMCC, speak the correlation: “The product-moment correlation coefficient r = 0.82 suggests a strong positive linear relationship between hours studied and exam score.” Use words like “suggests” or “indicates” to maintain statistical cautiousness.

    从 PMCC 得出结论时,说出相关性:“积矩相关系数 r = 0.82 表明学习时间与考试分数之间存在强正线性关系。” 使用“表明”或“显示”等词以保持统计上的谨慎。


    7. Probability Statements: Making Sense of Results | 概率陈述:让结果有意义

    Translate a probability value into everyday language: “P(X > 5) = 0.12 means there is a 12% probability of observing more than 5 successes under the assumed model.” This verbal translation helps both you and the examiner follow your logic.

    将概率值转化为日常语言:“P(X > 5) = 0.12 表示在假设模型下观察到超过 5 次成功的概率为 12%。” 这种口头转化有助于你和考官都理解你的逻辑。


    8. Normal Distribution “Listening” for Conditions | 正态分布的“听力”:检查条件

    Before applying the normal distribution, listen carefully to whether the question mentions approximate normality, a large sample size, or known population variance. Use the “hearing” skill to spot these clues.

    在应用正态分布之前,仔细倾听(审题)题目是否提及近似正态、大样本量或已知的总体方差。运用“听力”技巧来发现这些线索。


    9. Sampling and Bias: Explaining as You Would to a Friend | 抽样与偏差:像对朋友解释一样

    If a question asks about sampling method limitations, speak plainly: “Using a convenience sample may introduce selection bias because it does not represent all groups.” Practice saying your answer quietly; this reinforces clear writing.

    如果题目询问抽样方法的局限性,请直白地说:“使用便利样本可能会引入选择偏差,因为它不代表所有群体。” 练习小声说出你的答案;这能强化清晰的书面表达。


    10. The Central Limit Theorem in Plain Speech | 中心极限定理的通俗讲述

    Simplify the CLT: “For a large sample size, the sampling distribution of the sample mean becomes approximately normal, even if the population is not normal.” Phrase it as if you are teaching someone who has never heard of it before.

    简化中心极限定理:“对于大样本量,样本均值的抽样分布近似正态,即使总体不是正态。” 就像在教一个从未听说过的人一样去表述。


    11. Error Types: Pronouncing Your Verdict | 两类错误:宣告你的裁决

    Distinguish Type I and Type II errors with a clear spoken logic: “A Type I error occurs if we reject H₀ when it is actually true — like a false alarm. A Type II error is failing to reject H₀ when it is false — a missed detection.” Writing this out mimics a confident oral explanation.

    用清晰的口头逻辑区分第一类错误和第二类错误:“如果我们拒绝了真实的 H₀,就发生了第一类错误——就像误报。第二类错误是没有拒绝错误的 H₀——属于漏报。” 把这写出来就像进行一次自信的口头解释。


    12. Practice with Your Inner Voice | 用内心声音进行练习

    During revision, read each problem aloud (or under your breath) and answer verbally before writing. This “speaking” rehearsal speeds up response fluency and reduces the chance of writing ambiguous statements during the actual exam.

    在复习时,大声(或用气息)读出每道题,并在书写前口头作答。这种“说话”排练能加快答题流畅度,并减少在实际考试中写出模糊陈述的可能性。


    Published by TutorHao | Statistics Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Year 13 Edexcel Statistics: Formula and Theorem Quick Reference Handbook | Year 13 Edexcel 统计:公式定理速查手册

    📚 Year 13 Edexcel Statistics: Formula and Theorem Quick Reference Handbook | Year 13 Edexcel 统计:公式定理速查手册

    This quick reference handbook compiles the essential formulas, theorems, and key concepts for the Year 13 Edexcel Statistics course. It serves as a concise revision aid, covering probability distributions, sampling theory, hypothesis testing, and regression analysis. Every effort has been made to present the material in a clear, bilingual format to support both conceptual understanding and exam preparation.

    本速查手册汇集了 Year 13 Edexcel 统计课程的核心公式、定理和重要概念,是一份精炼的复习工具,覆盖概率分布、抽样理论、假设检验和回归分析。内容以中英双语清晰呈现,助力同学在理解概念的同时高效备考。


    1. Probability Foundations & Expectation | 概率基础与期望

    For a discrete random variable X, the expectation is the probability‑weighted average of its possible values.

    对于离散随机变量 X,期望是其所有可能取值的概率加权平均。

    E(X) = Σ x·P(X = x)

    The variance measures the spread of the distribution and can be calculated using the alternative formula.

    方差衡量分布的离散程度,可用简便公式计算。

    Var(X) = E(X2) − [E(X)]2

    Two events A and B are independent if and only if P(AB) = P(A)·P(B). Independence is crucial when combining random variables.

    当且仅当 P(AB) = P(A)·P(B) 时,事件 AB 独立。独立性的概念在随机变量的组合中至关重要。

    The expectation is linear, so for any constants a and b,

    期望具有线性性,对任意常数 ab

    E(aX + b) = aE(X) + b

    and the variance transforms as

    方差变换规律为

    Var(aX + b) = a2Var(X).

    For independent random variables X and Y, Var(X ± Y) = Var(X) + Var(Y), noting that the variance of a difference is still the sum of variances.

    若随机变量 XY 独立,Var(X ± Y) = Var(X) + Var(Y),注意差的方差仍然是方差相加。


    2. Binomial Distribution | 二项分布

    A binomial distribution models the number of successes in a fixed number n of independent Bernoulli trials, each with the same success probability p.

    二项分布描述在 n 次独立伯努利试验中成功的次数,每次试验成功的概率均为 p

    X ~ B(n, p)

    The probability of obtaining exactly k successes is

    恰好获得 k 次成功的概率为

    P(X = k) = nCk pk (1 − p)n−k

    where nCk = n!/(k!(n−k)!). The mean and variance are simple to compute:

    其中 nCk = n!/(k!(n−k)!)。其均值与方差便于计算:

    E(X) = np  Var(X) = np(1 − p)

    The binomial distribution is symmetric when p = 0.5 and becomes skewed for extreme p. It forms the foundation for many hypothesis tests involving proportions.

    p = 0.5 时二项分布是对称的,p 趋近极端值时则呈现偏态。它为许多关于比例的假设检验奠定基础。


    3. Poisson Distribution | 泊松分布

    The Poisson distribution models the number of events occurring in a fixed interval of time or space when events happen independently at a constant average rate λ.

    泊松分布用于模拟在固定时间或空间区间内事件发生的次数,事件独立且以恒定平均速率 λ 发生。

    X ~ Po(λ)

    The probability mass function is

    概率质量函数为

    P(X = k) = (e−λ λk) / k!

    The mean and variance are both equal to λ, a distinctive feature of the Poisson distribution.

    均值和方差均等于 λ,这是泊松分布的一个显著特征。

    E(X) = λ  Var(X) = λ

    When n is large and p is small (typically n > 20 and p < 0.1), the binomial distribution B(n, p) can be approximated by Poisson(λ = np).

    n 较大且 p 较小(一般 n > 20 且 p < 0.1)时,二项分布 B(n, p) 可用泊松分布 Poisson(λ = np) 近似。

    Similarly, for large λ (typically λ > 15), the Poisson distribution can be approximated by a normal distribution N(λ, λ).

    类似地,当 λ 较大(一般 λ > 15)时,泊松分布可用正态分布 N(λ, λ) 近似。


    4. Normal Distribution | 正态分布

    The normal distribution is the most important continuous distribution in statistics, characterized by its bell‑shaped curve.

    正态分布是统计学中最重要的连续分布,以其钟形曲线为特征。

    X ~ N(μ, σ2)

    The probability density function is not required in the formula booklet, but standardisation is essential:

    概率密度函数不需要记忆,但标准化是关键:

    Z = (X − μ) / σ ~ N(0, 1)

    Using standard normal tables, we can find probabilities such as

    使用标准正态分布表可查得以下常用概率

    • P(−1.96 < Z < 1.96) ≈ 0.95
    • P(−2.576 < Z < 2.576) ≈ 0.99
    • P(Z > 1.6449) ≈ 0.05

    For any normal variable, 68% of values lie within μ ± σ, 95% within μ ± 2σ, and 99.7% within μ ± 3σ (the empirical rule).

    对于任何正态变量,约有 68% 的值落在 μ ± σ 范围内,95% 落在 μ ± 2σ,99.7% 落在 μ ± 3σ(经验法则)。

    The sum or linear combination of independent normal variables is also normally distributed. This property underlies the central limit theorem.

    独立正态随机变量的和或线性组合仍服从正态分布,这一性质是中心极限定理的基础。


    5. Continuous Uniform Distribution | 连续均匀分布

    A continuous uniform distribution describes a variable that is equally likely to take any value within a specified interval [a, b].

    连续均匀分布描述在特定区间 [a, b] 内取值可能性均等的变量。

    X ~ U(a, b)

    The probability density function is constant over the interval:

    概率密度函数在区间内为常数:

    f(x) = 1 / (b − a)  for a ≤ x ≤ b

    The cumulative distribution function is

    累积分布函数为

    F(x) = (x − a) / (b − a)

    The mean and variance are

    均值与方差为

    E(X) = (a + b) / 2  Var(X) = (b − a)2 / 12

    This distribution is often used as a simple model for random number generation and as a testbed for theoretical results.

    该分布常被用作随机数生成的简单模型,也用于验证理论结果。


    6. Sampling Distributions & Central Limit Theorem | 抽样分布与中心极限定理

    When we take a random sample of size n from a population with mean μ and variance σ2, the sample mean is a random variable with

    从均值为 μ、方差为 σ2 的总体中抽取大小为 n 的随机样本,样本均值 是一个随机变量,满足

    E(X̄) = μ  Var(X̄) = σ2 / n

    If the population is normally distributed, then is exactly normal:

    若总体正态,则样本均值精确服从正态分布:

    X̄ ~ N(μ, σ2 / n)

    The Central Limit Theorem (CLT) states that, even when the population is not normal, the distribution of becomes approximately normal as n increases (typically n ≥ 30 is sufficient).

    中心极限定理 (CLT) 指出,即使总体不服从正态分布,当样本量 n 足够大时(通常 n ≥ 30),样本均值的分布也近似正态。

    For a sample proportion = X / n from a binomial setting, the sampling distribution satisfies

    对于来自二项分布总体的样本比例 = X / n,其抽样分布满足

    E(p̂) = p  Var(p̂) = p(1 − p) / n

    and by the CLT, is approximately normal for large n.

    根据中心极限定理,当样本量较大时 近似服从正态分布。


    7. Confidence Intervals | 置信区间

    A confidence interval provides a range of plausible values for a population parameter, based on a sample statistic.

    置信区间基于样本统计量给出总体参数的一个可能取值范围。

    When the population variance is known (or for large samples), a 100(1 − α)% confidence interval for the mean μ is

    当总体方差已知(或大样本)时,均值 μ 的 100(1 − α)% 置信区间为

    x̄ ± zα/2 × σ / √n

    If σ is unknown and the sample size is small, we use the t‑distribution with n − 1 degrees of freedom:

    若 σ 未知且样本较小,则采用自由度为 n − 1 的 t 分布:

    x̄ ± tn−1, α/2 × s / √n

    Published by TutorHao | Year 13 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Year 13 Edexcel Statistics: Experimental & Practical Assessment Essentials | 高三 Edexcel 统计:实验与实践考核要点

    📚 Year 13 Edexcel Statistics: Experimental & Practical Assessment Essentials | 高三 Edexcel 统计:实验与实践考核要点

    In Year 13 Edexcel Statistics, the ability to design robust experiments and evaluate practical investigations is a critical assessment focus. Whether you are tackling questions on clinical trials, agricultural studies, or quality control, examiners expect you to demonstrate a thorough understanding of experimental principles, data collection methods, and the mitigation of bias. This article distils the essential concepts, common pitfalls, and effective exam strategies to help you secure top marks.

    在高三 Edexcel 统计学中,设计严谨的实验并评估实践调查是关键考核点。无论是解答临床试验、农业研究还是质量控制问题,考官都期望你展示对实验原则、数据收集方法以及减少偏差的透彻理解。本文提炼了核心概念、常见陷阱和高效的考试策略,助你斩获高分。


    1. Core Principles of Experimental Design | 实验设计核心原则

    Sound experimental design rests on three pillars: randomisation, replication, and control. Randomisation ensures that treatments are assigned without bias, balancing out unknown confounding variables. Replication involves repeating the experiment on multiple independent units to estimate variability and increase precision. Control, often achieved through control groups, provides a baseline against which treatment effects are measured. In Edexcel exams, you should be able to explain why each principle is essential and recognise flaws when they are absent.

    稳健的实验设计建立在三大支柱之上:随机化、重复和对照。随机化确保处理分配无偏,平衡未知的混杂变量。重复指在多个独立单元上重复实验,以估计变异并提高精确度。对照(通常通过对照组实现)为衡量处理效应提供了基线。在 Edexcel 考试中,你应能解释每个原则的重要性,并在它们缺失时识别出缺陷。

    For example, in a crop fertiliser trial, random allocation of fertiliser types to plots prevents soil fertility gradients from skewing results. Replication with many plots allows estimation of natural variation. A control plot with no fertiliser reveals the baseline yield.

    例如,在作物肥料试验中,将肥料类型随机分配给地块可以防止土壤肥力梯度扭曲结果。通过多个地块重复可估计自然变异。不施肥的对照地块则揭示基线产量。


    2. Control Groups and Placebos | 对照组与安慰剂

    A control group is a standard of comparison that receives no treatment or a standard treatment. In medical experiments, a placebo — a

    Published by TutorHao | Year 13 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Teaching AQA Year 12 Statistics: Strategies, Lesson Plans & Insights | AQA 12年级统计教学:策略、教案与心得

    📚 Teaching AQA Year 12 Statistics: Strategies, Lesson Plans & Insights | AQA 12年级统计教学:策略、教案与心得

    Teaching AQA Year 12 Statistics requires blending GCSE probability foundations with new rigorous concepts such as the binomial and normal distributions, and the logic of hypothesis testing. This article shares practical teaching strategies, ready-to-use lesson ideas, and diagnostic approaches to help students build confidence and avoid common pitfalls. The focus is on active learning, conceptual understanding, and exam success.

    教授AQA 12年级统计需要将GCSE概率基础与新的严谨概念(如二项分布、正态分布和假设检验逻辑)相融合。本文分享实用的教学策略、可以直接使用的教案思路和诊断方法,帮助学生建立信心、规避常见错误。重点在于主动学习、概念理解和考试成功。


    1. Building on GCSE: Statistical Measures with Deeper Insight | 在GCSE基础上深化:统计度量

    Begin by assessing prior knowledge using a quick diagnostic quiz on mean, median, mode, and IQR. Use real data like daily temperatures or shoe sizes. Then introduce the concept of spread more rigorously through the standard deviation, connecting to the idea of ‘average distance from the mean’.

    通过一个关于均值、中位数、众数和四分位距的快速诊断测验来评估先备知识。使用真实数据,如每日气温或鞋码。然后通过标准差更严格地引入离散度概念,将其与“与均值的平均距离”联系起来。

    Teach the formula step by step, using a small dataset of 5 values. Have students calculate deviations, square them, sum, divide by (n-1) and take the square root. Emphasise why we use n-1 for a sample.

    逐步讲授公式,使用一个包含5个数值的小数据集。让学生计算离差、平方、求和、除以(n-1)再开方。强调为什么用n-1作为样本分母。

    s = √( Σ(x – x̄)² / (n – 1) )

    Then contrast with the population standard deviation σ = √[ Σ(x – μ)² / N ]. This distinction is crucial for later inference.

    然后对比总体标准差 σ = √( Σ(x – μ)² / N )。这一区别对后续推断至关重要。

    Have students compare the standard deviation of two datasets with the same mean but different spread, building intuitive understanding.

    让学生比较两个均值相同但离散度不同的数据集的标准差,建立直观理解。


    2. From Venn to Tree Diagrams: Mastering Conditional Probability | 从韦恩图到树状图:掌握条件概率

    Students often struggle with conditional probability. Start with a concrete two-way table, then translate into a Venn diagram. Define P(A|B) = P(A ∩ B) / P(B). Use ‘given that’ language consistently.

    学生常对条件概率感到困难。从一个具体的双向表入手,然后转换为韦恩图。定义 P(A|B) = P(A ∩ B) / P(B),并始终使用“已知……的条件下”的表述。

    Use tree diagrams for sequential events. Have students highlight the second branch paths that correspond to conditional probabilities. A common mistake is using the wrong denominator; always remind them to restrict the sample space.

    使用树状图处理序贯事件。让学生高亮对应于条件概率的第二分支路径。常见错误是使用错误的分母;始终提醒他们限制样本空间。

    Provide problems such as ‘Given that a randomly chosen student studies Maths, find the probability they also study Physics’. Model the extraction of numbers from a table.

    给出问题,如“已知随机选出的学生学习数学,求他们也学习物理的概率”。示范如何从表格中提取数字。

    Challenge students to create their own conditional probability questions from a given dataset and swap with a partner to solve. This deepens understanding.

    挑战学生从给定数据集创建自己的条件概率问题,并与同学交换解答,以加深理解。


    3. Discrete Random Variables: Expectation and Variance Formulae | 离散随机变量:期望与方差公式

    Introduce a discrete random variable as a function mapping outcomes to numbers. Use a simple example like the score on a biased die. Construct a probability distribution table.

    将离散随机变量介绍为将结果映射为数值的函数。使用一个简单的例子,如一枚不均匀骰子的得分。构建概率分布表。

    Define E(X) = Σ x·P(X=x). Show that it is a weighted mean. Calculate manually, then verify using a spreadsheet. Then introduce the shortcut formula for variance.

    定义 E(X) = Σ x·P(X=x),展示它是一个加权平均值。手动计算后用电子表格验证。然后推导方差的简化公式。

    E(X) = Σ x·P(X=x), Var(X) = Σ x² P(X=x) – [E(X)]²

    Use a table with columns: x, P(X=x), x·P, x²·P. Highlight that the sum of P(X=x) must be 1. This is a good point to discuss modelling assumptions.

    使用一个包含列 x, P(X=x), x·P, x²·P 的表格。强调 P(X=x) 之和必须为1。这是讨论建模假设的好时机。

    Give a context like a game where a player wins amounts with certain probabilities. Ask ‘Is the game fair?’ by checking E(X) = 0, or calculating expected profit for the organiser.

    给出一个情境,如一个游戏,玩家以特定概率赢得金额。通过检查 E(X) = 0 或计算组织者的期望利润来问“游戏公平吗?”


    4. Binomial Distribution: Linking Theory and Real Scenarios | 二项分布:连接理论与实际场景

    Define the conditions for a binomial distribution: fixed number of trials n, two possible outcomes, constant probability of success p, independent trials. Use a mnemonic BINS (Binary, Independent, Number fixed, Same probability).

    定义二项分布的条件:试验次数 n 固定、两种可能结果、成功概率 p 恒定、试验独立。使用助记符 BINS(二元、独立、次数固定、相同概率)。

    Derive the formula P(X = r) = ⁿCᵣ pʳ (1-p)ⁿ⁻ʳ. Show how the combination counts the number of ways. Use tree diagrams for small n to illustrate the coefficient.

    推导公式 P(X = r) = ⁿCᵣ pʳ (1-p)

    Published by TutorHao | Year 12 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Year 12 AQA Statistics: Key Terminology Quick Guide | 12 年级 AQA 统计:核心术语速记指南

    📚 Year 12 AQA Statistics: Key Terminology Quick Guide | 12 年级 AQA 统计:核心术语速记指南

    Mastering the vocabulary of Statistics is a vital first step towards exam success in AQA Year 12. This guide breaks down the key terms you’ll encounter, pairing clear definitions with memory aids to help you recall them quickly during revision and in the exam hall.

    掌握统计学的词汇是 AQA 12 年级考试成功的关键第一步。本指南详细拆解你将遇到的核心术语,为每个概念提供清晰的定义和巧妙的记忆辅助,帮助你在复习和考试中精准快速地提取知识。

    1. Types of Data | 数据类型

    In statistics, data fall into two broad categories: qualitative (categorical) and quantitative (numerical). The type of data dictates which statistical techniques are appropriate.

    在统计学中,数据分为两大类别:定性(分类)数据和定量(数值)数据。数据的类型决定了适合使用哪些统计分析方法。

    Nominal data consist of categories with no natural order. Examples include hair colour, favourite sport, or nationalities. A memorable trick: NOminal = No Order.

    名义数据由没有自然顺序的类别组成,例如发色、最喜欢的运动或国籍。记忆窍门:名义(Nominal)中的 ‘N’ 和 ‘O’ 代表 ‘No Order’(无顺序)。

    Ordinal data have a meaningful order but the differences between values are not necessarily equal. Examples: education level (GCSE, A-Level, Degree) or a satisfaction rating. Think ORdinal = ORdered.

    有序数据具有有意义的顺序,但值之间的差异不一定等距,例如教育水平(GCSE、A-水平、学位)或满意度评级。记忆:有序(Ordinal)的 ‘OR’ 开头代表 ‘ORdered’(有顺序)。

    Discrete numerical data can only take certain values, usually whole numbers. Countable items like number of students or shoe sizes (limited set) are discrete. Link ‘discrete’ to ‘distinct and countable’.

    离散数值数据只能取特定的值,通常是整数。如学生人数、鞋码(有限集合)都是离散的。将离散(discrete)联想到“distinct and countable”(可区分的、可数的)。

    Continuous data can take any value within a given range, such as height, weight or time. You can measure them to any level of precision. The word ‘continuous’ hints at a continuum of infinite possibilities.

    连续数据可以在给定范围内取任何值,例如身高、体重或时间。你可以以任意精度测量。连续(continuous)一词暗示着连续谱上无限的可能取值。


    2. Populations and Samples | 总体与样本

    A population is the entire group of individuals or objects that you are interested in studying. A sample is a subset of the population, chosen to represent it. The language used here must be precise.

    总体是你感兴趣研究的全部个体或对象的集合。样本是从总体中选出的一个子集,用于代表总体。这里的术语必须使用精确。

    A census attempts to collect data from every member of the population. While it gives the most accurate picture, it is usually time-consuming and expensive.

    普查试图收集总体中每一个成员的数据。虽然能给出最准确的图景,但通常耗时且昂贵。

    The sample frame is a list of all members of the population from which the sample is drawn. If the frame is not complete, the sample may be biased.

    抽样框是总体中所有成员的名单,样本从此名单中抽取。如果抽样框不完整,样本可能产生偏差。

    A sampling unit is each individual member of the population. The sampling fraction is the sample size n divided by the population size N.

    总体中的每一个个体都是一个抽样单位抽样比是样本容量 n 除以总体容量 N。

    A parameter is a numerical characteristic of a population (e.g. population mean μ). A statistic is a numerical characteristic of a sample (e.g. sample mean x̄). We use statistics to estimate parameters. Remember: Parameter = Population, Statistic = Sample.

    参数是总体的数值特征(如总体均值 μ)。统计量是样本的数值特征(如样本均值 x̄)。我们用统计量来估计参数。记忆:Parameter 对应 Population,Statistic 对应 Sample(均以 ‘P’ 和 ‘S’ 开头)。


    3. Measures of Central Tendency | 集中趋势的度量

    These statistics describe the centre or typical value of a data set. The three main measures are the mean, median and mode.

    这些统计量用来描述数据集中心或典型值。三种主要度量是均值、中

    Published by TutorHao | Year 12 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • AQA Statistics for International Competitions: A Strategic Guide | AQA 统计学国际竞赛备战攻略

    📚 AQA Statistics for International Competitions: A Strategic Guide | AQA 统计学国际竞赛备战攻略

    In Year 12 AQA Statistics, students build foundational skills in data handling, probability, and statistical inference. These concepts are not only essential for A-level exams but also form the core of many international mathematics and statistics competitions. This guide will show you how to leverage your classroom knowledge to excel in contests such as the UKMT Senior Mathematical Challenge, national olympiads, or even the International Statistics Olympiad, with practical strategies and expert insights.

    在Year 12 AQA统计学课程中,学生们建立起数据处理、概率和统计推断的基础能力。这些概念不仅是A-level考试的关键,也构成了许多国际数学与统计竞赛的核心。本攻略将向你展示如何运用课堂所学知识,在英国数学信托基金会高级数学挑战赛(UKMT SMC)、国家级奥林匹克乃至国际统计奥林匹克等竞赛中脱颖而出,并提供实用策略和专业洞见。


    1. Why Statistics Matters in Competitions | 统计学在竞赛中的重要性

    Many students underestimate the role of statistics in mathematical contests, assuming that pure algebra or geometry dominate. However, data interpretation, probability puzzles, and logical inference questions frequently appear and can be decisive. For instance, the UKMT SMC often includes problems on averages, counting probabilities, and expected values. A solid grasp of Year 12 AQA statistics gives you a competitive edge.

    许多学生低估了统计学在数学竞赛中的作用,认为纯粹代数和几何占主导地位。然而,数据解读、概率谜题和逻辑推断问题频繁出现,并且可能成为决胜关键。例如,UKMT SMC常包含关于平均数、计数概率和期望值的问题。扎实掌握Year 12 AQA统计学能给你带来竞争优势。

    Moreover, the reasoning skills developed through hypothesis testing and critical evaluation of data help you tackle unfamiliar contest problems systematically. Understanding concepts like sampling bias and variability allows you to spot flawed arguments or misleading graphs quickly.

    此外,通过假设检验和数据批判性评价培养的推理能力,能帮助你系统地解决竞赛中不熟悉的问题。理解抽样偏差和变异性等概念,让你能迅速发现有缺陷的论证或误导性图表。


    2. Mastering Data Representation | 掌握数据表示

    AQA Statistics Year 12 covers histograms, cumulative frequency diagrams, and box plots. In competitions, you may be asked to interpret complex charts or extract key information rapidly. Practice sketching quick histograms with unequal class widths and calculating frequency density = frequency ÷ class width. Use the formula: frequency density = frequency / class width.

    AQA统计学Year 12涵盖直方图、累积频率图和箱形图。在竞赛中,你可能需要解读复杂图表或快速提取关键信息。练习快速绘制不等组距的直方图,并计算频率密度 = 频数 ÷ 组距。使用公式:频率密度 = 频数 / 组距。

    Cumulative frequency graphs are powerful for estimating medians, quartiles, and percentiles without raw data. Contest problems might give you a cumulative frequency table and ask for the interquartile range. Remember: Q₁ is at 25%, Q₂ at 50%, Q₃ at 75% of total frequency. Using linear interpolation between points can yield accurate estimates.

    累积频率图对于在没有原始数据的情况下估计中位数、四分位数和百分位数非常有用。竞赛题目可能会给出累积频率表,要求计算四分位距。记住:Q₁在总频率的25%处,Q₂在50%处,Q₃在75%处。在点之间使用线性插值可以得到准确的估计。

    Box plots provide a concise view of spread and skew. If given multiple box plots in a contest, compare medians and interquartile ranges to assess central tendency and variability. Outliers, defined as points beyond 1.5 × IQR from the quartiles, can reveal interesting anomalies

    Published by TutorHao | Year 12 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • Year 12 AQA Statistics: Quick Reference Formula & Theorem Handbook | AQA 统计:公式定理速查手册

    📚 Year 12 AQA Statistics: Quick Reference Formula & Theorem Handbook | AQA 统计:公式定理速查手册

    This quick reference handbook collects all the essential formulas, theorems, and definitions needed for the Year 12 AQA Statistics course. It is designed to help you efficiently review probability rules, common distributions, hypothesis testing procedures, and correlation/regression techniques before your exams. Each section presents the core mathematical expressions with paired English and Chinese explanations.

    本速查手册汇集了 Year 12 AQA 统计课程所需的所有关键公式、定理和定义。它旨在帮助你在考试前高效地复习概率规则、常见分布、假设检验程序以及相关与回归技术。每个小节都以英中对照的方式呈现核心数学表达式。


    1. Measures of Central Tendency & Spread | 集中趋势与离散度量

    The sample mean is the arithmetic average of the data, providing a measure of central location. The sample variance quantifies the average squared deviation from the mean, using n-1 to give an unbiased estimate of the population variance.

    样本均值是数据的算术平均值,用来度量数据的中心位置。样本方差量化了各数据点与均值之差的平方的平均值,使用 n-1 是为了得到总体方差的无偏估计。

    x̄ = (Σxᵢ) / n

    x̄ = (Σxᵢ) / n

    s² = Σ(xᵢ – x̄)² / (n – 1)

    s² = Σ(xᵢ – x̄)² / (n – 1)

    The standard deviation is the positive square root of the variance, s = √s². Other key measures include the median (middle value when ordered) and the interquartile range IQR = Q₃ – Q₁.

    标准差是方差的正平方根,s = √s²。其他重要度量包括中位数(排序后中间的值)以及四分位距 IQR = Q₃ – Q₁。

    Statistic Formula 中文
    Mean x̄ = Σxᵢ / n 均值
    Variance (sample) s² = Σ(xᵢ – x̄)² / (n-1) 样本方差
    Standard deviation s = √[Σ(xᵢ – x̄)²/(n-1)] 标准差
    Range max – min 极差
    IQR Q₃ – Q₁ 四分位距

    2. Basic Probability Rules | 概率基本规则

    For any two events A and B, the probability that at least one of them occurs is given by the general addition rule. When events are mutually exclusive, the intersection term becomes zero.

    对于任何两个事件 A 和 B,至少有一个发生的概率由一般加法规则给出。当事件互斥时,交集项为零。

    P(A ∪ B) = P(A) + P(B) – P(A ∩ B)

    P(A ∪ B) = P(A) + P(B) – P(A ∩ B)

    For independent events, the joint probability is simply the product of the individual probabilities. The complement rule states that the probability of an event not occurring is 1 minus the probability of the event.

    对于独立事件,联合概率就是个别的概率之积。互补规则指出,一个事件不发生的概率等于 1 减去该事件发生的概率。

    P(A ∩ B) = P(A) × P(B) (if independent / 如果独立)

    P(A ∩ B) = P(A) × P(B) (如果独立)

    P(A’) = 1 – P(A)

    P(A’) = 1 – P(A)


    3. Conditional Probability & Independence | 条件概率与独立性

    The conditional probability of B given A is the probability that B occurs under the condition that A has already occurred. It is defined only when P(A) > 0. This formula forms the basis of Bayes’ Theorem when dealing with reversed conditions.

    在 A 已发生的条件下 B 发生的条件概率记为 P(B|A)。它仅在 P(A) > 0 时有定义。该公式是处理条件反转问题时贝叶斯定理的基础。

    P(B|A) = P(A ∩ B) / P(A)

    P(B|A) = P(A ∩ B) / P(A)

    Two events A and B are independent if and only if P(A ∩ B) = P(A) P(B) or equivalently P(B|A) = P(B). A tree diagram is a powerful tool for visualising sequences of conditional probabilities.

    两个事件 A 和 B 独立的充要条件是 P(A ∩ B) = P(A) P(B) 或等价地 P(B|A) = P(B)。树图是可视化一系列条件概率的强大工具。

    P(A|B) = [P(B|A) P(A)] / P(B) (Bayes’ Theorem / 贝叶斯定理)

    P(A|B) = [P(B|A) P(A)] / P(B) (贝叶斯定理)


    4. Discrete Random Variables | 离散随机变量

    A discrete random variable X takes a countable set of values with associated probabilities. The expected value E(X) gives the long-run average, while Var(X) measures the spread of the distribution.

    离散随机变量 X 取可数个可能值,每个值都有对应的概率。期望值 E(X) 给出了长期的平均值,而方差 Var(X) 衡量分布的离散程度。

    E(X) = Σ x P(X = x)

    E(X) = Σ x P(X = x)

    Var(X) = E(X²) – [E(X)]² = Σ x² P(X = x) – μ²

    Var(X) = E(X²) – [E(X)]² = Σ x² P(X = x) – μ²

    For a linear transformation of a random variable, the expectation and variance follow simple rules. Adding a constant shifts the mean but does not affect the variance; multiplying by a constant scales both the mean and the variance (squared).

    对于随机变量的线性变换,期望和方差遵循简单的规则。加一个常数平移均值但不影响方差;乘一个常数同时缩放均值和方差(方差要乘以常数的平方)。

    E(aX + b) = a E(X) + b

    E(aX + b) = a E(X) + b

    Var(aX + b) = a² Var(X)

    Var(aX + b) = a² Var(X)


    5. Binomial Distribution | 二项分布

    The binomial distribution models the number of successes in n independent trials, each with constant probability of success p. The distribution is denoted by X ~ B(n, p). The probability of exactly r successes is given by the binomial formula.

    二项分布用于描述在 n 次独立试验中成功的次数,每次成功的概率恒为 p。该分布记作 X ~ B(n, p)。恰好得到 r 次成功的概率由二项式公式给出。

    P(X = r) = ⁿCᵣ pʳ (1 – p)ⁿ⁻ʳ, r = 0,1,…,n

    P(X = r) = ⁿCᵣ pʳ (1 – p)ⁿ⁻ʳ, r = 0,1,…,n

    The mean and variance of a binomial random variable are simple multiples of n and p. When n is large and p is close to 0.5, the binomial distribution is approximately symmetric; it can also be approximated by a normal distribution under certain conditions.

    二项随机变量的均值和方差都是 n 与 p 的简单乘积。当 n 很大且 p 接近 0.5 时,二项分布近似对称;在某些条件下还可用正态分布近似。

    E(X) = np

    E(X) = np

    Var(X) = np(1 – p)

    Var(X) = np(1 – p)


    6. Normal Distribution | 正态分布

    The normal distribution is a continuous probability distribution defined by two parameters: the mean μ and the standard deviation σ. Its probability density function

    Published by TutorHao | Year 12 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)

  • AQA Year 12 Statistics: In-depth Analysis of Past Papers | AQA 12年级统计:历年真题深度解析

    📚 AQA Year 12 Statistics: In-depth Analysis of Past Papers | AQA 12年级统计:历年真题深度解析

    Success in AQA Year 12 Statistics requires more than just memorising formulas; it demands a deep understanding of key concepts and the ability to apply them to exam-style questions. This article provides an in-depth analysis of past papers, highlighting common question types, essential techniques, and strategies to help you secure top marks.

    在AQA 12年级统计考试中取得高分,仅仅记住公式是不够的;你需要深入理解核心概念,并能够将其应用到真题题型中。本文通过对历年真题的深度解析,着重分析常见题型、关键解题技巧以及帮助你获得高分的策略。

    1. Understanding Data: Measures of Central Tendency | 理解数据:集中趋势的度量

    AQA frequently tests the calculation and interpretation of mean, median, and mode from raw data, frequency tables, and grouped data. In a typical past paper question (e.g., June 2019 Q2), you are given a frequency table and asked to estimate the mean using midpoints. The mean of grouped data is estimated as x̄ = Σ(f × midpoint) / Σf. Remember that the median for grouped data requires linear interpolation: Median = L + [(n/2 – F) / f] × w, where L is the lower boundary of the median class, F is the cumulative frequency before the median class, f is the frequency of the median class, and w is the class width.

    AQA经常考查从原始数据、频数表和分组数据中计算和解释平均数、中位数和众数。在典型的历年真题中(例如2019年6月第2题),会给出一个频数表,要求使用组中值估算平均数。分组数据的平均数估算为 x̄ = Σ(f × 组中值) / Σf。须注意分组数据的中位数需要使用线性插值:中位数 = L + [(n/2 – F) / f] × w,其中L为中位数所在组的下限,F为小于该组的累积频数,f为所在组的频数,w为组距。

    A common mistake is to forget that the median and quartiles for discrete data may need careful positioning. When data are listed, the median is the (n+1)/2 th value. In past papers, examiners often award marks for stating the correct position before identifying the value. Always show your working clearly.

    一个常见错误是忘记离散数据的中位数和四分位数需要谨慎定位。当数据列出时,中位数是第 (n+1)/2 个值。在历年真题中,考官经常会给确定位置这一步骤打分,然后再识别数值。务必清晰地展示计算过程。


    2. Measures of Spread: Variance and Standard Deviation | 离散程度的度量:方差与标准差

    Variance and standard deviation are frequent topics. AQA expects you to use the formula Var(X) = Σ(x – μ)²P(X=x) for discrete random variables and the computational formula s² = Σx²/n – (Σx/n)² for sample data. In past papers (e.g., 2020 Q4), candidates are given summary statistics Σx and Σx² and asked to calculate the standard deviation. Standard deviation is the square root of variance: σ = √Var. Remember to distinguish between population variance (dividing by n) and sample variance (dividing by n-1). AQA usually specifies which one to use, but when dealing with a sample, use s².

    方差和标准差是高频考点。AQA要求你掌握离散随机变量的方差公式 Var(X) = Σ(x – μ)²P(X=x) 以及样本数据的计算式 s² = Σx²/n – (Σx/n)²。在历年真题中(如2020年第4题),考生会得到汇总统计量 Σx 和 Σx²,然后要求计算标准差。标准差是方差的平方根:σ = √Var。务必区分总体方差(除以n)和样本方差(除以n-1)。AQA通常会指明使用哪一种,但当处理样本时,应使用s²。

    Interpreting standard deviation is just as important as calculating it. Some questions ask you to compare the dispersion of two datasets using mean and standard deviation. When the means differ, you may also need to calculate the coefficient of variation (CV = σ/μ × 100%) to compare relative variability, though AQA may not explicitly require it, it is good practice.

    解释标准差与计算同样重要。有些题目要求你使用平均数和标准差比较两个数据集的离散程度。当平均数不同时,你还需要计算变异系数(CV = σ/μ × 100%)来比较相对波动性,尽管AQA不一定明确要求,但这是一种良好做法。


    3. Probability Basics: Tree Diagrams and Venn Diagrams | 概率基础:树状图与维恩图

    AQA probability questions often involve tree diagrams, especially for conditional probability. A standard past paper question (e.g., 2018 Q5) presents two bags with coloured counters and asks you to draw a tree diagram, label probabilities, and find the probability of specific outcomes. Always multiply along branches for ‘and’ and add between branches for ‘or’. For conditional probability, use P(A|B) = P(A ∩ B) / P(B).

    AQA的概率题常常涉及树状图,尤其是条件概率。标准的历年真题(如2018年第5题)会给出两个装有彩色计数器的袋子,要求画出树状图、标注概率并求特定结果的概率。始终沿着分支相乘得到“且”的概率,在分支之间相加得到“或”的概率。对于条件概率,使用 P(A|B) = P(A ∩ B) / P(B)。

    Venn diagrams are commonly used to illustrate events and their intersections. You may be given a Venn diagram with probabilities and asked to test for independence using P(A) × P(B) = P(A ∩ B) or prove that events are mutually exclusive if P(A ∩ B) = 0. AQA past papers also test complement rule P(A’) = 1 – P(A). Ensure you are comfortable shading regions like A ∪ B and A ∩ B’.

    维恩图常用于说明事件及其交集。题目可能给出带有概率的维恩图,要求使用 P(A) × P(B) = P(A ∩ B) 检验独立性,或者当 P(A ∩ B) = 0 时证明事件互斥。AQA历年真题也考查补集规则 P(A’) = 1 – P(A)。确保你能熟练地用阴影表示A ∪ B 和 A ∩ B’ 等区域。


    4. Discrete Random Variables: Expected Value and Variance | 离散随机变量:期望与方差

    A discrete random variable X takes a set of values with corresponding probabilities. AQA typically provides a probability distribution table and asks for E(X) = Σ x·P(X=x) and Var(X) = E(X²) – [E(X)]². A common exam question (e.g., 2021 Q3) gives a table with one unknown probability k, requiring you to first use Σ P(X=x) = 1 to find k, then calculate E(X) and Var(X).

    离散随机变量X取一组值,并有相应的概率。AQA通常会给出一个概率分布表,要求计算 E(X) = Σ x·P(X=x) 以及 Var(X) = E(X²) – [E(X)]²。一道常见的考题(如2021年第3题)会给出一个含有一个未知概率k的表格,需要先利用 Σ P(X=x) = 1 求出k,然后计算E(X)和Var(X)。

    After finding E(X) and Var(X), you may need to answer questions about the expected profit in a game or the standard deviation. For linear transformations, remember that E(aX + b) = aE(X) + b and Var(aX + b) = a²Var(X). These transformation rules appear regularly in past papers, often as part of a larger modelling question.

    在求出E(X)和Var(X)之后,可能需要回答关于游戏中的期望利润或标准差的问题。对于线性变换,记住 E(aX + b) = aE(X) + b 以及 Var(aX + b) = a²Var(X)。这些变换规则在历年真题中频繁出现,通常是某个更大的建模题目的一部分。


    5. The Binomial Distribution: Conditions and Calculations | 二项分布:条件与计算

    The binomial distribution X ~ B(n, p) requires four conditions: a fixed number of trials n, two possible outcomes (success/failure), constant probability p, and independence of trials. AQA past papers often begin by asking you to explain why a situation can be modelled by a binomial distribution. You must mention these conditions explicitly to earn method marks.

    二项分布 X ~ B(n, p) 需要满足四个条件:固定试验次数n、两种可能结果(成功/失败)、概率p恒定以及各次试验相互独立。AQA历年真题常常先要求你解释为什么某个情形可以用二项分布来建模。你必须明确提及这些条件,才能获得方法分。

    To find probabilities, use P(X = r) = ⁿCᵣ pʳ (1-p)ⁿ⁻ʳ. However, in many questions, you need to compute cumulative probabilities like P(X ≤ r) or P(X ≥ r). AQA provides binomial cumulative distribution tables, but you must know how to use them. For example, P(X ≥ 5) = 1 – P(X ≤ 4). When p > 0.5, the table may require you to use the complementary event. Always check if your calculator can provide exact values; AQA accepts calculator use, but you must show the working.

    要计算概率,使用 P(X = r) = ⁿCᵣ pʳ (1-p)ⁿ⁻ʳ。然而,在很多问题中,你需要计算累积概率,如 P(X ≤ r) 或 P(X ≥ r)。AQA会提供二项分布累积概率表,但你必须知道如何使用。例如,P(X ≥ 5) = 1 –

    Published by TutorHao | Year 12 统计 Revision Series | aleveler.com

    更多咨询请联系16621398022(同微信)