Year 9 AQA Statistics: Formula & Theorem Quick Reference Guide | Year 9 AQA 统计:公式定理速查手册

📚 Year 9 AQA Statistics: Formula & Theorem Quick Reference Guide | Year 9 AQA 统计:公式定理速查手册

This quick reference handbook gathers the essential formulas and theorems you will encounter in Year 9 AQA Statistics. From measures of central tendency to probability rules and data representation, each entry is presented as a clear bilingual pair so you can revise confidently in both English and Chinese.

本速查手册汇总了 Year 9 AQA 统计课程中的核心公式与定理。从集中趋势度量到概率规则与数据展示,每一条知识点都以清晰的中英文对照呈现,帮助你自信地使用双语进行复习。

1. Measures of Central Tendency | 集中趋势度量

The mean is the sum of all data values divided by the number of values. It is the arithmetic average and is affected by every data point.

平均数是所有数据值的总和除以数据的个数。它是算术平均值,受每一个数据点的影响。

Mean = Σx / n

The median is the middle value when data are arranged in order. For an odd number of data points, it is the central one; for an even number, it is the mean of the two central values.

中位数是将数据按顺序排列后处于中间位置的值。数据个数为奇数时取正中间的那个数;为偶数时取中间两个数的平均数。

The mode is the value that appears most frequently. A data set can have one mode (unimodal), two modes (bimodal) or more (multimodal).

众数是出现频率最高的数值。数据集可以有一个众数(单峰)、两个众数(双峰)或更多(多峰)。


2. Range and Interquartile Range | 极差与四分位距

The range is the difference between the largest and smallest values. It gives a simple measure of spread but is sensitive to extreme values.

极差是最大值与最小值之间的差值。它给出了一个简单的离散程度度量,但对极端值非常敏感。

Range = Maximum − Minimum

The interquartile range (IQR) is the difference between the upper quartile (Q₃) and the lower quartile (Q₁). It describes the spread of the middle 50% of the data and is resistant to outliers.

四分位距 (IQR) 是上四分位数 (Q₃) 与下四分位数 (Q₁) 的差。它描述了中间 50% 数据的离散程度,不易受异常值影响。

IQR = Q₃ − Q₁

To find quartiles: Q₂ is the median; Q₁ is the median of the lower half; Q₃ is the median of the upper half (excluding the overall median if odd count).

寻找四分位数的方法:Q₂ 即中位数;Q₁ 是下半部分数据的中位数;Q₃ 是上半部分数据的中位数(若总数为奇数则不包括总中位数)。


3. Frequency Tables and Estimated Mean | 频数表与估算平均数

For grouped data, we estimate the mean using the midpoint of each class interval. Multiply each midpoint by its frequency, sum the products and divide by the total frequency.

对于分组数据,我们使用每个组距的组中值来估算平均数。将每个组中值乘以对应的频数,求和后除以总频数。

Estimated Mean = Σ(f × x) / Σf

Here, f is the frequency and x is the midpoint of the interval. The modal class is the interval with the highest frequency; the median class is found by cumulative frequency.

此处 f 为频数,x 为区间的组中值。众数组是频数最高的区间;中位数组通过累计频数来寻找。

The midpoint is calculated as (lower bound + upper bound) ÷ 2. Be careful with boundaries where data are continuous.

组中值计算公式为 (下限 + 上限) ÷ 2。处理连续数据时,需注意边界的确定。


4. Cumulative Frequency and Box Plots | 累积频数与箱线图

Cumulative frequency is the running total of frequencies. Plotting cumulative frequency against the upper class boundary gives an S-shaped curve useful for estimating medians and quartiles.

累积频数是频数的累计总和。将累积频数对组上限描点,可得到 S 形曲线,用于估算中位数和四分位数。

From a cumulative frequency graph, locate the position (n/2 for median, n/4 for Q₁, 3n/4 for Q₃) on the cumulative frequency axis, then draw a horizontal line to the curve and down to the data axis.

在累积频数图中,在累积频数轴上找到对应位置(中位数为 n/2,Q₁ 为 n/4,Q₃ 为 3n/4),作水平线交曲线,再垂直向下读取数据值。

A box plot (or box-and-whisker diagram) displays the minimum, Q₁, median, Q₃ and maximum. It visually summarises the spread and symmetry of the data.

箱线图(箱须图)展示最小值、Q₁、中位数、Q₃ 和最大值。它直观地概括了数据的离散程度与对称性。

The box represents the IQR; the line inside marks the median. Whiskers extend to the minimum and maximum unless outliers are defined separately.

箱体表示 IQR,箱内线标记中位数。须线延伸至最小值和最大值,除非单独标出异常值。


5. Probability Basics | 概率基础

Probability measures how likely an event is to happen, ranging from 0 (impossible) to 1 (certain). It can be expressed as a fraction, decimal or percentage.

概率衡量事件发生的可能性,范围从 0(不可能)到 1(必然)。可以用分数、小数或百分数表示。

Probability of an event A = Number of favourable outcomes / Total number of equally likely outcomes

The sum of probabilities of all possible outcomes is 1. If the probability of an event is p, the probability it does not happen is 1 − p.

所有可能结果的概率之和为 1。若某事件概率为 p,则不发生的概率为 1 − p。

Events can be placed on a probability scale to compare likelihoods. A probability close to 0 means the event is unlikely; close to 1 means it is very likely.

可以将事件置于概率标尺上比较可能性。概率接近 0 表示不太可能发生;接近 1 表示极可能发生。


6. Mutually Exclusive and Independent Events | 互斥事件与独立事件

Mutually exclusive events cannot occur at the same time. For two such events A and B, P(A or B) = P(A) + P(B). This is the addition rule for mutually exclusive events.

互斥事件不能同时发生。对于两个互斥事件 A 和 B,P(A 或 B) = P(A) + P(B)。这是互斥事件的加法法则。

Independent events are those where the occurrence of one does not affect the probability of the other. For independent events A and B, P(A and B) = P(A) × P(B).

独立事件是指一个事件的发生不影响另一个事件发生的概率。对于独立事件 A 和 B,P(A 且 B) = P(A) × P(B)。

Always check whether events are mutually exclusive or independent before applying formulas. Use Venn diagrams or tree diagrams to visualise relationships.

应用公式前,务必检查事件是互斥还是独立。可以使用维恩图或树形图来可视化事件关系。


7. Scatter Graphs and Correlation | 散点图与相关

A scatter graph displays the relationship between two variables. Each point represents a pair of values (x, y). The pattern of points suggests the type of correlation.

散点图展示两个变量之间的关系。每一个点代表一对数值 (x, y)。点的分布特征暗示相关的类型。

Positive correlation means as one variable increases, the other tends to increase. Negative correlation means as one increases, the other tends to decrease. No correlation means no clear pattern.

正相关表示一个变量增加时,另一个也趋于增加。负相关表示一个增加时,另一个趋于减少。无相关则没有明显模式。

Correlation does not imply causation. Even strong correlation may be due to a third factor or coincidence.

相关并不意味着因果。即便是强相关,也可能是由第三个因素或巧合造成的。


8. Line of Best Fit and Interpolation | 最佳拟合线与插值

A line of best fit (or trend line) is drawn on a scatter graph to model the relationship. It should pass through the mean point (x̄, ȳ) and have roughly equal numbers of points above and below it.

最佳拟合线(或趋势线)画在散点图上以模拟关系。它应通过均值点 (x̄, ȳ),并且线上下的点数大致相等。

Interpolation is estimating a value within the range of the data using the line of best fit. Extrapolation is predicting beyond the data range and is less reliable.

插值是利用最佳拟合线在数据范围内估计数值。外推是预测数据范围之外的值,可信度较低。

Use the equation of the line y = mx + c if given, where m is the gradient and c is the y-intercept. The gradient shows the rate of change between the variables.

如果已知直线方程 y = mx + c,可以直接使用,其中 m 是斜率,c 是 y 轴截距。斜率显示了变量之间的变化率。


9. Sampling Methods | 抽样方法

Random sampling gives every member of the population an equal chance of being selected. It helps to avoid bias but may be impractical for large populations.

随机抽样让总体中的每个成员都有相等的机会被选中。它有助于避免偏差,但对于大总体可能不切实际。

Systematic sampling selects members at regular intervals from a list. For example, choosing every 10th name. It is simple but can introduce bias if there is a hidden pattern.

系统抽样是按固定间隔从名单中选取样本。例如,每 10 个人选一个。此方法简单,但若存在隐藏模式可能引入偏差。

Stratified sampling divides the population into groups (strata) and randomly selects from each in proportion to its size. It ensures representation of all subgroups.

分层抽样将总体分成若干层,然后按各层占比随机抽取样本。这能确保所有子群体都被代表。

Opportunity (convenience) sampling uses people who are easiest to reach. It is quick but often biased and not representative of the whole population.

机会(便利)抽样选择最容易接触到的人。此方法快捷,但常有偏差,不能代表整个总体。


10. Data Collection and Questionnaires | 数据收集与问卷设计

Primary data is collected first-hand for a specific purpose (e.g., experiments, surveys). Secondary data is data already gathered by others (e.g., government statistics, internet databases).

一手数据是为特定目的直接收集的(如实验、调查)。二手数据是他人已经收集的数据(如政府统计、网络数据库)。

Questionnaires must use clear, unbiased language. Avoid leading questions, double-barrelled questions or overlapping response options.

问卷必须使用清晰、无偏向的语言。避免引导性问题、双重问题或重叠的答案选项。

Response boxes should be exhaustive and mutually exclusive. Pilot testing a questionnaire helps identify misunderstandings before full distribution.

答案选项应当穷尽且互斥。在大规模发放前进行问卷预测试有助于发现理解上的问题。

Data can be categorical (qualitative) or numerical (quantitative). Numerical data may be discrete (countable) or continuous (measurable). Knowing the data type guides the choice of graph and statistical measure.

数据可以是分类(定性)数据或数值(定量)数据。数值数据又可分为离散(可数)或连续(可测)数据。了解数据类型有助于选择合适的图表和统计度量。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading