Year 10 CIE Statistics: Formula & Theorem Quick Reference | CIE 10 年级统计公式定理速查手册

📚 Year 10 CIE Statistics: Formula & Theorem Quick Reference | CIE 10 年级统计公式定理速查手册

This quick reference handbook brings together the essential formulas, theorems and key concepts required for Year 10 CIE Statistics. It covers measures of central tendency, dispersion, probability, data representation and bivariate analysis. Use it to review definitions, memorise equations and build confidence before exams.

本速查手册汇总了 CIE 十年级统计课程最核心的公式、定理与关键概念,涵盖集中趋势、离散度、概率、数据展示及二元分析等内容。可用于复习定义、记忆公式,帮助你在考试前自信备战。


1. Measures of Central Tendency | 集中趋势的度量

The mean (average) is calculated by summing all data values and dividing by the number of values.

平均数(均值)等于所有数据值之和除以数据个数。

Mean: x̄ = Σx / n

The median is the middle value when the data are arranged in order. For an even number of values, it is the average of the two middle numbers.

中位数是将数据排序后位于中间位置的数值。若数据个数为偶数,则取中间两位数的平均值。

The mode is the value that occurs most frequently in a data set.

众数是一组数据中出现次数最多的数值。

For grouped data, the modal class is the class interval with the highest frequency.

对于分组数据,众数所在组是频率最高的组距。


2. Measures of Dispersion | 离散程度的度量

The range is the difference between the largest and smallest values.

全距是最大值与最小值的差。

Range = xmax – xmin

Quartiles divide ordered data into four equal parts. Q1 is the lower quartile (25th percentile), Q2 is the median (50th percentile) and Q3 is the upper quartile (75th percentile).

四分位数将有序数据等分为四份。下四分位数 Q1(第 25 百分位数),中位数 Q2(第 50 百分位数),上四分位数 Q3(第 75 百分位数)。

The interquartile range (IQR) measures the spread of the middle 50% of the data.

四分位距(IQR)衡量中间 50% 数据的分散程度。

IQR = Q3 – Q1

Population variance and standard deviation describe how data points spread around the mean.

总体方差和标准差描述数据点围绕均值的散布程度。

σ² = Σ(x – μ)² / N    and    σ = √(Σ(x – μ)² / N)

For a sample, the variance uses n-1 to give an unbiased estimate.

对于样本,方差除以 n-1 以得到无偏估计。

s² = Σ(x – x̄)² / (n – 1)

An alternative computational formula is often faster.

另一种计算式常可加快运算。

σ² = Σx² / n – μ²


3. Frequency Distributions | 频率分布

For a frequency table with values x and frequencies f, the mean is given by the weighted sum.

对于含取值 x 与频率 f 的频数表,均值为加权和。

x̄ = Σ(fx) / Σf

The cumulative frequency of a class is the sum of frequencies for that class and all previous classes.

一组的累积频率为该组及之前所有组的频率之和。

When estimating the median from grouped data, use the cumulative frequency graph or linear interpolation.

利用分组数据估计中位数时,可使用累积频率图或线性插值法。


4. Probability Basics | 概率基础

Probability measures the chance that an event occurs, always lying between 0 and 1.

概率衡量事件发生的可能性,其值总是在 0 到 1 之间。

P(A) = Number of favourable outcomes / Total number of equally likely outcomes

The sum of probabilities of all possible outcomes of an experiment is 1.

试验中所有可能结果的概率之和为 1。

P(S) = 1, where S is the sample space

The complement of event A, denoted A’, is the event that A does not happen.

事件 A 的补事件记为 A’,表示 A 不发生的事件。

P(A’) = 1 – P(A)


5. Mutually Exclusive & Independent Events | 互斥事件与独立事件

Two events are mutually exclusive if they cannot occur at the same time.

若两事件不可能同时发生,则它们是互斥事件。

P(A or B) = P(A) + P(B)    (for mutually exclusive events)

If events are not mutually exclusive, we subtract the overlap to avoid double-counting.

若事件并非互斥,则需减去重叠部分以避免重复计算。

P(A or B) = P(A) + P(B) – P(A and B)

Two events are independent if the outcome of one does not affect the probability of the other.

若一个事件的发生不影响另一事件的概率,则两事件独立。

P(A and B) = P(A) × P(B)    (for independent events)


6. Tree Diagrams & Conditional Probability | 树状图与条件概率

A tree diagram shows all possible outcomes of a sequence of events, with probabilities written along the branches.

树状图展示事件序列的所有可能结果,并在分支上标注概率。

The probability of a path is the product of the probabilities along that path.

一条路径的概率等于该路径上各分支概率的乘积。

Conditional probability is the probability of event A given that event B has already occurred.

条件概率指在事件 B 已发生的前提下事件 A 发生的概率。

P(A|B) = P(A and B) / P(B)

Rearranging gives the general multiplication rule.

由此可变形得到一般乘法法则。

P(A and B) = P(A|B) × P(B)


7. Cumulative Frequency & Percentiles | 累积频率与百分位数

A cumulative frequency curve plots the running total of frequencies against the upper class boundary.

累积频率图将各组的频率累计总数对该组上界作图。

The median and quartiles can be read directly from the cumulative frequency curve.

中位数和四分位数可直接从累积频率曲线上读出。

To find a percentile, locate the corresponding cumulative frequency value and draw a horizontal line across to the curve, then drop vertically to the axis.

要找某百分位数,先定位对应的累积频率值,画水平线交曲线,再垂直投射到坐标轴上。

The interpercentile range, such as the 10th to 90th percentile range, describes the spread of the central 80% of data.

百分位数间距(如第 10 至第 90 百分位数距离)描述中间 80% 数据的分布宽度。


8. Box-and-Whisker Plots | 箱线图

A box-and-whisker plot (box plot) visually displays the minimum, Q1, median, Q3 and maximum.

箱线图(盒须图)直观显示最小值、Q1、中位数、Q3 和最大值。

The box represents the interquartile range, and the line inside the box marks the median.

箱体代表四分位距,箱内横线标示中位数。

Whiskers extend to the smallest and largest values that are not outliers, often defined as values within 1.5 × IQR from the quartiles.

须线延伸至不是异常值的最大和最小值,异常值通常定义为距四分位数超过 1.5 × IQR 的数据点。

Outliers are plotted as individual points beyond the whiskers.

异常值作为独立点绘制在须线之外。

Box plots are useful for comparing spreads and central tendencies between different data sets.

箱线图适合比较不同数据集之间的离散程度和集中趋势。


9. Histograms & Frequency Density | 直方图与频率密度

In a histogram, the area of each bar is proportional to the frequency of that class.

在直方图中,每根柱子的面积与该组的频率成正比。

When class widths are unequal, we use frequency density instead of frequency on the vertical axis.

当组距不等时,纵轴应使用频率密度而非频率。

Frequency density = Frequency / Class width

To find the frequency from a histogram, multiply the frequency density by the class width.

若要从直方图反推频率,将频率密度乘以组距即可。

The total area of all bars equals the total frequency.

所有矩形的总面积等于总频率。


10. Scatter Diagrams & Correlation | 散点图与相关性

A scatter diagram plots paired data (x, y) to show the relationship between two variables.

散点图将成对数据 (x, y) 绘制出来,以展示两个变量之间的关系。

Positive correlation means that as x increases, y tends to increase. Negative correlation means that as x increases, y tends to decrease.

正相关表示 x 增大时 y 也倾向增大;负相关表示 x 增大时 y 倾向减小。

No correlation indicates the points show no clear upward or downward trend.

零相关表示点的分布无明显上升或下降趋势。

The strength of correlation is judged by how closely the points cluster around a straight line.

相关性强度由点围绕一条直线的紧密程度来判断。

A line of best fit can be drawn by eye to summarise the trend and to make predictions.

可通过目测画出一条最佳拟合线,用以总结趋势并进行预测。


11. Regression Line | 回归线

The least-squares regression line minimises the sum of the squares of the vertical distances from the points to the line.

最小二乘回归线使各点到直线的竖直距离平方之和最小。

The equation of the regression line is usually written as y = a + bx.

回归线方程通常写作 y = a + bx。

The slope b and intercept a are calculated using the following formulas.

斜率 b 与截距 a 可按以下公式计算。

b = Σ((x – x̄)(y – ȳ)) / Σ(x – x̄)²

a = ȳ – b x̄

(Here ȳ represents y-bar, the mean of y-values, and x̄ represents the mean of x-values.)

(其中 ȳ 表示 y 的均值,x̄ 表示 x 的均值。)

The regression line can be used to estimate the value of y for a given x (interpolation) or beyond the data range (extrapolation), although extrapolation may be unreliable.

回归线可用于给定 x 值估计 y 值(内插),或预测数据范围之外的值(外推),但外推可能不可靠。


12. Summary of Key Symbols | 关键符号总结

Symbol Meaning 中文含义
x̄ Sample mean 样本平均数
μ Population mean 总体平均数
σ² Population variance 总体方差
s² Sample variance 样本方差
σ Population standard deviation 总体标准差
Q1, Q2, Q3 Quartiles 四分位数
IQR Interquartile range 四分位距
P(A) Probability of event A 事件 A 的概率
P(A|B) Conditional probability 条件概率
Σ Sum 求和

Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading