Year 9 CAIE Statistics: Formula & Theorem Quick Reference Handbook | Year 9 CAIE 统计:公式定理速查手册

📚 Year 9 CAIE Statistics: Formula & Theorem Quick Reference Handbook | Year 9 CAIE 统计:公式定理速查手册

Welcome to your one-stop quick reference for all key formulas, definitions, and theorems in Year 9 CAIE Statistics. This handbook organises essential content into short, exam-focused sections so you can revise efficiently. Each entry presents the concept in English first, followed immediately by its Chinese counterpart for bilingual clarity.

欢迎使用 Year 9 CAIE 统计学的关键公式、定义和定理速查手册。本手册将核心内容组织成简短、聚焦考试的章节,让你高效复习。每个条目先用英文呈现概念,紧接着用中文对应解释,实现双语清晰理解。


1. Mean (Arithmetic Average) | 算术平均值

The mean is the sum of all data values divided by the number of values. It is the most common measure of central tendency.

平均值是所有数据值之和除以数值的个数。它是最常用的集中趋势度量。

For ungrouped data, the formula is:

对于未分组数据,公式为:

Mean = (Σx) / n

where Σx is the sum of all observations and n is the total number of observations.

其中 Σx 是所有观测值的总和,n 是观测值的总数。

For a frequency table, the mean is calculated as:

对于频数表,平均值计算方式为:

Mean = (Σfx) / Σf

where f represents the frequency of each value x.

其中 f 表示每个值 x 的频数。


2. Median (Middle Value) | 中位数(中间值)

The median is the middle value when the data set is arranged in ascending order. It divides the data into two equal halves.

中位数是将数据集按升序排列后位于中间的值,它将数据分成两个相等的部分。

For an odd number of observations, the median is the value at position (n + 1) / 2.

如果观测值个数为奇数,中位数是第 (n + 1) / 2 个位置的值。

For an even number of observations, the median is the average of the two middle values, found at positions n/2 and (n/2 + 1).

如果观测值个数为偶数,中位数是中间两个值的平均数,这两个值位于第 n/2 和 (n/2 + 1) 个位置。


3. Mode (Most Frequent Value) | 众数(最常见值)

The mode is the value or category that occurs most frequently in a data set. A data set may have one mode (unimodal), more than one mode (bimodal or multimodal), or no mode if all values appear equally often.

众数是数据集中出现频率最高的值或类别。一个数据集可能有一个众数(单峰)、多个众数(双峰或多峰),或者如果没有值重复出现则没有众数。

For grouped frequency data, the modal class is the class interval with the highest frequency.

对于分组频数数据,众数类别是频数最高的组距区间。


4. Range (Measure of Spread) | 极差(离散度量)

The range is the simplest measure of dispersion, showing how spread out the data are.

极差是最简单的离散量数,反映数据的分散程度。

Range = Highest value − Lowest value

It is easy to compute but can be heavily affected by outliers.

它易于计算,但极易受异常值影响。


5. Frequency Tables & Cumulative Frequency | 频数表与累积频数

A frequency table organises raw data by counting how many times each value or class interval occurs. The cumulative frequency is the running total of frequencies up to that class.

频数表通过统计每个值或组距区间出现的次数来整理原始数据。累积频数是到该类别为止的频数累计和。

Cumulative frequency helps in finding the median and quartiles from grouped data.

累积频数有助于从分组数据中求中位数和四分位数。

  • To construct a cumulative frequency column, add each frequency to the sum of previous frequencies.
  • 要构建累积频数列,将每个频数加到前面所有频数之和上。

6. Bar Charts, Pie Charts & Histograms | 条形图、饼图和直方图

Bar charts represent categorical data with rectangular bars whose lengths are proportional to the frequencies. Bars are usually separated by gaps.

条形图用矩形条表示分类数据,条的长度与频数成比例。条之间通常留有间隙。

Pie charts display data as slices of a circle where the angle of each slice is proportional to the frequency: Angle = (Frequency / Total frequency) × 360°.

饼图将数据显示为圆形的扇区,每个扇区的角度与频数成比例:角度 = (频数 / 总频数) × 360°。

Histograms are used for continuous grouped data. In a histogram, the frequency is represented by the area of the bar. For equal class widths, frequency is proportional to height; for unequal class widths, we use frequency density:

直方图用于连续的分组数据。在直方图中,频率由条的面积表示。对于等宽组距,频数与高度成比例;对于不等宽组距,我们使用频数密度:

Frequency density = Frequency / Class width


7. Pictograms & Stem-and-Leaf Diagrams | 象形图与茎叶图

A pictogram uses symbols or pictures to represent a certain number of units. A key must always be provided to show the symbol’s value.

象形图用符号或图片表示一定数量的单位。必须始终提供图例来说明符号的价值。

A stem-and-leaf diagram organises numerical data while preserving each original value. The ‘stem’ represents the leading digit(s), and the ‘leaf’ represents the final digit. This plot quickly shows the shape of the distribution and is useful for finding the median and mode.

茎叶图组织数值数据,同时保留每个原始值。“茎”代表前导数字,“叶”代表最后一位数字。这种图形能快速显示分布形状,便于求中位数和众数。


8. Quartiles & Interquartile Range (IQR) | 四分位数与四分位距

Quartiles divide an ordered data set into four equal parts. The lower quartile (Q1) is the median of the lower half of the data, the median (Q2) is the second quartile, and the upper quartile (Q3) is the median of the upper half.

四分位数将有序数据集分成四个相等部分。下四分位数(Q1)是数据下半部分的中位数,中位数(Q2)是第二个四分位数,上四分位数(Q3)是数据上半部分的中位数。

The interquartile range measures the spread of the middle 50% of the data:

四分位距度量中间50%数据的离散程度:

IQR = Q3 − Q1

The IQR is not affected by extreme values, making it a more robust measure than the range.

四分位距不受极端值影响,因此比极差更稳健。


9. Basic Probability | 基础概率

Probability measures how likely an event is to occur, expressed as a number between 0 (impossible) and 1 (certain) or as a fraction, decimal, or percentage.

概率度量事件发生的可能性,用0(不可能)到1(必然)之间的数字表示,也可用分数、小数或百分比表示。

Probability of an event = (Number of favourable outcomes) / (Total number of possible outcomes)

The complement rule states that the probability of an event not occurring is 1 minus the probability that it does occur:

互补规则指出,事件不发生的概率等于1减去它发生的概率:

P(not A) = 1 − P(A)

For mutually exclusive events A and B, the probability that A or B occurs is the sum of their individual probabilities:

对于互斥事件 A 和 B,A 或 B 发生的概率等于各自概率之和:

P(A or B) = P(A) + P(B)


10. Probability from Experimental Data | 实验数据的概率

When outcomes are not equally likely, we estimate probability using relative frequency from an experiment or survey:

当结果不是等可能时,我们通过实验或调查中的相对频数来估计概率:

Estimated probability = (Frequency of event) / (Total number of trials)

As the number of trials increases, the experimental probability usually approaches the theoretical probability (Law of Large Numbers).

随着试验次数增加,实验概率通常趋近于理论概率(大数定律)。


11. Scatter Graphs & Correlation | 散点图与相关性

A scatter graph displays the relationship between two variables. Each point represents a pair of values (x, y).

散点图展现两个变量之间的关系。每个点代表一对值 (x, y)。

Correlation describes the direction and strength of the relationship:

相关性描述关系的方向和强度:

  • Positive correlation: as x increases, y tends to increase.
  • 正相关:当 x 增加时,y 也倾向于增加。
  • Negative correlation: as x increases, y tends to decrease.
  • 负相关:当 x 增加时,y 倾向于减少。
  • No correlation: no clear pattern between x and y.
  • 无相关:x 和 y 之间没有明确的模式。

We can draw a line of best fit (a straight line roughly passing through the points) to make predictions, but only for the range of data available (interpolation). Extrapolation outside the data range should be avoided unless trends are reliable.

我们可以绘制最佳拟合线(一条大致穿过各点的直线)来进行预测,但仅限于现有数据范围(内插)。除非趋势可靠,否则应避免在数据范围外外推。


12. Averages from Grouped Data (Estimates) | 分组数据的平均值(估算)

When data are grouped into class intervals, we cannot calculate the exact mean; instead we use the midpoint of each class as an estimate.

当数据被分成组距区间时,无法计算精确平均值;我们使用每个区间的中点作为估计值。

For each class, let the midpoint be m and the frequency be f. The estimated mean is:

对于每个组,设中点为 m,频数为 f。估算的平均值为:

Estimated mean = (Σfm) / Σf

Similarly, the modal class is the class with the highest frequency; the median class can be located using cumulative frequency by finding the class containing the (n/2)th value.

类似地,众数类别是频数最高的组;中位数组可以通过累积频数找到包含第 (n/2) 个值的组来确定。

Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version