IGCSE OCR Statistics: Vocabulary & Terminology Quick Recall Guide | IGCSE OCR 统计:词汇术语速记指南

📚 IGCSE OCR Statistics: Vocabulary & Terminology Quick Recall Guide | IGCSE OCR 统计:词汇术语速记指南

Mastering the language of statistics is half the battle in your IGCSE OCR Statistics exam. This guide provides a focused, bilingual glossary of essential terms, designed to help you quickly recall definitions, spot common misconceptions, and confidently interpret exam questions. Each term is paired with its Chinese equivalent and a clear explanation in both languages, ensuring deep understanding rather than rote memorisation.

掌握统计学的语言是攻克 IGCSE OCR 统计考试的关键。本书面双语指南精选了核心术语,旨在帮助你快速记忆定义、识别常见误区,并自信地解读试题。每个术语都配有中文对照和清晰的双语解释,确保你真正理解而非死记硬背。

1. Data Types & Collection | 数据类型与数据收集

Primary data is information you collect yourself for a specific purpose, such as conducting a survey or experiment. Secondary data is data that has already been collected by someone else, like information from the internet or a textbook.

原始数据 是你自己出于特定目的而收集的信息,例如进行调查或实验。二手数据 是他人已经收集好的数据,比如来自互联网或教科书的信息。

A population is the entire set of individuals or items that we are interested in studying. A sample is a subset of the population, chosen to represent the whole. A sampling frame is a list of all members of the population from which the sample is drawn.

总体 是我们感兴趣的整个个体或项目集合。样本 是从总体中选出的一个子集,用以代表整体。抽样框 是列出总体所有成员的清单,样本从该清单中抽取。

Qualitative data (or categorical data) describes qualities or categories, such as eye colour or favourite sport. Quantitative data is numerical and can be discrete (countable, e.g., number of cars) or continuous (measurable, e.g., height, weight).

定性数据(或称分类数据)描述的是性质或类别,例如眼睛颜色或最喜爱的运动。定量数据 是数值型数据,可分为离散型(可数,如汽车数量)或连续型(可测量,如身高、体重)。


2. Sampling Methods | 抽样方法

Random sampling means every member of the population has an equal chance of being selected. This reduces bias. A simple random sample can be generated using random number tables or a calculator.

随机抽样 是指总体的每个成员都有相同被选中的机会。这能减少偏差。简单随机样本 可使用随机数表或计算器生成。

Stratified sampling involves dividing the population into distinct groups (strata) and then taking a random sample from each group proportional to its size. This guarantees representation of key subgroups.

分层抽样 是先将总体分成不同的组(层),然后从每一层中按比例随机抽取样本。这保证了关键子群体的代表性。

Systematic sampling selects individuals at regular intervals from a list after a random start. For example, choosing every 10th name. Quota sampling is a non-random method where interviewers select a fixed number of individuals with certain characteristics.

系统抽样 是从清单中以固定间隔选取个体,起始点随机。例如,每 10 个名字选一个。配额抽样 是一种非随机方法,采访者按照某些特征选取固定数量的个体。

Convenience sampling (or opportunity sampling) uses people who are readily available, which often leads to biased results. A voluntary response sample is where participants choose themselves, also prone to bias.

便利抽样(或称机会抽样)使用最容易接触到的人,这常常导致结果有偏。自愿回应样本 是参与者自己选择加入,也容易出现偏差。


3. Frequency & Graphical Terms | 频数与图表术语

Frequency is the number of times a value or category occurs. The class interval is the range of values grouped together in a frequency table, e.g., 10 ≤ x < 20. The class width is the difference between the upper and lower limits of a class interval (here, 10).

频数 是某个数值或类别出现的次数。组距 是在频数表中将数值分组后的范围,例如 10 ≤ x < 20。组宽度 是组距上限与下限之差(此处为 10)。

A histogram is a diagram for grouped continuous data where the area of each bar is proportional to the frequency. If class widths are unequal, frequency density is plotted on the vertical axis, calculated as: frequency ÷ class width.

直方图 是用于分组连续数据的图表,其中每一条形的面积与频数成正比。如果各组宽度不等,纵轴使用频数密度,计算公式为:频数 ÷ 组宽度。

A frequency polygon is formed by joining the midpoints of the tops of histogram bars with straight lines. A cumulative frequency graph (or ogive) plots the running total of frequencies against the upper class boundary; it helps find medians and percentiles.

频数折线图 是通过线段连接每个直方条顶部中点而形成的。累积频数图(或称肩形图)以累积频数对组上限值作图,可用于查找中位数和百分位数。


4. Averages & Measures of Central Tendency | 平均数与集中趋势量数

The mode is the value that occurs most often. A data set can have one mode (unimodal), two modes (bimodal), or more. The median is the middle value when data is ordered; for n values, its position is (n+1)/2.

众数 是出现次数最多的数值。数据集可以有一个众数(单峰)、两个众数(双峰)或更多。中位数 是将数据排序后位于中间的数值;对于 n 个数据,其位置为 (n+1)/2。

The mean (arithmetic average) is calculated by summing all values and dividing by the number of values. For grouped data, use the midpoints: Mean = Σ(f × x) / Σf.

平均数(算术平均值)的计算方法是将所有数值之和除以数值个数。对于分组数据,使用组中点:平均数 = Σ(f × x) / Σf。

These three measures are often compared: the mean uses all data but is affected by outliers; the median is robust to outliers; the mode is the only average suitable for qualitative data.

这三个量数常被比较:平均数用了所有数据但受异常值影响;中位数对异常值稳健;众数是唯一适用于定性数据的平均数。


5. Measures of Spread & Dispersion | 离散程度量数

Range is the simplest measure: Range = maximum value – minimum value. It ignores the distribution of middle values. The interquartile range (IQR) is the difference between the upper quartile (Q₃) and the lower quartile (Q₁). It describes the spread of the middle 50% of data.

极差 是最简单的量数:极差 = 最大值 – 最小值。它忽略了中间值的分布情况。四分位距 (IQR) 是上四分位数 (Q₃) 与下四分位数 (Q₁) 之差,描述了中间 50% 数据的离散程度。

Quartiles are found from ordered data or cumulative frequency graphs. The lower quartile (Q₁) is at position (n+1)/4, the median at (n+1)/2, and the upper quartile (Q₃) at 3(n+1)/4.

四分位数可通过排序数据或累积频数图求得。下四分位数 (Q₁) 位于 (n+1)/4 的位置,中位数位于 (n+1)/2,上四分位数 (Q₃) 位于 3(n+1)/4 的位置。

Percentile is a value below which a given percentage of observations fall. For example, the 90th percentile is the value below which 90% of the data lies. The median is the 50th percentile.

百分位数 是某一给定百分比的观测值所落在其下的那个数值。例如,第 90 百分位数表示有 90% 的数据低于该值。中位数即为第 50 百分位数。


6. Correlation & Scatter Graphs | 相关与散点图

Bivariate data involves two variables measured for each individual. A scatter graph (scatter plot) displays the relationship. Correlation describes the strength and direction of a linear relationship.

双变量数据 涉及对每个个体测量两个变量。散点图 用于展示二者关系。相关 描述的是线性关系的强度和方向。

Positive correlation means as one variable increases, the other tends to increase. Negative correlation means as one increases, the other tends to decrease. If no pattern exists, it is zero correlation or no correlation.

正相关 表示一个变量增大时,另一个也倾向于增大。负相关 表示一个变量增大时,另一个倾向于减小。如果没有规律,则为零相关或无相关。

Correlation does not imply causation. Even a strong correlation may be due to a third factor (lurking variable) or coincidence. The line of best fit is drawn roughly through the middle of points, balancing the distances on both sides; it can be used to estimate values (interpolation/extrapolation).

相关不代表因果。 即使相关性很强,也可能是由第三因素(潜在变量)或巧合造成。最佳拟合线 大致穿过点阵中间,平衡两侧的偏差距离,可用于估算数值(内插/外推)。


7. Probability Basics | 概率基础

Probability is a measure of how likely an event is to happen, expressed as a number between 0 (impossible) and 1 (certain). It can be written as a fraction, decimal, or percentage. The sample space is the set of all possible outcomes.

概率 衡量某一事件发生的可能性大小,用 0(不可能)到 1(必然)之间的数字表示,可以写成分数、小数或百分比。样本空间 是所有可能结果的集合。

For equally likely outcomes: P(Event) = Number of favourable outcomes / Total number of outcomes. The complement rule states: P(not A) = 1 – P(A).

对于等可能结果:P(事件) = 有利结果的数量 / 结果总数。补集规则:P(非 A) = 1 – P(A)。

Events are mutually exclusive if they cannot occur at the same time (e.g., getting a head and tail in one coin toss). They are independent if the occurrence of one does not affect the probability of the other. For independent events A and B: P(A and B) = P(A) × P(B).

若两个事件不能同时发生(如掷一枚硬币同时得到正面和反面),则称它们互斥。若一个事件的发生不影响另一个事件发生的概率,则它们是独立的。对于独立事件 A 和 B:P(A 且 B) = P(A) × P(B)。

Venn diagrams and tree diagrams are powerful tools for organising probabilities, especially for combined events and conditional probability.

维恩图和树形图是整理概率的强大工具,尤其适用于组合事件和条件概率。


8. Probability Distributions & Expectation | 概率分布与期望

A probability distribution lists all possible outcomes of a discrete random variable and their associated probabilities. The sum of all probabilities must equal 1.

概率分布 列出了离散型随机变量的所有可能结果及其对应概率。所有概率之和必须等于 1。

The expected value (mean) of a discrete random variable X is E(X) = Σ [x × P(X=x)]. It represents the theoretical long-run average.

离散型随机变量 X 的期望值(平均数)为 E(X) = Σ [x × P(X=x)],代表理论上的长期平均值。

Experimental probability (relative frequency) is based on actual trials: Number of times event occurred / Total number of trials. As the number of trials increases, the experimental probability tends to stabilise around the theoretical probability (Law of Large Numbers).

实验概率(相对频数)基于实际试验:事件发生的次数 / 总试验次数。随着试验次数增加,实验概率会趋于理论概率(大数定律)。


9. Statistical Diagrams & Interpretation | 统计图表与解读

A bar chart is used for categorical or discrete data with gaps between bars; bar height represents frequency. A pie chart shows proportions of a whole, with sector angles calculated as (frequency / total) × 360°.

条形图 用于分类或离散数据,条形之间有间隙;条的高度代表频数。饼图 展示各部分在整体中的比例,扇形角度 = (频数 / 总数) × 360°。

A stem-and-leaf diagram shows the shape of the distribution while preserving the original data. The key explains how to read the digits. A box-and-whisker plot (box plot) displays five-number summary: minimum, Q₁, median, Q₃, maximum, helping to visualise central tendency and spread.

茎叶图 在保留原始数据的同时展示分布形态。通过“图例”说明如何解读数字。箱线图(箱须图)展示五数概括:最小值、Q₁、中位数、Q₃、最大值,有助于直观呈现集中趋势和离散程度。

Outliers are extreme values that lie far from the majority of data; commonly a value is considered an outlier if it is below Q₁ – 1.5 × IQR or above Q₃ + 1.5 × IQR. Outliers should be investigated but not removed without reason.

异常值 是远离数据主体的极端值;通常若数值小于 Q₁ – 1.5 × IQR 或大于 Q₃ + 1.5 × IQR,则视为异常值。异常值需探究原因,但不可无故删除。


10. Time Series & Index Numbers | 时间序列与指数

A time series is data collected at regular intervals over time. It may show trend (long-term movement), seasonal variation (regular pattern within fixed periods), and random fluctuations. The trend can be found using moving averages.

时间序列 是每隔固定时间间隔收集的数据。它可能显示出趋势(长期变动)、季节变动(固定周期内的规律模式)和随机波动。趋势可通过移动平均求得。

When calculating a moving average, take the average of a consecutive block of seasons (e.g., for quarterly data, a 4-point moving average) and centre it if needed. Seasonal variation = actual value – trend value.

计算移动平均时,取连续数个季节的数据求平均(如季度数据用 4 点移动平均),必要时使其居中。季节变动 = 实际值 – 趋势值。

Index numbers are used to compare prices or quantities over time. The base year index is usually 100. The index for a given year is calculated as (Value in given year / Value in base year) × 100. Composite index combines several weighted items.

指数 用于比较不同时期的价格或物量。基年指数通常设为 100。某年指数 = (该年数值 / 基年数值) × 100。综合指数 加权合并了多个项目。


11. Common Exam Pitfalls & Quick Tips | 常见考试陷阱与速记贴士

Check the axes: Is a histogram using frequency or frequency density? Are class boundaries marked correctly? Remember context: When interpreting mean/median, use units (e.g., cm, £). Skewness: In a right-skewed distribution: mean > median > mode; left-skewed: mean < median < mode.

检查坐标轴: 直方图的纵轴是频数还是频数密度?组界标注是否正确?注意情境: 解释平均数/中位数时要带单位(如 cm、£)。偏态: 右偏分布中平均数 > 中位数 > 众数;左偏分布中平均数 < 中位数 < 众数。

Probability wording: ‘Given that’ signals conditional probability. Tree diagrams: Multiply along branches for ‘and’, add probabilities from different branches for ‘or’. Always label probabilities and check they sum to 1 at each set of branches.

概率措辞: “已知……”提示条件概率。树形图: 沿分支相乘计算“且”,不同分支概率相加计算“或”。务必标注概率,并检查每组分支的概率之和为 1。

Don’t confuse: Mutually exclusive (cannot happen together) vs. independent (one does not affect the other’s probability). Bar chart vs. histogram: Bar charts have gaps and are for discrete/categorical data; histograms have no gaps (unless a class is empty) and are for continuous data.

勿混淆: 互斥(不能同时发生)与独立(一个的发生不影响另一个的概率)。条形图 vs. 直方图: 条形图有间隙,用于离散/分类数据;直方图无间隙(除非某组频数为零),用于连续数据。


12. Summary Table of Key Notation | 关键符号速查表

Familiarity with standard notation saves time in the exam. Below is a quick reference of commonly used symbols in IGCSE OCR Statistics.

熟悉标准符号能节省考试时间。以下是 IGCSE OCR 统计中常用符号的快速索引。

Notation Meaning 中文含义
n Sample size, number of data points 样本量,数据个数
Σ Sum 求和
x̄ (x-bar) Sample mean 样本平均数
μ (mu) Population mean 总体平均数
σ (sigma) Population standard deviation 总体标准差
s Sample standard deviation 样本标准差
Q₁, Q₂, Q₃ Quartiles (Q₂ = median) 四分位数(Q₂ = 中位数)
IQR Interquartile range = Q₃ – Q₁ 四分位距
f Frequency 频数
cf Cumulative frequency 累积频数
P(A) Probability of event A 事件 A 的概率
P(A | B) Probability of A given B (conditional) 在 B 发生的条件下 A 的概率

Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading