📚 AS CCEA Statistics: Vocabulary Memorisation Guide | AS CCEA 统计:词汇术语速记指南
Mastering the language of statistics is the first step towards confident problem-solving in your AS CCEA examinations. This guide breaks down over 80 essential terms into 12 focused sections, each packed with definitions, examples, mnemonics, and dual-language explanations so you can learn actively and retain key concepts. Whether you are revising sampling methods, probability distributions, or regression analysis, this resource is designed to turn abstract jargon into clear, exam-ready knowledge.
掌握统计学的语言是在 AS CCEA 考试中自信解题的第一步。本指南将 80 多个核心术语划分为 12 个专题小节,每节都配有定义、示例、记忆口诀和中英双语解释,帮助你主动学习并牢固掌握关键概念。无论你是在复习抽样方法、概率分布还是回归分析,这份资料都能将抽象的术语转化为清晰、应试必备的知识。
1. Core Statistical Concepts | 核心统计概念
A population is the entire group of individuals or items of interest. A sample is a subset of the population selected for study. A parameter is a numerical summary describing a population (e.g., population mean μ), whereas a statistic is a numerical summary calculated from a sample (e.g., sample mean x̄). A census attempts to collect data from every member of the population, while a sample survey collects information from a sample only.
总体是研究对象的全部个体或项目集合。样本是从总体中选出用于研究的子集。参数是描述总体的数值概括(如总体均值 μ),统计量则是由样本计算出的数值概括(如样本均值 x̄)。普查试图从总体每个成员收集数据,而抽样调查仅从样本中获取信息。
Memory aid: The letters ‘P’ and ‘S’ link Population ↔ Parameter and Sample ↔ Statistic. A parameter is fixed but usually unknown; a statistic varies from sample to sample and can be used to estimate a parameter.
记忆窍门:字母 “P” 和 “S” 分别对应 Population/Parameter 与 Sample/Statistic。参数是固定的但通常未知;统计量会随样本变化并可用于估计参数。
2. Data Types and Measurement Scales | 数据类型与测量尺度
Data can be qualitative (categorical) when they describe qualities or labels, or quantitative (numerical) when they represent counts or measurements. Qualitative data are either nominal (categories with no natural order, e.g., hair colour) or ordinal (categories with a logical order, e.g., satisfaction ratings). Quantitative data are discrete (take only distinct, countable values, e.g., number of cars) or continuous (can take any value within an interval, e.g., height). Primary data are collected first-hand by the researcher, whereas secondary data already exist.
数据可以是定性(分类)的,描述属性或标签;也可以是定量(数值)的,表示计数或测量值。定性数据分为名义数据(无自然顺序的类别,如头发颜色)和有序数据(有逻辑顺序的类别,如满意度评分)。定量数据分为离散数据(只能取有限个可数值,如汽车数量)和连续数据(可在区间内取任意值,如身高)。原始数据由研究者亲自收集,二手数据则已存在。
| English | 中文 | Quick Example / 速例 |
|---|---|---|
| Nominal | 名义 | Blood type / 血型 |
| Ordinal | 有序 | Exam grade / 考试等级 |
| Discrete | 离散 | Number of pets / 宠物数量 |
| Continuous | 连续 | Time / 时间 |
3. Sampling Methods | 抽样方法
A simple random sample ensures every member of the population has an equal chance of selection, usually obtained by a lottery or random number generator. Stratified sampling divides the population into distinct groups (strata) and takes a random sample from each in proportion to its size. Systematic sampling selects members at regular intervals from a list. Quota sampling selects a pre‑determined number of individuals with specified characteristics, often used in market research. Opportunity (convenience) sampling uses subjects who are readily available.
简单随机抽样保证总体中每个成员被抽中的机会相等,通常通过抽签或随机数生成器实现。分层抽样先将总体分为不同的组群(层),然后按比例从每层中随机抽取样本。系统抽样按固定间隔从名单中抽取成员。配额抽样选取预先设定数量的具有特定特征的个体,常用于市场调研。便利抽样则使用最易获得的受试者。
Remember: Random and stratified methods reduce bias and allow inference; quota and convenience sampling are non‑random and may produce biased samples.
记忆要点:随机和分层抽样可减少偏差并能进行统计推断;配额与便利抽样是非随机的,可能产生偏差样本。
4. Frequency Distributions and Graphs | 频率分布与统计图
For grouped data, class boundaries separate classes without gaps, and the class midpoint is the average of the upper and lower boundaries. In a histogram, the area of each bar is proportional to frequency: frequency density = frequency ÷ class width. A cumulative frequency graph helps estimate medians and quartiles. A box plot (box‑and‑whisker plot) displays the minimum, Q₁, median, Q₃ and maximum, easily revealing skewness and potential outliers.
对于分组数据,组界使各组之间无缝隙衔接,组中点是上下组界的平均数。在直方图中,每个长方条的面积与频数成正比:频数密度 = 频数 ÷ 组距。累积频数图可用于估算中位数和四分位数。箱线图(箱须图)展示最小值、Q₁、中位数、Q₃ 和最大值,能直观显示偏度和可能的异常值。
The five-number summary (min, Q₁, median, Q₃, max) forms the backbone of a box plot and gives a robust picture of the data’s spread and centre.
五数概括(最小值、Q₁、中位数、Q₃、最大值)是箱线图的骨架,能稳健地刻画数据的散布和中心。
5. Measures of Central Tendency | 集中趋势度量
The mean (x̄ for a sample, μ for a population) is the arithmetic average: for ungrouped data x̄ = Σx/n, and for grouped data x̄ = Σfx/Σf. The median is the middle value when data are ordered; its position is (n+1)/2. The mode is the most frequently occurring value. The mean uses all values and is sensitive to outliers, while the median is resistant to extreme values.
均值(样本记为 x̄,总体记为 μ)是算术平均数:未分组数据为 x̄ = Σx/n,分组数据为 x̄ = Σfx/Σf。中位数是排序后居于中间位置的值,位置为 (n+1)/2。众数是出现次数最多的值。均值利用了全部数据,易受异常值影响;中位数则对极端值具有抗干扰性。
In a symmetric distribution, mean = median = mode. In a positively skewed distribution, mean > median > mode; in a negatively skewed distribution, mean < median < mode.
在对称分布中,均值 = 中位数 = 众数。正偏态时均值 > 中位数 > 众数;负偏态时均值 < 中位数 < 众数。
6. Measures of Dispersion | 离散程度度量
Range = maximum − minimum. The interquartile range (IQR) = Q₃ − Q₁, measuring the spread of the middle 50% of the data. Variance uses squared deviations: sample variance s² = Σ(x − x̄)²/(n−1); population variance σ² = Σ(x − μ)²/N. The standard deviation is the square root of variance: s or σ. A larger standard deviation indicates greater variability.
极差 = 最大值 − 最小值。四分位距 (IQR) = Q₃ − Q₁,衡量中间 50% 数据的散布。方差使用离差平方和:样本方差 s² = Σ(x − x̄)²/(n−1),总体方差 σ² = Σ(x − μ)²/N。标准差为方差的平方根:s 或 σ。标准差越大表示变异性越强。
When comparing data sets, use mean and standard deviation if the distribution is symmetric; use median and IQR when it is skewed or contains outliers.
比较数据集时,若分布对称则使用均值与标准差;若分布偏斜或含异常值则使用中位数与四分位距。
7. Probability Fundamentals | 概率基础
An experiment is a repeatable process with observable outcomes. The sample space (S) is the set of all possible outcomes. An event is a subset of the sample space. The probability of an event A, P(A), satisfies 0 ≤ P(A) ≤ 1. The complement of A is A′ and P(A′) = 1 − P(A). Two events are mutually exclusive if they cannot occur together: P(A ∩ B) = 0. They are independent if the occurrence of one does not affect the probability of the other: P(A ∩ B) = P(A) × P(B).
试验是一个可重复且有可观测结果的过程。样本空间 (S) 是所有可能结果的集合。事件是样本空间的子集。事件 A 的概率 P(A) 满足 0 ≤ P(A) ≤ 1。补事件 A′ 的概率 P(A′) = 1 − P(A)。若两个事件不能同时发生,则它们互斥:P(A ∩ B) = 0。若一个事件的发生不影响另一个事件的概率,则它们相互独立:P(A ∩ B) = P(A) × P(B)。
P(A|B) = P(A ∩ B) / P(B), P(B) > 0
Conditional probability P(A|B) is the probability of A given that B has occurred. It helps analyse dependent events and leads to Bayes’ theorem in more advanced study.
条件概率 P(A|B) 表示在 B 已发生的条件下 A 发生的概率。它有助于分析相关事件,并在更高阶学习中引出贝叶斯定理。
8. Discrete Random Variables | 离散随机变量
A discrete random variable X takes a countable number of possible values, each with an associated probability P(X = x). The probability distribution lists all values of X and their probabilities; the total probability must equal 1. The expected value E(X) is the long‑run average: E(X) = Σ x P(X = x). The variance Var(X) = E[(X − μ)²] = Σ x² P(X = x) − [E(X)]².
离散随机变量 X 可取可数个可能值,每个值对应一个概率 P(X = x)。概率分布列出 X 的所有取值及其概率,总概率必须等于 1。期望值 E(X) 是长期平均值:E(X) = Σ x P(X = x)。方差 Var(X) = E[(X − μ)²] = Σ x² P(X = x) − [E(X)]²。
Linear transformations follow simple rules: E(aX + b) = a E(X) + b and Var(aX + b) = a² Var(X).
线性变换遵循简单规则:E(aX + b) = a E(X) + b,Var(aX + b) = a² Var(X)。
9. Binomial Distribution | 二项分布
A binomial experiment has a fixed number n of independent
Published by TutorHao | AS 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导