Year 11 Cambridge Statistics: Vocabulary & Terminology Quick Memorisation Guide | 剑桥11年级统计:词汇术语速记指南

📚 Year 11 Cambridge Statistics: Vocabulary & Terminology Quick Memorisation Guide | 剑桥11年级统计:词汇术语速记指南

In Cambridge IGCSE Statistics, mastering key vocabulary is the first step to understanding data, graphs, probability, and inference. This guide presents essential terms paired with clear explanations in English and Chinese, helping you recall definitions quickly and apply them confidently in exams. Use the numbered sections to revise topic by topic, and notice how each English definition is immediately followed by its Chinese equivalent for rapid memorisation.

在剑桥IGCSE统计中,掌握关键术语是理解数据、图形、概率和推断的第一步。本指南呈现了重要的词汇,并配以英文和中文的清晰解释,帮助你快速记忆定义并在考试中自信应用。你可以按编号的章节逐主题复习,同时注意每个英文定义后都紧接中文对应,以便速记。

1. Types of Data | 数据类型

Categorical data: non‑numerical information grouped into categories, e.g. eye colour or favourite subject.

分类数据:非数字信息,按类别分组,例如眼睛颜色或最喜欢的科目。

Numerical data: data consisting of numbers. It is subdivided into discrete (countable values like number of students) and continuous (measurable values like height or time).

数值数据:由数字组成的数据。细分为离散数据(可数的值,如学生人数)和连续数据(可测量的值,如身高或时间)。

Qualitative data: describes qualities or characteristics; it is categorical in nature.

定性数据:描述性质或特征;本质上是分类数据。

Quantitative data: describes numerical quantities; it can be either discrete or continuous.

定量数据:描述数量的多少;可以是离散的或连续的。

Ordinal data: categorical data with a meaningful order but unequal intervals, e.g. survey ratings (poor, average, good).

有序分类数据:有排列顺序但间隔不等的分类数据,例如调查评分(差、一般、好)。

Nominal data: categorical data without any order, e.g. types of fruit.

名义分类数据:没有先后次序的分类数据,例如果汁的种类。


2. Sampling Methods & Bias | 抽样方法与偏差

Random sampling: every member of the population has an equal chance of being selected, reducing selection bias.

随机抽样:总体中每个成员被选中的机会相等,从而减少选择偏差。

Stratified sampling: the population is divided into subgroups (strata) and a random sample is taken from each in proportion to its size.

分层抽样:将总体分成若干子群(层),然后按各层的大小比例从中抽取随机样本。

Systematic sampling: individuals are chosen at regular intervals from a list, e.g. every 10th person.

系统抽样:从名单中按固定间隔选取个体,例如每隔10个人抽取一名。

Convenience sampling: a sample taken from people who are easy to reach; often leads to bias.

便利抽样:从容易接触到的人群中取样;通常会导致偏差。

Quota sampling: the interviewer selects a preset number of people from different groups, but not randomly.

配额抽样:访问员按预设数量从不同群组中选择受访者,但并非随机选取。

Bias: a systematic error that over‑ or underestimates the true value, often caused by poor sampling or misleading questions.

偏差:一种系统性误差,会高估或低估真实值,通常由抽样不当或引导性问题引起。


3. Measures of Central Tendency | 集中趋势度量

Mean: the arithmetic average, found by summing all values and dividing by the number of values. It is sensitive to outliers.

均值:算术平均数,将所有数值求和后除以数值个数。易受离群值影响。

Median: the middle value when the data are arranged in order. It splits the dataset into two equal halves and is resistant to outliers.

中位数:数据按大小排序后位于中间的值。它将数据集平分成两半,且不易受离群值影响。

Mode: the most frequently occurring value in a data set. There can be one mode, more than one (bimodal or multimodal) or no mode at all.

众数:数据集中出现次数最多的值。可以有一个众数、多个众数(双峰或多峰)或者没有众数。

Weighted mean: an average where each value is assigned a weight reflecting its importance, used in index numbers.

加权平均数:每个数值按其重要性赋予权重后计算的平均值,常用于指数编制。


4. Measures of Dispersion | 离差度量

Range: the difference between the largest and smallest values. Simple but heavily affected by extreme values.

极差(全距):最大值与最小值之差。计算简单但受极端值影响极大。

Interquartile range (IQR): the difference between the upper quartile (Q₃) and lower quartile (Q₁). It measures the spread of the middle 50% of the data.

四分位距 (IQR):上四分位数 (Q₃) 与下四分位数 (Q₁) 的差值,衡量中间50%数据的散布程度。

Variance: the average of the squared deviations from the mean. It gives the spread in squared units.

方差:各数据与均值离差平方的平均数。单位为原始单位的平方。

Standard deviation: the square root of the variance. It shows the typical distance of a value from the mean, expressed in the original units.

标准差:方差的平方根,表示一个典型值与均值的距离,单位为原始单位。

Percentile range: the difference between two specified percentiles, e.g. 10th to 90th percentile.

百分位数间距:两个指定百分位数之差,例如第10百分位数到第90百分位数。


5. Statistical Diagrams | 统计图表

Bar chart: uses equal‑width bars with gaps between them to display categorical data; the height represents frequency.

条形图:用等宽且有间隙的条形表示分类数据,条的高度代表频数。

Histogram: a diagram for continuous data where the area of each bar is proportional to frequency. Class width may vary, so frequency density is plotted.

直方图:用于连续数据的图形,各长方形的面积与频数成比例。组距可能不等,因此纵轴为频率密度。

Frequency density: calculated as frequency ÷ class width, essential for constructing histograms with unequal intervals.

频率密度:计算公式为频数÷组距,是绘制不等距直方图的关键量。

Pie chart: a circular chart divided into sectors, each representing a proportion of the whole.

饼图:划分为扇形的圆形图,每个扇形代表整体的一部分。

Cumulative frequency curve (ogive): a line graph showing the running total of frequencies; used to estimate medians and quartiles.

累积频数曲线(肩形图):显示频数累计总和的折线图,用于估计中位数和四分位数。

Box plot: displays a five‑number summary (minimum, Q₁, median, Q₃, maximum) and highlights outliers.

箱线图:显示五数概括(最小值、Q₁、中位数、Q₃、最大值)并突出离群值。


6. Frequency Distributions | 频数分布

Frequency: the number of times a value or observation occurs in a data set.

频数:某个数值或观测值在数据集中出现的次数。

Class interval: a range of values into which raw data are grouped for a frequency table.

组距区间:将原始数据分组时所用的数值范围。

Class boundaries: the precise limits between classes, ensuring no gaps between intervals in histograms.

组界限:各组之间精确的分界值,确保证直方图中间没有间隙。

Midpoint: the central value of a class interval, used to represent the group in calculations of mean or standard deviation.

组中值:组距区间的中心值,在计算均值或标准差时代替该组。

Relative frequency: frequency divided by total frequency, often expressed as a decimal or percentage.

相对频数:频数除以总频数,通常用小数或百分数表示。


7. Probability Basics | 概率基础

Probability: a measure of the likelihood that an event will occur, expressed on a scale from 0 (impossible) to 1 (certain).

概率:衡量事件发生可能性的度量,范围从0(不可能)到1(必然)。

Experiment: a repeatable process that gives rise to outcomes, e.g. rolling a die.

试验:一个可重复的过程,产生若干结果,例如掷骰子。

Outcome: a possible result of an experiment.

结果:试验的一种可能效果。

Event: a set of one or more outcomes, e.g. rolling an even number.

事件:由一个或多个结果组成的集合,例如掷出偶数。

Sample space: the list of all possible outcomes of an experiment.

样本空间:一个试验所有可能结果的清单。

Mutually exclusive events: events that cannot happen at the same time; P(A ∩ B) = 0.

互斥事件:不可能同时发生的事件;P(A ∩ B) = 0。

Independent events: the occurrence of one event does not affect the probability of the other; P(A ∩ B) = P(A) × P(B).

独立事件:一个事件的发生不影响另一个事件的概率;P(A ∩ B) = P(A) × P(B)。

Conditional probability: the probability of an event A given that B has happened, written as P(A|B).

条件概率:在事件B已发生的条件下,事件A发生的概率,记为P(A|B)。


8. Correlation and Regression | 相关与回归

Correlation: describes the strength and direction of a linear relationship between two variables.

相关:描述两个变量之间线性关系的强度和方向。

Positive correlation: as one variable increases, the other also tends to increase.

正相关:当一个变量增加时,另一个变量也趋向增加。

Negative correlation: as one variable increases, the other tends to decrease.

负相关:当一个变量增加时,另一个变量趋向减少。

Line of best fit: a straight line drawn through a scatter graph that best represents the trend in the data, used for prediction.

最佳拟合线:在散点图上绘制的直线,最佳地表示数据趋势,用于预测。

Interpolation: predicting a value within the range of the given data. It is usually reliable.

内插法:在已知数据范围内进行预测,通常较为可靠。

Extrapolation: predicting a value outside the range of the data. It can be unreliable because the trend might change.

外推法:在已知数据范围外进行预测,可能不够可靠,因为趋势可能改变。

Causation: a cause‑and‑effect relationship. Importantly, correlation does not imply causation.

因果关系:一种因果联系。重要的是,相关未必意味着因果。


9. Cumulative Frequency & Percentiles | 累积频数与百分位数

Cumulative frequency: the running total of frequencies up to the end of each class interval.

累积频数:截至每个组距区间末端的频数累加总和。

Percentile: a value below which a given percentage of observations fall. The median is the 50th percentile.

百分位数:某个数值,使得一定百分比的观测值低于它。中位数即第50百分位数。

Lower quartile (Q₁): the 25th percentile; the median of the lower half of the data.

下四分位数 (Q₁):第25百分位数,即数据下半部分的中位数。

Upper quartile (Q₃): the 75th percentile; the median of the upper half.

上四分位数 (Q₃):第75百分位数,即数据上半部分的中位数。

Interpercentile range: the difference between two percentiles, often the 10th and 90th, used to describe spread.

百分位数间距:两个百分位数之差,常用第10和第90百分位数,用于描述散布。


10. Time Series & Index Numbers | 时间序列与指数

Time series: a sequence of data points recorded at regular time intervals, used to identify trends and seasonal patterns.

时间序列:按固定时间间隔记录的一系列数据点,用于识别趋势和季节波动模式。

Trend: the underlying long‑term movement of a time series after removing short‑term fluctuations.

趋势:剔除短期波动后时间序列的潜在长期变化方向。

Moving average: a series of averages calculated from successive groups of time‑series data, used to smooth out irregularities and highlight the trend.

移动平均:由时间序列中连续分组数据计算的平均值序列,用于抹平不规则波动并凸显趋势。

Index number: a statistical measure that compares the value of a variable with a base value, typically set to 100. It simplifies comparison over time.

指数:将变量值与基期值(通常设为100)比较的统计度量,简化不同时间的对比。

Base year: the reference year in index number construction, assigned an index of 100.

基年:指数编制中的参考年份,其指数定为100。


11. Essential Support Vocabulary | 必备辅助词汇

Population: the entire set of individuals or items of interest in a statistical study.

总体:统计研究中感兴趣的全体个体或项目的集合。

Sample: a subset of the population selected for investigation, intended to represent the population.

样本:为调查研究而选出的总体子集,旨在代表总体。

Parameter: a numerical characteristic of a population, e.g. population mean (μ).

参数:总体的数值特征,例如总体均值 (μ)。

Statistic: a numerical characteristic calculated from a sample, used to estimate a parameter, e.g. sample mean (x̄).

统计量:根据样本计算出的数值特征,用于估计参数,例如样本均值 (x̄)。

Census: a study that gathers data from every member of the population. It eliminates sampling error but can be expensive and time‑consuming.

普查:从总体每个成员收集数据的研究。它消除了抽样误差,但可能昂贵且耗时。

Sampling frame: a list of all members of the population from which a sample is drawn. A poor frame can lead to bias.

抽样框:总体所有成员的名单,从中抽取样本。不完整的抽样框会导致偏差。

Outlier: an observation that lies far from the rest of the data. It may be caused by error or natural variability and can strongly affect the mean and range.

离群值:远离其他数据的观测值。可能由误差或自然变异引起,会严重影响均值和极差。

Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading