Mastering IGCSE AQA Statistics Vocabulary: A Quick Memorization Guide | IGCSE AQA 统计词汇术语速记指南

📚 Mastering IGCSE AQA Statistics Vocabulary: A Quick Memorization Guide | IGCSE AQA 统计词汇术语速记指南

The IGCSE AQA Statistics course demands a solid understanding of key terminology to interpret data, construct graphs, and perform calculations accurately. This guide breaks down essential statistical terms into logical groups with bilingual explanations and memory aids, helping you master the vocabulary quickly and efficiently.

IGCSE AQA 统计学课程要求你扎实掌握关键术语,才能准确解读数据、构建图表并进行计算。本指南将核心统计词汇按逻辑分组,提供双语解释和记忆辅助,助你快速高效掌握术语。

1. Core Concepts: Population, Sample & Variable | 核心概念:总体、样本与变量

Statistics is the science of collecting, organizing, analysing, and interpreting numerical data to draw conclusions and make informed decisions.

统计学是收集、整理、分析并解释数值数据,从而得出结论并作出明智决策的科学。

Population refers to the entire set of individuals, items, or measurements that you are interested in studying. A parameter (e.g., population mean μ) describes a characteristic of a population.

总体指的是你感兴趣研究的所有个体、项目或测量值的全集。参数(如总体均值 μ)描述总体的一种特征。

Sample is a subset of the population selected to represent the whole. A statistic (e.g., sample mean x̄) is a numerical summary calculated from a sample.

样本是从总体中选出的一个子集,用于代表全体。统计量(如样本均值 x̄)是从样本计算出的数值概括。

Variable: any characteristic, number, or quantity that can be measured or counted and can take different values. Variables can be categorical or numerical.

变量:任何可以测量或计数且可取不同值的特征、数字或数量。变量可以是分类的或数值的。

Memory trick: Think of ‘population’ as the whole pie, ‘sample’ as a slice you taste, and ‘parameter vs statistic’ like the true recipe (p) vs your estimate (s).

记忆技巧:将“总体”想作整个馅饼,“样本”是你品尝的一小块;“参数与统计量”就如真实的配方(p)与你的估算(s)。


2. Data Types: Qualitative, Quantitative, Discrete & Continuous | 数据类型:定性、定量、离散与连续

Qualitative (categorical) data describes qualities or categories that cannot be measured numerically, such as colours, brands, or types of animal. Think ‘quality’.

定性(分类)数据描述无法用数值测量的性质或类别,如颜色、品牌或动物种类。牢记“性质/质量”。

Quantitative (numerical) data consists of numbers representing counts or measurements. It is split into discrete and continuous.

定量(数值)数据由代表计数或测量结果的数字组成,可进一步分为离散数据和连续数据。

Discrete data can only take specific, separate values – usually whole numbers that you can count (number of students, shoe size). One memory aid: ‘Discrete – distinct, counting steps’.

离散数据只能取特定的、分离的数值——通常是可以计数的整数(如学生人数、鞋码)。助记:“离散——独立可数的步阶”。

Continuous data can take any value within a given range and is measured, not counted (height, temperature, time). Think of a continuous smooth line.

连续数据在给定范围内可取任何数值,是通过测量而非计数得到的(身高、温度、时间)。想象一条连续平滑的线。

Discrete / 离散 Continuous / 连续
Counted (计数) Measured (测量)
Gaps between values (数值间有间隔) No gaps, can be any decimal (无间隔,可为任何小数)
e.g. goals, pages (例:进球数、页数) e.g. weight, speed (例:重量、速度)

3. Data Collection Methods: Census, Survey & Sampling | 数据收集方法:普查、调查与抽样

Census: a survey that collects data from every member of the population. It is accurate but expensive and time-consuming. ‘Census = Complete count’.

普查:从总体中每一个成员收集数据的调查。结果准确但昂贵耗时。助记:“普查 = 全面清点”。

Sample survey: collecting data from only a portion (sample) of the population. Must be carefully designed to avoid bias.

抽样调查:只从总体的一部分(样本)收集数据。必须精心设计以避免偏差。

Random sampling: every member of the population has an equal chance of being selected. This gives an unbiased representative sample.

随机抽样:总体的每个成员被选中的机会均等,可获得无偏的代表性样本。

Stratified sampling: the population is divided into distinct groups (strata), and a random sample is taken from each group in proportion to its size. This ensures all subgroups are represented.

分层抽样:将总体分成不同的层(strata),按各层大小比例从每层随机抽取样本。这确保所有子群体都被代表。

Systematic sampling: choosing every k‑th item from a list after a random start. Quick to use but can introduce pattern bias.

系统抽样:从名单中随机起点后每隔 k 个抽取一个单位。使用便捷,但可能引入模式偏差。

Bias: a systematic error that causes a sample to be unrepresentative of the population. Sources include non‑response, leading questions, or convenience sampling.

偏差:导致样本无法代表总体的系统性误差。来源包括无回答、诱导性问题或便利抽样。


4. Frequency Distributions: Tables & Grouped Data | 频率分布:表格与分组数据

Frequency is the count of how many times a particular value or category occurs in a data set. A frequency table organises these counts alongside the data values or categories.

频数是指数据集中某个特定值或类别出现的次数。频数表将这些计数与数据值或类别一起整理列出。

Grouped frequency table: used for continuous or large discrete data sets. Data are placed into class intervals (bins) such as 0 ≤ x < 10. The class width is the difference between the upper and lower boundaries.

分组频数表:用于连续数据或大型离散数据集。数据被放入组距(区间),如 0 ≤ x < 10。组宽是上下边界之差。

Modal class is the class interval with the highest frequency – it is the ‘mode’ for grouped data.

众数组是频数最高的组距——即分组数据的“众数”所在组。

Class boundaries are the exact limits used to separate classes without gaps. For example, if intervals are 0–9 and 10–19, the boundary is 9.5.

组边界是用来无间隙分隔各组的精确界限。例如,若区间为 0–9 和 10–19,边界为 9.5。

Memory aid: ‘When data get large, group with a bin – the tallest bin wins.’

记忆要点:“数据一多就分组装箱——最高的箱子就是众数组。”


5. Graphical Representation: Bar Charts, Histograms & Frequency Polygons | 图形表示:条形图、直方图与频数多边形

Bar chart: displays categorical data with separate bars of equal width. Gaps between bars highlight that categories are distinct. The height shows frequency or frequency density.

条形图:用等宽分离的条形展示分类数据。条间间隙突显类别的独立性,高度表示频数或频率密度。

Histogram: used for continuous (or grouped discrete) data. No gaps between bars because the data scale is continuous. The area of each bar is proportional to frequency – so bar height = frequency density = frequency ÷ class width.

直方图:用于连续(或分组离散)数据。条间无间隙,因为数据尺度是连续的。每块的面积与频数成正比——因此条高 = 频率密度 = 频数 ÷ 组宽。

Frequency polygon: formed by joining the midpoints of the tops of histogram bars or by plotting midpoints against frequencies. Useful for comparing distributions.

频数多边形:连接直方图条顶中点,或绘制中点对应频数的点并连线而成。便于比较分布。

Quick check: ‘Bars apart → bar chart (categories); Bars joined → histogram (continuous).’

快速辨别:“条分开 → 条形图(分类数据);条紧挨 → 直方图(连续数据)。”


6. Cumulative Frequency & Box Plots | 累积频率与箱线图

Cumulative frequency is a running total of frequencies up to the end of each class. A cumulative frequency table and graph use the upper class boundary to plot points.

累积频率是截至每个组末的频数累计总和。累积频率表和图使用上组边界来描点。

Median (Q₂) is estimated from the cumulative frequency graph at the 50th percentile position (n/2). Lower quartile Q₁ at n/4, upper quartile Q₃ at 3n/4.

中位数(Q₂)从累积频率图上 50% 位置(n/2)估计。下四分位数 Q₁ 在 n/4 位置,上四分位数 Q₃ 在 3n/4 位置。

Interquartile range (IQR) = Q₃ – Q₁. It measures the spread of the middle 50% of the data and is not affected by outliers.

四分位距(IQR)= Q₃ – Q₁。衡量中间 50% 数据的分散程度,且不受异常值影响。

Box plot (box‑and‑whisker plot) displays minimum, Q₁, median, Q₃, and maximum. It gives a visual summary of symmetry, spread, and potential outliers.

箱线图(盒须图)展示最小值、Q₁、中位数、Q₃ 和最大值。它提供了对称性、离散度和潜在异常值的可视化摘要。

Memory: ‘CF graph for the middle family, box plot for the five‑number story.’

记忆:“累积频率图给出中位家族,箱线图讲述五数故事。”


7. Measures of Central Tendency: Mean, Median & Mode | 集中趋势度量:平均数、中位数与众数

Mean (often x̄ for a sample, μ for a population) is the average: sum of all values ÷ number of values. Sensitive to extreme values (outliers).

平均数(样本常记为 x̄,总体常记为 μ)即均值:所有数值之和 ÷ 数值个数。它对极端值(异常值)敏感。

Median is the middle value when data are ordered. It splits the data into two equal halves and is robust against outliers.

中位数是数据排序后位于中间的值。它将数据平分为两半,且对异常值稳健。

Mode is the value that occurs most often. A data set can have no mode, one mode (unimodal), or several modes. For grouped data, use the modal class.

众数是出现频率最高的值。数据集可无众数、单众数或多众数。对于分组数据,使用众数组。

Memory phrase: ‘Mean – the balancing point (affected by extremes); Median – the middle seat; Mode – the most popular.’

记忆口诀:“平均数——平衡点(受极端值影响);中位数——中间座位;众数——人气王。”


8. Measures of Dispersion: Range, IQR & Standard Deviation | 离散程度度量:极差、四分位距与标准差

Range = maximum value – minimum value. It gives the full width of the data but is heavily influenced by single outliers.

极差 = 最大值 – 最小值。给出数据的完整宽度,但极易受单个异常值影响。

Interquartile range (IQR) = Q₃ – Q₁. It covers the central half of the data and is resistant to extreme values. Often used with the median.

四分位距 (IQR) = Q₃ – Q₁。它覆盖数据的中间一半,对极端值不敏感,常与中位数搭配使用。

Variance measures the average squared distance from the mean. For a population, variance σ² = Σ(x – μ)²/N; for a sample, s² = Σ(x – x̄)²/(n–1).

方差衡量各数据点与均值之差的平方的平均值。总体方差 σ² = Σ(x – μ)²/N;样本方差 s² = Σ(x – x̄)²/(n–1)。

Standard deviation is the square root of variance. It has the same units as the original data. Usually denoted σ (population) or s (sample). A smaller SD means data are tightly packed around the mean.

标准差是方差的平方根,单位与原数据相同,通常记为 σ(总体)或 s(样本)。标准差越小,数据越紧密围绕均值分布。

Memory link: ‘Range = full stretch; IQR = middle half; SD = typical distance from the centre.’

联系记忆:“极差 = 全幅伸展;IQR = 中间一半;标准差 = 距离中心的典型距离。”


9. Scatter Diagrams & Correlation | 散点图与相关性

A scatter diagram (scatter graph) displays paired numerical data as points to reveal a relationship. If points trend upwards, there is positive correlation; downwards, negative correlation; no clear pattern suggests zero or no correlation.

散点图将成对数值数据以点的形式展示,以揭示关系。若点向上趋势,存在正相关;向下趋势为负相关;无明显模式则表明零相关或无相关。

Line of best fit is a straight line drawn through the points to model the relationship. It should have roughly equal numbers of points above and below the line.

最佳拟合线是一条穿过各点的直线,用于模拟关系。线上下方点数应大致相等。

Interpolation: predicting a value inside the range of the given data by reading from the line of best fit. Considered more reliable.

内插:在给定数据范围内,利用最佳拟合线预测一个值。被认为更可靠。

Extrapolation: extending the line beyond the data range to make a prediction. Riskier because the trend may not continue.

外推:将拟合线延伸到数据范围之外进行预测。风险更大,因为趋势可能不延续。

Be careful: correlation does not imply causation – two variables may rise together without one causing the other.

注意:相关不等于因果——两个变量可能同时上升,但并非一个导致另一个。


10. Probability Basics: Experiment, Outcome & Event | 概率基础:试验、结果与事件

A random experiment is a process with uncertain individual outcomes but a known set of all possible results. The set of all possible outcomes is the sample space (S).

随机试验是一个过程,单独结果不确定,但所有可能结果的集合已知。所有可能结果的集合称为样本空间(S)。

An outcome is a single result of an experiment. An event is a subset of the sample space, containing one or more outcomes. Events are equally likely if they have the same chance of occurring.

结果是试验的单个产物。事件是样本空间的一个子集,包含一个或多个结果。若各事件发生机会相同,则称为等可能事件。

The probability of event A, P(A), is a number between 0 and 1: P(A) = number of favourable outcomes / total number of possible outcomes, when outcomes are equally likely.

事件 A 的概率 P(A) 是 0 到 1 之间的一个数:当结果为等可能时,P(A) = 有利结果数 / 可能结果总数。

Complementary event A’ is ‘not A’; P(A’) = 1 – P(A). Mutually exclusive events cannot happen at the same time; then P(A or B) = P(A) + P(B).

互补事件 A’ 即“非 A”;P(A’) = 1 – P(A)。互斥事件不能同时发生;此时 P(A 或 B) = P(A) + P(B)。


11. Probability Diagrams & Expectation | 概率图与期望值

A Venn diagram uses overlapping circles to show relationships between events. The overlap (intersection) represents ‘A and B’. The whole region covered by circles is ‘A or B’.

韦恩图Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading