GCSE WJEC Statistics: Vocabulary and Terminology Quick Recall Guide | GCSE WJEC 统计:词汇术语速记指南

📚 GCSE WJEC Statistics: Vocabulary and Terminology Quick Recall Guide | GCSE WJEC 统计:词汇术语速记指南

Mastering the specialised language of Statistics is half the battle in your GCSE WJEC exam. This bilingual guide breaks down the essential vocabulary with clear definitions, contrasts, and memory shortcuts, so you can navigate data, probability, and inference with confidence.

掌握统计学的专门语言是攻克 GCSE WJEC 考试的一半。这份双语指南通过清晰的定义、对比和记忆捷径,帮你拆解核心词汇,让你从容应对数据、概率和推断。

1. Types of Data and Collection | 数据类型与收集

Data are categorised as qualitative (categorical) or quantitative (numerical). Qualitative data describe attributes, like hair colour or favourite genre, while quantitative data involve numerical measurements, such as test scores or temperatures.

数据分为定性(类别)数据和定量(数值)数据。定性数据描述属性,如发色或最喜欢的类型;定量数据涉及数值测量,如考试成绩或温度。

Quantitative data can be discrete or continuous. Discrete data can only take distinct, separate values (e.g. number of books, shoe size in half‑sizes). Continuous data can assume any value within a given range (e.g. height, weight, time).

定量数据可以是离散的或连续的。离散数据只能取分离的特定值(例如书本数量、以半码记的鞋码);连续数据可以在给定范围内取任意值(例如身高、体重、时间)。

Data Type Key Feature Examples
Qualitative Non‑numerical labels Colour, gender, type of pet
Quantitative discrete Countable, gaps between values Number of students, goals scored
Quantitative continuous Measurable, any value in an interval Length, volume, temperature

Primary data are collected first‑hand by the researcher (e.g. through a survey or experiment). Secondary data are obtained from existing sources such as government reports or websites. Primary data are usually more tailored and reliable but costly; secondary data are cheaper but may be outdated or biased.

主要数据由研究者亲自收集(例如通过调查或实验)。次要数据则来自政府报告或网站等已有来源。主要数据通常更贴合需求且更可靠,但成本高;次要数据成本低,但可能过时或存在偏差。


2. Sampling Methods | 抽样方法

Simple random sampling gives every member of the population an equal chance of being selected. It is like drawing names from a hat: unbiased but difficult with a large population.

简单随机抽样使总体中每个成员被抽中的概率相等,如同从帽子中抽名字:无偏性但用于大总体时操作困难。

Stratified sampling divides the population into distinct strata (e.g. age groups) and then takes a random sample from each stratum in proportion to its size. This guarantees representation of key subgroups.

分层抽样将总体分成不同层(例如年龄组),然后按各层在总体中的比例随机抽样。这能保证关键子群体的代表性。

Systematic sampling selects every kth item after a random start. It is simple to implement but can introduce periodicity bias if the list has a hidden pattern.

系统抽样随机确定起点后每隔 k 个抽取一个。它易于实施,但如果名单存在隐藏模式,可能引入周期性偏差。

Quota sampling is a non‑probability method where interviewers fill quotas for certain characteristics (e.g. 50 men and 50 women). It is quick and cheap but prone to interviewer bias because the selection within quotas is not random.

配额抽样是一种非概率方法,调查员按特定特征(如 50 名男性和 50 名女性)完成配额。它快速且成本低,但由于配额内的选择并非随机,容易产生调查员偏差。

Memory aid: remember the four main probability‑based methods as SSSQ – Simple, Stratified, Systematic, Quota.

记忆小窍门:将四种主要概率/非概率方法记为 SSSQ – 简单、分层、系统、配额。


3. Presenting Data: Charts and Graphs | 数据展示:图表

A bar chart displays categorical data using bars with gaps between them. A compound or dual bar chart can compare multiple categories. In contrast, a histogram displays grouped continuous data with touching bars; the area of each bar is proportional to the frequency.

条形图用带有间隙的矩形条展示分类数据。复合或双条形图可以比较多组类别。与之不同,直方图展示分组的连续数据,条与条之间无间隙,条形的面积与频数成比例。

Frequency polygons are drawn by joining the mid‑points of the tops of histogram bars with straight lines. They are useful for showing the shape of a distribution and for comparing two data sets on the same graph.

频率多边形是通过连接直方图各矩形顶端中点而成的折线图,适合展示分布形状并在同一图中比较两个数据集。

A cumulative frequency curve (ogive) plots the running total of frequencies against the upper class boundary. It lets you estimate the median, quartiles and percentiles directly from the graph.

累计频率曲线(折形图)将累计频数对上组上界作图,能直接从图中估算中位数、四分位数和百分位数。

A box plot (box‑and‑whisker diagram) displays the five‑number summary: minimum, lower quartile Q₁, median Q₂, upper quartile Q₃, and maximum. It highlights the spread and skewness of the data at a glance.

箱线图(箱形图)展示五数概括:最小值、下四分位数 Q₁、中位数 Q₂、上四分位数 Q₃ 和最大值,能一目了然地呈现数据的离散程度与偏态。

A pie chart is used for qualitative data to show proportions of a whole. Each slice represents a category, and the angle is proportional to the frequency.

饼图用于定性数据,展示各部分占整体的比例。每个扇形代表一个类别,其角度与频数成比例。


4. Measures of Central Tendency | 集中趋势的度量

The mean is calculated by adding all values and dividing by the number of values. It is sensitive to extreme values (outliers).

平均数由所有数值之和除以数值个数计算得出,它对极端值(异常值)敏感。

Mean x̄ = Σx / n

The median is the middle value when data are ordered. If there is an even number of observations, it is the mean of the two middle values. The median is robust against outliers.

中位数是数据排序后位于中间的数值。如果观察值为偶数个,则是中间两个数的平均数。中位数对异常值具有稳健性。

The mode is the most frequently occurring value. A set of data can have more than one mode (bimodal) or no mode at all.

众数是出现频率最高的值。一组数据可能有一个以上的众数(双峰)或根本就没有众数。

Memory trick: think of the ‘3 M’s’ – Mean (the average), Median (the middle), Mode (the most).

记忆诀窍:记住“3 M”——Mean(均值)、Median(中位)、Mode(众数)。


5. Measures of Spread | 离散程度的度量

The range is the difference between the maximum and minimum values. It is easy to compute but affected heavily by a single outlier.

极差是最大值与最小值之差,计算简单但极易受单个异常值影响。

The interquartile range (IQR) is the difference between the upper quartile Q₃ and the lower quartile Q₁. IQR = Q₃ − Q₁. It measures the spread of the middle 50% and is resistant to outliers.

四分位距(IQR)是上四分位数 Q₃ 与下四分位数 Q₁ 之差,IQR = Q₃ − Q₁。它衡量中间 50% 数据的离散程度且能抵抗异常值的影响。

Quartiles are often found from a cumulative frequency graph. The lower quartile is the 25th percentile, the median is the 50th percentile, and the upper quartile is the 75th percentile.

四分位数常通过累计频率图求得。下四分位数即第 25 百分位数,中位数是第 50 百分位数,上四分位数是第 75 百分位数。

Standard deviation (σ or s) describes the average distance of data points from the mean. A small standard deviation means data cluster tightly around the mean; a large one indicates they are more spread out.

标准差(σ 或 s)描述数据点与平均数之间的平均距离。标准差小表示数据紧密聚集在均值附近;标准差大则表明数据更为分散。

Quick recall: RIQ – Range, Interquartile range, and then the advanced Standard deviation.

快速记忆:RIQ – Range(极差)、Interquartile range(四分位距),再进阶是 Standard deviation(标准差)。


6. Probability Terminology | 概率术语

An experiment is any repeatable process that produces outcomes. The set of all possible outcomes is the sample space. An event is a subset of the sample space.

试验是任何能产生结果的可重复过程。所有可能结果的集合称为样本空间。一个事件是样本空间的一个子集。

The probability of an event A is written as P(A) and satisfies 0 ≤ P(A) ≤ 1. Probability can be expressed as a fraction, decimal, or percentage.

事件 A 的概率记作 P(A),且满足 0 ≤ P(A) ≤ 1。概率可以用分数、小数或百分数表示。

Two events are mutually exclusive if they cannot occur at the same time. For mutually exclusive events A and B, P(A or B) = P(A) + P(B).

两个事件互斥,如果它们不能同时发生。对于互斥事件 A 和 B,有 P(A 或 B) = P(A) + P(B)。

Two events are independent if the occurrence of one does not affect the probability of the other. For independent events, P(A and B) = P(A) × P(B).

两个事件独立,如果一个事件的发生不影响另一个事件发生的概率。对于独立事件,P(A 且 B) = P(A) × P(B)。

Relative frequency is the ratio of the number of times an event occurs to the total number of trials. It can be used to estimate an unknown probability from experiment data.

相对频率是事件发生次数与总试验次数之比。它可用于从实验数据估算未知概率。

Conditional probability, P(A given B), is the probability that event A occurs given that B has already occurred. It is calculated as P(A and B) / P(B).

条件概率 P(A 给定 B) 是已知事件 B 已经发生的情况下事件 A 发生的概率,计算公式为 P(A 且 B) / P(B)。

Mnemonics: ‘MInT’– Mutually exclusive, Independent, and Total probability rules.

助记词:“MInT”——Mutually exclusive(互斥)、Independent(独立)、Total probability(全概率)。


7. Correlation and Regression | 相关性与回归

A scatter graph (scatter plot) displays the relationship between two numerical variables. Correlation describes the direction and strength of a linear relationship.

散点图展示两个数值变量之间的关系。相关描述线性关系的方向和强度。

Positive correlation: as one variable increases, the other also tends to increase. Negative correlation: as one variable increases, the other tends to decrease. No correlation means no clear linear pattern.

正相关:一个变量增加,另一个也倾向于增加。负相关:一个变量增加,另一个倾向于减少。无相关意味着没有明显的线性模式。

The correlation coefficient, often denoted r, measures the strength of linear correlation on a scale from -1 to +1. Values close to ±1 indicate strong correlation; values near 0 indicate weak or no linear correlation.

相关系数通常记作 r,衡量线性相关的强度,取值范围从 -1 到 +1。接近 ±1 表示强相关,接近 0 表示弱相关或没有线性相关。

A line of best fit (regression line) is drawn through the scatter plot to model the trend. It can be used for interpolation (predicting within the range of data) but extrapolation (predicting outside the range) can be unreliable.

最佳拟合线(回归线)穿过散点图以模拟趋势。它可用于插值(在数据范围内预测),但外推(在范围外预测)可能不可靠。

When interpreting correlation, remember: correlation does not imply causation. A hidden third variable may be involved.

解读相关时请牢记:相关不代表因果。可能存在隐藏的第三变量。


8. Distribution Shapes and Skewness | 分布形状与偏度

A symmetric distribution has mirror‑image sides. The normal distribution is a bell‑shaped symmetric curve where the mean, median and mode are equal.

对称分布两侧呈镜像。正态分布是钟形对称曲线,平均数、中位数和众数相等。

Skewness measures asymmetry. A distribution is positively skewed when the right tail is longer; the mass of data is concentrated on the left. For positive skew, mean > median > mode.

偏度衡量不对称性。当右尾更长时,分布为正偏态(右偏),数据主体集中在左侧。在正偏态中,平均数 > 中位数 > 众数。

A distribution is negatively skewed when the left tail is longer; the mean < median < mode. To remember, the 'tail drags the mean' in its direction.

当左尾更长时为负偏态(左偏),此时平均数 < 中位数 < 众数。记忆方法:“尾巴将平均数拖向自己的方向”。

Skew Type Tail Direction Order of M’s
Positive (right) skew Long tail to the right Mean > Median > Mode
Negative (left) skew Long tail to the left Mean < Median < Mode
Symmetric (no skew) Tails are roughly equal Mean ≈ Median ≈ Mode

Box plots quickly reveal skewness: if the median is closer to the left whisker, the data are positively skewed; if it is closer to the right whisker, the data are negatively skewed.

箱线图可快速显示

Published by TutorHao | GCSE 统计 Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading