GCSE WJEC Statistics: Core Knowledge Compilation | GCSE WJEC 统计:核心知识点梳理

📚 GCSE WJEC Statistics: Core Knowledge Compilation | GCSE WJEC 统计:核心知识点梳理

GCSE WJEC Statistics equips you with the skills to collect, analyse, and interpret data in a wide variety of contexts. This article distils the essential knowledge areas – from types of data and sampling to probability and time series – into a clear, exam-focused revision guide. Each section provides a bilingual breakdown to reinforce both English and Chinese terminology, helping you build confidence for your assessments.

GCSE WJEC 统计学帮助你掌握在各类情境中收集、分析和解读数据的技能。本文提炼了从数据类型、抽样到概率和时间序列等核心知识领域,形成一份清晰、紧扣考试的复习指南。每个部分都提供中英双语拆解,巩固英文和中文术语,助你自信应对评估。

1. Types of Data | 数据类型

In WJEC Statistics, data is first classified as qualitative or quantitative. Qualitative (categorical) data describes qualities – like eye colour or vehicle type – and can be nominal (no natural order) or ordinal (ordered categories, e.g. satisfaction ratings).

在WJEC统计学中,数据首先分为定性数据和定量数据。定性(分类)数据描述品质——如眼睛颜色或车辆类型——可以是名义数据(无自然顺序)或有序数据(有序类别,如满意度评分)。

Quantitative data involves numbers and is either discrete or continuous. Discrete data can only take specific, separate values – for example, the number of students in a class or shoe sizes (where halves and wholes are allowed but not any value). Continuous data can take any value within a given range, such as height, weight, or time.

定量数据涉及数字,分为离散数据和连续数据。离散数据只能取特定的、分离的值——例如班级学生人数或鞋码(允许半码和整码,但不是任意值)。连续数据可以在给定范围内取任意值,如身高、体重或时间。

Knowing the data type is crucial because it determines which diagrams and statistical measures are appropriate. For instance, you cannot calculate a meaningful mean for nominal data, and a pie chart is unsuitable for continuous raw data unless grouped.

了解数据类型至关重要,因为它决定了适合使用哪些图表和统计量。例如,对于名义数据,计算有意义的均值是不可能的;除非对连续原始数据进行分组,否则饼图也不适用。

WJEC exam questions often start by asking you to identify whether a dataset is qualitative/quantitative and discrete/continuous, so make sure you can justify your classification.

WJEC的考题经常首先要求你判断一个数据集是定性还是定量、离散还是连续,所以务必能够为你的分类给出合理理由。


2. Data Collection and Sampling | 数据收集与抽样

Primary data is collected firsthand for a specific purpose – through experiments, surveys, or observations. Secondary data is obtained from existing sources, such as government publications or databases. You need to discuss the advantages and disadvantages of each: primary data is more tailored but time-consuming; secondary data is quicker and often cheaper but may not be perfectly suited to your investigation.

一手数据是为特定目的直接收集的——通过实验、调查或观察获得。二手数据来自现有来源,如政府出版物或数据库。你需要讨论每种方法的优缺点:一手数据更贴合需求但耗时;二手数据更快、通常更便宜,但可能不完全适合你的研究。

Sampling methods are essential to learn. Random sampling gives every member of the population an equal chance of selection, reducing bias. Stratified sampling divides the population into distinct groups (strata) and samples proportionally from each, ensuring fair representation. Systematic sampling selects every nth individual from a list, which is easy to implement but can introduce bias if there is a hidden pattern. Cluster sampling and quota sampling may also appear, but WJEC focuses on the first three.

抽样方法是必须掌握的内容。随机抽样让总体中每个成员有均等的被选中的机会,从而减少偏差。分层抽样将总体划分为不同组(层),并按照比例从每一层抽样,确保公平代表。系统抽样从名单中每隔 n 个抽取一个,易于实施,但如果存在隐藏的模式则可能引入偏差。整群抽样和配额抽样也可能出现,但WJEC重点在前三种。

Understanding bias is critical. Selection bias, non-response bias, and measurement bias can all distort findings. When evaluating a sampling technique in an exam, comment on both its practicality and its potential for bias.

理解偏差至关重要。选择偏差、无应答偏差和测量偏差都可能扭曲研究结果。在考试中评价一种抽样技术时,既要评论其实用性,也要评论其可能产生偏差的风险。


3. Presenting Data: Charts and Graphs | 数据呈现:图表

Choosing the correct diagram depends on your data type and what you want to show. Bar charts are used for categorical or discrete data, with gaps between bars. Pie charts show proportions of a whole and work best for nominal categorical data with relatively few categories.

选择正确的图表取决于你的数据类型和你想展示的信息。条形图用于分类数据或离散数据,条形之间有间隙。饼图展示整体中的比例,最适合类别较少的名义分类数据。

For continuous data, histograms are essential. In a histogram, frequency is represented by the area of the bar, not its height, so when class widths are unequal, you must use frequency density = frequency / class width. This is a key WJEC skill: calculating frequency densities and plotting correct histograms.

对于连续数据,直方图是必须的。在直方图中,频数由条形的面积而非高度表示,因此当组距不相等时,必须使用频数密度 = 频数 / 组距。这是WJEC的一项关键技能:计算频数密度并绘制正确的直方图。

Frequency polygons are created by joining the midpoints of the tops of histogram bars with straight lines. They are useful for comparing two distributions because multiple polygons can be drawn on the same axes. Cumulative frequency curves (ogives) are also vital for finding medians and quartiles.

频数多边形是通过用直线连接直方图各条形顶端中点绘制的。它们对于比较两个分布很有用,因为可以在同一坐标系中绘制多条频数多边形。累积频数曲线(肩形图)对于求中位数和四分位数也至关重要。

Stem-and-leaf diagrams keep the original data values visible while showing the shape of the distribution. Always include a key and order the leaves. Back-to-back stem-and-leaf diagrams are excellent for comparing two related datasets.

茎叶图在展示数据分布形状的同时保留了原始数值。始终包含一个图例并使叶有序排列。背靠背茎叶图非常适合比较两个相关的数据集。


4. Measures of Central Tendency | 集中趋势度量

The three common measures are the mean, median, and mode. The mean (x̄) is calculated as Σx / n for raw data or Σfx / Σf for grouped data. It uses all values, which makes it sensitive to outliers.

常见的三种度量是均值、中位数和众数。对于原始数据,均值 (x̄) 计算为 Σx / n;对于分组数据,计算为 Σfx / Σf。它使用了所有数值,因此对异常值敏感。

The median is the middle value when data is ordered. For n values, its position is (n + 1)/2. In grouped data, linear interpolation is required within the median class interval. The median is unaffected by extreme values, so it’s better for skewed distributions or datasets with outliers.

中位数是数据排序后的中间值。对于 n 个值,其位置为 (n + 1)/2。在分组数据中,需要在中位数所在组距内进行线性插值。中位数不受极端值影响,因此对于偏态分布或有异常值的数据集更为适用。

The mode is simply the most frequent value or class. A dataset can have more than one mode (bimodal or multimodal) or no mode at all. For grouped data, the modal class is the class with the highest frequency density, not necessarily the highest frequency.

众数只是出现频率最高的值或组。一个数据集可以有多于一个众数(双众数或多众数),也可以没有众数。对于分组数据,众数组是频数密度最高的组,而不一定是频数最高的组。

WJEC questions often ask you to choose the most appropriate average for a given context – for example, the median for house prices where a few luxury homes would inflate the mean – and to justify your choice.

WJEC的题目常让你为给定情境选择最合适的平均数——例如,对于房价,少数豪宅会拉高均值,因此中位数更适合——并要求你证明所选的理由。


5. Measures of Dispersion | 离散度量

Dispersion tells us how spread out the data is. The range (max – min) is the simplest measure but only uses two values, so it can be misleading. The interquartile range (IQR = Q₃ – Q₁) captures the spread of the middle 50% of data, making it robust against outliers.

离散度告诉我们数据的分散程度。极差(最大值 – 最小值)是最简单的度量,但只使用了两个值,因此可能会产生误导。四分位距(IQR = Q₃ – Q₁)捕捉了中间50%数据的分散度,使其对异常值具有稳健性。

Quartiles are found from a list or via a cumulative frequency graph: Q₁ is at the 25th percentile, Q₂ the median (50th percentile), Q₃ the 75th percentile. From a stem-and-leaf or ordered set, use the (n+1)/4 and 3(n+1)/4 positions (or a consistent method specified in the exam).

四分位数可以从列表中或通过累积频数图求得:Q₁ 位于第 25 百分位数,Q₂ 为中位数(第 50 百分位数),Q₃ 为第 75 百分位数。对于茎叶图或有序数据集,使用 (n+1)/4 和 3(n+1)/4 位置(或考试指定的统一方法)。

Standard deviation measures the typical distance of data points from the mean. For a population, σ = √(Σ(x – μ)² / n). For a sample, WJEC typically uses s = √(Σ(x – x̄)² / (n – 1)), which distinguishes it from variance. You must be able to calculate it using a table of values or your calculator’s statistics functions and interpret it in context: a larger standard deviation means more variability.

标准差衡量数据点与均值的典型距离。对于总体,σ = √(Σ(x – μ)² / n)。对于样本,WJEC通常使用 s = √(Σ(x – x̄)² / (n – 1)),这与方差相区别。你必须能够使用数值表或计算器的统计功能来计算标准差,并结合情境解释:标准差越大,意味着变异性越大。

Comparing dispersion often involves using the coefficient of variation when the means differ significantly, though basic comparison focuses on standard deviation and IQR.

当均值差异较大时,比较离散度常会用到变异系数,但基本比较的重点是标准差和四分位距。


6. Cumulative Frequency and Box Plots | 累积频数与箱形图

A cumulative frequency table adds frequencies successively. Plotting the upper class boundary against cumulative frequency gives the S-shaped ogive. From this graph, you can estimate the median, quartiles, and percentiles without raw data.

累积频数表是将频数逐次累加。将上限组界与累积频数描点作图,就得到 S 形肩形图。从该图中,你可以在没有原始数据的情况下估算中位数、四分位数和百分位数。

To construct a box plot (box-and-whisker diagram), you need the five-number summary: minimum, Q₁, median, Q₃, and maximum. Outliers are often defined as values more than 1.5 × IQR below Q₁ or above Q₃. In WJEC, outliers may be plotted as separate points, and the whiskers extend to the most extreme non-outlier values.

要构建箱形图(箱线图),你需要五数汇总:最小值、Q₁、中位数、Q₃ 和最大值。异常值通常定义为低于 Q₁ – 1.5 × IQR 或高于 Q₃ + 1.5 × IQR 的值。在WJEC中,异常值可以绘制为单独的点,而触须延伸到最极端的非异常值。

Box plots allow quick visual comparison of distributions – their central tendency, spread, and skewness. Two box plots side by side make it easy to comment on overlaps and differences in median and IQR.

箱形图可以快速直观地比较分布——它们的集中趋势、分散度和偏斜程度。并排的两个箱形图便于对中位数和四分位距的重叠与差异进行评述。

When interpreting an ogive and a box plot together, remember that the steeper sections of the ogive indicate high frequency density, while the box plot simplifies the whole picture into a five-number summary that is easy to communicate.

在同时解读肩形图和箱形图时,请记住肩形图中较陡的部分表示频率密度高,而箱形图则将整体信息简化为易于传达的五数汇总。


7. Probability Basics | 概率基础

Probability is a measure of how likely an event is, on a scale from 0 (impossible) to 1 (certain). The probability of an event A is P(A) = number of favourable outcomes / total number of equally likely outcomes. The complement rule states P(not A) = 1 – P(A).

概率是衡量事件发生可能性的度量,范围从 0(不可能)到 1(必然)。事件 A 的概率为 P(A) = 有利结果数 / 等可能的结果总数。互补规则为 P(非 A) = 1 – P(A)。

For mutually exclusive events (they cannot happen at the same time), P(A or B) = P(A) + P(B). For independent events, P(A and B) = P(A) × P(B). WJEC questions will test whether you can identify when events are independent or mutually exclusive.

对于互斥事件(它们不能同时发生):P(A 或 B) = P(A) + P(B)。对于独立事件:P(A 和 B) = P(A) × P(B)。WJEC的题目会测试你是否能判断事件何时独立或互斥。

Relative frequency and expected frequency also feature. Expected frequency = n × P(A). If an experiment is repeated many times, the relative frequency of an event tends to its theoretical probability – this is the long-run relative frequency interpretation of probability.

相对频率和期望频率也会出现。期望频率 = n × P(A)。如果实验重复多次,事件的相对频率会趋近于其理论概率——这就是概率的长期相对频率解释。

Sample space diagrams and two-way tables help you list all possible outcomes systematically, ensuring no event is missed when calculating combined probabilities.

样本空间图和双向表帮助你有条理地列出所有可能结果,确保在计算组合概率时不会遗漏任何事件。


8. Probability Trees and Conditional Probability | 概率树与条件概率

Tree diagrams are used to model two or more successive events. At each branch, probabilities sum to 1. Multiply along branches for combined events and add relevant final probabilities for alternative outcomes.

树形图用于模拟两个或多个相继的事件。每个分支上的概率之和为 1。沿着分支相乘得到组合事件的概率,然后将相关最终结果的概率相加。

When events are not independent, the probabilities on the second set of branches change depending on the first outcome. These are conditional probabilities, written as P(B|A) – the probability of B given that A has occurred.

当事件不独立时,第二组分枝上的概率会根据第一个结果而改变。这些就是条件概率,记作 P(B|A)——在 A 已发生的条件下 B 的概率。

The multiplication rule for dependent events is P(A and B) = P(A) × P(B|A). WJEC exams often set problems involving picking objects from a bag without replacement – a classic context where probabilities change after each selection.

相关事件的乘法规则是 P(A 和 B) = P(A) × P(B|A)。WJEC考试常设置从袋子中无放回地抽取物品的问题——这是每次选择后概率都会变化的经典情境。

To calculate conditional probability from a two-way table or a Venn diagram, use P(A|B) = P(A and B) / P(B). You must be comfortable rearranging this relationship and solving problems that ask for a different probability.

若要通过双向表或韦恩图计算条件概率,使用 P(A|B) = P(A 和 B) / P(B)。你必须能熟练地对该关系进行变形,并解决要求计算不同概率的问题。

Check your tree diagrams carefully: all branch pairs must sum to 1, and all end-node probabilities should sum to 1 as a final check.

仔细检查你的树形图:所有分支对的概率之和必须为 1,所有末端节点的概率之和也应作为最终检查等于 1。


9. Scatter Diagrams and Correlation | 散点图与相关性

A scatter diagram (scatter plot) shows the relationship between two numerical variables. Each point represents a pair (x, y). The pattern reveals correlation: positive (as x increases, y tends to increase), negative (as x increases, y tends to decrease), or no correlation.

散点图(散点图)展示两个数值变量之间的关系。每个点代表一对 (x, y)。点的分布模式显示出相关性:正相关(x 增加,y 趋向增加)、负相关(x 增加,y 趋向减少)或无相关。

Correlation does not imply causation. An observed association might be due to a third, lurking variable. WJEC expects you to be cautious when interpreting results.

相关并不意味着因果。观察到的关联可能是由于第三个潜在变量造成的。WJEC要求你在解释结果时保持谨慎。

The strength of correlation can be described using terms like strong, moderate, or weak. You can also draw a line of best fit by eye, passing through the mean point (x̄, ȳ). This line is used to make predictions (interpolation) but extending beyond the data range (extrapolation) may be unreliable.

相关性的强度可以用强、中等或弱等术语描述。你还可以通过目测绘制一条经过均值点 (x̄, ȳ) 的最佳拟合线。这条线可用于进行预测(内插法),但扩展至数据范围之外(外推法)可能不可靠。

Spearman’s rank correlation coefficient is occasionally introduced at GCSE level as a numerical measure of correlation for ranked data, but in WJEC Statistics, the focus is mainly on qualitative interpretation and the line of best fit.

斯皮尔曼等级相关系数偶尔在GCSE层级引入,作为排序数据相关性的数值度量,但在WJEC统计学中,重点主要在于定性解释和最佳拟合线。

When describing scatter diagrams, always comment on the direction, strength, and any outliers that do not follow the general trend.

在描述散点图时,始终要评论其方向、强度以及任何不遵循总体趋势的异常点。


10. Time Series and Moving Averages | 时间序列与移动平均

A time series is a sequence of data points recorded at successive points in time, usually at regular intervals. Plotting the data reveals trends (long-term movement), seasonal variations (regular patterns within fixed periods), and random fluctuations.

时间序列是按连续时间点(通常等间隔)记录的一系列数据点。将数据绘制成图可揭示趋势(长期变动)、季节性变化(固定周期内的规律模式)和随机波动。

Moving averages smooth out short-term fluctuations to reveal the underlying trend. For a 4-point moving average, you calculate the mean of the first four values, then drop the first and include the fifth, and so on. These averages are plotted at the midpoint of the time intervals they cover.

移动平均能平滑短期波动,以揭示潜在的趋势。对于四点移动平均,你先计算前四个值的均值,然后去掉第一个值并纳入第五个值,依次类推。这些平均值绘制在所覆盖时间区间的中点上。

In WJEC, you may need to find the moving average and then calculate seasonal variation as the difference (actual minus trend). Seasonal effects are assumed to average out to zero over a complete cycle, so adjustments can be made for forecasting.

在WJEC中,你可能需要求出移动平均,然后计算季节性变化,即实际值减去趋势值的差。假定一个完整周期内的季节性效应平均为零,因此可进行调整以用于预测。

Time series analysis helps in understanding past behaviour and making informed predictions. Remember that forecasts become less reliable the further ahead they go, as external factors may change.

时间序列分析有助于理解过去的行为并做出有根据的预测。请记住,预测的前瞻期越长,可靠性越低,因为外部因素可能发生变化。


11. Comparing Data Sets | 数据集比较

Comparison questions are common. You must use measures of central tendency (mean, median) and spread (range, IQR, standard deviation) together. Commenting on just one measure is insufficient – a full comparison requires at least an average and a measure of dispersion, linked to context.

比较类题目很常见。你必须同时使用集中趋势度量(均值、中位数)和离散度量(极差、IQR、标准差)。仅评论一种度量是不够的——完整的比较至少需要一个平均数和一个离散度量,并结合情境。

When two datasets have similar means but different standard deviations, the one with the smaller standard deviation is more consistent. If medians differ noticeably, you can infer that one group tends to have higher values.

当两个数据集的均值相似但标准差不同时,标准差较小的那个更具一致性。如果中位数差异显著,你可以推断某一组的值往往更高。

Use diagrams to support your comparison – side-by-side box plots, back-to-back stem-and-leaf plots, or dual frequency polygons. Referring to the shape of a distribution (symmetry, skewness) adds depth to your analysis.

使用图表来支持你的比较——并排箱形图、背靠背茎叶图或双频数多边形。提及分布的形状(对称性、偏斜度)能为你的分析增色。

In exam answers, always state the values you are comparing (e.g., “The median mark for girls was 72, compared to 65 for boys, suggesting girls performed better on average”), and explain what the comparison means in the context of the problem.

在考试答案中,一定要陈述你所比较的数值(例如,“女生的中位数分数为 72,而男生为 65,表明女生平均表现更好”),并解释该比较在问题情境下的含义。

Finally, always consider whether a difference is meaningful. A small difference in means might not be significant if the standard deviations are large and the sample sizes small.

最后,始终考虑差异是否有意义。如果标准差很大且样本量很小,均值的微小差异可能并不显著。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version