📚 Year 10 WJEC Statistics: Core Knowledge Review | Year 10 WJEC 统计:核心知识点梳理
This article provides a comprehensive summary of the key topics covered in Year 10 of the WJEC GCSE Statistics course. From types of data to probability and time series, we review the essential concepts, techniques and formulas you need to master. Each section is presented in both English and Chinese to support bilingual learners and to help you build a solid foundation for the exam.
本文全面梳理了 WJEC GCSE 统计课程 Year 10 的核心知识点,涵盖数据类型、概率、时间序列等重要内容。每个小节都采用中英双语对照的形式,帮助双语学生理解并巩固考试必备的基础概念、方法和公式。
1. Types of Data | 数据类型
Data can be classified as qualitative (categorical) or quantitative (numerical). Qualitative data describe qualities, such as eye colour or car brands. Quantitative data involve numbers and can be further divided into discrete data (countable, e.g. number of siblings) and continuous data (measurable, e.g. height in cm). Understanding the data type is crucial because it determines which statistical diagrams and measures can be used.
数据可分为定性(分类)数据和定量(数值)数据。定性数据描述性质,如眼睛颜色或汽车品牌。定量数据涉及数字,并进一步分为离散数据(可数,如兄弟姐妹数量)和连续数据(可测量,如以厘米为单位的身高)。了解数据类型至关重要,因为它决定了可以使用哪些统计图表和度量指标。
We also distinguish between primary data (collected by the researcher for a specific purpose) and secondary data (obtained from existing sources, such as government publications or the internet). Primary data offer better control over accuracy, while secondary data are quicker and cheaper to obtain but may be less reliable.
我们还区分一手数据(研究者为特定目的自行收集)和二手数据(从现有来源获取,如政府出版物或互联网)。一手数据能更好地控制准确性,而二手数据获取更快、成本更低,但可能可靠性较差。
2. Sampling Methods | 抽样方法
A population is the entire group we are interested in, while a sample is a subset of that population. Sampling is necessary because surveying a whole population is often impractical. Random sampling gives each member an equal chance of being selected, reducing bias. Common random methods include simple random sampling (e.g. using a random number generator) and stratified sampling.
总体是我们感兴趣的整个群体,样本是其中的一个子集。抽样是必要的,因为普查整个总体通常不切实际。随机抽样让每个成员被抽中的概率相等,从而减少偏差。常见的随机方法有简单随机抽样(如使用随机数生成器)和分层抽样。
Stratified sampling divides the population into mutually exclusive groups (strata) based on a characteristic, then randomly selects from each stratum in proportion to its size. This ensures representation of all key subgroups. Non-random methods, such as convenience sampling or quota sampling, are easier but likely to introduce bias, so their results must be interpreted with caution.
分层抽样根据某一特征将总体分成互斥的组(层),然后按各层的大小比例从中随机抽取样本。这保证了所有关键小组都被代表。非随机方法,如便利抽样或配额抽样,操作更方便但容易引入偏差,因此对其结果的解释必须谨慎。
3. Charts and Diagrams | 图表与图示
Selecting the right diagram depends on the data type. For qualitative data, we use bar charts, pie charts and pictograms. For discrete quantitative data, vertical line charts and bar charts are suitable. For continuous data, we rely on histograms, frequency polygons and cumulative frequency curves.
选择合适的图表取决于数据类型。对于定性数据,我们使用条形图、饼图和象形图。对于离散定量数据,适用垂线图和条形图。对于连续数据,我们依赖直方图、频数多边形和累积频数曲线。
In a histogram, the frequency is represented by the area of the bar, so the frequency density (frequency ÷ class width) is plotted on the vertical axis. Stem-and-leaf diagrams keep the original data values while showing the shape of the distribution. Box plots (box-and-whisker plots) display the minimum, lower quartile, median, upper quartile and maximum, making them ideal for comparing distributions and identifying outliers.
在直方图中,频数用柱形面积表示,因此纵轴是频率密度(频数 ÷ 组距)。茎叶图在保留原始数据值的同时显示分布形状。箱线图显示最小值、下四分位数、中位数、上四分位数和最大值,非常适合比较不同分布和识别异常值。
| Data Type | Recommended Charts |
|---|---|
| Qualitative | Bar chart, pie chart, pictogram |
| Discrete | Vertical line chart, bar chart, stem-and-leaf |
| Continuous | Histogram, frequency polygon, cumulative frequency curve, box plot |
4. Measures of Central Tendency | 集中趋势的度量
The three main measures of average are the mean, median and mode. The mean (x̄) is calculated by summing all data values and dividing by the number of items. It uses every value but is sensitive to outliers. The median is the middle value when data are ordered; it is unaffected by extreme values and is therefore preferred for skewed distributions. The mode is the most frequent value and can be used for both numerical and categorical data.
三种主要的平均数度量指标是均值、中位数和众数。均值(x̄)通过将所有数据值相加再除以数据个数计算得出。它使用了每一个值,但容易受异常值影响。中位数是数据排序后的中间值;它不受极端值影响,因此更适合偏态分布。众数是出现频率最高的值,可用于数值数据和分类数据。
For grouped frequency tables, the mean is estimated using midpoints of classes. The median class can be found from cumulative frequency, and the modal class is the one with the highest frequency density in a histogram. Choosing the most appropriate average is a key skill: for symmetrical data without outliers, the mean is best; for skewed data or data with outliers, the median is more representative.
对于分组频数表,均值使用组中值来估算。中位数所在组可通过累积频数找到,众数所在组是直方图中频率密度最高的组。选择合适的平均数是关键技能:对于对称且没有异常值的数据,均值最佳;对于偏态或含有异常值的数据,中位数更具代表性。
5. Measures of Spread | 离散程度的度量
Spread tells us how consistent or varied the data are. The simplest measure is the range (maximum – minimum), but it only uses two extreme values and can be distorted by outliers. The interquartile range (IQR = upper quartile – lower quartile) covers the middle 50% of the data and is resistant to extreme values, making it a robust measure of spread.
离散程度告诉我们数据的一致性如何,变化有多大。最简单的度量是全距(最大值 – 最小值),但它只用到两个极端值,容易被异常值扭曲。四分位距(IQR = 上四分位数 – 下四分位数)覆盖中间50%的数据,且不受极端值影响,因此是一种稳健的离散度指标。
Standard deviation measures the average distance of data values from the mean. For a sample, we often use the formula with n-1 (the sample standard deviation, s). A low standard deviation indicates that the data cluster closely around the mean, while a high standard deviation shows greater variability. When comparing two sets of data, both the mean and the standard deviation should be considered together.
标准差度量数据值与均值的平均距离。对于样本,我们通常使用除以 n-1 的公式(样本标准差 s)。标准差较小表明数据紧密聚集在均值周围,标准差较大则显示变异性较大。在比较两组数据时,应同时考虑均值和标准差。
6. Probability Basics | 概率基础
Probability measures the chance of an event occurring, expressed as a number between 0 (impossible) and 1 (certain). It can be written as a fraction, decimal or percentage. The sum of probabilities of all mutually exclusive outcomes of an experiment equals 1. The complement rule states that P(not A) = 1 – P(A).
概率度量事件发生的可能性,用一个介于 0(不可能)和 1(必然)之间的数字表示。它可以写成分数、小数或百分比。一个试验中所有互斥结果的概率之和等于 1。补集规则是 P(非 A) = 1 – P(A)。
For combined events, we use sample space diagrams, tree diagrams and Venn diagrams. The addition rule P(A ∪ B) = P(A) + P(B) – P(A ∩ B) applies to any two events. For independent events, multiplication is used: P(A ∩ B) = P(A) × P(B). Conditional probability P(A|B) represents the probability of A given that B has occurred, and is linked by the formula P(A|B) = P(A ∩ B) / P(B).
对于组合事件,我们使用样本空间图、树状图和文氏图。对于任意两个事件,加法法则为 P(A ∪ B) = P(A) + P(B) – P(A ∩ B)。独立事件使用乘法:P(A ∩ B) = P(A) × P(B)。条件概率 P(A|B) 表示在 B 已发生的条件下 A 发生的概率,并由公式 P(A|B) = P(A ∩ B) / P(B) 关联。
Tree diagrams are especially useful for multi-stage events, where probabilities are multiplied along branches and added across different paths. Remember to check that probabilities on branches from the same point add up to 1.
树状图特别适用于多阶段事件,沿分支概率相乘,不同路径概率相加。请记住,同一点发出的分支概率之和必须等于 1。
7. Scatter Graphs and Correlation | 散点图与相关
A scatter graph displays the relationship between two quantitative variables. Correlation describes the strength and direction of this relationship: positive correlation means both variables increase together; negative correlation means one variable increases as the other decreases. If there is no pattern, we say there is zero correlation. Correlation does not imply causation – an observed association may be due to a third lurking variable.
散点图展示两个定量变量之间的关系。相关描述这种关系的强度和方向:正相关意味着两个变量同时增加;负相关意味着一个变量增加而另一个减少。如果没有规律,则称为零相关。相关不代表因果关系——观察到的关联可能由第三个潜在变量造成。
We can assess correlation by drawing a line of best fit (linear regression line) ‘by eye’ or using more formal methods. Spearman’s rank correlation coefficient (rₛ) is a non-parametric measure used when data are ranked or when the relationship is monotonic but not necessarily linear. The formula is rₛ = 1 – (6 Σ d²) / [n (n² – 1)], where d is the difference in ranks for each pair and n is the number of pairs. The value of rₛ ranges from -1 (perfect negative correlation) to +1 (perfect positive correlation).
我们通过目测绘制最佳拟合直线(线性回归线)或使用更正式的方法来评估相关。斯皮尔曼等级相关系数(rₛ)是一种非参数度量,适用于秩次数据或单调但不必线性的关系。公式为 rₛ = 1 – (6 Σ d²) / [n (n² – 1)],其中 d 是每对数据的秩次差,n 是对数。rₛ 的值介于 -1(完全负相关)到 +1(完全正相关)之间。
8. Time Series and Index Numbers | 时间序列与指数
A time series records data at regular time intervals. It typically shows four components: trend (long-term movement), seasonal variation (regular pattern within each year), cyclical variation (fluctuations longer than a year) and random variation. Moving averages help to smooth out short-term fluctuations and reveal the underlying trend.
时间序列记录等间隔时间点上的数据,通常包含四个成分:趋势(长期变动)、季节变动(每年内的规律模式)、循环变动(超过一年的波动)和随机变动。移动平均有助于消除短期波动,揭示潜在的趋势。
To calculate a moving average for quarterly data, we usually use a 4-point moving average and then centre it if an even number of points is used. Once the trend is identified, seasonal effects can be estimated by subtracting the trend from the actual values (additive model) or dividing (multiplicative model).
对于季度数据,通常使用4点移动平均,若点数偶数则需进行居中处理。识别出趋势后,可通过从实际值中减去趋势(加法模型)或相除(乘法模型)来估计季节效应。
Index numbers compare the value of a variable over time relative to a base period. The formula is Index = (Value in current period / Value in base period) × 100. They widely appear in economics, e.g. consumer price indices, and are useful for making percentage comparisons. Understanding weighting is important when different items have different importance in a composite index.
指数用于衡量变量相对于基期随时间的变化。公式为 指数 = (当期数值 / 基期数值)× 100。指数广泛用于经济学,例如消费者价格指数,并有助于进行百分比比较。当不同项目在总指数中重要性不同时,理解加权的概念很重要。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导