📚 Year 10 Edexcel Statistics: Core Knowledge Summary | 核心知识点梳理
This article provides a comprehensive revision guide for Year 10 students following the Edexcel GCSE Statistics specification. It summarises the core topics, from data types and sampling to probability and diagrams, with clear explanations and key formulas. Mastering these foundations will help you tackle exam questions with confidence.
本文为学习 Edexcel GCSE 统计学的 10 年级学生提供全面的复习指南。文章梳理了从数据类型、抽样方法到概率和统计图表等核心主题,附有清晰的解释和关键公式。扎实掌握这些基础知识将有助于您自信地应对考试题目。
1. Types of Data | 数据类型
Data can be classified as qualitative (categorical) or quantitative (numerical). Qualitative data describe qualities or categories, such as eye colour or type of vehicle. Quantitative data involve numbers and are split into discrete (counted, e.g. number of siblings) and continuous (measured, e.g. temperature).
数据可分为定性(分类)数据与定量(数值)数据。定性数据描述性质或类别,如眼睛颜色或车辆类型。定量数据涉及数字,细分为离散型(可数,如兄弟姐妹数量)和连续型(测量值,如温度)。
Another important distinction is between primary data, collected directly by the researcher, and secondary data, gathered from existing sources. Primary data are specific but time‑consuming, while secondary data are quicker to obtain but may be less reliable.
另一重要区分是一手数据(研究者直接收集)与二手数据(来自现有来源)。一手数据针对性强但耗时,二手数据获取更快但可能较不可靠。
Data types: qualitative vs quantitative, discrete vs continuous, primary vs secondary.
数据类型:定性/定量,离散/连续,一手/二手。
2. Sampling Methods | 抽样方法
Sampling is used to select a subset of a population. A random sample gives every member an equal chance of being chosen, reducing bias. Stratified sampling divides the population into groups and samples proportionally from each, ensuring representation of subgroups.
抽样用于从总体中选取子集。随机抽样使每个成员有均等被选中的机会,减少偏差。分层抽样将总体分组并按比例从每组取样,确保各子群的代表性。
Systematic sampling selects every k‑th member after a random start, while convenience sampling uses easily available participants but often leads to bias. Quota sampling sets fixed numbers for categories but does not involve random selection.
系统抽样在随机起点后每隔 k 个选取一个成员;便利抽样使用容易获得的个体,但易产生偏差。配额抽样为各类别设定固定数量,但并非随机选取。
| Sampling method | Key feature | 抽样方法 |
|---|---|---|
| Simple random | Equal chance, unbiased | 简单随机:等概率,无偏 |
| Stratified | Proportional representation | 分层:按比例代表 |
| Systematic | Fixed interval selection | 系统:固定间隔选取 |
| Convenience | Easy access, often biased | 便利:易获取,常有偏 |
3. Frequency Tables and Diagrams | 频数表与统计图
Data are often organised into frequency tables. For grouped data, class intervals must not overlap. The class midpoint is used when calculating estimates. A histogram displays grouped continuous data: the area of each bar is proportional to frequency. Frequency density = frequency ÷ class width.
数据通常整理到频数表中。对于分组数据,组区间不得重叠。计算估计值时使用组中点。直方图用于分组连续数据:每个条形的面积与频数成比例。频数密度 = 频数 ÷ 组距宽度。
Frequency density = Frequency ÷ Class width
频数密度 = 频数 ÷ 组距宽度
Bar charts are for categorical or discrete data, with gaps between bars. Pie charts show proportions of a whole, with sector angle = (frequency ÷ total) × 360°. Frequency polygons join the midpoints of each class interval.
条形图用于分类或离散数据,条形之间有间隔。饼图展示整体中各部分比例,扇区角度 = (频数 ÷ 总数) × 360°。频数多边形连接每个组区间的中点。
4. Measures of Central Tendency | 集中趋势的度量
The mean, median and mode summarise the centre of a dataset. The mean is the arithmetic average, the median is the middle value when data are ordered, and the mode is the most frequent value.
均值、中位数和众数概括数据集的中心。均值是算术平均值,中位数是排序后中间的值,众数是最常出现的值。
Mean = Σx ÷ n
均值 = 总计 ÷ 个数
For a frequency table, use Σfx ÷ Σf, where x is the data value (or midpoint). The median position is (n+1)÷2. The mean is affected by outliers, while the median is resistant.
在频数表中,使用 Σfx ÷ Σf,其中 x 是数据值(或组中点)。中位数位置为 (n+1)÷2。均值受异常值影响,中位数则具有抗干扰能力。
5. Measures of Spread | 离散程度的度量
Spread shows how varied the data are. The range = maximum – minimum. The interquartile range (IQR) = Upper quartile (Q₃) – Lower quartile (Q₁). IQR focuses on the middle 50% and is not affected by extreme values.
离散程度表示数据的变异性。极差 = 最大值 − 最小值。四分位距 (IQR) = 上四分位数 (Q₃) − 下四分位数 (Q₁)。IQR 着眼于中间 50% 的数据,不受极端值影响。
IQR = Q₃ − Q₁
四分位距 = Q₃ − Q₁
Standard deviation (SD) is a more advanced measure of spread. It indicates how closely values cluster around the mean. A smaller SD means data are more consistent. For Year 10, you need to interpret SD, not compute it manually for large datasets.
标准差 (SD) 是更高级的离散度量。它显示数据值围绕均值的紧密程度。标准差越小,数据越稳定。对于 10 年级,您需要解读标准差,但不必手动计算大型数据集的标准差。
6. Box Plots and Cumulative Frequency | 箱线图与累积频数
A cumulative frequency table adds frequencies row by row. The cumulative frequency graph (ogive) plots upper class boundaries against cumulative frequency. Quartiles and the median can be read directly from the graph.
累积频数表逐行累加频数。累积频数图(拱形图)以组上界为横坐标、累积频数为纵坐标绘图。可从图上直接读出四分位数和中位数。
Median ≈ value at 50% of total frequency; Q₁ at 25%; Q₃ at 75%.
中位数 ≈ 总频数 50% 处的值;Q₁ 在 25% 处;Q₃ 在 75% 处。
A box plot (box‑and‑whisker diagram) uses the five‑number summary: minimum, Q₁, median, Q₃, maximum. Outliers can be detected using fences: Lower fence = Q₁ − 1.5 × IQR; Upper fence = Q₃ + 1.5 × IQR. Values beyond the fences are potential outliers, often marked with asterisks.
箱线图(盒须图)采用五数综合:最小值、Q₁、中位数、Q₃、最大值。离群值可通过界限进行检测:下界限 = Q₁ − 1.5 × IQR;上界限 = Q₃ + 1.5 × IQR。超出界限的值为潜在离群值,通常用星号标记。
7. Scatter Graphs and Correlation | 散点图与相关性
A scatter graph displays paired data. Correlation describes the relationship: positive (as one variable increases, so does the other), negative (as one increases, the other decreases), or none. The strength can be strong, moderate or weak.
散点图展示成对数据。相关性描述关系:正相关(一个变量增加,另一个也增加),负相关(一个增加,另一个减少),或无相关。相关强度可为强、中等或弱。
A line of best fit is drawn to pass through as many points as possible, balancing points above and below. It is used to estimate unknown values. Be aware that correlation does not imply causation – another factor could be responsible.
绘制最佳拟合线时,应使其尽可能多地穿过数据点,并使线上下的点数大致平衡。该线用于估算未知值。请注意,相关性并不意味着因果关系——可能存在其他因素。
Correlation ≠ causation
相关性 ≠ 因果关系
8. Time Series and Trend Lines | 时间序列与趋势线
A time series is a sequence of observations taken over time. Plotting values against time reveals underlying trend, seasonal variation, and irregular fluctuations. The trend line smooths out short‑term movements, making long‑term changes visible.
时间序列是按时间顺序记录的一系列观测值。将数值对时间作图可揭示潜在趋势、季节性变化和不规则波动。趋势线平滑了短期变动,使长期变化清晰可见。
Moving averages are used to identify the trend. For a 3‑point moving average, the average of three consecutive points is plotted at the centre. This reduces random noise but leads to some loss of data points at the ends.
移动平均用于识别趋势。对于三点移动平均,将三个连续点的平均值绘制于中间位置。这样可以减少随机噪声,但两端会损失一些数据点。
3‑point moving average: (yₜ₋₁ + yₜ + yₜ₊₁) ÷ 3
三点移动平均:(前一值 + 当前值 + 后一值) ÷ 3
9. Probability Basics | 概率基础
Probability measures how likely an event is to occur. It is always between 0 (impossible) and 1 (certain). The probability of an event not occurring is 1 minus the probability of the event: P(not A) = 1 – P(A).
概率衡量事件发生的可能性,取值在 0(不可能)到 1(必然)之间。事件不发生的概率等于 1 减去事件发生的概率:P(非 A) = 1 – P(A)。
Two events are mutually exclusive if they cannot happen at the same time. For mutually exclusive events, P(A or B) = P(A) + P(B). Independent events have no influence on each other; for independent events, P(A and B) = P(A) × P(B).
若两个事件不能同时发生,则互斥。对于互斥事件,P(A 或 B) = P(A) + P(B)。独立事件互不影响;对于独立事件,P(A 且 B) = P(A) × P(B)。
Expected frequency predicts the number of times an event should happen in n trials: Expected frequency = n × P(event).
期望频数预计在 n 次试验中事件发生的次数:期望频数 = n × P(事件)。
10. Probability Tree Diagrams | 概率树图
Tree diagrams show all possible outcomes of multi‑stage events. Branches are labelled with probabilities (summing to 1 at each set of branches). Multiply along branches for combined probability, then add relevant branch probabilities for “or” conditions.
树图展示多阶段事件的所有可能结果。分支上标注概率(每一组分叉的概率和为 1)。沿分支相乘得到组合概率,再将相关分支的概率相加即得“或”条件的概率。
Always check that the ends of branches cover the complete sample space. Replacement or non‑replacement scenarios change the probabilities for second stages. For non‑replacement, denominators decrease accordingly.
务必确保分支末端涵盖整个样本空间。有放回与无放回情境会改变第二阶段的概率。无放回时,分母相应减小。
P(outcome A and outcome B) = P(A) × P(B|A)
P(结果 A 和结果 B) = P(A) × P(B|A)
In simple independent cases, P(A and B) = P(A) × P(B). For conditional problems, adjust the second branch probability accordingly.
在简单的独立情形中,P(A 和 B) = P(A) × P(B)。对于条件概率问题,需相应调整第二个分支的概率。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导