📚 Year 11 Eduqas Statistics: Core Knowledge Summary | Year 11 Eduqas 统计:核心知识点梳理
Eduqas GCSE Statistics brings together all the essential techniques for collecting, representing, analysing, and interpreting data. This revision guide covers the core topics you need for your Year 11 exam, with clear explanations and key formulae.
Eduqas GCSE 统计学将收集、表示、分析和解读数据的基本技术融于一体。本复习指南涵盖了你为 Year 11 考试需要掌握的核心主题,配以清晰的解释和重要公式。
1. Data Collection Methods | 数据收集方法
Data can be gathered through surveys, questionnaires, experiments, interviews, or observations. Primary data is information you collect yourself for a specific purpose, while secondary data has already been collected by someone else.
数据可以通过调查、问卷、实验、访谈或观察来收集。一手数据是你为了特定目的亲自收集的信息,而二手数据是已经被其他人收集过的数据。
A well-designed questionnaire should avoid leading questions, offer mutually exclusive response categories, and include an option for every possible answer. Pilot studies are small-scale trial runs that help refine the data collection process before the main study.
设计良好的问卷应避免诱导性问题,提供互斥的回答类别,并为每种可能的答案都设置选项。试点研究是小规模的试运行,有助于在主研究之前完善数据收集过程。
Data can be classified as categorical (qualitative) or numerical (quantitative). Numerical data can be further split into discrete (counted) and continuous (measured) types. Understanding variable types helps you choose the right graph or summary statistic.
数据可以分为类别(定性)或数值(定量)。数值数据可进一步分为离散型(计数的)和连续型(测量的)。理解变量类型有助于你选择合适的图表或汇总统计量。
2. Sampling Techniques | 抽样技术
A population is the entire set of individuals or items you are interested in. A sample is a subset selected from the population. Sampling is used because it is often impractical or too expensive to study every member.
总体是你感兴趣的整个个体或项目的集合。样本是从总体中选出的一个子集。抽样之所以被使用,是因为研究每一个成员往往不切实际或成本过高。
Simple random sampling gives every member an equal chance of being chosen, reducing bias. Stratified sampling divides the population into meaningful strata and then takes a random sample from each, often in proportion to its size. This ensures key groups are represented.
简单随机抽样给予每个成员均等的被选中的机会,减少了偏差。分层抽样将总体划分为有意义的层,然后从每一层中按比例随机抽样,常常与其规模成比例。这确保了关键群体的代表性。
Systematic sampling selects every kth member from a list after a random start. Quota sampling is a non‑probability method where interviewers select participants to fill set quotas, commonly used in market research. Each sampling method has trade‑offs between cost, convenience, and how representative the sample is.
系统抽样从列表中随机起点开始,每隔 k 个成员选取一个。配额抽样是一种非概率方法,由访问员选取参与者来填满设定的份额,常用于市场调研。每种抽样方法在成本、便利性和样本代表性之间都有取舍。
3. Presenting Data: Charts and Diagrams | 数据展示:图表与图示
Bar charts display categorical data using separated bars of equal width. The height of each bar represents the frequency or frequency density. Pie charts show how a total is divided into sectors, with each slice proportional to its frequency.
条形图用等宽且分开的条形来展示分类数据。每个条形的高度代表频数或频率密度。饼图则展示总体如何被划分为扇形,每块扇形与其频数成比例。
Pictograms use simple symbols to represent a certain number of units. A key is essential. Population pyramids show the age and sex distribution of a population back-to-back, revealing demographic structure. Stem‑and‑leaf diagrams preserve the actual data values while showing their shape.
象形图用简单符号代表一定数量的单位,必须附有图例。人口金字塔背靠背地展示了一个人口在年龄和性别上的分布,揭示了人口结构。茎叶图在显示数据形状的同时保留了原始数值。
Two‑way tables organise bivariate categorical data, allowing you to calculate row or column percentages and spot associations. Always label axes clearly, give a title, and use an appropriate scale when drawing diagrams.
双向表对双变量分类数据进行整理,可计算行或列的百分比并发现关联。绘制图表时务必清晰地标记坐标轴、给出标题并使用合适的刻度。
4. Averages and Measures of Spread | 平均数与离散程度
An average is a measure of central tendency, telling you where the centre of a data set lies. The three common averages are the mean, median, and mode. The mean is calculated by summing all values and dividing by the number of values; it is the most familiar but can be distorted by outliers.
平均数是集中趋势的度量,告诉你数据集的中心在哪里。常见的三种平均数是平均数、中位数和众数。平均数通过将所有数值相加再除以数值个数计算得到;它最为人熟知,但可能被异常值扭曲。
The median is the middle value when data are ordered. If there is an even number of values, take the mean of the two middle numbers. The mode is the value that occurs most often. A data set can have more than one mode or no mode at all.
中位数是数据排序后的中间值。若数值个数为偶数,则取中间两个数值的平均数。众数是出现频率最高的数值。一个数据集可以有多个众数,也可以没有众数。
Measures of spread describe how data is dispersed. Range = maximum – minimum. The interquartile range (IQR) = upper quartile (Q₃) – lower quartile (Q₁) and covers the middle 50% of data. The IQR is resistant to outliers.
离散程度的度量描述了数据的分散情况。极差 = 最大值 – 最小值。四分位距 (IQR) = 上四分位数 (Q₃) – 下四分位数 (Q₁),涵盖了中间 50% 的数据。IQR 不受异常值的影响。
Mean x̄ = Σx / n | Standard deviation σ = √[ Σ(x – μ)² / n ]
For a sample, the sample variance uses n – 1. The standard deviation is the square root of the variance and gives a typical distance from the mean. In Eduqas GCSE, you may be expected to calculate it using a table or calculator.
对于样本,样本方差使用 n – 1。标准差是方差的平方根,给出了与平均数之间的典型距离。在 Eduqas GCSE 中,可能会要求你使用表格或计算器来计算它。
5. Cumulative Frequency and Box Plots | 累积频率与箱线图
Cumulative frequency is a running total of frequencies as you move through ordered data classes. Plotting the upper class boundary against cumulative frequency gives a cumulative frequency curve, often an S‑shape.
累积频率是在有序的数据组中逐步累加的频率总和。将组距上界相对于累积频率描点就可得到累积频率曲线,通常呈S形。
From the curve you can estimate quartiles and the median by drawing horizontal lines across from the appropriate cumulative frequencies (n/4, n/2, 3n/4) and reading down to the data axis. Percentiles can be found in a similar way.
从曲线中可以估计四分位数和中位数:从相应的累积频率(n/4, n/2, 3n/4)处画水平线,再向下读取数据轴。百分位数也可用类似方法求出。
A box plot (box‑and‑whisker diagram) shows the minimum, Q₁, median, Q₃, and maximum. The box spans the IQR, with a line for the median. Whiskers extend to the extreme values within 1.5 × IQR from the box; points beyond are possible outliers.
箱线图(盒须图)显示了最小值、Q₁、中位数、Q₃ 和最大值。箱子覆盖 IQR,并用一条线标出中位数。须线从箱子延伸至距箱子 1.5 × IQR 范围内的最远值;超出此范围的点可能就是异常值。
Box plots are excellent for comparing distributions side‑by‑side and highlighting skewness. A longer upper whisker indicates positive skew.
箱线图非常适合并排比较分布并突出偏态。较长的上须线表明正偏态。
6. Histograms and Frequency Density | 直方图与频率密度
Unlike bar charts, histograms are used for continuous data and have no gaps between bars. The area of each bar represents the frequency of that class. When class widths are equal, bar height can simply be the frequency.
与条形图不同,直方图用于连续数据,条形之间没有间隙。每个条形的面积代表该组的频率。当组距等宽时,条形高度可以直接使用频率。
For unequal class widths, you must use frequency density to keep the area proportional to frequency. Frequency density = frequency ÷ class width. The vertical axis is labelled ‘Frequency density’.
对于不等宽的组距,必须使用频率密度,以使面积与频率成比例。频率密度 = 频率 ÷ 组距。纵轴标记为“频率密度”。
Frequency density = Frequency / Class width
To find a frequency from a histogram bar, multiply the frequency density by the class width. Always check that the total area corresponds to the total number of observations.
要从直方图的条形求出频率,用频率密度乘以组距。务必检查总面积是否与观测总数相对应。
7. Scatter Graphs and Correlation | 散点图与相关性
A scatter graph displays the relationship between two numerical variables. Each point represents a pair of values (x, y). The pattern of points reveals the type of correlation: positive, negative, or none.
散点图展示了两个数值变量之间的关系。每个点代表一对数值 (x, y)。点的分布模式揭示了相关性的类型:正相关、负相关或无相关。
A line of best fit (regression line) can be drawn by eye, passing as close as possible to all points with roughly equal numbers on each side. It can be used for interpolation (predicting within the data range) but extrapolation outside the range is unreliable.
最佳拟合线(回归线)可通过目测画出,尽可能靠近所有点并使两侧点数大致相等。它可用于内插(在数据范围内预测),但外推超出范围则不可靠。
Correlation does not imply causation; a strong relationship could be due to a third lurking variable. Spearman’s rank correlation coefficient (rₛ) measures the strength of a monotonic relationship for ranked data and is suitable when the relationship is not linear.
相关性不意味因果关系;强相关可能是由于第三个潜变量所致。斯皮尔曼等级相关系数 (rₛ) 衡量排序数据的单调关系强度,适用于非线性关系的情况。
Spearman’s rₛ always falls between –1 and +1. The closer to ±1, the stronger the correlation. The formula uses the differences in ranks (d) as: rₛ = 1 – (6Σd²) / [n(n² – 1)].
斯皮尔曼 rₛ 总是在 –1 到 +1 之间。越接近 ±1,相关性越强。公式使用了秩差 (d):rₛ = 1 – (6Σd²) / [n(n² – 1)]。
8. Probability Basics | 概率基础
Probability quantifies how likely an event is, on a scale from 0 (impossible) to 1 (certain). Theoretical probability of event A is P(A) = number of favourable outcomes / total number of equally likely outcomes.
概率量化了事件发生的可能性,范围从 0(不可能)到 1(必然)。事件 A 的理论概率为 P(A) = 有利结果数量 / 等可能结果总数。
Experimental probability is based on observed data: relative frequency = number of successful trials / total number of trials. The more trials, the more stable the relative frequency tends to become.
实验概率基于观测数据:相对频率 = 成功试验次数 / 总试验次数。试验次数越多,相对频率往往越稳定。
For combined events, the AND rule (multiply) applies to independent events: P(A and B) = P(A) × P(B). The OR rule (add) for mutually exclusive events: P(A or B) = P(A) + P(B). When events are not mutually exclusive, subtract the intersection: P(A or B) = P(A) + P(B) – P(A and B).
对于组合事件,乘法规则 (AND) 适用于独立事件:P(A and B) = P(A) × P(B)。加法规则 (OR) 适用于互斥事件:P(A or B) = P(A) + P(B)。当事件不互斥时,需减去交集:P(A or B) = P(A) + P(B) – P(A and B)。
Venn diagrams and tree diagrams are powerful tools for organising outcomes and calculating probabilities. Conditional probability P(A|B) is the probability of A given that B has occurred and is especially important when events are not independent.
维恩图和树状图是组织结果和计算概率的强大工具。条件概率 P(A|B) 是在事件 B 已发生的情况下 A 发生的概率,当事件不独立时尤为重要。
9. Probability Distributions – the Binomial Distribution | 概率分布 – 二项分布
Some situations involve a fixed number of independent trials, each with the same probability of success p. The number of successes, X, follows a binomial distribution, written as X ~ B(n, p).
有些情景涉及固定次数的独立试验,每次试验成功的概率相同为 p。成功次数 X 服从二项分布,记作 X ~ B(n, p)。
The formula for exactly r successes is: P(X = r) = ⁿCᵣ pʳ (1 – p)ⁿ⁻ʳ, where ⁿCᵣ = n! / [r! (n – r)!] is the binomial coefficient. The mean of a binomial distribution is μ = np, and the variance is σ² = np(1 – p).
恰好取得 r 次成功的概率公式为:P(X = r) = ⁿCᵣ pʳ (1 – p)ⁿ⁻ʳ,其中 ⁿCᵣ = n! / [r! (n – r)!] 是二项式系数。二项分布的均值是 μ = np,方差是 σ² = np(1 – p)。
You can use tables, calculators, or the formula to find probabilities. Understanding the assumptions (fixed n, constant p, independence) is crucial for determining when to apply this distribution.
你可以使用表格、计算器或公式来求概率。理解这些假设(固定的 n、恒定的 p、独立性)对于判断何时应用此分布至关重要。
10. Time Series and Moving Averages | 时间序列与移动平均
A time series is a sequence of data points recorded at regular time intervals, such as daily temperatures or quarterly sales. It often contains a trend (long‑term movement) and seasonal variation (pattern that repeats at fixed periods).
时间序列是以固定时间间隔记录的一系列数据点,例如每日温度或季度销售额。它通常包含趋势(长期变动)和季节变动(以固定周期重复出现的模式)。
Moving averages smooth out short‑term fluctuations to make the underlying trend clearer. For seasonal data with an even number of periods, a centred moving average is used so that the smoothed value lines up with a time point.
移动平均可以消除短期波动,使潜在趋势更加清晰。对于具有偶数个周期的季节性数据,使用中心化移动平均,使平滑后的值对准一个时间点。
Once a trend line is established, you can estimate seasonal effects by subtracting the trend from the actual values. Predictions can be made by extrapolating the trend and then adding back the average seasonal effect, but forecasts become increasingly uncertain further ahead.
一旦确立了趋势线,可以通过实际值减去趋势来估计季节效应。预测可以通过外推趋势然后加回平均季节效应来进行,但越往后的预测不确定性越大。
11. Index Numbers | 指数
Index numbers measure relative change in a variable, such as price or quantity, compared with a base period. The base period is typically given an index of 100. The formula for a simple price index is: index = (current price / base price) × 100.
指数衡量变量(如价格或数量)相对于基期的相对变化。基期通常被赋予指数 100。简单价格指数的公式为:指数 = (当前价格 / 基期价格) × 100。
When items have different importance, a weighted index is constructed. Each item is given a weight w, and the weighted index = Σ(w × price relative) / Σw, where price relative = (current price / base price) × 100.
当项目具有不同的重要性时,需构建加权指数。每个项目被赋予权重 w,加权指数 = Σ(w × 价格比) / Σw,其中价格比 = (当前价格 / 基
Published by TutorHao | Year 11 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导