📚 Year 9 AQA Statistics: Core Knowledge Points Review | 九年级AQA统计:核心知识点梳理
Statistics is the branch of mathematics that deals with collecting, organising, analysing, interpreting and presenting data. In Year 9, students following the AQA specification develop essential skills that form the backbone of the GCSE Statistics course. This review walks through the core knowledge points, from understanding different types of data to drawing box plots and calculating probabilities, ensuring you have a solid grasp of every key concept.
统计学是数学的一个分支,涉及数据的收集、整理、分析、解释和呈现。对于学习AQA课程的九年级学生来说,掌握这些基本技能是GCSE统计课程的重要基础。这篇文章梳理了所有核心知识点,从理解不同类型的数据,到绘制箱线图和计算概率,帮助你牢固掌握每一个关键概念。
1. Types of Data and Data Collection | 数据类型与数据收集
Data can be classified as primary or secondary. Primary data is collected firsthand by the researcher for a specific purpose, such as conducting a survey or an experiment. Secondary data has been collected by someone else for a different purpose, like using census information or published statistics. Both types have advantages; primary data is tailored to the investigation but time‑consuming to gather, while secondary data is cheaper and quicker to obtain but may not perfectly fit the research needs.
数据可以分为一手数据和二手数据。一手数据是由研究者为了特定目的亲自收集的,比如进行问卷调查或实验。二手数据则是由他人为了其他目的收集的,例如使用人口普查信息或已发布的统计数据。两种类型各有优点:一手数据能精确满足调查需求,但收集起来耗时;二手数据获取成本低、速度快,但可能无法完全契合研究需要。
Data is also classified as qualitative or quantitative. Qualitative data describes qualities or categories, such as favourite colour or type of pet; it is non‑numerical. Quantitative data involves numbers and can be further split into discrete and continuous. Discrete data can only take certain values, often whole numbers, like the number of students in a class. Continuous data can take any value within a range, such as height, weight or time. Recognising the correct data type helps you choose appropriate diagrams and calculations.
数据还可以分为定性数据和定量数据。定性数据描述的是性质或类别,例如最喜欢的颜色或宠物的种类;它是非数值的。定量数据涉及数字,可以进一步分为离散数据和连续数据。离散数据只能取特定的值,通常是整数,比如班级里的学生人数。连续数据可以在一个范围内取任意值,例如身高、体重或时间。正确识别数据类型有助于选择合适的统计图表和计算方法。
| Data type | Description | Example |
|---|---|---|
| Qualitative | Non‑numerical categories | Eye colour, car brand |
| Quantitative discrete | Numerical, countable values | Number of siblings, shoe size |
| Quantitative continuous | Numerical, any value in a range | Temperature, length of a leaf |
2. Sampling Methods | 抽样方法
When it is impractical to study an entire population, we select a sample—a smaller group chosen to represent the whole. The way a sample is selected greatly influences the reliability of conclusions. Random sampling gives every member of the population an equal chance of being chosen, often using a random number generator or drawing names from a hat. This helps to avoid bias but can be difficult to organise with large populations.
当研究整个总体不切实际时,我们会选择一个样本——一个用来代表整体的小群体。样本的选择方式极大地影响着结论的可靠性。随机抽样让总体中的每一个成员都有均等的机会被选中,通常使用随机数生成器或从帽子中抽签的方式。这有助于避免偏差,但在总体很大时可能难以组织。
Stratified sampling divides the population into groups (strata) based on a characteristic, such as age or gender, and then takes a random sample from each stratum in proportion to its size. This method ensures that the sample accurately reflects the structure of the population. For example, if 40% of a school are girls, a stratified sample of size 50 should contain 20 girls. Systematic sampling selects items at regular intervals from an ordered list, while convenience sampling uses easily available individuals, which often leads to bias. AQA questions frequently ask you to identify the most suitable method and explain why it reduces bias.
分层抽样根据某一特征(如年龄或性别)将总体分成若干组(层),然后从每一层中按比例随机抽取样本。这种方法确保样本能够精确反映总体的结构。例如,如果一所学校有40%的女生,那么一个容量为50的分层样本应包含20名女生。系统抽样是从一个有序列表中每隔固定间隔选取项目,而便利抽样则使用最容易获得的个体,这往往会导致偏差。AQA考试经常要求你识别最合适的抽样方法,并解释为什么它能减少偏差。
3. Frequency Tables and Grouped Data | 频率表与分组数据
Organising raw data into frequency tables is one of the first steps in statistical analysis. A tally chart helps record data as it is collected; each observation is marked with a tally stroke, and the frequency is the total count for each category. Once data is tallied, we can see patterns more clearly. For discrete data with few different values, an ungrouped frequency table works well.
将原始数据整理到频率表中是统计分析的第一步。计数表有助于在收集数据时进行记录;每个观测值用计数符号标记,频率就是属于每个类别的总数。将数据制成计数表后,我们可以更清晰地看到模式。对于不同值较少的离散数据,使用未分组的频率表就很合适。
When dealing with a large range of continuous data, or discrete data with many different values, we use a grouped frequency table. Data is sorted into class intervals, such as 0 ≤ h < 10, 10 ≤ h < 20, and so on. It is important that class intervals do not overlap and have equal width where possible. With grouped data, we cannot find the exact mean, but we can estimate it using the midpoints of intervals. The modal class is the interval with the highest frequency.
当处理范围较大的连续数据,或具有很多不同值的离散数据时,我们会使用分组频率表。数据被整理到组距区间中,如 0 ≤ h < 10,10 ≤ h < 20 等等。组距不能重叠,并且尽可能等宽。对于分组数据,我们无法求出准确的均值,但可以利用组中值进行估计。众数所在的组就是频率最高的区间,称为众数组。
4. Statistical Diagrams: Bar Charts, Pie Charts, Pictograms | 统计图表:条形图、饼图、象形图
Bar charts are used to display categorical or discrete data. The height of each bar represents the frequency, and there are equal gaps between the bars to show the categories are separate. A dual bar chart allows you to compare two datasets side by side, for instance, the favourite sports of boys and girls. The key must clearly label which bar represents which group.
条形图用于展示分类数据或离散数据。每个条形的高度代表频率,条形之间留有等间距的空隙,以表明各个类别是独立的。复式条形图则可以将两个数据集并排比较,例如男生和女生最喜欢的运动。图例必须清楚地标明哪个条形代表哪个组。
Pie charts show proportions of a whole. Each category’s sector angle is calculated using the formula: Angle = (Frequency / Total frequency) × 360°. AQA expects you to measure and draw these angles accurately with a protractor. Pictograms represent data using symbols, where each symbol stands for a certain number of items. A key is essential, and a fraction of a symbol can be used to show values that are not multiples of the key quantity.
饼图用于显示整体中各部分的比例。每个类别扇区的角度通过以下公式计算:角度 = (该类别的频率 / 总频率)× 360°。AQA考试要求你用量角器精确测量并绘制这些角度。象形图用符号来表示数据,每个符号代表一定数量的项目。图例至关重要,当数据值不是图例数量的整数倍时,可以使用部分符号来表示。
Pie chart angle = (Category frequency ÷ Total frequency) × 360°
5. Stem-and-Leaf Diagrams | 茎叶图
A stem-and-leaf diagram keeps the original data values visible while organising them into a compact shape. The ‘stem’ consists of all but the last digit of each number, and the ‘leaf’ is the final digit. For example, 34 would have a stem of 3 and a leaf of 4. All leaves must be written in ascending order from the stem, and a key must be provided, e.g. ‘3 | 4 means 34’. This diagram makes it easy to find the median, mode and range directly.
茎叶图既能将数据整理成紧凑的形式,又能保留原始数据值。“茎”由每个数字除最后一位外的所有位数构成,“叶”则是最后一位数字。例如,34的茎是3,叶是4。所有叶必须从茎开始按升序书写,并需要提供图例,比如“3 | 4 表示 34”。这种图表非常便于直接找出中位数、众数和全距。
Back‑to‑back stem‑and‑leaf diagrams allow you to compare two related datasets, such as test scores of two different classes. The stem runs down the middle, with leaves for one dataset extending to the left and leaves for the other dataset extending to the right. Both sides should be ordered from the stem outwards. This visual layout highlights differences in spread and central tendency between the two groups.
背靠背茎叶图可以用来比较两个相关的数据集,比如两个不同班级的测验分数。茎在中间延伸,一个数据集的叶向左延伸,另一个数据集的叶向右延伸。两侧都应从茎向外按顺序排列。这种可视化布局可以突出两组数据在离散程度和集中趋势上的差异。
6. Averages: Mean, Median, Mode | 平均数:均值、中位数、众数
An average is a single value that summarises the centre of a dataset. The three measures of central tendency are the mean, median and mode. The mean is calculated by adding all values and dividing by the number of values. It uses every data point, so it is sensitive to outliers. The median is the middle value when the data is arranged in order; it is not affected by extreme values, making it a better choice for skewed distributions. The mode is the value that appears most often, and it is the only average that can be used with qualitative data.
平均数是一个概括数据集中心位置的单一数值。三种集中趋势的度量分别是均值、中位数和众数。均值是通过将所有数值相加再除以数据个数计算得出的。它考虑了每一个数据点,因此对异常值很敏感。中位数是将数据按顺序排列后处于中间位置的值;它不受极端值的影响,因此在数据分布偏斜时是更好的选择。众数是出现次数最多的值,并且是唯一可用于定性数据的平均数。
For a frequency table, the mean is estimated using the formula ∑fx / ∑f, where f is the frequency and x is the data value (or midpoint for grouped data). To find the median from an ungrouped frequency table, you can list the values in order or use cumulative frequency. The modal class is the group with the highest frequency when data is grouped.
对于频率表,均值的估算公式为 ∑fx / ∑f,其中 f 是频率,x 是数据值(或分组数据的组中值)。要从未分组频率表中找出中位数,可以将值按顺序列出,或者使用累积频率。在数据分组的情况下,众数组是频率最高的那个组。
Mean = (Sum of all data values) ÷ (Number of data values) = ∑x / n
Estimated mean from a frequency table = ∑fx / ∑f
7. Range and Quartiles | 范围与四分位数
A single measure of spread cannot describe a dataset fully, but combined with an average it gives a better picture. The range is the simplest measure: Range = maximum value − minimum value. Although easy to calculate, the range is strongly affected by outliers. A more robust measure is the interquartile range (IQR), which focuses on the middle 50% of the data.
单一的离散程度度量无法全面描述一个数据集,但与平均数结合使用可以给出更完整的图像。全距是最简单的度量:全距 = 最大值 − 最小值。尽管计算简单,全距极易受异常值的影响。一种更稳健的度量是四分位距(IQR),它聚焦于中间50%的数据。
To find quartiles, first arrange the data in ascending order. The lower quartile (Q₁) is the median of the lower half of the data; the upper quartile (Q₃) is the median of the upper half. When the number of data points, n, is odd, the overall median is excluded from both halves. The interquartile range is calculated as IQR = Q₃ − Q₁. Knowing how to extract quartiles from a stem-and-leaf diagram or a cumulative frequency graph is a vital skill.
要找出四分位数,首先将数据按升序排列。下四分位数(Q₁)是数据下半部分的中位数;上四分位数(Q₃)是数据上半部分的中位数。当数据个数 n 为奇数时,总中位数不包含在上下两半中。四分位距的计算公式为 IQR = Q₃ − Q₁。掌握如何从茎叶图或累积频率图中提取四分位数是一项关键技能。
Range = Maximum − Minimum
Interquartile range (IQR) = Upper quartile (Q₃) − Lower quartile (Q₁)
8. Box Plots | 箱线图
A box plot (or box-and-whisker diagram) is a clear visual representation of the five-number summary: the minimum, lower quartile (Q₁), median (Q₂), upper quartile (Q₃) and maximum. It is drawn on a scale with a rectangular box from Q₁ to Q₃, a vertical line inside the box at the median, and ‘whiskers’ extending to the minimum and maximum, provided there are no outliers. Outliers are usually defined as values more than 1.5 × IQR beyond the quartiles.
箱线图(或称箱形图)是五数总结的清晰可视化表示:最小值、下四分位数(Q₁)、中位数(Q₂)、上四分位数(Q₃)和最大值。它绘制在一个带有刻度的数轴上,从 Q₁ 到 Q₃ 画一个矩形箱体,箱体内在中位数处画一条竖线,“须线”从箱体两端延伸至最小值和最大值(假设没有异常值)。异常值通常定义为超出四分位数1.5倍IQR的值。
Box plots are extremely useful for comparing the distributions of two or more datasets. By placing box plots for different groups above the same scale, you can quickly comment on the median, spread (IQR and range) and skewness. For example, ‘Class B has a higher median score than Class A, but Class A’s IQR is smaller, indicating more consistent performance.’ Remember that a box plot does not show the mean or the detailed shape of the distribution.
箱线图在比较两个或多个数据集的分布时非常有用。将不同组的箱线图放在同一刻度上方,你可以迅速对中位数、离散程度(IQR 和全距)以及偏态做出评论。例如,“B班的中位数分数高于A班,但A班的IQR更小,表明成绩更稳定”。请记住,箱线图不显示均值,也不展示分布的详细形状。
9. Scatter Graphs and Correlation | 散点图与相关性
A scatter graph displays the relationship between two variables. Each point on the graph represents a pair of values. The pattern of the points indicates the type of correlation: positive correlation means as one variable increases, the other tends to increase; negative correlation means as one variable increases, the other tends to decrease; no correlation means there is no apparent pattern. The strength of correlation can be described as strong, moderate or weak.
散点图展示两个变量之间的关系。图上的每个点代表一对数值。点的分布模式表明了相关的类型:正相关意味着当一个变量增加时,另一个变量也倾向于增加;负相关意味着当一个变量增加时,另一个变量倾向于减少;零相关表示没有明显的模式。相关的强度可以描述为强、中等或弱。
A line of best fit (or trend line) can be drawn through the points to model the relationship. It should pass through the mean of the x‑values and the mean of the y‑values, and have roughly the same number of points above and below the line. The line can be used to estimate unknown values: interpolation (predicting within the range of data) is reliable; extrapolation (predicting outside the data range) is less reliable. Critically, correlation does not imply causation—just because two variables increase together does not mean one causes the other.
可以通过数据点绘制一条最佳拟合线(或趋势线)来建立关系模型。该线应经过x值的均值和y值的均值,并且线上方和线下方的点数大致相同。这条线可以用来估算未知值:内插(在数据范围内进行预测)是可靠的;外推(在数据范围之外进行预测)则不太可靠。关键的一点是,相关性并不意味着因果关系——两个变量同时增加并不代表一个导致了另一个。
10. Introduction to Probability | 概率入门
Probability measures how likely an event is to happen, expressed as a number between 0 (impossible) and 1 (certain), or as a fraction, decimal or percentage. For equally likely outcomes, the theoretical probability of an event A is P(A) = number of favourable outcomes / total number of possible outcomes. This forms the foundation for calculating probabilities in games of chance, dice rolling and card drawing.
概率衡量一个事件发生的可能性,用一个0(不可能)到1(必然)之间的数字表示,也可以用分数、小数或百分数表示。对于等可能的结果,事件A的理论概率为 P(A) = 有利结果的数量 / 所有可能结果的总数。这是计算博彩游戏、掷骰子和抽牌等概率的基础。
When outcomes are not equally likely, or when we have observed data, we can use experimental probability (relative frequency): Relative frequency = number of times event occurs / total number of trials. The more trials we conduct, the closer the experimental probability tends to get to the theoretical probability. The expected number of occurrences of an event can be calculated by multiplying the probability by the number of trials: Expected number = P(event) × number of trials. This concept links probability back to data analysis.
当结果不是等可能时,或者当我们有观测数据时,可以使用实验概率(相对频率):相对频率 = 事件发生的次数 / 总试验次数。我们进行的试验次数越多,实验概率往往越接近理论概率。一个事件的预期发生次数可以通过概率乘以试验次数来计算:预期次数 = P(事件) × 试验次数。这个概念将概率与数据分析联系了起来。
P(Event) = Number of favourable outcomes ÷ Total number of possible outcomes
Expected frequency = Probability × Number of trials
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导