📚 Year 10 OCR Statistics: Core Knowledge Review | Year 10 OCR 统计:核心知识点梳理
This article provides a comprehensive overview of the essential topics covered in Year 10 OCR Statistics. It is designed to help you consolidate your understanding, revise key concepts, and build confidence for assessments. Each section pairs an English explanation with its Chinese counterpart, ensuring clarity for bilingual learners.
本文全面梳理了 Year 10 OCR 统计课程的核心知识点,旨在帮助你巩固理解、复习关键概念,并为评估建立信心。每个小节都将英文解释与中文对照呈现,确保双语学习者能够清晰掌握。
1. Types of Data | 数据类型
Data is the raw information we collect, and it can be classified into two broad categories: qualitative and quantitative.
数据是我们收集的原始信息,可分为两大类:定性数据和定量数据。
Qualitative (categorical) data describes qualities or attributes that cannot be measured with numbers. Examples include favourite colours, types of pet, or responses like ‘agree’, ‘neutral’, ‘disagree’.
定性(分类)数据描述的是无法用数字衡量的性质或属性。例如最喜欢的颜色、宠物种类,或“同意”“中立”“不同意”等回答。
Quantitative data is numerical and can be further subdivided into discrete and continuous data.
定量数据是数字型数据,可进一步细分为离散数据和连续数据。
Discrete data can only take specific, separate values – often whole numbers. For instance, the number of students in a class or the roll of a die are discrete.
离散数据只能取特定的、分离的值——通常是整数。例如,班级里的学生人数或掷骰子的点数都是离散的。
Continuous data can take any value within a given range and is usually measured. Height, weight, time, and temperature are continuous data because they can be infinitely divided.
连续数据可以在给定范围内取任何值,并且通常是测量得到的。身高、体重、时间和温度都是连续数据,因为它们可以无限细分。
2. Data Collection and Sampling | 数据收集与抽样
Primary data is information you collect yourself for a specific purpose, such as conducting a survey or an experiment. While it can be time-consuming and expensive, it ensures the data is directly relevant and up to date.
一手数据是你为特定目的亲自收集的信息,例如开展问卷调查或进行实验。虽然可能耗时且成本较高,但能确保数据直接相关且是最新的。
Secondary data is information that has already been collected by someone else, such as government statistics, internet databases, or historical records. It is often cheaper and quicker to obtain, but you must check its reliability and whether it truly matches your needs.
二手数据是别人已经收集好的信息,例如政府统计数据、网络数据库或历史记录。获取它通常更便宜、更快捷,但你必须检查其可靠性以及是否真正符合你的需求。
When you cannot collect data from an entire population, you use a sample. A good sample should be representative and free from bias.
当你无法从整个总体中收集数据时,就需要使用样本。一个好的样本应具有代表性且无偏差。
Random sampling gives every member of the population an equal chance of being selected. This helps reduce bias but may still produce an unrepresentative sample by chance in small populations.
随机抽样让总体中的每个成员都有相等的机会被选中。这有助于减少偏差,但在小总体中仍可能偶然产生不具代表性的样本。
Stratified sampling divides the population into distinct groups (strata) and then takes a random sample from each group in proportion to its size. This guarantees that all subgroups are fairly represented.
分层抽样将总体划分为不同的组(层),然后按比例从每组中随机抽取样本。这可以保证所有子群体都得到公平的代表。
Other methods like systematic sampling (selecting every nth individual) and quota sampling (selecting a fixed number from each category) also appear in OCR, but be aware that convenience and volunteer samples often introduce bias and should be avoided in rigorous statistical work.
其他方法如系统抽样(每第 n 个个体抽取一个)和配额抽样(从每个类别中选取固定数量)也会在 OCR 考试中出现,但要注意便利抽样和自愿抽样往往会引入偏差,在严谨的统计工作中应当避免。
3. Frequency Tables and Grouped Data | 频数表与分组数据
A frequency table organises raw data by listing each distinct value (or group) alongside how often it occurs. This makes patterns easier to spot.
频数表通过列出每个不同的值(或组)及其出现的次数来整理原始数据,这使规律更容易发现。
For discrete data, you simply tally each outcome. For example, recording the shoe sizes of 20 students gives a table with size categories and their frequencies.
对于离散数据,只需记数每个结果即可。例如,记录 20 名学生的鞋码,会得到一个包含鞋码类别及其频数的表格。
When dealing with continuous data or a large range of discrete values, you group the data into class intervals. It is important that these intervals are of equal width where possible, and they must not overlap.
当处理连续数据或离散数据的范围很大时,需要将数据分组为“组距”。尽可能让组距宽度相等,并且组与组之间不得重叠。
OCR often tests the use of inequality notation to define class intervals: for example, 0 ≤ x < 10 means all values from 0 up to but not including 10. This avoids ambiguity at the boundaries.
OCR 考试中常会考查用不等式符号定义组距:例如,0 ≤ x < 10 表示所有从 0 开始到 10(不含 10)的数值,这避免了边界处的歧义。
From a grouped frequency table you can estimate the total frequency, find the modal class (the group with the highest frequency), and calculate estimates of the mean and median using midpoints.
从分组频数表中,你可以得到总频数、找出众数所在组(频数最高的组),并使用组中点来估算均值和中位数。
4. Charts and Diagrams: Bar Charts, Pie Charts and Stem-and-Leaf | 图表:条形图、饼图和茎叶图
Bar charts represent categorical or discrete data using rectangular bars whose heights are proportional to the frequencies. The bars are separated by equal gaps to show that the categories are distinct.
条形图用矩形条表示分类或离散数据,条的高度与频数成比例。条与条之间留有相等的间距,以显示类别是独立区分的。
A vertical bar chart is often used for comparing frequencies across categories, while a horizontal bar chart can be easier to read when category labels are long.
垂直条形图常用于比较各类别的频数,而当类别标签较长时,水平条形图可能更易读。
Pie charts show proportions of a whole. The angle of each sector is calculated using: sector angle = (frequency / total frequency) × 360°.
饼图显示各部分在整体中的比例。每个扇形的角度计算公式为:扇形角度 = (频数 / 总频数) × 360°。
Stem-and-leaf diagrams retain the original data values while showing their distribution. The ‘stem’ represents the leading digit(s), and the ‘leaf’ is the final digit. An ordered stem-and-leaf diagram sorts the leaves to make it easy to find medians and quartiles.
茎叶图既能展示数据的分布状况,又保留了原始数据值。“茎”代表前导数字,“叶”是最后一位数字。有序茎叶图将叶子排序,便于找出中位数和四分位数。
A key must always be included in a stem-and-leaf diagram, for example, ‘5 | 2 means 52’ or ‘5 | 2 represents 5.2’, so the reader knows how to interpret the digits.
在茎叶图中必须包含一个图例,例如“5 | 2 表示 52”或“5 | 2 表示 5.2”,以便读者知道如何解读这些数字。
5. Measures of Central Tendency | 集中趋势度量
An average is a single value that summarises the centre of a data set. The three most common averages are the mode, median, and mean.
平均数是一个概括数据中心趋势的单一数值。最常见的三种平均数是众数、中位数和均值。
The mode (or modal value) is the value that occurs most frequently. A data set can have one mode (unimodal), two modes (bimodal), or more. For grouped data, you identify the modal class.
众数(或众数值)是出现最频繁的值。数据集可以有一个众数(单峰的)、两个众数(双峰的)或更多。对于分组数据,你确定的是众数所在组。
The median is the middle value when the data is ordered from smallest to largest. If there are n values and n is odd, the median is the value at position (n+1)/2. If n is even, it is the average of the two central numbers.
中位数是将数据从小到大排序后位于中间的值。如果有 n 个值且 n 为奇数,中位数是位于 (n+1)/2 位置上的值;如果 n 为偶数,则是中间两个数的平均值。
The mean (often called the arithmetic average) is found by adding all the values together and dividing by the total number of values. In symbols, for a set of n numbers, mean x̄ = Σx / n.
均值(通常称为算术平均数)的计算方法是将所有数值相加,再除以数值的总个数。用符号表示,对于有 n 个数值的数据集,均值 x̄ = Σx / n。
Each average has strengths and weaknesses. The mean uses all data but is sensitive to outliers. The median is robust to outliers but ignores the actual values beyond the middle. The mode is the only average suitable for qualitative data.
每种平均数都有优缺点。均值使用了所有数据,但对异常值敏感;中位数不受异常值影响,但忽略了中间以外的实际数值;众数是唯一适用于定性数据的平均数。
6. Measures of Spread | 离散趋势度量
Spread tells us how consistent or varied the data is. The simplest measures are the range and the interquartile range (IQR).
离散度告诉我们数据的一致性或差异程度。最简单的度量是极差和四分位距 (IQR)。
The range is the difference between the largest and smallest values: Range = maximum – minimum. It is easy to calculate but highly affected by extreme values.
极差是最大值与最小值的差:极差 = 最大值 – 最小值。计算简单,但极易受极端值的影响。
The interquartile range focuses on the middle 50% of the data. It is the difference between the upper quartile (Q3) and the lower quartile (Q1): IQR = Q3 – Q1.
四分位距关注的是中间 50% 的数据。它是上四分位数 (Q3) 与下四分位数 (Q1) 的差值:IQR = Q3 – Q1。
Quartiles split an ordered data set into four equal parts. Q1 is the median of the lower half, Q2 is the overall median, and Q3 is the median of the upper half. OCR often expects you to find these from stem-and-leaf diagrams or cumulative frequency graphs.
四分位数将有序数据集分成四个等份。Q1 是下半部分的中位数,Q2 是整体的中位数,Q3 是上半部分的中位数。OCR 考试常要求你从茎叶图或累积频数图中找出这些值。
Standard deviation is a more sophisticated measure of spread that shows how far, on average, each value deviates from the mean. For a population, the formula is:
标准差是一种更精细的离散度量,它反映的是各数值与均值之间的平均距离。对于总体,其公式为:
σ = √[ Σ(x – μ)² / N ]
For a sample, we usually divide by (n – 1) instead of N to give an unbiased estimate:
对于样本,我们通常除以 (n – 1) 而非 N,以得到无偏估计值:
s = √[ Σ(x – x̄)² / (n – 1) ]
A smaller standard deviation indicates that the data points are clustered closely around the mean, while a larger one shows they are more spread out.
标准差越小,表示数据点紧密聚集在均值周围;标准差越大,表示数据分布得越分散。
7. Cumulative Frequency and Box Plots | 累积频数与箱线图
A cumulative frequency table adds up the frequencies as you move through the classes, giving a running total. This helps you see how many data points lie below a certain upper class boundary.
累积频数表在逐组浏览时,将频数依次累加,形成一个“运行总数”。这有助于你看清有多少数据点低于某个组上限。
You plot a cumulative frequency graph (or ogive) by marking the upper class boundary against the cumulative frequency and joining the points with a smooth curve. The curve should start at the lower boundary of the first class with a cumulative frequency of 0.
绘制累积频数图(或称为肩形图)时,将各组的上限对应累积频数标出点,然后用平滑曲线连接这些点。曲线应从第一组的组下限、累积频数为 0 处开始。
From the graph, you can estimate the median by going to half the total frequency and reading down to the horizontal axis. You can also estimate Q1 at a quarter of the total frequency and Q3 at three-quarters.
从图中,你可以通过找到总频数的一半位置并向下读横坐标来估算中位数。同样,在总频数四分之一处可以估算 Q1,四分之三处可估算 Q3。
A box plot (or box-and-whisker diagram) displays the five-number summary: minimum value, Q1, median, Q3, and maximum value. The box is drawn from Q1 to Q3 with the median marked inside, and whiskers extend to the minimum and maximum, or to 1.5 × IQR beyond the quartiles to identify outliers.
箱线图(或盒须图)展示了五数概括:最小值、Q1、中位数、Q3 和最大值。盒子从 Q1 画到 Q3,内部标出中位数;须线延伸至最小值和最大值,或者延伸至超出四分位数 1.5 倍 IQR 的位置,以便识别异常值。
Box plots are excellent for comparing distributions side by side, as they immediately reveal the central tendency, spread, and skewness of data sets.
箱线图非常适合并排比较不同数据集的分布,因为它能一目了然地呈现数据的集中趋势、离散程度和偏斜情况。
8. Scatter Graphs and Correlation | 散点图与相关
A scatter graph displays the relationship between two continuous variables. Each point on the graph represents a pair of values (x, y).
散点图显示两个连续变量之间的关系。图上的每个点代表一对数值 (x, y)。
Correlation describes the strength and direction of the linear relationship between the variables. Positive correlation means that as one variable increases, the other tends to increase. Negative correlation means as one increases, the other tends to decrease.
相关性描述了两个变量之间线性关系的强度和方向。正相关意味着一个变量增加时,另一个也倾向于增加。负相关意味着一个变量增加时,另一个倾向于减少。
No correlation indicates there is no clear pattern. Be careful not to confuse correlation with causation – a strong correlation does not prove that one variable causes the change in the other.
零相关表示没有明显的规律。小心不要把相关与因果混淆——强相关并不证明一个变量的变化是由另一个变量引起的。
If the points roughly follow a straight line, you can add a line of best fit. This line should pass through the mean of the x-values and the mean of the y-values, and it must have roughly equal numbers of points above and below it.
如果图中的点大致沿一条直线分布,你可以添加一条最佳拟合线。该直线应穿过 x 值的均值和 y 值的均值,并且线上方和线下方的点数应大致相等。
You can use the line of best fit to make estimates: interpolation (predicting within the range of the data) is usually reliable, but extrapolation (predicting outside the range) can be unreliable because the trend may not continue.
你可以利用最佳拟合线进行估计:内插法(在数据范围内预测)通常可靠,但外推法(在范围之外预测)可能不可靠,因为趋势可能不会延续。
9. Introduction to Probability | 概率入门
Probability measures the chance that a specific event will happen. It is always a number between 0 and 1 inclusive. A probability of 0 means the event is impossible, and 1 means it is certain.
概率衡量某个特定事件发生的可能性。它总是介于 0 和 1 之间(含 0 和 1)。概率为 0 表示事件不可能发生,概率为 1 表示事件必然发生。
For equally likely outcomes, the probability of an event A is: P(A) = (number of favourable outcomes) / (total number of possible outcomes). For example, when rolling a fair six-sided die, P(rolling a 4) = 1/6.
对于等可能结果,事件 A 的概率为:P(A) = (有利结果的数量) / (所有可能结果的总数)。例如,抛掷一枚公平的六面骰子,P(掷出 4 点) = 1/6。
The sum of probabilities of all mutually exclusive outcomes in a sample space is 1. Therefore, the probability of an event not happening is given by: P(not A) = 1 – P(A).
样本空间中所有互斥结果的概率之和为 1。因此,一个事件不发生的概率为:P(非 A) = 1 – P(A)。
Two events are mutually exclusive if they cannot occur at the same time. For such events, P(A or B) = P(A) + P(B). If events can happen together, this simple addition does not work and we must use Venn diagrams or two-way tables to avoid double-counting.
若两个事件不能同时发生,则它们互斥。对于互斥事件,P(A 或 B) = P(A) + P(B)。若事件可以同时发生,就不能简单地相加,必须使用文氏图或双向表来避免重复计数。
Experimental probability (or relative frequency) is found by conducting an experiment or survey: relative frequency = (number of times the event occurs) / (total number of trials). The more trials you do, the closer the experimental probability gets to the theoretical probability – this is the law of large numbers.
实验概率(或相对频率)通过进行实验或调查来获得:相对频率 = (事件发生的次数) / (试验总次数)。试验次数越多,实验概率就越接近理论概率——这就是大数定律。
10. Venn Diagrams and Set Notation | 文氏图与集合符号
A Venn diagram uses overlapping circles to show how sets relate to each other. The rectangle represents the universal set (usually labelled ξ), which contains everything under consideration.
文氏图用相互重叠的圆来表示集合之间的关系。矩形代表全集(通常标记为 ξ),它包含了所考虑范围内的所有事物。
The intersection of two sets A and B, written as A ∩ B, contains the elements that belong to both A and B. The union A ∪ B contains elements that are in A, or in B, or in both.
两个集合 A 和 B 的交集,记作 A ∩ B,包含同时属于 A 和 B 的元素。并集 A ∪ B 包含属于 A、或属于 B、或同时属于两者的元素。
The complement of set A, denoted A’, consists of all elements in the universal set that are not in A.
集合 A 的补集,记作 A’,由全集中所有不属于 A 的元素组成。
When filling in a Venn diagram, always start with the intersection region if it is given, then work outwards to the other regions. This prevents errors in counting that appear when elements belong to multiple sets.
在填写文氏图时,如果给出了交集区域,一定要先填交集,再向外填其他区域。这样能避免因元素属于多个集合而导致的计数错误。
You can use a Venn diagram to solve probability problems involving ‘or’ and ‘and’: for any two sets A and B, P(A ∪ B) = P(A) + P(B) – P(A ∩ B). This formula ensures the overlap is not counted twice.
你可以用文氏图来解决包含“或”和“且”的概率问题:对于任何两个集合 A 和 B,P(A ∪ B) = P(A) + P(B) – P(A ∩ B)。这个公式确保重叠部分不被重复计算。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导