📚 Year 8 OCR Statistics: Core Knowledge Review | Year 8 OCR 统计:核心知识点梳理
Statistics is about collecting, organising and interpreting data to make sense of the world around us. In Year 8 OCR Mathematics, you will build on your earlier work with charts and averages to handle a wider range of data types, analyse spread and begin to explore probability in a structured way. This article summarises the core knowledge you need, from data collection and frequency tables to probability tree diagrams.
统计学的核心在于收集、整理和解读数据,从而理解我们周围的世界。在 Year 8 OCR 数学中,你将在之前学习过的图表和平均数基础上,处理更广泛的数据类型,分析数据的离散程度,并开始系统地探索概率。本文将梳理你需要掌握的核心知识点,从数据收集、频率表一直延伸到概率树状图。
1. Types of Data | 数据类型
Data can be classified as qualitative (descriptive) or quantitative (numerical). Quantitative data is further divided into discrete data, which can only take specific values (often counts, like number of siblings), and continuous data, which can take any value within a range (like height or mass). Knowing the data type helps you choose the right graph and statistical measure.
数据可以分为定性数据(描述性)和定量数据(数值型)。定量数据又分为离散数据,只能取特定值(通常是计数,例如兄弟姐妹的数量),以及连续数据,可以在一个范围内取任意值(如身高或质量)。了解数据类型有助于你选择合适的图表和统计量。
| Data Type | Description | Example |
|---|---|---|
| Qualitative | Non-numerical, categories | Eye colour, favourite film |
| Quantitative discrete | Can be counted, gaps between values | Number of pets: 0, 1, 2, … |
| Quantitative continuous | Measured, can take any value in an interval | Height: 152.4 cm |
数据类型表格:定性、定量离散、定量连续。
2. Collecting Data | 数据收集
Data can be collected through a census (surveying every member of a population) or a sample (surveying a subset). A sample must be representative and, ideally, random to avoid bias. In Year 8, you will learn to design simple questionnaires and appreciate that poorly worded questions can produce misleading results.
数据可以通过普查(调查总体中的每一个成员)或样本(调查一个子集)来收集。样本必须具有代表性,理想情况下还应随机抽取以避免偏差。在 Year 8,你将学习设计简单的问卷,并认识到问题表述不当可能会产生误导性的结果。
A random sample means every member of the population has an equal chance of being chosen. You might use a random number generator or names from a hat to achieve this. Bias occurs when the sample is not fair, for example, only asking your friends in a survey about school meals.
随机样本意味着总体中的每个成员都有平等的机会被选中。你可以使用随机数生成器或从帽子里抽名字来实现。如果样本不公平,比如在学校午餐调查中只询问自己的朋友,就会产生偏差。
3. Frequency Tables | 频率表
A frequency table organises raw data into a table showing each data value (or category) and how often it occurs. Tally marks are often used to count frequencies before recording the final number. The total of the frequency column gives the number of data items.
频率表将原始数据整理成表格,显示每个数据值(或类别)及其出现的次数。在记录最终次数之前,通常会用划记符号来计数。频率列的总和就是数据项的个数。
For discrete data with few distinct values, an ungrouped frequency table is sufficient. For continuous data or data with many different values, a grouped frequency table (covered later) is more appropriate.
对于只有少数不同值的离散数据,使用未分组的频率表就足够了。对于连续数据或有很多不同值的数据,更适合使用分组频率表(稍后介绍)。
4. Bar Charts and Pie Charts | 条形图和饼图
Bar charts display categorical or discrete data using rectangular bars. The height or length of each bar represents the frequency. Bars must be of equal width and separated by equal gaps. A bar chart can be drawn vertically or horizontally, and you should always label axes and give the chart a title.
条形图使用矩形条来展示分类数据或离散数据。每个条的高度或长度代表频数。条的宽度必须相等,并且条与条之间要有相等的间隔。条形图可以垂直或水平绘制,并且必须标注坐标轴并为图表添加标题。
Pie charts show proportions of a whole. The size of each sector angle is calculated using the formula: (frequency ÷ total frequency) × 360°. In Year 8 you learn to interpret pie charts and to construct them using a protractor and compasses.
饼图用于显示整体中各个部分的比例。每个扇区的角度大小使用公式计算:(频数 ÷ 总频数)× 360°。在 Year 8,你将学习解读饼图,并使用量角器和圆规绘制饼图。
5. Line Graphs and Scatter Graphs | 折线图和散点图
A line graph is used to show how a quantity changes over time. Plot points for each time period and connect them with straight lines. Time always goes on the horizontal axis. Line graphs are useful for spotting trends, such as increasing or decreasing patterns.
折线图用于展示某个量随时间变化的情况。为每个时间段描点,然后用直线连接。时间总是放在横轴上。折线图有助于识别趋势,比如上升或下降的模式。
A scatter graph (scatter plot) shows the relationship between two sets of continuous data. Each point represents a pair of values. If the points show an upward pattern, there is a positive correlation; a downward pattern indicates negative correlation. If points are spread randomly, there is no correlation. You will also learn to draw a line of best fit to make predictions.
散点图显示两组连续数据之间的关系。每个点代表一对数值。如果点的分布呈现上升模式,说明存在正相关;下降模式则表明负相关。如果点随机分布,则没有相关性。你还将学习绘制最佳拟合线来进行预测。
Correlation is not causation!
相关性并不意味着因果关系!
6. Mean, Median and Mode | 平均数、中位数和众数
These are measures of central tendency (averages) that summarise a data set with a single typical value.
- Mode: the value that occurs most often. A set can have one mode, more than one mode (bimodal) or no mode.
- Median: the middle value when data is ordered from smallest to largest. If there are two middle values, the median is their mean.
- Mean: calculated by adding all values together and dividing by the number of values.
它们是集中趋势(平均数)的度量,可以用一个典型值来概括数据集。
- 众数:出现次数最多的值。一个数据集可以有一个众数、多个众数(双众数)或没有众数。
- 中位数:将数据从小到大排序后中间位置的值。如果中间位置有两个值,中位数是这两个数的平均数。
- 平均数:将所有数据值相加,然后除以数据值的个数。
Formula for the mean: Sum of all data values ÷ Number of values. When using a frequency table, the mean = (Σ value × frequency) ÷ Σ frequency.
平均数的公式:所有数据值的总和 ÷ 数据值的个数。使用频率表时,平均数 = (Σ 值 × 频数) ÷ Σ 频数。
7. Range and Spread | 范围和离散程度
The range is a simple measure of spread: it is the difference between the largest and smallest values in the data set. A larger range means the data is more spread out. The range tells you nothing about the values in between, so it can be affected by outliers.
范围是一种简单的离散程度度量:它是数据集中最大值与最小值的差。范围越大,说明数据越分散。范围不能告诉你中间值的情况,因此容易受到异常值的影响。
An outlier is a value that lies very far from the rest of the data. You might spot an outlier on a scatter graph or in a list. When calculating the mean, outliers can pull the mean up or down significantly, whereas the median is less affected.
异常值是远离其他数据的一个值。你可以在散点图或列表中识别出异常值。在计算平均数时,异常值可能会显著拉高或拉低平均数,而中位数受到的影响较小。
8. Grouped Frequency Tables | 分组频率表
When dealing with continuous data or a large number of different values, we group the data into class intervals. For example, heights can be grouped as 150 ≤ h < 160 cm. A grouped frequency table shows each class interval and its frequency. You must be careful to use inequalities correctly so that no data value falls into two groups.
处理连续数据或大量不同数值时,我们会将数据分组为组距。例如,身高可以分组为 150 ≤ h < 160 cm。分组频率表显示每个组距及其频数。你必须正确使用不等式,确保没有数据值同时落入两个组。
To estimate the mean from a grouped frequency table, you find the midpoint of each class interval, multiply by the frequency, sum these products, then divide by the total frequency. Because we use midpoints, this gives an estimated mean, not the exact mean.
要从分组频率表估计平均数,你需要找出每个组距的中点值,乘以频数,将这些乘积相加,然后除以总频数。因为我们使用的是中点值,所以得到的是估计平均数,而不是精确的平均数。
| Class interval | Frequency | Midpoint | Midpoint × Frequency |
|---|---|---|---|
| 0 ≤ x < 10 | 5 | 5 | 25 |
| 10 ≤ x < 20 | 8 | 15 | 120 |
估计平均数 = 总乘积和 ÷ 总频数。
9. Introduction to Probability | 概率基础
Probability is a measure of how likely an event is to happen. It is given as a number between 0 and 1, where 0 means impossible and 1 means certain. You can write probability as a fraction, decimal or percentage. In Year 8, you will often use fractions in simplest form.
概率衡量一个事件发生的可能性大小,用 0 到 1 之间的一个数字表示,0 表示不可能,1 表示必然发生。你可以用分数、小数或百分比来表示概率。在 Year 8,你通常会使用最简分数。
For equally likely outcomes, theoretical probability is:
P(event) = Number of favourable outcomes ÷ Total number of outcomes
对于等可能的结果,理论概率是:
P(事件) = 有利结果数 ÷ 总结果数
The probability of an event not happening is 1 minus the probability of the event happening. Mutually exclusive events cannot happen at the same time; the sum of their probabilities is the probability that either occurs.
事件不发生的概率等于 1 减去事件发生的概率。互斥事件不能同时发生;它们概率的和就是其中任一事件发生的概率。
10. Sample Spaces and Listing Outcomes | 样本空间与列举结果
A sample space is the set of all possible outcomes of an experiment. Listing outcomes systematically, for example using a table or a diagram, helps ensure you count each outcome exactly once. For rolling two dice, a 6 × 6 table is a very effective sample space diagram.
样本空间是一次试验所有可能结果的集合。系统性地列举结果,例如使用表格或图表,有助于确保每个结果恰好计数一次。对于掷两个骰子的情况,一个 6 × 6 的表格是一个非常有效的样本空间图。
You can use sample space diagrams to find probabilities of combined events, such as ‘sum of the scores is 7’. Count the number of outcomes that satisfy the condition and divide by 36.
你可以使用样本空间图来求组合事件的概率,例如“点数之和为 7”。数出满足条件的结果个数,然后除以 36。
For experiments where outcomes are not equally likely, you can carry out an experiment and use relative frequency to estimate probability: Relative frequency = Number of times event occurs ÷ Total number of trials. The more trials you carry out, the closer the relative frequency tends to get to the theoretical probability.
对于结果不等可能的试验,你可以通过实验使用相对频率来估计概率:相对频率 = 事件发生的次数 ÷ 总试验次数。进行的试验次数越多,相对频率往往越接近理论概率。
11. Venn Diagrams | 维恩图
A Venn diagram uses circles (or other shapes) to show sets and the relationships between them. The universal set ξ (or E) contains all elements under consideration. Overlapping regions represent elements that belong to both sets (intersection). You will use Venn diagrams to organise data and to solve probability problems involving ‘AND’ and ‘OR’.
维恩图用圆圈(或其他形状)来表示集合以及它们之间的关系。全集 ξ(或 E)包含所考虑的所有元素。重叠区域表示同时属于两个集合的元素(交集)。你将使用维恩图来整理数据,并解决涉及“且”和“或”的概率问题。
For two sets A and B, the probability of A or B happening is P(A ∪ B) = P(A) + P(B) − P(A ∩ B). You can often find these probabilities easily by placing numbers directly on the Venn diagram.
对于两个集合 A 和 B,A 或 B 发生的概率为 P(A ∪ B) = P(A) + P(B) − P(A ∩ B)。通常,你可以直接在维恩图上放置数字,从而轻松求出这些概率。
12. Tree Diagrams | 树状图
Tree diagrams show all possible outcomes of two or more events happening in sequence. Branches are labelled with probabilities, and the probabilities on branches from the same point must sum to 1. To find the probability of a combination of events, multiply the probabilities along the relevant branches.
树状图展示了两个或多个事件按顺序发生的所有可能结果。分支上标注着概率,并且从同一点出发的各分支概率之和必须为 1。要求出一系列事件组合的概率,沿相应分支将概率相乘。
Tree diagrams are especially useful when events are independent (the outcome of one does not affect the other). For example, flipping a coin and rolling a die: the coin’s result does not change the die’s probabilities. You will also meet tree diagrams for conditional probability in later years, but in Year 8 the focus is on independent events.
当事件相互独立(一个事件的结果不影响另一个事件)时,树状图特别有用。例如,抛一枚硬币并掷一个骰子:硬币的结果不会改变骰子的概率。后续年级你会遇到条件概率的树状图,但在 Year 8,重点是独立事件。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导