Core Statistics Revision for KS3 Cambridge | KS3剑桥统计:核心知识点梳理

📚 Core Statistics Revision for KS3 Cambridge | KS3剑桥统计:核心知识点梳理

Welcome to your ultimate KS3 Cambridge Statistics revision guide. Whether you are learning how to collect data, draw graphs, or calculate averages and probabilities, this article covers every core idea you need to understand. Work through each section carefully, and you will build a strong foundation in statistics for Checkpoint and beyond.

欢迎阅读你的终极 KS3 剑桥统计复习指南。无论你正在学习如何收集数据、绘制图表,还是计算平均值和概率,本文都涵盖了你需要掌握的每一个核心概念。请仔细阅读每一节,你将打下扎实的统计基础,从容应对 Checkpoint 考试及后续学习。


1. Types of Data | 数据的类型

Data is the raw information we collect to answer questions. It falls into two broad types: categorical data and numerical data. Categorical data records qualities or groups – for example, eye colour, favourite sport, or car brands.

数据是我们为回答问题而收集的原始信息。它分为两大类:分类数据和数值数据。分类数据记录的是品质或组别——例如眼睛颜色、最喜爱的运动或汽车品牌。

Numerical data, on the other hand, records quantities that can be measured or counted. Height, mass, temperature and test scores are all numerical. Numerical data is further divided into discrete data, which can only take certain values (like the number of students in a class), and continuous data, which can take any value within a range (like the mass of an apple).

另一方面,数值数据记录的是可以测量或计数的数量。身高、质量、温度和考试分数都是数值数据。数值数据又分为离散型数据(只能取特定数值,如班级学生人数)和连续型数据(在一个范围内可以取任意数值,如一个苹果的质量)。

Understanding data types helps you decide how to display and analyse information appropriately. For example, you would use a bar chart for categorical data but a histogram or line graph for continuous data.

理解数据类型有助于你恰当地展示和分析信息。例如,你用条形图展示分类数据,而用直方图或折线图展示连续数据。


2. Collecting Data | 数据收集

Before we can analyse anything, we must collect data. Two important questions to ask are: ‘Where does the data come from?’ and ‘How was it collected?’ Data can be primary (collected by you, for your own purpose) or secondary (taken from existing sources like websites, books or databases).

在分析任何东西之前,我们必须收集数据。要问两个重要的问题:“数据从哪里来?”和“数据是如何收集的?”数据可以是初级数据(由你为自己的目的而收集)或次级数据(取自现有来源,如网站、书籍或数据库)。

A survey or an experiment produces primary data. For instance, measuring the hand span of every student in your class gives you primary data. Using meteorological records from the internet to study rainfall is using secondary data.

调查或实验产生初级数据。例如,测量你班上每个学生的手掌宽度就得到了初级数据。而利用互联网上的气象记录来研究降雨量就是在使用次级数据。

When designing a data collection sheet, think carefully about what you want to record. Use a clear table with headings, and always include the units of measurement. When you ask people questions, avoid leading or ambiguous wording to keep your data fair and reliable.

在设计数据收集表时,要仔细考虑你想要记录什么。使用有标题的清晰表格,并且始终注明计量单位。当你向人们提问时,避免诱导性或含糊不清的措辞,以保持数据的公正和可靠。


3. Frequency Tables and Grouped Data | 频率表与分组数据

Raw data is often messy and hard to read. A frequency table organises data by showing each value or category alongside how many times it occurs (its frequency). For small sets of numerical data, we simply list the values and tally up the frequencies.

原始数据通常杂乱而难以阅读。频率表通过展示每一个数值或类别及其出现的次数(频数),将数据组织起来。对于较小的数值数据集,我们只需列出数值并用计数标记来累计频数。

When there are many different values, we use grouped frequency tables. We split the data into equal class intervals, such as 0–9, 10–19, etc. Each group must not overlap, and the intervals should cover the entire range of the data.

当数值种类很多时,我们使用分组频率表。我们将数据分成等距的组距,例如 0–9, 10–19 等等。每个组不得重叠,而且组距应覆盖数据的整个范围。

From a grouped frequency table, we can estimate the mean, draw a histogram, or find the modal class (the group with the highest frequency). Remember that in a grouped table we lose some detail about individual values.

利用分组频率表,我们可以估算平均数、绘制直方图,或找出众数组(频数最高的组)。请记住,在分组表中我们会丢失关于个别数值的一些详细信息。


4. Bar Charts and Pictograms | 条形图与象形图

A bar chart displays categorical data using rectangular bars. The height or length of each bar represents the frequency. Bars are drawn with equal width and gaps between them to show that the categories are separate.

条形图用矩形条来展示分类数据。每个条的高度或长度代表频数。条形的宽度相等,并且在它们之间留有间隙,以表明各个类别是独立的。

Pictograms use simple pictures or symbols to represent data. Each symbol stands for a certain number of items. A key must be provided to explain what one symbol means. Pictograms are engaging but can be tricky when a frequency is not an exact multiple of the symbol value – in that case you may need to show half a symbol.

象形图使用简单的图画或符号来表示数据。每一个符号代表一定数量的项目。必须提供图例来解释每个符号的含义。象形图很吸引人,但当频数不是符号值的精确倍数时可能比较棘手——这时你可能需要画出半个符号。

When drawing bar charts or pictograms, always label your axes (for bar charts) and give your chart a clear title. For bar charts, the vertical axis must start from zero to avoid creating a misleading impression of the differences between categories.

在绘制条形图或象形图时,始终要标注坐标轴(条形图)并给图表一个清晰的标题。对于条形图,纵轴必须从零开始,以免在类别差异上造成误导性的视觉印象。


5. Pie Charts | 饼图

A pie chart is a circular graph divided into sectors. Each sector represents a category, and its angle is proportional to the frequency of that category. The entire circle (360°) represents the total data set.

饼图是一种被分成多个扇形的圆形图表。每个扇形代表一个类别,其圆心角与该类别的频数成比例。整个圆(360°)代表整个数据集。

To draw a pie chart, first calculate the angle for each category using the formula: angle = (frequency ÷ total frequency) × 360°. Then use a protractor to measure and draw each sector. Use a different colour for each sector and include a legend or labels.

要绘制饼图,首先使用公式计算每个类别的角度:角度 = (频数 ÷ 总频数)× 360°。然后使用量角器测量并画出每个扇形。每个扇形用不同的颜色,并要包含图例或标签。

Pie charts are excellent for showing proportions at a glance, but they are less useful when comparing many categories or when categories have very similar frequencies. The sectors should always add up to 360°.

饼图非常适合一目了然地显示比例,但在比较许多类别或类别频数非常相近时,它的作用就小一些。所有扇形的角度总和必须为 360°。


6. Line Graphs and Time Series | 折线图与时间序列

A line graph is used to show how a quantity changes over time or in relation to another continuous variable. Points are plotted and joined with straight lines. Time is usually placed on the horizontal axis.

折线图用于展示一个量如何随时间变化,或与另一个连续变量的关系。描出数据点并用直线连接起来。时间通常放在横轴上。

When the data is recorded at regular time intervals, we call it a time series. Examples include daily temperatures, monthly rainfall or a company’s annual profit. Time series graphs help you spot trends, seasonal patterns and sudden changes.

当数据在有规律的时间间隔上记录时,我们称之为时间序列。例子包括每日气温、每月降雨量或一家公司的年利润。时间序列图能帮助你发现趋势、季节性模式和突发的变动。

When reading a line graph, check the scales on both axes carefully. A broken line or different starting point on the vertical axis can be used to draw attention, but you should always look at the numbers to avoid being misled by the visual steepness of the line.

阅读折线图时,要仔细检查两条坐标轴的刻度。纵轴上使用截断线或不同的起始值可以吸引注意力,但你始终应该看数字,以免被线段视觉上的陡峭程度所误导。


7. Scatter Graphs and Correlation | 散点图与相关性

A scatter graph (or scatter plot) shows the relationship between two numerical variables. Each point on the graph represents a pair of values. One variable goes on the horizontal axis, the other on the vertical axis.

散点图(散点图)展示两个数值变量之间的关系。图上的每一个点代表一对数值。一个变量放在横轴上,另一个放在纵轴上。

If the points cluster roughly along a straight line with a positive slope, we say there is a positive correlation (as one variable increases, the other tends to increase). A negative correlation means that as one variable increases, the other tends to decrease. If the points are spread out with no clear pattern, there is no correlation.

如果这些点大致沿着一条具有正斜率的直线聚集,我们说存在正相关(当一个变量增大时,另一个也趋于增大)。负相关意味着一个变量增大时,另一个趋于减小。如果点分散且没有清晰的规律,则没有相关性。

Correlation does not imply causation. Even if two variables show a strong correlation, it does not prove that one causes the other. A line of best fit can be drawn through the points to summarise the relationship, and this line can be used to estimate unknown values.

相关性并不意味着因果关系。即使两个变量表现出强相关,也不能证明一个导致另一个。可以通过数据点画一条最佳拟合线来总结这种关系,并且可以用这条线来估计未知的数值。


8. Averages: Mean, Median, Mode | 平均值:平均数、中位数、众数

An average is a single value that summarises the centre of a data set. The three most common averages are the mean, the median and the mode. Each one gives a different type of ‘typical’ value.

平均值是一个概括数据中心趋势的单一数值。三种最常见的平均值是平均数、中位数和众数。每一种都给出不同类型的“典型”数值。

The mean is calculated by adding up all the values and then dividing by the number of values. If we have values x₁, x₂, …, xₙ, then

Mean = (x₁ + x₂ + … + xₙ) ÷ n

The median is the middle value when the data is arranged in order. If there are two middle numbers, the median is the mean of those two. The mode is simply the value that occurs most often. A data set may have one mode, more than one mode, or no mode at all.

平均数是通过将所有数值相加,然后除以数值的个数来计算。如果我们有数值 x₁, x₂, …, xₙ,那么

平均数 = (x₁ + x₂ + … + xₙ) ÷ n

中位数是将数据按顺序排列后位于中间位置的数值。如果中间有两个数,中位数就是这两个数的平均数。众数就是出现次数最多的那个数值。一个数据集可能有一个众数、多个众数,或者根本没有众数。

Choosing the right average is important. The mean is affected by extreme values (outliers), while the median is resistant to them. The mode is useful for non-numerical data, like finding the most popular crisp flavour.

选择正确的平均值很重要。平均数会受极端值(异常值)的影响,而中位数不受它们影响。众数对非数值数据很有用,例如找出最受欢迎的薯片口味。


9. Range and Spread | 极差与数据分散程度

The range is a simple measure of how spread out the data is. It is the difference between the largest and the smallest values.

Range = Largest value – Smallest value

极差是一个衡量数据分散程度的简单指标。它是最大值和最小值之间的差。

极差 = 最大值 – 最小值

A small range means the data points are clustered closely together; a large range means they are widely spread. The range is easy to calculate but can be greatly influenced by a single unusually high or low value.

极差小意味着数据点紧密地聚集在一起;极差大意味着数据分布很分散。极差容易计算,但可能会受到单个异常高或异常低数值的很大影响。

Together with an average, the range gives a quick summary of a data set. For example, if Class A has a mean of 60% with a range of 30%, and Class B has the same mean but a range of 80%, we know Class B’s scores are much more varied.

结合平均值,极差能给出对数据集的快速总结。例如,如果 A 班的平均分为 60%,极差为 30%,而 B 班平均分相同,极差却是 80%,我们就知道 B 班的分数差异要大得多。


10. Introduction to Probability | 概率导论

Probability is a measure of how likely an event is to happen. The probability of an event can be written as a fraction, decimal or percentage. It always lies between 0 (impossible) and 1 (certain).

概率是衡量一个事件发生可能性的量度。一个事件的概率可以用分数、小数或百分数表示。它总是在 0(不可能)和 1(必然)之间。

The probability of an event is found by dividing the number of ways the event can happen by the total number of possible outcomes – provided all outcomes are equally likely.

P(event) = Number of favourable outcomes ÷ Total number of outcomes

事件的概率通过将该事件可以发生的方式数除以所有可能结果的总数来求得——前提是所有结果等可能。

P(事件) = 有利结果数 ÷ 总结果数

Words such as ‘likely’, ‘unlikely’, ‘evens’ (50-50 chance), ‘certain’ and ‘impossible’ are used in everyday language to describe probability. The probability scale places these words on a line from 0 to 1.

像“很可能”、“不太可能”、“对等机会”(五五开)、“必然”和“不可能”这样的词语在日常语言中用来描述概率。概率尺度将这些词语放在从 0 到 1 的数轴上。


11. Experimental vs Theoretical Probability | 实验概率与理论概率

Theoretical probability is what we expect to happen based on equally likely outcomes. For example, when tossing a fair coin, the theoretical probability of getting a head is 1/2.

理论概率是我们基于所有结果等可能而期望发生的情况。例如,抛一枚均匀硬币时,得到正面的理论概率是 1/2。

Experimental probability (or relative frequency) comes from actually carrying out a trial or experiment. It is calculated as:

Experimental probability = Number of times the event occurs ÷ Total number of trials

实验概率(或相对频率)来自实际进行的试验或实验。其计算公式为:

实验概率 = 事件发生的次数 ÷ 总试验次数

With a small number of trials, the experimental probability may differ from the theoretical probability. However, as the number of trials increases, the experimental probability usually gets closer to the theoretical value. This is called the Law of Large Numbers.

当试验次数较少时,实验概率可能与理论概率有差异。但随着试验次数的增加,实验概率通常会越来越接近理论值。这被称为大数定律。


12. Sample Spaces and Tree Diagrams | 样本空间与树形图

A sample space is the set of all possible outcomes of an experiment. Listing the sample space systematically helps you calculate probabilities accurately. For a single dice roll, the sample space is {1, 2, 3, 4, 5, 6}.

样本空间是一个实验的所有可能结果的集合。系统地列出样本空间有助于你准确地计算概率。对于掷一次骰子,样本空间是 {1, 2, 3, 4, 5, 6}。

For two or more events happening together, a tree diagram is a very useful tool. Each branch represents a possible outcome, and you write the probability on the branch. To find the probability of a combination of events along a path, you multiply the probabilities on the branches.

对于两个或多个事件一起发生的情况,树形图是一个非常有用的工具。每个分支代表一个可能的结果,你需要在分支上写出概率。要找出沿某一路径组合事件的概率,将各分支上的概率相乘。

Always check that the probabilities on branches from a single point add up to 1. Tree diagrams can be used for independent events (where one outcome does not affect another) and for situations with or without replacement.

始终要检查从一个节点出发的各分支概率之和为 1。树形图可用于独立事件(一个结果不影响另一个)以及有放回或无放回的情形。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading