Year 8 CCEA Statistics: Core Knowledge Summary | Year 8 CCEA 统计:核心知识点梳理

📚 Year 8 CCEA Statistics: Core Knowledge Summary | Year 8 CCEA 统计:核心知识点梳理

Statistics is about collecting, organising, presenting, and interpreting data to answer questions and make decisions. In Year 8, you build a foundation by learning different data types, how to gather reliable data, and how to display it using charts and graphs. You also begin to summarise data with averages and range, and explore the basics of probability. Mastering these core ideas will help you think critically about the information you see every day and prepare you for more advanced statistical work in later years.

统计学是关于收集、整理、展示和解读数据,以回答问题并做出决策的学科。在 Year 8,你将建立基础,学习不同的数据类型、如何收集可靠的数据,以及如何用图表展示它们。你还将开始用平均数和极差来概括数据,并探索概率的基础知识。掌握这些核心概念能帮助你批判性地看待日常信息,并为日后更深入的统计学习做好准备。

1. Understanding Data Types | 理解数据类型

Data can be classified into two broad types. Qualitative data, also called categorical data, describes qualities or categories, such as eye colour, favourite subject, or car brand. Quantitative data deals with numerical values and can be further split into discrete data and continuous data. Discrete data can only take specific, separate values, usually whole numbers from counting, like the number of books in a bag or goals scored in a match. Continuous data can take any value within a range, like height, mass, or temperature, because it is measured on a scale.

数据可分为两大类。定性数据,也叫分类数据,描述品质或类别,例如眼睛颜色、最喜欢的科目或汽车品牌。定量数据涉及数值,可以进一步分为离散数据和连续数据。离散数据只能取特定、分离的值,通常是计数得出的整数,比如书包里的书本数或比赛中的进球数。连续数据可以在一个范围内取任意值,例如身高、质量或温度,因为它是通过测量尺度得到的。

Recognising the data type is essential because it influences which charts you can use and which averages are meaningful. For instance, you cannot calculate a meaningful mean for qualitative data like colours, but you can find the mode. For continuous data, the mean and median make sense, but the mode might be less useful if all values are different.

识别数据类型至关重要,因为它会影响你可以使用哪些图表以及哪些平均数有意义。例如,你无法为颜色等定性数据计算有意义的均值,但你可以找到众数。对于连续数据,均值和中位数有意义,但如果所有值都不同,众数可能就没那么有用了。


2. Collecting Data | 数据收集

Data collection is the first step in a statistical investigation. You can gather data through surveys, observations, or experiments. When designing a survey, it is vital to write clear, unbiased questions that do not lead the respondent towards a particular answer. Closed questions with specific options often produce data that is easier to analyse, while open questions capture more detail but are harder to summarise.

数据收集是统计调查的第一步。你可以通过调查、观察或实验来收集数据。在设计问卷时,至关重要的是要编写清晰、无偏见的问题,不能引导受访者给出特定答案。带有具体选项的封闭式问题通常能产生更易于分析的数据,而开放式问题能捕捉更多细节,但难以概括。

You also need to decide whether to conduct a census or a sample. A census collects data from every member of the population, giving highly accurate results, but it can be time-consuming and expensive. A sample selects only a part of the population. A good sample should be representative, meaning it fairly reflects the whole population. Random sampling helps avoid bias, but convenience samples can lead to misleading conclusions.

你还需要决定是进行普查还是抽样。普查从总体中的每一个成员收集数据,能给出高度准确的结果,但可能耗时且昂贵。抽样只选取总体的一部分。一个好的样本应该具有代表性,即它能公平地反映整个总体。随机抽样有助于避免偏差,但便利抽样可能导致误导性的结论。


3. Organising Data: Frequency Tables | 数据整理:频数表

Before drawing graphs, raw data needs to be organised. A frequency table lists each distinct data value or category alongside its frequency, which is the number of times it occurs. Tally marks are often used to count frequencies efficiently; every fifth mark is drawn diagonally across the previous four to make groups of five, speeding up the final total.

在绘制图表之前,需要对原始数据进行整理。频数表会列出每个不同的数据值或类别,及其频数,即它出现的次数。计数符号通常用于高效地统计频数;每当数到第五个时,就画一条对角线穿过前四个,从而形成五个一组,加快最终总数的计算。

For continuous data, values are often grouped into equal-sized class intervals, such as 0 ≤ h < 10 cm. This makes the data easier to manage and display, although some detail is lost because you no longer know the exact individual values. When building grouped frequency tables, choose intervals that cover the full range without overlapping.

对于连续数据,数值通常被归入等宽的组距,比如 0 ≤ h < 10 厘米。这使数据更易于管理和展示,尽管会丢失一些细节,因为无法知道每个值的精确大小。在构建分组频数表时,要选择互不重叠且能覆盖整个范围的区间。


4. Bar Charts and Pictograms | 条形图和象形图

A bar chart displays categorical or discrete data using rectangular bars. The categories go on the horizontal axis, and frequency on the vertical axis. The bars must be of equal width and should have equal gaps between them to show that the categories are separate. Always label both axes, give the chart a clear title, and use an appropriate scale that makes differences easy to see.

条形图用矩形条来展示分类或离散数据。类别放在横轴上,频数放在纵轴上。条形必须等宽,并且条形之间应有相等的间距,以表明类别是分开的。务必给两个坐标轴贴上标签,为图表配上清晰的标题,并使用能让差异一目了然的合适刻度。

A pictogram uses simple pictures or symbols to represent a certain number of items. A key must be provided to show what one symbol stands for. Pictograms can be more engaging than bar charts, but they are less precise when fractions of symbols are needed. For example, if one symbol represents 4 people, half a symbol represents 2 people.

象形图用简单的图片或符号来代表一定数量的项目。必须提供一个图例来说明一个符号代表什么。象形图可以比条形图更吸引人,但一旦需要用到部分符号,精确度就会降低。例如,如果一个符号代表4人,那么半个符号就代表2人。


5. Pie Charts | 饼图

A pie chart shows how a whole is divided into its parts. Each slice represents a category, and the size of the slice corresponds to its proportion of the total frequency. To construct a pie chart by hand, you must calculate the angle for each sector using the formula:

Angle = (Frequency ÷ Total frequency) × 360°

饼图展示一个整体如何被分割成各个部分。每个扇形代表一个类别,扇形的大小与其占总频数的比例相对应。要手工绘制饼图,你必须用以下公式计算每个扇形的角度:

角度 = (频数 ÷ 总频数) × 360°

You can also express each slice as a percentage using Percentage = (Frequency ÷ Total frequency) × 100%. When interpreting pie charts, look for the largest and smallest sectors to identify which categories are most or least common. Remember that pie charts do not show actual frequencies unless the total sample size is given.

你也可以用 百分比 = (频数 ÷ 总频数) × 100% 来表示每个扇形的占比。在解读饼图时,要寻找最大和最小的扇形,以确定哪些类别最常见或最不常见。请记住,除非给出总样本量,否则饼图不会显示实际频数。

Below is an example of a calculation table for a pie chart:

Favourite fruit Frequency Calculation Angle
Apple 12 12/30 × 360 144°
Banana 8 8/30 × 360 96°
Orange 10 10/30 × 360 120°

Always check that your angles sum to 360° before drawing.

在绘制前,务必检查你的角度总和是否为360°。


6. Line Graphs and Time Series | 折线图与时间序列

A line graph is particularly useful for showing how a variable changes over time. Time is plotted on the horizontal x-axis, and the other variable, such as temperature or sales, is plotted on the vertical y-axis. A time series is simply a set of data collected at regular time intervals. Points are joined by straight line segments, which makes it easy to spot trends, seasonal patterns, and periods of growth or decline.

折线图特别适用于显示一个变量如何随时间变化。时间绘制在水平的x轴上,另一个变量(如温度或销售额)绘制在垂直的y轴上。时间序列就是一组按固定时间间隔收集的数据。点之间用直线段连接,这有助于轻松发现趋势、季节性模式以及增长或下降的时期。

When you read a line graph, pay close attention to the scales on the axes. A broken scale or a scale that does not start at zero can sometimes exaggerate the changes and mislead the viewer. Always look for the highest and lowest points, and try to describe the overall story the graph is telling. Common descriptions include ‘increasing gradually’, ‘sharply decreasing’, or ‘staying constant’.

阅读折线图时,要密切注意坐标轴上的刻度。断续刻度或不是从零开始的刻度有时会夸大变化,误导读者。始终寻找最高点和最低点,并尝试描述图表所讲述的整体故事。常见的描述有“逐渐增加”、“急剧下降”或“保持恒定”。


7. Scatter Graphs and Correlation | 散点图与相关性

A scatter graph plots two numerical variables to see if there is a relationship between them. One variable is shown on the horizontal axis, the other on the vertical. Each point represents one observation or item. The pattern of the points indicates correlation: positive correlation means that as one variable increases, the other tends to increase; negative correlation means that as one variable increases, the other tends to decrease.

散点图将两个数值变量绘制出来,以观察它们之间是否存在关系。一个变量显示在横轴上,另一个在纵轴上。每个点代表一个观测值或项目。点的分布模式表明了相关性:正相关意味着一个变量增加时,另一个也趋于增加;负相关意味着一个变量增加时,另一个趋于减少。

If the points are scattered randomly with no clear upward or downward pattern, we say there is no correlation. It is important to remember that correlation does not imply causation. Even if two variables move together, it does not necessarily mean that one causes the other. For example, ice cream sales and sunburn cases are positively correlated, but this is because both are related to hot weather, not because ice cream causes sunburn.

如果点随机散布,没有明显的上升或下降模式,我们就说没有相关性。重要的是要记住,相关性并不意味着因果关系。即使两个变量一起变化,也不一定意味着一个导致了另一个。例如,冰激凌销量和晒伤病例呈正相关,但这是因为两者都与炎热天气有关,而不是因为冰激凌导致晒伤。


8. Averages: Mean, Median, Mode | 平均数:均值、中位数、众数

Three main averages summarise a set of data with a single typical value. The mean is what many people call the average; it is calculated by adding up all the values and dividing by how many there are:

Mean = Sum of all data values ÷ Number of values

三种主要的平均数用一个典型值来概括一组数据。均值就是许多人所说的平均数;它的计算方法是将所有数值相加,再除以数值的个数:

均值 = 所有数据值的总和 ÷ 数值的个数

The median is the middle value when the data is arranged in ascending order. If there is an odd number of values, the median is the middle one. If there is an even number, it is the mean of the two middle numbers. The median is not affected by extremely high or low outliers, making it useful for data like house prices or incomes.

中位数是将数据按升序排列后位于中间的值。如果有奇数个值,中位数就是正中间的那个;如果有偶数个值,中位数就是中间两个数的均值。中位数不受极端高或极端低异常值的影响,因此在房价或收入等数据中非常有用。

The mode is the value that occurs most frequently. A data set can have no mode, one mode (unimodal), or more than one mode (bimodal or multimodal). The mode works for both numerical and categorical data, which makes it the only average you can use for things like favourite colour. Choosing the best average depends on the data type and what you want to show.

众数是出现频率最高的值。一个数据集可能没有众数,可能有一个众数(单峰),也可能有多个众数(双峰或多峰)。众数既可用于数值数据也可用于分类数据,这使它成为你唯一可以用于最喜欢的颜色等事物的平均数。选择最佳平均数取决于数据类型和你想要展示的目的。


9. Range and Spread | 极差与离散程度

Averages alone do not tell the full story, because two very different data sets can have the same mean or median. The range measures how spread out the data is. You find it by subtracting the smallest value from the largest value.

Range = Largest value – Smallest value

仅仅有平均数并不能说明全部问题,因为两个截然不同的数据集可能具有相同的均值或中位数。极差衡量的是数据的离散程度。你可以用最大值减去最小值来求得。

极差 = 最大值 – 最小值

A smaller range means the data values are clustered closely together, showing consistency. A larger range indicates greater variability. However, the range only uses the two extreme values and ignores the distribution of the rest of the data. A single outlier can make the range very large, so it is often used together with the median to give a more robust summary.

较小的极差意味着数据值紧密地聚集在一起,显示出较高的一致性。较大的极差则表明变异性更大。但是,极差只用了两个极端值,而忽略了其余数据的分布情况。一个异常值就能使极差变得非常大,因此极差通常与中位数一起使用,以提供更为稳健的概括。


10. Introduction to Probability | 概率入门

Probability is a branch of mathematics closely linked to statistics. It measures how likely events are to happen, using a scale from 0 to 1. A probability of 0 means the event is impossible, 1 means it is certain, and 0.5 means it has an even chance. Probabilities can also be expressed as fractions, decimals, or percentages.

概率是与统计学紧密相连的一个数学分支。它用0到1之间的数值来衡量事件发生的可能性大小。概率为0意味着事件不可能发生,1意味着事件一定会发生,0.5意味着事件有均等的机会。概率也可以用分数、小数或百分数表示。

For equally likely outcomes, the theoretical probability of an event is:

Probability = Number of favourable outcomes ÷ Total number of possible outcomes

对于等可能的结果,一个事件的理论概率为:

概率 = 有利结果的数量 ÷ 所有可能结果的总数

Experimental probability is based on actually carrying out trials, such as rolling dice or spinning spinners. It is calculated using Relative frequency = Number of successful trials ÷ Total number of trials. With more trials, the experimental probability usually gets closer to the theoretical probability. This is known as the law of large numbers. Probability helps us quantify risk and make informed predictions.

实验概率是基于实际进行的试验,比如掷骰子或转指针。它的计算公式为 相对频率 = 成功试验次数 ÷ 总试验次数。试验次数越多,实验概率通常会越接近理论概率。这就是所谓的大数定律。概率帮助我们量化风险,并做出有依据的预测。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading