📚 Year 10 Eduqas Statistics: Summer Prep and Bridging Course | 暑期预习与衔接课程
Welcome to your Year 10 Statistics bridging course! This resource is designed to help you move smoothly from Year 9 mathematics into the world of GCSE Statistics with Eduqas. You will revisit key ideas and gain a head start on topics such as data handling, probability and correlation, setting you up for a confident and successful year ahead.
欢迎来到十年级统计衔接课程!本资源旨在帮助你从九年级数学平稳过渡到 Eduqas 的 GCSE 统计学世界。你将回顾关键概念,提前接触数据处理、概率和相关等主题,为即将到来的一年打下自信而成功的基础。
1. The Statistical Enquiry Cycle | 统计探究循环
In GCSE Statistics, you will think and work like a statistician. Every investigation follows a structured framework called the statistical enquiry cycle, often remembered as PPDAC: Problem, Plan, Data, Analysis, Conclusion. Understanding this cycle is the first step towards mastering the subject.
在 GCSE 统计学中,你将像统计学家一样思考和工作。每一个调查都遵循一个结构化的框架,称为统计探究循环,通常记为 PPDAC:问题、计划、数据、分析、结论。理解这一循环是掌握这门学科的第一步。
The process begins with defining a clear problem or question – for example, ‘Do students who eat breakfast perform better in tests?’ Next comes the plan: deciding what data to collect, how to sample, and what methods to use. Then you collect your data following your plan carefully.
这一过程从明确问题或疑问开始——例如“吃早餐的学生考试成绩是否更好?”接下来是计划:决定收集哪些数据、如何抽样以及使用什么方法。然后你根据计划仔细收集数据。
Once you have your data, you move to analysis – creating graphs, calculating averages and measures of spread, and interpreting patterns. Finally, you write a conclusion that answers the original problem, discusses any limitations and suggests improvements. This cycle is iterative, meaning you might go back and refine your question or plan after initial findings.
获得数据后,你进入分析阶段——制作图表、计算平均数和离散程度,并解读模式。最后,你撰写结论来回答最初的问题,讨论任何局限性并提出改进建议。这个循环是迭代的,意味着你可能根据初步发现回过头来完善问题或计划。
2. Types of Data | 数据的类型
Data is at the heart of statistics, and we classify it in different ways. The first important distinction is between qualitative data (categories or labels, like favourite colour or type of pet) and quantitative data (numbers that measure or count something). Understanding this helps you choose the right tools for analysis.
数据是统计学的核心,我们以不同方式对其进行分类。第一个重要区分是定性数据(类别或标签,如最喜欢的颜色或宠物类型)与定量数据(测量或计数的数字)。理解这一点有助于你选择合适的分析工具。
Quantitative data can be further split into discrete and continuous. Discrete data can only take specific, separate values – for example, the number of students in a class or the roll of a dice. Continuous data, on the other hand, can take any value within a range and is usually measured: height, time, temperature. Remember, if you can measure it to more and more decimal places, it is continuous.
定量数据进一步分为离散型和连续型。离散数据只能取特定的、分开的值——例如班级学生人数或掷骰子的点数。而连续数据可以取一个范围内的任何值,并且通常是测量得到的:身高、时间、温度。请记住,如果你能将它测量到越来越多的小数位,它就是连续数据。
3. Sampling Methods | 抽样方法
Because it is rarely possible to collect information from an entire population, we select a sample. The way you choose your sample has a huge effect on how trustworthy your conclusions are. In Year 10, you will learn about several sampling techniques, each with its strengths and weaknesses.
由于很少有可能从整个总体中收集信息,我们会选取一个样本。选择样本的方式对你结论的可信度有着巨大影响。在十年级,你将学习几种抽样技术,每种都有其优点和缺点。
A simple random sample gives every member of the population an equal chance of being picked, often using a random number generator. It avoids bias but can be impractical for large, spread‑out populations. Stratified sampling splits the population into groups (strata) based on a characteristic like age, then takes a random sample from each group in proportion to its size. This guarantees a representative mix.
简单随机抽样让总体中的每个成员都有同等被选中的机会,通常使用随机数生成器。它避免了偏差,但对于庞大、分散的总体可能不切实际。分层抽样根据某个特征(如年龄)将总体分成几个组(层),然后按比例从每组中随机抽样。这确保了代表性的混合。
Systematic sampling selects individuals at regular intervals from a list, for instance every 10th name. It is simple to carry out but can introduce bias if there is a hidden pattern. Convenience sampling uses people who are easy to reach, like friends or the first passers‑by; it is quick but very likely to be biased. Always think about which method suits the context and how to minimise bias.
系统抽样按固定间隔从一个列表中选取个体,例如每第10个名字。它操作简单,但如果存在隐藏模式则可能引入偏差。便利抽样使用容易接触到的人,如朋友或第一批路人;它很快,但极有可能产生偏差。要始终思考哪种方法适合情境以及如何尽量减少偏差。
4. Organising and Displaying Data | 数据的整理与展示
Raw data are messy and hard to interpret. Organising them into tables and drawing appropriate diagrams bring patterns to life. A frequency table is often the starting point: it records how many times each value or group of values occurs.
原始数据杂乱且难以解读。将它们整理成表格并绘制合适的图表能让模式生动起来。频数表通常是起点:它记录了每个值或每组值出现的次数。
For categorical data, bar charts and pie charts are common. A bar chart uses equal‑width bars with heights proportional to frequency, while a pie chart shows proportions of a whole. When dealing with continuous data, you will create grouped frequency tables and draw histograms. Remember, in a histogram the area of each bar represents frequency, so you will often work with frequency density (frequency ÷ class width).
对于分类数据,条形图和饼图很常用。条形图使用等宽的长条,其高度与频数成比例;饼图则显示整体中的比例。处理连续数据时,你会创建分组频数表并绘制直方图。请记住,直方图中每个长条的面积代表频数,因此你经常要用到频率密度(频数 ÷ 组距)。
Other useful diagrams include stem‑and‑leaf plots, which keep all the original data values, and cumulative frequency curves, which help you find medians and quartiles. Your choice of display should make the data’s story clear and honest.
其他有用的图表包括能保留所有原始数据值的茎叶图,以及帮助你找到中位数和四分位数的累积频数曲线。你选择的图表应清晰而诚实地讲述数据的故事。
5. Measures of Central Tendency | 集中趋势量数
Averages summarise the centre of a data set. The three main ones are the mean, the median and the mode. The mean (often just called the average) is calculated by adding all values and dividing by how many there are:
Mean = Σx ÷ n
It uses all the data but can be distorted by extreme values (outliers).
平均数概括了数据集的中心。三个主要量数是平均数、中位数和众数。平均数(通常就称为均值)通过将所有数值相加再除以数值的个数来计算:
平均数 = Σx ÷ n
它用到了所有数据,但可能被极端值(离群值)扭曲。
The median is the middle value when data are ordered. If there is an even number of values, take the mean of the two middle numbers. It is unaffected by outliers, so it is often preferred for skewed distributions or when data contain errors. The mode is simply the most frequent value. A data set can have no mode, one mode, or multiple modes (bimodal). Modes are especially useful for categorical data, where means make no sense.
中位数是将数据排序后的中间值。如果有偶数个数值,则取中间两个数的平均数。它不受离群值影响,因此对于偏态分布或数据包含错误的情况,常被优先选用。众数是最频繁出现的值。一个数据集可以没有众数、有一个众数或多个众数(双众数)。众数对于分类数据尤其有用,因为分类数据求平均数没有意义。
6. Measures of Spread | 离散程度量数
An average on its own can be misleading if you do not know how spread out the data are. The simplest measure of spread is the range (largest value minus smallest value). It is quick to compute but tells you nothing about the shape of the distribution.
如果你不知道数据的分散程度,单看平均数可能产生误导。最简单的离散量数是极差(最大值减去最小值)。它计算迅速,但无法告诉你分布的形状。
A more robust measure is the interquartile range (IQR). The lower quartile (Q₁) marks the 25th percentile, the upper quartile (Q₃) the 75th percentile, and the IQR = Q₃ – Q₁. It covers the middle 50% of the data and is resistant to outliers. Box plots use these five‑number summaries (minimum, Q₁, median, Q₃, maximum) to compare distributions visually.
一个更稳健的量数是四分位距(IQR)。下四分位数(Q₁)标记第 25 百分位数,上四分位数(Q₃)标记第 75 百分位数,IQR = Q₃ – Q₁。它覆盖了中间 50% 的数据,且不受离群值影响。箱线图利用这五项数据概括(最小值、Q₁、中位数、Q₃、最大值)来直观地比较分布。
For a detailed picture, we use variance and standard deviation. The population standard deviation σ measures how much, on average, data points deviate from the mean:
σ = √[ Σ(x – μ)² / N ]
where μ is the population mean and N is the population size. A small standard deviation means data are tightly clustered around the mean; a large one shows wider dispersion. In GCSE Statistics, you will learn to calculate these step‑by‑step.
为了获得更详细的信息,我们使用方差和标准差。总体标准差 σ 衡量数据点平均偏离均值的程度:
σ = √[ Σ(x – μ)² / N ]
其中 μ 为总体均值,N 为总体大小。标准差小表示数据紧密围绕在均值周围;标准差大则显示更广的离散。在 GCSE 统计学中,你将逐步学会计算这些量数。
7. Introduction to Probability | 概率入门
Probability is the branch of mathematics that deals with chance and uncertainty. It is written as a number between 0 and 1, where 0 means impossible and 1 means certain. A probability can also be expressed as a fraction, decimal or percentage.
概率是处理随机和不确定性的数学分支。它用 0 到 1 之间的数字表示,0 代表不可能,1 代表必然。概率也可以用分数、小数或百分比表示。
The theoretical probability of an event is the number of favourable outcomes divided by the total number of possible outcomes, assuming all outcomes are equally likely. Experimental probability comes from doing trials or looking at historical data; it may not match the theoretical value exactly, especially with small sample sizes, but it should get closer as more trials are performed (the law of large numbers).
一个事件的理论概率是“有利结果的数量”除以“所有可能结果的总数”,假设所有结果等可能。实验概率来源于进行试验或查看历史数据;它可能并不精确等于理论值,特别是在样本量较小时,但随着试验次数增加,它应越来越接近(大数定律)。
Tree diagrams are powerful tools for listing outcomes of two or more events. You multiply along branches to find the probability of a combination, and add probabilities of different branches to find the probability of an overall event. They help you handle ‘and’ and ‘or’ situations correctly.
树状图是列出两个或多个事件结果的有力工具。你沿着分支相乘求得组合的概率,将不同分支的概率相加求得一个整体事件的概率。它们能帮你正确处理“且”和“或”的情形。
8. Discrete Random Variables | 离散随机变量
When an experiment produces numerical outcomes that are discrete and governed by chance, we define a discrete random variable. For instance, let X be the score when a fair six‑sided die is rolled; X can take values 1, 2, 3, 4, 5 or 6, each with probability 1/6.
当一个实验产生的数值结果是离散的且受机遇支配时,我们定义一个离散随机变量。例如,设 X 为投掷一枚公平六面骰子的点数;X 可取 1、2、3、4、5 或 6,每个概率为 1/6。
A probability distribution table lists all possible values of X alongside their probabilities. The probabilities must add up to 1. You can calculate the expected value of X, denoted E(X), which is the long‑run average if you repeated the experiment many times. It is found by multiplying each value by its probability and summing the results:
E(X) = Σ x · P(X = x)
This concept is important in games of chance, risk analysis and decision‑making.
概率分布表列出 X 的所有可能取值及其概率,这些概率之和必须为 1。你可以计算 X 的期望值 E(X),它表示若重复实验许多次后的长期平均值。计算方法是将每一个值乘以其概率,然后求和:
E(X) = Σ x · P(X = x)
这一概念在机会游戏、风险分析和决策中都很重要。
9. Scatter Diagrams and Correlation | 散点图与相关
Often we want to know whether two variables are linked – for example, does the number of hours spent revising relate to the exam score? A scatter diagram plots paired data on a graph, with one variable on the x‑axis and the other on the y‑axis. Each point represents one item or person.
我们常常想知道两个变量是否存在关联——例如,复习的小时数是否与考试成绩有关?散点图将成对数据绘制在坐标系中,一个变量在 x 轴,另一个在 y 轴。每个点代表一个物品或一个人。
The pattern of points can reveal correlation: positive correlation means as one variable increases, the other tends to increase; negative correlation means as one increases, the other tends to decrease. If the points are scattered randomly, there is no correlation. Correlation does not imply causation – just because two things move together does not mean one causes the other.
点的分布模式可以揭示相关性:正相关意味着一个变量增加时,另一个也倾向于增加;负相关意味着一个增加时,另一个倾向于减少。如果点随机分布,则没有相关性。相关性并不意味着因果关系——仅仅因为两个事物一起变化,并不说明一个导致了另一个。
You will learn to draw a line of best fit by eye and to use it to make estimates. This line can also be calculated precisely using the method of least squares in Year 10–11. The strength of a linear correlation is measured by the correlation coefficient, r, which ranges from –1 to 1. Values close to 1 or –1 show a strong linear link; values around 0 show a weak or no linear link.
你将学会凭目测画出最佳拟合线,并用它进行估计。这条线也可以在十到十一年级用最小二乘法精确计算。线性相关的强度用相关系数 r 来衡量,其取值范围从 –1 到 1。接近 1 或 –1 的值表示强线性关联;接近 0 的值表示弱线性关联或无线性关联。
10. Introduction to Time Series | 时间序列简介
A time series is a set of data collected at regular time intervals – for instance, daily temperatures, monthly sales or yearly population figures. Time series analysis helps us spot patterns and make forecasts. Two key components are trend and seasonality.
时间序列是按固定时间间隔收集的一组数据——例如每日气温、月度销售额或年度人口数字。时间序列分析帮助我们识别模式并做出预测。两个关键组成部分是趋势和季节性。
The trend is the long‑term movement of the series, ignoring the short‑term ups and downs. You can smooth out erratic fluctuations by calculating moving averages. For example, a three‑point moving average replaces each point with the average of itself, the previous point and the next point, producing a steadier line.
趋势是序列的长期走向,忽略短期的起伏。你可以通过计算移动平均来平滑不规则的波动。例如,三点移动平均将每个数据点替换为它自身、前一点和后一点的平均值,从而得到一条更平稳的线。
Seasonality refers to regular, repeating patterns within a fixed period, such as higher ice‑cream sales every summer. By separating trend and seasonal effects, you can make more accurate predictions. In GCSE Statistics, you will draw time‑series graphs, add trend lines, and estimate seasonal variations.
季节性是指在一个固定周期内规律性重复的模式,例如每年夏天冰淇淋销量上升。通过将趋势和季节效应分离开来,你可以做出更准确的预测。在 GCSE 统计学中,你将绘制时间序列图、添加趋势线,并估计季节性变化。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply