📚 Year 10 SQA Statistics: Summer Preparation and Bridging Course | Year 10 SQA 统计:暑期预习与衔接课程
Welcome to your essential summer bridge into SQA Statistics. Whether you are moving from Year 9 or simply refreshing your skills, this article covers the core concepts you will need. We will explore data, probability, and the tools that turn numbers into meaningful stories. Let’s build your confidence before the new term starts.
欢迎来到你不可或缺的 SQA 统计暑期衔接课程。无论你是从 Year 9 升入还是只是巩固技能,本文涵盖了你所需的核心概念。我们将探索数据、概率以及将数字转化为有意义故事的工具。让我们在新学期开始前建立你的信心。
1. What is Statistics? | 什么是统计?
Statistics is the science of collecting, organising, analysing, and interpreting data. It helps us make sense of the world around us, from predicting weather to evaluating medical treatments. In your SQA course, you will learn to handle data responsibly, identify patterns, and draw conclusions in the presence of uncertainty. You will discover that statistics is not just about numbers—it is about gaining insight and making evidence-based decisions.
统计是收集、整理、分析和解释数据的科学。它帮助我们理解周围的世界,从预测天气到评估医疗方案。在 SQA 课程中,你将学会负责任地处理数据、识别模式,并在存在不确定性时得出结论。你将发现统计不仅仅是关于数字——它更是关于获取洞察并做出基于证据的决策。
Statistics is broadly divided into descriptive and inferential statistics. Descriptive statistics summarise data using measures like the mean and graphs; inferential statistics allow us to make predictions or test hypotheses using sample data. In Year 10, the focus is on descriptive statistics and the foundations of probability, setting the stage for more advanced work later.
统计大致分为描述性统计和推断性统计。描述性统计使用均值和图形来概括数据;推断性统计则允许我们使用样本数据做出预测或检验假设。在 Year 10,重点在于描述性统计和概率基础,为后续更高阶的学习奠定基础。
2. Types of Data | 数据类型
Data can be classified as qualitative (categorical) or quantitative (numerical). Categorical data represent qualities or groups, such as hair colour, favourite subject, or yes/no responses. These are often displayed in bar charts or pie charts. Numerical data are measurements or counts, and they can be further split into discrete and continuous types.
数据可分为定性(分类)数据和定量(数值)数据。分类数据表示性质或组别,如头发颜色、最喜爱的科目或是/否回答。这类数据通常用条形图或饼图展示。数值数据是度量值或计数,并可以进一步分为离散型和连续型。
Discrete data arise from counting and take only certain isolated values. Examples include the number of students in a class (25, 26, 27…) or the score on a dice. Continuous data come from measuring and can take any value within a range. Examples include height, mass, and time taken to finish a race. Recognising the data type is crucial because it determines which statistical tools and graphs are appropriate.
离散数据来自计数,只能取某些孤立的值。例如班级人数(25, 26, 27…)或骰子的点数。连续数据来自测量,可以取某个范围内的任何值。例如身高、质量和完成比赛所需的时间。识别数据类型至关重要,因为它决定了哪些统计工具和图形是合适的。
3. Collecting Data | 数据收集
High-quality statistics begin with reliable data. Primary data are collected directly by the researcher for a specific purpose. Methods include questionnaires, interviews, and experiments. Secondary data are obtained from existing sources such as government databases, academic papers, or historical records. While secondary data save time, you must check their reliability and suitability for your own investigation.
高质量的统计始于可靠的数据。一手数据是研究者为特定目的直接收集的。方法包括问卷调查、访谈和实验。二手数据来自现有来源,如政府数据库、学术论文或历史记录。虽然二手数据节省时间,但必须检查其可靠性以及对自己研究的适用性。
When designing a survey, questions must be clear, unbiased, and easy to answer. Avoid leading questions such as “Don’t you agree that statistics is the most exciting subject?” Instead, ask neutral questions. Sampling will be covered later, but remember that well-planned data collection reduces errors and increases the validity of your conclusions.
在设计调查时,问题必须清晰、无偏见且易于回答。避免诱导性问题,如“难道你不认为统计是最令人兴奋的学科吗?”应改为提出中立的问题。抽样将在后面讨论,但请记住,精心规划的数据收集能减少误差并提高结论的有效性。
4. Organising Data: Frequency Tables | 数据整理:频数表
A frequency table is a simple way to organise raw data. For a list of test scores or survey responses, we can tally each occurrence and record the frequency. For grouped data, we define class intervals and count how many observations fall into each interval. This is especially useful for continuous data or large data sets.
频数表是整理原始数据的简单方法。对于一列测试分数或调查回答,我们可以用划记统计每次出现并记录频数。对于分组数据,我们定义组距并计量落入每个区间的观测值个数。这对连续数据或大数据集尤其有用。
| Score (x) | Tally | Frequency (f) |
|---|---|---|
| 5 | || | 2 |
| 6 | |||| | 5 |
| 7 | ||| | 3 |
| 8 | | | 1 |
From a frequency table, we can also compute cumulative frequency (running total) and relative frequency (proportion of the total). Cumulative frequency helps in finding medians and quartiles, while relative frequency can be expressed as a fraction, decimal, or percentage.
从频数表中,我们还可以计算累积频数(滚动总和)和相对频数(占总体的比例)。累积频数有助于求中位数和四分位数,而相对频数可以用分数、小数或百分比表示。
5. Displaying Data: Charts and Graphs | 数据展示:图表
Different graphs suit different data types. Bar charts are ideal for categorical data and discrete numerical data; each bar’s height represents the frequency. Pie charts display proportions of a whole, but avoid them when there are many categories. Histograms are used for grouped continuous data: the area of each bar is proportional to frequency, and bars touch to show continuity.
不同的图表适用于不同的数据类型。条形图适合分类数据和离散数值数据;每个条形的高度代表频数。饼图展示整体中的比例,但在类别较多时应避免使用。直方图用于分组连续数据:每个条形的面积与频数成比例,且条形相连以显示连续性。
Line graphs show trends over time, such as monthly rainfall. Stem-and-leaf diagrams keep the original data visible while showing shape—perfect for small data sets. Box plots (box-and-whisker diagrams) summarise a distribution using the five-number summary: minimum, lower quartile, median, upper quartile, and maximum. These will be revisited when you study measures of spread.
折线图显示随时间变化的趋势,如每月降雨量。茎叶图在显示分布形态的同时保留了原始数据——非常适合小数据集。箱线图(箱须图)利用五数概括法:最小值、下四分位数、中位数、上四分位数和最大值来总结分布。在学习离散程度度量时我们还会再次提及。
6. Measures of Central Tendency | 集中趋势度量
Central tendency tells us where the centre of a data set lies. The three main measures are the mean, median, and mode. The mean (average) is calculated by summing all values and dividing by the number of values. We use the symbol x̄ for a sample mean:
集中趋势告诉我们数据集的中心位于哪里。三种主要度量是均值、中位数和众数。均值(平均数)通过将所有值求和再除以值的个数来计算。我们用符号 x̄ 表示样本均值:
x̄ = Σx / n
The median is the middle value when data are ordered. If n is even, take the mean of the two middle values. The mode is the most frequently occurring value; there can be more than one mode or none at all. Each measure has strengths: the mean uses all data but is affected by outliers, the median is robust to outliers, and the mode is useful for categorical data.
中位数是数据排序后的中间值。如果 n 为偶数,则取中间两个值的平均数。众数是出现频率最高的值;可以有多个众数,也可以没有。每种度量各有优势:均值使用了所有数据但受异常值影响,中位数对异常值稳健,众数对分类数据有用。
7. Measures of Spread | 离散程度度量
Spread describes how varied the data are. The simplest measure is the range: maximum value minus minimum value. However, the range only considers extremes. The interquartile range (IQR) is more robust: IQR = Q₃ − Q₁, where Q₁ is the lower quartile (25th percentile) and Q₃ is the upper quartile (75th percentile). The IQR captures the middle 50% of the data.
离散程度描述数据的变异程度。最简单的度量是全距:最大值减最小值。但是,全距只考虑极端值。四分位距(IQR)更稳健:IQR = Q₃ − Q₁,其中 Q₁ 是下四分位数(第25百分位数),Q₃ 是上四分位数(第75百分位数)。IQR 涵盖了中间50%的数据。
To find quartiles, order the data. The median splits the list into two halves; the lower quartile is the median of the lower half, and the upper quartile is the median of the upper half. When combined with the median, the IQR helps you construct box plots and spot potential outliers. A common rule flags any value below Q₁ − 1.5 × IQR or above Q₃ + 1.5 × IQR as an outlier.
要找到四分位数,先对数据排序。中位数将列表分成两半;下四分位数是下半部分的中位数,上四分位数是上半部分的中位数。与中位数结合使用时,IQR 有助于构建箱线图并识别潜在的异常值。常用规则将任何低于 Q₁ − 1.5 × IQR 或高于 Q₃ + 1.5 × IQR 的值标记为异常值。
8. Introduction to Probability | 概率入门
Probability measures how likely an event is to happen. It is expressed as a number between 0 (impossible) and 1 (certain). For equally likely outcomes, probability is:
概率衡量事件发生的可能性。它用介于0(不可能)和1(必然)之间的数字表示。对于等可能结果,概率为:
P(Event) = Number of favourable outcomes / Total number of outcomes
In SQA statistics, you will work with experimental probability (based on actual trials) and theoretical probability (based on expected outcomes). The basics include listing outcomes systematically using sample space diagrams, such as two-way tables for combined events. Probability can be written as a fraction, decimal, or percentage.
在 SQA 统计中,你将使用实验概率(基于实际试验)和理论概率(基于预期结果)。基础知识包括使用样本空间图系统地列出所有结果,例如用双向表处理组合事件。概率可以写成分数、小数或百分比。
Key rules: probabilities of all possible outcomes sum to 1. For mutually exclusive events, P(A or B) = P(A) + P(B). For independent events, P(A and B) = P(A) × P(B). Understanding these rules prepares you for more complex problems involving tree diagrams and conditional probability.
关键规则:所有可能结果的概率之和为1。对于互斥事件,P(A 或 B) = P(A) + P(B)。对于独立事件,P(A 且 B) = P(A) × P(B)。理解这些规则为你解决涉及树形图和条件概率的更复杂问题做好了准备。
9. Probability Trees and Diagrams | 概率树图与图表
Probability tree diagrams are a powerful tool for displaying sequences of events. Each branch is labelled with a probability, and the probabilities on branches from the same point must add to 1. To find the probability of a combined sequence of events, multiply along the branches. If multiple sequences satisfy the outcome, add their probabilities.
概率树图是展示事件序列的强大工具。每个分枝都标有概率,且从同一点出发的分枝概率之和必须为1。要计算一系列组合事件的概率,沿分枝相乘。如果有多个序列满足该结果,则将它们的概率相加。
Consider two independent events: drawing a red sweet and then a blue sweet from a bag with replacement. The tree shows both stages, and the probability of “red then blue” is P(red) × P(blue). Without replacement, probabilities change after each draw, so they must be updated on the second set of branches. Tree diagrams also help with conditional probability when combined with Venn diagrams or two-way tables.
考虑两个独立事件:从一个袋子里有放回地先取出一颗红色糖果再取出一颗蓝色糖果。树形图显示两个阶段,“先红后蓝”的概率为 P(红) × P(蓝)。无放回时,每次抽取后的概率会改变,因此必须在第二组分枝上更新。当与韦恩图或双向表结合时,树形图也有助于处理条件概率。
10. Correlation and Scatter Plots | 相关性与散点图
Scatter plots (also called scatter graphs) are used to investigate the relationship between two numerical variables. Each point represents a pair of values (x, y). By looking at the pattern of points, we can describe the correlation as positive (as one variable increases, so does the other), negative (as one increases, the other decreases), or no correlation.
散点图(也称散点图)用于研究两个数值变量之间的关系。每个点代表一对值(x, y)。通过观察点的分布模式,我们可以将相关性描述为正相关(一个变量增大,另一个也增大)、负相关(一个增大,另一个减小)或无相关。
Correlation does not imply causation. Two variables may move together because of a third hidden factor or simply by chance. In SQA tasks, you will be expected to draw a line of best fit through the points (by eye or using a ruler) and use it to make predictions. Interpolation (predicting within the data range) is generally reliable; extrapolation (predicting outside the range) is riskier.
相关不一定意味着因果。两个变量可能因第三个隐藏因素或只是偶然而一起变动。在 SQA 任务中,你需要通过点画出最佳拟合线(凭眼力或用直尺),并用它进行预测。内插法(在数据范围内预测)通常可靠;外推法(在范围外预测)风险较大。
11. Sampling Methods | 抽样方法
In statistics, the population is the entire group we want to know about. Since it is often impractical to survey everyone, we select a sample. The goal is for the sample to be representative of the population. A biased sample leads to unreliable conclusions. Common sampling techniques include simple random sampling, where every member has an equal chance of selection, and stratified sampling, where the population is divided into subgroups (strata) and a random sample is taken from each in proportion to its size.
在统计中,总体是我们想了解的整个群体。由于通常不可能调查每个人,我们选择样本。目标是让样本能够代表总体。有偏的样本会导致不可靠的结论。常用的抽样技术包括简单随机抽样,每个成员被选中的机会相等;以及分层抽样,先将总体划分为子组(层),然后按比例从每一层中随机抽取样本。
Other methods include systematic sampling (select every k-th member after a random start) and convenience sampling (choosing easily available individuals, which often introduces bias). Understanding how samples are collected helps you critique survey results and design your own investigations. The size of the sample also matters: larger samples generally give more precise estimates, but they must still be randomly chosen.
其他方法包括系统抽样(随机起点后每隔 k 个成员选一个)和便利抽样(选择容易获取的个体,这往往会引入偏差)。了解样本的收集方式有助于你评判调查结果并设计自己的研究。样本大小也很重要:较大的样本通常给出更精确的估计,但仍然必须随机选择。
12. Bridging from Summer to SQA Success | 从暑期衔接到 SQA 成功
Use the summer weeks to consolidate these foundations. Practice drawing frequency tables for everyday data, calculate means and medians from small data sets, and sketch simple bar charts and scatter plots by hand. Work through probability puzzles and try designing a mini-survey with a friend to test your sampling knowledge. Engage with real data: sports statistics, weather records, or social media polls often provide excellent material.
利用暑期这几周巩固这些基础。练习为日常数据绘制频数表,从小数据集计算均值和中位数,并手绘简单的条形图和散点图。解决概率谜题,尝试与朋友一起设计一个小型调查以检验你的抽样知识。接触真实数据:体育统计、天气记录或社交媒体投票通常提供了极好的素材。
When you return to school, you will build on this platform with more formal hypothesis testing, standard deviation, and the normal distribution. A confident grasp of the basics will ensure you can focus on reasoning and interpretation rather than getting stuck on arithmetic. Remember, statistics is a skill that improves with practice—and it is a lifelong tool for thinking critically about the world.
当你回到学校时,你将在这个平台上进一步学习更正式的假设检验、标准差和正态分布。对基础知识的自信掌握将确保你能专注于推理和解释,而不是被算术所困。请记住,统计是一项通过练习来提高的技能——它是你终身批判性思考世界的工具。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导