📚 Year 10 SQA Statistics: Core Knowledge Review | Year 10 SQA 统计:核心知识点梳理
This article provides a comprehensive review of the key topics in statistics for Year 10 students following the SQA curriculum. It covers data types, collection methods, sampling, graphical representation, measures of central tendency and spread, correlation, probability, and the normal distribution. Each concept is explained with clear examples to help you build a solid foundation and prepare effectively for assessments.
本文全面梳理了针对 SQA 课程的 Year 10 统计核心知识点,涵盖数据类型、收集方法、抽样、图表表示、集中趋势与离散程度的度量、相关性、概率以及正态分布。每个概念都配有清晰的解释和示例,帮助你打下扎实基础,高效备考。
1. Types of Data | 数据类型
Data in statistics is broadly classified as quantitative or qualitative. Quantitative data deals with numbers and can be measured or counted. It is further split into discrete data, which can only take certain values (e.g., number of siblings), and continuous data, which can take any value within a range (e.g., mass, temperature). Qualitative data, also called categorical data, describes qualities or categories, such as hair colour or type of pet. Additionally, data can be primary (collected first‑hand by the researcher) or secondary (obtained from existing sources like websites or books).
统计学中的数据通常分为定量数据和定性数据。定量数据涉及数字,可测量或可计数,进一步分为离散数据(只能取特定值,如兄弟姐妹数量)和连续数据(可在一定范围内取任意值,如质量、温度)。定性数据也称类别数据,描述性质或类别,如头发颜色或宠物类型。此外,数据还可分为一手数据(由研究者亲自收集)和二手数据(来自现有来源,如网站或书籍)。
2. Data Collection Methods | 数据收集方法
Reliable conclusions depend on how data is gathered. Common methods include surveys, experiments, and observational studies. A survey uses questionnaires or interviews to obtain opinions or facts from a sample. Experiments involve manipulating one variable to see its effect on another, while observational studies record data without interference. It is crucial to design questions that are unbiased and clear, and to consider ethical issues such as confidentiality and consent.
可靠的结论取决于数据的收集方式。常用方法有调查、实验和观察研究。调查通过问卷或访谈从样本中获取意见或事实;实验通过操控一个变量来观察其对另一个变量的影响;观察研究则在无干预的情况下记录数据。设计问题时应确保无偏且清晰,并考虑保密和知情同意等伦理问题。
3. Sampling Techniques | 抽样技术
When it is impractical to survey an entire population, a sample is selected. Random sampling gives every member of the population an equal chance of being chosen, reducing bias. Stratified sampling divides the population into subgroups (strata) and takes a random sample from each in proportion to its size. Systematic sampling selects every k‑th individual from a list. Convenience sampling uses readily available participants, but it often leads to bias. The larger the sample size, the more representative it is likely to be.
当调查整个总体不切实际时,就会选取样本。简单随机抽样让总体中每个成员被选中的机会均等,从而减少偏差。分层抽样将总体分成若干子群(层),并按比例从各层随机抽取个体。系统抽样从名单中每隔 k 个抽取一个个体。便利抽样使用容易获得的参与者,但容易产生偏差。样本容量越大,通常代表性越强。
4. Frequency Tables and Diagrams | 频数表与图表
A frequency table organises raw data into a clear summary, showing how often each value or category occurs. The table can include tally marks, frequencies, and sometimes cumulative frequencies. For grouped continuous data, class intervals are used, and we record the frequency of data falling within each interval. The midpoint of a class interval is often used for further calculations.
频数表将原始数据整理成清晰的摘要,显示每个值或类别出现的次数。表中可包含划记、频数,有时还有累计频数。对于分组的连续数据,会使用组距区间,并记录落入每个区间的频数。组中值常用于后续计算。
5. Bar Charts and Pie Charts | 条形图和饼图
Bar charts display categorical data with rectangular bars whose lengths are proportional to the frequencies. Bars should be of equal width and separated by gaps to emphasise that the categories are distinct. Pie charts represent proportions by dividing a circle into sectors; each sector’s angle is (frequency / total) × 360°. A key or labels help identify the categories. These charts make it easy to compare parts of a whole at a glance.
条形图用矩形条表示类别数据,条的长度与频数成正比。条形应等宽并留有间隙,以强调类别是离散的。饼图通过将圆分割成扇形来表示比例,每个扇形的角度为(频数 / 总数)× 360°。图例或标签有助于识别类别。这些图表让人能够一目了然地比较各部分的占比。
6. Histograms and Frequency Polygons | 直方图和频数多边形
Histograms are used for continuous data or grouped discrete data. Unlike bar charts, the bars touch each other to reflect the continuous scale. The area of each bar is proportional to the frequency, so if class widths are unequal, we use frequency density = frequency ÷ class width. A frequency polygon is drawn by plotting the midpoints of class intervals against frequency and joining the points with straight lines, often with the polygon closed at both ends at zero frequency.
直方图用于连续数据或分组离散数据。与条形图不同,直方图的条形紧挨在一起,以体现连续尺度。每个条形的面积与频数成正比,因此如果组距宽度不相等,需要使用频数密度 = 频数 ÷ 组距宽度。频数多边形通过绘制组距中点对应的频数,并用直线连接各点而成,通常两端会连至频数为零的位置。
7. Scatter Graphs and Correlation | 散点图与相关性
A scatter graph plots bivariate data (pairs of values) on coordinate axes to reveal relationships. If points show an upward trend, there is positive correlation; a downward trend indicates negative correlation. When points are randomly scattered, there is no correlation. Correlation does not imply causation. A line of best fit can be drawn through the points to model the trend and make predictions, either by eye or using the mean point.
散点图将双变量数据(成对值)绘制在坐标轴上以揭示关系。如果点呈上升趋势,则存在正相关;下降趋势表示负相关。如果点随机分布,则无相关性。相关性并不意味着因果关系。可以通过点群画一条最佳拟合线来建立趋势模型并进行预测,可凭目测绘制,也可利用均值点绘制。
8. Measures of Central Tendency | 集中趋势的度量
The three main averages are the mean, median, and mode. The mean is the sum of all values divided by the number of values. The median is the middle value when data is ordered; if there is an even number of values, it is the mean of the two middle numbers. The mode is the most frequently occurring value. Each measure has its strengths: the mean uses all data but is affected by outliers; the median is robust against outliers; the mode is useful for categorical data.
三种主要的平均数是均值、中位数和众数。均值为所有数值之和除以数值个数。中位数是将数据排序后位于中间的值;如果数值个数为偶数,则为中间两个数的均值。众数是出现频率最高的值。每种度量方式各有优点:均值利用了所有数据,但易受异常值影响;中位数对异常值稳健;众数适用于类别数据。
Mean x̄ = Σx / n
均值 x̄ = Σx / n
9. Measures of Spread | 离散程度的度量
Spread tells us how scattered the data is. The range is the simplest measure: maximum value minus minimum value. The interquartile range (IQR) is more resistant to outliers: IQR = Q₃ – Q₁, where Q₁ is the lower quartile (25th percentile) and Q₃ is the upper quartile (75th percentile). The standard deviation measures how far values deviate from the mean on average. For a sample, it is calculated as:
离散程度反映数据的分散情况。极差是最简单的度量:最大值减去最小值。四分位距 (IQR) 对异常值更具抗干扰性:IQR = Q₃ – Q₁,其中 Q₁ 为下四分位数(第25百分位数),Q₃ 为上四分位数(第75百分位数)。标准差衡量各数值与均值的平均偏差。对于样本,其计算公式为:
s = √[ Σ(x – x̄)² / (n – 1) ]
s = √[ Σ(x – x̄)² / (n – 1) ]
A larger standard deviation indicates more spread. The five‑number summary (minimum, Q₁, median, Q₃, maximum) is often used to construct box plots, which visually display the centre and spread of a dataset.
标准差越大,数据越分散。五数概括法(最小值、Q₁、中位数、Q₃、最大值)常用于绘制箱线图,直观展示数据集的中心位置和离散程度。
10. Introduction to Probability | 概率入门
Probability measures the chance of an event occurring, expressed as a number between 0 (impossible) and 1 (certain). It can be written as a fraction, decimal, or percentage. Theoretical probability is based on equally likely outcomes: P(event) = number of favourable outcomes / total number of outcomes. Experimental probability comes from trials or experiments: P(event) = frequency of event / total number of trials. The more trials conducted, the closer experimental probability tends to get to theoretical probability – this is the law of large numbers.
概率衡量事件发生的可能性,用一个介于 0(不可能)和 1(必然)之间的数字表示,可写成分数、小数或百分比。理论概率基于等可能结果:P(事件) = 有利结果数 / 总结果数。实验概率来自试验或实验:P(事件) = 事件发生的频数 / 总试验次数。进行的试验次数越多,实验概率往往越接近理论概率——这就是大数定律。
11. Standard Deviation and the Normal Distribution | 标准差与正态分布
Many natural datasets follow a bell‑shaped curve known as the normal distribution. It is symmetric about the mean, and the spread is controlled by the standard deviation. In a normal distribution, about 68% of data lies within 1 standard deviation of the mean, 95% within 2 standard deviations, and 99.7% within 3 standard deviations. This is known as the empirical rule. Knowing the mean and standard deviation allows you to estimate proportions and make predictions about normally distributed data.
许多自然数据集服从钟形曲线,即正态分布。它关于均值对称,其离散程度由标准差控制。在正态分布中,大约 68% 的数据落在均值 ±1 个标准差的范围内,95% 落在 ±2 个标准差内,99.7% 落在 ±3 个标准差内。这被称为经验法则。知道均值和标准差,就可以估计比例并对正态分布的数据做出预测。
| Interval | 区间 | Approximate percentage | 近似百分比 |
|---|---|---|---|
| μ ± σ | μ ± σ | 68% | 68% |
| μ ± 2σ | μ ± 2σ | 95% | 95% |
| μ ± 3σ | μ ± 3σ | 99.7% | 99.7% |
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导