📚 Core Statistics Knowledge for Year 10 Eduqas | Year 10 Eduqas 统计:核心知识点梳理
This article provides a comprehensive review of the essential statistical topics for Year 10 students following the Eduqas specification. From data collection to probability models, mastering these foundations will help you analyse information, draw valid conclusions and succeed in your exams.
本文全面梳理了Eduqas 考试局 Year 10 统计课程的核心知识点。从数据收集到概率模型,掌握这些基础内容将帮助你分析信息、得出有效结论并在考试中取得好成绩。
1. Types of Data | 数据的类型
Data can be classified as qualitative (categorical) or quantitative (numerical). Quantitative data is further split into discrete (countable, such as number of students) and continuous (measurable, such as height).
数据可分为定性(分类)数据和定量(数值)数据。定量数据又可细分为离散数据(可数,如学生人数)和连续数据(可测量,如身高)。
Qualitative data describes qualities or categories, for example eye colour or types of transport. It is often displayed using bar charts or pie charts.
定性数据描述属性或类别,例如眼睛颜色或交通方式。它通常用条形图或饼图展示。
Discrete data can only take specific values, usually integers. Continuous data can take any value within a range and is measured rather than counted.
离散数据只能取特定值,通常是整数。连续数据可以在一个范围内取任意值,是测量得到的而不是计数得到的。
2. Data Collection: Primary and Secondary Sources | 数据收集:一手资料与二手资料
Primary data is collected first-hand by the researcher for a specific purpose. Secondary data is information that already exists, gathered by someone else.
一手数据是研究者为了特定目的亲自收集的数据。二手数据是已经存在的、由他人收集的信息。
Examples of primary data include questionnaires, interviews and experiments. Secondary data can come from websites, books or government statistics.
一手数据的例子包括问卷、访谈和实验。二手数据可以来自网站、书籍或政府统计数据。
Primary data is often more reliable for a specific investigation but can be time-consuming to collect. Secondary data is quicker to obtain but may not perfectly fit the research question.
对于特定调查,一手数据通常更可靠,但收集起来费时费力。二手数据获取更快,但可能与研究问题不完全匹配。
3. Sampling Methods | 抽样方法
A sample is a subset of a population used to draw conclusions about the whole group. A good sample should be representative and free from bias.
样本是总体中的一个子集,用于对整体做出推断。一个好的样本应具有代表性且没有偏差。
Random sampling gives every member of the population an equal chance of being selected. Systematic sampling selects individuals at regular intervals from an ordered list.
随机抽样让总体中的每个成员都有均等的机会被选中。系统抽样从一个有序列表中按固定间隔选取个体。
Stratified sampling divides the population into groups (strata) and takes a proportional random sample from each group. This ensures all subgroups are fairly represented.
分层抽样将总体划分为多个层,并从每一层按比例随机抽样。这确保了所有子群体都能得到公平代表。
Convenience sampling involves choosing individuals who are easy to reach, which often introduces bias and should be avoided in rigorous statistical work.
便利抽样选取容易接触到的个体,这往往会引入偏差,在严谨的统计工作中应避免使用。
4. Organising and Representing Data | 数据的整理与表示
Frequency tables help organise raw data by showing how many times each value or category occurs. Tally marks are often used during data collection.
频率表通过显示每个数值或类别出现的次数来整理原始数据。收集数据时常使用计数符号。
Bar charts represent categorical data with rectangular bars whose heights are proportional to the frequencies. The bars are separated by gaps to show distinct categories.
条形图用矩形条表示分类数据,条的高度与频率成正比。条与条之间留有间隙以表示不同类别。
Pie charts show proportions as slices of a circle. The angle of each slice is calculated as (frequency ÷ total frequency) × 360°.
饼图用圆的扇形表示比例。每个扇形的角度计算公式为:(该类频数 ÷ 总频数)× 360°。
Pictograms use symbols to represent data. A key must state what each symbol represents, and partial symbols can show fractional amounts.
象形图用符号表示数据。必须附上图例说明每个符号代表什么,部分符号可用于表示不完整的数量。
5. Stem-and-Leaf Diagrams | 茎叶图
A stem-and-leaf diagram displays data while preserving the original values. The stem represents all digits except the last, and the leaf is the final digit.
茎叶图既能展示数据又能保留原始数值。茎代表除最后一位外的所有数位,叶是最后一位数字。
To construct the diagram, list stems in a vertical column, then write the leaves next to the corresponding stem in ascending order. Always include a key explaining place values.
制作茎叶图时,将茎垂直列成一列,然后在对应茎旁边按升序写下叶。必须包含一个说明位值的图例。
Back-to-back stem-and-leaf diagrams compare two related data sets, with one set’s leaves extending to the left and the other to the right of the central stem.
背靠背茎叶图可以比较两组相关数据,一组数据的叶向左延伸,另一组向右延伸,共享中间的茎。
6. Measures of Central Tendency | 集中趋势的度量
The mean is the sum of all data values divided by the number of values. It uses every piece of data but is affected by extreme values (outliers).
均值是所有数据值的总和除以数据个数。它利用了每一个数据,但会受到极端值(异常值)的影响。
For a frequency distribution, mean = Σ(f × x) ÷ Σf, where f is the frequency and x is the data value or class midpoint.
对于频率分布,均值 = Σ(f × x)÷ Σf,其中 f 是频数,x 是数据值或组中点。
The median is the middle value when the data is ordered. For an odd number of values, it is the (n+1)/2ᵗʰ value. For an even number, it is the average of the n/2ᵗʰ and (n/2 +1)ᵗʰ values.
中位数是将数据排序后的中间值。数据个数为奇数时,中位数是第 (n+1)/2 个值;为偶数时,是第 n/2 个与第 (n/2 +1) 个值的平均数。
The mode is the most frequently occurring value. A data set can have one mode, more than one mode (bimodal or multimodal), or no mode at all.
众数是出现次数最多的数值。一个数据集可能有一个众数、多个众数(双众数或多众数),或根本没有众数。
7. Measures of Spread | 离散程度的度量
The range is the difference between the largest and smallest values. It is simple to calculate but ignores the distribution of the middle data.
极差是最大值与最小值之差。计算简单,但忽略了中间数据的分布情况。
Quartiles divide an ordered data set into four equal parts. The lower quartile (Q₁) is the median of the lower half, and the upper quartile (Q₃) is the median of the upper half.
四分位数将有序数据分成四等份。下四分位数(Q₁)是下半部分的中位数,上四分位数(Q₃)是上半部分的中位数。
The interquartile range (IQR) = Q₃ − Q₁. It measures the spread of the middle 50% of data and is resistant to outliers, making it more reliable than the range in many situations.
四分位距(IQR)= Q₃ − Q₁。它衡量中间 50% 数据的分散程度,不受异常值影响,因此在许多情况下比极差更可靠。
8. Box Plots and Cumulative Frequency | 箱形图与累积频率
A box plot (box-and-whisker plot) displays the minimum, Q₁, median, Q₃ and maximum. The box spans Q₁ to Q₃, with a line at the median, and whiskers extend to the extremes.
箱形图(盒须图)展示最小值、Q₁、中位数、Q₃ 和最大值。矩形盒从 Q₁ 延伸到 Q₃,中位数处有一条线,须线延伸到两端极值。
Box plots are excellent for comparing distributions and identifying skewness. A longer whisker on one side suggests the data is skewed in that direction.
箱形图非常适合比较分布和识别偏态。某一侧的须线较长意味着数据向该方向偏斜。
A cumulative frequency graph shows the running total of frequencies. The median and quartiles can be estimated by drawing horizontal lines at ½, ¼ and ¾ of the total frequency.
累积频率图展示频率的累计总和。通过在总频数的 ½、¼ 和 ¾ 处画水平线,可以估算中位数和四分位数。
9. Scatter Graphs and Correlation | 散点图与相关性
A scatter graph plots paired numerical data to investigate the relationship between two variables. The independent variable goes on the x-axis, the dependent on the y-axis.
散点图通过描绘成对数值数据来探究两个变量之间的关系。自变量放在 x 轴,因变量放在 y 轴。
Correlation describes the strength and direction of the linear relationship. Positive correlation means as one variable increases, the other tends to increase; negative correlation means as one increases, the other tends to decrease.
相关性描述线性关系的强度和方向。正相关意味着一个变量增大,另一个也趋于增大;负相关意味着一个变量增大,另一个趋于减小。
A line of best fit (trend line) can be drawn by eye to pass through the middle of the points. It can be used to make predictions, but extrapolation beyond the data range should be treated with caution.
最佳拟合线(趋势线)可以用目测方法画出,穿过数据点的中心位置。它可用于做出预测,但超出数据范围的外推应谨慎对待。
10. Basic Probability Concepts | 基本概率概念
Probability measures the chance that an event will happen, expressed as a number between 0 (impossible) and 1 (certain). It can be written as a fraction, decimal or percentage.
概率衡量一个事件发生的可能性,用一个介于 0(不可能)和 1(必然)之间的数字表示。可以用分数、小数或百分数书写。
The theoretical probability of an event A is:
P(A) = Number of favourable outcomes ÷ Total number of possible outcomes
事件 A 的理论概率为:
P(A) = 有利结果的数量 ÷ 所有可能结果的总数
Experimental probability is based on data from trials or experiments. It is calculated as the number of times the event occurs divided by the total number of trials.
实验概率基于试验或实验的数据。它的计算方法是事件发生的次数除以总试验次数。
If an experiment is repeated many times, the experimental probability tends to approach the theoretical probability. This is known as the law of large numbers.
如果实验重复多次,实验概率会趋近理论概率。这就是大数定律。
11. Probability Trees and Combined Events | 概率树与组合事件
A probability tree diagram shows all possible outcomes for two or more events. Branches are labelled with probabilities, which must sum to 1 at each branching point.
概率树图展示两个或多个事件的所有可能结果。分枝上标注概率,每个分枝点上的概率之和必须为 1。
To find the probability of two independent events both occurring, multiply the probabilities along the branches. For example, P(A and B) = P(A) × P(B).
要求得两个独立事件同时发生的概率,将沿着分枝的概率相乘。例如,P(A 且 B) = P(A) × P(B)。
If an event can occur in more than one way, add the probabilities of the separate paths. For instance, the probability of at least one success in two trials can be found by adding the probabilities of the relevant end points.
如果一个事件可以通过多种方式发生,则将各条路径的概率相加。例如,两次试验中至少有一次成功的概率可以通过将相关终点的概率相加得到。
Tree diagrams can also handle conditional probability, where the probability of an event depends on the outcome of a previous event. In this case, the probabilities on the second set of branches change.
概率树图还可以处理条件概率,即一个事件的概率取决于前一事件的结果。此时,第二层分支上的概率会发生变化。
12. Introduction to Discrete Random Variables | 离散随机变量简介
A discrete random variable is a variable that can take a countable number of values, each with an associated probability. The probability distribution lists all values and their probabilities, which must sum to 1.
离散随机变量是一个可以取可数个值的变量,每个值都有对应的概率。概率分布列出所有取值及其概率,这些概率之和必须为 1。
The expectation or expected value E(X) of a discrete random variable is the long-run average. It is calculated as:
E(X) = Σ [x × P(X = x)]
离散随机变量的期望值 E(X) 是长期平均值。计算公式为:
E(X) = Σ [x × P(X = x)]
A very simple special case is the discrete uniform distribution, where all outcomes are equally likely. For example, rolling a fair six-sided die gives each outcome a probability of ⅙.
一个非常简单的特例是离散均匀分布,所有结果等可能发生。例如,掷一枚公平的六面骰子,每个结果的概率为 ⅙。
Another common distribution met at this level is the binomial distribution for a fixed number n of independent trials, each with the same probability of success p. The probability of exactly x successes is given by a formula involving combinations.
在这个阶段遇到的另一种常见分布是二项分布,适用于固定的独立试验次数 n,每次试验成功概率相同 p。恰好获得 x 次成功的概率由一个组合公式给出。
Understanding these models helps students begin to make inferences about populations based on sample data, a skill that will be developed further in Year 11 and beyond.
理解这些模型有助于学生开始根据样本数据推断总体特征,这项技能将在 Year 11 及以后的学习中进一步发展。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply