Year 9 AQA Statistics: Vocabulary & Terminology Quick Memorization Guide | 九年级AQA统计:词汇术语速记指南

📚 Year 9 AQA Statistics: Vocabulary & Terminology Quick Memorization Guide | 九年级AQA统计:词汇术语速记指南

Welcome to the essential vocabulary guide for Year 9 AQA Statistics. Mastering key terms is the first step towards confidence in data handling, probability, and statistical diagrams. This article provides clear definitions, paired English-Chinese explanations, and mnemonics to help you remember them quickly.

欢迎阅读九年级AQA统计必学词汇指南。掌握关键术语是自信处理数据、概率和统计图表的第一步。本文提供清晰的定义、英中对照解释和记忆法,帮助你快速牢记。


1. Types of Data | 数据类型

Data can be classified as qualitative or quantitative. Qualitative data describes qualities and is non-numerical, like colours or hair types. Quantitative data is numerical and can be measured.

数据可分为定性数据和定量数据。定性数据描述性质,非数值,如颜色或发型。定量数据是数值型,可以测量。

Quantitative data splits further into discrete and continuous. Discrete data can only take specific values, often counted (e.g. number of siblings). Continuous data can take any value within a range and is measured (e.g. height, time). Remember: ‘Discrete = countable, Continuous = measurable’.

定量数据又分为离散型和连续型。离散数据只能取特定值,通常可数(如兄弟姐妹的数量)。连续数据可以取区间内任何值,可测量(如身高、时间)。记住:“离散可数,连续可测”。

Tip: Associate ‘discrete’ with ‘distinct’ — both start with ‘dis’ and refer to separate items. ‘Continuous’ sounds like ‘continue’, implying no breaks.

提示:将‘discrete’与‘distinct’联系——都以‘dis’开头,指分离的个体。‘Continuous’听起来像‘continue’,意味着不间断。


2. Populations and Samples | 总体与样本

A population is the entire group you want to study. A sample is a smaller subset selected from the population. Why sample? It saves time and money while still giving useful information if chosen correctly.

总体是你要研究的整个群体。样本是从总体中选取的一个较小子集。为什么要抽样?它省时省钱,如果选取得当,仍能提供有用信息。

A sampling frame is a list of all members in the population, such as a school register. The sampling unit is each individual element that could be selected. A representative sample mirrors the characteristics of the population, avoiding bias.

抽样框是总体所有成员的名单,例如学校花名册。抽样单元是可能被选中的每一个个体单元。有代表性的样本能反映总体的特征,从而避免偏差。

Bias means a systematic error that favours certain outcomes. For example, only asking students who like sports about PE lessons introduces selection bias. Random sampling helps combat this.

偏差指一种系统性的错误,倾向于某些结果。例如,只询问喜欢运动的学生对体育课的看法,会引入选择偏差。随机抽样有助于克服这一点。


3. Measures of Central Tendency | 集中趋势的度量

The mean is calculated by summing all data values and dividing by the number of values. It is commonly called the average, but it is sensitive to extreme outliers.

平均数是将所有数据值相加再除以数值个数得出的。它通常被称为平均值,但对极端异常值敏感。

The median is the middle value when the data are ordered from smallest to largest. If there are two middle numbers, the median is their mean. It is robust against outliers because it depends only on position.

中位数是将数据从小到大排序后的中间值。如果有两个中间数,中位数就是它们的平均数。它对异常值不敏感,因为它只依赖于位置。

The mode is the value that appears most frequently. A data set can have one mode (unimodal), two modes (bimodal), or more. Mode is particularly useful for categorical data.

众数是出现频率最高的值。一个数据集可以有一个众数(单峰)、两个众数(双峰)或更多。众数对于分类数据特别有用。

For grouped data, the mean is estimated using

Mean ≈ Σ(f × midpoint) ÷ Σf

, where f is the frequency and midpoint is the class centre.

对于分组数据,平均数用公式 Mean ≈ Σ(f × midpoint) ÷ Σf 估算,其中 f 是频数,midpoint 是组中值。


4. Measures of Spread | 离散程度的度量

The range is the simplest measure of spread: largest value minus smallest value. It shows the total spread but can be distorted by a single outlier.

极差是最简单的离散度量:最大值减最小值。它显示总分布范围,但可能被单个异常值扭曲。

The interquartile range (IQR) is the difference between the upper quartile (Q₃) and the lower quartile (Q₁). It represents the range of the middle 50% of the data and is not affected by outliers.

四分位距(IQR)是上四分位数 Q₃ 与下四分位数 Q₁ 的差。它代表中间50%数据的范围,且不受异常值影响。

Standard deviation measures the average distance of data points from the mean. A larger standard deviation indicates more spread. While more common at GCSE level, understanding the concept early helps.

标准差衡量数据点与平均值的平均距离。标准差越大,数据越分散。虽然这个概念在 GCSE 阶段更常见,但提前理解会有帮助。


5. Frequency Distributions | 频数分布

A frequency table organises data by listing values or groups and how often they occur. Tally marks are often used to record raw data.

频数表通过列出数值或组及其出现的次数来整理数据。计数记号常被用来记录原始数据。

Cumulative frequency is the running total of frequencies. Adding each frequency to the sum of previous ones builds the cumulative frequency column, which is essential for drawing cumulative frequency curves.

累积频数是频数的累加总和。将每个频数加到前面所有频数的总和上,就得到了累积频数列,这对绘制累积频数曲线至关重要。

Relative frequency is the frequency of a value divided by the total number of observations. It is often expressed as a fraction, decimal, or percentage and can be used as an experimental probability estimate.

相对频数是某个值的频数除以观测总数。它通常表示为分数、小数或百分数,并可用作实验概率的估计值。


6. Data Representation Charts | 数据呈现图表

Bar charts represent categorical or discrete data using rectangular bars. The bars are usually separated by equal gaps, and the height of each bar is proportional to the frequency.

条形图用矩形条表示分类或离散数据。条形之间通常有相等的间隔,每个条形的高度与频数成正比。

Pie charts show the proportions of a whole. Each sector angle is calculated as

Angle = (Category frequency ÷ Total) × 360°

. Always check that the sum of angles equals 360°.

饼图展示整体各部分的比例。每个扇形的角度计算公式为:角度 = (类别频数 ÷ 总数) × 360°。务必检查角度之和等于 360°。

Pictograms use simple pictures or symbols to represent data. A key, e.g. one symbol = 2 people, ensures the chart is read correctly. Partial symbols can show fractional amounts.

象形图使用简单的图片或符号来表示数据。图例(例如一个符号 = 2 人)确保图表被正确读取。不完整的符号可以表示部分数量。

Histograms are used for continuous data. Unlike bar charts, the bars touch and it is the area of the bar that is proportional to frequency. Frequency density = frequency ÷ class width is a key concept for unequal class intervals.

直方图用于连续数据。与条形图不同,直方图的条形相连,并且条形面积与频数成比例。频数密度 = 频数 ÷ 组距,这是处理不等组距时的关键概念。


7. Cumulative Frequency and Box Plots | 累积频数与箱线图

A box plot (or box-and-whisker diagram) summarises data using five key values: minimum, lower quartile Q₁, median, upper quartile Q₃, and maximum. It quickly shows the centre, spread, and skewness of the data.

箱线图(或盒须图)用五个关键值来总结数据:最小值、下四分位数 Q₁、中位数、上四分位数 Q₃ 和最大值。它能快速显示数据的中心、离散度及偏态。

The ‘box’ extends from Q₁ to Q₃, covering the IQR. The line inside the box marks the median. The ‘whiskers’ reach out to the minimum and maximum, unless extreme outliers are plotted as separate points.

“箱”从 Q₁ 延伸到 Q₃,涵盖了 IQR。箱内的线标示中位数。“须”则延伸到最小值和最大值,除非极端异常值被单独标为点。

A cumulative frequency graph is drawn by plotting cumulative frequency against the upper class boundary for each group. Points are joined by a smooth curve. You can then use this curve to estimate the median, quartiles, and percentiles.

累积频数图通过将每组的上限值对应累积频数描点来绘制,用光滑曲线连接各点。然后可以利用该曲线估计中位数、四分位数和百分位数。


8. Probability Terms | 概率术语

Probability measures how likely an event is to happen. It lies on a scale from 0 (impossible) to 1 (certain). A fair coin has P(Heads) = 0.5.

概率衡量事件发生的可能性,范围在 0(不可能)到 1(必然)之间。一枚公平硬币的 P(正面) = 0.5。

The sample space is the set of all possible outcomes. An event is one or more outcomes. For example, rolling a die, the sample space is {1, 2, 3, 4, 5, 6} and the event ‘an even number’ is {2, 4, 6}.

样本空间是所有可能结果的集合。事件是一个或多个结果。例如,掷骰子的样本空间是 {1, 2, 3, 4, 5, 6},事件“偶数”是 {2, 4, 6}。

Mutually exclusive events cannot occur at the same time. If A and B are mutually exclusive,

P(A or B) = P(A) + P(B)

. Think ‘OR means add for exclusive events’.

互斥事件不能同时发生。如果 A 和 B 互斥,则 P(A or B) = P(A) + P(B)。注意:“互斥情况下,或运算用加法”。

Independent events have no influence on each other. For independent events A and B,

P(A and B) = P(A) × P(B)

. The ‘AND’ rule multiplies probabilities when events are independent.

独立事件彼此没有影响。对于独立事件 A 和 B,P(A and B) = P(A) × P(B)。独立情况下,“且运算用乘法”。

Expected frequency predicts how many times an event will occur in a number of trials:

Expected frequency = Probability × Number of trials

.

期望频数预测在若干次试验中事件发生的次数:期望频数 = 概率 × 试验次数。


9. Scatter Graphs and Correlation | 散点图与相关

A scatter graph displays paired numerical data to reveal a relationship. Each axis represents a variable, and each point is a pair (x, y).

散点图显示成对的数值数据,以揭示关系。每个轴代表一个变量,每个点是一对 (x, y)。

Correlation describes the trend. Positive correlation means that as x increases, y tends to increase. Negative correlation means as x increases, y tends to decrease. No correlation implies no clear pattern.

相关描述趋势。正相关意味着随着 x 增加,y 也倾向增加。负相关意味着随着 x 增加,y 倾向减少。无相关表示没有明显模式。

A line of best fit (drawn by eye) follows the general direction of the points. It should have roughly equal numbers of points above and below the line. Interpolation is estimating inside the data range; extrapolation is outside the range and can be unreliable.

最佳拟合线(目测绘制)沿数据点的总体方向画出。它应该使线上的点和线下的点数量大致相等。内插是在数据范围内估计;外推是在范围外估计,可能不可靠。

Strong correlation does not mean causation. Two variables might both be influenced by a third factor. Always ask: ‘Is there a common cause?’

强相关并不意味着因果关系。两个变量可能同时受到第三个因素的影响。请始终思考:“是否存在一个共同原因?”


10. Key Terms for Statistical Investigations | 统计调查关键词

Primary data is original information collected directly for your investigation. Secondary data is gathered from existing sources, like websites or textbooks.

一手数据是直接为你的调查收集的原始信息。二手数据是从现有来源收集的,如网站或教科书。

A hypothesis is a statement that can be tested by data. In Year 9, a comparative hypothesis might be ‘Girls in Year 9 read more books per month than boys’.

假设是一个能够用数据检验的陈述。在九年级,一个比较性假设可能是“九年级女生每月阅读的书比男生多”。

Sampling methods include simple random sampling (every member equally likely), systematic (choose every nth), stratified (divide into groups and sample proportionally), and convenience (non-random, easy access).

抽样方法包括简单随机抽样(每个成员等概率入选)、系统抽样(每第 n 个选取)、分层抽样(分成组并按比例抽样)和便利抽样(非随机,易于接近)。

Reliability and validity are important. Reliable data can be repeated with consistent results. Valid data measures what it intends to measure. A poorly worded question may yield unreliable or invalid responses.

信度和效度很重要。信度高的数据可以重复得到一致结果。效度高的数据能测量其想测量的东西。措辞不当的问题可能产生不可信或无效的回答。


11. Quick Mnemonics and Memory Hacks | 快速记忆窍门

Measures rap: ‘Mean adds and divides, it’s fair but hides extremes. Median is the middle, easy to find, robust and keen. Mode is most often seen, the popular teen!’

度量口诀:“平均数加加减减,公平却藏极端;中位数在中间,容易找又稳健;众数出现最频繁,就像人气少年!”

Data types trick: ‘QuaNtitative has an N for Numbers; QuaLitative has an L for Labels’. This keeps them straight.

数据类型窍门:‘QuaNtitative 中有 N 代表 Numbers(数字);QuaLitative 中有 L 代表 Labels(标签)。’这能让你分清楚。

IQR visual: Think of a birthday box — the box holds the middle 50

Published by TutorHao | Year 9 统计 Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading