WJEC Year 11 Statistics: Terminology Rapid Recall Guide | WJEC 11年级统计:术语速记指南

📚 WJEC Year 11 Statistics: Terminology Rapid Recall Guide | WJEC 11年级统计:术语速记指南

Mastering the language of statistics is half the battle. This guide provides clear, memorisation-friendly definitions of the key terms you will encounter in the WJEC Year 11 Statistics course, organised by topic. Read each English explanation, then reinforce it with the paired Chinese version – perfect for bilingual learners and last-minute revision.

掌握统计学的语言是成功的一半。本指南针对WJEC 11年级统计课程中的核心术语,提供清晰且便于记忆的定义,并按主题分类。每一条都配有英文解释和对应的中文说明,非常适合双语学习者以及考前快速复习。

1. Types of Data | 数据类型

Qualitative data describe non‑numerical qualities, characteristics or categories. They are often words or labels, such as hair colour (‘brown’, ‘black’) or types of fruit. These cannot be measured with numbers meaningfully but can be counted in categories.

定性数据描述非数值的性质、特征或类别。它们通常是词语或标签,例如头发的颜色(“棕色”“黑色”)或水果的种类。这些数据不能用数字进行有意义的测量,但可以按类别计数。

Quantitative data are numerical and arise from counting or measuring. They are split into discrete data, which can only take certain exact values (e.g. number of students in a class, shoe size), and continuous data, which can take any value within a range (e.g. height, mass, temperature).

定量数据是数值型的,来自计数或测量。它们分为离散型数据(只能取某些确切的值,例如班级学生人数、鞋码)和连续型数据(可以取某一范围内的任意值,例如身高、质量、温度)。

Primary data are collected first-hand by the researcher for a specific investigation. Secondary data have already been collected by someone else for a different purpose, such as government statistics or internet databases.

原始数据是由研究者为特定调查亲自收集的。二手数据则是已经由他人出于其他目的收集好的,例如政府统计数据或网络数据库。

Categorical (nominal) data are a type of qualitative data where the categories have no natural order, e.g. favourite colour. Ordinal data have categories that can be ordered or ranked, such as satisfaction ratings (poor, fair, good).

分类(名义)数据是一种定性数据,其类别没有自然顺序,例如最喜欢的颜色。有序数据则具备可排序或可排名的类别,例如满意度评分(差、一般、好)。

Quick memory tip: ‘Qualitative = Quality (descriptions), Quantitative = Quantity (numbers). Discrete = distinct points, Continuous = connected range.’

速记提示:“Qualitative 看品质(描述),Quantitative 看数量(数字)。Discrete 是分散的点,Continuous 是连续的范围。”


2. Sampling Terminology | 抽样术语

A population is the entire group you want to draw conclusions about. A sample is a subset of that population, selected to represent it. A sampling frame is a list of all members of the population from which the sample is drawn.

总体是你想要得出结论的整个群体。样本是从总体中选出的一个子集,用以代表总体。抽样框是列出总体中所有成员的名单,样本从中抽取。

A random sample gives every member of the population an equal chance of being chosen. A stratified sample divides the population into distinct groups (strata) and takes a proportional random sample from each stratum, ensuring representation.

随机样本使总体中的每个成员都有相等的机会被选中。分层抽样先将总体分成不同的层,再从每一层中按比例抽取随机样本,从而确保代表性。

Systematic sampling selects every k‑th member from the sampling frame after a random start. Cluster sampling divides the population into clusters, randomly selects some clusters, and then uses all members within those selected clusters.

系统抽样在随机起点后,每隔k个成员从抽样框中选取一个。整群抽样将总体分成群,随机选择若干群,然后调查这些被选群中的所有成员。

Bias occurs when a sample is not representative of the population, for instance through convenience sampling (choosing friends) or voluntary response (TV phone-ins). A biased sample leads to unreliable conclusions.

当样本无法代表总体时便会产生偏差,例如便利抽样(选择朋友)或自愿回应(电视电话访问)。有偏差的样本会导致不可靠的结论。

Mnemonic: ‘SCaRS’ – Stratified, Cluster and Random, Systematic. Always check the frame to avoid bias!

助记口诀:“分层、整群、随机、系统”——牢记检查抽样框,避开偏差!


3. Measures of Central Tendency | 集中趋势的度量

The mean is the arithmetic average, calculated as sum of all data values divided by the number of values: x̄ = Σx / n. It uses every piece of data but is affected by extreme outliers.

均值是算术平均数,计算公式为:所有数据值的总和除以数据个数,x̄ = Σx / n。它利用了每一个数据点,但会受到极端离群值的影响。

The median is the middle value when the data are arranged in order. For an even number of data, it is the average of the two middle numbers. The median is resistant to outliers and is useful for skewed distributions.

中位数是将数据排序后位于中间的值。如果数据个数为偶数,则取中间两个数的平均值。中位数不受离群值的影响,适用于偏态分布。

The mode is the value that appears most frequently. A data set can have one mode (unimodal), two modes (bimodal) or many modes (multimodal). The mode is the only average suitable for qualitative data.

众数是出现频率最高的值。数据集可以有一个众数(单峰)、两个众数(双峰)或多个众数(多峰)。众数是唯一适用于定性数据的平均数。

For grouped data, we estimate the mean using the midpoints of class intervals. The modal class is the interval with the highest frequency, and the median class is found using cumulative frequency.

对于分组数据,我们使用组中值来估算平均数。众数组是频数最高的区间,中位数组则通过累积频数来确定。

Memory aid: ‘Mean is sensitive, Median is resistant, Mode is most frequent.’

记忆口诀:“均值怕极端,中位数稳如山,众数看谁出现最频繁。”


4. Measures of Dispersion | 离散程度的度量

The range is the simplest measure of spread: highest value minus lowest value. It is quick to calculate but only uses two data points, making it sensitive to outliers.

极差是最简单的离散度量:最大值减去最小值。它计算简便,但只用到了两个数据点,因此极易受离群值影响。

The interquartile range (IQR) is the difference between the upper quartile (Q3) and lower quartile (Q1): IQR = Q3 − Q1. It represents the spread of the middle 50 % of the data and is not skewed by extreme values.

四分位距(IQR)是上四分位数(Q3)与下四分位数(Q1)之差:IQR = Q3 − Q1。它反映了中间50%数据的分散程度,并且不受极端值影响。

Variance measures the average squared deviation from the mean. The standard deviation is the square root of the variance: s = √[ Σ(x − x̄)² / (n − 1) ] for a sample. A smaller standard deviation indicates data are clustered closely around the mean.

方差衡量各数据点与均值之差的平方的平均值。标准差是方差的平方根,样本标准差公式为 s = √[ Σ(x − x̄)² / (n − 1) ]。标准差越小,表明数据越紧密地聚集在均值周围。

When comparing data sets, always discuss both a measure of central tendency and a measure of dispersion, e.g. “On average, apples weigh more, but their weights are also more variable.”

在比较数据集时,务必同时讨论集中趋势和离散程度,例如:“平均而言,苹果更重,但其重量的差异也更大。”


5. Charts and Diagrams | 图表与图示

A bar chart uses bars of equal width to represent frequencies of categorical or discrete data, with gaps between bars. A pie chart displays proportions as sectors of a circle, where each sector angle = (frequency / total) × 360°.

条形图使用等宽的条形表示分类或离散数据的频数,条形之间留有空隙。饼图则以圆的扇区来呈现比例,每个扇区的角度 = (频数 / 总频数) × 360°。

A histogram is used for continuous grouped data. Unlike a bar chart, the bars touch and the area of each bar is proportional to the frequency. Frequency density = frequency / class width, which is plotted on the vertical axis when class widths are unequal.

直方图用于连续分组数据。与条形图不同的是,直方图的条形紧贴在一起,每个条形的面积与频数成正比。当组距不等时,纵轴需使用频数密度,频数密度 = 频数 / 组距。

A frequency polygon is formed by joining the midpoints of the tops of histogram bars with straight lines. It is often used to compare two distributions on the same axes.

频数多边形是用直线连接直方图各条形顶端中点而形成的图形。它通常用于在同一坐标轴上比较两个分布。

A scatter graph displays the relationship between two variables. Each point represents a pair of values (x, y). Scatter graphs help identify correlation, trends and outliers.

散点图展现两个变量之间的关系。每个点代表一对数值 (x, y)。散点图有助于识别相关性、变化趋势和异常点。


6. Cumulative Frequency and Box Plots | 累积频数与箱线图

Cumulative frequency is a running total of the frequencies. Plotting upper class boundaries against cumulative frequency gives an S‑shaped (ogive) curve, used to estimate the median, quartiles and percentiles.

累积频数是频数的逐次累加。以组上限为横坐标、累积频数为纵坐标绘图,可得到一条 S 形曲线(累积频数曲线),用于估算中位数、四分位数和百分位数。

A box plot (box‑and‑whisker diagram) displays the five‑number summary: minimum, lower quartile (Q1), median (Q2), upper quartile (Q3) and maximum. The box spans Q1 to Q3, with a line at the median; whiskers extend to the extremes within 1.5 × IQR.

箱线图(盒须图)展示五数概括:最小值、下四分位数(Q1)、中位数(Q2)、上四分位数(Q3)和最大值。箱子从 Q1 延伸到 Q3,中间有一条中线表示中位数;须线伸展到不超过 1.5 倍 IQR 范围内的最远点。

Outliers are data values that lie more than 1.5 × IQR below Q1 or above Q3. On a box plot they are plotted as individual points beyond the whiskers.

离群值是落在 Q1 − 1.5×IQR 以下,或 Q3 + 1.5×IQR 以上的数据值。在箱线图中,它们被绘制成须线之外的单独的点。

A box plot is extremely useful for comparing distributions side‑by‑side. You can instantly compare medians, spreads and skewness.

箱线图非常适合将多个分布并排比较,可以立刻对比它们的中位数、离散程度和偏态。


7. Probability Language | 概率语言

Experiment, outcome, event. An experiment is a repeatable process giving results. An outcome is a single result. An event is one or more outcomes. Probability of an event = number of favourable outcomes / total number of possible outcomes, assuming all outcomes are equally likely.

试验、结果、事件。试验是一个可重复产生结果的过程。结果是单个结果。事件则是一个或多个结果的集合。在所有可能结果等可能的前提下,事件的概率 = 有利结果数目 / 可能结果总数。

Two events are mutually exclusive if they cannot happen at the same time, e.g. rolling a 3 and a 5 on a single die. For mutually exclusive events, P(A or B) = P(A) + P(B).

如果两个事件不可能同时发生,则它们互斥,例如掷一次骰子同时出现 3 点和 5 点。对于互斥事件,P(A 或 B) = P(A) + P(B)。

Independent events are those where the occurrence of one does not affect the probability of the other, e.g. flipping a fair coin twice. For independent events, P(A and B) = P(A) × P(B).

独立事件是指一个事件的发生不影响另一个事件的发生概率,例如抛两次公平硬币。对于独立事件,P(A 且 B) = P(A) × P(B)。

Conditional probability is the probability of an event given that another event has occurred, written as P(A|B). Tree diagrams are powerful tools for solving sequential probability problems; multiply along the branches.

条件概率是在另一事件已经发生的条件下某事件发生的概率,记作 P(A|B)。树形图是解决序贯概率问题的有力工具;沿着分支相乘即可。

Expected frequency is the theoretical number of times an event would occur in a large number of trials: Expected frequency = probability × number of trials.

期望频数是指如果在大量试验中,某一事件理论上应发生的次数:期望频数 = 概率 × 试验次数。


8. Correlation and Regression | 相关与回归

Correlation describes the strength and direction of a linear relationship between two variables. Positive correlation means as one variable increases, the other tends to increase. Negative correlation means as one increases, the other tends to decrease.

相关性描述的是两个变量之间线性关系的强度和方向。正相关意味着一个变量增大时另一个也趋于增大;负相关则是一个增大时另一个趋于减小。

Correlation does not imply causation. Just because two variables are associated does not mean one causes the other; a third (lurking) variable may be responsible.

相关性并不意味着因果关系。两个变量有关联并不代表一个引起另一个,可能存在第三个(隐匿)变量在起作用。

Spearman’s rank correlation coefficient (rₛ) measures the strength of monotonic association between two ranked variables. rₛ ranges from –1 (perfect negative) to +1 (perfect positive). A value near 0 indicates no monotonic correlation.

斯皮尔曼等级相关系数(rₛ)衡量两个排序变量之间的单调关联强度。rₛ 的取值范围从 –1(完全负相关)到 +1(完全正相关),接近 0 表示没有单调相关。

The line of best fit on a scatter graph can be drawn by eye (‘informal regression’) or calculated using the least‑squares regression line. Regression lines are used to make predictions, but extrapolation beyond the data range is unreliable.

散点图上的最佳拟合线可以通过目测绘制(“非正式回归”),也可以使用最小二乘回归线计算。回归线用于进行预测,但超出数据范围的外推不可靠。

The equation of a regression line is often written as y = a + bx, where b is the slope (change in y per unit change in x) and a is the intercept (value of y when x = 0).

回归线方程通常写作 y = a + bx,其中 b 是斜率(x 每变化一个单位 y 的变化量),a 是截距(x = 0 时的 y 值)。


9. Time Series Analysis | 时间序列分析

A time series is a set of data recorded at regular time intervals (e.g. monthly sales). Time series graphs plot the data against time to reveal underlying patterns: trend, seasonal variation and random fluctuations.

时间序列是按固定时间间隔记录的一组数据(例如月度销售额)。时间序列图将数据按时间绘制,揭示潜在的模式:趋势、季节性变动和随机波动。

The trend is the long‑term, smooth movement of the series. Seasonal variations are regular, short‑term fluctuations that repeat over fixed periods (e.g. higher ice‑cream sales in summer).

趋势是序列长期的、平滑的移动方向。季节性变动则是规律的、短期波动,在固定周期内重复出现(例如夏季冰淇淋销量增高)。

A moving average smooths out short‑term fluctuations to help highlight the trend. For quarterly data, a 4‑point moving average is typical. The moving average is plotted at the centre of the time points used.

移动平均通过平滑短期波动来突出趋势。对于季度数据,通常使用 4 点移动平均。移动平均值绘制在所用时间点的中心位置。

Seasonal effects can be estimated by subtracting the moving average (trend) from the actual data. The mean seasonal variation can then be used to adjust predictions or make forecasts.

季节性影响可通过实际数据减去移动平均(趋势)来估算。进而可以用平均季节性变动来调整预测或做出预报。

Remember: “Trend + Seasonal + Random = Observed value.” Decompose carefully and always plot the moving average on the same graph.

记住:“观测值 = 趋势 + 季节 + 随机。” 仔细分解,并始终把移动平均绘制在同一张图上。


10. Index Numbers and Weighted Means | 指数与加权平均数

An index number measures the relative change in a quantity (e.g. price, quantity, income) compared to a base period. The base period index is usually set to 100. Index = (Value / Base value) × 100.

指数用来衡量某个量(如价格、数量、收入)相对于基期的变化。基期的指数通常设为 100。指数 = (数值 / 基期数值) × 100。

A weighted index gives different importance to different items. For example, a consumer price index (CPI) weights goods according to typical household spending. Weighted aggregate index = Σ(weight × item index) / Σ(weight).

加权指数对不同项目赋予不同的重要性。例如,消费者价格指数(CPI)根据典型家庭支出对各商品进行加权。加权综合指数 = Σ(权重 × 项目指数) / Σ(权重)。

A weighted mean is used when some values contribute more than others. Weighted mean = Σ(w × x) / Σw, where w are the weights and x are the values. This is essential for calculating overall grades or composite scores.

当某些值的贡献程度不同时,使用加权平均数。加权平均数 = Σ(w × x) / Σw,其中 w 为权重,x 为数值。这在计算总评成绩或综合分数时至关重要。

Chain base index numbers link each period’s value directly to the previous period, useful for tracking period‑to‑period growth rates. In contrast, fixed base indices relate everything back to a single base year.

链基指数将每一期的值直接与前一期相联系,便于追踪逐期增长率。而固定基期指数则将所有时期都与同一个基年进行比较。

Trick: “Base is 100, weight multiplies.” Always check whether the question asks for a simple or weighted index.

技巧:“基期总是100,权重直接相乘。” 答题时务必先看清是求简单指数还是加权指数。

Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version