Year 11 Edexcel Statistics: Key Terminology Memorisation Guide | Edexcel 统计学核心术语速记指南

📚 Year 11 Edexcel Statistics: Key Terminology Memorisation Guide | Edexcel 统计学核心术语速记指南

Mastering statistical vocabulary is essential for tackling Edexcel Year 11 Statistics with confidence. This bilingual guide breaks down the most frequently examined terms into clear, paired explanations, helping you memorise their meanings and apply them correctly in context.

掌握统计词汇是自信应对 Edexcel 十一年级统计学考试的基础。本双语指南将高频术语拆解为清晰的中英对照解释,帮助你记忆含义并在具体情境中准确应用。

1. Population and Sample | 总体与样本

A population is the entire set of individuals or items that you want to investigate. For example, all students in a school or every light bulb produced by a factory.

总体是你想调查的全体个体或物件的集合。例如,一所学校的所有学生或一家工厂生产的每一个灯泡。

A sample is a smaller group selected from the population to represent it. Sampling saves time and money, but the sample must be representative to avoid bias.

样本是从总体中选出的一个较小的组,用以代表总体。抽样可以省时省钱,但样本必须具有代表性以避偏差。

A census attempts to collect data from every member of the population. While it gives a true picture, it is often impractical for large populations.

普查试图从总体的每一个成员那里收集数据。虽然它能提供真实情况,但对于大总体往往不切实际。

The sampling frame is the list of all members of the population from which the sample is drawn, such as a register or database.

抽样框是总体所有成员的名单,样本从这个名单中抽取,如登记册或数据库。


2. Types of Data | 数据类型

Qualitative data describe qualities or categories and are non‑numerical, like eye colour or types of vehicle. Quantitative data are numerical and represent quantities.

定性数据描述性质或类别,是非数值型的,例如眼睛颜色或车辆类型。定量数据是数值型的,表示数量。

Quantitative data can be discrete (countable, whole numbers – e.g. number of books) or continuous (measurable, can take any value in a range – e.g. height, time).

定量数据可以是离散型(可数的整数,如书本数量)或连续型(可测量的,可取区间内任意值,如身高、时间)。

Nominal data are categories with no natural order, such as favourite colours. Ordinal data have a clear order, like satisfaction ratings (poor, good, excellent).

名义数据是无自然顺序的类别,如最喜欢的颜色。有序数据则有明确顺序,如满意度评分(差、好、优秀)。

Primary data are collected directly by the researcher for the specific purpose (e.g. conducting a survey). Secondary data are obtained from existing sources (e.g. government statistics).

一手数据由研究者为特定目的直接收集(如开展问卷调查)。二手数据取自现有来源(如政府统计数据)。


3. Sampling Methods | 抽样方法

Simple random sampling gives every member of the population an equal chance of being selected, often using random number generators. It removes selection bias but requires a full sampling frame.

简单随机抽样使总体中每个成员被选中的机会均等,常使用随机数生成器。它消除了选择偏差,但需要一个完整的抽样框。

Stratified sampling divides the population into distinct groups (strata) and then randomly samples from each group in proportion to its size. This ensures representation of all relevant subgroups.

分层抽样将总体分成不同的组(层),然后按比例从每一层中随机抽样。这确保了所有相关子群体都有代表。

Systematic sampling selects members at regular intervals from an ordered list, e.g. every 10th person. It is simple to use, but can introduce bias if the list has a hidden pattern.

系统抽样从有序的名单中按固定间隔选取成员,例如每10人选1人。这种方法简单易行,但如果名单存在隐藏规律,可能会引入偏差。

Quota sampling involves interviewing a fixed number of people from set categories (e.g. age and gender) without random selection. It is inexpensive, but relies on interviewer judgement and may be biased.

配额抽样是在没有随机选择的情况下,从设定的类别(如年龄和性别)中采访固定数量的人。这种方法成本低,但依赖访问员的判断,可能有偏差。


4. Measures of Central Tendency | 集中趋势度量

The mean is the average found by adding all values and dividing by the number of values. The formula is:

Mean = Σx / n

均值是通过将所有数值相加并除以数据个数得出的平均数。公式为:均值 = Σx / n。

The median is the middle value when data are ordered. For an even number of data points, it is the midpoint of the two central values. It is less affected by extreme values than the mean.

中位数是将数据排序后中间的值。如果数据个数为偶数,则为中间两个值的平均数。中位数受极端值影响比均值小。

The mode is the value that appears most often. A data set can have one mode (unimodal), more than one mode (bimodal/multimodal), or no mode at all.

众数是出现次数最多的值。数据集可以有一个众数(单峰)、多个众数(双峰/多峰)或无众数。


5. Measures of Spread | 离散程度度量

The range is the difference between the largest and smallest values. It is easy to calculate but heavily influenced by outliers.

极差是最大值与最小值之差。它容易计算,但强烈受离群值影响。

The interquartile range (IQR) is the difference between the upper quartile (Q₃) and the lower quartile (Q₁): IQR = Q₃ − Q₁. It describes the spread of the middle 50% of the data and is robust to outliers.

四分位距 (IQR) 是上四分位数 (Q₃) 与下四分位数 (Q₁) 之差:IQR = Q₃ − Q₁。它描述中间50%数据的离散程度,对离群值稳健。

The standard deviation measures how much the data deviate from the mean. The sample standard deviation is calculated as:

s = √[ Σ(x − x̄)² / (n − 1) ]

标准差衡量数据偏离均值的程度。样本标准差的计算公式为:s = √[ Σ(x − x̄)² / (n − 1) ]。


6. Charts and Diagrams | 图表与图形

Bar charts display categorical data with rectangular bars whose lengths are proportional to the frequencies. The bars are separated to show distinct categories.

条形图用矩形条表示分类数据,条的长度与频数成比例。条形之间留有空隙以显示不同类别。

Histograms are used for continuous data grouped into intervals. The area of each bar represents frequency, so the vertical axis shows frequency density: Frequency density = Frequency ÷ Class width. Bars touch to indicate continuous scale.

直方图用于分组连续数据。每个矩形的面积代表频数,因此纵轴表示频数密度:频数密度 = 频数 ÷ 组距。矩形之间紧靠以表示连续尺度。

Cumulative frequency curves (ogives) plot the running total of frequencies against the upper class boundary. They allow you to estimate the median and quartiles directly.

累积频数曲线(拱形图)以累积频数对上组界绘图。可直接从中估计中位数和四分位数。

Box plots (box‑and‑whisker diagrams) show the minimum, lower quartile, median, upper quartile and maximum. They are excellent for comparing distributions and identifying outliers.

箱线图(箱形图)显示最小值、下四分位数、中位数、上四分位数和最大值。它特别适合比较分布和识别离群值。


7. Correlation and Regression | 相关与回归

Correlation describes the strength and direction of a linear relationship between two variables. It can be positive (as one increases, the other tends to increase), negative (one increases, the other decreases) or zero (no linear pattern).

相关描述两个变量之间线性关系的强度和方向。可以是正相关(一个增加,另一个也趋向增加)、负相关(一个增加,另一个减少)或零相关(无线性模式)。

A scatter graph displays paired data points. The line of best fit is drawn to summarise the trend and can be used to make predictions.

散点图展示成对的数据点。最佳拟合线用来概括趋势,并可用于做出预测。

Interpolation is predicting a value within the range of the given data. Extrapolation predicts outside the data range and is less reliable because the trend may not continue.

内插法是在已知数据范围内进行预测。外推法是预测数据范围外的值,它较不可靠,因为趋势可能不会延续。

Spearman’s rank correlation coefficient (rₛ) measures the strength of association between two ranked variables. Values close to +1 indicate strong positive correlation, and −1 strong negative correlation.

斯皮尔曼等级相关系数 (rₛ) 衡量两个排序变量之间的关联强度。接近+1表示强正相关,接近−1表示强负相关。


8. Probability Basics | 概率基础

An experiment is a repeatable process that gives outcomes. An event is a set of one or more outcomes. The probability of an event is between 0 (impossible) and 1 (certain).

试验是可重复的过程,产生结果。事件是由一个或多个结果组成的集合。事件的概率介于0(不可能)和1(必然)之间。

Relative frequency is the proportion of times an event occurs when an experiment is repeated many times. Theoretical probability is based on equally likely outcomes without performing the experiment.

相对频率是试验重复多次时事件发生的比例。理论概率是基于等可能结果而不需实际试验。

Mutually exclusive events cannot happen at the same time. The probability of either one or the other occurring is the sum of their individual probabilities: P(A or B) = P(A) + P(B).

互斥事件不能同时发生。其中一个或另一个发生的概率是它们各自概率之和:P(A 或 B) = P(A) + P(B)。

Independent events are ones where the outcome of one does not affect the probability of the other. For independent events, P(A and B) = P(A) × P(B).

独立事件是指一个事件的结果不影响另一个事件的概率。对于独立事件,P(A 且 B) = P(A) × P(B)。

Tree diagrams help list all possible outcomes for multi‑stage experiments and calculate compound probabilities by multiplying along branches.

树状图有助于列出多阶段试验的所有可能结果,并通过沿分支相乘来计算复合概率。


9. Index Numbers | 指数

An index number measures the change in a quantity over time compared with a base period. The base period is usually assigned an index of 100.

指数衡量一个量相对于基期的随时间变化。基期通常被赋予指数100。

The simple index formula is:

Index = (Value in current period / Value in base period) × 100

简单指数公式为:指数 = (现期值 / 基期值) × 100。

A weighted index takes account of the relative importance of different items. For example, the Retail Price Index (RPI) weights goods according to typical household spending.

加权指数考虑了不同项目的相对重要性。例如,零售价格指数 (RPI) 根据典型家庭支出对商品进行加权。


10. Time Series and Data Quality | 时间序列与数据质量

A time series is a sequence of data points collected at regular time intervals. The main components are trend (long‑term movement), seasonal variation (regular pattern within a year) and random fluctuation.

时间序列是按固定时间间隔收集的数据序列。主要构成包括趋势(长期变动)、季节波动(年内规律性模式)和随机波动。

A moving average smooths out short‑term fluctuations to show the underlying trend. For quarterly data, a four‑point moving average is often used.

移动平均可以平滑短期波动,揭示潜在趋势。对于季度数据,常使用四项移动平均。

Bias occurs when a sample or survey systematically over‑ or under‑represents certain groups. Using random sampling and carefully wording questions can reduce bias.

当样本或调查系统性地高估或低估某些群体时,便产生偏差。使用随机抽样和谨慎设计问题措辞可以减少偏差。

Reliability refers to how consistent the results are if the study is repeated. Validity concerns whether the data actually measure what they are intended to measure. A pilot study helps test the design before the main investigation.

信度指如果重复研究,结果的一致性如何。效度则关注数据是否真正测量了想要测量的东西。试点研究有助于在正式调查前检验设计。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading