GCSE OCR Statistics: Key Terms Quick Memorisation Guide | GCSE OCR 统计:关键术语速记指南

📚 GCSE OCR Statistics: Key Terms Quick Memorisation Guide | GCSE OCR 统计:关键术语速记指南

Statistics can feel overwhelming when you are first introduced to all the new vocabulary. This guide breaks down the most important terms from the OCR GCSE Statistics specification into simple, memorable definitions. Use it as a quick reference to build your confidence before the exam.

统计学在刚接触时可能会让人感到不知所措,因为有大量新词汇。本指南将 OCR GCSE 统计规范中最重要的术语分解为简单易记的定义。你可以将其作为考前快速参考,增强信心。


1. Types of Data | 数据类型

Qualitative data: non‑numerical information that describes qualities or categories. Examples include eye colour, types of pet, or favourite subjects.

定性数据:描述性质或类别的非数值信息。例如眼睛颜色、宠物种类或最喜欢的科目。

Quantitative data: numerical data that can be measured or counted. It is split into discrete and continuous types.

定量数据:可以测量或计数的数值数据,分为离散型和连续型两种。

Discrete data: quantitative data that can only take specific, separate values—often whole numbers. Shoe size and number of siblings are good examples.

离散数据:只能取特定、分离值的定量数据,通常为整数。鞋码和兄弟姐妹数量就是很好的例子。

Continuous data: quantitative data that can take any value within a given range, including decimals. Height, mass, and time are typical continuous variables.

连续数据:可以在给定范围内取任意值(包括小数)的定量数据。身高、质量和时间都是典型的连续变量。

Ordinal data: a type of qualitative data where categories have a natural order or rank, such as satisfaction ratings (poor, fair, good, excellent).

顺序数据:一种定性数据,其类别具有自然顺序或等级,例如满意度评分(差、一般、好、优秀)。


2. Sampling Methods | 抽样方法

Population: the entire set of individuals or items that you are interested in studying. For example, all students in a school.

总体:你感兴趣研究的整个个体或项目集合。例如一所学校的所有学生。

Sample: a smaller group selected from the population to represent it. Studying a sample saves time and money.

样本:从总体中选出的、用于代表总体的较小群体。研究样本可以节省时间和成本。

Sampling frame: a complete list of all members of the population from which the sample is drawn, e.g. a school register.

抽样框:总体所有成员的一份完整清单,样本从中抽取,例如学校点名册。

Simple random sampling: every member of the population has an equal chance of being selected. Often uses random number generators or names in a hat.

简单随机抽样:总体中每个成员都有相等的机会被选中。常使用随机数生成器或抽签法。

Stratified sampling: the population is divided into groups (strata) based on a characteristic, then a random sample is taken from each group in proportion to its size.

分层抽样:根据某一特征将总体划分为若干层,然后按比例从每层中随机抽取样本。

Systematic sampling: members are chosen at regular intervals from the sampling frame, e.g. every 10th person on a list after a random start.

系统抽样:从抽样框中按固定间隔选取成员,例如从列表中随机确定起点后,每隔9人取一个。

Quota sampling: interviewers select a fixed number of people from different categories until quotas are filled. It is non‑random and prone to bias.

配额抽样:访问员从不同类别中选择固定数量的人,直到配额填满。这是非随机方法,容易产生偏差。

Opportunity sampling: the sample is taken from people who are available and willing to take part at the time of the study. Also called convenience sampling.

机会抽样:从研究时在场且愿意参与的人中选取样本。也称为便利抽样。


3. Data Collection Terms | 数据收集术语

Primary data: data that you collect yourself for a specific purpose, such as conducting a survey or an experiment.

原始数据:你为了特定目的亲自收集的数据,例如开展调查或实验。

Secondary data: data that has been collected by someone else for a different purpose, like government statistics, newspaper articles, or websites.

二手数据:由他人为其他目的收集的数据,如政府统计数据、报纸文章或网站信息。

Census: a survey that collects data from every member of the population. It gives very accurate results but is expensive and time‑consuming.

普查:从总体每个成员那里收集数据的调查。结果非常准确,但昂贵且耗时。

Pilot survey: a small trial run of a questionnaire or data collection method to identify problems before the main study.

试点调查:在主研究之前,对问卷或数据收集方法进行小规模试运行,以发现问题。

Open question: a question that allows respondents to answer in their own words, producing qualitative, detailed data.

开放式问题:允许受访者用自己的话回答的问题,产生详细、定性的数据。

Closed question: a question that provides a set of answer options to choose from, making data easier to process and analyse.

封闭式问题:提供一组答案选项供选择的问题,使数据更易于处理和分析。


4. Charts and Diagrams | 图表与图示

Bar chart: uses bars of equal width to represent frequencies or values for categorical data. The height of each bar shows the frequency.

条形图:用等宽条形表示分类数据的频数或数值。每个条形的高度表示频数。

Pie chart: a circular chart divided into sectors, where each sector’s angle is proportional to the frequency it represents.

饼图:一个划分成扇区的圆形图,每个扇区的角度与其所代表的频数成正比。

Histogram: a diagram for grouped continuous data, where the area of each bar is proportional to the frequency. Bar widths can vary if class intervals are unequal.

直方图:用于分组连续数据的图形,每个条形的面积与频数成正比。如果组距不等,条形的宽度也会不同。

Frequency polygon: a line graph formed by joining the midpoints of the tops of histogram bars (or class midpoints vs frequency). Often used to compare distributions.

频数多边形:通过连接直方图各条形顶端中点(或组中点与频数)形成的折线图。常用于比较分布。

Cumulative frequency curve: a graph showing the running total of frequencies. Used to find medians, quartiles, and percentiles.

累积频数曲线:显示频数累积总数的图形。用于求中位数、四分位数和百分位数。

Box plot: a diagram displaying the five‑number summary: minimum, lower quartile, median, upper quartile, and maximum. It highlights spread and outliers.

箱线图:展示五数概括的图形:最小值、下四分位数、中位数、上四分位数和最大值。能凸显离散程度和异常值。


5. Measures of Central Tendency | 集中趋势度量

Mean: the average found by adding all data values and dividing by the number of values. Sensitive to extreme values.

平均数(均值):将所有数据值相加后除以数值个数得到的平均值。易受极端值影响。

Median: the middle value when data is ordered from smallest to largest. Not affected by outliers, so it is useful for skewed distributions.

中位数:将数据从小到大排序后位于中间的值。不受异常值影响,因此对偏态分布很有用。

Mode: the value that occurs most frequently. A data set can have one mode, more than one mode, or no mode at all.

众数:出现次数最多的值。数据集可以有一个众数、多个众数或没有众数。

Modal class: for grouped data, the class interval with the highest frequency. Used when individual values are not known.

众数组:对于分组数据,频数最高的那个组距区间。在无法知道具体数值时使用。

Weighted mean: a mean where some values contribute more than others based on assigned weights. Formula: ∑(w × x) ÷ ∑w, where w is the weight and x is the data value.

加权平均数:根据所赋权重,有些值贡献更大的平均数。公式:∑(w × x) ÷ ∑w,其中 w 为权重,x 为数据值。


6. Measures of Spread | 离散程度度量

Range: the difference between the largest and smallest values. It is a quick but rough measure of spread.

极差(全距):最大值与最小值的差。是一个快速但粗略的离差度量。

Quartiles: values that split the ordered data into four equal parts. The lower quartile (Q₁) is the median of the lower half, the upper quartile (Q₃) is the median of the upper half.

四分位数:将排序后的数据分成四等份的值。下四分位数 Q₁ 是下半部分的中位数,上四分位数 Q₃ 是上半部分的中位数。

Interquartile range (IQR): IQR = Q₃ – Q₁. It measures the spread of the middle 50% of data and is resistant to outliers.

四分位距 (IQR):IQR = Q₃ – Q₁。衡量中间50%数据的离散程度,不受异常值影响。

Standard deviation: a measure of how spread out the data values are around the mean. A low standard deviation means values are close to the mean. The formula for a population is σ = √(∑(x – μ)² / N). For a sample we use s = √(∑(x – x̄)² / (n – 1)).

标准差:衡量数据值围绕平均值的离散程度。标准差小表示数值靠近平均值。总体标准差公式为 σ = √(∑(x – μ)² / N),样本标准差为 s = √(∑(x – x̄)² / (n – 1))。

Variance: the square of the standard deviation. It is often used in more advanced statistical calculations.

方差:标准差的平方。常用于更高级的统计计算。

Outlier: an extreme value that lies far away from the rest of the data. Outliers can be identified using 1.5 × IQR rule or by checking data values more than 2 standard deviations from the mean.

异常值:远离其他数据点的极端值。可通过 1.5×IQR 法则,或检查超过平均值正负两个标准差的数据值来识别。


7. Probability Basics | 概率基础

Experiment: a repeatable process that gives rise to a set of outcomes, such as rolling a dice or flipping a coin.

试验:可重复的、产生一系列结果的过程,例如掷骰子或抛硬币。

Outcome: a possible result of an experiment. For a dice roll, outcomes are 1, 2, 3, 4, 5, 6.

结果:试验的一个可能结果。掷骰子的结果是 1、2、3、4、5、6。

Event: a set of one or more outcomes. An event could be ‘rolling an even number’ which includes the outcomes 2, 4, 6.

事件:一个或一组结果的集合。例如“掷到偶数”是一个事件,包括结果 2、4、6。

Sample space: the list of all possible outcomes of an experiment. It is often shown in a table or a list.

样本空间:试验所有可能结果的清单。通常用表格或列表表示。

Mutually exclusive events: events that cannot happen at the same time. P(A and B) = 0 if A and B are mutually exclusive.

互斥事件:不可能同时发生的事件。若 A、B 互斥,则 P(A 且 B) = 0。

Independent events: the outcome of one event does not affect the probability of the other. P(A and B) = P(A) × P(B) when independent.

独立事件:一个事件的结果不影响另一个事件的概率。独立时 P(A 且 B) = P(A) × P(B)。

Relative frequency: an estimate of probability calculated from experimental data: number of times an event occurs ÷ total number of trials. It tends towards theoretical probability with more trials.

相对频率:由实验数据估算的概率:事件发生次数 ÷ 总试验次数。随着试验次数增多,它会趋近理论概率。

Theoretical probability: the probability of an event based on equally likely outcomes, calculated as (number of favourable outcomes) ÷ (total number of possible outcomes).

理论概率:基于等可能结果的事件概率,计算公式为(有利结果数)÷(可能结果总数)。


8. Correlation and Regression | 相关与回归

Correlation: a statistical measure of the relationship between two variables. It can be positive, negative, or zero (no correlation).

相关:两个变量之间关系的统计度量。可以是正相关、负相关或零相关(无相关)。

Positive correlation: as one variable increases, the other also tends to increase. The scatter graph slopes upwards.

正相关:一个变量增大时,另一个也趋于增大。散点图呈向上倾斜。

Negative correlation: as one variable increases, the other tends to decrease. The scatter graph slopes downwards.

负相关:一个变量增大时,另一个趋于减小。散点图呈向下倾斜。

Line of best fit: a straight line drawn through the middle of a scatter graph to model the relationship. It should have roughly the same number of points above and below the line.

最佳拟合线:穿过散点图中部的直线,用于模拟变量关系。线上下的点数应大致相等。

Interpolation: estimating a value within the range of the given data using the line of best fit. This is usually reliable.

内插法:利用最佳拟合线,在给定数据范围内估计一个值。通常比较可靠。

Extrapolation: estimating a value outside the range of the data. It can be unreliable because the trend may not continue beyond the data points.

外推法:在数据范围之外估计一个值。可能不可靠,因为趋势在数据点之外不一定延续。

Spearman’s rank correlation coefficient: a value between –1 and +1 that measures the strength of a monotonic relationship between two variables, without assuming a linear relationship.

斯皮尔曼等级相关系数:介于 -1 到 +1 之间的值,衡量两个变量间单调关系的强度,不假设线性关系。


9. Time Series Analysis | 时间序列分析

Time series: a set of data points recorded in time order, often at regular intervals, such as monthly sales figures or annual temperatures.

时间序列:按时间顺序记录的一组数据点,通常间隔相等,如月销售额或年温度。

Trend: the long‑term movement or general direction in a time series, ignoring short‑term fluctuations. It can be upward, downward, or flat.

趋势:时间序列中长期的运动或大方向,忽略短期波动。可以是上升、下降或持平。

Seasonal variation: regular and predictable changes that occur in the same period each year or each week. For example, ice cream sales rising in summer.

季节性变动:每年或每周同一时期发生的规律性、可预测的变化。例如冰淇淋销量在夏季上升。

Moving average: a series of averages calculated from successive fixed‑size subsets of the time series data. It smooths out short‑term fluctuations to reveal the trend.

移动平均:从时间序列数据的连续固定大小子集计算出的平均值序列。它能平滑短期波动,突显趋势。

Cyclic fluctuation: patterns that repeat over more than one year, often linked to economic cycles. Unlike seasonal variation, the period is not fixed.

循环波动:超过一年重复出现的模式,通常与经济周期相关。与季节性变动不同,其周期不固定。

Random fluctuation: unpredictable, irregular movements in a time series that remain after trend, seasonal, and cyclic components are removed.

随机波动:剔除趋势、季节性和循环因素后,时间序列中不可预测的不规则变动。


10. Bias and Reliability | 偏差与可靠性

Bias: any factor that causes the results of a survey or experiment to be systematically different from the true population value. Biased data leads to invalid conclusions.

偏差:导致调查或实验结果与真实总体值系统性偏离的任何因素。有偏差的数据会得出无效结论。

Sampling bias: occurs when the sample is not representative of the population, often because some groups are over‑ or under‑represented.

抽样偏差:当样本不能代表总体时发生,常因为某些群体被过度代表或代表不足。

Non‑response bias: happens when people selected for the sample do not respond and their views differ from those who do respond.

无回应偏差:当被选入样本的人不回应,且他们的看法与回应者不同时产生的偏差。

Response bias: occurs when respondents give inaccurate answers, often because of poorly worded questions or sensitive topics.

回答偏差:受访者给出不准确答案,常因问题措辞不当或话题敏感而产生。

Leading question: a question that encourages a particular answer, e.g. ‘Don’t you agree that the park is beautiful?’ This introduces response bias.

诱导性问题:鼓励特定回答的问题,例如“你难道不认为这个公园很美吗?”这种问题会引入回答偏差。

Reliability: the extent to which a measurement or survey produces consistent results under the same conditions. Reliable data is repeatable.

可靠性:测量或调查在相同条件下得出结果的一致性程度。可靠数据是可重复的。

Validity: the extent to which data measures what it is intended to measure. An experiment may be reliable but not valid if it measures the wrong variable.

效度:数据测量其预定测量内容的程度。如果一个实验测量了错误的变量,它可能可靠但无效。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading