📚 Year 11 OCR Statistics: Rapid Terminology Memorisation Guide | Year 11 OCR 统计:词汇术语速记指南
In OCR GCSE Statistics, a firm grasp of terminology is just as important as being able to perform calculations. Examiners expect you to use precise language when describing data, probability, sampling and statistical relationships. This guide breaks down the key vocabulary into thematic groups, provides a clear definition for every term, and includes simple memory hooks to help you recall them under exam pressure. Work through each section, cover the definitions, and test your recall both ways – from English to Chinese, and back again.
在 OCR GCSE 统计学中,扎实掌握术语与掌握计算同样重要。考官要求你在描述数据、概率、抽样和统计关系时使用精确的语言。本指南将核心词汇按主题分组,为每个术语提供清晰定义,并附上简洁的记忆挂钩,帮助你在考试压力下快速回忆。请逐节学习,遮挡定义后自测,并尝试英汉互译,确保彻底记住。
1. Core Statistical Terms | 核心统计术语
The foundation of every statistical investigation rests on a few key words. If you confuse ‘population’ with ‘sample’, your whole interpretation can collapse. Commit these to memory first.
每一项统计调查的基础都建立在几个关键词上。如果你混淆了“总体”和“样本”,整个解读都可能崩塌。请优先记牢这些词。
Population – the entire collection of individuals, objects or measurements that you are investigating. Think of it as the whole country’s population: everyone is included.
总体 —— 你所研究的所有个体、对象或测量值的整体集合。想象它是一个国家的全体人口:所有人都包含在内。
Sample – a subset of the population selected to represent the whole. Remember that a ‘sample’ of food only gives a taste, not the entire meal.
样本 —— 从总体中选出的一部分,用来代表整体。联想一下:品尝食物的小样无法替代整份餐点。
Variable – any characteristic that can vary from one member of the population to another. For example, ‘height’ is a variable because it changes from person to person.
变量 —— 在总体中不同成员间会发生变化的任何特征。例如“身高”就是一个变量,因人而异。
Data – the actual values collected for the variables you are studying. The word comes from the Latin for ‘something given’, so treat your data as the facts given to you by the investigation.
数据 —— 为所研究变量收集到的实际值。该词源自拉丁语“被给予之物”,因此请始终将数据视作调查提供给你的事实。
2. Types of Data | 数据类型
Classifying data correctly determines which calculations and diagrams are appropriate in OCR exam questions. Many students lose marks by confusing discrete and continuous data when choosing a histogram or a bar chart.
正确分类数据决定了你在 OCR 考题中应选用哪种计算方法和图表。许多学生因在选择直方图或条形图时混淆离散与连续数据而丢分。
Qualitative (categorical) data – non‑numerical information that describes qualities, such as hair colour or favourite film. To remember this, think ‘quality’ not ‘quantity’.
定性(分类)数据 —— 描述品质、类别等非数字信息,例如发色或最爱的电影。记忆时联想“性质”而非“数量”。
Quantitative data – numerical information that can be measured or counted. The word ‘quant’ inside it hints at quantity.
定量数据 —— 可测量或可计数的数字信息。单词中的“quant”暗示着数量。
Discrete data – quantitative data that can only take specific, separate values, such as the number of students in a class. Think of distinct blocks that you can count with your fingers.
离散数据 —— 只能取特定、分隔数值的定量数据,例如班级里的学生人数。想象成一个个能用手指点清的不同方块。
Continuous data – quantitative data that can take any value along a scale, like height or temperature. It flows continuously, like water, without sudden breaks.
连续数据 —— 能在尺度上取任意值的定量数据,如身高或温度。它像水流般连续,没有突然断裂。
3. Measures of Central Tendency | 集中趋势度量
Central tendency measures attempt to summarise a whole data set with a single ‘typical’ value. The three principal ones – mean, median, mode – each have strengths and weaknesses that examiners love to test.
集中趋势度量试图用一个“典型”值概括整个数据集。均值、中位数和众数这三大指标各有优缺点,考官特别喜欢就此出题。
Mean, x̄ = Σx / n
Mean – the arithmetic average, found by summing all values and dividing by the number of values. It is sensitive to extreme values (outliers). Memory tip: the ‘mean’ teacher averages everyone’s grade.
均值 —— 算术平均数,将所有数值求和再除以数值个数。它对极端值(离群值)很敏感。记忆法:“均值”老师把所有人的分数平均了。
Median – the middle value when the data are arranged in order. Panic about the middle? ‘Median’ has the same first three letters as ‘middle’. It is robust against outliers.
中位数 —— 将数据按顺序排列后正中间的值。担心中间?Median 的前三个字母与 middle 相同。中位数不易受离群值影响。
Mode – the most frequently occurring value. Mode and ‘most’ both start with ‘mo’, and the shirt you wear ‘most’ often is in your ‘mode’ wardrobe.
众数 —— 出现频率最高的值。Mode 和 most 都以 mo 开头;你“最”常穿的那件衬衫就是你衣橱里的“时尚 (mode)”。
4. Measures of Spread | 离散程度度量
Central tendency alone cannot describe how varied the data are. Two sets can have the same mean but completely different spreads. OCR questions frequently ask you to compare distributions using spread measures.
仅靠集中趋势无法描述数据的变异程度。两组数据可以有相同的均值,但离散程度截然不同。OCR 题目经常要求你使用离散度量比较分布。
Range = max value – min value
Range – the difference between the largest and smallest values. Quick to compute but affected by only two points. Picture the spread of a mountain ‘range’ from peak to valley.
全距 —— 最大值与最小值之差。计算快捷,但仅取决于两个点。想象山脉从峰顶到谷底的全距。
Interquartile range, IQR = Q₃ – Q₁
Interquartile range (IQR) – the range of the middle 50% of the data, calculated as the upper quartile (Q₃) minus the lower quartile (Q₁). It resists extreme values. Memory trick: the ‘I’ in IQR stands for ‘inner’ 50%.
四分位距 (IQR) —— 中间 50% 数据的全距,等于上四分位数 (Q₃) 减下四分位数 (Q₁)。它能抵抗极端值。记忆窍门:IQR 中的 ‘I’ 代表 ‘inner 50%’。
Variance – the average of the squared deviations from the mean. It forms the foundation of more advanced statistics. Think of it as the ‘variance’ menu: everything varies around the mean.
方差 —— 各数值与均值之差的平方的平均数。它是更高级统计的基础。想象一个“方差”菜单:一切围绕均值变化。
Sample standard deviation, s = √[ Σ(xᵢ – x̄)² / (n-1) ]
Standard deviation – the square root of the variance, bringing the spread back to the original units. It measures the typical distance of a data point from the mean. If the numbers deviate a lot, your standard of living deviates!
标准差 —— 方差的平方根,将离散度拉回原始单位。它衡量数据点与均值的典型距离。若数字偏差很大,你的“标准”生活就偏离了。
5. Charts and Graphs | 图表
Selecting and interpreting diagrams is a core skill in OCR Statistics. The examiner will expect you to know when a histogram is required and how it differs from a bar chart.
选择并解读图表是 OCR 统计学的核心技能。考官要求你知晓何时该用直方图,以及它与条形图的区别。
Bar chart – used for qualitative data or discrete quantitative data. Bars have equal width and gaps between them. Think ‘bar’ as in separate counters in a bar game.
条形图 —— 用于定性数据或离散定量数据。条形等宽,条形间有空隙。想象酒吧 (bar) 里分开的一张张高脚凳。
Pie chart – a circular chart divided into sectors proportional to frequency. It shows how a whole is divided. The ‘pie’ has slices that together make the full dessert.
饼图 —— 被按频率比例分割成扇区的圆形图表,展示整体如何被切分。整个饼 (pie) 分成多块,合起来就是完整的甜点。
Histogram – used for continuous data or grouped discrete data. Bars have varying width (if class widths differ) and no gaps; area represents frequency. Memory aid: no gaps, so the ‘histo’ry flows without interruption.
直方图 —— 用于连续数据或分组离散数据。条形宽度可能不同(若组距不等),条形间无空隙;面积代表频率。记忆法:无空隙,历史 (history) 如流。
Frequency polygon – formed by joining the mid‑points of the tops of histogram bars (or class mid‑points) with straight lines. Think of a polygon as a shape that connects many points.
频数折线图 —— 用直线连接直方图条顶中点(或组中值)形成。把它看作连接多个点的多边形 (polygon)。
Cumulative frequency curve (ogive) – a running total of frequencies plotted against the upper class boundaries. Useful for finding medians and quartiles. The S‑shaped curve is like an ‘oh‑give’ me the middle value.
累积频数曲线(肩形图) —— 将频数累积和对应上组界绘制而成,用于求中位数和四分位数。S 形曲线仿佛在说“哦,给我中位值吧”。
6. Probability Language | 概率语言
Probability terms must be used with absolute precision in exam answers. Statements like ‘mutually exclusive’ or ‘independent’ carry specific conditions that you must be able to describe and apply.
在考试答案中必须绝对精确地使用概率术语。“互斥”或“独立”等表述带有特定条件,你必须能够描述并应用它们。
Random experiment – a process whose outcome cannot be predicted with certainty, such as tossing a fair coin. The key is ‘random’ means every possible outcome is equally likely if the process is fair.
随机试验 —— 结果无法事先确定的过程,例如投掷公平硬币。关键在于,“随机”意味着如果是公平过程,所有可能结果等可能发生。
Outcome – a single possible result of an experiment, e.g. getting a ‘5’ on a die.
结果 —— 试验的某个可能出现的结果,例如掷骰子得到“5”。
Event – a collection of one or more outcomes that share a specified property, e.g. rolling an even number.
事件 —— 一个或多个具有指定属性的结果的集合,例如掷出偶数。
P(A ∪ B) = P(A) + P(B) – P(A ∩ B)
Mutually exclusive events – events that cannot happen at the same time, so P(A ∩ B) = 0. ‘Exclusive’ means they exclude each other; you cannot be in two exclusive clubs at once.
互斥事件 —— 不可能同时发生的事件,因此 P(A ∩ B) = 0。’Exclusive’ 表示互相排斥;你无法同时加入两个互斥俱乐部。
For independent events: P(A and B) = P(A) × P(B)
Independent events – the occurrence of one event does not affect the probability of the other. ‘Independent’ people don’t influence each other’s decisions.
独立事件 —— 一个事件的发生不影响另一个事件的概率。“独立”的人在做决定时互不影响。
7. Correlation and Regression | 相关与回归
When studying two variables together (bivariate data), you need language to describe their relationship and to make predictions from a line of best fit. OCR often presents scatter graphs and asks you to comment on correlation.
当同时研究两个变量(双变量数据)时,你需要用语言描述它们的关系,并利用最佳拟合线进行预测。OCR 常常给出散点图并要求你对相关性作出评论。
Bivariate data – data that involves two variables measured on the same individuals, such as hours of revision and exam score. ‘Bi-‘ means two: bicycle has two wheels; bivariate has two variables.
双变量数据 —— 在同一对象上测量两个变量得到的数据,如复习时间与考试成绩。“Bi-”意为二:自行车有两个轮子;双变量有两个变量。
Scatter graph (scatter plot) – a graph on which two variables are plotted as points along two axes. It reveals patterns of association.
散点图 —— 将两个变量沿两轴绘制为点对的图形,能揭示关联模式。
Positive correlation – as one variable increases, the other also tends to increase. Think of a positive attitude that lifts both variables upward.
正相关 —— 当一个变量增加时,另一个也倾向于增加。想象积极的态度把两个变量一起推高。
Negative correlation – as one variable increases, the other tends to decrease. A negative sign points downward.
负相关 —— 一个变量增加而另一个趋于减少。负号指向下方。
Line of best fit – a straight line drawn through a scatter graph to model the relationship. It can be used to estimate unknown values. Always remember: interpolation (predicting within the data range) is safer than extrapolation (predicting outside).
最佳拟合线 —— 穿过散点图的直线,用于建模关系并估计未知值。始终记住:内插(在数据范围内预测)比外推(超出数据范围预测)更可靠。
Spearman’s rank correlation coefficient (rₛ) – a number between -1 and 1 that measures the strength and direction of monotonic association. It uses ranked data rather than raw values.
斯皮尔曼等级相关系数 (rₛ) —— 介于 -1 和 1 之间的数值,衡量单调关系的强度与方向。它使用排秩数据而非原始值。
8. Sampling Techniques | 抽样方法
Choosing a sample wisely is vital to avoid bias. OCR expects you to describe, compare and critique different sampling methods. Learn their names and the mechanics behind each one.
明智地选择样本对避免偏差至关重要。OCR 要求你描述、比较并评析不同抽样方法。请记住它们的名称和背后的机制。
Simple random sampling – every member of the population has an equal chance of being selected, often using random number generators. Like drawing names from a hat where everyone’s name truly is in the hat.
简单随机抽样 —— 总体中每个成员被选中的机会相等,通常借助随机数生成器。就像从一顶帽子里抽名字,而且每个人的名字确确实实都在帽子里。
Stratified sampling – the population is split into distinct groups (strata), and a random sample is taken from each in proportion to its size. It guarantees representation. Think of ‘strata’ as layers of a cake – you sample from every layer.
分层抽样 —— 将总体分成互不重叠的组(层),然后按比例从各层随机抽取样本,确保代表性。把“层”想象成蛋糕的各层——每层都要取一点。
Systematic sampling – items are chosen at regular intervals from a list, e.g. every 10th person. The system saves time but can miss patterns hidden in the order.
系统抽样 —— 从名单中按固定间隔选取对象,例如每隔10人抽取。该方法省时,但可能遗漏顺序中隐藏的模式。
Quota sampling – the interviewer selects a fixed number (quota) of people from predefined categories, but convenience often replaces randomness. It is fast but prone to interviewer bias.
配额抽样
Published by TutorHao | Year 11 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导