Year 10 OCR Statistics: Vocabulary and Terminology Quick Reference Guide | Year 10 OCR 统计:词汇术语速记指南

📚 Year 10 OCR Statistics: Vocabulary and Terminology Quick Reference Guide | Year 10 OCR 统计:词汇术语速记指南

Welcome to your essential quick-reference guide for mastering statistical vocabulary in the Year 10 OCR Statistics course. This article breaks down key terms into bite-sized explanations, each paired with a memorable mnemonic or shortcut so that you can recall definitions accurately under exam pressure. From data types and sampling methods to probability and correlation, we cover every term you are likely to encounter. The bilingual format allows you to reinforce understanding in both English and Chinese, making revision more effective and less stressful.

欢迎来到 Year 10 OCR 统计课程的核心术语速记指南。本文将关键概念拆解成一口大小的解释,并为每个术语配上易记的口诀或捷径,帮助你在考试压力下准确回忆定义。从数据类型、抽样方法到概率与相关关系,我们涵盖了你可能会遇到的所有术语。中英双语的形式让你能用两种语言巩固理解,使复习更加高效、轻松。


1. Types of Data | 数据类型

Qualitative data describes qualities or categories, such as eye colour or favourite sport. Think of ‘quality’ to remember it deals with descriptions, not numbers.

定性数据描述的是性质或类别,例如眼睛颜色或最喜爱的运动。想到“性质”这个词,就能记住它处理的是描述而非数字。

Quantitative data is numerical and can be measured. The word ‘quantity’ points to numbers. It splits further into discrete (countable, whole numbers like number of students) and continuous (measurable, any value like height).

定量数据是数值型的、可以测量的数据。“数量”一词指向数字。它进一步分为离散数据(可数的,整数,如学生人数)和连续数据(可测量的,任意值,如身高)。

Quick tip: Discrete has a ‘t’ like ‘count’, continuous has an ‘o’ like ‘flow’ — it can take any value in a range.

速记提示:Discrete(离散)里有个“t”像“count”(数数),continuous(连续)里有个“o”像“flow”(流动)——它可以在一个范围内取任意值。


2. Sampling Methods | 抽样方法

Random sampling gives every member of the population an equal chance of being selected. Imagine picking names out of a hat — pure luck, no bias.

随机抽样让总体中的每个成员都有相等的被选中的机会。想象从帽子里抽名字——纯粹的运气,没有偏差。

Stratified sampling divides the population into groups (strata) and then randomly selects from each group in proportion to its size. Remember ‘strata’ layers, like a layered cake — each layer is represented.

分层抽样把总体分成若干个组(层),然后按比例从每组中随机选取。记住“strata”是层的意思,像千层蛋糕——每一层都被代表到了。

Systematic sampling selects every kth individual after a random start. It is like clockwork — systematic and regular. Use ‘system’ to recall the fixed interval.

系统抽样在随机起点之后每隔k个个体选取一个。就像钟表一样有规律——系统且规则。用“系统”这个词来联想固定间隔。

Convenience sampling uses people who are easy to reach, like asking friends. It is quick but often biased. The word ‘convenience’ signals ‘easy for me’, not necessarily representative.

便利抽样使用容易接触到的人,比如问朋友。它快速但常有偏差。“便利”这个词暗示“方便我”,但不一定有代表性。


3. Measures of Central Tendency | 集中趋势的度量

Mean is the arithmetic average: sum of all values divided by the number of values. Think of ‘mean’ as the ‘sharing equally’ measure — if everyone got the same amount, that is the mean.

平均数是算术平均值:所有数值之和除以数值的个数。把“mean”想象成“公平分享”的度量——如果每个人都得到一样的量,那就是平均数。

Median is the middle value when data are ordered. The word ‘median’ has ‘mid’ inside it. For an even number of values, it is the mean of the two middle numbers.

中位数是数据排序后位于中间的数值。“median”一词里含有“mid”(中间)。对于偶数个数值,它是中间两个数的平均数。

Mode is the most frequent value. ‘Mode’ sounds like ‘most’ — it is the value that appears most often. A data set can have more than one mode (bimodal, multimodal) or none at all.

众数是出现频率最高的数值。“Mode”听起来像“most”(最多)——它就是出现次数最多的值。一个数据集可以有多个众数(双众数、多重众数),也可以没有众数。

Mnemonic: Mean = Everyone shares, Median = Middle when lined up, Mode = Most seen.

记忆口诀:平均数人人分,中位数中间站,众数最多见。


4. Measures of Spread | 离散程度的度量

Range is the difference between the largest and smallest values. It gives a quick sense of spread but ignores the middle. ‘Range’ from smallest to largest — like the range of a mountain chain.

极差是最大值与最小值之差。它能快速反映数据的分散程度,但忽略了中间部分。“Range”从最小到最大——像山脉的范围。

Interquartile range (IQR) is the range of the middle 50% of data: upper quartile (Q3) minus lower quartile (Q1). It is resistant to outliers. ‘Inter’ means between, so it is the spread between quartiles.

四分位距(IQR)是中间50%数据的范围:上四分位数(Q3)减去下四分位数(Q1)。它对异常值有抵抗力。“Inter”是之间的意思,所以它是四分位数之间的差距。

Standard deviation measures how far, on average, data points are from the mean. A small standard deviation means data are clustered closely; a large one means they are spread out. Think ‘standard’ as the typical distance from the centre.

标准差衡量数据点平均离均值有多远。标准差小意味着数据紧密聚集;标准差大意味着数据分散。把“standard”想象成与中心之间的典型距离。

Quick recall: Range = max − min; IQR = Q3 − Q1; Standard deviation = typical deviation from mean.

速记:极差 = 最大 − 最小;四分位距 = Q3 − Q1;标准差 = 与均值的典型偏差。


5. Charts and Graphs | 图表

Bar chart uses bars of equal width to show frequency for categories. The bars do not touch, which helps you remember it is for discrete or categorical data. ‘Bar’ like separate blocks.

条形图用等宽的条形来显示各类别的频数。条形之间有空隙,这提醒你它是用于离散数据或分类数据。“Bar”就像分离的积木。

Histogram also uses bars, but they touch and the area of each bar represents frequency. It is used for continuous data. Think ‘history’ — continuous timeline, bars touching.

直方图也使用条形,但条形紧密相连,并且每个条形的面积代表频数。它用于连续数据。联想“history”(历史)——连续的时间线,条形相连。

Pie chart shows proportions as slices of a circle. The whole circle is 360°, so each category gets an angle. ‘Pie’ reminds you it is about parts of a whole.

饼图以圆形切片来显示比例。整个圆是360°,所以每个类别得到一个角度。“Pie”让你想起它是整体的一部分。

Scatter graph plots two variables as points on a grid. It shows correlation. ‘Scatter’ because points are scattered across the plane, not joined.

散点图将两个变量以点的形式绘制在网格上。它显示相关关系。“Scatter”(散布)因为点分散在平面上,不连线。

Line graph connects data points with lines, often used for time series. ‘Line’ implies trend over time.

折线图用线段连接数据点,常用于时间序列。“Line”(线)暗示随时间变化的趋势。


6. Frequency and Cumulative Frequency | 频数与累积频数

Frequency is simply the count of how often something occurs. ‘Frequent’ means often — so frequency is the number of occurrences.

频数就是某事发生的计数。“Frequent”(频繁)意味着经常——所以频数就是发生的次数。

Cumulative frequency is the running total of frequencies up to each point. ‘Cumulative’ shares a root with ‘accumulate’ — keep adding up. A cumulative frequency graph plots these running totals and is useful for finding medians and quartiles directly.

累积频数是截至每个点的频数累计总和。“Cumulative”与“accumulate”(积累)同源——不断累加。累积频数图绘制这些累计总数,可用于直接查找中位数和四分位数。

Tip: The cumulative frequency curve is always going up — it never decreases. If it flattens, no new data was added in that interval.

提示:累积频数曲线始终上升——永不下降。如果它变平,说明在该区间没有新增数据。


7. Probability Basics | 概率基础

Probability measures how likely an event is, on a scale from 0 (impossible) to 1 (certain). Probability of event A is written P(A). The formula: P(A) = number of favourable outcomes ÷ total number of possible outcomes.

概率衡量一个事件发生的可能性,范围从0(不可能)到1(必然)。事件A的概率记为P(A)。公式:P(A) = 有利结果的数量 ÷ 所有可能结果的总数。

Mutually exclusive events cannot happen at the same time. If one happens, the other cannot. Think ‘mutually’ — they refuse to co-exist. P(A or B) = P(A) + P(B) for mutually exclusive events.

互斥事件不能同时发生。一个发生,另一个就不能发生。想到“互相排斥”——它们拒绝共存。对于互斥事件,P(A 或 B) = P(A) + P(B)。

Independent events are those where the outcome of one does not affect the other. ‘Independent’ like independent people — one’s action does not change the other’s probability. P(A and B) = P(A) × P(B).

独立事件是一个事件的结果不影响另一个事件。’Independent’像独立的人——一个人的行为不会改变另一个人的概率。P(A 和 B) = P(A) × P(B)。

Sample space is the set of all possible outcomes. Imagine a ‘space’ containing all that can happen. It is often shown in a list or a two-way table.

样本空间是所有可能结果的集合。想象一个“空间”,里面装着所有可能发生的事情。它通常用列表或双向表格表示。

Venn diagram uses overlapping circles to show relationships between events. The overlap shows intersection (A ∩ B). ‘Venn’ like ‘when’ — when they overlap.

维恩图用重叠的圆来表示事件之间的关系。重叠部分表示交集(A ∩ B)。”Venn”听起来像“when”——当它们重叠时。


8. Correlation and Relationships | 相关与关系

Correlation describes the strength and direction of a linear relationship between two variables. It does NOT imply causation. Remember: ‘correlation is not causation’ — just because two things vary together does not mean one causes the other.

相关性描述两个变量之间线性关系的强度和方向。它并不意味着因果关系。记住:“相关不等于因果”——两个事物一起变化并不代表一个导致了另一个。

Positive correlation: as one variable increases, the other tends to increase. Points on a scatter graph slope upward. Think ‘positive slope’.

正相关:一个变量增大,另一个也趋向增大。散点图上的点向上倾斜。想到“正坡度”。

Negative correlation: as one variable increases, the other tends to decrease. Points slope downward. Think ‘negative slope’.

负相关:一个变量增大,另一个趋向减小。点向下倾斜。想到“负坡度”。

No correlation: points are scattered randomly with no clear pattern. Like a cloud — no trend.

无相关:点随机散布,没有清晰的模式。像一朵云——没有趋势。

Line of best fit is a straight line drawn through the points to model the relationship. It is used to estimate values (interpolation within the data range, extrapolation outside). ‘Best fit’ means it minimises the total distance from points.

最佳拟合线是一条穿过点群的直线,用来模拟关系。它用于估计数值(数据范围内的插值,范围外的外推)。”最佳拟合”意味着它使得点与线的总距离最小。


9. Bias and Reliability | 偏差与可靠性

Bias occurs when a sample or method systematically over- or under-estimates a value. Think of ‘biased’ as ‘unfair’. Common sources: leading questions, non‑response, convenience sampling.

偏差当样本或方法系统地高估或低估某个值时出现。把“biased”想成“不公平”。常见来源:诱导性问题、无回答、便利抽样。

Reliability refers to consistency. If an experiment or survey is repeated, reliable results are similar each time. ‘Reliable’ like a reliable friend — you can count on them to give the same support.

可靠性指一致性。如果重复实验或调查,可靠的结果每次都很相似。“Reliable”像一个可靠的朋友——你可以指望他们每次给出同样的支持。

Validity is about whether you are measuring what you intend to measure. A test can be reliable but not valid (e.g., a weight scale that consistently gives the wrong reading). ‘Valid’ as in ‘true to purpose’.

效度是关于你是否测量了你打算测量的东西。一个测试可以可靠但无效(比如一个体重秤每次都给出错误的读数)。“Valid”即“符合目的”。

Outlier is an extreme value that lies far from the rest of the data. It can distort the mean and range. ‘Outlier’ — lies ‘out’ of the main cluster. Decide carefully whether to include or exclude it.

异常值是一个远离其他数据的极端值。它会扭曲平均数和极差。“Outlier”——位于主群“之外”。要谨慎决定是否包含或排除它。


10. Key Notation and Symbols | 关键符号与记号

Σ (sigma) means ‘sum of’. The Greek capital letter S reminds you to Sum all values. e.g., Σx means sum all x values.

Σ(西格玛)表示“总和”。这个希腊大写字母S提醒你要把所有值相加。例如,Σx 表示所有x值的总和。

x̄ (x-bar) represents the sample mean. The bar over x is like the ‘average bar’ — a common symbol used in statistics. Say ‘x bar’.

x̄(x拔)表示样本平均数。x上的横线就像“平均线”——统计学中常用的符号。读作“x bar”。

n is the sample size (number of data points). ‘n’ for ‘number’.

n是样本容量(数据点的数量)。“n”代表“number”(数量)。

P(A) denotes probability of event A. ‘P’ for probability, parentheses enclose the event.

P(A)表示事件A的概率。“P”代表概率,括号里放入事件。

Q1, Q2, Q3 are the first, second (median) and third quartiles. Q1 splits off the lowest 25%, Q3 the highest 25%. ‘Quart’ like ‘quarter’.

Q1, Q2, Q3是第一、第二(中位数)和第三四分位数。Q1分掉最低的25%,Q3分掉最高的25%。“Quart”像“quarter”(四分之一)。

IQR = Q3 − Q1, the interquartile range. It shows the spread of the middle 50%. ‘Inter’ means between quartiles.

IQR = Q3 − Q1,四分位距。它显示中间50%数据的分散程度。“Inter”表示四分位数之间。

μ (mu) is the population mean (used in formulas for standard deviation). ‘mu’ looks a bit like ‘u’ in ‘population’.

μ(缪)是总体平均数(用于标准差公式)。“mu”有点像“population”里的“u”。

σ (sigma) is the population standard deviation. Lower case σ for spread.

σ(小写西格玛)是总体标准差。小写σ表示离散程度。


11. Statistical Diagrams and Their Uses | 统计图表及其用途

Box plot (box-and-whisker plot) displays minimum, Q1, median, Q3 and maximum. The ‘box’ holds the middle 50% (IQR), while the ‘whiskers’ extend to the extremes. Useful for comparing distributions and spotting outliers.

箱线图(盒须图)显示最小值、Q1、中位数、Q3和最大值。“箱”包含着中间50%的数据(IQR),“须”延伸到两端。它便于比较分布和发现异常值。

Stem-and-leaf diagram shows the shape of data while retaining individual values. The ‘stem’ is the leading digit(s), the ‘leaf’ is the final digit. It is a quick way to sort data and find medians and modes.

茎叶图在展示数据形态的同时保留每个数值。“茎”是前导数字,“叶”是最后一位数字。它是快速排序数据并找出中位数和众数的方法。

Frequency polygon is created by joining the midpoints of histogram bars with straight lines. It helps visualise the shape of a distribution. Think of it as the ‘outline’ of a histogram.

频数多边形通过用直线连接直方图条形顶端的中点而得到。它有助于呈现分布的形状。可以把它想象成直方图的“轮廓”。

Two-way table organises data for two categorical variables. Each cell shows the frequency for a combination. It helps calculate probabilities, especially conditional probabilities.

双向表格为两个分类变量整理数据。每个单元格显示一个组合的频数。它有助于计算概率,尤其是条件概率。


12. Common Exam Terms and Command Words | 常见考试术语与指令词

Describe means say what you see — patterns, trends, shape. Do not explain reasons. For a graph, say ‘as x increases, y tends to increase’.

描述意味着说出你看到的内容——模式、趋势、形状。不要解释原因。对于图表,要说“随着x增加,y倾向于增加”。

Compare means point out similarities and differences, often using statistics like mean, median, range. Use comparative phrases: ‘higher than’, ‘less spread out than’.

比较意味着指出相似和不同之处,通常使用平均数、中位数、极差等统计量。用比较性的短语:“高于”、“比……分散程度小”。

Interpret means explain the meaning of results in the context of the problem. For a probability, say ‘there is a 0.7 chance that a randomly selected person…’

解释意味着在问题情境中说明结果的含义。对于概率,要说“随机选择一个……的概率是0.7”。

Estimate means find an approximate value using a graph or calculation. Read from a line of best fit or use rounding.

估算意味着利用图形或计算找出近似值。从最佳拟合线上读取,或使用四舍五入。

Calculate means work out precisely using a formula or method. Show steps. For mean: add then divide.

计算意味着用公式或方法精确地求出答案。要展示步骤。对于平均数:相加然后除以。

Comment means give an observation based on evidence, often with a statistical justification. Use numbers to back up your point.

评论意味着基于证据给出观察,通常需要统计依据。用数字来支持你的观点。

Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version