📚 Year 10 Eduqas Statistics: A Quick Guide to Memorising Key Vocabulary | Eduqas 十年级统计:关键术语速记指南
Building a strong vocabulary is the foundation of success in Eduqas Year 10 Statistics. When you can instantly recall the meaning of terms like ‘discrete data’ or ‘interquartile range’, reading a question becomes a checklist rather than a puzzle. This guide breaks down the most frequently examined terms into memorable chunks, using associations, contrasts and simple memory hacks to help you lock each definition into your long-term memory.
扎扎实实的词汇功底是攻克 Eduqas 十年级统计的基石。当你能够瞬间回想起“离散数据”或“四分位距”的含义时,解读题目就从猜谜变成了按图索骥。本指南将最高频的考试术语拆解成容易记忆的模块,借助联想、对比和简单的记忆窍门,帮你把每一条定义牢牢锁定在长期记忆中。
1. Types of Data | 数据类型
Data comes in two fundamental families: qualitative and quantitative. Qualitative data describes qualities, categories or labels that cannot be measured with numbers – examples include eye colour, car brands or a rating of ‘poor, good, excellent’. Quantitative data is numerical and can be measured; it splits further into discrete data, which can only take certain isolated values (like shoe size 5, 6, 7 or the number of goals scored), and continuous data, which can take any value within a range (like height, mass or time). Think ‘discrete means you can count the possible jumps’, while ‘continuous means you can always zoom in’.
数据有两大基本类别:定性数据和定量数据。定性数据描述的是无法用数字衡量的品质、类别或标签——例如眼睛颜色、汽车品牌或者“差、好、优秀”的评级。定量数据则是可以用数值衡量的;它又进一步分为离散数据和连续数据。离散数据只能取某些孤立的特定值(比如鞋码 5、6、7 或者进球数),而连续数据则可以在一个区间内取任何值(如身高、质量或时间)。记住:离散数据就是你能“数出一个个台阶”,连续数据则是“可以一直放大细分”。
Within qualitative data, you may also see the terms nominal and ordinal. Nominal data has categories with no natural order (e.g. hair colour), whereas ordinal data suggests a ranking or order but the gaps between ranks are not equal (e.g. satisfaction survey: very dissatisfied, dissatisfied, neutral, satisfied, very satisfied). A quick memory aid: ordinal = order, both start with ‘o’.
在定性数据内部,可能还会见到名词型数据和顺序型数据这两个术语。名词型数据的类别没有天然的顺序(比如发色),而顺序型数据则隐含了等级或排序,但等级之间的间隔并不相等(例如满意度调查:非常不满、不满、一般、满意、非常满意)。一个快速记忆技巧:ordinal(顺序)与 order(顺序)首字母相同,都有“序”。
2. Population vs Sample | 总体与样本
A population is the entire set of individuals or items you are interested in studying – every single Year 10 student in Wales, all the lightbulbs from a factory, or the complete set of yesterday’s tweets about a football match. A sample is a smaller, manageable subset drawn from that population to make data collection practical. The key is that the sample must be representative of the population, otherwise your conclusions will be biased. Use the mnemonic POPulation = the whole bunch.
总体是你有兴趣研究的全部个体或项目的总和——威尔士每一名十年级学生、某工厂生产的所有灯泡,或者昨天所有关于某场足球赛的推文。样本则是从总体中抽出的、便于处理的较小子集,让数据收集变得切实可行。关键在于样本必须对总体具有代表性,否则你的结论就会产生偏倚。用口诀记忆:POPulation(总体)是“全体出动”的那个总。
A census attempts to collect data from every member of the population. It is accurate but often expensive, time-consuming or impossible. A sample survey collects data from a sample and uses it to estimate characteristics of the population. In Eduqas assessments, you will often be asked to recommend whether a census or a sample is more appropriate and to justify your choice.
普查试图从总体中的每一个成员那里收集数据。它准确,但往往成本高昂、耗时甚长甚至不可能实现。抽样调查则从样本中收集数据,并以此估计总体的特征。在 Eduqas 的考核中,经常会要求你建议普查还是抽样更合适,并说明理由。
3. Measures of Central Tendency | 集中趋势度量
Central tendency is a way of describing the ‘typical’ or ‘average’ value in a data set. The three main measures are mean, median and mode. The mean is the sum of all values divided by the number of values – it uses every piece of data but is easily pulled by extreme outliers. The median is the middle value when data are sorted in order; it resists outliers, making it ideal for skewed distributions like house prices. The mode is the most frequent value; it is the only measure you can use for non-numerical data, such as the most popular crisp flavour. Think: Mean = fair share, Median = middle seat, Mode = most wanted.
集中趋势是一种描述数据集中“典型”值或“平均”值的方法。三个主要指标是平均数、中位数和众数。平均数是所有数值之和除以数值个数——它用到了每一个数据点,但极易受极端值拉扯。中位数则是排序后处于正中间的那个数;它不受极端值影响,非常适合描述如房价这样偏斜的数据分布。众数是出现次数最多的值;它是唯一可用于非数值型数据的指标,比如最受欢迎的薯片口味。联想记忆:平均数像分蛋糕(公平分享),中位数像选中间座位,众数就是最想要的那个。
From a grouped frequency table you can only estimate the mean using the midpoints of class intervals, and you can identify the modal class interval, which is the group with the highest frequency. For the median you determine the class interval containing the median position, which is (n + 1) ÷ 2.
在分组频率表中,只能使用组中值来估算平均数,而众数则只能是频数最高的那个组,称为众数组。至于中位数,你需要找出包含第 (n + 1) ÷ 2 个数据位置的那个组距。
4. Measures of Spread | 离散程度度量
Spread tells us how consistent or variable a dataset is. The range is the simplest measure: maximum minus minimum. However, range is highly sensitive to outliers. A more robust alternative is the interquartile range (IQR), which is the difference between the upper quartile (Q₃) and the lower quartile (Q₁). The IQR covers the middle 50% of data. Visualise a box plot: the box itself represents the IQR. To memorise, think ‘IQR = Inside the box, cutting off the extremes’.
离散程度告诉我们一组数据的一致性如何或者变化程度有多大。全距是最简单的度量:最大值减去最小值。但全距对异常值极为敏感。一个更稳健的替代指标是四分位距(IQR),即上四分位数(Q₃)与下四分位数(Q₁)之差。IQR 涵盖了中间百分之五十的数据。想象箱线图:箱子本身代表的就是 IQR。记忆可以抓住:IQR = 箱内的距,切掉极端两边的距。
Standard deviation is another measure of spread that becomes more prominent in later stages of GCSE Statistics. It quantifies how far, on average, each data value is from the mean. A small standard deviation means data points are tightly clustered around the mean; a large one indicates they are widely scattered. In Year 10, you might start by comparing distributions using range, IQR or standard deviation.
标准差是另一种离散度量,在 GCSE 统计的后续阶段会更加突出。它量化了平均来说每个数据值离平均数有多远。标准差小意味着数据点紧密聚集在平均数周围;标准差大则表明数据点分布很广。在十年级,你可以先从比较全距、四分位距或标准差入手分析数据分布。
5. Probability Terminology | 概率术语
Probability is measured on a scale from 0 to 1, where 0 means impossible and 1 means certain. An event with a probability of 0.5 has an even chance. Key words you must know: trial (one attempt or observation), outcome (a possible result), event (a set of one or more outcomes), sample space (the list of all possible outcomes). A sample space diagram or a two-way table helps you systematically count outcomes. To recall the difference between ‘experiment’, ‘trial’ and ‘outcome’: an experiment is rolling a die ten times; each roll is a trial; getting a 4 is an outcome.
概率用 0 到 1 之间的数值来衡量,0 表示不可能,1 表示必然。概率为 0.5 的事件具有均等的机会。必须掌握的术语:试验(一次尝试或观察)、结果(一个可能发生的情况)、事件(一个或多个结果的集合)、样本空间(所有可能结果的列表)。使用样本空间图表或双向表可以帮助你系统地数出全部结果。如何区分“实验”、“试验”和“结果”?实验是掷十次骰子;每一次掷是一次试验;掷出 4 点就是一个结果。
The word mutually exclusive describes events that cannot happen at the same time – like flipping a coin and getting both heads and tails on the same flip. The probability of mutually exclusive events adding together is simply the sum of their individual probabilities. Independent events are those where the outcome of one does not affect the other; for example, rolling a die twice. You multiply probabilities along branches of a tree diagram to find the probability of combined independent events.
互斥事件指不能同时发生的事件——比如掷一枚硬币,不可能同时得到正面和反面。互斥事件的概率相加就是它们各自概率之和。独立事件是指一个事件的结果不影响另一个事件的情形;例如掷两次骰子。在树状图中,沿分支将概率相乘就可以得到独立合并事件发生的概率。
6. Sampling Methods | 抽样方法
When a census is not feasible, statisticians use a sample. Simple random sampling gives every member of the population an equal chance of being selected – often using random number generators or drawing names from a hat. It is fair but can accidentally produce an unrepresentative sample in small populations. Stratified sampling divides the population into distinct subgroups (strata) that are relevant to the study, like age bands or year groups, and then takes a random sample from each stratum in proportion to its size. This guarantees representation but requires knowing the population structure. To remember: Stratified = Strata = structure inside.
当普查不可行时,统计学家会用样本。简单随机抽样给总体中的每一个个体均等的被选中机会——通常利用随机数生成器或抽签。它很公平,但在小总体中可能偶然抽出不具代表性的样本。分层抽样先将总体划分为与研究相关的不同子群(称为层),如年龄段或年级,再从每一层按该层占总体的比例随机抽取样本。这保证了代表性,但需要事先了解总体的构成。记忆:分层 Stratified 中的 Strat 就是层(strata),结构先分清。
Other methods like systematic sampling (selecting every kth member from a list) and convenience sampling (picking individuals who are easy to reach) also appear. Systematic sampling can be quick and even, but beware hidden patterns. Convenience sampling is likely to produce bias. In an exam, justify why a given method may lead to a biased or unrepresentative sample.
其他方法如系统抽样(从名单中每隔 k 个抽取一个)和便利抽样(选择容易接触到的对象)也会出现。系统抽样快捷均匀,但要警惕隐藏的周期性规律。便利抽样很可能产生偏倚。在考试中,要能够论证为何某种抽样方法可能导致样本有偏倚或不具代表性。
7. Data Collection: Primary and Secondary | 数据收集:一手与二手
Primary data is information you collect yourself, first-hand, for a specific purpose. Conducting a questionnaire, measuring plant growth in a lab, or timing your journey to school are all examples. It gives you full control but can be time-consuming. Secondary data is information that already exists, collected by someone else – government statistics, newspaper archives, past experimental results. It is often cheaper and quicker to obtain, but you must check its reliability and to see whether it matches your exact needs. Memory cue: Primary = you are the pioneer; Secondary = second-hand story.
一手数据是你亲自为一特定目的而收集的第一手信息。设计问卷、在实验室测量植物生长、或是记录自己上学所需的时间,都属于一手数据。它让你可以完全掌控,但可能比较耗时。二手数据则是已经存在、由他人收集好的信息——政府统计数据、报纸档案、过往实验结果等。它通常更省钱、更快捷,但你需要检查其可信度,并看它是否完全契合自己的需求。记忆暗示:一手 Primary 就是“排头兵”,二手 Secondary 是“第二手听来的”。
Data can also be classified by how it is collected: through observation, experiment, questionnaire or interview. Observation produces data without interfering; an experiment deliberately changes one variable to see the effect on another. Recognising these contexts will help you answer questions about designing data-collection sheets and identifying possible sources of bias.
数据还可以按收集方式分类:观察、实验、问卷或访谈。观察是在不干预的情况下取得数据;实验则是有意改变一个变量来观察它对另一个变量的影响。认清这些情境,将有助于你回答有关设计数据收集表以及识别可能的偏倚来源的问题。
8. Graphs and Charts | 图表类型
Choosing the right diagram is just as important as calculating the numbers. Bar charts are used for discrete data or categorical data; the gaps between bars remind you that the categories are separate. Histograms resemble bar charts but are for continuous data: the bars touch because the number line flows without gaps, and the area of each bar is proportional to frequency. A frequency polygon joins the midpoints of histogram bars with straight lines; it is excellent for comparing two distributions on the same axes.
选对图形和计算出数据同等重要。条形图用于离散数据或分类数据;条块之间的空隙在提醒你这些类别是相互独立的。直方图形似条形图,但用于连续数据:直条紧挨在一起,因为数轴是连续延伸的,且每个直条的面积与频数成比例。频数多边形将直方图各直条的组中点用直线连接起来;非常适合在同一坐标轴上比较两个分布。
Scatter graphs display paired numerical data to reveal relationships. A line of best fit, drawn by eye or using a computer, must follow the general trend, passing as close to as many points as possible with roughly equal numbers of points above and below the line. Outliers on a scatter graph are points that lie far from the general pattern; they may be errors or genuinely unusual values and should be commented on.
散点图展示成对的数值型数据以揭示关系。一条最佳拟合线(凭目测或计算机画出)必须遵循总体趋势,尽可能地从靠近更多的点,且线两侧的点数大致相等。散点图中的异常值是那些远离整体模式的点;它们可能是错误,也可能是真正特殊的数值,需要加以说明。
Pie charts show proportions of a whole; each slice’s angle equals (frequency ÷ total frequency) × 360°. Cumulative frequency curves require you to add frequencies step by step and plot against the upper class boundary; they allow you to estimate medians and quartiles directly from the graph.
饼图展示各部分占整体的比例;每个扇形的角度等于(该组频数 ÷ 总频数)× 360°。累积频数曲线需要你逐步累加频数并对着组上限描点;利用它可以直接从图上估算中位数和四分位数。
9. Correlation and Regression | 相关与回归
Correlation describes the strength and direction of a linear relationship between two variables. Positive correlation means as one variable increases, the other tends to increase (e.g. hours of revision and test scores). Negative correlation means as one increases, the other tends to decrease (e.g. time spent on social media and concentration). If no pattern emerges, we say there is no correlation. Correlation does not imply causation – just because two things move together does not prove one causes the other.
相关描述了两个变量之间线性关系的强度和方向。正相关意味着当一个变量增加时,另一个变量也趋于增加(比如复习小时数和测验成绩)。负相关意味着一个变量增加,另一个趋于减少(例如花在社交媒体的时间与注意力集中程度)。如果看不出任何规律,我们就说没有相关性。相关关系并不意味着因果关系——两个事物一起动,并不能证明一个导致了另一个。
The line of best fit can be used to make estimates. Interpolation is estimating a value within the range of the original data; it is generally reliable. Extrapolation is estimating a value outside the range of the data, which is much riskier because you assume the trend continues unchanged. A clear warning: extrapolation = extra risk, outside the known range.
最佳拟合线可以用来做估算。内插法是在原始数据范围内估算一个值;通常比较可靠。外推法则是在数据范围之外进行估算,这风险要大得多,因为你假定了趋势毫无变化地延续。一条明确的警告:外推 extrapolation = 额外的风险 extra,超出了已知范围。
You may meet Spearman’s rank correlation coefficient later, but for Year 10, understanding visual interpretation and describing correlation as ‘strong positive’, ‘weak negative’ etc. is essential. Always use context: ‘The more hours a car engine has run, the lower its resale price, showing a negative correlation.’
以后你可能会接触到斯皮尔曼等级相关系数,但对十年级来说,理解目视解读并且用“强正相关”“弱负相关”等词语来描述相关关系就非常重要。一定要结合情境:“汽车发动机运转的小时数越多,其转售价格就越低,显示出负相关。”
10. Hypothesis Testing Basics | 假设检验基础
A hypothesis is a statement that you can test with data. In GCSE Statistics, you often start with a question like ‘Is there a link between hours of sleep and reaction time?’ and rephrase it as a prediction: ‘As hours of sleep increase, reaction time decreases.’ A proper hypothesis includes what you plan to measure, the variables and a direction of change. The null hypothesis is the default position that there is no relationship or no difference; your investigation then tries to find evidence against this null.
假设是一个你可以用数据来检验的说法。在 GCSE 统计中,通常会从一个诸如“睡眠时间和反应时间之间是否存在联系?”的问题入手,并将其改写为一个预测:“随着睡眠时间增加,反应时间会减少。”一个恰当假设会包含你计划测量什么、涉及的变量以及变化的方向。零假设是默认立场,即没有关系或没有差异;你的调查研究再试图找到否定零假设的证据。
After collecting and analysing data, you write a conclusion. Always refer back to your original hypothesis: do your findings support it or not? Use cautious language such as ‘The data suggests…’ or ‘There is some evidence for…’ because a sample can never give absolute proof. A useful checklist is H-I-C: Hypothesis, Investigation, Conclusion.
在收集并分析数据之后,你要写出结论。一定要回溯到最初的假设:你的调查结果支持它吗?使用谨慎的语言,如“数据表明……”或“有一些证据支持……”,因为样本永远无法给出绝对的证明。一个有用的快捷清单是 H-I-C:假设、调查、结论。
11. Common Vocabulary Pitfalls and Memory Helpers | 常犯词汇误区与记忆帮手
Many students confuse ‘range’ with ‘interquartile range’. Remember: range takes the extremes, IQR takes the trunk. The word ‘quartile’ itself helps – ‘quart’ means four, so quartiles split data into four equal parts. ‘Discrete’ data often comes from counting items; think of ‘discrete = distinct units, like stepping stones’. ‘Continuous’ data comes from measuring; think of ‘continuous = continuum, a flowing line’.
很多学生混淆“全距”和“四分位距”。记住:全距取的是两极,IQR 取的是躯干。“四分位数”(quartile)这个词本身就有提示——“quart”意味着四,所以四分位数将数据等分为四份。“离散”数据常源自计数;联想“离散就是一个个分立的单元,像踏脚石”。“连续”数据源自测量;联想“连续就是一条连续的线,流淌不断”。
For ‘biased sample’, picture a choir being surveyed only from the bass section; it would not represent all voices. Bias means the sample systematically over- or under-represents a group. The mnemonic Bias = Built-in Slant might help. When you see ‘representative’, it indicates the sample mirrors the population in key characteristics.
对于“有偏样本”,想象只从合唱团的低音声部进行调查;它就无法代表所有声部。偏倚意味着样本系统性地过度或过少代表了某个群体。助记口诀 Bias = 内置的倾斜 (Built-in Slant) 也许有帮助。看到“代表性”这个词,它意味着样本在关键特征上反映了总体。
Finally, always read the word ‘estimate’ carefully. You estimate a mean from grouped data, you estimate a median from a cumulative frequency graph. The word acknowledges that the true value is unknown and the calculated value is an approximation – never call it the ‘exact mean’ unless you have the raw data.
最后,始终仔细辨析“估算”这个词。从分组数据中估算平均数,从累积频数图中估算中位数。“估算”一词承认真实值未知,计算值是个近似值——除非你掌握原始数据,否则绝不要称其为“精确平均数”。
12. Quick Terminology Self-Check | 词汇速查自查
Use this rapid-fire checklist to test your recall before an assessment. Cover the definitions and try to say them aloud. Qualitative, discrete, continuous, primary, secondary, census, sample, mean, median, mode, modal class, range, IQR, standard deviation, random sample, stratified sample, bias, representative, bar chart, histogram, frequency polygon, scatter graph, line of best fit, interpolation, extrapolation, positive correlation, negative correlation, hypothesis, null hypothesis, conclusion. If you can define each one in a single sentence and give an example, you are exam-ready.
用这份快速自检清单在评估前测试你的记忆。盖住定义,试着大声说出来。定性、离散、连续、一手、二手、普查、样本、平均数、中位数、众数、众数组、全距、IQR、标准差、随机样本、分层样本、偏倚、代表性、条形图、直方图、频数多边形、散点图、最佳拟合线、内插、外推、正相关、负相关、假设、零假设、结论。如果你能用一个句子定义每一个术语并举出例子,你就已经为考试做好了准备。
Pair this list with a set of flashcards: write the term on one side and the one-sentence definition plus an example on the other. Shuffle and practise daily for five minutes. Repetition builds recognition, and soon these statistical words will feel as familiar as your own name.
将这个清单和一套闪卡配合使用:一面写上术语,另一面写上单句定义加一个例子。每天打乱卡片练习五分钟。重复构建识别,很快这些统计词汇就会像你自己的名字一样熟悉。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导