AS OCR Statistics: Vocabulary Terminology Quick Guide | AS OCR 统计:词汇术语速记指南

📚 AS OCR Statistics: Vocabulary Terminology Quick Guide | AS OCR 统计:词汇术语速记指南

Welcome to the ultimate quick guide for AS OCR Statistics terminology. Statistics is a language in itself — mastering the precise vocabulary is half the battle. This guide breaks down essential terms from the OCR AS specification into clear, bilingual explanations. Use it to sharpen your interpretation, avoid common pitfalls, and communicate your reasoning with confidence in exams.

欢迎使用 AS OCR 统计学词汇术语速记指南。统计学本身就是一门语言,掌握精准的术语是成功的一半。本指南将 OCR AS 考纲中的核心词汇拆解为清晰的中英双语解释,帮助你快速理解题意、避开常见误区,并在考试中自信地表达推理过程。


1. Statistical Vocabulary: Population, Sample, Census | 统计词汇:总体、样本、普查

Population: The complete set of all individuals, items, or data points we wish to draw conclusions about. It is the whole group under investigation.

总体:我们想要得出结论的全部个体、物件或数据点的集合。它是研究中感兴趣的整个群体。

Sample: A subset of the population, carefully selected to represent the whole. We analyse samples to make inferences about the population without examining every member.

样本:从总体中选出的一部分,需具有代表性。通过分析样本,我们可以对总体进行推断,而不必检查每个成员。

Census: A data collection exercise that measures every single member of the population. A census gives the true value of a parameter but is expensive, time‑consuming, and usually impractical for large populations.

普查:对总体中每一个成员都进行测量的数据收集方式。普查能给出参数的真值,但成本高、耗时长,对于大型总体通常不切实际。

Sampling frame: A list, map, or database of all the members of the population from which a sample is drawn. A poor sampling frame leads to selection bias.

抽样框:包含总体所有成员的清单、地图或数据库,用于抽取样本。抽样框不完善会导致选择性偏差。

Sampling unit: Each individual element that can be selected in the sample. In a survey of students, each student is a sampling unit.

抽样单位:样本中可被选中的每一个单独元素。在一项学生调查中,每位学生即为一个抽样单位。


2. Data Types: Qualitative, Quantitative, Discrete, Continuous | 数据类型:定性、定量、离散、连续

Qualitative (categorical) data: Non‑numerical information that describes qualities or categories — for example, eye colour, type of vehicle, or favourite subject. It can be nominal (no order) or ordinal (natural order).

定性(分类)数据:描述性质或类别的非数值信息,例如眼睛颜色、车辆类型、最喜欢的科目。它可以是名义的(无序)或有顺序的(等级)。

Quantitative data: Numerical information that can be measured or counted. This type allows meaningful arithmetic. Quantitative data is split into discrete and continuous.

定量数据:可以测量或计数的数值信息,能进行有意义的算术运算。定量数据分为离散型和连续型。

Discrete data: Takes specific, separate whole‑number values, usually obtained by counting. Examples: number of goals scored, number of students in a class. There are clear gaps between values.

离散数据:取值是特定的、分离的整数值,通常通过计数获得。例如进球数、班级学生人数。值之间存在明确的间隔。

Continuous data: Can take any value within a range and is obtained by measuring. Examples: height, time taken, temperature. The data may be rounded to a certain degree of accuracy.

连续数据:在一个区间内可以取任何值,通过测量得到。例如身高、所用时间、温度。数据可能被四舍五入到一定精度。


3. Sampling Techniques | 抽样方法

Simple random sampling: Every member of the population has an equal chance of being chosen. Usually involves random number generators or lottery methods. It is free of investigator bias but requires a full sampling frame.

简单随机抽样:总体的每个成员被选中的概率相等。通常使用随机数生成器或抽签法。该方法无调查者偏差,但需要完整的抽样框。

Systematic sampling: Select every k‑th member from the sampling frame, starting from a random point. Quick to use and works well when the population is ordered in some way, but can introduce periodicity bias.

系统抽样:从抽样框中每隔第k个成员抽取,起点随机。使用快捷,当总体有某种顺序时效果较好,但可能引入周期性偏差。

Stratified sampling: Divide the population into distinct groups (strata) that share a characteristic, then take a random sample from each stratum proportional to its size. It guarantees representation of all subgroups.

分层抽样:将总体按某一特征划分为不同群组(层),然后在每层中按比例随机抽样。能确保所有子群体都被代表。

Quota sampling: Interviewers are told how many people with certain characteristics to interview, then they select anyone who fits those quotas until full. It does not require a sampling frame, but selection is subjective and can lead to bias.

定额抽样:调查员被告知需要采访多少具备某些特征的人,然后自行挑选符合条件的个体直到满额。不需要抽样框,但选择过程主观,可能产生偏差。

Opportunity (convenience) sampling: Simply use whoever is available at the time of the study. It is cheap and quick, but unlikely to be representative of the wider population.

便利抽样:直接选取在研究时容易接触到的人。成本低、速度快,但很可能无法代表更广泛的总体。


4. Measures of Central Tendency: Mean, Median, Mode | 集中趋势度量:均值、中位数、众数

Mean (x̄): The arithmetic average, calculated as Σx ÷ n for a sample. It uses all data values, making it sensitive to extreme outliers. The population mean is denoted by μ (mu).

均值(x̄):算术平均数,样本计算为 Σx ÷ n。它使用了所有数据值,因此对极端异常值敏感。总体均值用 μ 表示。

Median: The middle value when data is ordered. For n data points, the position is (n+1)/2. It is unaffected by outliers and is often preferred for skewed distributions.

中位数:将数据排序后位于中间的值。对于n个数据点,位置为 (n+1)/2。它不受异常值影响,常用于偏态分布。

Mode: The value that occurs most often, or the data class with the highest frequency. Some data sets have no mode, one mode (unimodal), or several modes (bimodal/multimodal).

众数:出现次数最多的值,或频率最高的数据类别。有些数据集没有众数,有的有一个(单峰)或多个(双峰/多峰)。


5. Measures of Spread: Range, IQR, Variance, Standard Deviation | 离散度量:极差、四分位距、方差、标准差

Range: The difference between the maximum and minimum values. It is quick to compute but heavily influenced by a single outlier.

极差:最大值与最小值之差。计算简便,但极易受单个异常值影响。

Interquartile range (IQR): IQR = Q₃ − Q₁, the range of the middle 50% of the data. Q₁ is the lower quartile (25th percentile) and Q₃ is the upper quartile (75th percentile). IQR is robust to outliers.

四分位距(IQR):IQR = Q₃ − Q₁,是中间 50% 数据的范围。Q₁ 为下四分位数(第25百分位数),Q₃ 为上四分位数(第75百分位数)。IQR 对异常值稳健。

Variance: The average of the squared deviations from the mean. Sample variance s² = Σ(x − x̄)²/(n−1); population variance σ² = Σ(x − μ)²/N. It measures overall spread in squared units.

方差:各数据与均值之差的平方的平均数。样本方差 s² = Σ(x − x̄)²/(n−1);总体方差 σ² = Σ(x − μ)²/N。它以平方单位衡量整体离散程度。

Standard deviation: The square root of the variance, s or σ. It returns dispersion to the original units and is the most widely used measure of spread.

标准差:方差的平方根,记作 s 或 σ。它使离散度量回到原始单位,是最常用的离散指标。

Outlier: A data point that lies an abnormal distance from the rest of the data. A common rule is any value 1.5 × IQR below Q₁ or above Q₃.

异常值:明显偏离其他数据点的数值。常用规则是小于 Q₁ − 1.5×IQR 或大于 Q₃ + 1.5×IQR 的值。


6. Probability Terminology and Laws | 概率术语与法则

Experiment: A repeatable process that produces one of several possible outcomes. Tossing a fair coin is an experiment.

试验:可重复的、会产生若干可能结果之一的过程。抛掷一枚公平硬币就是一个试验。

Sample space: The set of all possible outcomes of an experiment, usually denoted by S or Ω. For a die, S = {1,2,3,4,5,6}.

样本空间:试验所有可能结果的集合,通常用 S 或 Ω 表示。掷一枚骰子,S = {1,2,3,4,5,6}。

Event: A subset of the sample space. An event can be simple (one outcome) or compound (more than one outcome).

事件:样本空间的一个子集。事件可以是简单的(一个结果)或复合的(多个结果)。

Mutually exclusive events: Two events cannot happen at the same time. P(A ∩ B) = 0. The addition rule simplifies to P(A ∪ B) = P(A) + P(B).

互斥事件:两个事件不可能同时发生。P(A ∩ B) = 0。加法法则简化为 P(A ∪ B) = P(A) + P(B)。

Independent events: The occurrence of one event does not affect the probability of the other. P(A ∩ B) = P(A) × P(B). Do not confuse ‘mutually exclusive’ with ‘independent’.

独立事件:一个事件的发生不影响另一个事件发生的概率。P(A ∩ B) = P(A) × P(B)。注意不要混淆“互斥”与“独立”。

Conditional probability: The probability that event A occurs given that event B has already occurred, written as P(A | B). It is calculated by P(A ∩ B) / P(B).

条件概率:已知事件B已发生的条件下,事件A发生的概率,记作 P(A | B),计算公式为 P(A ∩ B) / P(B)。

P(A | B) = P(A ∩ B) / P(B)


7. The Binomial Distribution: Conditions, Notation & Expectation | 二项分布:条件、记法与期望

Discrete random variable: A variable whose value depends on the outcome of a random event and can take only separate, countable values. The sum of probabilities in its distribution equals 1.

离散随机变量:其取值取决于随机事件的结果,且只能取分离的可数值。其概率分布之和为1。

Binomial conditions: There are n independent trials, each trial has two possible outcomes (success and failure), the probability of success p remains constant, and we count the number of successes X.

二项分布条件:进行 n 次独立试验,每次试验只有两种可能结果(成功和失败),成功概率 p 保持不变,计数成功次数 X。

Notation: We write X ~ B(n, p). The random variable X is the number of successes in n trials. The probability of exactly r successes is given by the probability mass function.

记法:写作 X ~ B(n, p)。随机变量 X 表示 n 次试验中成功的次数。恰好 r 次成功的概率由概率质量函数给出。

P(X = r) = nCr pr (1 − p)n−r

Expectation and variance: For X ~ B(n, p), E(X) = np and Var(X) = np(1 − p). These summarise the centre and spread of the distribution without listing all probabilities.

期望与方差:对 X ~ B(n, p),E(X) = np,Var(X) = np(1 − p)。它们从整体上概括了分布的中心和离散程度,无需列出所有概率。


8. Hypothesis Testing with the Binomial Distribution | 基于二项分布的假设检验

Null hypothesis H₀: The default assumption about the population parameter, usually a statement of no effect or no difference, e.g. H₀: p = 0.5. It is the hypothesis we assume to be true initially.

原假设 H₀:关于总体参数的默认假设,通常表示无效或无差异,例如 H₀: p = 0.5。我们最初假定其为真。

Alternative hypothesis H₁: The statement we are testing for, which contradicts H₀. It can be one‑tailed (p < ... or p > …) or two‑tailed (p ≠ …).

备择假设 H₁:我们要检验的、与 H₀ 相矛盾的陈述。可以是单尾(p < ... 或 p > …)或双尾(p ≠ …)。

Significance level α: The probability of rejecting H₀ when it is actually true (Type I error). Common values are 0.05 (5%) or 0.01. The critical region is chosen so that P(in critical region | H₀ true) ≤ α.

显著性水平 α:当H₀实际为真时拒绝它的概率(第一类错误)。常用 0.05(5%)或 0.01。临界区域的选择应满足 P(落入临界区域 | H₀为真) ≤ α。

Test statistic: In binomial hypothesis testing, the test statistic is the observed number of successes, X. We calculate the probability of getting this observed value or a more extreme one under H₀.

检验统计量:在二项假设检验中,检验统计量是观测到的成功次数 X。我们计算在 H₀ 成立下得到该观测值或更极端情况的概率。

p‑value: The probability of obtaining results at least as extreme as the observed result, assuming H₀ is true. If p‑value ≤ α, we reject H₀; otherwise we do not reject H₀.

p 值:在原假设为真的条件下,得到与观测结果同样极端或更极端结果的概率。若 p 值 ≤ α,拒绝 H₀;否则不拒绝 H₀。

Critical region and critical value: The critical region is the set of values of X that lead to rejection of H₀. The boundary of this region is called the critical value. For a lower‑tail test H₁: p < ..., find the largest r such that P(X ≤ r) ≤ α.

临界区域与临界值:临界区域是导致拒绝 H₀ 的 X 取值集合,其边界称为临界值。对于下尾检验 H₁: p < ...,找出满足 P(X ≤ r) ≤ α 的最大 r。

Conclusion wording: Always phrase your conclusion in the context of the problem. ‘There is sufficient evidence to reject H₀…’ or ‘There is insufficient evidence to reject H₀…’ Never say ‘accept H₀’.

结论措辞:始终在问题背景下表述结论。“有足够证据拒绝 H₀……”或“证据不足以拒绝 H₀……”。绝不要说“接受 H₀”。


Published by TutorHao | Statistics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version