📚 IGCSE CAIE Statistics: Core Vocabulary & Terminology Quick Guide | IGCSE CAIE 统计:核心词汇速记指南
Welcome to your essential vocabulary cheat sheet for IGCSE CAIE Statistics. Understanding the precise meaning of statistical terms is critical for interpreting questions, using correct notation, and avoiding common mistakes. This guide will walk you through the key terminology section by section, with clear definitions in English and Chinese to help you remember faster.
欢迎查阅 IGCSE CAIE 统计必备词汇速记表。准确理解统计术语对于解读题目、使用正确符号和避免常见错误至关重要。本指南将按小节带你梳理关键术语,提供清晰的英文和中文释义,帮助你快速记忆。
1. Types of Data | 数据类型
Discrete data: Data that can only take specific, separate values, usually obtained by counting. Examples: number of students in a class, shoe sizes.
离散数据:只能取特定、分离数值的数据,通常通过计数获得。例如班级人数、鞋码。
Continuous data: Data that can take any value within a given range, usually obtained by measuring. Examples: height, time, temperature.
连续数据:可在一定范围内取任何数值的数据,通常通过测量获得。例如身高、时间、温度。
Qualitative data: Non‑numerical data that describe categories or attributes. Also called categorical data. Examples: colours, types of transport.
定性数据:描述类别或属性的非数值数据,也称分类数据。例如颜色、交通方式。
Quantitative data: Numerical data that can be measured or counted. It can be either discrete or continuous.
定量数据:可测量或计数的数值数据,可以是离散或连续数据。
Primary data: Data collected directly by the researcher for a specific purpose. It tends to be more reliable but time‑consuming to gather.
原始数据:研究者为特定目的直接收集的数据,通常更可靠但收集耗时。
Secondary data: Data obtained from existing sources such as published reports or databases. It is quicker to access but may be less suited to the exact research question.
二手数据:从已有来源(如公开报告或数据库)获取的数据。获取快速但可能不完全契合研究问题。
Variable: Any characteristic that can vary from one member of a population to another. In statistics, variables are often denoted by letters like x or y.
变量:总体中不同个体之间可变化的任何特征,统计中常用 x 或 y 等字母表示。
2. Sampling and Populations | 抽样与总体
Population: The entire set of individuals, items, or data that you wish to draw conclusions about.
总体:你希望得出结论的全部个体、项目或数据的集合。
Sample: A subset of the population selected for investigation. A well‑chosen sample should represent the population.
样本:从总体中选出用于调查的子集,精心选取的样本应能代表总体。
Census: A survey or investigation that includes every member of the population. It gives precise results but is often expensive and time‑consuming.
普查:涵盖总体中每一个个体的调查,结果精确但通常昂贵且耗时。
Random sample: A sample where each member of the population has an equal chance of being selected. Often produced using a random number generator or drawing lots.
随机样本:总体中每个成员被选中的机会均等的样本,通常通过随机数生成器或抽签获得。
Stratified sample: A sample in which the population is divided into distinct groups (strata) and a random sample is taken from each group in proportion to its size.
分层样本:将总体分成不同组别(层),并按各组所占总体的比例从中抽取随机样本。
Bias: A systematic error that causes results to be consistently distorted in one direction. Sampling bias can occur when the sample does not adequately represent the population.
偏差:导致结果单方向持续偏离的系统性误差。当样本未能充分代表总体时会出现抽样偏差。
Representative sample: A sample that accurately mirrors the characteristics of the population, allowing valid conclusions to be drawn.
代表性样本:准确反映总体特征的样本,据此得出的结论才有效。
3. Measures of Central Tendency | 集中趋势度量
Mean: The arithmetic average, calculated by summing all values and dividing by the number of observations. It is sensitive to extreme values (outliers).
平均数:算术平均值,将所有数值加总后除以观测数量。易受极端值(离群值)影响。
Mean (x̄) = ∑x / n
Median: The middle value when the data are arranged in order. For an even number of observations, it is the average of the two central values. It is not affected by outliers.
中位数:数据排序后位于中间的值。观测数量为偶数时,取中间两个值的平均数。不受离群值影响。
Mode: The value that occurs most frequently in a data set. A set may have one mode, more than one mode (bimodal or multimodal), or no mode at all.
众数:数据集中出现频率最高的值。一组数据可以有一个众数、多个众数(双峰或多峰)或没有众数。
Modal class: For grouped data, the class interval with the highest frequency.
众数类:在分组数据中,频率最高的那个组距。
4. Measures of Spread | 离散程度度量
Range: The difference between the largest and smallest values in a data set. It gives a basic measure of spread but is heavily influenced by outliers.
极差:数据中最大值与最小值之差,提供基本的离散程度度量,但极易受离群值影响。
Quartiles: Values that divide an ordered data set into four equal parts. Q₁ (lower quartile) is the median of the lower half, Q₂ is the median, and Q₃ (upper quartile) is the median of the upper half.
四分位数:将有序数据分为四等份的值。Q₁(下四分位数)是下半部分的中位数,Q₂即中位数,Q₃(上四分位数)是上半部分的中位数。
Interquartile range (IQR): The range of the middle 50% of the data, calculated as IQR = Q₃ – Q₁. It is a robust measure of spread resistant to outliers.
四分位距:中间50%数据的范围,计算公式为 IQR = Q₃ – Q₁,是能够抵抗离群值影响的稳健离散度量。
IQR = Q₃ – Q₁
Standard deviation: A measure of the average distance of data values from the mean. A small standard deviation indicates data are clustered near the mean; a large one shows wide spread. The sample standard deviation is denoted by s.
标准差:衡量数据值偏离平均值的平均距离。标准差小说明数据聚集在均值附近,大则说明离散度大。样本标准差记为 s。
s = √[∑(x – x̄)² / (n – 1)]
Variance: The square of the standard deviation. It represents the average squared deviation from the mean.
方差:标准差的平方,表示数据偏离平均值的平方的平均值。
Variance = s² = ∑(x – x̄)² / (n – 1)
Outlier: An observation that lies an abnormal distance from other values. Typically defined as a value smaller than Q₁ – 1.5×IQR or larger than Q₃ + 1.5×IQR.
离群值:与其他数值距离异常的观测值,通常定义为小于 Q₁ – 1.5×IQR 或大于 Q₃ + 1.5×IQR 的值。
5. Frequency Distributions | 频率分布
Frequency: The number of times a particular value or category occurs in a data set.
频数:数据集中某个特定值或类别出现的次数。
Grouped frequency distribution: A table that summarises data by grouping values into class intervals, showing the frequency for each interval. Useful for continuous data or large sets of discrete data.
分组频率分布:将数据值按组距分组并显示每组频数的表格,适用于连续数据或大量离散数据。
Class interval: A range of values into which data are grouped, e.g., 10 ≤ x < 20.
组距:数据分组所依据的数值范围,例如 10 ≤ x < 20。
Class width: The size of a class interval, found by subtracting the lower boundary from the upper boundary.
组宽:组距的大小,由上界减下界得出。
Class boundaries: The precise values that separate one class from another, used when plotting histograms. For the interval 10–19, boundaries are 9.5 and 19.5.
组界:区分相邻组的精确值,用于绘制直方图。如组距 10–19,其组界为 9.5 和 19.5。
Frequency density: An adjusted frequency used in histograms to ensure area corresponds to frequency. It is calculated as frequency ÷ class width.
频率密度:用于直方图的调整后频率,以保证面积与频数对应,计算公式为频数 ÷ 组宽。
Frequency density = Frequency / Class width
Cumulative frequency: The running total of frequencies as you move through the classes in order. It helps determine the number of observations below a certain value.
累积频率:按顺序通过各组时频数的累计总和,用于确定低于某一值的观测数量。
6. Graphical Representations | 图表表征
Bar chart: A diagram that uses rectangular bars of equal width with heights proportional to frequency. Bars are separated by gaps to show distinct categories. Used for qualitative or discrete data.
条形图:用等宽矩形条、高度与频数成比例的图表。条形之间有间隙,用于表示不同类别,适用于定性或离散数据。
Histogram: A graphical display for continuous or grouped data where the area of each bar is proportional to frequency. There are no gaps between bars, and vertical axis often shows frequency density.
直方图:用于连续或分组数据的图形,每块矩形的面积与频数成比例。矩形之间没有间隙,纵轴通常为频率密度。
Pie chart: A circular chart divided into sectors, each representing a proportion of the whole. Sector angle = (frequency ÷ total frequency) × 360°.
饼图:将圆形划分为扇形的图表,每个扇形代表整体的一个部分。扇形角度 =(频数 ÷ 总频数)× 360°。
Pictogram: A chart that uses pictures or symbols to represent frequencies. Each symbol stands for a certain number of items. A key must be provided.
象形图:用图片或符号表示频数的图表,每个符号代表一定数量的项目,必须配有图例。
Stem‑and‑leaf diagram: A method of displaying raw data while preserving individual values. The ‘stem’ represents the leading digit(s) and the ‘leaf’ the trailing digit. A key is essential.
茎叶图:在展示数据分布的同时保留原始数值的图形。 ‘茎’ 表示前导数字, ‘叶’ 表示尾随数字,必须提供图例。
Dot plot: A simple chart where each data value is represented by a dot above a number line. Very useful for small data sets to see the shape of the distribution.
点图:每个数据值用数轴上方的点表示的简单图表,特别适用于小数据集以观察分布形态。
7. Cumulative Frequency and Box Plots | 累积频率与箱线图
Cumulative frequency curve (ogive): A smooth curve plotted from cumulative frequency against the upper class boundaries. It allows estimation of medians, quartiles, and percentiles.
累积频率曲线:以累积频率对组上界描点连成的光滑曲线,可用于估计中位数、四分位数和百分位数。
Median from cumulative frequency curve: The value on the horizontal axis corresponding to half the total frequency on the cumulative frequency curve.
从累积频率曲线求中位数:累积频率曲线上对应总频数一半处所对应的的水平轴数值。
Percentile: A value below which a given percentage of observations fall. The median is the 50th percentile, Q₁ the 25th, and Q₃ the 75th.
百分位数:某一百分比观测值低于该值。中位数是第50百分位数,Q₁ 是第25百分位数,Q₃ 是第75百分位数。
Box‑and‑whisker plot (box plot): A diagram that displays the five‑number summary: minimum, Q₁, median, Q₃, and maximum. Possible outliers are often shown as separate points.
箱线图:展示五数概括(最小值、Q₁、中位数、Q₃、最大值)的图形。可能的离群值通常单独标出。
Interpreting the box plot: The box shows the IQR; the line inside the box marks the median. The whiskers extend to the smallest and largest values that are not outliers. It provides a visual comparison of distributions.
解读箱线图:矩形箱表示 IQR,箱内横线为中位数,触须延伸至非离群的最小值与最大值。箱线图可直观比较多个分布。
8. Scatter Diagrams, Correlation and Line of Best Fit | 散点图、相关与最佳拟合线
Scatter diagram: A graph plotting bivariate data as points on a coordinate grid. Each axis represents one variable. It reveals whether a relationship exists between the two variables.
散点图:将双变量数据以点形式标在坐标网格上的图形,每轴代表一个变量,用于探查两变量之间是否存在关系。
Correlation: A measure of the strength and direction of a linear relationship between two variables. It can be positive, negative, or zero. Correlation does not imply causation.
相关:衡量两变量间线性关系强弱的指标,可为正相关、负相关或零相关。相关不代表因果关系。
Positive correlation: As one variable increases, the other tends to increase. The points slope upwards.
正相关:一个变量增大时,另一个也倾向于增大,散点向上倾斜。
Negative correlation: As one variable increases, the other tends to
Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply