Types of Data | 数据的类型

📚 Types of Data | 数据的类型

In A-Level Statistics, every investigation begins with data. Data are the raw facts and figures collected from observations, measurements, or responses. Understanding the type of data you are working with is fundamental because it determines which statistical diagrams, calculations, and hypothesis tests are appropriate. Getting the classification wrong can lead to flawed conclusions and lost marks in exams. This article will guide you through the essential data types covered in the Edexcel A-Level Mathematics specification, using clear definitions, examples, and paired English-Chinese explanations.

在A-Level统计中,每一次调查研究都从数据开始。数据是通过观察、测量或问卷收集到的原始信息和数字。弄清楚你正在处理的是哪种类型的数据至关重要,因为它决定了应使用哪种统计图表、计算方法和假设检验。一旦分类出错,就可能导致错误结论,并在考试中丢分。本文将根据 Edexcel A-Level 数学考纲,带你系统梳理各类数据,提供清晰的定义、示例和中英文对照讲解。


1. Introduction to Data Types | 数据类型简介

Data can be classified in several ways depending on their nature, how they were collected, and what they represent. The most fundamental split is between qualitative and quantitative data. From there, quantitative data branch into discrete and continuous types. We also consider whether data are primary or secondary, grouped or ungrouped, and whether they carry intrinsic order (ordinal). Edexcel exam questions frequently ask students to identify the data type in a given scenario, so a thorough understanding is key.

数据可以根据其性质、收集方式和所代表的内容进行多种分类。最基本的划分是定性数据和定量数据。定量数据又进一步分为离散型和连续型。我们还要考虑数据是一手数据还是二手数据、是已分组数据还是未分组数据,以及它们是否具有内在顺序(有序数据)。Edexcel 考试题目经常要求考生判断给定情景下的数据类型,因此透彻理解是拿分关键。


2. Qualitative and Quantitative Data | 定性与定量数据

Qualitative data describe qualities or characteristics that cannot be measured with numbers in a meaningful arithmetic way. They are often called categorical data. Examples include eye colour (blue, brown, green), types of vehicle (car, van, lorry), or a person’s favourite subject. Although we might code these categories with numbers (e.g. 1 for male, 2 for female), the numbers do not have numerical meaning — you cannot calculate an average eye colour meaningfully.

定性数据描述的是无法用有意义的算术方法测量的品质或特征,通常被称为分类数据。例如:眼睛颜色(蓝、棕、绿)、车辆类型(小轿车、厢式货车、卡车),或一个人最喜欢的科目。尽管我们可以用数字对这些类别进行编码(如1代表男性、2代表女性),但这些数字没有数值意义——你无法有意义地计算眼睛颜色的“平均值”。

Quantitative data consist of numerical values on which arithmetic operations make sense. They arise from counting or measuring. Examples include the height of students, the number of pets owned, or the temperature of a liquid. Quantitative data allow us to calculate summary statistics such as the mean, median, and standard deviation, and to draw box plots, histograms, and scatter diagrams.

定量数据由可以进行有意义的算术计算的数值组成。它们来自计数或测量。例子包括学生的身高、拥有的宠物数量或液体的温度。定量数据允许我们计算平均数、中位数和标准差等汇总统计量,并可以绘制箱形图、直方图和散点图等。


3. Discrete and Continuous Data | 离散与连续数据

Discrete data are quantitative data that can only take specific, separate values. Typically, they arise from counting and there are clear gaps between one value and the next. The number of students in a class (29, 30, 31) and the score on a dice (1, 2, 3, 4, 5, 6) are discrete. In Edexcel exams, you might be asked why a variable such as ‘the number of goals scored in a match’ is discrete: because it can only be whole numbers — you cannot score 2.3 goals.

离散数据是只能取特定、可分离数值的定量数据。它们通常来自计数,数值之间存在明确的间隔。班级里的学生人数(29, 30, 31)和掷骰子的点数(1, 2, 3, 4, 5, 6)都是离散的。在 Edexcel 考试中,你可能会被问到为什么“一场比赛的进球数”是离散变量:因为它只能是整数——你不可能进 2.3 个球。

Continuous data can take any value within a given range. These are obtained by measuring rather than counting, and there are no gaps between possible values — any value on a continuous scale is theoretically possible. Examples include the mass of a chocolate bar (which could be 49.5 g, 49.52 g, etc.), the time taken to run 100 m, or the length of a leaf. When handling continuous data, we often need to consider class intervals and the concept of rounding.

连续数据可以在给定范围内取任意值。它们来自测量而非计数,可能的取值之间没有间隔——在连续尺度上理论上任何一个值都可能出现。例子包括巧克力棒的质量(可能是 49.5 g、49.52 g 等)、跑 100 米所用的时间,或一片叶子的长度。在处理连续数据时,我们经常需要考虑组距和舍入的概念。


4. Primary and Secondary Data | 一手与二手数据

Primary data are data collected directly by the researcher for the specific purpose of the investigation. Examples include conducting your own experiment to measure the reaction times of classmates, or sending out a questionnaire designed by you. The advantage is that you have control over the data collection process — you can ensure the sample is suitable and the measurements are accurate. The disadvantage is that it can be time-consuming and expensive.

一手数据是研究人员为了特定的调查目的而亲自收集的数据。例如,自己做实验测量同学们的反应时间,或者发放由你设计的问卷。其优势在于你能控制数据收集过程——可以确保样本合适、测量准确。缺点则是可能耗时且成本较高。

Secondary data are data that have already been collected by someone else for a different purpose, but can be used in your investigation. Examples include government census data, results from published scientific studies, or stock market records available online. Secondary data are often cheaper and quicker to obtain, but you must be cautious about their reliability, accuracy, and whether they exactly fit the context of your own research.

二手数据是由他人出于其他目的收集的、但可以用于你调查的数据。例子包括政府人口普查数据、已发表科学研究的结果,或在网上可以获取的股市记录。二手数据通常获取成本更低、速度更快,但你必须谨慎对待其可靠性、准确性,以及它们是否完全符合你自己的研究背景。


5. Categorical and Numerical Data | 分类数据与数值数据

Within qualitative data, we often distinguish between nominal and ordinal categorical data. Nominal data are categories without any natural order. For instance, types of fruit (apple, banana, cherry) or blood groups (A, B, AB, O) are nominal — there is no meaningful way to rank them. Numerical codes may be assigned, but any calculations with these codes are invalid.

在定性数据中,我们常常区分名义分类数据和有序分类数据。名义数据是没有自然顺序的类别。例如,水果种类(苹果、香蕉、樱桃)或血型(A、B、AB、O)就是名义数据——没有有意义的方式给它们排序。可以为它们指定数字代码,但用这些代码进行任何计算都是无效的。

Numerical data is another term for quantitative data, where the values represent a measurable quantity. However, not all numbers are numerical in the statistical sense. A telephone number is a string of digits but is categorical (nominal) because you cannot perform arithmetic on it. Edexcel examiners love to test this subtle point.

数值数据是定量数据的另一种说法,其数值代表一个可测量的量。但并非所有数字在统计学意义上都是数值数据。电话号码是一串数字,但它属于分类数据(名义),因为你不能对其进行算术运算。Edexcel 的考官很喜欢考查这个细微差别。


6. Ordinal Data | 有序数据

Ordinal data are categorical data where the categories have a clear, meaningful order or ranking. The intervals between ranks are not necessarily equal. Common examples include satisfaction ratings (very dissatisfied, dissatisfied, neutral, satisfied, very satisfied), grades (A*, A, B, C, D, E), or the size of a drink (small, medium, large). While we can say ‘large’ is greater than ‘medium’, we cannot quantify exactly how much greater.

有序数据是具有明确、有意义顺序或排名的分类数据。等级之间的间隔并不一定相等。常见的例子包括满意度评分(非常不满意、不满意、一般、满意、非常满意)、成绩等级(A*、A、B、C、D、E),或饮料的尺寸(小、中、大)。虽然我们可以说“大”大于“中”,但我们无法精确量化大多少。

In A-Level statistics, ordinal data are often treated with non-parametric techniques. For example, Spearman’s rank correlation coefficient is used to measure the strength of association between two sets of ordinal data. Students should recognise that while ordinal data can be coded with numbers, the underlying variable is not necessarily quantitative.

在A-Level统计中,有序数据通常使用非参数方法处理。例如,斯皮尔曼等级相关系数用于测量两组有序数据之间的关联强度。学生应该认识到,虽然有序数据可以用数字编码,但潜在变量不一定是定量的。


7. Grouped and Ungrouped Data | 分组与未分组数据

Ungrouped data are presented as a list or a simple frequency table where each individual value is shown. For example, the exact marks of 10 students in a test: 56, 67, 74, 74, 80, 81, 85, 90, 92, 98. Ungrouped data allow precise calculations of the mean, median, and interquartile range without loss of information. They are typical when the number of distinct values is small.

未分组数据以列表或简单频数表的形式呈现,每一个单独的值都被列出来。例如,10 名学生的精确测试分数:56、67、74、74、80、81、85、90、92、98。未分组数据允许我们精确计算平均数、中位数和四分位数间距,不会丢失信息。当不同取值的数量较少时,通常采用未分组数据。

Grouped data arise when observations are organised into intervals or classes. This often happens with continuous data or when there is a large range of values. For instance, the heights of 100 students may be recorded in 5 cm intervals: 140 ≤ h < 145, 145 ≤ h < 150, and so on. When data are grouped, we lose the exact individual values and must use the midpoint of each interval as an estimate. This introduces a small degree of approximation in calculations such as the mean and standard deviation.

当观测值被整理放入区间或组中时,就形成了分组数据。这种情况常见于连续数据或取值跨度较大的情形。例如,100 名学生的身高可能按 5 厘米的间隔记录:140 ≤ h < 145,145 ≤ h < 150 等等。数据一旦分组,我们就失去了精确的个别数值,必须使用每个区间的中点作为估计值。这会在计算平均数、标准差等统计量时引入一定程度的近似。

The distinction between grouped and ungrouped data is critical for choosing the correct statistical diagram. Ungrouped discrete data can be shown on a bar chart or stem-and-leaf diagram, whereas grouped continuous data are represented by histograms with frequency density on the vertical axis — a key concept in the Edexcel specification.

区分分组与未分组数据对于选择正确的统计图表至关重要。未分组离散数据可以用条形图或茎叶图表示,而分组连续数据则用直方图表示,其垂直轴为频率密度——这是 Edexcel 考纲中的一个关键概念。


8. Binary Data | 二元数据

Binary data are a special type of categorical data where there are exactly two possible outcomes. Examples include yes/no responses, pass/fail results, or whether a component is defective or not. In statistics, binary data are often coded as 0 and 1, which makes them suitable for probability models such as the binomial distribution. The Edexcel A-Level syllabus makes extensive use of binary data when introducing hypothesis tests for the probability of success p in a binomial setting.

二元数据是一类特殊的分类数据,仅有两种可能的结果。例子包括是/否的回答、通过/未通过的结果,或零件是否故障。在统计学中,二元数据经常用 0 和 1 来编码,这使得它们适用于二项分布等概率模型。Edexcel A-Level 考纲在介绍二项分布中成功概率 p 的假设检验时,大量使用了二元数据。


9. Data Collection Methods and Their Impact | 数据收集方法及其影响

The type of data we obtain is closely linked to how we collect it. A survey using closed questions with a limited set of answers produces categorical (often ordinal or binary) data. An experiment measuring reaction times gives continuous numerical data. Observational studies might yield both qualitative and quantitative data. In an Edexcel exam, you might be asked to suggest a suitable data collection method and then describe the resulting data type.

我们获得的数据类型与收集数据的方式密切相关。使用选项有限的封闭式问卷会产生分类数据(通常是有序或二元数据)。测量反应时间的实验给出连续数值数据。观察研究则可能同时产生定性和定量数据。在 Edexcel 考试中,你可能会被要求建议一种合适的数据收集方法,然后描述由此产生的数据类型。

Sampling techniques also interact with data type. For instance, a stratified sample might be used to ensure proportional representation of categories (qualitative). When collecting primary quantitative data, it is essential to use calibrated instruments to avoid measurement error, which could affect the reliability of your findings.

抽样方法也与数据类型相互作用。例如,分层抽样可用于确保各类别(定性)的比例代表性。在收集一手定量数据时,必须使用校准过的仪器以避免测量误差,这会影响研究结果的可靠性。


10. Choosing the Right Statistical Diagram and Analysis | 选择正确的统计图表和分析方法

Visual representation depends on data type. Qualitative data are best displayed with pie charts or bar charts (for frequencies). Ordinal data can be shown with bar charts where order is preserved. Ungrouped discrete data can be displayed in a stem-and-leaf diagram or a box plot. For grouped continuous data, the histogram with frequency density is mandatory. Scatter diagrams are used to explore relationships between two quantitative variables.

图表的选择取决于数据类型。定性数据最好用饼图或条形图(用于展示频数)呈现。有序数据可以用保留顺序的条形图表示。未分组离散数据可以用茎叶图或箱形图显示。对于分组连续数据,使用频率密度直方图是必须的。散点图用于探索两个定量变量之间的关系。

Calculations also depend on data type. With ungrouped quantitative data, we can compute the exact mean using Σx / n. For grouped data, we estimate the mean using Σfx / Σf where x is the class midpoint. The modal class is used instead of the mode for grouped data. Measures of dispersion like standard deviation can be calculated exactly for discrete ungrouped data, but only estimated for grouped data unless raw values are given.

计算同样依赖于数据类型。对于未分组定量数据,我们可以使用 Σx / n 计算精确平均数。对于分组数据,我们使用 Σfx / Σf 来估计平均数,其中 x 是组中点。在分组数据中,我们使用众数类而不是众数本身。离散未分组数据的标准差等离散度量可以精确计算,但分组数据只能估算,除非给出了原始数值。


11. Common Misconceptions in Data Classification | 数据分类的常见误区

A frequent mistake is to assume that all numbers are quantitative. A postcode like ‘NW1 4SA’ contains numbers but is categorical. Similarly, a rating scale of 1 to 5 used for satisfaction is ordinal, not truly continuous — students should not treat it as continuous in hypothesis testing unless instructed otherwise.

一个常见错误是认为所有数字都是定量的。像 ‘NW1 4SA’ 这样的邮政编码包含数字,但它是分类数据。同样地,用于满意度的 1 到 5 评分量表是有序的,并非真正的连续数据——学生除非另有说明,不应在假设检验中将其当作连续数据处理。

Another pitfall is misidentifying discrete data as continuous. ‘Shoe size’ is a classic example — although shoe size can take half values like 6.5, the possible values are limited to specific increments, making it discrete. Age, when recorded as ‘age last birthday’, is discrete; but the exact age (including fractional years) is continuous. Always examine the context of how the data were recorded.

另一个陷阱是将离散数据误判为连续数据。“鞋码”就是一个典型例子——尽管鞋码可以取半值,如 6.5,但可能的取值仅限于特定的增量,因此它是离散数据。年龄如果记录为“上次生日年龄”,则是离散的;而精确年龄(含小数年)则是连续的。务必根据数据的记录方式审视上下文。


12. Summary and Exam Tips | 总结与考试技巧

To master data types for Edexcel A-Level Mathematics, remember the hierarchy: qualitative vs quantitative, then discrete vs continuous for quantitative, nominal vs ordinal for qualitative. Understand the differences between primary and secondary, and grouped vs ungrouped data. In the exam, always read the question carefully — look for clues such as ‘measured’, ‘counted’, ‘categories’, or ‘intervals’. When asked to justify your classification, use precise language: ‘the variable is continuous because it is measured on a scale with no gaps between values’.

要掌握 Edexcel A-Level 数学的数据类型,请记住分类层次:定性 vs 定量,定量再分为离散 vs 连续,定性再分为名义 vs 有序。理解一手与二手数据、分组与未分组数据的区别。考试中务必仔细审题——寻找诸如“测量”、“计数”、“类别”或“区间”等线索。当被要求说明分类理由时,使用严谨的语言:“该变量是连续的,因为它是在没有间隔的尺度上测量的”。

Finally, ensure you can link data type to the appropriate diagram and calculation. Create a quick reference table in your revision notes classifying variables from past papers under the headings: variable, type, suitable diagram, and statistical measure. This practice will build confidence and accuracy for the data-handling questions that appear regularly in the Edexcel Statistics component.

最后,确保你能将数据类型与合适的图表和计算方法联系起来。在复习笔记中制作一个快速对照表,把历年真题中的变量分类归纳在以下标题下:变量名、类型、合适图表和统计测度。这样的练习能帮助你建立信心、提高准确性,从容应对 Edexcel 统计部分中经常出现的数据处理题。

Published by TutorHao | Mathematics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading