Types of Data | 数据类型

📚 Types of Data | 数据类型

In Edexcel A Level Mathematics, the topic “types of data” is a statistical foundation. Before choosing a diagram, calculating an average, or carrying out a hypothesis test, you must identify the nature of the variable. Misclassifying data can lead to an invalid method and lost marks in the examination.

在 Edexcel A Level 数学中,”数据类型” 是统计学的基础。在选择图表、计算平均数或进行假设检验之前,必须先确定变量的性质。错误地分类数据可能导致方法无效,并在考试中失分。


1. Why Data Types Matter | 为什么数据类型重要

Data classification controls every later statistical decision. A bar chart may be suitable for categorical data but is not appropriate for continuous data. Similarly, calculating a mean is meaningful for quantitative data but has no meaning for labels such as “red”, “blue”, and “green”.

数据分类控制着之后所有的统计决策。条形图可能适用于分类数据,但不适用于连续数据。同样,计算平均数对定量数据有意义,但对 “红”、”蓝”、”绿” 等标签没有意义。

Edexcel exam questions often ask you to state whether a variable is qualitative or quantitative, and if quantitative, whether it is discrete or continuous. This classification then helps justify the chosen summary statistic or diagram.

Edexcel 试题经常要求说明变量是定性还是定量,如果是定量,还要说明是离散还是连续。这一分类随后有助于证明所选汇总统计量或图表的合理性。


2. Qualitative and Quantitative Data | 定性数据与定量数据

Qualitative data describe a quality or characteristic that cannot be measured numerically. They are also called categorical data. Examples include eye colour, gender, type of vehicle, and postcode area.

定性数据描述无法用数值测量的性质或特征。它们也称为分类数据。例子包括眼睛颜色、性别、车辆类型和邮编区域。

Quantitative data are numerical and arise from counting or measuring. Examples include height, time, temperature, marks in a test, and the number of customers entering a shop.

定量数据是数值型数据,来源于计数或测量。例子包括身高、时间、温度、考试分数和进入商店的顾客人数。

A quick test is to ask whether finding an average makes sense. If the values can be ordered and an average is meaningful, the data are usually quantitative. If the values are just labels, the data are qualitative.

一个快速判断方法是问求平均值是否有意义。如果数值可以排序,并且求平均值有意义,则数据通常是定量数据;如果数值只是标签,则数据是定性数据。


3. Discrete and Continuous Quantitative Data | 离散型与连续型定量数据

Quantitative data are further divided into discrete and continuous. Discrete data can only take exact, separate values, usually obtained by counting. Examples include the number of sisters, the number of goals scored, and the number of cars in a car park.

定量数据进一步分为离散型和连续型。离散数据只能取确切、分离的值,通常通过计数获得。例子包括姐妹数量、进球数和停车场汽车数量。

Continuous data can take any value within a given interval, usually obtained by measuring. Examples include height, weight, time, temperature, and length.

连续数据可以在给定区间内取任意值,通常通过测量获得。例子包括身高、体重、时间、温度和长度。

Note that age is continuous even though it is often recorded in whole years. Rounding a continuous variable to whole numbers does not change its underlying type. Variables such as shoe size may look continuous but are usually treated as discrete because they take only a limited set of exact values.

注意年龄是连续数据,尽管通常按整岁记录。将连续变量取整为整数不会改变其本质类型。鞋码等变量看起来可能是连续的,但通常被视为离散数据,因为它们只取有限的一组确切值。

Data type | 数据类型 Definition | 定义 Examples | 示例
Discrete quantitative | 离散定量 Counted, exact separate values | 计数得到,确切分隔的值 Number of books, goals | 书本数量、进球数
Continuous quantitative | 连续定量 Measured, any value in an interval | 测量得到,区间内任意值 Height, time, temperature | 身高、时间、温度

4. Categorical Nominal and Ordinal Data | 名义分类与有序分类数据

Categorical data can also be divided into nominal and ordinal. Nominal data have categories with no natural order, such as favourite sport, blood type, or country of birth. The order of the categories can be changed without losing information.

分类数据还可进一步分为名义数据和有序数据。名义数据的类别没有自然顺序,例如最喜欢的运动、血型或出生国家。改变类别的顺序不会丢失信息。

Ordinal data have categories with a clear order, but the differences between successive categories may not be equal. Examples include satisfaction ratings such as “very dissatisfied”, “dissatisfied”, “neutral”, “satisfied”, and “very satisfied”, or grades A*, A, B, C, D, E.

有序数据具有明确顺序的类别,但相邻类别之间的差异可能不相等。例子包括满意度评分,如 “非常不满意”、”不满意”、”一般”、”满意”、”非常满意”,或成绩 A*、A、B、C、D、E。

In A Level questions, ordinal data are often treated as categorical even though the categories can be ordered. Do not calculate a mean from ordinal categories unless the question explicitly treats the scale as numerical by assigning scores.

在 A Level 题目中,有序数据通常被视为分类数据,尽管类别可以排序。除非题目通过赋予分数明确将等级视为数值,否则不要从有序类别计算平均值。


5. Primary and Secondary Data | 一手数据与二手数据

Primary data are collected by the person or team who will use them, through surveys, experiments, or observations. The collector has control over accuracy, relevance, and collection method.

一手数据由将要使用数据的人或团队通过调查、实验或观察收集。收集者可以控制数据的准确性、相关性和收集方法。

Secondary data are collected by someone else for a different purpose and then reused. Examples include government statistics, large data sets published by a weather agency, or past research results.

二手数据由他人为不同的目的收集,然后被再次使用。例子包括政府统计数据、气象机构发布的大样本数据集或过去的研究结果。

Primary data can be tailored to the specific question and may be more reliable, but they are often costly and time-consuming to collect. Secondary data are cheaper and quicker to obtain, but they may be less precise, incomplete, or out of date.

一手数据可以针对特定问题量身定制,而且可能更可靠,但收集起来通常成本高、耗时长。二手数据获取起来更便宜、更快捷,但可能不够精确、不完整或已经过时。


6. Raw and Grouped Data | 原始数据与分组数据

Raw data are the original unorganised observations as collected, such as a list of 20 test scores. Grouped data are organised into intervals or classes, for example 0 ≤ x < 10, 10 ≤ x < 20, and so on.

原始数据是收集到的未经整理的最初观测值,例如 20 个考试分数的列表。分组数据被整理成区间或组,例如 0 ≤ x < 10,10 ≤ x < 20 等。

Grouping loses the exact values, so any calculation from grouped data is an estimate. The class midpoint is used to represent each class when estimating the mean or standard deviation.

分组会丢失精确数值,因此由分组数据得到的任何计算都是估计值。在估计平均数或标准差时,使用组中点来表示每组。

class midpoint = (lower class boundary + upper class boundary) ÷ 2

For the interval 10 ≤ x < 20, the lower boundary is 10 and the upper boundary is 20, so the midpoint is (10 + 20) ÷ 2 = 15. This value is used as the representative value for every observation in that class.

对于区间 10 ≤ x < 20,下边界是 10,上边界是 20,因此组中点是 (10 + 20) ÷ 2 = 15。该值用作该组内每个观测值的代表值。


7. Measurement Scales: Nominal, Ordinal, Interval, Ratio | 测量尺度:名义、顺序、间隔、比率

A more formal way to classify data uses four measurement scales. Nominal and ordinal scales apply to qualitative data, while interval and ratio scales apply to quantitative data.

一种更正式的数据分类方法使用四种测量尺度。名义尺度和顺序尺度适用于定性数据,而间隔尺度和比率尺度适用于定量数据。

Interval data have equal gaps between values but no true zero. Temperature in degrees Celsius is a typical example: 20 °C is not twice as hot as 10 °C because 0 °C does not mean “no temperature”.

间隔数据的数值之间间隔相等,但没有真正的零。摄氏度温度就是一个典型例子:20 °C 并不是 10 °C 的两倍热,因为 0

Published by TutorHao | A-Level Mathematics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading