📚 Data Types and Data Structures in Sport | 体育中的数据类型与数据结构
In modern sport, data drives decisions about training load, technique correction, tactical planning and injury prevention. Understanding the different data types and the structures used to store them is essential for any sport scientist, coach or analyst. This article explains the key concepts of data types and data structures using real examples from physical education and sport performance analysis.
在现代体育中,数据驱动着训练负荷、技术纠正、战术规划和伤病预防等方面的决策。理解不同的数据类型以及用于存储它们的数据结构,对任何运动科学家、教练或分析师来说都至关重要。本文利用体育教育和运动表现分析中的真实案例,解释数据类型和数据结构的关键概念。
1. Why Data Matters in Sport | 为什么数据在体育中重要
Data in sport can be as simple as a stopwatch time or as complex as a three-dimensional motion capture file. The way we classify and organise data affects how accurately we can analyse performance and make recommendations. A clear understanding of data types and data structures allows analysts to choose the correct statistical test, design better experiments and avoid misleading conclusions.
体育中的数据可以简单到秒表计时,也可以复杂到三维动作捕捉文件。我们对数据进行分类和组织的方式,会影响我们分析表现和提出建议的准确程度。清楚地理解数据类型和数据结构,能让分析师选择正确的统计检验方法、设计更好的实验并避免误导性结论。
2. Quantitative and Qualitative Data | 定量数据与定性数据
Quantitative data are numerical measurements that can be counted or measured. Examples in sport include a 100 m sprint time of 11.2 s, a heart rate of 175 beats per minute, or the number of passes completed in a football match. Quantitative data allow for arithmetic operations and statistical analysis.
定量数据是可以计数或测量的数值。体育中的例子包括 100 米短跑时间 11.2 秒、每分钟 175 次的心率,或者足球比赛中完成的传球次数。定量数据可以进行算术运算和统计分析。
Qualitative data are non-numerical descriptions of qualities or characteristics. In sport, a coach might record that a gymnast’s landing was ‘unstable’ or that a tennis player’s serve had ‘good rotation’. Qualitative data are often converted into categories or rating scales so they can be analysed more systematically.
定性数据是描述性质或特征的非数值数据。在体育中,教练可能会记录体操运动员的落地 ‘不稳定’ 或网球运动员的发球 ‘旋转良好’。定性数据通常被转换为类别或评分量表,以便更系统地进行统计分析。
3. Discrete and Continuous Data | 离散数据与连续数据
Discrete data can only take certain separate values, usually whole numbers. They arise from counting. For example, the number of goals scored by a team in a season, the number of fouls committed, or the number of athletes in a squad are all discrete variables.
离散数据只能取某些分离的数值,通常是整数。它们来自计数。例如,一支球队在一个赛季中的进球数、犯规次数或一支队伍中的运动员人数都是离散变量。
Continuous data can take any value within a given range and are obtained by measuring. Examples include body mass (72.4 kg), vertical jump height (0.54 m), running speed (8.3 m/s) and blood lactate concentration (4.2 mmol/L). Continuous data often have decimal places and can be split into smaller and smaller intervals.
连续数据在给定范围内可以取任何数值,是通过测量获得的。例子包括体重(72.4 kg)、垂直跳高度(0.54 m)、跑步速度(8.3 m/s)和血乳酸浓度(4.2 mmol/L)。连续数据通常带有小数位,并且可以无限细分。
4. Levels of Measurement: Nominal, Ordinal, Interval, Ratio | 测量水平:定类、定序、定距、定比
Data can also be classified by their level of measurement. Nominal data are named categories with no order, such as playing position (goalkeeper, defender, midfielder, forward) or type of sport (swimming, running, cycling). Ordinal data have a rank order but the gaps between ranks are not equal, such as finishing positions in a race (1st, 2nd, 3rd) or a coach’s rating of effort as ‘low’, ‘medium’ or ‘high’.
数据还可以根据测量水平进行分类。定类数据是无顺序的命名类别,例如比赛位置(守门员、后卫、中场、前锋)或运动类型(游泳、跑步、骑行)。定序数据有等级顺序,但等级之间的差距不相等,例如比赛中的名次(第 1、第 2、第 3)或教练将努力程度评为 ‘低’、’中’、’高’。
Interval data have equal intervals between values but no true zero point. Temperature in degrees Celsius is an interval scale because 0°C does not mean ‘no temperature’. In sport, a performance rating scale from 1 to 10 is sometimes treated as interval. Ratio data have equal intervals and a true zero point, meaning zero represents the absence of the quantity. Examples include time, distance, mass, and force. With ratio data, it is meaningful to say that 20 m is twice as long as 10 m.
定距数据在数值之间有相等的间隔,但没有真正的零点。摄氏温度是定距量表,因为 0°C 不表示 ‘没有温度’。在体育中,从 1 到 10 的表现评分有时被视为定距数据。定比数据具有相等的间隔和真正的零点,意味着零点表示该数量的不存在。例子包括时间、距离、质量和力。对于定比数据,可以说 20 米是 10 米的两倍长,这是有意义的。
| Level | Properties | Sport Example |
|---|---|---|
| Nominal | Categories, no order | Shirt number category: forward or back |
| Ordinal | Rank order, unequal gaps | 1st, 2nd, 3rd in a 800 m race |
| Interval | Equal gaps, no true zero | Temperature of an ice bath in °C |
| Ratio | Equal gaps, true zero | Reaction time (0.32 s), jump height (0.45 m) |
5. Primary and Secondary Data Sources | 原始数据与二手数据来源
Primary data are collected directly by the researcher or coach for a specific purpose. For example, a strength and conditioning coach may conduct a 30 m sprint test to measure acceleration. Primary data are often more reliable for the specific question being asked, but they require time, equipment and standardised protocols.
原始数据是由研究者或教练为特定目的直接收集的数据。例如,体能教练可能会进行 30 米冲刺测试来测量加速度。原始数据通常对特定研究问题更可靠,但需要时间、设备和标准化流程。
Secondary data are data that already exist and were collected by someone else for another purpose. Examples include official league statistics, national fitness survey results, or data from a published research paper. Secondary data can save time and allow large-scale comparisons, but their quality and relevance must be checked carefully.
二手数据是已经存在、由他人为其他目的收集的数据。例子包括官方联赛统计、国家体能调查结果或已发表研究论文中的数据。二手数据可以节省时间并支持大规模比较,但必须仔细检查其质量和相关性。
6. Tabular Data Structures in Performance Analysis | 运动表现分析中的表格数据结构
The most common data structure in sport is the table, also called a dataset or data frame. In a table, each row represents a record, such as one athlete or one match event, and each column represents a variable, such as heart rate, distance covered, or number of sprints. Rows and columns can be sorted, filtered and aggregated to reveal patterns.
体育中最常见的数据结构是表格,也称为数据集或数据框。在表格中,每一行代表一条记录,如一名运动员或一次比赛事件;每一列代表一个变量,如心率、跑动距离或冲刺次数。行和列可以进行排序、筛选和汇总,以揭示规律。
For example, a performance analysis table might include columns for player ID, match date, total distance (m), high-speed running (m), and pass completion rate (%). The structure makes it easy to compare players or track changes over time. Most spreadsheets and statistical software store data in this rectangular format.
例如,一个表现分析表可能包括球员编号、比赛日期、总距离(米)、高速跑距离(米)和传球成功率(%)等列。这种结构便于比较球员或追踪随时间的变化。大多数电子表格和统计软件都以这种矩形格式存储数据。
7. Arrays and Matrices for Biomechanical Data | 生物力学数据的数组与矩阵
In biomechanics, data are often collected at high sampling rates from force plates, motion capture cameras or inertial sensors. These data are naturally stored as arrays: ordered collections of values indexed by time or sensor channel. A single force-time curve from a vertical jump may be stored as a one-dimensional array of force values sampled every 0.001 s.
在生物力学中,数据通常通过测力台、动作捕捉相机或惯性传感器以高采样率收集。这些数据很自然地存储为数组:按时间或传感器通道索引的有序值集合。垂直跳中的单条力-时间曲线可以存储为每 0.001 秒采样一次的力值一维数组。
A matrix is a two-dimensional array with rows and columns. For example, a matrix could store joint angles for 10 joints across 500 time frames during a running gait cycle. Matrices allow mathematical operations such as multiplying by a rotation matrix or calculating the inverse for inverse dynamics analysis.
矩阵是包含行和列的二维数组。例如,一个矩阵可以存储跑步步态周期中 10 个关节在 500 个时间帧上的关节角度。矩阵允许进行数学运算,例如乘以旋转矩阵或计算逆矩阵以进行逆向动力学分析。
F = m × a
For example, a force vector F can be stored as a column matrix with three components (Fₓ, Fᵧ, F₂) representing forces in the x, y and z directions.
例如,力矢量 F 可以存储为一个具有三个分量(Fₓ, Fᵧ, F₂)的列矩阵,分别表示 x、y 和 z 方向上的力。
8. Athlete Management Databases | 运动员管理数据库
Many professional teams and sports institutes use relational databases to manage athlete information. A relational database stores data in linked tables. For example, one table may hold athlete details (name, date of birth, position), another table may hold training sessions (date, duration, load), and a third may hold injury records (injury type, date, severity).
许多职业运动队和体育机构使用关系数据库来管理运动员信息。关系数据库将数据存储在相互关联的表中。例如,一个表可能保存运动员详细信息(姓名、出生日期、位置),另一个表保存训练课(日期、时长、负荷),第三个表保存伤病记录(伤病类型、日期、严重程度)。
The tables are linked by a unique identifier called a primary key. For instance, athlete ID can connect the athlete table to the training and injury tables. This structure reduces duplication and allows complex queries, such as ‘find all training sessions with load above 800 units for athletes who had a hamstring injury in the past 6 months’.
这些表通过称为主键的唯一标识符进行关联。例如,运动员编号可以将运动员表与训练表和伤病表连接起来。这种结构减少了重复,并允许进行复杂查询,例如 ‘查找过去 6 个月内有过腘绳肌损伤的运动员负荷超过 800 单位的所有训练课’。
- athlete table: athlete_id, name, sport, position – 运动员表:运动员编号、姓名、运动项目、位置
- training table: session_id, athlete_id, date, load – 训练表:课次编号、运动员编号、日期、负荷
- injury table: injury_id, athlete_id, injury_type, days_lost – 伤病表:伤病编号、运动员编号、伤病类型、缺训天数
9. Visual Data Structures: Heat Maps and Graphs | 可视化数据结构:热力图与图表
Data structures are not only for storage; they can also be visual. A heat map is a two-dimensional grid where each cell is coloured according to the value of a variable. In football, a player’s touch heat map shows which zones of the pitch they occupied most often. The underlying structure is a matrix of pitch zones with counts of touches in each zone.
数据结构不仅用于存储,还可以是可视化的。热力图是一个二维网格,其中每个单元格根据变量值进行着色。在足球中,球员的触球热力图显示了他们最常占据的球场区域。其底层结构是一个球场区域矩阵,每个区域中记录触球次数。
Graph structures are used for network analysis, such as passing networks. Each player is a node, and each pass between players is an edge. The thickness of an edge can represent the number of passes. This graph data structure helps coaches understand team connectivity and identify key playmakers.
图结构用于网络分析,例如传球网络。每名球员是一个节点,球员之间的每次传球是一条边。边的粗细可以代表传球次数。这种图数据结构帮助教练理解球队的连接性并识别关键组织者。
10. GPS Tracking Data Structures | GPS 跟踪数据结构
Global Positioning System (GPS) devices worn by athletes produce large amounts of time-series data. Each data point may include a timestamp, latitude, longitude, speed, acceleration and heart rate. The data are often structured as a sequence of records ordered by time, sometimes called an event stream or time series.
运动员佩戴的全球定位系统(GPS)设备会产生大量的时间序列数据。每个数据点可能包括时间戳、纬度、经度、速度、加速度和心率。这些数据通常按照时间排序为一系列记录,有时称为事件流或时间序列。
In practice, GPS data are often stored in nested structures. For example, a match file might contain a list of players, and each player has an array of time-stamped positions. This hierarchical structure allows efficient storage and fast retrieval of specific periods, such as the first 15 minutes of the second half.
在实践中,GPS 数据通常以嵌套结构存储。例如,一个比赛文件可能包含球员列表,每个球员都有一个带时间戳的位置数组。这种分层结构允许高效存储和快速检索特定时段,例如下半场前 15 分钟。
An example GPS data point might be stored as: time = 1250.4 s, speed = 7.8 m/s, distance = 5240.6 m, accelerations = 12.
一个 GPS 数据点示例可能存储为:时间 = 1250.4 s,速度 = 7.8 m/s,距离 = 5240.6 m,加速次数 = 12。
11. Reliability, Validity and Data Cleaning | 信度、效度与数据清理
The choice of data type and structure affects the reliability and validity of sport science measurements. Reliability refers to the consistency of a measurement when repeated under the same conditions. Validity refers to whether the data actually measure what they are intended to measure. For example, a GPS device may reliably record speed but may not validly measure very short accelerations.
数据类型和结构的选择会影响运动科学测量的信度和效度。信度是指测量在相同条件下重复进行时的一致性。效度是指数据是否真正测量了它们想要测量的内容。例如,GPS 设备可以可靠地记录速度,但可能无法有效地测量非常短时间的加速度。
Before analysis, raw data often need cleaning. This involves handling missing values, removing impossible outliers (such as a heart rate of 250 beats per minute in a resting athlete), and converting data types (for example, changing a text field ‘low’ into a numeric code). The structure of the dataset determines how easily these cleaning steps can be performed.
在分析之前,原始数据通常需要清理。这涉及处理缺失值、删除不可能的异常值(例如静息运动员心率 250 次/分),以及转换数据类型(例如将文本字段 ‘低’ 转换为数值代码)。数据集的结构决定了这些清理步骤执行的难易程度。
Coefficient of variation (CV) = (SD / Mean) × 100%
For example, if repeated 30 m sprint tests give a mean of 4.20 s and a standard deviation of 0.05 s, the CV is about 1.2%, which indicates high reliability.
例如,如果重复 30 米冲刺测试得到的平均值为 4.20 秒,标准差为 0.05 秒,则变异系数约为 1.2%,表明信度很高。
12. Summary and Exam Tips | 总结与考试提示
Data types describe the nature of the values: quantitative or qualitative, discrete or continuous, and the level of measurement (nominal, ordinal, interval, ratio). Data structures describe how values are organised: tables, arrays, matrices, databases, heat maps and graph networks. In sport and exercise science, selecting the correct data type and structure is crucial for valid analysis and clear communication.
数据类型描述数值的性质:定量或定性、离散或连续,以及测量水平(定类、定序、定距、定比)。数据结构描述数值的组织方式:表格、数组、矩阵、数据库、热力图和图网络。在体育与运动科学中,选择正确的数据类型和结构对于有效分析和清晰沟通至关重要。
When answering exam questions on this topic, always link the concept to a sport example. Define the term precisely, then give a relevant example and explain why it matters. For instance, if asked about continuous data, you could state that a swimmer’s 50 m freestyle time is continuous because it can take any value within a range, and this allows the coach to detect very small improvements.
在回答有关该主题的考试问题时,始终将概念与体育实例联系起来。准确给出术语定义,然后给出相关例子并解释其重要性。例如,如果被问到连续数据,你可以说游泳运动员 50 米自由泳时间是连续数据,因为它可以取某个范围内的任何值,这使得教练能够检测到非常微小的进步。
- Quantitative – numerical, counts or measures – 定量数据 – 数值,计数或测量
- Qualitative – descriptive, categories – 定性数据 – 描述性,类别
- Discrete – whole numbers from counting – 离散数据 – 计数的整数
- Continuous – any value from measuring – 连续数据 – 测量的任意值
- Table – rows as records, columns as variables – 表格 – 行是记录,列是变量
- Database – linked tables with primary keys – 数据库 – 带主键的关联表
Published by TutorHao | Physical Education Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导