Data Representation Exam Key Points | IB CIE 计算机:数据表示考点精讲

📚 Data Representation Exam Key Points | IB CIE 计算机:数据表示考点精讲

Data representation is the foundation of all computer science – how numbers, text, images and sound are encoded as bits. In IB and CIE assessments, you are expected to convert between binary, denary and hexadecimal, explain how integers and characters are stored, calculate file sizes, and understand compression techniques. This revision guide walks you through every key concept with bilingual explanations, ready for your exam.

数据表示是计算机科学的基础——数字、文字、图像和声音如何被编码为比特。在IB和CIE的考试中,你需要能在二进制、十进制和十六进制之间转换,解释整数和字符的存储方式,计算文件大小,并理解压缩技术。这份复习指南用中英双语为你梳理每一个核心概念,助你轻松应对考试。

1. Why Data Representation Matters | 数据表示的重要性

Computers use binary (base-2) because their logic circuits rely on two states: on and off. Everything from a keypress to a video stream is ultimately stored as a string of 0s and 1s. Understanding how these patterns represent real-world data allows you to predict storage needs, prevent errors like overflow, and choose efficient encoding methods – topics that appear in almost every IB and CIE paper.

计算机使用二进制(基数为2),因为其逻辑电路依赖于两种状态:开和关。从按键到视频流,所有内容最终都以一串0和1的形式存储。理解这些模式如何表示真实世界的数据,可以帮助你预测存储需求、防止溢出等错误,并选择高效的编码方法——这些主题几乎会出现在每一份IB和CIE试卷中。


2. Binary System Essentials | 二进制系统要点

Binary is a positional number system with only two digits: 0 and 1. Each position represents a power of 2. For example, the binary number 1101₂ can be expanded as 1×2³ + 1×2² + 0×2¹ + 1×2⁰ = 8+4+0+1 = 13 in denary. This weighted sum approach is the key to all binary-to-denary conversions.

二进制是一种只使用0和1两个数字的位值记数系统。每一位代表2的幂。例如,二进制数1101₂可以展开为1×2³ + 1×2² + 0×2¹ + 1×2⁰ = 8+4+0+1 = 13(十进制)。这种加权求和方法是所有二进制转十进制换算的关键。

When converting denary to binary, repeatedly divide by 2 and record the remainders. For 13: 13÷2=6 rem 1, 6÷2=3 rem 0, 3÷2=1 rem 1, 1÷2=0 rem 1. Reading the remainders backwards gives 1101₂. A common exam mistake is forgetting to include leading zeros when a fixed number of bits is specified.

将十进制转换为二进制时,不断除以2并记录余数。以13为例:13÷2=6余1,6÷2=3余0,3÷2=1余1,1÷2=0余1。反向读出余数即得1101₂。考试中常见的错误是在指定固定位数时忘记补齐前导零。


3. Hexadecimal Notation | 十六进制表示法

Hexadecimal (base-16) provides a compact way to display binary values. Digits 0-9 represent themselves, while A=10, B=11, C=12, D=13, E=14, F=15. One hex digit corresponds to exactly 4 bits (a nibble), so conversion is straightforward: group binary into nibbles and replace each group with its hex equivalent.

十六进制(基数为16)提供了一种紧凑显示二进制值的方式。数字0-9代表自身,而A=10,B=11,C=12,D=13,E=14,F=15。一个十六进制位恰好对应4个比特(一个半字节),因此转换非常简单:将二进制数分组为半字节,每组用相应的十六进制符号替换即可。

Example: 11010110₂ → 1101 0110 → D6₁₆

示例:11010110₂ → 1101 0110 → D6₁₆

To convert hex to denary, multiply each digit by 16 raised to the appropriate power. For D6₁₆: D×16¹ + 6×16⁰ = 13×16 + 6 = 214. In exams you often see hex used for colour codes (e.g. #FFA500) and memory addresses.

要将十六进制转为十进制,每位数字乘以16的相应次幂。D6₁₆:D×16¹ + 6×16⁰ = 13×16 + 6 = 214。考试中常以颜色编码(如#FFA500)和内存地址等形式考查十六进制。


4. Data Units and Prefixes | 数据单位与前缀

Data sizes are measured in bytes. One byte = 8 bits. The SI prefixes kilo, mega, giga, tera are based on powers of 10, whereas the binary-based IEC prefixes kibi, mebi, gibi use powers of 2. In CIE and IB, you must know both conventions and be able to convert between them.

数据大小以字节为单位衡量。1字节 = 8比特。国际单位制前缀千、兆、吉、太基于10的幂,而基于二进制的IEC前缀kibi、mebi、gibi使用2的幂。在CIE和IB考试中,你需要了解这两种约定并能在它们之间进行转换。

Decimal Binary (IEC) Bytes
1 kB = 10³ 1 KiB = 2¹⁰ 1,024 bytes
1 MB = 10⁶ 1 MiB = 2²⁰ 1,048,576 bytes
1 GB = 10⁹ 1 GiB = 2³⁰ 1,073,741,824 bytes

When calculating file sizes, always check whether the question uses decimal or binary multipliers. A typical trap is using 1 MB = 1,000,000 bytes when the image size formula produces a byte count that should be divided by 1,048,576.

在计算文件大小时,务必确认题目使用的是十进制乘法因子还是二进制乘法因子。常见陷阱是,当图像大小公式得出字节数后,本应除以1,048,576得到MiB,却错误地除以1,000,000得到MB。


5. Integer Representation: Sign-and-Magnitude & Two’s Complement | 整数表示:原码与补码

Signed integers can be represented in several ways. Sign-and-magnitude uses the most significant bit as a sign (0 = positive, 1 = negative) and the remaining bits as the magnitude. For example, in 8 bits, +9 is 00001001 and -9 is 10001001. This method still causes two zeros (+0 and -0) and complicates arithmetic.

有符号整数有多种表示方法。原码(Sign-and-magnitude)用最高位表示符号(0=正,1=负),其余位表示数值大小。例如8位下,+9为00001001,-9为10001001。这种方法仍存在+0和-0两个零,且使算术运算复杂化。

Two’s complement overcomes these issues and is the standard representation in modern computers. To obtain the two’s complement negative of a number, invert all bits and add 1. The most significant bit still indicates the sign, but the range is asymmetric. For n bits, values range from -2ⁿ⁻¹ to 2ⁿ⁻¹-1.

补码(Two’s complement)克服了这些缺点,是现代计算机中表示有符号整数的标准方法。要得到一个数的补码负数,将所有位取反后加1。最高位仍表示符号,但数值范围不再对称。对于n位,可表示的范围是从-2ⁿ⁻¹到2ⁿ⁻¹-1。

8-bit two’s complement range: -128 to 127 | 8位补码范围:-128 到 127

Example: -3 in 4-bit two’s complement: +3 = 0011 → invert → 1100 → add 1 → 1101

示例:4位补码表示-3:+3 = 0011 → 取反 → 1100 → 加1 → 1101


6. Binary Arithmetic & Overflow | 二进制算术与溢出

Addition in binary follows simple rules: 0+0=0, 0+1=1, 1+0=1, 1+1=0 with a carry of 1 to the next higher bit. Subtraction is performed by adding the two’s complement of the subtrahend. For instance, 5 – 3 becomes 5 + (-3) in two’s complement form.

二进制加法遵循简单规则:0+0=0,0+1=1,1+0=1,1+1=0并向高位进1。减法通过加上减数的补码来完成。例如,5 – 3 转换为补码形式下的5 + (-3)。

Overflow occurs when the result of an operation falls outside the representable range. In two’s complement arithmetic, overflow can be detected by comparing the carry into the sign bit with the carry out of it. If they differ, overflow has occurred. For example, adding 0111 (7) + 0001 (1) in 4-bit gives 1000 (-8), which is incorrect — the carries into and out of the sign bit were 1 and 0 respectively.

当运算结果超出可表示的范围时,就会发生溢出。在补码运算中,可比较进入符号位的进位与从符号位产生的进位来检测溢出;若二者不同,即发生溢出。例如,在4位下0111(7)加0001(1)得到1000(-8),结果错误——进入符号位的进位为1,符号位产生的进位为0,两者不同。


7. Character Encoding: ASCII & Unicode | 字符编码:ASCII与Unicode

Characters are stored as binary codes. The 7-bit ASCII table represents 128 characters, including control codes, digits, uppercase and lowercase letters. Extended ASCII uses 8 bits, giving 256 characters. However, ASCII cannot cover international scripts simultaneously.

字符以二进制码形式存储。7位ASCII表有128个字符,包括控制码、数字、大写和小写字母。扩展ASCII使用8位,提供256个字符。但ASCII无法同时覆盖各国的文字。

Unicode solves this with a much larger code space. Common encodings include UTF-8 (variable-length, 1 to 4 bytes) and UTF-16. In exam questions, you may be asked to explain why Unicode is needed for multilingual applications, or to calculate the storage for a given text string in a specific encoding.

Unicode用更大的码空间解决了这一问题。常见的编码有UTF-8(变长,1到4字节)和UTF-16。考题中可能会要求解释为什么多语言应用需要Unicode,或计算给定文本字符串在特定编码下的存储空间。

‘A’ = 65₁₀ = 01000001₂ in ASCII; ‘中’ = U+4E2D in Unicode, encoded in UTF-8 as E4 B8 AD.

‘A’ = 65₁₀ = 01000001₂(ASCII);’中’ = U+4E2D(Unicode),UTF-8编码为E4 B8 AD。


8. Floating-Point Numbers | 浮点数

Floating-point representation stores real numbers in the form ± mantissa × 2^exponent, similar to scientific notation. The IEEE 754 standard defines single (32-bit) and double (64-bit) precision formats, allocating bits for sign, exponent and mantissa. The exponent is typically stored with a bias to allow negative exponents.

浮点数表示法以 ± 尾数 × 2^指数 的形式存储实数,类似科学记数法。IEEE 754标准定义了单精度(32位)和双精度(64位)格式,将比特分配给符号、指数和尾数。指数通常通过偏移量存储,以支持负指数。

When converting a denary number to binary floating point, the number is first expressed in normalised form (leading 1 before the binary point). Bits are then packed into the sign-exponent-mantissa fields. Precision is limited, so rounding errors can occur, which is a frequent discussion point in IB and CIE papers.

将十进制数转换为二进制浮点数时,首先将其表示为规格化形式(小数点前为1)。然后将各部位分别填入符号、指数和尾数区。精度有限,因此可能产生舍入误差,这是IB和CIE试卷中常见的讨论点。


9. Representing Images | 图像表示

A bitmap image is a grid of pixels, each assigned a binary value for its colour. The colour depth (bits per pixel) determines how many distinct colours can be represented: a depth of n bits gives 2ⁿ colours. A 24-bit colour depth (8 bits each for red, green, blue) can display over 16 million colours.

位图图像是一个像素网格,每个像素被赋予一个表示颜色的二进制值。颜色深度(每像素位数)决定了能表示的不同颜色数量:n位可表示2ⁿ种颜色。24位颜色深度(红、绿、蓝各8位)可显示超过1600万种颜色。

Image file size (bytes) = width × height × colour depth / 8

图像文件大小(字节)= 宽度 × 高度 × 颜色深度 / 8

When metadata is included, the actual file size is slightly larger. Resolution refers to pixel density (e.g. 300 dpi). In exam calculations, be careful to convert all units consistently and clarify whether you are using decimal or binary prefixes for the final file size.

当包含元数据时,实际文件大小会稍大一些。分辨率指像素密度(如300 dpi)。在考试计算中,注意统一换算单位,并在给出最终文件大小时明确使用的是十进制还是二进制前缀。


10. Representing Sound | 声音表示

Sound is analogue, so it must be sampled to be stored digitally. The sampling rate (e.g. 44.1 kHz) is the number of samples taken per second, and the bit depth determines the number of possible amplitude levels per sample. Higher values for both give better fidelity but larger files.

声音是模拟信号,必须通过采样才能以数字形式存储。采样率(如44.1 kHz)是每秒采集的样本数,位深度决定了每个样本可能的振幅级数。二者数值越高,保真度越高,但文件也越大。

Sound file size (bits) = sampling rate × bit depth × duration (s) × number of channels

声音文件大小(位)= 采样率 × 位深度 × 时间(秒) × 声道数

Common CD-quality audio uses 44.1 kHz, 16-bit stereo. When you are asked to calculate the size of a short recording, remember to convert bits to bytes (divide by 8) and then to the required unit. A frequent mistake is forgetting to multiply by the number of channels.

CD品质的音频通常使用44.1 kHz、16位立体声。当被要求计算一段短录音的文件大小时,记得将位转换为字节(除以8)再换算为需要的单位。常见的错误是忘记乘以声道数。


11. Compression Techniques | 压缩技术

Compression reduces file size for storage and transmission. Lossless compression (e.g. ZIP, PNG, FLAC) preserves all original data and works by finding patterns or redundancies. Run-length encoding (RLE) is a simple lossless method that replaces repeated values with a count and value pair, effective for images with large uniform areas.

压缩可减小文件大小以便存储和传输。无损压缩(如ZIP、PNG、FLAC)保留所有原始数据,通过寻找模式或冗余来实现。游程编码是一种简单的无损方法,用计数和数值对替代重复内容,对于有大片均匀区域的图像很有效。

Lossy compression (e.g. JPEG, MP3) sacrifices some data that is less perceptible to humans to achieve much smaller sizes. In exams, you may be asked to compare the two or justify when a lossy codec would be inappropriate (e.g. for text or medical imaging) because it permanently removes detail.

有损压缩(如JPEG、MP3)通过丢弃人类不易察觉的部分数据来显著缩小文件。考试中可能要求比较这两种压缩方式,或说明在什么情况下不适合使用有损编码(如文本或医学影像),因为它会永久性地丢失细节。

Published by TutorHao | Computer Science Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version