IB Computer Science: Data Representation – Key Concepts | IB 计算机:数据表示 考点精讲

📚 IB Computer Science: Data Representation – Key Concepts | IB 计算机:数据表示 考点精讲

In IB Computer Science, understanding how data is represented inside a computer is fundamental. Computers operate using binary digits (bits), and every piece of information – numbers, text, images, sound – must be encoded as sequences of 0s and 1s. This article covers the key concepts you need to master for the examination, including number bases, negative number representations, floating point, character encoding, multimedia representation, and compression techniques.

在 IB 计算机科学中,理解计算机内部如何表示数据是基础。计算机使用二进制数字(位)运行,任何信息——数字、文本、图像、声音——都必须编码成 0 和 1 的序列。本文涵盖考试中必须掌握的核心概念,包括数制、负数表示法、浮点数、字符编码、多媒体表示和压缩技术。

1. Number Systems: Binary, Denary, Hexadecimal | 数制:二进制、十进制、十六进制

Computers use the binary (base‑2) number system because electronic components can reliably represent two states – on (1) and off (0). Humans naturally work with denary (base‑10). Hexadecimal (base‑16) provides a compact way to represent binary strings: one hex digit corresponds to four bits (a nibble). The digits are 0‑9 and A‑F (A=10, B=11, …, F=15). In IB Computer Science you must be able to convert freely between these three bases.

计算机使用二进制(基数为 2)数制,因为电子元件能够可靠地表示两种状态——开(1)和关(0)。人类习惯使用十进制(基数为 10)。十六进制(基数为 16)提供了一种简洁表示二进制串的方法:一个十六进制数字对应四个位(一个半字节)。数字为 0‑9 和 A‑F(A=10, B=11, …, F=15)。在 IB 计算机科学中,你必须能够在这三种进制之间自由转换。

Key terms: bit (binary digit), byte (8 bits), nibble (4 bits), word (a group of bits processed as a unit). The most significant bit (MSB) is the leftmost bit; the least significant bit (LSB) is the rightmost bit. Hexadecimal numbers are often prefixed with ‘0x’ or suffixed with ‘h’, e.g., 0x5A or 5Ah.

关键术语:位(二进制数字)、字节(8 位)、半字节(4 位)、字(作为一个单位处理的一组位)。最高有效位(MSB)是最左边的位;最低有效位(LSB)是最右边的位。十六进制数通常以 ‘0x’ 为前缀或以 ‘h’ 为后缀,例如 0x5A 或 5Ah。


2. Converting Between Bases | 进制转换

To convert binary to denary, sum the powers of 2 where a 1 appears. For example, 1101₂ = 1×2³ + 1×2² + 0×2¹ + 1×2⁰ = 8 + 4 + 0 + 1 = 13₁₀. Denary to binary uses successive division by 2, recording the remainders, then reading them from bottom to top. Hexadecimal to binary is easy: replace each hex digit with its 4‑bit binary equivalent. Binary to hexadecimal groups bits in fours from the right, then replaces each group with the hex digit.

将二进制转换为十进制,只需将出现 1 的位所对应的 2 的幂相加。例如,1101₂ = 1×2³ + 1×2² + 0×2¹ + 1×2⁰ = 8 + 4 + 0 + 1 = 13₁₀。十进制转二进制采用连续除以 2 并记录余数的方法,然后从下往上读取余数。十六进制转二进制很简单:将每个十六进制数字替换为其 4 位二进制等效值。二进制转十六进制则从右往左按四位一组分组,再将每组替换为十六进制数字。

Example table for binary‑hex conversion:

二进制‑十六进制转换示例表:

Binary 0000 0001 0010 0011 0100 0101 0110 0111
Hex 0 1 2 3 4 5 6 7
Binary 1000 1001 1010 1011 1100 1101 1110 1111
Hex 8 9 A B C D E F

3. Binary Addition and Overflow | 二进制加法与溢出

Binary addition follows simple rules: 0+0=0, 0+1=1, 1+0=1, 1+1=0 carry 1, 1+1+1=1 carry 1. When adding two binary numbers, you start from the LSB and move left. For example, 0110 (6) + 0111 (7) yields 1101 (13) in 4 bits. However, if the result requires more bits than the word size, an overflow error occurs. Overflow happens when the carry into the MSB and the carry out of the MSB are different – this flag is often checked by the CPU.

二进制加法遵循简单规则:0+0=0, 0+1=1, 1+0=1, 1+1=0 进位 1, 1+1+1=1 进位 1。相加两个二进制数时,从最低有效位开始向左移动。例如,0110 (6) + 0111 (7) 产生 1101 (13)(4 位)。然而,如果结果所需的位数超过字长,就会发生溢出错误。当进位进入最高有效位和从最高有效位出来的进位不同时,就会发生溢出——CPU 通常会检查这个标志。

Overflow is not the same as carry; carry is a normal part of addition across byte boundaries, while overflow indicates a result that cannot be represented in the given number of bits. In signed arithmetic, overflow corrupts the sign bit, leading to an incorrect result.

溢出与进位不同;进位是跨越字节边界的正常加法部分,而溢出表示结果无法用给定数量的位表示。在有符号运算中,溢出会破坏符号位,导致错误结果。


4. Sign-and-Magnitude Representation | 符号数值表示法

The simplest way to represent signed integers is to reserve the MSB for the sign (0 for positive, 1 for negative) and use the remaining bits for the magnitude. For example, in an 8‑bit system, +13 = 00001101 and -13 = 10001101. This method is easy for humans to read but creates two representations of zero (00000000 and 10000000) and complicates hardware arithmetic – the CPU would need separate circuits for addition and subtraction.

表示有符号整数的最简单方法是将最高有效位保留为符号位(0 表示正数,1 表示负数),并用剩余位表示数值大小。例如,在 8 位系统中,+13 = 00001101,-13 = 10001101。这种方法便于人类阅读,但产生了两个零的表示(00000000 和 10000000),并使硬件运算复杂化——CPU 需要为加法和减法设计不同的电路。

Because of these drawbacks, sign‑and‑magnitude is rarely used in practical computer arithmetic. It is, however, tested in IB to highlight the evolution towards two’s complement.

由于这些缺点,符号数值表示法在实际计算机运算中极少使用。然而,IB 会考察它以突出向补码的演变。


5. Two’s Complement Representation | 补码表示法

Two’s complement is the most common method for representing signed integers because it eliminates the dual‑zero issue and allows subtraction to be performed using addition circuits. To obtain the two’s complement of a binary number, invert all bits (one’s complement) and add 1 to the least significant bit. For example, the two’s complement of 00001101 (+13) is 11110011, which represents -13 in 8‑bit two’s complement.

补码是表示有符号整数最常用的方法,因为它消除了双零问题,并允许使用加法电路执行减法。要获得一个二进制数的补码,先将所有位取反(反码),然后在最低有效位加 1。例如,00001101 (+13) 的补码是 11110011,在 8 位补码中表示 -13。

To convert a negative two’s complement number back to denary, you can apply the same procedure: invert and add 1 to get the magnitude, then keep the negative sign. Alternatively, interpret the MSB as having a negative weight: the decimal value is -bₙ₋₁ × 2ⁿ⁻¹ + sum of bᵢ × 2ⁱ for i=0 to n‑2. For 11110011: -1×2⁷ + 1×2⁶ + 1×2⁵ + 1×2⁴ + 0×2³ + 0×2² + 1×2¹ + 1×2⁰ = -128 + 64+32+16+2+1 = -13.

要将负的补码数值转换回十进制,可以应用相同的步骤:取反加 1 得到数值大小,再保留负号。或者将最高有效位解释为负权:十进制值 = -bₙ₋₁ × 2ⁿ⁻¹ + Σ bᵢ × 2ⁱ(i 从 0 到 n‑2)。对于 11110011:-1×2⁷ + 1×2⁶ + 1×2⁵ + 1×2⁴ + 0×2³ + 0×2² + 1×2¹ + 1×2⁰ = -128 + 64+32+16+2+1 = -13。


6. Two’s Complement Arithmetic and Range | 补码运算与范围

Using two’s complement, subtraction A – B is performed as A + (two’s complement of B). Addition works exactly as binary addition, and any carry beyond the MSB is ignored. Overflow detection in two’s complement is crucial: overflow occurs if the sum of two positive numbers yields a negative result, or the sum of two negative numbers yields a positive result. Formally, overflow happens when the carry into the sign bit and the carry out of the sign bit differ.

使用补码,减法 A – B 可作为 A +(B 的补码)执行。加法与二进制加法完全相同,超出最高有效位的任何进位被忽略。补码中的溢出检测至关重要:若两个正数之和产生负结果,或两个负数之和产生正结果,则发生溢出。形式上,当进入符号位的进位与离开符号位的进位不同时发生溢出。

The range of an n‑bit two’s complement number is from -2ⁿ⁻¹ to 2ⁿ⁻¹ – 1. For example, an 8‑bit representation can hold values from -128 to 127. The smallest negative number does not have a positive counterpart in the same bit width, which can trap unwary programmers.

n 位补码数的范围是从 -2ⁿ⁻¹ 到 2ⁿ⁻¹ – 1。例如,8 位表示可以保存从 -128 到 127 的值。最小的负数在相同位宽下没有对应的正数,这可能会给粗心的程序员造成陷阱。


7. Floating Point Numbers | 浮点数

Numbers with fractional parts and very large or very small magnitudes are stored in floating‑point format, similar to scientific notation. A binary floating‑point number consists of a sign bit (S), a mantissa (M, also called significand), and an exponent (E). The general form is: value = (-1)^S × 1.M × 2^(E – bias). The mantissa is normalised so that its first bit after the binary point is always 1 (except for zero), maximising precision.

带有小数部分以及数量级非常大或非常小的数字以浮点格式存储,类似于科学记数法。一个二进制浮点数由符号位(S)、尾数(M,也称为有效数)和指数(E)组成。一般形式为:value = (-1)^S × 1.M × 2^(E – bias)。尾数被归一化,使得二进制小数点后的第一位始终为 1(零除外),从而最大化精度。

The IB exam expects familiarity with the IEEE 754 single‑precision (32‑bit) layout: 1 bit sign, 8 bits exponent (bias 127), 23 bits mantissa. To convert a denary number to floating point, normalise it to 1.xxx × 2^y, then store the fractional part of the mantissa (without the leading 1) and store the exponent as (y + 127) in binary. For example, the number 10.5 in binary is 1010.1 = 1.0101 × 2³. The sign=0, mantissa=010100… (23 bits), exponent=3+127=130 (10000010). The stored pattern is 0 10000010 01010000000000000000000.

IB 考试要求熟悉 IEEE 754 单精度(32 位)布局:1 位符号、8 位指数(偏移量 127)、23 位尾数。要将十进制数转换为浮点数,将其归一化为 1.xxx × 2^y,然后存储不含前导 1 的尾数小数部分,并以二进制将指数存储为 (y + 127)。例如,数字 10.5 二进制为 1010.1 = 1.0101 × 2³。符号=0,尾数=010100…(23 位),指数=3+127=130(10000010)。存储模式为 0 10000010 01010000000000000000000。

You should also be aware of special values: zero (all bits 0), infinity (exponent all 1s, mantissa 0), and NaN (exponent all 1s, mantissa non‑zero).

你还应了解特殊值:零(所有位为 0)、无穷大(指数全为 1,尾数为 0)和 NaN(指数全为 1,尾数非零)。


8. Character Encoding: ASCII and Unicode | 字符编码:ASCII 与 Unicode

Characters are represented by assigning a unique numeric code to each symbol. The American Standard Code for Information Interchange (ASCII) originally used 7 bits to encode 128 characters (including control characters, digits, uppercase and lowercase letters, and punctuation). Extended ASCII uses 8 bits (1 byte) to cover 256 characters, often including accented letters and line‑drawing characters. ASCII values: ‘A’=65, ‘a’=97, ‘0’=48, space=32.

字符通过为每个符号分配唯一的数字代码来表示。美国信息交换标准代码(ASCII)最初使用 7 位编码 128 个字符(包括控制字符、数字、大小写字母和标点符号)。扩展 ASCII 使用 8 位(1 字节)覆盖 256 个字符,通常包括带重音字母和制表符号。ASCII 值:’A’=65,’a’=97,’0’=48,空格=32。

Unicode is a modern standard that aims to encode all writing systems of the world. It defines over 140,000 characters across multiple planes. The most common encoding form is UTF‑8, which is backward‑compatible with ASCII: the first 128 Unicode code points are identical to ASCII and are stored in 1 byte. Additional bytes represent extended characters using multi‑byte sequences. UTF‑8 can use up to 4 bytes per character, making it efficient for predominantly ASCII text while supporting international scripts.

Unicode 是一个现代标准,旨在对世界上所有书写系统进行编码。它在多个平面中定义了超过 140,000 个字符。最常见的编码形式是 UTF‑8,它与 ASCII 向后兼容:前 128 个 Unicode 代码点与 ASCII 相同,并以 1 个字节存储。额外字节使用多字节序列表示扩展字符。UTF‑8 每个字符最多可使用 4 个字节,这使得

Published by TutorHao | IB Computer Science Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading