📚 IGCSE CIE Computer Science: Data Representation Key Points | IGCSE CIE 计算机:数据表示 考点精讲
Data representation is a foundational topic in IGCSE CIE Computer Science, covering how computers store and manipulate numbers, text, images, and sound. This article breaks down every essential concept you need to master for the exam, from binary arithmetic to compression techniques.
数据表示是 IGCSE CIE 计算机科学的基础主题,涵盖计算机如何存储和处理数字、文本、图像及声音。本文剖析考试必须掌握的每一个核心概念,从二进制运算到压缩技术。
1. Bits, Bytes and Binary Numbers | 位、字节与二进制数
A bit (binary digit) is the smallest unit of data, storing either 0 or 1. A group of 8 bits forms a byte, which can represent 2⁸ = 256 different values (0 to 255). Computers use binary because digital circuits have two stable states (on/off).
位(二进制数字)是最小的数据单位,存储 0 或 1。8 位组成一个字节,可表示 2⁸ = 256 个不同的值(0 到 255)。计算机使用二进制是因为数字电路只有两种稳定状态(开/关)。
A typical GCSE question asks you to convert a 5‑digit binary number to denary using place values (16s, 8s, 4s, 2s, 1s).
典型的 GCSE 题目要求用位权(16、8、4、2、1)将 5 位二进制数转换为十进制。
| Place value / 位权 | 16 | 8 | 4 | 2 | 1 |
|---|---|---|---|---|---|
| Binary / 二进制 | 1 | 0 | 1 | 1 | 0 |
10110₂ = (1×16) + (0×8) + (1×4) + (1×2) + (0×1) = 16 + 4 + 2 = 22.
10110₂ = (1×16) + (0×8) + (1×4) + (1×2) + (0×1) = 16 + 4 + 2 = 22。
2. Converting Between Denary and Binary | 十进制与二进制之间的转换
To convert denary to binary, repeatedly divide by 2 and record the remainders. Reading the remainders backwards gives the binary equivalent.
要将十进制转换为二进制,连续除以 2 并记录余数,反向读取余数即得二进制。
Example: Convert 29 to binary.
29 ÷ 2 = 14 remainder 1
14 ÷ 2 = 7 remainder 0
7 ÷ 2 = 3 remainder 1
3 ÷ 2 = 1 remainder 1
1 ÷ 2 = 0 remainder 1
Reading from bottom to top: 11101₂.
示例:将 29 转换为二进制。
29 ÷ 2 = 14 余 1
14 ÷ 2 = 7 余 0
7 ÷ 2 = 3 余 1
3 ÷ 2 = 1 余 1
1 ÷ 2 = 0 余 1
从下往上读:11101₂。
For values up to 255, you can also use the place‑value table, subtracting the largest power of 2 that fits.
对于不超过 255 的值,你也可以使用权值表,减去能容纳的 2 的最大次幂。
3. Binary Addition and Overflow | 二进制加法与溢出
Binary addition follows four simple rules: 0+0=0, 0+1=1, 1+0=1, 1+1=0 carry 1. When two 8‑bit numbers produce a result requiring 9 bits, an overflow error occurs because the processor’s register can only hold a fixed number of bits.
二进制加法遵循四条简单规则:0+0=0,0+1=1,1+0=1,1+1=0 进 1。当两个 8 位数相加产生需要 9 位的结果时,会发生溢出错误,因为处理器的寄存器只能保存固定数量的位。
Detecting overflow: if the carry into the most significant bit (MSB) differs from the carry out, overflow has occurred. This is essential when interpreting results as signed or unsigned.
溢出检测:如果进入最高位(MSB)的进位与输出进位不同,就发生了溢出。在以带符号或无符号方式解释结果时,这一点至关重要。
4. Logical Shifts | 逻辑移位
A logical left shift moves each bit one place to the left; a 0 is placed in the least significant bit (LSB) and the MSB is lost. Each left shift multiplies the unsigned value by 2.
逻辑左移将每一位向左移动一位;最低位(LSB)补 0,最高位(MSB)丢失。每左移一位,无符号数值乘以 2。
A logical right shift moves bits to the right; 0 enters the MSB, and the LSB is lost. Each right shift performs integer division by 2 on unsigned numbers.
逻辑右移将位向右移动;最高位补 0,最低位丢失。每右移一位,对无符号数进行整数除以 2。
Example: 00010110₂ (22) left shift → 00101100₂ (44)
示例:00010110₂ (22) 左移 → 00101100₂ (44)
5. Hexadecimal: Uses and Conversions | 十六进制:用途与转换
Hexadecimal (base 16) uses digits 0‑9 and letters A‑F. It is shorter and less error‑prone than binary, so programmers use it to represent memory addresses, colour codes, and MAC addresses. One hex digit exactly represents 4 bits.
十六进制(基数为 16)使用数字 0‑9 和字母 A‑F。它比二进制简短且不易出错,因此程序员用它表示内存地址、颜色代码和 MAC 地址。一个十六进制数字刚好对应 4 个比特。
To convert binary to hex, group bits into nibbles (4 bits) and convert each nibble. Example: 10111101₂ → 1011 1101 → B D → BD₁₆.
二进制转十六进制:将位分组为半字节(4 位),然后转换每个半字节。示例:10111101₂ → 1011 1101 → B D → BD₁₆。
To convert denary to hex, either go via binary or repeatedly divide by 16, recording remainders in reverse.
十进制转十六进制:可通过二进制中转,也可以反复除以 16,反向记录余数。
6. Representing Text: ASCII and Unicode | 文本表示:ASCII 与 Unicode
Characters are represented by numeric codes. ASCII uses 7 bits (0‑127) for English letters, digits, and symbols. Extended ASCII uses 8 bits (0‑255) for additional characters such as accented letters.
字符用数字编码表示。ASCII 用 7 位(0‑127)表示英文字母、数字和符号。扩展 ASCII 用 8 位(0‑255)表示带重音符号等额外字符。
Unicode was developed to cover all writing systems, using up to 32 bits per character. UTF‑8 is a variable‑length encoding that is backward‑compatible with ASCII. Unicode allows international text to be stored and transmitted consistently.
Unicode 为涵盖所有书写系统而开发,每个字符最多使用 32 位。UTF‑8 是一种与 ASCII 向后兼容的变长编码。Unicode 使得国际文本能够一致地存储和传输。
- ASCII advantage: requires less memory per character.
- Unicode advantage: supports many languages and emojis.
- ASCII 优点:每个字符占用更少内存。
- Unicode 优点:支持多种语言和表情符号。
7. Representing Images: Pixels, Resolution, Colour Depth | 图像表示:像素、分辨率与颜色深度
A bitmap image consists of a grid of pixels. Each pixel is assigned a binary code for its colour. The resolution is the number of pixels in the image, e.g. 1920×1080. Higher resolution gives more detail but larger file size.
位图图像由像素网格组成。每个像素被赋予其颜色的二进制编码。分辨率是图像中的像素数量,如 1920×1080。分辨率越高,细节越多,但文件也越大。
Colour depth (bit depth) is the number of bits used per pixel. A 1‑bit image can show 2 colours; an 8‑bit image shows 2⁸ = 256 colours; 24‑bit (True Colour) shows 224 ≈ 16.7 million colours.
颜色深度(位深度)是每个像素使用的位数。1 位图像可显示 2 种颜色;8 位图像显示 2⁸ = 256 种颜色;24 位(真彩色)显示 224 ≈ 1670 万种颜色。
Metadata stores additional information such as width, height, colour depth, and creation date, which increases the total file size.
元数据存储附加信息,如宽度、高度、颜色深度和创建日期,这会增加总文件大小。
8. Calculating Image File Size | 计算图像文件大小
Image file size is estimated by:
图像文件大小可按以下公式估算:
File size (bits) = width × height × colour depth
文件大小(比特) = 宽度 × 高度 × 颜色深度
Then convert bits to bytes (divide by 8), and further to kilobytes, megabytes if required. This calculation ignores metadata and compression; it is the raw size.
然后将位数转换为字节(除以 8),必要时再转换为千字节、兆字节。该计算忽略元数据和压缩,是原始大小。
Example: an image 1024×768 pixels with 24‑bit colour depth: 1024 × 768 × 24 = 18,874,368 bits = 2,359,296 bytes ≈ 2.3 MB.
示例:1024×768 像素、24 位颜色深度的图像:1024 × 768 × 24 = 18,874,368 位 = 2,359,296 字节 ≈ 2.3 MB。
9. Representing Sound: Sampling and Bit Depth | 声音表示:采样与位深度
Sound is an analogue signal. To store it digitally, the amplitude is measured at regular intervals (sampling) and converted to a binary value (quantisation). The sample rate is measured in hertz (Hz).
声音是模拟信号。为以数字形式存储,需以固定间隔测量振幅(采样)并转换为二进制值(量化)。采样率的单位是赫兹(Hz)。
A higher sample rate captures higher frequencies, improving quality. The bit depth determines the number of possible amplitude levels: with n bits you get 2ⁿ levels. Common CD quality is 44.1 kHz, 16‑bit.
较高的采样率能捕获更高的频率,提升音质。位深度决定了可能的振幅等级数量:n 位可得到 2ⁿ 种等级。常见的 CD 音质为 44.1 kHz,16 位。
- Increasing sample rate or bit depth → better quality, larger file size.
- 增加采样率或位深度 → 更好的音质,更大的文件。
10. Calculating Sound File Size | 计算声音文件大小
Mono sound file size (bits) = sample rate × bit depth × duration (seconds).
单声道声音文件大小(比特) = 采样率 × 位深度 × 时长(秒)。
Stereo file size = sample rate × bit depth × seconds × 2
立体声文件大小 = 采样率 × 位深度 × 秒 × 2
Example: 30 seconds of stereo sound at 44.1 kHz, 16‑bit: 44,100 × 16 × 30 × 2 = 42,336,000 bits = 5,292,000 bytes ≈ 5.05 MB.
示例:30 秒立体声,44.1 kHz,16 位:44,100 × 16 × 30 × 2 = 42,336,000 比特 = 5,292,000 字节 ≈ 5.05 MB。
Always show your working in exams and check unit conversions carefully.
考试中务必展示计算过程并仔细核对单位换算。
11. Data Compression: Lossy and Lossless | 数据压缩:有损与无损
Compression reduces file size for storage or transmission. Lossless compression preserves the original data exactly; it is essential for text, programs, and some images. Techniques include run‑length encoding and dictionary methods.
压缩可减小文件大小以利于存储或传输。无损压缩精确保留原始数据;对于文本、程序及某些图像必不可少。常用技术包括游程编码和字典方法。
Lossy compression discards some detail that is less noticeable to human senses, achieving much higher compression ratios. It is used for photographs (JPEG), music (MP3), and video (MPEG). Lost data cannot be recovered.
有损压缩丢弃人类感官不易察觉的一些细节,实现高得多的压缩比。它用于照片(JPEG)、音乐(MP3)和视频(MPEG)。丢失的数据无法恢复。
| Aspect / 方面 | Lossless / 无损 | Lossy / 有损 |
|---|---|---|
| Data integrity / 数据完整性 | Perfect original restored / 完美恢复原始数据 | Some information lost / 部分信息丢失 |
| Compression ratio / 压缩比 | Moderate / 中等 | High / 高 |
| Common use / 常见用途 | PNG, ZIP, GIF / PNG, ZIP, GIF | JPEG, MP3, MP4 / JPEG, MP3, MP4 |
12. Run‑Length Encoding (RLE) in Detail | 游程编码详解
RLE is a simple lossless compression method that replaces consecutive identical data values with a count and the value. For example, AAABBBBBCC can be stored as 3A 5B 2C, saving space when repetitions are frequent.
游程编码是一种简单的无损压缩方法,用计数和数值对来代替连续相同的数据值。例如,AAABBBBBCC 可存储为 3A 5B 2C,当重复频繁时可节省空间。
In bitmap images, RLE works well with large blocks of the same colour. However, if the data has few runs, encoding can increase the file size. CIE often asks students to apply RLE to a given string or pixel sequence.
在位图图像中,游程编码对大面积同色区域效果良好。但如果数据重复很少,编码可能会增大文件体积。CIE 常要求学生对给定字符串或像素序列应用游程编码。
Example: uncompressed 10 pixels: WWWWBBWWWW → RLE: (W,4) (B,2) (W,4). Each run is represented by two values: colour and count.
示例:未压缩的 10 个像素:WWWWBBWWWW → 游程编码:(W,4) (B,2) (W,4)。每次游程用两个值表示:颜色和计数。
Published by TutorHao | Computer Science Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导