📚 Data Representation | 数据表示 考点精讲
All data inside a computer system is stored and processed as sequences of 0s and 1s, known as binary digits. For the IGCSE OCR Computer Science exam, you must understand how numbers, text, images and sound are represented in binary, and how to perform conversions between denary, binary and hexadecimal. This revision guide walks through the essential concepts, including binary addition, two’s complement, character encoding, multimedia representation and data compression.
计算机系统内所有的数据都以一串串 0 和 1 (二进制位) 的形式存储和处理。对于 IGCSE OCR 计算机科学考试,你必须掌握数字、文字、图像和声音如何用二进制表示,以及如何在十进制、二进制和十六进制之间进行转换。本复习指南全面梳理二进制加法、二进制补码、字符编码、多媒体表示和数据压缩等核心考点。
1. Binary and Denary Numbers | 二进制与十进制
A denary (decimal) number uses base 10 with digits 0–9. A binary number uses base 2 with digits 0 and 1. Each position in a binary number represents a power of 2, starting from 2⁰ on the right. To convert from binary to denary, multiply each bit by its column value and sum the results. For an 8‑bit number, column values are 128, 64, 32, 16, 8, 4, 2, 1.
十进制数使用基数为 10,数码为 0–9。二进制数使用基数为 2,数码为 0 和 1。二进制数的每一位代表 2 的幂,从最右边的 2⁰ 开始。要将二进制转换为十进制,将每个位乘以其位权值,然后相加。对于 8 位二进制数,位权值分别是 128、64、32、16、8、4、2、1。
To convert denary to binary, repeatedly divide the denary number by 2 and record the remainders. The binary number is the sequence of remainders read from bottom to top. For example, denary 89 → 01011001 in 8‑bit binary (0b01011001). OCR often expects you to show an 8‑bit representation.
将十进制转换为二进制,反复将十进制数除以 2 并记录余数。从下往上读取余数序列即为二进制数。例如,十进制 89 转为 8 位二进制 01011001。OCR 通常要求展示 8 位二进制表示。
2. Hexadecimal Numbers | 十六进制
Hexadecimal (hex) uses base 16 with digits 0–9 and letters A–F (A=10, B=11, …, F=15). Hex is a convenient shorthand for binary, because one hex digit represents exactly four binary bits (a nibble). This reduces errors when reading or writing long binary strings.
十六进制使用基数为 16,数码为 0–9 以及字母 A–F (A=10, B=11, …, F=15)。十六进制是二进制的便捷简写形式,因为一位十六进制数码恰好代表四位二进制位(一个 nibble)。这可以减少读写长二进制串时的错误。
To convert binary to hex, split the binary number into groups of four bits from the right, then convert each group to its hex equivalent. Example: 1011 1010₂ → B A₁₆ → BA₁₆. To convert denary to hex, first convert to binary, or divide by 16 and use remainders.
将二进制转换为十六进制,从右边开始将二进制数分成四位一组,然后将每组转换为对应的十六进制数码。示例:1011 1010₂ → B A₁₆ → BA₁₆。将十进制转为十六进制,可先转换为二进制,或除以 16 并用余数转换。
3. Data Units | 数据单位
The smallest unit of data is a bit (binary digit). A group of 4 bits is a nibble, and a group of 8 bits is a byte. Larger units are formed by multiples of bytes: kilobyte (KB) ≈ 10³ bytes, megabyte (MB) ≈ 10⁶ bytes, gigabyte (GB) ≈ 10⁹ bytes, terabyte (TB) ≈ 10¹² bytes. In some contexts, binary prefixes (kibibyte, mebibyte) use powers of 2, but the OCR syllabus uses the decimal definitions.
数据的最小单位是 bit (二进制位)。4 位组成一个 nibble,8 位组成一个字节 (byte)。更大的单位以字节的倍数形成:千字节 (KB) 约 10³ 字节,兆字节 (MB) 约 10⁶ 字节,吉字节 (GB) 约 10⁹ 字节,太字节 (TB) 约 10¹² 字节。在某些上下文中会使用 2 的幂的二进制前缀 (kibibyte, mebibyte),但 OCR 考纲采用十进制定义。
You need to be able to calculate file sizes: e.g., a 5000‑byte file is 5 KB (5000 ÷ 1000). Always state the unit clearly.
你需要能够计算文件大小:例如,一个 5000 字节的文件是 5 KB (5000 ÷ 1000)。务必清晰注明单位。
4. Binary Addition | 二进制加法
Binary addition follows similar rules to denary addition, but with only two digits. The rules are: 0+0=0, 0+1=1, 1+0=1, 1+1=0 carry 1, 1+1+1=1 carry 1. When adding two 8‑bit binary numbers, work column by column from the right, carrying any overflow into the next column.
二进制加法遵循与十进制加法相似的规则,但只有两个数码。规则为:0+0=0, 0+1=1, 1+0=1, 1+1=0 进位 1, 1+1+1=1 进位 1。当两个 8 位二进制数相加时,从右向左逐列计算,将任何溢出进位到下一位。
Example: 01001101₂ + 00101001₂. After adding column by column, the result is 01110110₂. You must show the carries to gain full marks.
示例:01001101₂ + 00101001₂。逐列相加后结果为 01110110₂。必须展示进位过程才能获得满分。
5. Overflow in Binary Addition | 二进制加法中的溢出
Overflow occurs when the result of an addition requires more bits than the space available. In an 8‑bit register, if two numbers produce a sum that needs a 9th bit (a carry out of the most significant bit), this is an overflow error. The computer stores only the lower 8 bits, which gives an incorrect result.
当加法结果需要比可用空间更多的位时,就会发生溢出。在 8 位寄存器中,如果两个数相加产生的和需要第 9 位(最高有效位产生进位),这就是溢出错误。计算机只存储低 8 位,从而得到错误结果。
For example, adding 11111111₂ (255) and 00000001₂ (1) in an 8‑bit register yields 00000000₂ with a carry of 1, which is stored as 0. This overflow must be detected, and some processors set a flag to indicate it.
例如,在 8 位寄存器中将 11111111₂ (255) 与 00000001₂ (1) 相加,得到 00000000₂ 和进位 1,存储结果为 0。必须检测到这种溢出,某些处理器会设置标志位来指示。
6. Two’s Complement Representation | 二进制补码表示
Two’s complement is used to represent negative binary numbers. The most significant bit (MSB) acts as a sign bit: 0 for positive, 1 for negative. For an 8‑bit number, the range is -128 to +127. To find the two’s complement of a negative number: write the positive magnitude in binary, invert all bits (1’s complement), then add 1.
二进制补码用于表示负的二进制数。最高有效位 (MSB) 充当符号位:0 表示正,1 表示负。对于 8 位数,表示范围是 -128 到 +127。求一个负数的补码:先写出其正数的二进制形式,将所有位取反 (1 的补码),然后加 1。
Example: Represent -23 in 8‑bit two’s complement. Positive 23 is 00010111₂. Invert: 11101000₂. Add 1: 11101001₂. To convert a negative two’s complement number back to denary, note the negative sign, find its two’s complement again to get the magnitude, then apply the sign.
示例:用 8 位补码表示 -23。正数 23 为 00010111₂。取反:11101000₂。加 1:11101001₂。要将负的补码数转换回十进制,注意到负号后,再次求其补码得到绝对值,再加上负号。
Two’s complement simplifies binary arithmetic because subtraction can be done by adding the two’s complement of the subtrahend.
二进制补码简化了二进制运算,因为减法可以通过加上减数的补码来实现。
7. Character Encoding – ASCII and Unicode | 字符编码 – ASCII 与 Unicode
Characters are represented in binary using character sets. ASCII uses 7 bits, providing 128 codes for English letters, digits and symbols. Extended ASCII uses 8 bits (256 codes) to include additional characters like accented letters. Unicode was developed to cover characters from all writing systems, using either 16 bits (UTF‑16) or variable‑length encoding (UTF‑8).
字符通过字符集用二进制表示。ASCII 使用 7 位,为英文字母、数字和符号提供 128 个码位。扩展 ASCII 使用 8 位 (256 个码位),包括带重音字母等额外字符。Unicode 的开发旨在覆盖所有书写系统的字符,使用 16 位 (UTF‑16) 或变长编码 (UTF‑8)。
OCR expects you to understand the limitations of ASCII (no support for many languages) and the advantage of Unicode (universal coverage) as well as the trade‑off: Unicode files can be larger if using fixed‑width encoding.
OCR 要求你理解 ASCII 的局限性(不支持许多语言)和 Unicode 的优势(通用覆盖),以及权衡:若使用定长编码,Unicode 文件可能更大。
8. Representing Images | 图像表示
An image is represented as a grid of pixels (picture elements). Each pixel is assigned a binary code for its colour. The resolution is the number of pixels per row and column, e.g. 1920×1080. Higher resolution gives more detail but increases file size. The colour depth is the number of bits used per pixel; 1 bit gives 2 colours, 8 bits gives 256 colours, 24 bits (true colour) gives about 16.7 million colours.
图像表示为一个个像素 (picture elements) 组成的网格。每个像素被赋予一个表示其颜色的二进制码。分辨率是每行和每列的像素数量,例如 1920×1080。更高的分辨率带来更多细节,但会增加文件大小。色彩深度是每个像素使用的位数;1 位给出 2 种颜色,8 位给出 256 种颜色,24 位 (真彩色) 给出约 1670 万种颜色。
File size of an uncompressed image = resolution width × height × colour depth in bits. You may be asked to calculate this, converting bits to bytes or kilobytes. Metadata (e.g., date, camera settings) also adds to file size.
未压缩图像文件大小 = 分辨率宽 × 高 × 色彩深度 (以位计)。你可能会被要求计算,并将位转换为字节或千字节。元数据 (如日期、相机设置) 也会增加文件大小。
9. Representing Sound | 声音表示
Sound is analogue and must be converted to digital by sampling. The amplitude (loudness) of the sound wave is measured at regular intervals. The sample rate is the number of samples per second, measured in hertz (Hz). A higher sample rate captures more detail and gives better quality. The bit depth is the number of bits used to store each sample; 16‑bit audio can represent 65,536 amplitude levels.
声音是模拟信号,必须通过采样转换成数字信号。声波的振幅 (响度) 以固定的时间间隔进行测量。采样率是每秒采样次数,以赫兹 (Hz) 为单位。更高的采样率能获取更多细节,音质更好。位深度是存储每个样本所用的位数;16 位音频可表示 65 536 个振幅级别。
Uncompressed audio file size (bits) = sample rate × bit depth × duration in seconds × number of channels. For stereo, channels = 2. Typical CD quality uses 44.1 kHz sample rate, 16‑bit depth, stereo.
未压缩音频文件大小 (位) = 采样率 × 位深度 × 时长 (秒) × 声道数。立体声声道数为 2。典型 CD 音质使用 44.1 kHz 采样率、16 位深度、立体声。
10. Data Compression | 数据压缩
Compression reduces file size for storage and transmission. There are two main types: lossless and lossy. Lossless compression (e.g., ZIP, PNG, FLAC) reduces file size without losing any data; the original file can be perfectly reconstructed. It works by finding patterns and using run‑length encoding or dictionary encoding.
压缩可以减少文件大小以节省存储空间和传输带宽。主要有两种类型:无损压缩和有损压缩。无损压缩 (如 ZIP、PNG、FLAC) 在不丢失任何数据的情况下减小文件大小;原始文件可以完全重建。它通过寻找模式并使用行程长度编码或字典编码来工作。
Lossy compression (e.g., JPEG, MP3) permanently removes some data that is less noticeable to humans. This achieves much smaller file sizes but the original cannot be perfectly restored. It is acceptable for images and audio where perfect fidelity is not required.
有损压缩 (如 JPEG、MP3) 永久性移除人类不易察觉的某些数据。这可以大幅缩小文件大小,但无法完美复原原始数据。在不需要绝对保真的图像和音频中可以接受。
OCR exam questions often ask you to justify the choice of compression type for a given scenario, considering available bandwidth and quality requirements.
OCR 考试中常要求你针对特定场景合理选择压缩类型,并考虑带宽和质量需求。
Published by TutorHao | Computer Science Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导