Data Representation | 数据表示 考点精讲

📚 Data Representation | 数据表示 考点精讲

Data representation forms the foundation of every computer system. In the Edexcel A‑Level Computer Science syllabus, mastering how numbers, text, images, and sound are encoded in binary is essential for understanding processors, memory, and data communication. This guide covers all key topics, from basic binary arithmetic to floating‑point normalisation and error detection, with clear explanations, worked examples, and exam‑focused tips.

数据表示是每一套计算机系统的基石。在 Edexcel A‑Level 计算机科学大纲中,掌握数字、文字、图像和声音如何用二进制编码,是理解处理器、内存和数据通信的关键。本指南涵盖从基础二进制算术到浮点规格化和差错检测的所有核心主题,配有清晰的讲解、例题和应试技巧。

1. Number Bases: Binary, Denary and Hexadecimal | 数制基础:二进制、十进制与十六进制

The three number systems you must be fluent in are binary (base 2), denary (base 10), and hexadecimal (base 16). Binary uses only the digits 0 and 1; each position represents a power of 2. Hexadecimal condenses four binary digits into a single symbol (0‑9, A‑F), making long binary strings easier to read and write. Conversions between these bases are frequently examined, either directly or as part of larger questions on memory addresses or colour codes.

你必须熟练掌握的三种数制是二进制(基数为2)、十进制(基数为10)和十六进制(基数为16)。二进制只使用数字0和1,每一位代表2的幂。十六进制将四个二进制位压缩为一个符号(0‑9、A‑F),使长二进制串更易读写。这三种进制之间的转换经常被考查,无论是直接设问还是作为内存地址或颜色编码等大题的一部分。

  • Binary to Denary: Multiply each bit by its place value and sum. For example, 1011₂ = 1×8 + 0×4 + 1×2 + 1×1 = 11₁₀.
  • 二进制转十进制:每位乘以位权后求和。例如,1011₂ = 1×8 + 0×4 + 1×2 + 1×1 = 11₁₀。
  • Denary to Binary: Repeatedly divide by 2 and record remainders. 13₁₀ → 13 ÷ 2 = 6 r1; 6 ÷ 2 = 3 r0; 3 ÷ 2 = 1 r1; 1 ÷ 2 = 0 r1 → 1101₂.
  • 十进制转二进制:反复除以2并记录余数。13₁₀ → 13 ÷ 2 = 6 余1;6 ÷ 2 = 3 余0;3 ÷ 2 = 1 余1;1 ÷ 2 = 0 余1 → 1101₂。
  • Hexadecimal to Binary: Replace each hex digit with its 4‑bit equivalent. 3A₁₆ → 0011 1010₂. Do not omit leading zeros when forming the 4‑bit groups.
  • 十六进制转二进制:将每一位十六进制数字替换为对应的4位二进制。3A₁₆ → 0011 1010₂。分组时不要省略前导零。

2. Binary Arithmetic: Addition and Overflow | 二进制算术:加法与溢出

Binary addition follows the same principles as denary addition, but with only four rules: 0+0=0, 0+1=1, 1+0=1, 1+1=0 carry 1. When two unsigned binary numbers are added, a carry may propagate through the columns. If the final carry exceeds the available bit width, an overflow occurs. Overflow must be detected and handled, especially in fixed‑width registers, because it can lead to incorrect results.

二进制加法遵循与十进制加法相同的原则,但只有四条规则:0+0=0、0+1=1、1+0=1、1+1=0进1。当两个无符号二进制数相加时,进位可能会跨列传播。如果最终的进位超出了可用的位宽,就会发生溢出。溢出必须被检测和处理,尤其是在定宽寄存器中,因为它会导致错误的结果。

Example: 0101₂ (5) + 0111₂ (7) = 1100₂ (12). No overflow in 4 bits.

示例:0101₂ (5) + 0111₂ (7) = 1100₂ (12),在4位内无溢出。

Example with overflow: 1101₂ (13) + 1010₂ (10) = 1 0111₂ (23) but the 5th bit is lost in a 4‑bit register → result 0111₂ (7), incorrect.

溢出示例:1101₂ (13) + 1010₂ (10) = 1 0111₂ (23),但在4位寄存器中丢失第5位 → 结果为0111₂ (7),不正确。

In Edexcel exams, you may be asked to add binary numbers and comment on whether overflow has occurred by referring to the carry into and out of the most significant bit. Always check the carry flags.

在Edexcel考试中,可能会要求你将二进制数相加,并通过最高位的进位输入和输出来判断是否发生溢出。务必检查进位标志。


3. Binary Multiplication and Division: Logical Shifts | 二进制乘除:逻辑移位

Multiplying or dividing a binary number by powers of two can be achieved through bit shifts. A left logical shift moves every bit one place to the left, fills the rightmost bit with 0, and discards the leftmost bit that falls off. This operation multiplies the value by 2 for each shift. A right logical shift divides the value by 2 per shift, truncating any fractional part. These shifts are extremely fast in hardware and are used for efficient multiplication and division.

二进制数乘以或除以2的幂可以通过移位实现。逻辑左移将每一位向左移动一位,最右侧填0,最左侧移出的位被丢弃。每次左移相当于乘以2。逻辑右移每次将数值除以2,丢弃小数部分。移位在硬件中速度极快,常用于高效的乘除运算。

Left shift (×2): 0011₂ (3) << 1 → 0110₂ (6). Left shift twice (×4): 0011₂ → 1100₂ (12).

左移(×2):0011₂ (3) << 1 → 0110₂ (6)。左移两次(×4):0011₂ → 1100₂ (12)。

It is important to distinguish logical shifts from arithmetic shifts, which preserve the sign bit in two’s complement numbers. Logical shifts treat the bit pattern as a pure unsigned integer, while arithmetic right shifts replicate the sign bit to maintain the number’s sign.

重要的是区分逻辑移位与算术移位,后者在二进制补码数中保留符号位。逻辑移位将位模式视为纯无符号整数,而算术右移则复制符号位以保持数值的符号。


4. Two’s Complement Representation | 二进制补码表示

Computers use two’s complement to represent signed integers. In an n‑bit system, the most significant bit (MSB) carries the negative weight of −2ⁿ⁻¹. This allows a single representation for zero and simplifies subtraction to addition of a negative number. The range of an n‑bit two’s complement integer is from −2ⁿ⁻¹ to 2ⁿ⁻¹ − 1. For example, an 8‑bit two’s complement number ranges from −128 to +127.

计算机使用补码来表示有符号整数。在一个n位系统中,最高有效位(MSB)携带的负权为 −2ⁿ⁻¹。这种表示法为零提供了唯一的表示,并将减法简化为加一个负数。n位补码整数的范围是从 −2ⁿ⁻¹ 到 2ⁿ⁻¹ − 1。例如,8位补码数的范围是 −128 到 +127。

To find the two’s complement of a positive number, write the positive binary and then invert all bits (one’s complement) and add 1. To find the denary equivalent of a negative two’s complement number, the same procedure is applied backwards: recognise the sign from the MSB, perform two’s complement conversion to obtain the positive magnitude, and then apply the negative sign.

要找出一个正数的补码,先写出正数的二进制,然后将所有位取反(反码)再加1。要找出负数补码对应的十进制值,可以反向操作:根据MSB识别符号,执行补码转换得到正数大小,再加上负号。

Example: Represent −5 in 4‑bit two’s complement. +5 = 0101, invert → 1010, add 1 → 1011. Thus −5 is 1011₂.

示例:用4位补码表示 −5。+5 = 0101,取反 → 1010,加1 → 1011。因此 −5 为 1011₂。

One common pitfall is forgetting that the positive representation must have a leading zero to fit into the given number of bits. Always check that the most significant bit correctly indicates the sign.

一个常见的误区是忘记正数的表示必须有前导零,以适应给定的位数。始终检查最高有效位是否正确表示了符号。


5. Floating Point Representation: Mantissa and Exponent | 浮点数表示:尾数与阶码

Real numbers are stored in binary floating point format, consisting of a mantissa (fractional part) and an exponent. The mantissa holds the significant digits, while the exponent scales the number, analogous to scientific notation. In Edexcel, you will encounter a fixed‑size register split into mantissa and exponent fields, both using two’s complement where appropriate. The general form is mantissa × 2ᵉˣᵖᵒⁿᵉⁿᵗ, with the binary point placed just after the sign bit of the mantissa.

实数以二进制浮点格式存储,由尾数(小数部分)和阶码组成。尾数保存有效数字,阶码对数值进行缩放,类似于科学记数法。在Edexcel考试中,你会遇到将定宽寄存器划分为尾数域和阶码域的情况,两者在适当时候均使用补码。一般形式为 尾数 × 2ᵉˣᵖᵒⁿᵉⁿᵗ,小数点位于尾数符号位之后。

For example, a 12‑bit representation with an 8‑bit mantissa (including sign) and a 4‑bit exponent (including sign) might store the value as: mantissa = 0.1101000 (denary 0.8125), exponent = 0010 (denary +2). The number is 0.8125 × 2² = 3.25. The actual bit pattern would need to be interpreted carefully, using two’s complement for negative exponents and negative mantissas.

例如,一个12位表示,其中8位尾数(含符号)和4位阶码(含符号),可能存储的值为:尾数 = 0.1101000(十进制0.8125),阶码 = 0010(十进制+2)。该数为 0.8125 × 2² = 3.25。实际位模式需要仔细解读,对于负阶码和负尾数要使用补码。

Converting a denary number to floating point involves normalising the binary number, determining the exponent, and packing the bits into the specified format. This topic is frequently tested with a completed example and then a request to fill in the blanks or derive a value.

将十进制数转换为浮点数涉及对二进制数进行规格化、确定阶码,并将位填充到指定格式中。这个主题经常被考查,给出一个完整的示例,然后要求填空或推导数值。


6. Normalisation of Floating Point Numbers | 浮点数的规格化

Normalisation is the process of adjusting the mantissa so that its first two bits are different (01 for positive, 10 for negative) after the sign bit. A normalised representation maximises precision by ensuring no leading zeros or useless sign‑extension bits remain in the mantissa. For a positive mantissa, the normalised form starts with 0.1; for a negative mantissa, with 1.0. This rule is vital because it guarantees a unique representation for each number and avoids wasting bit storage.

规格化是调整尾数,使得符号位后两位不同的过程(正数为01,负数为10)。规格化表示通过确保尾数中无前导零或无效的符号扩展位来最大化精度。对于正尾数,规格化形式以0.1开头;对于负尾数,以1.0开头。这条规则至关重要,因为它保证了每个数的唯一表示,避免浪费位存储空间。

Example: The mantissa 0.0001101 is not normalised because it has leading zeros. Shifting left three places gives 0.1101000, and the exponent must be decreased by 3 to compensate.

示例:尾数0.0001101未规格化,因为它有前导零。左移三位得到0.1101000,阶码必须减去3以补偿。

When arithmetic operations produce an unnormalised mantissa, a re‑normalisation step is needed. This may involve shifting the mantissa left or right and adjusting the exponent accordingly. Normalisation errors are common in exam answers; always verify that the first two bits differ.

当算术运算产生未规格化的尾数时,需要一个重新规格化的步骤。这可能涉及将尾数左移或右移,并相应地调整阶码。规格化错误在考试答案中很常见,务必验证前两位不同。


7. Character Encoding: ASCII and Unicode | 字符编码:ASCII与Unicode

Characters are represented inside the computer using numeric codes. The American Standard Code for Information Interchange (ASCII) originally used 7 bits to represent 128 characters, including control codes, numerals, uppercase and lowercase letters, and common punctuation. Extended ASCII uses 8 bits (1 byte) for 256 characters, accommodating additional symbols and accented characters for Western languages. Unicode was developed to support the thousands of characters needed for global scripts and emojis; UTF‑8 and UTF‑16 are variable‑length encodings that maintain compatibility with ASCII while allowing representation of over a million code points.

字符在计算机内部用数字编码表示。美国信息交换标准码(ASCII)最初使用7位表示128个字符,包括控制码、数字、大小写字母和常用标点。扩展ASCII使用8位(1字节)表示256个字符,可容纳额外的符号和西方语言的重音字符。Unicode的开发是为了支持全球文字和表情符号所需的成千上万个字符;UTF‑8和UTF‑16是变长编码,既保持与ASCII的兼容,又能表示超过一百万个码点。

Key exam points: know the ASCII codes for ‘A’ (65), ‘a’ (97), ‘0’ (48), and space (32). Understand that Unicode provides a unique code point for every character, while the encoding form (UTF‑8, UTF‑16) defines how those code points are stored as bytes. The advantage of Unicode is its ability to represent all languages, reducing problems with legacy character sets.

考试重点:记住‘A’(65)、‘a’(97)、‘0’(48)和空格(32)的ASCII码。理解Unicode为每一个字符提供唯一的码点,而编码形式(UTF‑8、UTF‑16)定义了这些码点如何存储为字节。Unicode的优势在于能表示所有语言,减少了遗留字符集带来的问题。


8. Bitmaps and Image Representation | 位图与图像表示

A bitmap image is represented as a grid of pixels, each assigned a binary colour value. The number of bits per pixel (colour depth) determines how many colours can be encoded. For example, with 1 bit per pixel, only black and white are possible; with 8 bits, 256 colours or shades of grey can be represented. The resolution of an image is the number of pixels (width × height), directly impacting file size and detail. Monochrome bitmaps are often examined where you calculate the storage required from the grid dimensions.

位图图像表示为一个像素网格,每个像素被赋予一个二进制颜色值。每个像素的位数(颜色深度)决定可以编码多少种颜色。例如,每像素1位只能表示黑白;每像素8位可以表示256种颜色或灰度。图像的分辨率是像素数量(宽×高),直接影响文件大小和细节。单色位图经常被考查,要求根据网格尺寸计算所需存储空间。

Computing file size: file size (in bytes) = (width × height × colour depth) / 8. Metadata such as header information (file type, dimensions, colour table) must also be included, increasing the actual size. In Edexcel papers, you may be asked to calculate the number of bits required for a given image or to determine the colour depth needed for a certain range of colours.

计算文件大小:文件大小(字节)=(宽 × 高 × 颜色深度)/ 8。元数据,如头部信息(文件类型、尺寸、颜色表),也必须包含在内,这会增加实际大小。在Edexcel试卷中,可能会要求计算给定图像所需的位数,或确定某种颜色范围所需的颜色深度。


9. Sound Representation: Sampling | 声音表示:采样

Sound is an analogue waveform that must be digitised for storage and processing. Sampling measures the amplitude of the wave at discrete time intervals. The sample rate (measured in Hz) is the number of samples taken per second; higher rates capture higher frequencies, adhering to the Nyquist theorem (sample rate ≥ 2× highest frequency). The sample resolution (bit depth) determines the number of possible amplitude values—16 bits allow 65,536 levels, delivering CD‑quality audio.

声音是一种模拟波形,必须进行数字化才能存储和处理。采样是按离散的时间间隔测量波的振幅。采样率(单位为Hz)是每秒采集的样本数;采样率越高,可捕获的频率越高,遵循奈奎斯特定理(采样率 ≥ 2×最高频率)。采样分辨率(位深度)决定可能的振幅值数量——16位允许65,536个级别,可提供CD品质的音频。

Typical calculations involve determining the bit rate (bits per second) = sample rate × sample resolution × number of channels. For example, mono sound at 44.1 kHz with 16‑bit resolution produces a bit rate of 705,600 bps. Storage for a given duration = bit rate × time. You should be comfortable converting between bits, bytes, kilobytes, and megabytes using powers of two.

典型的计算涉及确定比特率(每秒位数)= 采样率 × 采样分辨率 × 通道数。例如,44.1 kHz、16位分辨率的单声道声音产生705,600 bps的比特率。给定持续时间的存储量 = 比特率 × 时间。你应熟练使用2的幂在比特、字节、千字节和兆字节之间进行转换。


10. Data Compression: Lossy and Lossless | 数据压缩:有损与无损

Compression reduces the number of bits needed to represent data, saving storage space and transmission time. Lossless compression achieves size reduction without any loss of original data; run‑length encoding (RLE) and dictionary‑based methods are common examples. RLE replaces repeated bytes with a count and the byte value. Lossy compression removes some data that is deemed less noticeable to humans, as in JPEG for images and MP3 for audio, achieving much higher compression ratios but irreversible loss of quality.

压缩可减少表示数据所需的比特数,节省存储空间和传输时间。无损压缩在减少大小的同时完全不丢失原始数据;游程编码(RLE)和基于字典的方法是常见例子。RLE将重复字节替换为计数值和字节值。有损压缩则移除一些对人类来说不太明显的数据,例如图像的JPEG和音频的MP3,可以实现高得多的压缩比,但质量损失不可逆。

In exams, be ready to apply RLE to simple bitmap patterns, or to explain why a file consisting of a long sequence of identical characters compresses well. Understand the trade‑offs: lossless is essential for text and executable files; lossy is acceptable for multimedia where perfect reproduction is not required.

在考试中,准备对简单的位图模式应用RLE,或解释为什么由长串相同字符组成的文件压缩效果好。理解权衡:无损压缩对文本和可执行文件至关重要;有损压缩适用于不需要完全还原的多媒体。


11. Error Detection and Correction: Parity and Checksums | 差错检测与纠正:奇偶校验与校验和

During data transmission or storage, bits can be flipped due to noise or hardware faults. Parity checking adds a single parity bit to a block of data to make the total number of 1s either even (even parity) or odd (odd parity). The receiver can detect single‑bit errors (and some multi‑bit errors) by re‑computing the parity. Parity is simple but cannot correct errors or detect all multi‑bit errors. Checksums extend this concept by summing blocks of data and attaching the sum; the receiver re‑calculates the checksum and compares.

在数据传输或存储过程中,比特可能因噪声或硬件故障而翻转。奇偶校验为数据块添加一个奇偶校验位,使1的总数为偶数(偶校验)或奇数(奇校验)。接收方通过重新计算奇偶性,可检测出单比特错误(以及某些多比特错误)。奇偶校验简单,但无法纠正错误,也不能检测所有多比特错误。校验和将这一概念扩展,对数据块求和并附加该和值;接收方重新计算校验和并进行比较。

More advanced techniques such as CRC (Cyclic Redundancy Check) use polynomial division to generate a checksum that can detect burst errors. The Edexcel specification focuses on basic parity and checksum methods; you should be able to calculate a parity bit or a simple arithmetic checksum from given byte values.

更先进的技术,如循环冗余校验(CRC),使用多项式除法生成能检测突发错误的校验和。Edexcel规范侧重于基本的奇偶校验和校验和方法;你应能根据给定的字节值计算奇偶校验位或简单的算术校验和。


12. Exam Tips and Common Pitfalls | 考试技巧与常见误区

Data representation questions reward precision and methodical working. Always show your steps when converting between bases, performing binary addition, or normalising floating point numbers. A frequent mistake is misaligning bits when extracting the mantissa and exponent fields; draw a vertical line to separate the two parts in your working. When calculating image or sound file sizes, check whether the question expects the result in bits, bytes, or multiples of kilobytes; remember that file size conversions use 1024, not 1000, unless specified.

数据表示类的题目奖赏精准和有条理的解题过程。在进制转换、二进制加法或浮点数规格化时,务必展示每一个步骤。一个常见的错误是在提取尾数和阶码域时没有对齐位;解题时画一条竖线将两部分分开。在计算图像或声音文件大小时,检查题目要求结果是以位、字节还是千字节为单位;记住,文件大小换算使用1024而非1000,除非另有说明。

  • In two’s complement arithmetic, watch the MSB carefully: a carry into the MSB that differs from the carry out indicates overflow for signed numbers.
  • 在二进制补码算术中,密切关注MSB:进入MSB的进位与从MSB输出的进位不同,表示有符号数溢出。
  • When shifting, be explicit about whether it is a logical or arithmetic shift; logical shifts are suitable for unsigned integers and bitwise operations, while arithmetic shifts preserve sign.
  • 进行移位操作时,明确是逻辑移位还是算术移位;逻辑移位适用于无符号整数和按位操作,而算术移位保留符号。
  • In normalisation, the mantissa must always be in the form 0.1xxx for positive or 1.0xxx for negative. Adjusting the exponent by the number of shifts left or right is essential and often forgotten.
  • 在规格化中,尾数必须始终是正数为0.1xxx、负数为1.0xxx的形式。通过左移或右移的位数调整阶码至关重要,却经常被遗忘。
  • Practice past paper questions on data compression and encoding; examiners often ask you to identify whether a given scenario describes lossy or lossless compression and justify your answer.
  • 练习历年真题中关于数据压缩和编码的题目;考官常要求你判断给定场景描述的是有损还是无损压缩,并给出理由。

Published by TutorHao | Computer Science Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading