📚 Waves 1.1.2 – Sound Part 2 | 波动 1.1.2 – 声音(二)
Sound, as an artistic medium, extends far beyond the acoustic vibrations explored in Part 1. Here we dive into the digital representation, manipulation, and spatial contextualisation of sound — essential knowledge for artists working with audio installations, multimedia performance, and sonic sculpture. Understanding how sound is captured, stored, synthesised, and perceived in space enables creative practitioners to sculpt immersive experiences that challenge the boundaries between science and art.
声音作为一种艺术媒介,远远超越了第一部分中探讨的声学振动。这里我们深入声音的数字化表示、操控及其空间语境化——这是从事音频装置、多媒体表演和声音雕塑创作的艺术家必备的知识。理解声音如何被捕捉、存储、合成及在空间中被感知,能使创作者塑造出挑战科学与艺术边界的沉浸式体验。
1. Digital Sound Representation | 数字声音表示
In digital art and music production, sound must be converted from continuous analogue pressure waves into discrete numerical data. This process, called analogue-to-digital conversion (ADC), involves measuring the amplitude of the waveform at regular intervals and quantising those values into binary code.
在数字艺术和音乐制作中,声音必须从连续的模拟压力波转换为离散的数字数据。这个过程称为模数转换 (ADC),包括以规整间隔测量波形的振幅,并将这些值量化为二进制代码。
The resulting digital signal is a stream of samples that, when reconstructed, can reproduce the original sound with remarkable fidelity. The quality of this representation depends on two critical parameters: sample rate and bit depth, which we will examine in the next section.
产生的数字信号是一串样本流,当被重建时,可以以极高的保真度再现原始声音。这种表示的质量取决于两个关键参数:采样率和位深度,我们将在下一节中详述。
2. Sample Rate and Bit Depth | 采样率与位深度
Sample rate defines how many times per second the amplitude is captured, measured in hertz (Hz) or kilohertz (kHz). According to the Nyquist–Shannon theorem, the sample rate must be at least twice the highest frequency to be reproduced. For human hearing (20 Hz to 20 kHz), a 44.1 kHz or 48 kHz sample rate is standard.
采样率定义了每秒捕捉振幅的次数,单位是赫兹 (Hz) 或千赫 (kHz)。根据奈奎斯特–香农定理,采样率必须至少是要再现最高频率的两倍。对于人类听觉(20 Hz 到 20 kHz),44.1 kHz 或 48 kHz 的采样率是标准。
Bit depth determines the number of possible amplitude values each sample can take. A 16‑bit system offers 65,536 levels, while 24‑bit provides over 16 million, dramatically reducing quantisation noise. In art installations, higher bit depths allow for greater dynamic range, preserving subtle details in soft passages or extreme crescendos.
位深度决定了每个样本可以取的振幅值的数量。16 位系统提供 65,536 个级别,而 24 位则提供超过 1600 万个级别,大大降低了量化噪声。在艺术装置中,更高的位深度可实现更大的动态范围,保留轻柔段落或极端渐强中的微妙细节。
-
CD quality: 44.1 kHz, 16‑bit stereo.
CD 质量:44.1 kHz,16 位立体声。
-
Professional audio: 48 kHz, 24‑bit or 96 kHz, 24‑bit for high‑resolution projects.
专业音频:48 kHz,24 位或 96 kHz,24 位用于高分辨率项目。
3. Synthesis Fundamentals | 合成基础
Electronic sound synthesis is a core technique in sound art, allowing artists to generate timbres impossible to produce acoustically. Common synthesis methods include subtractive, additive, FM (frequency modulation), and granular synthesis. Each starts with fundamental waveforms — sine, square, triangle, and sawtooth — which differ in harmonic content.
电子声音合成是声音艺术的核心技术,使艺术家能够产生声学上无法实现的音色。常见的合成方法包括减法、加法、FM(频率调制)和粒子合成。每种方法都从基本波形开始——正弦波、方波、三角波和锯齿波——它们在谐波含量上各不相同。
Subtractive synthesis filters a harmonically rich waveform to sculpt the desired tone. Additive synthesis builds complex sounds by layering multiple sine waves at integer multiples of a fundamental frequency. Artists like Ryoji Ikeda use raw sine tones and additive principles to create minimalist, data‑driven sonic sculptures that explore the aesthetics of pure frequency.
减法合成通过过滤谐波丰富的波形来塑造所需的音调。加法合成则通过在基频的整数倍上叠加多个正弦波来构建复杂的声音。艺术家如池田亮司使用原始正弦音和加法原理创造极简的、数据驱动的声音雕塑,探索纯频率的美学。
4. Harmonics, Overtones, and Timbre | 谐波、泛音与音色
Any sound can be described by its fundamental frequency (pitch) and its overtones — higher frequencies that are integer (harmonics) or non‑integer (partials) multiples of the fundamental. The specific mixture and amplitude envelope of these overtones create the unique timbre or “colour” of a sound.
任何声音都可以通过其基频(音高)和泛音来描述——泛音是基频的整数倍(谐波)或非整数倍(分音)的高频部分。这些泛音的特定组合和振幅包络创造了声音独特的音色或“色彩”。
In sound art, spectral manipulation — enhancing, suppressing, or dislocating overtones — allows for radical transformations of recorded material. A violin note can be made to sound like a bell or a metallic screech through real‑time spectral processing, a technique widely used in interactive installations by artists such as Carsten Nicolai (Alva Noto).
在声音艺术中,频谱操控——增强、抑制或错位泛音——允许对录制材料进行彻底变形。通过实时频谱处理,一个小提琴音符可以被处理成钟声或金属尖啸,这是艺术家如卡斯滕·尼古拉(艺名 Alva Noto)在互动装置中广泛使用的技术。
5. Envelopes and Temporal Shaping | 包络与时间塑形
A sound’s amplitude envelope describes how its volume changes over time, typically broken into the ADSR model: Attack, Decay, Sustain, Release. In artistic contexts, modifying the envelope is as expressive as pitch choice. A sharp attack creates percussive immediacy; a long, swelling attack evokes ethereal, evolving textures.
声音的振幅包络描述了其音量如何随时间变化,通常分为 ADSR 模型:起音 (Attack)、衰减 (Decay)、延持 (Sustain)、释音 (Release)。在艺术语境中,修改包络与音高选择一样富有表现力。尖锐的起音创造出打击乐的即时感;悠长的、渐强的起音则唤起空灵、演变的质感。
Time‑stretching — altering a sound’s duration without changing its pitch — and its opposite, pitch‑shifting without time change, are common in narrative soundscapes. These tools allow artists to stretch a single vocal syllable into a minute‑long drone or shift environmental noises into musical chords, blurring the line between sound and music.
时间拉伸——在不改变音高的情况下改变声音的持续时间——及其相反的操作,即不改变时间而改变音高,在叙事性音景中很常见。这些工具使艺术家能将单个语音音节拉伸成一分钟长的嗡嗡声,或将环境噪声转换为音乐和弦,模糊了声音与音乐之间的界限。
6. Spatial Hearing and Localisation | 空间听觉与定位
Human spatial hearing relies on several cues: interaural time difference (ITD), interaural level difference (ILD), and spectral filtering caused by the pinna (outer ear). ITD is dominant for frequencies below 1.5 kHz, while ILD becomes primary above 3 kHz. This knowledge is essential for designing 3‑D audio installations.
人类空间听觉依赖于几种线索:双耳时间差 (ITD)、双耳声级差 (ILD) 以及由耳廓(外耳)引起的频谱滤波。ITD 在低于 1.5 kHz 的频率上占主导,而 ILD 在高于 3 kHz 时成为主要线索。这些知识对于设计三维音频装置至关重要。
Artists can artificially recreate auditory space using binaural recording (a dummy head with microphones in the ear canals) or Ambisonics, a full‑sphere surround sound technique. Janet Cardiff’s “Forty‑Part Motet” uses a 40‑speaker array to place the listener inside a choir, making spatial position part of the artistic narrative.
艺术家可以使用双耳录音(一个带有耳道麦克风的仿真人头)或全场环绕声技术 Ambisonics 来人为地重建听觉空间。珍妮特·卡迪夫的《四十声部经文歌》使用 40 个扬声器阵列将听众置于合唱团内部,使空间位置成为艺术叙事的一部分。
7. Ambisonics and Wave Field Synthesis | Ambisonics 与波场合成
Ambisonics encodes sound direction and pressure into a mathematical representation (often using spherical harmonics), allowing rotation and decoding to arbitrary speaker layouts. First‑order Ambisonics captures three directional components plus omnidirectional pressure; higher‑order Ambisonics (HOA) increases spatial resolution dramatically.
Ambisonics 将声音方向和声压编码为数学表示(通常使用球谐函数),允许旋转并解码为任意扬声器布局。一阶 Ambisonics 捕捉三个方向分量加上全向声压;高阶 Ambisonics (HOA) 极大地提高了空间分辨率。
Wave Field Synthesis (WFS) recreates a sound field over a large area using hundreds of closely spaced loudspeakers. It reproduces the wavefront as if the original source were present, allowing multiple listeners to walk through the space without losing the spatial image. This technology has been used in large‑scale immersive art by groups like ZKM in Karlsruhe.
波场合成 (WFS) 使用数百个紧密排列的扬声器在大范围内重建声场。它再现了波前,就像原始声源在场一样,允许多个听众在空间中走动而不丢失空间影像。这项技术已被卡尔斯鲁厄的艺术与媒体中心 (ZKM) 等团体用于大型沉浸式艺术。
8. Psychoacoustics and Perceptual Tricks | 心理声学与感知技巧
Psychoacoustics explores how the brain processes sound, revealing phenomena that artists exploit for illusion and impact. The Shepard tone creates the perception of an ever‑ascending (or descending) pitch loop, used in films and sonic installations to evoke infinite progression or tension.
心理声学探索大脑如何处理声音,揭示了艺术家用于制造幻觉和冲击力的现象。谢泼德音调创造出一种不断上升(或下降)的音高循环的错觉,被用于电影和声音装置中,以唤起无限的进展或紧张感。
The precedence effect (or law of the first wavefront) helps localisation in reverberant spaces: if two identical sounds arrive within a short window (~1–40 ms), the brain fuses them and localises based on the first arrival. Artists can use this to steer attention or create phantom sources between real speakers.
优先效应(或第一波前定律)有助于在混响空间中进行定位:如果两个相同的声音在短时间内(大约 1–40 毫秒)到达,大脑会将它们融合,并根据最先到达的声音进行定位。艺术家可以利用这一点来引导注意力或在真实扬声器之间创造幻象声源。
| Phenomenon 现象 | Artistic application 艺术应用 |
| Auditory masking 听觉掩蔽 | Sculpting dense textures by hiding sounds 通过隐藏声音来塑造密集的质感 |
| Binaural beats 双耳节拍 | Inducing meditative states in interactive environments 在互动环境中诱导冥想状态 |
| Fletcher–Munson curves 弗莱彻–蒙森曲线 | Designing frequency balance for different listening volumes 为不同听音音量设计频率平衡 |
9. Acoustic Ecology and Soundscape Composition | 声学生态与音景创作
Acoustic ecology, pioneered by R. Murray Schafer, studies the relationship between living beings and their sonic environment. Soundscape composition treats environmental sound as primary material, arranging field recordings to reflect or critique ecological, social, and political narratives.
由 R. 默里·谢弗开创的声学生态学,研究生物与其声音环境之间的关系。音景创作将环境声音作为主要材料,编排实地录音以反映或批判生态、社会和政治叙事。
The concept of keynote sounds, sound signals, and soundmarks helps artists read a landscape sonically. A keynote sound is the background ambient (e.g., wind or traffic), signals are foregrounded warnings (sirens), and soundmarks are unique sounds anchored to a community (church bells). Composers like Hildegard Westerkamp blend these elements into narrative-driven soundwalks.
基调音、声音信号和声音地标的概念帮助艺术家从声音上解读一个地域。基调音是背景环境(例如风声或交通声),信号是前景化的警告(警笛),而声音地标是与某个社区紧密相连的独特声音(教堂钟声)。作曲家如希尔德加德·韦斯特坎普将这些元素融合成叙事驱动的声音漫步。
10. Interactive Audio Systems | 交互式音频系统
Modern sound art frequently involves interactivity, where sensors—cameras, microphones, pressure pads, or motion detectors—allow the audience to influence real‑time sound synthesis and processing. Max/MSP, Pure Data, and SuperCollider are common programming environments for such responsive systems.
现代声音艺术经常涉及交互性,通过传感器——摄像头、麦克风、压力垫或运动检测器——让观众影响实时的声音合成和处理。Max/MSP、Pure Data 和 SuperCollider 是这类响应系统常用的编程环境。
Mapping human movement to sonic parameters requires careful design: a visitor’s proximity might control reverb amount, while their speed alters playback rate. In Rafael Lozano‑Hemmer’s “Pulse Room,” participants’ heartbeats are converted into rhythmic flashes and synchronised sound, turning the gallery into a living, collective polyrhythm.
将人体运动映射到声音参数需要仔细设计:访客的距离可能控制混响量,而他们的速度则改变播放速率。在拉斐尔·洛萨诺-赫默的《脉搏室》中,参与者的心跳被转换为节奏性的闪光和同步声音,将画廊变成了有生命的集体复节奏。
11. Multisensory Integration: Sound and Vision | 多感官整合:声音与视觉
A central concern in contemporary sound art is the integration of auditory and visual stimuli. Artists often create visual scores, real‑time waveform projections, or light‑responsive sound events. The correspondence between colours and frequency, while subjective, can be systematically mapped: many practitioners link lower frequencies with darker, redder hues, and higher frequencies with brighter, bluer tones.
当代声音艺术的一个核心关注是听觉与视觉刺激的整合。艺术家经常创作视觉化乐谱、实时波形投影或光响应声音事件。颜色与频率之间的对应关系虽然具有主观性,但可以系统地映射:许多实践者将低频与较暗、偏红的色调关联,将高频与较亮、偏蓝的色调关联。
The study of chromesthesia, a form of synesthesia where sound involuntarily evokes colour perception, informs audiovisual works. Artists like Kandinsky (who painted musical compositions) and contemporary figures such as Ryoichi Kurokawa create audiovisual installations where sound and image are co‑dependent, neither subservient to the other.
色联觉(声响引发不自主色彩感知的一种联觉)的研究为视听作品提供了理论支持。从康定斯基(他画音乐作品)到当代人物如黑川良一,艺术家们创造的视听装置中声音与图像相互依存,彼此从属。
12. Preserving Sonic Artworks | 保存声音艺术作品
The ephemeral nature of sound presents unique challenges for conservation in the art world. Unlike a canvas, a sonic artwork may depend on specific hardware, software, and spatial arrangement that can become obsolete. Documentation must capture not only the audio files but also the performance instructions, equipment specifications, and room acoustics.
声音的瞬时性给艺术界的保存带来了独特挑战。与画布不同,声音艺术作品可能依赖于特定的硬件、软件和空间安排,而这些可能变得过时。记录不仅必须捕获音频文件,还要包括表演说明、设备规格和房间声学。
Formats like the Audio Engineering Society’s AES69 (SOFA) standard preserve spatial audio information for future playback. Institutions now treat sonic installations as time-based media, requiring detailed technical riders, video walk‑throughs, and interviews with the artist to ensure the work can be re‑staged decades later without losing its intended meaning.
像音频工程学会的 AES69 (SOFA) 标准这样的格式,为未来的回放保存空间音频信息。机构现在将声音装置视为时基媒体,要求详细的技术附加条款、视频导览以及与艺术家的访谈,以确保作品在数十年后能够重新演出而不失其本意。
Published by TutorHao | Art Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导