Digital audio is a list of numbers measured at a rate, each one stored with some precision. This reading steps back from any particular software to talk about what edits actually do to sound, and to introduce a vocabulary you'll reach for constantly: the envelope of a sound. That vocabulary is how you describe the sounds you want, find what you need in a sample library, and shape what you make.
Every sound has a shape over time. It starts somewhere, exists for some duration, and ends. The way it does each of these things is part of what makes the sound recognizable, and a property your ears use almost unconsciously to identify what you're hearing.
Consider three different sounds:
These sounds have very different shapes over time. We call this shape the envelope of the sound, and we describe it using three stages: attack, sustain, and release.
The attack is how a sound begins. A wood block struck with a mallet has a sharp attack: the sound goes from silence to full volume almost instantly. A bowed violin note has a slow attack: the bow catches the string gradually and the sound swells in. A note swelled with a volume pedal can have an even slower attack, taking several seconds to reach full volume.
Attack is one of the cues your ear uses to identify what kind of sound you're hearing. A piano and a violin playing the same note at the same volume sound completely different, largely because their attacks are different. Reverse a piano note (take the recording and play it backward) and it stops sounding like a piano almost entirely, because the attack now has the shape of a piano release.
The sustain is what happens during the body of the sound, after the attack and before the release. Some sounds have steady sustains: an organ note holds at roughly the same volume and timbre for as long as the key is pressed. Some have evolving sustains: a vocal "ahhh" wavers and shifts, or a bowed string note has subtle changes in pressure and bow speed that show up in the sound.
Many sounds have very short sustains, or none at all. A wood block has effectively no sustain, there's an attack and then almost immediately the release. A handclap is similar. Pure percussion is mostly attack-and-release with no sustain in between.
The release is how a sound ends. Some releases are abrupt: a recording stops, a finger lifts off a piano key with the damper engaged, a clap dies out instantly. Some are gradual: a piano note rings on after the key is released, decaying slowly. A reverberant space adds to the release of every sound in it. A single hand clap in a cathedral has a release of several seconds, even though the original clap was instant.
Release is also where you can hear the medium most clearly. A short tape loop's release is shaped by the splice; a digital file's release is exactly as long as the data. The same source sound can have a different release in different recordings or in different rooms.
Listen to each of the following with the envelope vocabulary in mind. Try to describe each one in terms of its attack, sustain, and release.
The first sound is mostly attack and release with no sustain in between. The second has a gradual swell into a sustained body, then a gentle decay. The third has no clear stages at all: the envelope is continuous and the sound is constantly changing within itself. Many sounds in the world fall into this third category, and it's worth noticing that the three-stage model isn't a perfect fit for every sound. It's a useful starting point, not a constraint.
You may have heard the acronym ADSR: attack, decay, sustain, release. This is the four-stage version of the envelope used in synthesizers. Decay is a new stage that sits between the attack peak and the held sustain level: it shapes how the sound settles down to the level it holds while a key is pressed. Here we're working with recorded sounds, so the three-stage version (attack, sustain, release) is enough.
Use the three sliders below to shape an envelope, then press play. The source is a fixed sine tone at 330 Hz; only the envelope changes. Try the extremes: a very short attack with no sustain (you'll get something percussive), a long attack with no release (it ends abruptly even though it began gradually), a balanced shape with all three stages clearly present. Notice how much the same source tone changes character based only on its envelope.
Keep your headphones at low volume before pressing Play. The source is a 330 Hz sine tone, played at a conservative level. Try setting attack to its minimum and sustain to zero for a percussive sound, or attack at 1500 ms and sustain at 1500 ms for a long swell.
A basic cut is a single move from a much larger vocabulary. Composers in the musique concrète tradition used these moves with razor blades and tape; you'll use them with a mouse and software. The principle is the same.
The fundamental insight is that recorded sound is material you can shape, not a fixed performance. Once a sound is recorded, you can cut it, repeat it, reverse it, slow it down, layer it, splice it next to other sounds. Each of these moves is a creative choice. The vocabulary below is what those moves are called.
Remove a selected region. The audio on either side of the cut closes the gap. Like deleting a phrase from the middle of a sentence.
Keep a selected region and remove everything else. The opposite of cut. Useful for isolating a single sound from a longer recording.
Place sounds next to each other in time. The fundamental musique concrète move: take a fragment of one recording, place a fragment of another after it, and a third after that. The arrangement is the composition.
A volume ramp at the boundary of an edit. A fade-in starts from silence and ramps up to full volume over some short duration; a fade-out is the reverse. Even very short fades (10 to 50 milliseconds) prevent click pops at edit boundaries. Use fades on every edit.
The overlap of a fade-out and a fade-in: one sound fades out as another fades in. The seam between two sounds becomes smooth instead of abrupt. Crossfades are how editors make hard cuts disappear.
A region played repeatedly. Useful as a working aid (loop a section while you adjust an effect) or as a compositional move (a tape-loop pattern that becomes the rhythmic base of a piece). In an audio editor like Audacity, you achieve this two ways: as a playback feature (the loop button in the transport, which repeats the selected region while you listen, without changing the audio), or by copying an audio clip and pasting it in succession within the track. Looping is far more central in DAWs built around it, like Ableton.
Play a sound backward. The recording is the same, but its envelope is flipped: the attack becomes the release, and the release becomes the attack. A piano note played backward swells in instead of striking; a struck cymbal played backward sounds like a swelling rush of air. Reverse changes the direction of the sound through time, not its duration or its pitch. It's one of the iconic moves of musique concrète: the composers had no software, so they reversed a sound by flipping the tape reel and playing it backward, and the resulting sounds became part of the language of the genre.
So far, every move in this list is fairly intuitive: you cut, you keep, you arrange, you fade, you blend, you repeat, you reverse. The next two moves are different. They both come out of a quirk of physics: in recorded sound, time and pitch aren't independent. To understand them you first have to understand that.
In the physical world, time and pitch are not independent. If you take a tape recording and play it back at half-speed, two things happen at once:
Speed it up to double-speed, and the opposite happens: it plays in half the time, and every pitch rises by an octave. This is a consequence of what sound is. A sound's pitch is determined by how often the waveform repeats per second; if you stretch the waveform out in time, each cycle takes longer, and the pitch drops in proportion.
The diagram below shows a sine wave played at three speeds, because a sine wave's cycles are visible at a glance. Notice that the cycles in the slowed version are wider, and the cycles in the sped-up version are narrower. The total length of the sound changes too: the slowed version takes twice as long; the sped-up version takes half as long.
All three are the same voice recording. The slow version is exactly 4/3 the duration of the source and exactly 3/4 of its pitch (about 5 semitones lower); the fast version is exactly 3/4 the duration and 4/3 of the pitch (about 5 semitones higher).
This is the world before signal processing. If you wanted a recording to be longer, you slowed it down, but it also went lower. If you wanted to raise its pitch, you sped it up, but it also got shorter. Schaeffer worked entirely in this world. So did every studio recording made before about 1990.
Modern software lets us decouple these two properties. With clever digital signal processing, you can change one without the other. The two terms below name the two halves of that decoupling:
Both are processing-intensive operations and both produce artifacts at extreme settings. Subtle shifts (a few semitones, a 10-20% time change) usually sound clean; dramatic ones (an octave up, doubling the duration) often sound digitally manipulated. The artifacts can be a bug or a feature, depending on what you're after.
The tradition is consistent across tools. Cut, splice, loop, reverse, time-and-pitch via tape speed: these moves all started with razor blades and tape machines. What software has added is speed (a click instead of a razor cut), reversibility (an undo button instead of starting over), and the ability to separate things, like time and pitch, that physical tape kept locked together.
The envelope vocabulary and the editing vocabulary are connected. Almost every edit you can make on a sound changes its envelope in some way. This is worth seeing explicitly.
Take a single source sound: a plucked string, lasting about three seconds. Sharp attack at the moment of the pluck, a sustained body that decays gradually, a natural release as the string's energy dissipates.
All four are the same recording. The source has a complete attack-sustain-release shape. The truncated version cuts off the release; the sound is the same up to 0.5 seconds, but instead of decaying naturally it ends abruptly. The reversed version flips the envelope: a sound that began with a sharp attack now begins with a slow swell. The fade-in version replaces the original sharp attack with a gradual one. None of these involves a different recording. They are all the same source, transformed by editing.
When you place two sounds next to each other in a piece, the seam between them is a small envelope problem. Sound A has a release; sound B has an attack. If you simply cut A off and start B, you might hear a click pop at the boundary, because the waveform of A doesn't end at zero and the waveform of B doesn't start at zero. The mismatch is audible.
Two solutions. The first: a small fade-out on A and a small fade-in on B. Each waveform now starts and ends at zero, so there's no jump. The second: a crossfade, where A and B overlap briefly. The result is smoother still: the seam becomes an active blend.
Listen for the moment of transition in each. The hard cut produces an audible click as the waveforms collide. The crossfade smooths the seam to the point that, depending on the source material, you might not notice it as a transition at all. The principle: in editing, the seams are where listeners catch you out. Fades and crossfades are how you make seams invisible.
Now turn this back on some real music. As you listen to Schaeffer's Étude aux chemins de fer and Henry's Variations pour une porte et un soupir, listen with two questions in mind:
For example: when you hear a sustained train whistle in Schaeffer suddenly cut off, ask yourself which it is. Is that a hard edit, with the recording chopped at that moment? A fast fade-out? Or a natural release the source happened to have? Sometimes you can tell, sometimes you can't, and being uncertain is part of the listening practice. The vocabulary doesn't promise certainty; it gives you a way to articulate what you're hearing and what you suspect, even when you're not sure.
You won't always be sure. The whole point of musique concrète is that the source sounds are transformed past the point of obvious recognition. But you can train your ear to notice these moves over time, and that ear will serve you when you start making your own pieces.
If shaping recorded sound is the part that grabbed you, there is a whole course in it. MUS 485, Advanced Sound Design I: Sampling, Editing, and Mixing, is built on these moves and takes them well past this first pass.
Links and further reading to come.