Outcome 4.1 — Science of Sound (Audio for Video Production)

Properties of Sound (4.1.1)

Sound is a mechanical wave—a pattern of pressure changes that travels through a material (air, water, a wall). That one sentence explains a lot about audio in video production: because sound is mechanical, it needs a medium (so it won’t travel in a vacuum), it interacts strongly with the environment (rooms, objects, people), and it can be described using wave properties that map directly to what you hear and what microphones capture.

Sound as pressure variation: compression and rarefaction

In air, a sound source (like a speaker cone or a vibrating guitar string) pushes and pulls on nearby air molecules. Regions where molecules are crowded are compressions (higher pressure), and regions where they are more spread out are rarefactions (lower pressure). Your ear—and a microphone diaphragm—responds to these pressure changes.

Why this matters: if you understand sound as moving pressure, you can predict what changes when you move a mic, change a room, or block a path. Many “mystery” audio problems (boomy voices, hollow tone, feedback) are just wave interactions.

Key wave properties and how they relate to what you hear
Frequency (pitch)

Frequency is how many wave cycles pass a point each second, measured in hertz (Hz). Higher frequency generally sounds higher in pitch.

  • A voice contains many frequencies at once (fundamental + harmonics), but the perceived pitch tends to follow the strongest repeating pattern.
  • In production, frequency is how you reason about EQ: “too much low end,” “harsh highs,” “mud in the low-mids.”
Wavelength (size of the wave)

Wavelength is the physical length of one cycle of a sound wave in a given medium.

The basic relationship between speed vv, frequency ff, and wavelength λ\lambda is:

v=fλv = f\lambda

Why this matters: wavelength tells you how sound interacts with objects and rooms.

  • Long wavelengths (low frequencies) wrap around objects more easily and are harder to block.
  • Short wavelengths (high frequencies) are easier to absorb and reflect more “directionally.”

Example (wavelength intuition): If you hear uneven bass in a room, it’s often because the room dimensions are comparable to bass wavelengths, creating strong room resonances (standing waves).

Amplitude (loudness) and sound pressure

Amplitude is the size of the pressure variation. Bigger amplitude generally means louder sound. In air, we often discuss this as sound pressure.

Because audio spans a huge range of pressures, we use the decibel (dB) scale, which is logarithmic.

For sound pressure level (SPL), a common expression is:

Lp=20log⁡10(pp0)L_p = 20\log_{10}\left(\frac{p}{p_0}\right)

where pp is the measured pressure and p0p_0 is a reference pressure (commonly 20 μPa20\,\mu\text{Pa} in air).

Why this matters in production:

  • Decibels compress big ranges into manageable numbers.
  • Many practical rules come from log behavior (for example, a small dB change can be clearly audible).
Phase (timing alignment)

Phase describes where you are within a wave cycle at a given moment—essentially timing alignment between waves of the same frequency.

Phase matters most when you combine signals (two mics on one source, a boom plus a lav, stereo pairs). If similar signals arrive slightly shifted in time, you can get partial cancellation at some frequencies and reinforcement at others.

A useful way to connect distance differences to phase shift is:

ϕ=2π(Δdλ)\phi = 2\pi\left(\frac{\Delta d}{\lambda}\right)

where Δd\Delta d is the path-length difference.

Speed of sound (propagation speed)

The speed of sound is how fast the pressure wave travels. In air, it depends mainly on temperature (and less on humidity/pressure). You don’t usually need an exact value for creative decisions, but you do need the concept:

  • Sound takes time to travel, so reflections arrive later than direct sound.
  • Mic spacing produces time differences that translate into phase differences.
Directionality and the “shape” of sound radiation

Real sources are not equally loud in all directions. A human voice, for example, is more directional at higher frequencies. This affects:

  • How “present” dialogue sounds when talent turns away from a mic.
  • How much room sound a mic captures (because the source-to-room ratio changes with direction and distance).
Worked examples (properties in action)

Example 1: Finding wavelength from frequency
If a tone has frequency f=1000 Hzf = 1000\,\text{Hz} and sound speed is about v≈343 m/sv \approx 343\,\text{m/s} (room temperature), then:

λ=vf\lambda = \frac{v}{f}

λ≈343 m/s1000 Hz\lambda \approx \frac{343\,\text{m/s}}{1000\,\text{Hz}}

λ≈0.343 m\lambda \approx 0.343\,\text{m}

So a 1 kHz wave is about 34 cm long—roughly the size of a laptop. That’s why moving a mic a few centimeters can audibly change tone when multiple mics are combined.

Example 2: Why doubling distance reduces level (free field intuition)
In open space, sound intensity drops roughly with the square of distance (inverse-square behavior). On a dB scale, doubling distance is commonly approximated as about a 6 dB drop. This is why getting a mic closer is often the best “audio fix” you can make: you increase direct sound much more than room sound.

What typically goes wrong (misconceptions)

A frequent misunderstanding is to treat frequency as “treble” and amplitude as “volume” in a simplistic way. Real sounds are mixtures of frequencies with time-varying amplitudes, so “loudness” and “brightness” are not single-knob properties—they depend on spectra, transients, and environment.

Exam Focus
  • Typical question patterns:
    • Identify which wave property corresponds to pitch, loudness, or timbre in a scenario.
    • Use v=fλv = f\lambda to reason about low vs high frequency behavior in a room.
    • Explain why phase issues appear when combining microphones.
  • Common mistakes:
    • Confusing frequency (Hz) with amplitude (level)—and describing one using the other.
    • Forgetting that wavelength depends on frequency and the medium (via speed).
    • Treating phase as “stereo width” rather than timing alignment that can cause cancellation.

Sound Transduction and Audio Signal Paths (4.1.2)

In video production, you’re constantly converting sound from one form into another—most importantly from air pressure changes into an electrical signal that can be recorded, transmitted, and edited. This conversion is called transduction.

What “sound transduction” means

Sound transduction is the process of converting acoustic energy (pressure variations in air) into electrical energy (a voltage/current signal) or vice versa.

  • Microphones transduce acoustic energy into electrical signals.
  • Speakers/headphones transduce electrical signals back into acoustic energy.

Why this matters: understanding transduction helps you troubleshoot low level, noise, distortion, and mismatched connections. Many field audio problems are not “the recorder is bad”—they’re “the signal is being degraded between the mic and the recorder.”

How microphones convert sound to electricity

Different microphone types use different physical mechanisms, but they share a big idea: a diaphragm moves in response to sound pressure, and that movement is converted into an electrical signal.

Dynamic (moving-coil) microphones

A dynamic microphone uses electromagnetic induction.

  • The diaphragm is attached to a coil of wire.
  • The coil moves within a magnetic field.
  • Motion induces a voltage corresponding to the sound waveform.

Why you care: dynamics are typically robust and handle loud sources well, but they may be less sensitive to very quiet detail than many condensers.

Condenser (capacitor) microphones

A condenser microphone uses a variable capacitor.

  • The diaphragm and backplate form a capacitor.
  • As the diaphragm moves, capacitance changes.
  • Electronics convert that change into a voltage signal.

Condenser mics require power for their internal electronics—often phantom power from a mixer/recorder (commonly 48 V in professional gear), though some use batteries.

Why you care: condensers are often chosen for detailed dialogue or ambience, but they are more sensitive to handling noise and environmental issues (wind, humidity) depending on design.

Ribbon microphones (conceptual)

A ribbon microphone uses a thin conductive ribbon suspended in a magnetic field. The ribbon moves with air particles, producing a voltage.

Why you care: ribbons can be smooth-sounding but are often more fragile and can be sensitive to wind blasts.

Resistance, impedance, and why “matching” matters

In DC circuits you learn resistance—opposition to current flow, measured in ohms. In audio, signals are AC (they vary over time), so we use impedance—the frequency-dependent “effective resistance” to AC.

  • A microphone output has an output impedance.
  • A recorder/mixer input has an input impedance.

In most modern audio systems you want impedance bridging:

  • Source output impedance relatively low.
  • Destination input impedance relatively high.

Why this matters: poor impedance relationships can reduce signal level, change frequency response, and increase noise susceptibility. A common real-world issue is plugging an instrument pickup, certain lav systems, or line-level outputs into the wrong kind of input.

Balanced vs unbalanced lines

The cable between your mic and recorder isn’t just “wire”—it’s part of the signal system. The two main connection types are balanced and unbalanced.

Unbalanced lines

An unbalanced cable uses two conductors:

  • Signal
  • Ground/shield (which serves as the return path)

Common connectors: TS 1/4-inch, RCA.

Main drawback: any noise induced into the cable tends to be added directly to the signal because the signal is measured relative to ground.

Balanced lines

A balanced line typically uses three conductors:

  • Hot (positive)
  • Cold (negative)
  • Shield/ground

Common connectors: XLR, TRS (when wired for balanced).

How it reduces noise (the key idea): the hot and cold carry the same audio but with opposite polarity. External interference tends to affect both conductors similarly (common-mode noise). The receiving device subtracts cold from hot, which cancels the common noise while reinforcing the desired signal.

Why this matters in video production:

  • Long cable runs on set are common.
  • Lighting dimmers, power cables, wireless devices, and video equipment can introduce electromagnetic interference.
  • Balanced cabling is one of the simplest, most effective noise-control tools.
Where resistance shows up practically (even if you don’t “calculate” it)

You may not compute resistance often, but it matters because:

  • Long or thin cables add resistance and can reduce level (more relevant for speakers than mics).
  • Poor connections add resistance intermittently—causing crackles, dropouts, or thin sound.
  • Shields rely on good continuity; corrosion or loose connectors can increase hum and RF problems.
Example: diagnosing a noisy connection

Imagine you have a lav mic feeding a camera input through an adapter, and you hear buzzing that changes when the cable moves.

  • If the run is unbalanced, it is far more likely to pick up interference.
  • A loose shield/ground can turn the cable into an antenna.
  • Using a balanced connection (or a proper balanced-to-unbalanced interface close to the camera) often reduces noise dramatically.

A common mistake is assuming “balanced is just a connector type.” It’s not. XLR often indicates balanced wiring, but what matters is how the system is wired end-to-end.

Comparison table: balanced vs unbalanced
FeatureBalanced lineUnbalanced line
ConductorsHot, cold, shieldSignal, shield/ground
Noise rejectionStrong (common-mode cancellation)Weak (noise adds to signal)
Best useLong runs, professional mics, set workShort runs, consumer gear
Common connectorsXLR, TRS (balanced)TS, RCA
Exam Focus
  • Typical question patterns:
    • Explain how a microphone converts sound energy to electrical energy.
    • Distinguish balanced vs unbalanced lines and predict which is better in a given scenario.
    • Interpret a noise/hum problem as likely grounding or interference and propose a fix.
  • Common mistakes:
    • Thinking “XLR automatically means clean audio” even if the equipment or adapter wiring is unbalanced.
    • Confusing resistance with impedance and ignoring frequency dependence.
    • Assuming all microphones need the same power (phantom power applies to many condensers, not dynamics).

Room Acoustics: Diffraction, Diffusion, Phase, and Harmonics (4.1.5)

Room acoustics is the reason dialogue sounds “cinematic” in one space and “cheap” in another, even with the same microphone. A room is not just a container—it’s an active filter that changes frequency balance, timing, and clarity through reflections and resonances.

Diffraction: sound bending around obstacles

Diffraction is the tendency of waves to bend around edges and spread into shadowed areas. The amount of diffraction depends strongly on wavelength:

  • Low frequencies (long wavelengths) diffract easily—bass wraps around objects and fills spaces.
  • High frequencies (short wavelengths) diffract less—treble is easier to block and more “line-of-sight.”

Why this matters on set:

  • A boom mic can lose high-frequency clarity if the actor’s mouth is partially blocked by an object (mask, book, or even head angle) because high frequencies don’t diffract as well.
  • Trying to “block” noise with thin barriers often fails for low-frequency rumble (traffic, HVAC) because those long waves diffract and transmit easily.
Diffusion: scattering reflections to reduce harshness

Diffusion is the scattering of sound reflections in many directions. A diffuser (or an irregular surface like bookshelves, textured walls, or purpose-built panels) breaks up strong, mirror-like reflections.

Why this matters:

  • Strong, focused reflections can cause flutter echo and comb filtering.
  • Diffusion can make a room sound more natural without deadening it as much as heavy absorption.

A practical way to think about it: absorption reduces the amount of reflected energy; diffusion changes the direction and density of reflections, making them less objectionable.

Phase interactions in rooms: comb filtering and cancellations

When direct sound and reflected sound combine at a microphone, they arrive with a time difference. That time difference corresponds to a phase difference that varies by frequency. The result is often comb filtering—a series of peaks and dips across the frequency response.

Mechanism (step by step):

  1. A sound leaves the source and reaches the mic directly.
  2. The same sound reflects off a surface and reaches the mic slightly later.
  3. At some frequencies, the delayed version is close to in-phase and reinforces the direct sound.
  4. At other frequencies, it is out-of-phase and cancels.

This is why dialogue can sound “hollow” or “phasey” in reflective rooms—especially when the mic is far from the mouth, making reflections relatively loud compared with direct sound.

Example: why small distance changes can change tone
If you move a mic, you change reflection path lengths. Since phase depends on Δd/λ\Delta d/\lambda, even a few centimeters can shift cancellations at mid and high frequencies.

A common misconception is to blame this on “bad EQ.” EQ can sometimes mask the worst peaks, but comb filtering is not a simple tonal tilt—it’s many narrow cancellations caused by timing.

Harmonics: why timbre changes in different spaces

Most real sounds are complex. Harmonics are frequency components at integer multiples of a fundamental frequency. They are a major contributor to timbre (tone color).

  • If the fundamental is f0f_0, harmonics occur at 2f02f_0, 3f03f_0, 4f04f_0, and so on.
  • The balance of these harmonics is what makes a voice sound like that person, and what makes a clarinet different from a flute.

Rooms affect harmonic balance because:

  • Surfaces absorb high frequencies more than low frequencies in many practical situations.
  • Reflections can reinforce or cancel certain frequency bands due to phase interactions.
  • Room resonances can exaggerate specific low and low-mid frequencies.
Standing waves (room modes) and “boomy” sound

In enclosed spaces, some frequencies fit between boundaries in a way that creates standing waves (also called room modes). At certain positions in the room, those frequencies become much louder (antinodes) or much quieter (nodes).

Why this matters:

  • A voice recorded near a wall or in a corner often sounds boomier because low-frequency energy builds up.
  • Moving the mic (or the talent) a small distance can significantly change bass response.

This is often misdiagnosed as “the mic has too much bass.” Sometimes it does—but room modes are frequently the bigger culprit.

Real-world applications: diagnosing a room quickly

A fast way to evaluate a room is to listen for:

  • Flutter echo: rapid, ringing repeats between parallel hard surfaces (clap test).
  • Boominess: persistent low-frequency buildup (often in corners or small rooms).
  • Hollowness/phase: comb filtering from nearby reflections (hard floors, bare walls).

On set, you often can’t acoustically redesign a room, but you can improve results by:

  • Getting the microphone closer to increase direct-to-room ratio.
  • Using soft furnishings, rugs, sound blankets, or curtains to reduce strong reflections.
  • Avoiding placing the talent right against reflective boundaries.
Exam Focus
  • Typical question patterns:
    • Explain why low frequencies “wrap around” obstacles (diffraction) and how that affects isolation.
    • Describe how diffusion differs from absorption and what each accomplishes.
    • Identify comb filtering or phase cancellation as the cause of “hollow” dialogue.
  • Common mistakes:
    • Treating diffraction as “echo” (diffraction is bending/spreading, not reflecting).
    • Assuming diffusion makes a room quieter (it usually redistributes energy rather than removing it).
    • Thinking phase problems only happen with two microphones—rooms create phase issues too.

Direct Sound, Early Reflections, and Reverberation (4.1.6)

When you record sound in a room, your mic captures a blend of arrivals over time. Understanding the timeline of arrivals—direct sound, early reflections, and reverberation—lets you predict clarity, intelligibility, and the sense of space.

Direct sound: the “clean” path

Direct sound is the sound that travels straight from the source to the microphone with no reflections.

Why it matters most: direct sound carries the clearest articulation for speech. If direct sound is strong compared with reflections, dialogue feels present and intelligible.

How you increase direct sound in practice:

  • Move the microphone closer (without ruining framing or sounding unnatural).
  • Use a microphone with appropriate directionality for that environment (but remember: directionality does not eliminate reflections—it mainly reduces off-axis pickup).

A common mistake is relying on a shotgun mic far away to “reach” dialogue. Distance doesn’t work that way—reflections and room tone rise in level relative to the voice as the mic moves away.

Early reflections: the first bounces that shape clarity

Early reflections are the first set of reflected sounds that arrive shortly after the direct sound, typically from nearby surfaces (floor, ceiling, walls, table tops).

Why they matter:

  • Early reflections can reinforce the direct sound and add a pleasing sense of fullness.
  • But if they are strong and delayed enough, they reduce speech clarity and can cause comb filtering.

Mechanism (what your mic “hears”):

  • Direct sound arrives first.
  • A strong early reflection arrives slightly later and mixes with the direct sound.
  • Depending on delay and level, the result ranges from “natural spaciousness” to “boxy/phasey.”

Example: floor reflection with a boom mic
A hard floor can create a strong reflection path from mouth to floor to mic. Adding a rug or sound blanket on the floor beneath the actor can significantly reduce that early reflection—often more effectively than changing EQ.

Reverberation: dense, decaying reflections (the room’s signature)

Reverberation (reverb) is the dense collection of many reflections arriving so close together that they blend into a smooth tail. It is not a single echo—it’s a decay over time.

Why it matters in video:

  • Dialogue recorded with excessive reverb can be hard to understand and difficult to “fix” in post.
  • Reverb tells the viewer where they are. Too much (or the wrong kind) can conflict with the visuals.

A common way to describe reverb time is RT60RT_{60}, the time it takes for reverberant sound level to decay by 60 dB. One classical estimation method (for certain room conditions) is Sabine’s formula:

RT60=0.161(VA)RT_{60} = 0.161\left(\frac{V}{A}\right)

where VV is room volume in m3\text{m}^3 and AA is total absorption in sabins.

You don’t need to calculate RT60RT_{60} on set often, but the concept helps you reason: more absorption (curtains, carpets, acoustic treatment, people) reduces reverb time; larger and more reflective rooms increase it.

Putting the three together: the direct-to-reverberant ratio

What you perceive as “close” or “far” in recorded dialogue is strongly influenced by the direct-to-reverberant ratio (how much direct voice compared with reflections and reverb).

  • Close mic placement increases direct sound much more than it increases reverb.
  • If the mic is far away, the direct sound weakens quickly, but the room’s reverberant field changes less dramatically—so the recording sounds roomy.

This is why the single most powerful dialogue improvement is usually mic proximity, not post-processing.

The precedence (Haas) effect: why small delays can still sound like one event

Your brain tends to localize a sound based on the first arriving wavefront (often the direct sound), and it merges very fast reflections into a single perceived event. This is helpful—small early reflections can add pleasant spaciousness without sounding like a separate echo.

But when reflections are strong and delayed enough, you start to hear them as distinct echoes or you experience loss of clarity.

Practical production strategies
  • Control early reflections first: Treat the nearest hard surfaces (floors, nearby walls, ceilings) because those reflections are strongest and earliest.
  • Choose mic placement to maximize direct sound: A boom just out of frame often beats a more distant mic even if it is “higher quality.”
  • Be careful combining mics: A boom and lav recorded together can create phase issues if mixed without alignment. Often you choose one as primary and use the other as backup or for specific moments.
Example: diagnosing a “reverby” interview

You record an interview in an empty office with hard walls.

  • If the voice sounds distant and roomy, the direct-to-reverberant ratio is likely low.
  • Moving the mic closer increases direct sound.
  • Adding absorption (sound blankets, curtains, even bringing in furniture) reduces reflections and reverb.
  • If the tone is hollow rather than simply roomy, strong early reflections and comb filtering are likely—treating a nearby reflective surface can help immediately.
Exam Focus
  • Typical question patterns:
    • Given a recording scenario, identify whether the main issue is weak direct sound, strong early reflections, or excessive reverberation.
    • Propose on-set changes (mic placement, absorption, diffusion) to improve intelligibility.
    • Explain why close miking improves clarity using the idea of direct-to-reverberant ratio.
  • Common mistakes:
    • Calling any roominess “echo” (echo is discrete; reverberation is dense decay).
    • Trying to solve a reflection problem only with EQ rather than addressing timing/geometry.
    • Mixing two dialogue mics together without considering time alignment and phase.