Comprehensive Study Guide for Fundamental Acoustics and Speech Science

Fundamentals of Acoustic Physics

  • Sound Wave (Ljudvåg): Variations in air pressure that propagate through the air.

  • Periodic Sound: A sound where the same waveform pattern repeats regularly over time.

  • Simple Periodic Sound: A sound consisting of a single frequency with no overtones, commonly referred to as a sine wave.

  • Compound Periodic Sound: A complex sound consisting of several frequencies occurring simultaneously. These sounds possess a fundamental tone and associated partials located above that fundamental frequency.

  • Frequency: The number of times a wave repeats per second. Low frequency is characterized by few oscillations per second, while high frequency is characterized by many oscillations per second.

  • Hertz (HzHz): The standard unit for frequency. 1Hz=1 oscillation per second1\,Hz = 1\text{ oscillation per second}. For example, 500Hz500\,Hz signifies 500500 oscillations per second.

  • Period: Refers to one complete oscillation of a wave.

  • Period Time: The duration required for one complete oscillation. A wave with a high frequency makes many oscillations quickly and thus has a short period time.

  • Amplitude: The magnitude of the wave from its center (zero point) to its peak. A larger amplitude results in a stronger (louder) sound, while a smaller amplitude results in a weaker sound.

  • Phase: Indicates the specific position of a wave within its oscillation cycle at a given time.

  • Fundamental Frequency (F0F_0): The frequency of the fundamental tone. It is the frequency that determines the perceived pitch of a fundamental tone.

  • Fundamental Tone (Grundton): The frequency located at the very bottom of the harmonic series. In a series containing 400Hz400\,Hz, 300Hz300\,Hz, and 200Hz200\,Hz as partials, the fundamental tone would be 100Hz100\,Hz.

  • Partials (Deltoner): The various other frequencies that exist alongside the fundamental frequency in a complex sound.

  • Aperiodic Sound: A sound where the waveform pattern does not repeat regularly. Examples include fricatives, which consist of noise and lack a regular periodic waveform.

  • White Noise (Vitt brus): Noise that contains approximately all frequencies with equal energy within the observed frequency range.

  • Colored Noise (Färgat brus): Noise where energy is not evenly distributed across the frequency spectrum.

Acoustic Signal Analysis and Measurement

  • Impulse Sound/Transient: A short, fast, and sudden sound progression or change, such as a snap, click, or clap.

  • Bandwidth: The size of the frequency range that a signal or system occupies.

  • Spectrogram: A visual representation of how a sound's frequencies and energy change over time. The x-axis represents time, and the y-axis represents frequency.

    • Narrowband Spectrogram: Used to clearly see the fundamental frequency (F0F_0) and the individual partials.

    • Broadband Spectrogram: Used to make formants and the temporal structure (time-based changes) more distinct.

    • Energy Representation: Darker areas indicate higher energy, while lighter areas indicate lower energy.

  • Spectrum: A display of which frequencies are present at a specific point in time.

  • Waveform: A visual representation showing how sound pressure changes over time.

  • Decibel (dBdB): A unit used to describe the level or difference in sound intensity/strength.

  • Sound Pressure Level: Sound pressure expressed on a logarithmic dBdB scale, indicating how strong the sound pressure is in dBdB.

  • Bark Scale: A psychoacoustic scale that corresponds more closely to how human hearing perceives frequencies rather than a linear scale.

Physiology and Resonances of Speech Production

  • Phonation: The process of the vocal folds vibrating to produce sound.

  • Vocal Tract (Ansatsrör): The space above the vocal folds through which sound passes before exiting the body, including the pharynx (svalg), oral cavity (munhåla), and nasal cavity (näshåla).

  • Harmonic Sound: A periodic sound where frequencies are built regularly based on the fundamental frequency (F0F_0).

  • Resonant Frequencies: Specific frequencies that the vocal tract reinforces or amplifies significantly, which become the formants F1F_1, F2F_2, and F3F_3.

  • Node (Nod): A point in a standing wave where the amplitude is at a minimum or near zero.

  • Antinode (Buk): A point in a standing wave where the amplitude is at its maximum.

  • Wavelength: The physical distance between two identical points in a wave cycle.

  • Speed of Sound: The speed at which a sound wave travels through a medium, such as water, steel, or air.

  • Formant: A frequency range where the vocal tract amplifies the sound. On a spectrogram, these appear as dark horizontal bands.

    • F1F_1: Associated with the openness/height of a vowel.

    • F2F_2: Associated with the frontness or backness of the tongue position.

  • Anti-formant: Features that occur primarily during the production of nasals and laterals. Anti-formants dampen sound energy. During the production of sounds like MM or NN, a pathway opens through the nasal cavity, creating extra resonances while simultaneously causing certain frequencies to weaken.

Theories of the Vocal Tract and Speech Production

  • Source-Filter Theory (Källa-filter teorin): This theory posits that speech production involves a source and a filter.

    • Vowels/Periodic Sounds: The source is the vocal folds; the filter is the vocal tract, which modifies the sound by reinforcing or dampening specific frequencies.

    • Nasals: The source is the vocal folds; the filter is the combination of the vocal tract and the nasal cavity.

    • Voiceless Plosives (e.g., T,p,kT, p, k): The source is air pressure/burst or explosion; the filter is the vocal tract.

    • Voiced Plosives (e.g., b,d,gb, d, g): The source is a combination of vocal fold vibration and a burst; the filter is the vocal tract.

    • Voiceless Fricatives (e.g., f,sf, s): The source is turbulence; the filter is the vocal tract.

    • Voiced Fricatives (e.g., v,zv, z): The source is a combination of vocal fold vibration and turbulence; the filter is the vocal tract.

  • Perturbation Theory: Relates to antinodes and nodes within the vocal tract modeled as a tube.

    • At the open end (front), there is a velocity antinode (hastighetsbuk) and a pressure node (trycknod). A constriction here lowers formant frequencies.

    • At the closed end (glottis), there is a pressure antinode (tryckbuk) and a velocity node (hastighetsnod). A constriction here raises formant frequencies.

  • Schwa: A neutral vowel used as a model for a half-open, half-thick tube. It serves as a neutral starting point for vocal tract modeling.

  • Vowel Categorization:

    • Open vs. Closed: Determined by F1F_1. A high F1F_1 indicates a greater degree of opening; a low F1F_1 indicates a smaller opening.

    • Front vs. Back: Determined by F2F_2. A high F2F_2 indicates the tongue is further forward; a low F2F_2 indicates the tongue is further back.

  • Voicing: Voiced sounds involve vocal fold vibration, while voiceless sounds do not. On a spectrogram, voicing is visible at the very bottom as a "voice bar."

  • Sublingual Cavity: The air space beneath the tongue. For certain sounds, this cavity can influence the resonances of the sound produced.

Voice Quality and Clinical Analysis

  • Perceptual Voice Analysis: Evaluating a voice by listening to its characteristics to judge qualities such as hoarseness, creakiness, pressed quality, strength, and pitch.

  • Voice Quality: Describes the specific sound of a person's voice. Examples include leaky, creaky, pressed, or modal.

  • Phonogram: A graph showing a person's vocal range, including the lowest and highest tones and the intensity (strength) at various pitches.

  • Perturbation Analysis/Measures: The analysis of small, irregular variations in the voice signal from one period to the next.

  • CPPS (Cepstral Peak Prominence Smoothed): A measurement used to determine how periodic or regular a voice is. It is utilized in assessing voice quality and the degree of dysphonia.

  • Specific Voice Qualities:

    • Modal Voice: The normal, neutral voice quality. Vocal folds vibrate regularly, providing a clear F0F_0 and partials. Vocal lips are typically short, thick, and relaxed.

    • Breathy/Leaky Voice: Occurs when the vocal folds do not close completely, allowing air to escape and creating significant noise.

    • Creak (Knarr): Characterized by short, thick, and relaxed vocal lips with tight contact (primarily the front part vibrates). Vibrations are slow and irregular, resulting in a low F0F_0 and high muscle tension.

    • Falsetto: Characterized by long, thin, and tense vocal lips. The glottis does not close completely. It features a high F0F_0 with regular vibrations but weak overtones.

    • Whisper: Lacks regular vocal fold vibration. Instead, noise is generated in the glottis, and there is no clear F0F_0.

    • Rough Voice: Caused by irregularities on the vocal lips combined with high tension. It results in irregular oscillations and noise components.

  • Subglottal Pressure: The air pressure beneath the vocal folds. As one breathes, pressure builds up against the vocal folds, which helps initiate vibration.

  • Larynx: The voice box or struphuvud.

  • Vocal Lips: Scientific term for the vocal folds (stämbanden).

Speech Perception

  • Speech Perception: The process of how we perceive and interpret speech, where speech signals are received, interpreted by the brain, and understood as linguistic sounds.

  • Invariance Problem: The phenomenon where the same linguistic sound can sound different depending on the speaker and the context. Invariance refers to the constant elements that remain despite this variation.

  • Segmentation Problem: In natural speech, there are no clear pauses between individual sounds; it sounds like a continuous stream. The brain must determine where the boundaries for each word are located.

  • Categorical Perception: The brain is less sensitive to small acoustic differences within the same category but perceives differences between categories much more sharply.

  • McGurk Effect: An illustration of how vision affects speech perception. If a person hears the sound /ba//ba/ but sees the mouth movement for /ga//ga/, they may perceive the sound as /da//da/ because the brain combines auditory and visual information.

  • Ganong Effect: The tendency for the brain to use words and linguistic context to interpret ambiguous or unclear speech sounds.

  • Phoneme Restoration: The process where the brain fills in a missing or unclear phoneme based on the surrounding context.

  • Indexical Properties: Characteristics in speech that provide information about the speaker’s identity, personality, or the specific situation.

Theories of Speech Perception

  • Motor Theory: Suggests that we use our knowledge of the motor and articulatory production of speech to help us perceive it.

  • H&H Theory (Hyper & Hypo): Posits that the speaker adapts to the listener and the level of background information available.

    • Hypoarticulation: Occurs when little information is needed by the listener.

    • Hyperarticulation: Occurs when a high amount of information is necessary for the listener to understand.

  • Prototype Theory: Suggests the existence of an idealized representation in the mental lexicon of how a specific sound should sound. Heard sounds are compared against this mental prototype.

  • Exemplar Theory: Suggests that instead of a single prototype, we store a vast number of specific memory traces (exemplars) of every time we have heard a sound or word before.

  • Statistically Based Models: Posit that the listener interprets speech by building on learned statistical relationships and probabilities found within linguistic and acoustic information.

Questions & Discussion

  • How does a spectrogram represent energy? In a spectrogram, the x-axis shows time and the y-axis shows frequency. The intensity of the energy is shown via darkness; darker areas indicate more energy and lighter areas indicate less energy.

  • What is the difference between narrowband and broadband spectrograms? A narrowband spectrogram allows for the clear visualization of F0F_0 and partials, whereas a broadband spectrogram makes formants and the temporal structure easier to see.

  • What determines vowel openness and frontness in terms of formants? Vowel openness is linked to F1F_1 (high F1F_1 is more open), and the frontness or backness of a vowel is linked to F2F_2 (high F2F_2 is more forward).