Comprehensive Study Guide for Fundamental Acoustics and Speech Science
Fundamentals of Acoustic Physics
Sound Wave (Ljudvåg): Variations in air pressure that propagate through the air.
Periodic Sound: A sound where the same waveform pattern repeats regularly over time.
Simple Periodic Sound: A sound consisting of a single frequency with no overtones, commonly referred to as a sine wave.
Compound Periodic Sound: A complex sound consisting of several frequencies occurring simultaneously. These sounds possess a fundamental tone and associated partials located above that fundamental frequency.
Frequency: The number of times a wave repeats per second. Low frequency is characterized by few oscillations per second, while high frequency is characterized by many oscillations per second.
Hertz (): The standard unit for frequency. . For example, signifies oscillations per second.
Period: Refers to one complete oscillation of a wave.
Period Time: The duration required for one complete oscillation. A wave with a high frequency makes many oscillations quickly and thus has a short period time.
Amplitude: The magnitude of the wave from its center (zero point) to its peak. A larger amplitude results in a stronger (louder) sound, while a smaller amplitude results in a weaker sound.
Phase: Indicates the specific position of a wave within its oscillation cycle at a given time.
Fundamental Frequency (): The frequency of the fundamental tone. It is the frequency that determines the perceived pitch of a fundamental tone.
Fundamental Tone (Grundton): The frequency located at the very bottom of the harmonic series. In a series containing , , and as partials, the fundamental tone would be .
Partials (Deltoner): The various other frequencies that exist alongside the fundamental frequency in a complex sound.
Aperiodic Sound: A sound where the waveform pattern does not repeat regularly. Examples include fricatives, which consist of noise and lack a regular periodic waveform.
White Noise (Vitt brus): Noise that contains approximately all frequencies with equal energy within the observed frequency range.
Colored Noise (Färgat brus): Noise where energy is not evenly distributed across the frequency spectrum.
Acoustic Signal Analysis and Measurement
Impulse Sound/Transient: A short, fast, and sudden sound progression or change, such as a snap, click, or clap.
Bandwidth: The size of the frequency range that a signal or system occupies.
Spectrogram: A visual representation of how a sound's frequencies and energy change over time. The x-axis represents time, and the y-axis represents frequency.
Narrowband Spectrogram: Used to clearly see the fundamental frequency () and the individual partials.
Broadband Spectrogram: Used to make formants and the temporal structure (time-based changes) more distinct.
Energy Representation: Darker areas indicate higher energy, while lighter areas indicate lower energy.
Spectrum: A display of which frequencies are present at a specific point in time.
Waveform: A visual representation showing how sound pressure changes over time.
Decibel (): A unit used to describe the level or difference in sound intensity/strength.
Sound Pressure Level: Sound pressure expressed on a logarithmic scale, indicating how strong the sound pressure is in .
Bark Scale: A psychoacoustic scale that corresponds more closely to how human hearing perceives frequencies rather than a linear scale.
Physiology and Resonances of Speech Production
Phonation: The process of the vocal folds vibrating to produce sound.
Vocal Tract (Ansatsrör): The space above the vocal folds through which sound passes before exiting the body, including the pharynx (svalg), oral cavity (munhåla), and nasal cavity (näshåla).
Harmonic Sound: A periodic sound where frequencies are built regularly based on the fundamental frequency ().
Resonant Frequencies: Specific frequencies that the vocal tract reinforces or amplifies significantly, which become the formants , , and .
Node (Nod): A point in a standing wave where the amplitude is at a minimum or near zero.
Antinode (Buk): A point in a standing wave where the amplitude is at its maximum.
Wavelength: The physical distance between two identical points in a wave cycle.
Speed of Sound: The speed at which a sound wave travels through a medium, such as water, steel, or air.
Formant: A frequency range where the vocal tract amplifies the sound. On a spectrogram, these appear as dark horizontal bands.
: Associated with the openness/height of a vowel.
: Associated with the frontness or backness of the tongue position.
Anti-formant: Features that occur primarily during the production of nasals and laterals. Anti-formants dampen sound energy. During the production of sounds like or , a pathway opens through the nasal cavity, creating extra resonances while simultaneously causing certain frequencies to weaken.
Theories of the Vocal Tract and Speech Production
Source-Filter Theory (Källa-filter teorin): This theory posits that speech production involves a source and a filter.
Vowels/Periodic Sounds: The source is the vocal folds; the filter is the vocal tract, which modifies the sound by reinforcing or dampening specific frequencies.
Nasals: The source is the vocal folds; the filter is the combination of the vocal tract and the nasal cavity.
Voiceless Plosives (e.g., ): The source is air pressure/burst or explosion; the filter is the vocal tract.
Voiced Plosives (e.g., ): The source is a combination of vocal fold vibration and a burst; the filter is the vocal tract.
Voiceless Fricatives (e.g., ): The source is turbulence; the filter is the vocal tract.
Voiced Fricatives (e.g., ): The source is a combination of vocal fold vibration and turbulence; the filter is the vocal tract.
Perturbation Theory: Relates to antinodes and nodes within the vocal tract modeled as a tube.
At the open end (front), there is a velocity antinode (hastighetsbuk) and a pressure node (trycknod). A constriction here lowers formant frequencies.
At the closed end (glottis), there is a pressure antinode (tryckbuk) and a velocity node (hastighetsnod). A constriction here raises formant frequencies.
Schwa: A neutral vowel used as a model for a half-open, half-thick tube. It serves as a neutral starting point for vocal tract modeling.
Vowel Categorization:
Open vs. Closed: Determined by . A high indicates a greater degree of opening; a low indicates a smaller opening.
Front vs. Back: Determined by . A high indicates the tongue is further forward; a low indicates the tongue is further back.
Voicing: Voiced sounds involve vocal fold vibration, while voiceless sounds do not. On a spectrogram, voicing is visible at the very bottom as a "voice bar."
Sublingual Cavity: The air space beneath the tongue. For certain sounds, this cavity can influence the resonances of the sound produced.
Voice Quality and Clinical Analysis
Perceptual Voice Analysis: Evaluating a voice by listening to its characteristics to judge qualities such as hoarseness, creakiness, pressed quality, strength, and pitch.
Voice Quality: Describes the specific sound of a person's voice. Examples include leaky, creaky, pressed, or modal.
Phonogram: A graph showing a person's vocal range, including the lowest and highest tones and the intensity (strength) at various pitches.
Perturbation Analysis/Measures: The analysis of small, irregular variations in the voice signal from one period to the next.
CPPS (Cepstral Peak Prominence Smoothed): A measurement used to determine how periodic or regular a voice is. It is utilized in assessing voice quality and the degree of dysphonia.
Specific Voice Qualities:
Modal Voice: The normal, neutral voice quality. Vocal folds vibrate regularly, providing a clear and partials. Vocal lips are typically short, thick, and relaxed.
Breathy/Leaky Voice: Occurs when the vocal folds do not close completely, allowing air to escape and creating significant noise.
Creak (Knarr): Characterized by short, thick, and relaxed vocal lips with tight contact (primarily the front part vibrates). Vibrations are slow and irregular, resulting in a low and high muscle tension.
Falsetto: Characterized by long, thin, and tense vocal lips. The glottis does not close completely. It features a high with regular vibrations but weak overtones.
Whisper: Lacks regular vocal fold vibration. Instead, noise is generated in the glottis, and there is no clear .
Rough Voice: Caused by irregularities on the vocal lips combined with high tension. It results in irregular oscillations and noise components.
Subglottal Pressure: The air pressure beneath the vocal folds. As one breathes, pressure builds up against the vocal folds, which helps initiate vibration.
Larynx: The voice box or struphuvud.
Vocal Lips: Scientific term for the vocal folds (stämbanden).
Speech Perception
Speech Perception: The process of how we perceive and interpret speech, where speech signals are received, interpreted by the brain, and understood as linguistic sounds.
Invariance Problem: The phenomenon where the same linguistic sound can sound different depending on the speaker and the context. Invariance refers to the constant elements that remain despite this variation.
Segmentation Problem: In natural speech, there are no clear pauses between individual sounds; it sounds like a continuous stream. The brain must determine where the boundaries for each word are located.
Categorical Perception: The brain is less sensitive to small acoustic differences within the same category but perceives differences between categories much more sharply.
McGurk Effect: An illustration of how vision affects speech perception. If a person hears the sound but sees the mouth movement for , they may perceive the sound as because the brain combines auditory and visual information.
Ganong Effect: The tendency for the brain to use words and linguistic context to interpret ambiguous or unclear speech sounds.
Phoneme Restoration: The process where the brain fills in a missing or unclear phoneme based on the surrounding context.
Indexical Properties: Characteristics in speech that provide information about the speaker’s identity, personality, or the specific situation.
Theories of Speech Perception
Motor Theory: Suggests that we use our knowledge of the motor and articulatory production of speech to help us perceive it.
H&H Theory (Hyper & Hypo): Posits that the speaker adapts to the listener and the level of background information available.
Hypoarticulation: Occurs when little information is needed by the listener.
Hyperarticulation: Occurs when a high amount of information is necessary for the listener to understand.
Prototype Theory: Suggests the existence of an idealized representation in the mental lexicon of how a specific sound should sound. Heard sounds are compared against this mental prototype.
Exemplar Theory: Suggests that instead of a single prototype, we store a vast number of specific memory traces (exemplars) of every time we have heard a sound or word before.
Statistically Based Models: Posit that the listener interprets speech by building on learned statistical relationships and probabilities found within linguistic and acoustic information.
Questions & Discussion
How does a spectrogram represent energy? In a spectrogram, the x-axis shows time and the y-axis shows frequency. The intensity of the energy is shown via darkness; darker areas indicate more energy and lighter areas indicate less energy.
What is the difference between narrowband and broadband spectrograms? A narrowband spectrogram allows for the clear visualization of and partials, whereas a broadband spectrogram makes formants and the temporal structure easier to see.
What determines vowel openness and frontness in terms of formants? Vowel openness is linked to (high is more open), and the frontness or backness of a vowel is linked to (high is more forward).