Speech Perception - Key Concepts and Examples
Speech Perception: Overview
- Speech perception is a key component of language processing that differs from other auditory perception (e.g., distinguishing a siren from a human voice even at the same loudness, pitch, and frequency).
- It happens incredibly fast: we perceive speech while someone is talking or singing, often with speed and efficiency that feel almost instantaneous.
- A central goal in psychology is to understand how humans do this and why some people find it easier than others.
- Today’s focus: main component parts of speech perception, features of speech perception, and specifically:
- categorical perception
- segmentation
- context/environment effects on perception
- The process is studied through varied examples and experiments to reveal underlying mechanisms and brain processes.
Key Concepts
- Speech perception vs. other auditory signals: even with similar acoustic properties, humans distinguish natural speech from environmental sounds or music, often automatically.
- The left hemisphere is heavily involved in language processing for most people (language centers typically left-lateralized).
- Speech processing may be akin to a specialized form of music in some respects (e.g., prosody and the way we modulate pitch and rhythm when communicating, especially with children).
- Child-directed speech (formerly called motherese) features higher pitch, slower rate, exaggerated vowel lengthening, and a sing-song quality that facilitates infant speech processing.
- The brain appears to integrate music-like processing and speech processing in overlapping regions.
Examples of Speech Perception in Action
- Auctioneers (International Auctioneer Championship): rapid, continuous, highly-contextual speech makes it hard to hear exact words, yet listeners can infer bids from context, gesture, and rhythmic cues.
- Example: a competition where contestants bid on lots (e.g., cutting boards) with rapid, overlapping numbers and phrases.
- Observations: listeners can still track intent and respond despite limited audible content; highlights the role of real-time context and expectations.
- Icelandic football commentary (Euro 2016): despite language barriers, listeners can understand excitement and gist through prosody and context.
- Note: even without understanding every word, the emotional and prosodic cues convey meaning.
- Overall point from these examples: speech perception relies on rapid decoding, pattern recognition, and top-down expectations to derive meaning from partial acoustic information.
Brain and Language Processing
- The left hemisphere often dominates language processing, including speech sounds.
- Context and language familiarity influence perception:
- Speech can be treated as a special kind of music, especially in how we modulate pitch, tempo, and intonation when addressing infants.
- Child-directed speech shows how prosodic features facilitate processing of language in early development.
- A looping auditory example: repeating phrases can start to sound like singing, illustrating overlapping mechanisms between speech and music processing.
- The perception of speech is influenced by context and statistical regularities, not just the raw acoustic signal.
- Categorical Perception (CP): we hear sounds as belonging to discrete categories (e.g., /b/ vs /d/) rather than along a continuous gradient.
- Classic demonstration: a continuum between /b/ and /d/ yields a crisp boundary where listeners report hearing either /b/ or /d/ with little or no perception of intermediate sounds.
- Example from class: a synthesized sound between ba and da is heard as one category or the other, not as a middle sound.
- Note: CP suggests our brains map continuous acoustic variation onto discrete linguistic categories.
- Phonemes vs. Syllables:
- Phonemes: the smallest units of sound that distinguish meaning in a language.
- Syllables: rhythmic units that may include one or more phonemes.
- Example: the word yacht has 3 phonemes but only 1 syllable.
- Example: psychology has variable phoneme counts (commonly eight, occasionally nine, depending on dialect and analysis).
- Phonetics and Phonology:
- Phonemes are the building blocks; their actual realization can vary (allophones) and may be mapped differently across languages.
- A phoneme can be realized by multiple letters (e.g., "ph" represents the /f/ sound but is written with two letters).
- The term phonetician refers to a specialist who studies phonetics.
- Spectrograms: a visual representation of the speech signal showing time on the horizontal axis, frequency on the vertical axis, and darkness (intensity) representing energy.
- Example: a sentence like "Joe took father's shoe bench out" shows stronger energy for vowels and certain consonants (e.g., /b/, /sh/) and weaker energy for others (e.g., /k/ in "took").
- Spectrograms help illustrate where syllables and phonemes are likely located and how coarticulation appears as overlapping energy patterns.
- Processing pipeline (in real time):
- Auditory input -> decode the features -> segment sounds -> perceive and categorize phonemes -> integrate into phonological representations -> determine word status and meaning -> use knowledge to interpret the message.
- This entire sequence unfolds in microseconds; some discussion touches on nano- to microsecond time scales.
- Individual differences that affect perception:
- Speaker differences: sex, dialect, speaking rate, accent.
- Rapid production can challenge perception, but humans remain highly adept at comprehension in many conditions.
- Adverse listening conditions:
- Environments with noise, reverberation, hard surfaces (cafés, large halls) can degrade perceptual clarity.
- Acoustic design considerations (e.g., fabric in seats at opera houses) can influence perception by shaping reverberation and noise.
- Segmentation: the ability to divide continuous speech into discrete words.
- Not taught explicitly; we infer word boundaries from phonotactics and context.
- Clues to segmentation include impossible or unlikely letter-sound combinations within a single word (e.g., /cf/ cannot occur inside a single English word, indicating a boundary between segments).
- Examples: fapple vs waffle; boundary after the syllable pattern of a word; stress patterns can signal word boundaries.
- Stress patterns and lexical boundaries:
- English stress shifts can change word meaning (e.g.,
- record (noun) vs record (verb)
- conduct (noun) vs conduct (verb)
- Misplaced stress can lead to perception differences, especially for non-native speakers.
- Mondegreens: misheard lyrics or phrases that differ from the actual wording (popularized by misheard songs).
- Common phenomenon: hearing a phrase that seems plausible but is not what was sung or spoken (e.g., misinterpretations in popular culture).
- Steven Pinker and others have discussed Mondegreens as a demonstration of perceptual interpretation vs. actual speech input.
- Coarticulation: overlap between adjacent sounds during natural speech, making segment boundaries less distinct.
- Examples:
- mashed potatoes: the /d/ in mashed can coarticulate with the start of "potatoes".
- handbag: often pronounced as "ham-bag" due to coarticulation of /b/ with preceding /d/ or /g/ sounds; more efficient speech
- six vs sixth: speakers may produce overlapping sounds that merge, producing a clipped or blended pronunciation
- Coarticulation is common at natural speech rates but can be reduced when speaking slowly.
- Context and top-down processing:
- Our knowledge of words and world knowledge influences perception earlier than you might expect.
- Example: Hambag vs handbag shows how top-down expectations bias perception toward a meaningful word.
- Eye-tracking and contextual experiments show that listeners use context rapidly to guide interpretation, sometimes before all auditory information is fully processed.
- A landmark 2014 study used an array with four objects and measured fixation times as sentences were heard; listeners rapidly matched heard nouns to seen objects if contextual cues supported the target.
- Phonemic restoration (speech perception under occlusion):
- When parts of a sentence are replaced by noise, listeners often report hearing a plausible completion consistent with context.
- Classic findings: listeners report hearing wheel when wheel is suggested by context, even if the actual sound was masked; this demonstrates strong top-down influence.
- Illustrative examples: hearing "eel on the axle" when the sentence would have included a word like "wheel"; similarly, other restored segments show the brain filling gaps based on expectations.
- Developmental and cross-linguistic considerations:
- Universal discriminators: until about age 8–9 months, all humans can discriminate sounds across languages; later, exposure tunes perception to the native language(s).
- Bilingualism: the linguistic community now uses bilingual to describe anyone who speaks or understands more than one language; contrast with earlier notions of bilingualism being strictly two languages.
- Tonal languages present additional perceptual demands: non-native listeners often struggle to distinguish tone-based phonemes if those tones are not part of their linguistic repertoire.
- Social and ethical issues: research on phoneme discrimination has faced controversy due to potential bias and stigma; the scientific goal remains understanding perception, not judging language varieties or speakers.
- Language and phonology basics refreshed:
- Phonemes are the smallest distinguishing sounds; syllables are rhythmic units often containing multiple phonemes.
- Graphemes: letters or letter groups that represent phonemes (e.g., ph is two graphemes representing one phoneme /f/).
- Syllables and phonemes are not always aligned with spelling; mapping between letters and sounds can be inconsistent (e.g., PSY in psychology).
- A spectrogram can reveal the distribution of energy across time and frequency and helps illustrate where phonemes and syllables occur within a stream of speech.
- Practical and ethical implications for everyday life:
- Recognizing the role of context and top-down processing helps explain why people occasionally misunderstand, especially in noisy settings or when listening to unfamiliar accents.
- Clinicians and educators should consider environmental acoustics and speech rate when assessing or teaching language perception.
- Awareness of biases: avoid judgments about speakers based on pronunciation or dialect; focus on understanding mechanisms rather than evaluating language variety.
Developmental Timelines and Language Diversity
- Universal discriminators up to about 8–9 months, after which tuning to home language begins.
- Monolingual vs bilingual speakers:
- In the past, monolinguals were more common; today, bilingualism is widespread and often expected in many communities.
- Bilingualism involves exposure to multiple phonemic repertoires and can alter perceptual boundaries and categorization across languages.
- Tonal languages:
- Pitch differences (tones) carry lexical meaning in many languages; non-native listeners may struggle to distinguish tones if not trained or exposed to those languages.
- Social considerations in research:
- Studies in speech perception should avoid stigmatizing judgments about language varieties; the science focuses on perceptual processes, not social judgments about speakers.
Phonemes, Syllables, and the Spectrogram: Quick Recall Questions
- How many phonemes in the word yacht?
- How many syllables in yacht?
- How many syllables in psychology?
- Answer: 4 (psych o l o gy)
- How many phonemes in psychology? Common answers range from 8 to 9 depending on dialect and analysis.
- What is a spectrogram showing in a sentence like "Joe took father’s shoe bench out"?
- The dark bars indicate higher energy (intensity) at certain frequencies; vowels and some consonants (e.g., /b/, /sh/) show stronger energy; other consonants (e.g., /k/) may appear as lighter energy due to lower acoustic energy.
- What is coarticulation? Give examples.
- Coarticulation is the overlap of articulatory gestures between adjacent sounds, which can blur boundaries between phonemes or syllables (e.g., "mashed potatoes", "handbag", or the pronunciation of "sixth" as a blend with the following sound).
- What is Mondegreen? Give an example.
- Mondegreen is when you mishear a line and reinterpret it plausibly (e.g., hearing "there’s a bathroom on the right" in Bad Moon Rising when that’s not the original lyric).
- What are the main factors that influence segmentation in everyday speech?
- Phonotactic constraints (which sound sequences are possible), stress patterns, word boundaries, and contextual knowledge.
- Phonemes per second in typical adult speech: approximately
- 10extto20extphonemespersecond
- Processing speed: perception occurs in microseconds, with a suggestion that some aspects may approach nanoseconds in real-time processing discussions. (Qualitative reference rather than a fixed numeric value.)
- Children’s phonemic discrimination development: universal discrimination up to roughly 8–9 months, then tuning to native language sounds.
Connections to Foundational Principles and Real-World Relevance
- Foundational principles:
- Perception combines bottom-up acoustic information with top-down knowledge (context, world knowledge, expectation-driven processing).
- Segmentation, CP, and coarticulation illustrate how the brain organizes continuous speech into meaningful units.
- Spectrograms provide a tangible link between acoustic signals and perceptual categories.
- Real-world relevance:
- Acoustic design in public spaces can support speech intelligibility (e.g., reducing reverberation, using soft furnishings for noise absorption).
- Language teaching and speech therapy benefit from understanding CP, segmentation cues, and coarticulation, as well as the impact of speech rate and prosody.
- Recognition that perspective-taking and bias-aware communication are important in research and education to avoid stigma around dialects or language varieties.
Summary Takeaways
- Speech perception is fast, context-sensitive, and relies on a blend of bottom-up acoustic cues and top-down knowledge.
- Key processes include decoding acoustic features, segmenting sounds into phonemes and words, and integrating with phonological knowledge to derive meaning.
- CP shows that we categorize sounds rather than perceiving a continuous spectrum, with development shaped by language exposure.
- Segmentation is guided by phonotactics, stress, and context; coarticulation reflects natural, efficient speech production and can complicate perception.
- Context, expectation, and phonemic restoration demonstrate that perception often fills in missing information using knowledge about the world and language.
- Environmental factors and individual differences influence speech perception, with practical implications for education, therapy, and acoustic design.
End-of-Lap Notes and Future Topics
- Tomorrow’s session will continue exploring CP and the limits of speech perception, including deeper dives into segmentation cues and experimental paradigms for understanding top-down vs bottom-up processing.
- Practical exercise: try listening to speech in noisy environments and observe how slowing speech or providing contextual cues can improve intelligibility.