Speech Perception - Key Concepts and Examples

Speech Perception: Overview

  • Speech perception is a key component of language processing that differs from other auditory perception (e.g., distinguishing a siren from a human voice even at the same loudness, pitch, and frequency).
  • It happens incredibly fast: we perceive speech while someone is talking or singing, often with speed and efficiency that feel almost instantaneous.
  • A central goal in psychology is to understand how humans do this and why some people find it easier than others.
  • Today’s focus: main component parts of speech perception, features of speech perception, and specifically:
    • categorical perception
    • segmentation
    • context/environment effects on perception
  • The process is studied through varied examples and experiments to reveal underlying mechanisms and brain processes.

Key Concepts

  • Speech perception vs. other auditory signals: even with similar acoustic properties, humans distinguish natural speech from environmental sounds or music, often automatically.
  • The left hemisphere is heavily involved in language processing for most people (language centers typically left-lateralized).
  • Speech processing may be akin to a specialized form of music in some respects (e.g., prosody and the way we modulate pitch and rhythm when communicating, especially with children).
  • Child-directed speech (formerly called motherese) features higher pitch, slower rate, exaggerated vowel lengthening, and a sing-song quality that facilitates infant speech processing.
  • The brain appears to integrate music-like processing and speech processing in overlapping regions.

Examples of Speech Perception in Action

  • Auctioneers (International Auctioneer Championship): rapid, continuous, highly-contextual speech makes it hard to hear exact words, yet listeners can infer bids from context, gesture, and rhythmic cues.
    • Example: a competition where contestants bid on lots (e.g., cutting boards) with rapid, overlapping numbers and phrases.
    • Observations: listeners can still track intent and respond despite limited audible content; highlights the role of real-time context and expectations.
  • Icelandic football commentary (Euro 2016): despite language barriers, listeners can understand excitement and gist through prosody and context.
    • Note: even without understanding every word, the emotional and prosodic cues convey meaning.
  • Overall point from these examples: speech perception relies on rapid decoding, pattern recognition, and top-down expectations to derive meaning from partial acoustic information.

Brain and Language Processing

  • The left hemisphere often dominates language processing, including speech sounds.
  • Context and language familiarity influence perception:
    • Speech can be treated as a special kind of music, especially in how we modulate pitch, tempo, and intonation when addressing infants.
    • Child-directed speech shows how prosodic features facilitate processing of language in early development.
  • A looping auditory example: repeating phrases can start to sound like singing, illustrating overlapping mechanisms between speech and music processing.
  • The perception of speech is influenced by context and statistical regularities, not just the raw acoustic signal.

Perceptual Phenomena and Tools

  • Categorical Perception (CP): we hear sounds as belonging to discrete categories (e.g., /b/ vs /d/) rather than along a continuous gradient.
    • Classic demonstration: a continuum between /b/ and /d/ yields a crisp boundary where listeners report hearing either /b/ or /d/ with little or no perception of intermediate sounds.
    • Example from class: a synthesized sound between ba and da is heard as one category or the other, not as a middle sound.
    • Note: CP suggests our brains map continuous acoustic variation onto discrete linguistic categories.
  • Phonemes vs. Syllables:
    • Phonemes: the smallest units of sound that distinguish meaning in a language.
    • Syllables: rhythmic units that may include one or more phonemes.
    • Example: the word yacht has 3 phonemes but only 1 syllable.
    • Example: psychology has variable phoneme counts (commonly eight, occasionally nine, depending on dialect and analysis).
  • Phonetics and Phonology:
    • Phonemes are the building blocks; their actual realization can vary (allophones) and may be mapped differently across languages.
    • A phoneme can be realized by multiple letters (e.g., "ph" represents the /f/ sound but is written with two letters).
    • The term phonetician refers to a specialist who studies phonetics.
  • Spectrograms: a visual representation of the speech signal showing time on the horizontal axis, frequency on the vertical axis, and darkness (intensity) representing energy.
    • Example: a sentence like "Joe took father's shoe bench out" shows stronger energy for vowels and certain consonants (e.g., /b/, /sh/) and weaker energy for others (e.g., /k/ in "took").
    • Spectrograms help illustrate where syllables and phonemes are likely located and how coarticulation appears as overlapping energy patterns.
  • Processing pipeline (in real time):
    • Auditory input -> decode the features -> segment sounds -> perceive and categorize phonemes -> integrate into phonological representations -> determine word status and meaning -> use knowledge to interpret the message.
    • This entire sequence unfolds in microseconds; some discussion touches on nano- to microsecond time scales.
  • Individual differences that affect perception:
    • Speaker differences: sex, dialect, speaking rate, accent.
    • Rapid production can challenge perception, but humans remain highly adept at comprehension in many conditions.
  • Adverse listening conditions:
    • Environments with noise, reverberation, hard surfaces (cafés, large halls) can degrade perceptual clarity.
    • Acoustic design considerations (e.g., fabric in seats at opera houses) can influence perception by shaping reverberation and noise.
  • Segmentation: the ability to divide continuous speech into discrete words.
    • Not taught explicitly; we infer word boundaries from phonotactics and context.
    • Clues to segmentation include impossible or unlikely letter-sound combinations within a single word (e.g., /cf/ cannot occur inside a single English word, indicating a boundary between segments).
    • Examples: fapple vs waffle; boundary after the syllable pattern of a word; stress patterns can signal word boundaries.
  • Stress patterns and lexical boundaries:
    • English stress shifts can change word meaning (e.g.,
    • record (noun) vs record (verb)
    • conduct (noun) vs conduct (verb)
    • Misplaced stress can lead to perception differences, especially for non-native speakers.
  • Mondegreens: misheard lyrics or phrases that differ from the actual wording (popularized by misheard songs).
    • Common phenomenon: hearing a phrase that seems plausible but is not what was sung or spoken (e.g., misinterpretations in popular culture).
    • Steven Pinker and others have discussed Mondegreens as a demonstration of perceptual interpretation vs. actual speech input.
  • Coarticulation: overlap between adjacent sounds during natural speech, making segment boundaries less distinct.
    • Examples:
    • mashed potatoes: the /d/ in mashed can coarticulate with the start of "potatoes".
    • handbag: often pronounced as "ham-bag" due to coarticulation of /b/ with preceding /d/ or /g/ sounds; more efficient speech
    • six vs sixth: speakers may produce overlapping sounds that merge, producing a clipped or blended pronunciation
    • Coarticulation is common at natural speech rates but can be reduced when speaking slowly.
  • Context and top-down processing:
    • Our knowledge of words and world knowledge influences perception earlier than you might expect.
    • Example: Hambag vs handbag shows how top-down expectations bias perception toward a meaningful word.
    • Eye-tracking and contextual experiments show that listeners use context rapidly to guide interpretation, sometimes before all auditory information is fully processed.
    • A landmark 2014 study used an array with four objects and measured fixation times as sentences were heard; listeners rapidly matched heard nouns to seen objects if contextual cues supported the target.
  • Phonemic restoration (speech perception under occlusion):
    • When parts of a sentence are replaced by noise, listeners often report hearing a plausible completion consistent with context.
    • Classic findings: listeners report hearing wheel when wheel is suggested by context, even if the actual sound was masked; this demonstrates strong top-down influence.
    • Illustrative examples: hearing "eel on the axle" when the sentence would have included a word like "wheel"; similarly, other restored segments show the brain filling gaps based on expectations.
  • Developmental and cross-linguistic considerations:
    • Universal discriminators: until about age 8–9 months, all humans can discriminate sounds across languages; later, exposure tunes perception to the native language(s).
    • Bilingualism: the linguistic community now uses bilingual to describe anyone who speaks or understands more than one language; contrast with earlier notions of bilingualism being strictly two languages.
    • Tonal languages present additional perceptual demands: non-native listeners often struggle to distinguish tone-based phonemes if those tones are not part of their linguistic repertoire.
    • Social and ethical issues: research on phoneme discrimination has faced controversy due to potential bias and stigma; the scientific goal remains understanding perception, not judging language varieties or speakers.
  • Language and phonology basics refreshed:
    • Phonemes are the smallest distinguishing sounds; syllables are rhythmic units often containing multiple phonemes.
    • Graphemes: letters or letter groups that represent phonemes (e.g., ph is two graphemes representing one phoneme /f/).
    • Syllables and phonemes are not always aligned with spelling; mapping between letters and sounds can be inconsistent (e.g., PSY in psychology).
    • A spectrogram can reveal the distribution of energy across time and frequency and helps illustrate where phonemes and syllables occur within a stream of speech.
  • Practical and ethical implications for everyday life:
    • Recognizing the role of context and top-down processing helps explain why people occasionally misunderstand, especially in noisy settings or when listening to unfamiliar accents.
    • Clinicians and educators should consider environmental acoustics and speech rate when assessing or teaching language perception.
    • Awareness of biases: avoid judgments about speakers based on pronunciation or dialect; focus on understanding mechanisms rather than evaluating language variety.

Developmental Timelines and Language Diversity

  • Universal discriminators up to about 8–9 months, after which tuning to home language begins.
  • Monolingual vs bilingual speakers:
    • In the past, monolinguals were more common; today, bilingualism is widespread and often expected in many communities.
    • Bilingualism involves exposure to multiple phonemic repertoires and can alter perceptual boundaries and categorization across languages.
  • Tonal languages:
    • Pitch differences (tones) carry lexical meaning in many languages; non-native listeners may struggle to distinguish tones if not trained or exposed to those languages.
  • Social considerations in research:
    • Studies in speech perception should avoid stigmatizing judgments about language varieties; the science focuses on perceptual processes, not social judgments about speakers.

Phonemes, Syllables, and the Spectrogram: Quick Recall Questions

  • How many phonemes in the word yacht?
    • Answer: 3
  • How many syllables in yacht?
    • Answer: 1
  • How many syllables in psychology?
    • Answer: 4 (psych o l o gy)
  • How many phonemes in psychology? Common answers range from 8 to 9 depending on dialect and analysis.
  • What is a spectrogram showing in a sentence like "Joe took father’s shoe bench out"?
    • The dark bars indicate higher energy (intensity) at certain frequencies; vowels and some consonants (e.g., /b/, /sh/) show stronger energy; other consonants (e.g., /k/) may appear as lighter energy due to lower acoustic energy.
  • What is coarticulation? Give examples.
    • Coarticulation is the overlap of articulatory gestures between adjacent sounds, which can blur boundaries between phonemes or syllables (e.g., "mashed potatoes", "handbag", or the pronunciation of "sixth" as a blend with the following sound).
  • What is Mondegreen? Give an example.
    • Mondegreen is when you mishear a line and reinterpret it plausibly (e.g., hearing "there’s a bathroom on the right" in Bad Moon Rising when that’s not the original lyric).
  • What are the main factors that influence segmentation in everyday speech?
    • Phonotactic constraints (which sound sequences are possible), stress patterns, word boundaries, and contextual knowledge.

Formulas and Numerical References

  • Phonemes per second in typical adult speech: approximately
    • 10extto20extphonemespersecond10 ext{ to } 20 ext{ phonemes per second}
  • Processing speed: perception occurs in microseconds, with a suggestion that some aspects may approach nanoseconds in real-time processing discussions. (Qualitative reference rather than a fixed numeric value.)
  • Children’s phonemic discrimination development: universal discrimination up to roughly 8–9 months, then tuning to native language sounds.

Connections to Foundational Principles and Real-World Relevance

  • Foundational principles:
    • Perception combines bottom-up acoustic information with top-down knowledge (context, world knowledge, expectation-driven processing).
    • Segmentation, CP, and coarticulation illustrate how the brain organizes continuous speech into meaningful units.
    • Spectrograms provide a tangible link between acoustic signals and perceptual categories.
  • Real-world relevance:
    • Acoustic design in public spaces can support speech intelligibility (e.g., reducing reverberation, using soft furnishings for noise absorption).
    • Language teaching and speech therapy benefit from understanding CP, segmentation cues, and coarticulation, as well as the impact of speech rate and prosody.
    • Recognition that perspective-taking and bias-aware communication are important in research and education to avoid stigma around dialects or language varieties.

Summary Takeaways

  • Speech perception is fast, context-sensitive, and relies on a blend of bottom-up acoustic cues and top-down knowledge.
  • Key processes include decoding acoustic features, segmenting sounds into phonemes and words, and integrating with phonological knowledge to derive meaning.
  • CP shows that we categorize sounds rather than perceiving a continuous spectrum, with development shaped by language exposure.
  • Segmentation is guided by phonotactics, stress, and context; coarticulation reflects natural, efficient speech production and can complicate perception.
  • Context, expectation, and phonemic restoration demonstrate that perception often fills in missing information using knowledge about the world and language.
  • Environmental factors and individual differences influence speech perception, with practical implications for education, therapy, and acoustic design.

End-of-Lap Notes and Future Topics

  • Tomorrow’s session will continue exploring CP and the limits of speech perception, including deeper dives into segmentation cues and experimental paradigms for understanding top-down vs bottom-up processing.
  • Practical exercise: try listening to speech in noisy environments and observe how slowing speech or providing contextual cues can improve intelligibility.