Learning Sound Patterns I
Sound Pattern Acquisition and the Challenges of Early Language Perception
The Problem of Continuous Acoustic Input: Infants face significant perceptual hurdles when acquiring language from acoustic signals:
Absence of Acoustic Gaps: Spoken language contains no physical pauses or gaps between words, unlike spaces in written text.
Lack of Abstract Concepts: Infants possess no initial conceptual understanding of what a "word" is.
Phonetic Variability: Speech sounds vary considerably depending on speaker voice, pitch, context, and time of production.
Contrastive Identification: Infants must determine which phonetic differences carry meaningful semantic distinction (e.g., versus ) and which variations are non-contrastive.
Sensory Overload: William James (1890) famously described early perceptual experience: "The baby, assailed by eyes, ears, nose, skin, and entrails at once, feels it all as one great blooming, buzzing confusion."
Timeline of Early Perceptual and Acoustic Achievements:
In Utero Perception: Fetal auditory experience captures the rhythmic properties, prosody, and melodic contours of native speech.
At Birth (Within days): Newborns discriminate their native language from unfamiliar languages. For example, French newborns exhibit higher non-nutritive sucking rates when listening to French compared to Russian (Mehler et al., 1988).
Pre-Verbal Auditory Segmentation: Long before productive word speech emerges (around \,months), infants learn to parse continuous speech into distinct sound units, categorize contrastive phonemes, and learn native phonotactic probabilities.
Infant Speech Segmentation Strategies
The Necessity of Words (Lexical Units):
Without segmentation into discrete words, every distinct utterance would require memorization as a holistic sound block (e.g., mapping a unique unanalyzed complex sound to "the dog bit the dog").
Words enable infinite recombination and linguistic productivity (Hockett).
Acoustic Realities of Word Boundaries:
Physical silence does not mark word boundaries in continuous speech streams.
In acoustic spectrograms (e.g., for the phrase "nineteenth century"), the primary silent gap occurs internally during the stop closure of the sound in "nineteenth", whereas zero acoustic gap exists between the words "nineteenth" and "century".


Head-Turn Preference Paradigm (HTPP):
Experimental Design: Used with infants aged \,to \,months (older toddlers fail to sit still).
Familiarization Phase: Infants listen to specific auditory passages or artificial streams while looking at central/side lights.
Test Phase: Familiar versus novel auditory stimuli are played from directional speakers. Looking time toward the speaker is measured.
Interpretation: Statistically significant differences in orientation duration (looking time) indicate successful stimulus discrimination and auditory memory categorization.

Empirical Demonstration of Word Extraction:
Jusczyk & Aslin (1995): Tested whether infants can isolate individual words from fluent passages.
Method: Exposed infants to passages containing target words (e.g., "His bike had big black wheels. The girl rode her big bike…"). Subsequently tested isolated target words ("bike") versus novel words ("dog").
Findings: -month-old infants listened significantly longer to isolated familiar words, demonstrating successful extraction from continuous speech streams. -month-old infants failed this task.
Segmentation Mechanisms and Environmental Cues:
Familiar Word Anchors:
Approximately of infant-directed speech consists of isolated single words (e.g., "Mommy", "Ella").
Infants leverage highly familiar anchor words to segment adjacent novel material (e.g., breaking
bankiritubendudifinintoban-kiri-tuben-dudi-fin).Bortfeld et al. (2005): Showed -month-old infants segment unfamiliar target words (e.g., "feet" or "bike") when placed adjacent to familiar name anchors (e.g., "Maggie's bike…", "Mommy's feet").
Phonotactic Constraints:
Definition: Language-specific rules defining permissible sequence combinations of sound segments at word beginnings, endings, or internal transitions.
English Example: Consonant cluster occurs internally (bandage) or word-finally (hand), but cannot occur word-initially. Thus, a continuous sequence like
Elgandokuis parsed by native English speakers asElgan dokurather thanElga ndoku.Cross-Linguistic Variations: Word-initial is legal in Czech; word-initial is legal in Swahili (ndela); English word-initial clusters are illegal in Spanish (requiring prosthetic vowel additions like espanish).
Developmental Timeline: Jusczyk et al. (1993) demonstrated that by \,months, American infants prefer legal English nonwords (cubeb, dudgeon) over legal Dutch nonwords (zampljes, vlatke), while Dutch infants show inverse preferences. Infants utilize these constraints for word segmentation by \,months (Mattys & Jusczyk, 2001).
Metrical Stress Patterns (Prosody):
English stress is predominantly trochaic (strong-weak/stressed-unstressed pattern, e.g., DOC-tor) rather than iambic (weak-strong/unstressed-stressed pattern, e.g., gui-TAR), appearing at a ratio of approximately .
Jusczyk et al. (1999): -month-old English-learning infants correctly segment trochaic words (DOC-tor) from continuous speech, but missegment iambic target phrases like "guiTAR is" as TAR-is.
Statistical Learning and Transitional Probabilities
The Bootstrap Dilemma: Relying exclusively on phonotactic rules or stress patterns creates a circular dependency: discovering structural rules requires a pre-existing vocabulary, yet constructing a vocabulary requires segmentation tools.
Mathematical Definition of Transitional Probability (TP):
Transitional probability measures the statistical likelihood of syllable following syllable :



TP Properties across Lexical Boundaries:
Within Words: Syllables internal to a word follow each other with high predictability (). For instance, inside the word pretty, the probability of ty following pre is high ().
Across Word Boundaries: Syllables spanning word boundaries are significantly less predictable ( drops dramatically). The syllable ty at the end of pretty can precede any initial syllable of subsequent words (e.g., bay, bo, nay), yielding low individual transition probabilities ().
Landmark Study: Saffran, Aslin, & Newport (1996):
Artificial Language Construction: Designed a controlled language comprising four 3-syllable nonsense words: bidaku, padoti, golabu, tutaba (or tupiro).
Auditory Familiarization Phase: -month-old infants were presented with a continuous -minute stream of synthesized computer speech containing each word repeated \,times in random order (bidakugolabudutabagolabupadotibidakudutabapadoti…).
Experimental Controls: Flat monotonic voice, absence of pauses, uniform pitch, absence of stress variations, uniform consonant-vowel (CV) syllable structure. This eliminated all prosodic, acoustic, and phonotactic cues, isolating statistical patterns.
Statistical Parameters: Within-word $TP = 1.0$; across-word-boundary $TP = 0.33$.
Testing & Results: HTPP evaluated listening times for intact "words" (bidaku, golabu) versus cross-boundary "part-words" (dakugo, buduta).
-month-olds demonstrated significantly longer mean looking times for part-words (\,seconds) compared to intact words (\,seconds).
Conclusion: Infants detect word boundaries based strictly on tracking conditional statistical probabilities between adjacent speech units.

Domain Generality and Evolutionary Implications of Statistical Learning
Cross-Species Statistical Tracking:
Cotton-Top Tamarins (Saguinus oedipus): Hauser et al. (2001) exposed tamarins to \,minutes of Saffran's artificial language stream. Tamarins exhibited significantly greater orienting head-turns toward speakers broadcasting part-word sequences.
Rats (Rattus norvegicus): Toro & Trobalón (2005) established that laboratory rats successfully discriminate words from part-words using transitional probabilities despite lacking language processing modules.

Cross-Modal and Early Developmental Pervasiveness:
Non-Linguistic Modalities: Humans apply identical statistical tracking mechanisms to musical tone sequences (Saffran et al., 1999) and visual geometric shapes presented in temporal streams (Fiser & Aslin, 2001).
Neonatal Statistical Learning During Sleep: Teinonen et al. (2009) played \,minutes of tri-syllabic artificial words to sleeping newborns under \,hours old. Event-Related Potentials (ERPs) revealed enhanced neural responses to unpredictable word-initial syllables relative to predictable internal syllables.
Theoretical Implications:
Statistical learning represents an ancient, domain-general cognitive computation mechanism rather than a human-unique, language-specific module.
Species Differences & Human Biases: Non-human species (e.g., rats) miss subtle prosodic/perceptual cues that human infants automatically integrate, and humans possess distinct perceptual biases regarding which structural patterns are easily learned.
Concept Verification and Empirical Applications
Contextual Word Extraction Check:
Scenario: A -month-old infant hears the string "Look at Mommy's book".
Inference: The infant is most likely to infer that "book" is a distinct word unit because it follows the highly familiar, pre-stored anchor word "Mommy's".
Phonotactic Parsing Check:
Scenario: English speakers presented with the novel string Elgandoku parse it as Elgan doku instead of Elga ndoku.
Mechanism: Native English phonotactic rules disallow word-initial clusters.
Methodological Isolation Control in Saffran et al. (1996):
Scenario: The artificial speech stream was delivered in a monotonic synthetic voice without pauses or stress.
Rationale: Eliminates prosodic, acoustic, and metrical cues, ensuring that transitional probability is the sole variable driving segmentation performance.
Transitional Probability Calculations in Continuous Speech:
Phrase: "the-ve-ry-red-car"
Highest TP Pair: The syllable transition ve ry possesses the highest transitional probability because it occurs internally within the word very (), whereas ry red spans a word boundary ().
Cross-Species Generalizability:
Finding: Laboratory rats successfully differentiate words from part-words based on statistical regularities.
Conclusion: Proves that transitional probability tracking is neither unique to human beings nor specific to linguistic domains.