Learning Processes, Behaviour and Performance - Recap (Video Notes)
What is Behaviour?
- Behaviour defined as any outward or inward response to a stimulus
- Includes actions, speech/language, feelings, and thoughts
- All behaviour is elicited by a stimulus, whether observed or unseen, conscious or unconscious
Part 3: Changes in Behaviour from Experience
- Repeated Stimulation
- Repeated exposure to the same stimulus has two effects:
- Habituation: decline in responding
- Sensitisation: increase in responding
- These changes are considered non-associative learning
Habituation
- Startle magnitude experiments
- Trials: 1 per day, every 3 seconds
- Observations: long-term habituation vs. short-term habituation
- Spontaneous recovery observed
- Examples of habituation
- Startle response: part of an organism’s defensive reaction to threat
- Used to study fear and learning
- Commonly studied in rats, mice, rabbits, humans using loud tones or bright lights
- Additional examples
- Salivation and hedonic ratings
- Stimulus-specific and attention-dependent
- Relevant for controlling eating behaviour
- Involve trials and dishabituation
- Infants and facial recognition
- Infants habituate to familiar faces
- Preferential, longer viewing time for novel faces, even when faces are shown at different orientations
- Habituation parameters
- With the same exposure to the stimulus, habituation occurs more rapidly with shorter interstimulus intervals
- With the same number of presentations and interstimulus interval, habituation occurs more rapidly if the stimulus is presented for longer periods at each presentation
- With the same interstimulus interval and duration, habituation occurs more rapidly if the stimulus is presented more frequently
Part 5: Emotions and Motivated Behaviour
- Emotions can be elicited by stimuli in addition to external responses
- Habituation and sensitisation extend to changes in emotions and forms of motivated behaviour (e.g., feeding, exploration, aggression)
- Drug addictions are a particular area of interest
- Solomon & Corbit (1974): biphasic intense emotional reactions
- One emotion during elicited stimulus; the opposite emotion after stimulus termination
- Emotional reactions change with experience
- Primary reaction weakens; after-reaction strengthens
- Habituation = drug tolerance
- Extends to other emotion-arousing stimuli (e.g., love, attachment)
- Opponent Process Theory (OP) (neurophysiological homeostasis)
- Emotions are linked to neurophysiological mechanisms that maintain homeostasis
- Behaviour is the net result of:
- Primary process (direct effect of the emotion-arousing stimulus)
- Opponent process (counteracts the direct effect)
- Primary process (P-process) and Opponent process (O-process)
- P-process: quality of the emotional state in presence of the stimulus
- O-process (b-process): generates the opposite emotional reaction; lags behind the primary disturbance
- Underlying opponent processes produce the manifest affective response observed during trials
- After extensive stimulus exposure
- Primary reaction weakens; after-reaction strengthens
- Similar to habituation: primary reaction wanes, after-reaction grows
- Addictive implications
- Frequent users may not experience high highs
- Addictions may serve to reduce aversiveness of after-reaction (aversive withdrawal states)
- Neurophysiological evidence: reduced activation in reward circuits; increased activation in anti-reward circuits
- Conditioning and drug tolerance
- Can be applied to conditioned drug tolerance (a form of associative learning)
- CS (cue) paired with US (drug effects)
- After repeated CS–US pairings, the B (opponent) response activates upon CS alone
- Tolerance forms as the body anticipates the US (needs more of the drug to achieve initial UR levels)
- Novel stimuli (e.g., different environment) can disrupt homeostasis and increase overdose risk
BASIC CONCEPTS OF CLASSICAL CONDITIONING (CC)
Classical Conditioning Paradigm
- Before conditioning
- Unconditioned stimulus (US) naturally provokes an unconditioned response (UR)
- Neutral stimulus (NS) is present but produces no conditioned response (CR)
- During conditioning
- CS is paired with US, leading to acquisition of a conditioned response (CR)
- After conditioning
- CS elicits CR similar to UR
- Common examples
- Food → Salivation (UR); Salivation becomes CR to CS (e.g., bell) over conditioning
- Sexual arousal as CR to CS
- Shock → Startle as UR; CS paired with shock leads to CR (startle) to CS
Extinction (Inhibitory Conditioning)
- Extinction = stop reinforcing a previously excitatory CS
- Inhibitory associations = CS predicts the absence of the outcome
- Extinction learning is slower than acquisition
Inhibitory Conditioning Procedures
- Pavlovian/CS-based procedures
- Variants include CI (inhibition of delay US), Explicitly unpaired, Backward conditioning, etc.
- A+ / AX- style designs illustrate how one cue can inhibit another when trained in compound
- Differential and CI (inhibition) approaches demonstrate inhibitory learning
Factors that Influence the Effectiveness of Conditioning
- Contiguity (closeness in time) between CS and US
- Contingency (predictive relationship) between CS and US
- CS pre-exposure (Latent inhibition): pre-exposure to CS slows acquisition
- Overshadowing: more salient CS dominates learning over less salient CS in a compound
- Blocking: a previously learned CS blocks learning to a new CS when both are paired with the US
- Salience: the conspicuousness or prominence of a CS affects learning rate
- Latent inhibition and US pre-exposure show context-specific effects
Higher-Order Conditioning
- Higher-order conditioning (SOC) allows learning without a direct CS–US experience
- Second-order conditioning: CS1 → CS2 → US chain where CS2 can elicit CR via CS1
- Sensory Preconditioning (SPC): Phase 1 stimulus pairings create S–S associations that later support CS–US conditioning
- Shoulder of debate: extinction of first-order stimuli can affect second-order responding; SPC may rely on S–S rather than S–R associations
Rescorla and Wagner Model (1972)
Key idea: learning is driven by total error reduction; all cues present on a trial contribute to prediction
Model accounts for cue competition phenomena (e.g., blocking, overshadowing) better than contiguity alone
Learning rule (simplified):
\Delta V = \alpha\beta\big(\lambda - \sum V_n\big)
Where:
- \Delta V: change in associative strength for a CS on a given trial
- \lambda: maximum associative strength (salience of US)
- \sum V_n: total associative strength of all CSs present on that trial
- \alpha: learning rate/associability of the CS (0 < a < 1)
- \beta: learning rate/associability of the US (0 < b < 1)
The model can be written without explicit a and b in some presentations as a form of total error reduction
Example (blocking scenario):
- If cue A has value 0.8 (A = 0.8) and X is introduced in compound with A
- Trial-wise ΔV calculations show predictive error becomes zero after a few trials, blocking X from gaining associative value
- Demonstrates that the total associative strength cannot exceed the US value
Example (overshadowing):
- If cue A is more salient than X, A gains most of the US value early, leaving little for X
- Demonstrates that a more salient cue can overshadow a less salient cue in learning
Problems with RW Model
- Spontaneous recovery, latent inhibition, etc. indicate limitations of a simple error-c-correction framework
Reinforcement and Punishment
- Role of the outcome: according to the Law of Effect, rewarding outcomes strengthen S-R associations; punishing outcomes weaken them
- Reinforcement: increasing operant responding (e.g., food)
- Punishment: decreasing operant responding (e.g., shock)
- Both reinforcement and punishment can involve giving or removing an outcome
- Outcomes defined:
- Positive reinforcement: producing the operant response yields appetitive outcome
- Negative reinforcement: producing the operant response results in the removal of an aversive outcome
- Positive punishment: producing the operant response yields an aversive outcome
- Negative punishment: producing the operant response yields removal of an appetitive outcome
Instrumental Conditioning Procedures
- Positive reinforcement: instrumental response produces outcome; strengthens behaviour
- Negative reinforcement: instrumental response removes a punishing outcome
- Punishment: instrumental response produces an aversive outcome (reduces behaviour)
- Negative punishment: instrumental response removes an appetitive outcome
- Omitted or Differential Reinforcement of Other behaviour (DRO) as a form of negative punishment
Simple Schedules of Reinforcement
- Fixed Ratio (FR): reinforcement after every nth response
- Variable Ratio (VR): reinforcement after an average of n responses; variability around the mean
- Fixed Interval (FI): reinforcement available after a fixed amount of time has elapsed; the first response after that interval is reinforced
- Variable Interval (VI): reinforcement available after a variable/time-average interval; the time to reinforcement varies
- Schedule effects
- Steady, high response rates under ratio schedules with post-reinforcement pauses (ratio runs)
- The length of the post-reinforcement pause correlates with the number of required responses
- Continuous Reinforcement (CRF): every instance of the behaviour is rewarded; effectively FR1
- Leads to steady, moderate response with brief pauses; relates to compulsive behaviours
Motivational Mechanisms
- Three-term contingency: Context (S) – Instrumental response (R) – Outcome (O)
- S-R association: driven by the Law of Effect; reinforcer stamps in the S-R link
- S-R-O: the basic motivation to perform the instrumental response via contextual cues
- S-O or S-S associations: learning that outcomes are expected or anticipated by cues
- Two-process theory (Hull, Spence): instrumental responses increase because S evokes the response directly via S-R; or via S-O expectancy that motivates the response
- Two-process theory suggests both S-R and S-O/motivational expectancy contribute to instrumental performance
Premack Principle and Response Deprivation
- Premack Principle: higher-probability (more likely) activities can reinforce lower-probability (less likely) ones
- If H (high probability) follows L (low probability) after an opportunity, L is reinforced
- If H follows L, H does not reinforce L; essentially, the more probable activity can serve as a reinforcer for the less probable one
- Applications
- Eating, sex, etc., are not inherently special beyond their likelihood
- Explains individual differences (e.g., videogaming as a reinforcer)
- Response-Deprivation Hypothesis
- Restriction of access to a normally available activity creates reinforcement for other activities
- Instrumental procedures inherently restrict access to some reinforcer, affecting its reinforcing value
Extinction (Revisited)
- Extinction in CC: CS no longer predicts the US; CR diminishes
- Extinction in Instrumental Conditioning: CR in the presence of CS no longer leads to US
- Interference effects include outcome interference, cue interference, punishment, etc.
- Extinction is not erasure but new inhibitory learning specific to context
Recovery from Extinction
- Extinction is not permanent; recovery of response can occur after extinction
- Mechanisms include:
- Spontaneous recovery: return of responding after time without training
- Renewal: recovery when testing occurs in a different context (ABA, ABC, AAB)
- Reinstatement: return after re-exposure to US
- Bouton’s Theory of Retrieval (1993)
- Extinction reflects retroactive outcome interference; second-learned X–O2 disrupts retrieval of X–O1
- Proactive interference: test context not similar to extinction context can lead to proactive interference
- Retrieval theory links context, time, and space to recovery
Generalisation vs Discrimination
- Discrimination: determining whether instrumental behaviour is controlled by a specific stimulus; different aspects of the same stimulus control different responses
- Generalisation: responding to similar stimuli; the extent of generalisation is shown by gradient steepness
- Generalisation gradients illustrate stimulus control: steeper gradient indicates strong control; shallow gradient indicates weaker control
Evaluative Conditioning
Introduction to Evaluative Conditioning
- Evaluative Conditioning (EC): conditioning of likes and dislikes via association with positive or negative stimuli
- Similar to classical conditioning but the outcome is a change in attitude/valence rather than a physiological reflex
- Examples involve preference changes (e.g., tea preference) rather than reflexive responses
Generality of Evaluative Conditioning
- Acquisition of likes/dislikes via associative processes
- Classical conditioning: neutral CS paired with US elicits preparatory/defensive responses (reflexes)
- Evaluative conditioning: the valence of a CS changes due to US associations; CS becomes liked or disliked
- CS = neutral stimulus; US = affective stimulus; DV = changes in valence or liking
Models of Evaluative Conditioning
- Conceptual-Categorisation Account (Davey, 1994)
- EC is not purely associative learning but concept learning
- CS contains both likeable and unlikable elements; pairing with US highlights congruent features
- CS recategorised based on features, not solely on affective valence
- Problems: cross-modal conditioning and US revaluation effects
- Holistic Account (Martin & Levey, 1970s)
- EC is a basic form of learning shared across species
- CS-US presentation yields a holistic representation including the evaluative nature of the US
- CS evokes US representation, explaining resistance to extinction and contingency awareness
- Problems: inconsistent with sensory preconditioning
- Referential Account (Baeyens et al., 1992)
- Classical conditioning: associative changes in preparatory responses to CS
- Evaluative conditioning: associative changes to the valence of the CS
- Distinguishes signal/expectancy learning (classical) from affective evaluative learning (referential)