Week 6: Mechanism of Learning -> Cue Competition
Compound Stimuli and Cue Competition
Compound Stimuli: Formed by presenting multiple cues at once
Cues presented in compound will often have identical levels of correlation and contiguity with the US
May not be learned about equally!
Cue Competition: When a compound CS is paired with a US, compound elements are competing with each other to be associated with the US
Elements split the association between themselves
One element might absorb ALL of the association
There might be equal association split
There might be unequal association split
Factors that determine how associations will split between compound elements
Overshadowing and the intensity of a CS
Informational Value and Relative Validity
Prior learning and blocking
Usually in learning irl there are rare times where only one CS is presented at a time, mostly there are a lot of cues happening at a time → important to learn this idea since its applicable to irl scenarios
Experimental Approach: Training an association as compound and then testing the elements individually and compare them
Test: Each stimulus is presented on its own to see the level of CR
Same contiguity and contingency, different learning across elements and compound groups

In the compound groups the associations are split between the elements(each would get 50% of the association for this example) when we present them at the same time
The element group shows that when the element is on its own it doesn’t have to compete to form the association
Two cues that are in compound predict the US, on their own as element they will predict a certain portion of the CR
Can be the case where one CS will not elicit a response and the other will take the entire response
Sensory Intensity and Overshadowing
What if one is more obvious to the animal when elements are in compound?
Experimental Example:
CS A: 40 db tone (quiet)
CS B: 80 db tone (loud)
Conditioning:
Group 1: A+ (CS A alone)
Will not have to compete
Group 2: AB+ (compound group)
Competition
One will strongly outcompete the other with different intensities
B+ alone elicits a STRONG response and A+ alone will have a weak response
Maybe the tone is so quiet that a response isn’t strongly gain but shows that the animal is still capable of learning the association
We know that there is competition going on since when A is alone then the CR is very strong
We would need to have different groups to understand the type of competition that is happening in the compound
Degree of overshadowing is easiest to predict when the stimuli have the same modalities (presentations) but have different intensities
The loud buzzer OVERSHADOWED the quieter tone conditioning
Overshadowing comes from relative stimulus intensity
Overshadowing: Conditioned Responding evoked by a CS is lower when the CS is paired as a compound CS than when paired as a element
A cause of overshadowing can be the different in the salience of the elemental CSs
This happened regardless of the contiguity or the contigency of both of the CSs
The degree of overshadowing is the easiest to predict when the stimuli have the same modality and are of different intensities
ex. the loud tone will overshadow, the brighter image will overshadow
We wouldn’t be able to readily guess which would overshadow if the modalities aren’t the same!
Important to have the quieter tone show examples of learning since it would show evidence that learning can still be made when the less contributing CS can still support normal pavlovian conditioning when it doesn’t have to compete
“Informational value” and relative validity
How much information does a CS have about a US to determine how the relationship will work
What happens when one element of a compound is reliable predictor of the US and another element of the compound is less reliable predictor of the US?
Two groups of animals (uncorrelated and correlated groups)
Correlated Group: the excitatory us is a 100% reliable predictor of the US
Only one US will predict the us 100%, one will predict 0%, other will predict 50%
Uncorrelated Group: Truly random design where no CS is a reliable predictor of the US
all CS’s will predict the US the same amount of times
no element will predict with 100% that the US will not be delivered

Correlated Group: A and B provide a lot of information regarding the US. C doesn’t
Uncorrelated Group: No CS is no better or no worse at predicting the US than another CS.
looking deeper at the information
the Conditioned response to X is different across the correlated and uncorrelated groups
When a CS is relatively BETTER predictors are the ones that are going to be associated with the US and creates the CR
Despite X have the same prediction for the correlated and uncorrelated groups, it will have a higher relative validity in the uncorrelated
Relative Validity refer to how good a CS is at predicting the US in comparison to other CSs for the scenario
For the uncorrelated version, since all of the CSs have the same amount of prediction they will all elicit the same response
When a CS has a higher relative validity it will OUTCOMPETE other CSs with lower validities
Overshadowing and “informational value”
Relative validity: If cues in a compound differ in how well they predict the US, the best predictor will outcompete the less-relaible predictors to associate with US
the strength of one CS’s association with the US depends on how good or bad OTHER cues are at predicting the US
Overshadowing and “informational value” in the real world
Symptoms in this example would be considered the CS
fever is a symptom of many illness (not very informative or a valid predictor)
if we see a certain rash it is highly predictive of only particular cause
Symptoms will be generic → provides little information
Specific symptoms → provide a lot of information about what illness it is (highly-predictive)
Experiment relies on the relative validity effect- poor predictos of a disease only INFLUENCE the diagnositc decisions when no better predictor of the disease exsited
Relative Validity: helps figure out which CSs to pay attention to based on how well the CS will predict the US
Blocking and prior training history
How does prior learning affect learning when the CS is added to a compound?
Blocking Group: learns simple Pavlovian Association (tone will produce a strong conditioned response
Both groups experiences the compound association
Blocking group has already learned something about tone
Results:
Control: Each element will elicit a conditioned response (both are learned about)
Blocking: Only the tone (the prior trained CS) will elicit learning and block the learning that could be done by the other CS
Prior learning of a CS will determine the learning that could be done by another CS when the CS is added in a compound
Blocking: one cues” prior training hisotry prevents a novel cue from being learnign about when the two are trained as compound
Will interfere with the learning during compound training
Kamin’s blocking experiment
Results show that little/no learning is done for the light (novel element) in the blocking group when added in a compound
Overshadowing group will show how the competition of the two CS’s will behave if there was no prior training of any of the CSs
we know that the CS is still capable of learning but the learning changing in the blocking group or when it’s in the overshadow group
Blocking: what’s going on here?
Previous learning to one cue can prevent other cues from being learned about later (get it BLOCKING)
Redundant cues aren’t learning about
Once a CS is predicted, there is no need to learn about more things that will predict
Learning system is using the most simple way to learn an association
Magnitude UNBlocking
When the addition of a new cue predict an increase in the intensity of the US, the novel cue in the compound is LEARNED → UNBLOCKED

The light in phase 2, group 2, will be learned about and unblocking since it is a novel stimulus that will be used to predict an increase in the intensity of the US
Unblocking happens when a novel stimulus is able to predict an increase of the magititude of the US
What happens when we decrease the magnititue of the US?…
Novel learning of the CS that will predict the decrease will undergo inhibitory learning
ex. the light cue will predict that the US that would’ve normally been learned about is NOT going to be learned
The prediction made is going to be for LESS US → will be the same as compound training for inhibitory learning
WE ARE CHANGING THE MAGNTITUDE
Identity unblocking
When the addition of a new cue predict a qualitative change in the US, the novel cue in the compound will be learned about → UNBLOCKED

Similar to magnitute unblocking but the cause is different
Results: in the unblocking group (group two) there will be a higher conditioned response
WE ARE CHANGING THE QUALITATIVE NATURE/IDENTITY OF THE CS BUT NOT THE MAGNITUDE
Trans-reinforcer blocking
Trans-reinforcer blocking: Blocking that persists across a qualitative change in the US
When we make a qualitive change to the US, the group that doesn’t undergo learning will have a higher conditioned response to the CS than the group that underwent prior learning
Interpretation: Showed that the learning system thinks that the change isn’t worth learning about so it didn’t learn a different association
Some association are more about the motivational content of the US than its specific identity
Learning is not big enough for learning to be done
Interpreting blocking/unblocking results PRACTICE THISSS!!

Blocking design will follow the same general scheme
Blocking: No learning is done for a novel CS
Unblocking: Learning is done a for a novel CS
Summary: Blocking
Modern framework for studying associative learning has its genesis in three major experiments
Resocrls correlation work _> showed that contiguit is not enough for learnign
Associations will not work if its not reliable
Wagners” relative validity → showed that cues will compete with each other based on information value
Associations depend on how good or how bad a predictor is which will overtake the associative relationship
kamins blcoking → redundant cues will not be learned about
Learn minimal amount of information they need to build an association
Effect are foundation for modern theories of associative learning
Cue Competition Summary
Multiple CSs paired in compound with an outcome will not create the same equal conditioned response
More Salient cues will outcompete less salients ones (overshadwoing by relative stimulus intensity)
Better predictors of the US outcompete poor predictions (informationtional value/relative validity)
Redundant cues are not usually learned about (Kamin’s Blocking)
Theoretical Implications
Simple Pavlovian learning isn’t passive
External factors: correlation/contiguity, relative validity, etc.
Internal factors: previous learning, sensory perception, etc.
In real-life, most learning involves compound stimuli, so cue competition effects show
up everywhere.
Understanding the computational basis of learning
What is a model of learning: A set of instructions for learning assocation (rules for how learning works)
Why? if we understand associative learning, we should be able to build an artificial model for it
Allows us to be specific of the model and think critically about each process of the step
Think throughouly
Allows us to make new predictions and check if the predictions are true
Allows us to identify a knowledge gap
The RW Model
Lamda:
Vcs associative strenght between a particulat x and a US in an experiement
How strongly an indivudal CS will [redict a US
Tells the indiviual how much US they should expect from a specific CS
Every CS will have its own associative strength
Changes over time as learning occur since the associaitve strength of a CS will change
Reflects the understanding of what the CS means
Tells us how much the associative strength of a particular Cs should changed given the experimental trial
second equation is used to get a new value for the associative strength
might need to change the associative strength by the end of the trial
adding the changing to the associative strength
Lambda represents the amount of US delivered on a trial ( the size or magnitude of a US)
Same units as the unconditioned stimulus
Lambda sets the assymptote of learning
point the learning curve will actually get
We expects the associative strengths to reach the value of lambda at the peak of the learning curve
We expect this to be the peak of learning
Vsum:the total amount of US predicted by all of the CSs present on a current trial
adding up the associative strength of all the cues in a trial to get the total associative strengths
we would add up indidivdual associative strengths to get the Vsum
gives us the total unconditioned stimulus expectation
all we care about the stimuli that are presented (the cues) regardless if the US is delivered or not! j
Lambda - Vsum (prediction error)
uncondiionted stimulus (what actually happened) - total associative strength (expectation)
quantitfies how good or how bad a subjects prediction was for that trial
when error is 0 = no error in prediction
when error is less than error = subjects were expecting more than what was delivered
when error is greater than error = subjects were expecting less that what was delivered
Prediction error can be thought of how surpising the US was to the subject during that specific tria;
If the outcome was as expected, nothing is surprsing
if the outcome was surprising then the prediction should be adjusted
We should adjust the associative strengths to make the predictions more accurated
Alpha cs
How intense the CS actually is
ex. loud and quiet tones being used
A more obvious CS is going to have a higher alpha value
value ranges from 0 to 1
Alpha controls the learning rate! → it will scale the prediction error
The alpha value will usually be given
We can predict that the alpha from a loud tone will be a higher value than the alpha of a quiet tone
We could also experimentally determine this alpha value (reasoning generally based on the salience of the CSs)
The amount of learning per trial for each CS (delta CS) is set by two things:
The prediction error
how good or how bad a prediction of a US was during that trial
will affect every CS that was present in the same way
The salience of each CS that was present on the trial
CSs might differ in their alpha value
Acquistion: Wagner Model
Plotting the change in the associative strength
strength will start at zero since there is no association being made
increases in the level of prediction through each trial
By the end of learning, the associative strength will be close to the value of lambda
Each point on the graph is (lambda - V) (prediction error)
Looks like a typical learning curve during Palovian Learning
Conditioned responding on every trial is propportion to the associative strengh of the cue on the trial
We assume that behavior is driven by the associative strength of the cue
Acquisition Example
Experiment: tone → food
Lambda value: 1
a tone = 0.35 (based on the sensory salience of the CS)
V old: associative strength before the trial
V new: associative strength after the trial
Assume that a novel CS will have an associative strength of zero before any learning has occurred
can round two decimal place and doesn’t really change the results
Chart how the associative strength of the CS changes through the different trials!


We will use the equations to compute any missing values
steps:
Consider what is happening during the trial
predict v sum
look at the current associaitve strenthg at the end of the previous trial
Find the prediction error which is the difference between lambda and our Vsum
the change in associative strength will get smaller as the learning approaches the assymptote for the conditioned stimulus
We want the prediction error to get smaller with each pairing so that associative strength reaches the value of lambda
Extinction in the Wagner Model
if we deliver no US then our lambda value will be 0 since nothing is being delivered
We would get a negative prediction error since we would have predicted a large US and nothing occurred
The alpha value did not change since the salience of the US didn’t change
We will get a decrease in associative strength for these trials
Delta V will decrease during each trial → the associative strength decreases
We would see a drop off like what expect during Pavlovs trial when we stop presented the US

when cues are presented, people are going to make a prediction of what’s going to happen
use cues to make prediction of how much US is going to happen
Wagner model gives us a set a of equations to use
Working out R-W Prediction → Minimal Math
We can make experimental predictions without making actual calculations
If we know the CSs associative strength and where it ened we can make prediction without making computations
The RW algorithm: we can either compute this or reason it
Use CSs to compute an expectation of the US (V)
add all to associative strenghts to make a prediction based on how much US they expect to observe
Observe the actual US (lambda)
wait and see what actually happens
Compute the prediction error for the trial (lambda - v)
used to update the associative strength of the trial
Use the PE to update the value for V considering the present cues during the trials
update the associative strenght, we can either compute this or reason it
if its a new CS then we can reason that the associative strength is 0
We can also predict that if one CS is paired until assymptote then we could predict that it’s associative strength is equal to lambda
this prediction of associative strength being equal to lambda isn’t always true but its our safest guess
only CSs that were presented during the trail could have their associative strength be changed during the trial
No change mechanism if it wasn’t in the trial
Every CS that was presented during a trial must change
Associative strenght oculd only change if their prediction error was not equal to 0
If the prediction error is zero then the associative strength will not change since the delta v value isn’t going to change
All associative strengths will change in the trial in the same direction
If the prediction error is positive then all the CSs during the trial will have an increase in their associaitve strenght
If the preidiction error on a trial is negative then all the CSs present will have a decrease in their associative
We could never have one cue increase and hte other cue decrease in their associative strength during the same trial
most enduring mathematical model of associative learning!
We can simulate trial by trial how the associative strength changes of the course of the experiment
OverShadowing
When multiple CSs are trained in compoud they will compete for associative strength
We split up the delta v, v old, and v new separately for each of the cues that are present during the experiment
We would compute the values on their own
The more salient CS will outcompete the other CS in their change in associative strength
When we calculate the Vsum we would add both of the associative strengths and compute our prediction error
For these values we would need the values for both of the CSs in the experiment
If we were to present the CSs on their own after they have faced competition we would see that the response for the less salient CS will have a lower response compared to the more salient CS!
