Week 6: Mechanism of Learning -> Cue Competition


Compound Stimuli and Cue Competition

  • Compound Stimuli: Formed by presenting multiple cues at once

    • Cues presented in compound will often have identical levels of correlation and contiguity with the US

      • May not be learned about equally!


  • Cue Competition: When a compound CS is paired with a US, compound elements are competing with each other to be associated with the US

    • Elements split the association between themselves

    • One element might absorb ALL of the association

    • There might be equal association split

    • There might be unequal association split


  • Factors that determine how associations will split between compound elements

    • Overshadowing and the intensity of a CS

    • Informational Value and Relative Validity

    • Prior learning and blocking


  • Usually in learning irl there are rare times where only one CS is presented at a time, mostly there are a lot of cues happening at a time → important to learn this idea since its applicable to irl scenarios


  • Experimental Approach: Training an association as compound and then testing the elements individually and compare them


  • Test: Each stimulus is presented on its own to see the level of CR

  • Same contiguity and contingency, different learning across elements and compound groups

  • In the compound groups the associations are split between the elements(each would get 50% of the association for this example) when we present them at the same time

  • The element group shows that when the element is on its own it doesn’t have to compete to form the association

  • Two cues that are in compound predict the US, on their own as element they will predict a certain portion of the CR

  • Can be the case where one CS will not elicit a response and the other will take the entire response



Sensory Intensity and Overshadowing

  • What if one is more obvious to the animal when elements are in compound?

    

  • Experimental Example:

    • CS A: 40 db tone (quiet)

    • CS B: 80 db tone (loud)


  • Conditioning:

    • Group 1: A+ (CS A alone)

      • Will not have to compete

    • Group 2: AB+ (compound group)

      • Competition


  • One will strongly outcompete the other with different intensities

    • B+ alone elicits a STRONG response and A+ alone will have a weak response

      • Maybe the tone is so quiet that a response isn’t strongly gain but shows that the animal is still capable of learning the association

  • We know that there is competition going on since when A is alone then the CR is very strong

  • We would need to have different groups to understand the type of competition that is happening in the compound


  • Degree of overshadowing is easiest to predict when the stimuli have the same modalities (presentations) but have different intensities

    • The loud buzzer OVERSHADOWED the quieter tone conditioning

  • Overshadowing comes from relative stimulus intensity


  • Overshadowing: Conditioned Responding evoked by a CS is lower when the CS is paired as a compound CS than when paired as a element

    • A cause of overshadowing can be the different in the salience of the elemental CSs

  • This happened regardless of the contiguity or the contigency of both of the CSs


  • The degree of overshadowing is the easiest to predict when the stimuli have the same modality and are of different intensities

    • ex. the loud tone will overshadow, the brighter image will overshadow

  • We wouldn’t be able to readily guess which would overshadow if the modalities aren’t the same!


  • Important to have the quieter tone show examples of learning since it would show evidence that learning can still be made when the less contributing CS can still support normal pavlovian conditioning when it doesn’t have to compete


“Informational value” and relative validity

  • How much information does a CS have about a US to determine how the relationship will work


  • What happens when one element of a compound is reliable predictor of the US and another element of the compound is less reliable predictor of the US?


  • Two groups of animals (uncorrelated and correlated groups)

    • Correlated Group: the excitatory us is a 100% reliable predictor of the US

      • Only one US will predict the us 100%, one will predict 0%, other will predict 50%

    • Uncorrelated Group: Truly random design where no CS is a reliable predictor of the US

      • all CS’s will predict the US the same amount of times

      • no element will predict with 100% that the US will not be delivered


  • Correlated Group: A and B provide a lot of information regarding the US. C doesn’t

  • Uncorrelated Group: No CS is no better or no worse at predicting the US than another CS.


  • looking deeper at the information

    • the Conditioned response to X is different across the correlated and uncorrelated groups

    • When a CS is relatively BETTER predictors are the ones that are going to be associated with the US and creates the CR

    • Despite X have the same prediction for the correlated and uncorrelated groups, it will have a higher relative validity in the uncorrelated

  • Relative Validity refer to how good a CS is at predicting the US in comparison to other CSs for the scenario

  • For the uncorrelated version, since all of the CSs have the same amount of prediction they will all elicit the same response


  • When a CS has a higher relative validity it will OUTCOMPETE other CSs with lower validities



Overshadowing and “informational value”


  • Relative validity: If cues in a compound differ in how well they predict the US, the best predictor will outcompete the less-relaible predictors to associate with US

    • the strength of one CS’s association with the US depends on how good or bad OTHER cues are at predicting the US



Overshadowing and “informational value” in the real world

  • Symptoms in this example would be considered the CS

    • fever is a symptom of many illness (not very informative or a valid predictor)

    • if we see a certain rash it is highly predictive of only particular cause

  • Symptoms will be generic → provides little information

  • Specific symptoms → provide a lot of information about what illness it is (highly-predictive)

  • Experiment relies on the relative validity effect- poor predictos of a disease only INFLUENCE the diagnositc decisions when no better predictor of the disease exsited


  • Relative Validity: helps figure out which CSs to pay attention to based on how well the CS will predict the US



Blocking and prior training history

  • How does prior learning affect learning when the CS is added to a compound?


  • Blocking Group: learns simple Pavlovian Association (tone will produce a strong conditioned response

  • Both groups experiences the compound association

    • Blocking group has already learned something about tone

  • Results:

    • Control: Each element will elicit a conditioned response (both are learned about)

    • Blocking: Only the tone (the prior trained CS) will elicit learning and block the learning that could be done by the other CS


  • Prior learning of a CS will determine the learning that could be done by another CS when the CS is added in a compound


  • Blocking: one cues” prior training hisotry prevents a novel cue from being learnign about when the two are trained as compound

    • Will interfere with the learning during compound training



Kamin’s blocking experiment

  • Results show that little/no learning is done for the light (novel element) in the blocking group when added in a compound

  • Overshadowing group will show how the competition of the two CS’s will behave if there was no prior training of any of the CSs

  • we know that the CS is still capable of learning but the learning changing in the blocking group or when it’s in the overshadow group



Blocking: what’s going on here?

  • Previous learning to one cue can prevent other cues from being learned about later (get it BLOCKING)

  • Redundant cues aren’t learning about

  • Once a CS is predicted, there is no need to learn about more things that will predict

  • Learning system is using the most simple way to learn an association



Magnitude UNBlocking

  • When the addition of a new cue predict an increase in the intensity of the US, the novel cue in the compound is LEARNED → UNBLOCKED


  • The light in phase 2, group 2, will be learned about and unblocking since it is a novel stimulus that will be used to predict an increase in the intensity of the US

  • Unblocking happens when a novel stimulus is able to predict an increase of the magititude of the US


  • What happens when we decrease the magnititue of the US?…

  • Novel learning of the CS that will predict the decrease will undergo inhibitory learning

    • ex. the light cue will predict that the US that would’ve normally been learned about is NOT going to be learned

  • The prediction made is going to be for LESS US → will be the same as compound training for inhibitory learning

  • WE ARE CHANGING THE MAGNTITUDE



Identity unblocking

  • When the addition of a new cue predict a qualitative change in the US, the novel cue in the compound will be learned about → UNBLOCKED


  • Similar to magnitute unblocking but the cause is different

  • Results: in the unblocking group (group two) there will be a higher conditioned response

  • WE ARE CHANGING THE QUALITATIVE NATURE/IDENTITY OF THE CS BUT NOT THE MAGNITUDE



Trans-reinforcer blocking

  • Trans-reinforcer blocking: Blocking that persists across a qualitative change in the US

  • When we make a qualitive change to the US, the group that doesn’t undergo learning will have a higher conditioned response to the CS than the group that underwent prior learning

  • Interpretation: Showed that the learning system thinks that the change isn’t worth learning about so it didn’t learn a different association


  • Some association are more about the motivational content of the US than its specific identity

  • Learning is not big enough for learning to be done


Interpreting blocking/unblocking results PRACTICE THISSS!!

  • Blocking design will follow the same general scheme

    • Blocking: No learning is done for a novel CS

    • Unblocking: Learning is done a for a novel CS



Summary: Blocking

  • Modern framework for studying associative learning has its genesis in three major experiments

    • Resocrls correlation work _> showed that contiguit is not enough for learnign

      • Associations will not work if its not reliable

    • Wagners” relative validity → showed that cues will compete with each other based on information value

      • Associations depend on how good or how bad a predictor is which will overtake the associative relationship

    • kamins blcoking → redundant cues will not be learned about

      • Learn minimal amount of information they need to build an association


  • Effect are foundation for modern theories of associative learning



Cue Competition Summary

  • Multiple CSs paired in compound with an outcome will not create the same equal conditioned response


  • More Salient cues will outcompete less salients ones (overshadwoing by relative stimulus intensity)

  • Better predictors of the US outcompete poor predictions (informationtional value/relative validity)

  • Redundant cues are not usually learned about (Kamin’s Blocking)


  • Theoretical Implications

    • Simple Pavlovian learning isn’t passive

      • External factors: correlation/contiguity, relative validity, etc.

        Internal factors: previous learning, sensory perception, etc.

        In real-life, most learning involves compound stimuli, so cue competition effects show

        up everywhere.



Understanding the computational basis of learning

  • What is a model of learning: A set of instructions for learning assocation (rules for how learning works)

  • Why? if we understand associative learning, we should be able to build an artificial model for it

    • Allows us to be specific of the model and think critically about each process of the step

    • Think throughouly

    • Allows us to make new predictions and check if the predictions are true

    • Allows us to identify a knowledge gap



The RW Model

  • Lamda:

  • Vcs associative strenght between a particulat x and a US in an experiement

    • How strongly an indivudal CS will [redict a US

    • Tells the indiviual how much US they should expect from a specific CS

    • Every CS will have its own associative strength

    • Changes over time as learning occur since the associaitve strength of a CS will change

    • Reflects the understanding of what the CS means

    • Tells us how much the associative strength of a particular Cs should changed given the experimental trial

  • second equation is used to get a new value for the associative strength

    • might need to change the associative strength by the end of the trial

    • adding the changing to the associative strength


  • Lambda represents the amount of US delivered on a trial ( the size or magnitude of a US)

    • Same units as the unconditioned stimulus

  • Lambda sets the assymptote of learning

    • point the learning curve will actually get


  • We expects the associative strengths to reach the value of lambda at the peak of the learning curve

    • We expect this to be the peak of learning


  • Vsum:the total amount of US predicted by all of the CSs present on a current trial

    • adding up the associative strength of all the cues in a trial to get the total associative strengths

    • we would add up indidivdual associative strengths to get the Vsum

    • gives us the total unconditioned stimulus expectation

    • all we care about the stimuli that are presented (the cues) regardless if the US is delivered or not! j


  • Lambda - Vsum (prediction error)

    • uncondiionted stimulus (what actually happened) - total associative strength (expectation)

    • quantitfies how good or how bad a subjects prediction was for that trial

    • when error is 0 = no error in prediction

    • when error is less than error = subjects were expecting more than what was delivered

    • when error is greater than error = subjects were expecting less that what was delivered


  • Prediction error can be thought of how surpising the US was to the subject during that specific tria;

    • If the outcome was as expected, nothing is surprsing

    • if the outcome was surprising then the prediction should be adjusted

      • We should adjust the associative strengths to make the predictions more accurated


  • Alpha cs

    • How intense the CS actually is

    • ex. loud and quiet tones being used

      • A more obvious CS is going to have a higher alpha value

    • value ranges from 0 to 1

    • Alpha controls the learning rate! → it will scale the prediction error


  • The alpha value will usually be given

    • We can predict that the alpha from a loud tone will be a higher value than the alpha of a quiet tone

    • We could also experimentally determine this alpha value (reasoning generally based on the salience of the CSs)


  • The amount of learning per trial for each CS (delta CS) is set by two things:

    • The prediction error

      • how good or how bad a prediction of a US was during that trial

      • will affect every CS that was present in the same way

    • The salience of each CS that was present on the trial

      • CSs might differ in their alpha value



Acquistion: Wagner Model

  • Plotting the change in the associative strength

    • strength will start at zero since there is no association being made

    • increases in the level of prediction through each trial

  • By the end of learning, the associative strength will be close to the value of lambda

  • Each point on the graph is (lambda - V) (prediction error)

  • Looks like a typical learning curve during Palovian Learning


  • Conditioned responding on every trial is propportion to the associative strengh of the cue on the trial


  • We assume that behavior is driven by the associative strength of the cue




Acquisition Example

  • Experiment: tone → food

    • Lambda value: 1

    • a tone = 0.35 (based on the sensory salience of the CS)


  • V old: associative strength before the trial

  • V new: associative strength after the trial

  • Assume that a novel CS will have an associative strength of zero before any learning has occurred

    • can round two decimal place and doesn’t really change the results


  • Chart how the associative strength of the CS changes through the different trials!


  • We will use the equations to compute any missing values

steps:

  • Consider what is happening during the trial

  • predict v sum

    • look at the current associaitve strenthg at the end of the previous trial

  • Find the prediction error which is the difference between lambda and our Vsum


  • the change in associative strength will get smaller as the learning approaches the assymptote for the conditioned stimulus

  • We want the prediction error to get smaller with each pairing so that associative strength reaches the value of lambda




Extinction in the Wagner Model

  • if we deliver no US then our lambda value will be 0 since nothing is being delivered

  • We would get a negative prediction error since we would have predicted a large US and nothing occurred

  • The alpha value did not change since the salience of the US didn’t change

  • We will get a decrease in associative strength for these trials

  • Delta V will decrease during each trial → the associative strength decreases

  • We would see a drop off like what expect during Pavlovs trial when we stop presented the US




when cues are presented, people are going to make a prediction of what’s going to happen

use cues to make prediction of how much US is going to happen

Wagner model gives us a set a of equations to use



Working out R-W Prediction → Minimal Math

  • We can make experimental predictions without making actual calculations

  • If we know the CSs associative strength and where it ened we can make prediction without making computations


The RW algorithm: we can either compute this or reason it

  • Use CSs to compute an expectation of the US (V)

    • add all to associative strenghts to make a prediction based on how much US they expect to observe

  • Observe the actual US (lambda)

    • wait and see what actually happens

  • Compute the prediction error for the trial (lambda - v)

    • used to update the associative strength of the trial

  • Use the PE to update the value for V considering the present cues during the trials

    • update the associative strenght, we can either compute this or reason it


  • if its a new CS then we can reason that the associative strength is 0

  • We can also predict that if one CS is paired until assymptote then we could predict that it’s associative strength is equal to lambda

    • this prediction of associative strength being equal to lambda isn’t always true but its our safest guess

  • only CSs that were presented during the trail could have their associative strength be changed during the trial

    • No change mechanism if it wasn’t in the trial

  • Every CS that was presented during a trial must change


  • Associative strenght oculd only change if their prediction error was not equal to 0

  • If the prediction error is zero then the associative strength will not change since the delta v value isn’t going to change


All associative strengths will change in the trial in the same direction

  • If the prediction error is positive then all the CSs during the trial will have an increase in their associaitve strenght

  • If the preidiction error on a trial is negative then all the CSs present will have a decrease in their associative

    • We could never have one cue increase and hte other cue decrease in their associative strength during the same trial




  • most enduring mathematical model of associative learning!

  • We can simulate trial by trial how the associative strength changes of the course of the experiment




OverShadowing

  • When multiple CSs are trained in compoud they will compete for associative strength

  • We split up the delta v, v old, and v new separately for each of the cues that are present during the experiment

  • We would compute the values on their own

  • The more salient CS will outcompete the other CS in their change in associative strength

  • When we calculate the Vsum we would add both of the associative strengths and compute our prediction error

    • For these values we would need the values for both of the CSs in the experiment

  • If we were to present the CSs on their own after they have faced competition we would see that the response for the less salient CS will have a lower response compared to the more salient CS!