1/95
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
The typical instrumental conditioning experiment…
requires organisms to discover that a response in a stimulus situation produces a reinforcement.
What is the main difference between classical and instrumental conditioning?
The reinforcer in classical conditioning is contingent on solely the stimulus, while the reinforcer in instrumental conditioning is contingent on the stimulus and the response.
What are the similarities with instrumental and classical conditioning?
Same effects of practice, same extinguishing when contingency removed, same spontaneous recovery etc.
What did Tinklepaugh discover?
The reinforcer can be a part of a learned association. Monkeys were given banana as a reinforcer. Later, they were given lettuce, and the monkeys showed disappointment. aka They formed an expectation for the reinforcer.
What did Corwell and Rescorla discover?
They learned that organisms that expect specific reinforcers to specific responses. With the left/right rod experiment, the stimulus and response had to be associated with the reinforcement since they stopped pushing the rod for the bad tasting food. The stimulus had to be associated with the reinforcer. (rod and food type).
In instrumental conditioning, organisms are learning a ___ term contigency
3 term": that a response in a particular stimulus will be followed by a reinforcement.
What did Claire-Smith and MacLaren discover?
Secondary reinforcement. All rats learned that lever produced noise. G1 learned that noise produced food. G2 learned light produced noise. G1 pressed the lever more since noise might produce food. Noise is secondary/conditioned reinforcer, food is primary reinforcer.
Organisms can learn that certain responses produce neutral…
outcomes and combine this information with other experiences to get reinforcement. like maze running and noise-lever experiment.
What was Saltzman’s experiment/discovery?
T maze training with white and black goal box. First, Saltzman gave rats food in the white goal box. Then, put them in a white + black box, and rats went to the white box even though there was no food. (White box is reinforcing the behavior)
A classic example of secondary reinforcement is ?
money for humans
A secondary reinforcer is a previously neutral stimulus that has acquired the ability to...?
reinforce behavior as a consequence of being paired with a primary reinforcer.
The extension of a conditioned response to new stimuli is called…?
generalization
What did Guttman and Kalish discover?
wavelengths with bird pecking. response rate decreases as the target wavelength differed from the testing wavelength. Created a generalization gradient.

A generalization gradient is
Y = number of responses. X = whatever is tested in a range.
What did Jenkins and Harrison discover? with Generalization
Pigeons pecking at a lit key at a 1000 HZ tone. At varied tones, the response didn’t really change. It only changed significantly with light. They either didn’t learn the association between the food and tone or were ignoring the tone. So, they changed the contingency by only giving food when light and tone together, not with light alone. This gave a sharp generalization gradient to the tone.
What did Jenkins and Harrison discover? with Discrimination
In a follow up experiment, the key was lit. The 1000 Hz tone was sometimes presented. Pigeons were reinforced when NO tone was presented. Pigeons were then discrimination trained to respond to a 1000 Hz tone. After training, response gradient was much steeper and a peak shift moved from 1000 Hz to 1050 Hz.
Organisms spontaneously generalize the CS,…
ignoring dimensions and certain differences in other dimensions.
What is Spence’s Theory of Discrimination Learning?
He developed a theory about how positive and negative stimuli combined produce a net generalization gradient.
How do you build a positive generalization gradient?
Reinforcement of a response in presence of a stimulus
How do you build a negative generalization gradient?
Lack of reinforcement of a response in presence of a stimulus
Discrimination learning is a simple combo of
positive and negative generalization gradients
What is a peak shift?
The stimulus that gets the most response, isn’t the positive stimulus but one shifted away from it and the negative stimulus.
Attentional learning involves
learning multiple, successive discriminations
What are the steps to attentional learning?
One dimension is reinforced like the color red (circle). Then, transfer training, so the color yellow is reinforced. This is called a reversal, since the opposite color is being reinforced. Or it could be a triangle, which is called a non-reversal transfer, since it’s a different dimension.
What is the blocking phenomena?
One dimension becomes so strongly associated that it blocks out other dimensions.
What is a non-reversal transfer?
Changing the dimension for attentional learning.
What type of transfer is easier to learn for humans and primates and why?
Humans/primates find reversal transfers easier, due to the prefrontal cortex. Young kids and non-primates are better with non-reversal transfers.
What is a intradimensional shift?
Discriminate between two values like blue and green within a relevant dimension. (easier)
What is a extradimensional shift?
Learning/discriminating between values on different dimensions.
Organisms tend to not discriminate in
responses with the same effect in the environment.
Attempts to shape behavior in organisms may be frustrated by
species-specific response patterns. (instinctive drift)
What did Brown and Jenkins discover?
Autoshaping, A stimulus starts evoking a species-specific behavior because it has been associated with a reinforcer. like birds pecking at key w open/closed beak. different stimuli select different aspects of species behavior.
Behavior systems analysis
approach that emphasizes the natural, unlearned organization of behavior for a species.
Hammond
Contiguity vs Contingency. Rats respond when reinforcement is contingent on their response. When reinforcement becomes equally likely with or without the response, responding decreases, even though response and reinforcement can still occur close together in time.
Organisms display conditioning when
there’s a contingency between response and reinforcement.
Response Shaping
selectively reinforcing ever close approximations to the target
response. A means for training responses that are initially unlikely to
occur. Skinner believed that all behavior was the result of reinforcement schedules. E.g., Ping-Pong Pigeons
Timberlake & Grant (1975) showed that…
rats responded differently depending on the nature of the stimulus:
If the CS was a block of wood, rats tended to gnaw it.
If the CS was another rat, they tended to socialize with it.
So even though both stimuli can become associated with a reinforcer, the rat’s species-typical response to that object influences what the conditioned response looks like.
What is superstitious learning?
Skinner. Animals will spontaneously produce behaviors even when there is no actual contingency between that behavior and reinforcement. (accidental contingencies like a bird was hopping when it got food so now it hops for food)
Staddon & Simmelhag
builds off superstitious behaviors.
Two classes of behavior:
1.) interim behaviors: appear to be superstitious, but serve other
purposes. (filler behaviors while waiting)
2.) terminal behaviors: are related to obtaining the reinforcement (closer to reinforcement, more related, like pecking near it)
What is partial reinforcement?
The target response is reinforced only sometimes. Conditioning takes longer, but resistance to extinction is stronger. and their resistance to extinction increases as the reinforcement rate is lowered. (paradox)
What is Learned Helplessness?
Organisms that have received repeated unavoidable aversive stimuli come to ignore the relationship between their behavior and environmental outcomes.
Associative Bias
organisms are biologically predisposed to learn some response–reinforcer associations more easily than others.
Species-specific defense reactions
related to associative bias. when the outcome is avoiding danger, its easier for rats to flee vs pressing the lever to avoid danger.
The hippocampus is surrounded by?
temporal cortex
Hippocampus is important for
cognitive map and spatial learning.
Morris Water Maze
rats with hippocampus lesions did poorly in the “place” condition. n the place condition, the platform is hidden, so the rat cannot simply see it. It has to use the surrounding environment to form a cognitive map of where the platform, but the rats with lesions couldn’t form a map.
Learning provides knowledge of the reinforcement contingencies of actions, and organisms…
generally select the most beneficial action given their knowledge.
Rational theory states that…
organism will select the behavior with the highest “expected value”
What is expected value?
Value of a future gain should be directly proportional to the chance of getting it.
What is value?
A reward or a benefit, or a cost
How is expected value calculated?
Multiply the probability of each outcome by its value and then take the sums. EV = f(value, P(value))
Oi
Probability of different outcomes

Vi
Value of different outcomes

What did Loftus research?
The number of points assigned to a picture was positively related to the probability of recognizing it later. Subjects also fixated more on pictures worth more points.
The value of a picture is clearly…
related to eye fixations. unrelated to correct recognition. so, rewards direct attention, not learning.
What is positive reinforcement? (reward training)
A desirable stimulus is made contingent on a response.
What is negative reinforcement? (escape or avoidance)
An aversive stimulus is omitted or prevented if the person makes a specific response, which increases the rate of that behavio
What is omission training?
A desirable stimulus is made contingent on the omission of a behavior/response. A positive or desired consequence is taken away (or omitted) if a specific behavior happens, which decreases the future rate of that behavior
What is punishment?
An aversive stimulus is made contingent on a behavior or a response.
What did Camp, Raymond, and Church do?
Studied effects of delay in punishment. Found that punishment is most effective when given asap.
What did Church discover?
Increased magnitude of punishment caused greater suppression of lever pressing.
What did N.E. Miller discover?
Low to high punishment was less effective than just high punishment.
Punishment is most effective when…
administered immediately, is extreme, and is consistent. and an alternative behavior is available.
What is the effect of noncontingent punishment on punishment?
Rats that were shocked regardless of behavior developed learned helplessness and continue pressing the lever.
If a punishment is to be effective, it must be contingent…
solely on the response it is supposed to suppress.
What is Azrin and Holtz discover?
Found that punishment is more effective if there is an alternative behavior available. Pigeons suppressed previous pecking on shock key when given another key to peck at.
What is the assumption of the drive reduction theory?
Positive reinforcers are good for the organism and negative reinforcers are bad from a evolutionary point of view.
What are drives?
States of deprivation energize or motivate certain behaviors like food, sex, pain
Behaviors that reduce drives are…
reinforcing.
What is the major two problems of the drive reduction theory?
Organisms can be reinforced by events that have no evolutionary benefits. Behaviors can be reinforced by things that do not reduce drives or even increase them.
What is the assumption of Premack’s Theory?
All behaviors can be assigned a value, and stimuli are neutral. so, the responses are rewarding, not the stimuli.
If kissing your sweety is contingent on going to your sweety’s parents’ house, then going to your sweety’s parent’s house will be reinforced.
What is the prediction of Premack’s theory?
More valued behaviors should be performed more often than less valued behaviors. More-preferred behavior contingent on less-preferred behavior → increases the less-preferred behavior.
What is a bliss point?
Desired baseline rate for events. Allison and Timberlake.
What is the equilibrium theory?
Allison and Timberlake. Change behaviors to move close towards the bliss point. Anything toward is rewarding and anything away is punishing.
When the hypothalamus is removed…
many drives are extinguished
Electrical stimulation to the hypothalamus can…
produce drive related behaviors like eating and drinking.
Olds and Milner discovered that
electrical stimulation to the hypothalamus can be a reinforcer. Rats pressed lever for the stimulation. Feeling of drunk and sexually aroused for patients.
What did Stein discover?
Organisms liked pharmalogical stimulation to the hypothalamus too through drugs. He thought drugs and cocaine would stimulate these areas. evidence linking neurotransmitter systems in hypothalamic reward areas to the reinforcing effects of drugs like opiates.
Since electrical stimulation and drugs serve no obvious biological function, reduce any natural drives, or involve behaviors.
this evidence is problematic for premack theory and drive reduction theory.
What is the process of neural basis for reinforcement in the hypothalamus?
The hippocampus projects to the entorhinal cortex.
The entorhinal cortex project to the amygdala.
The amygdala project to the hypothalamus.
What is fixed-ratio?
reinforcement is given after a fixed number of responses
What is variable ratio?
reinforcement is given after a variable number of responses
What is fixed interval?
reinforcement is given after a fixed amount of time
What is variable interval?
reinforcement is given after a variable amount of time
Response rates tend to be greater for ___ than ___ schedules.
Ratio over interval.
Variable schedules produce ____ rates of responding,
whereas fixed schedules produce ____ rates of
responding
fixed, variable.
What is the Matching Law?
The proportion of different behaviors equals the reinforcement proportion.

Faced with two variable-interval schedules,
an organism divides its responses between them in proportion to their two rates of reinforcement (matching law)
What is momentary maximizing?
a method to adapting to different reinforcement schedules.
momentary maximizing in a variable interval schedule
behavior reaches a stable state when the probability of the reinforcer equal the probability of the behavior.
momentary maximizing in a variable-ratio schedule of reinforcement.
the optimal behavior is to always choose the behavior that provides the greatest rate of reward.
What is probability matching?
For example, suppose:
Option A wins 80% of the time
Option B wins 20% of the time
A person who probability matches will tend to choose A about 80% of the time and B about 20% of the time. That matches the reward probabilities.
If you always choose A:
80% correct
If you probability match:
(.80 × .80) + (.20 × .20) = .68
So you'd only be correct 68% of the time. It gives up some potential rewards.
Discounting the Future:
The potential for gain or loss in the future is not as important as
the same game or loss in the present.
Subjective value
is a personal, idiosyncratic representation of the worth of some object, person, or event. One’s subjective value influences decisions.