Comprehensive Study Guide for Instrumental Conditioning and Rescorla-Wagner Limitations

Recap: The Rescorla-Wagner Model

  • The Rescorla-Wagner model is a mathematical framework used to describe the change in associative strength between a conditioned stimulus (CS) and an unconditioned stimulus (US) on any given trial.

  • The fundamental equation for the model is:     ΔV=α×β×(λ−ΣV)\Delta V = \alpha \times \beta \times (\lambda - \Sigma V)

  • Definition of terms:

    • ΔV\Delta V: The change in associative strength of the stimulus that occurs on any one trial.

    • α\alpha: A learning rate parameter determined by the salience of the conditioned stimulus.

    • β\beta: A learning rate parameter determined by the properties (e.g., intensity) of the unconditioned stimulus.

    • λ\lambda: The maximum associative strength that can be supported between the CS and the US.

    • ΣV\Sigma V: The sum of the associative strengths of all stimuli present on that trial.

Limitations of the Rescorla-Wagner Model

  • While the Rescorla-Wagner model predicts phenomena such as acquisition, extinction, blocking, over-expectation, conditioned inhibition, and super-conditioning, it fails to account for several complex learning behaviors:

    • Extinction Nuances: Extinction is not simply the mirror opposite of acquisition. The model struggles to account for spontaneous recovery, renewal, and reinstatement.

    • Configural Learning: The model is an elemental model, assuming animals learn about individual stimuli independently. It does not account for animals learning about the unique combination of stimuli (CS1CS_1 and CS2CS_2) as a distinct entity.

    • Negative Patterning: This is a specific problem for the model. In negative patterning, CS1CS_1 and CS2CS_2 are reinforced individually (CS1+CS_1+ and CS2+CS_2+), but the combination of the two is not reinforced (CS1CS2−CS_1CS_2-). Rescorla-Wagner predicts the highest expectation of a US when both are present, but animals actually learn to distinguish the compound as a signal for no US.

    • Latent Inhibition: Learning in an initial stage (Stage 1: A−A-) inhibits learning to the same stimulus in a later stage (Stage 2: A+A+). Because of the mixed predictive value of stimulus A, learning is slower when it is eventually reinforced. The Rescorla-Wagner model does not predict this.

    • Learned Irrelevance: Similar to latent inhibition, where exposure to uncorrelated stimuli hinders later learning.

    • Relearning Effects: Rescorla-Wagner assumes that if an association is extinguished, the associative strength (VV) returns to 00. This implies that the re-acquisition curve should be identical to the initial acquisition curve. However, in reality, re-acquisition is typically much more rapid.

    • Evolutionary Preparedness: Animals are innately predisposed to learn certain associations helping survival (e.g., monkeys fearing snakes more easily than flowers, Mineka & Cook, 1993). While the model accounts for salience (α\alpha) and intensity (β\beta), it ignores biological significance.

    • Other Issues: Includes configural learning, compound stimuli, and various relearning effects. Note that adaptations to the model have been proposed to address some of these.

Evidence for Relearning and Rapid Re-Acquisition

  • Ricker & Bouton (1996) Study: This research demonstrated that re-acquisition follows a more rapid trajectory than original training.

    • Methodology:

      • Groups R and L: Previously trained and then extinguished to associate food with a specific CS (the CS differed between the groups).

      • Group U: Received the food US without the predictive CS.

      • Group C: Exposed only to the context without the specific learning parameters applied to other groups.

    • Because associations are re-learned faster, it suggests the original learning was not completely erased (V≠0V \neq 0), contradicting the basic Rescorla-Wagner assumption.

Summary of Pavlovian (Classical) Conditioning Concepts

  • US Devaluation: A technique used to test whether associations are stimulus-response (S-R) or stimulus-stimulus (S-S).

  • Opponent Process Model of Addiction: This model predicts withdrawal and drug tolerance. When combined with Pavlovian conditioning, it helps explain drug relapse triggers.

  • Higher Order Conditioning: Involves the association of a CS with a US via an intermediary CS, extending learning beyond direct pairing.

Introduction to Instrumental (Operant) Conditioning

  • Definition: Instrumental conditioning is the learning of associations between voluntary responses and their consequences. It is also known as operant conditioning.

  • Nature of the Response: Unlike Pavlovian conditioning where behavior is elicited (involuntary), in instrumental conditioning, the response is voluntary and emitted by the subject.

  • The ABC Model of Behavior Analysis:

    • A (Antecedents): The conditions or environmental stimuli (SS) leading to a behavior.

    • B (Behaviour): The specific response (RR) produced by the subject.

    • C (Consequences): The outcome (OO), which can be a reinforcer or a punisher.

  • Association Types:

    • S-O links: Primarily associated with Pavlovian (classical) conditioning.

    • R-O links: The core of operant/instrumental conditioning.

    • S-R links: Characteristic of ingrained habits; historically, this was thought to be the primary mechanism of operant conditioning.

The Response (RR) and The Outcome (OO)

  • The Response: The specific, observable behavior of interest. It produces and is controlled by its consequences. It must be clearly distinguished from the outcome it produces.

  • Dependent Variable (DVDV): In instrumental conditioning research, the DVDV is usually the Response Rate (e.g., lever presses per minute).

    • Response rate indicates how strongly behavior is controlled by reinforcement.

    • Higher rates typically signal stronger learning or higher motivation.

  • The Outcome: The consequence following the behavior. Its primary function is to change the likelihood of the behavior occurring again.

    • Thorndike's Law of Effect: Animals and humans repeat actions resulting in favorable consequences and avoid those resulting in aversive consequences.

Determinants of Outcomes: Positive/Negative and Reinforcement/Punishment

  • Clarification of Terms: In the context of learning theory, "positive" and "negative" do not refer to the emotional quality (valence) of the event.

    • Positive (++): Something is added to the environment.

    • Negative (−-): Something is removed or subtracted from the environment.

    • Reinforcement: Increases the future likelihood of the response. Skinner (1953) noted the only defining characteristic of a reinforcer is that it reinforces.

    • Punishment: Decreases the future likelihood of the response.

The Four Quadrants of Instrumental Conditioning

  • Positive Reinforcement: Adding a reward to increase behavior.

    • Example: A child helps a sibling; the parents provide a hug and praise. The child is more likely to help in the future.

  • Negative Reinforcement: Removing a negative consequence to increase behavior.

    • Example: Taking ibuprofen for a headache. The headache (aversive stimulus) is removed, making the subject more likely to take ibuprofen for future headaches.

  • Positive Punishment: Adding a negative consequence to decrease behavior.

    • Example: A teenager stays out late; parents assign extra chores. The chores are added, making staying out late less likely in the future.

  • Negative Punishment: Removing a positive consequence to decrease behavior.

    • Example: A child swears; parents take away phone privileges. The removal of the phone makes swearing less likely in the future.

Pitfalls and Problems in Reinforcement and Punishment

  • Accidental Positive Reinforcement (Martin & Pear, 1999): It is easy to accidentally reinforce undesirable behaviors.

    • Scenario 1: A child begins fiddling with TV dials; the mother immediate suggests a walk. The attention reinforces the TV fiddling.

    • Scenario 2: A man yells loudly because he cannot find a shirt; the wife immediately finds it. The yelling is reinforced.

    • Scenario 3: A child hits his brother; the father stops what he is doing to play with the child. The aggression is reinforced by the father's attention.

  • Problems with Punishment:

    • Subjects may learn to avoid the punisher (doing the behavior when the punisher is absent).

    • Punishment can inhibit all behavior, not just the targeted one.

    • Subjects may develop fear or dislike for the punisher (Pavlovian conditioning).

    • The subject might copy the punisher's behavior (observational learning).

    • If punishment works, the person administering it is reinforced for using potentially violent or aggressive behavior.

Shaping and Complex Behaviors

  • Shaping: A technique used when a subject does not naturally perform the desired response. It involves reinforcing increasingly close approximations to the desired behavior.

  • Example: Using shaping to teach a cat to ring a bell by reinforcing looking at the bell, then touching it, then eventually ringing it.

Factors Affecting Acquisition: Contiguity and Contingency

  • Contiguity: The temporal proximity between the behavior and the outcome. Associations are stronger when the outcome follows the behavior closely in time. If delayed, learning is slower/less robust. Note: Contiguity alone is not sufficient for learning.

  • Contingency: The predictive relationship. It measures how much more likely the outcome is when the behavior occurs compared to when it does not. High contingency is required for strong learning; low contingency results in poor learning even if contiguity is high.

Reinforcement Schedules

  • Foundational Types:

    • Continuous: Every response results in reinforcement/punishment.

    • Intermittent: Only a subset of responses are reinforced/punished.

  • Schedules Defined:

    • Fixed-Ratio (FR): Reinforcement occurs after a set number of responses (e.g., every 5th press). Results in a high response rate with a post-reinforcement pause.

    • Fixed-Interval (FI): The first response after a fixed time period is reinforced (e.g., every 30 seconds). Results in a "scalloped" response pattern (pauses, then increasing rates as time approaches).

    • Variable-Ratio (VR): Reinforcement after an average number of responses, varying unpredictably. This produces the highest, steadiest response rates and is very resistant to extinction.

    • Variable-Interval (VI): The first response after an unpredictable time interval is reinforced (average of e.g., 30 seconds). Results in a moderate, steady response rate.

  • Skinner's Discovery: B.F. Skinner found that animals respond for longer and more persistently when reinforcement is not delivered every time (intermittent reinforcement).

Intermittent Reinforcement in Social and Clinical Contexts

  • Dating Strategy: Often discussed in "pickup artistry" or pop psychology. Tactics include "hot and cold" behavior, being mysterious, or flirting then pulling away to make a partner crave unpredictable attention.

  • Abuse and Trauma Bonding: In clinical psychology and narcissistic abuse literature, intermittent reinforcement is seen as a damaging dynamic. Periods of intense affection ("love bombing") alternate with neglect, rage, or gaslighting, creating a psychological bond that is hard to break for the survivor.

Choice and the Matching Law

  • Concurrent Schedules: When two or more reinforcement schedules are available simultaneously, allowing the subject to choose.

  • Matching Law (Herrnstein, 1961): Animals match their rate of responding to the rate of reinforcement provided by each option.

    • Formula:         B1B1+B2=R1R1+R2\frac{B_1}{B_1 + B_2} = \frac{R_1}{R_1 + R_2}

    • Where B1,B2B_1, B_2 are response rates and R1,R2R_1, R_2 are reinforcement rates. If a lever provides 60%60\% of total reinforcement, the animal will spend roughly 60%60\% of its time responding to that lever.

Extinction in Operant Conditioning

  • Definition: Extinction occurs when a previously reinforced response no longer produces the outcome.

  • Extinction Burst: An initial, temporary increase in the rate of responding immediately after reinforcement stops, often accompanied by frustration or aggression.

  • Partial Reinforcement Extinction Effect (PREE): Behaviors maintained by intermittent reinforcement take significantly longer to extinguish than behaviors that were continuously reinforced.

Theoretical Perspectives on Reinforcement

  • Drive Reduction Theory (Hull): Biological needs (food, water) create internal states of tension called "drives." Behavior is reinforced specifically because it reduces these drives and restores homeostasis.

  • Intrinsic Rewards: Harry Harlow (1955) demonstrated with monkeys that puzzles can provide intrinsic rewards. Monkeys would solve puzzles even when the solution did not lead to food, water, or sex.

  • The Premack Principle: A high-probability behavior (more preferred activity) can reinforce a low-probability behavior (less preferred activity).

    • Example: "You can play video games if you do your homework first."

  • Paradoxical Reward Effects: Cases where rewards contradict standard theory.

    • Overjustification Effect: Providing extrinsic rewards for an activity that is already intrinsically motivating can actually decrease the behavior once the reward is removed.

    • Negative Contrast Effect: Moving from a high-value reward to a lower-value reward causes performance to drop below the baseline of subjects who only ever received the low-value reward.

Goal-Directed vs. Habitual Behavior

  • Goal-Directed Behavior: Driven by outcomes, involves deliberate planning, is flexible, and sensitive to outcome devaluation.

  • Habitual Behavior: Based on S-R associations, automatic, not flexible, and insensitive to outcome devaluation.

  • Outcome Devaluation Tests: Used to distinguish between the two. Forms include individual satiety (eating a specific food until full) or taste aversion (associating a food with illness).

  • Christopher Adams Experiment (cited in Haselgrove, 2016):

    • Rats with moderate training (100100 lever presses) showed goal-directed behavior—they stopped responding when the outcome was devalued.

    • Rats with extensive training (500500 lever presses) became habitual—they continued to press the lever even after the outcome was devalued.

  • Smokers Study (Hogart and Chase): Human subjects with buttons for cigarettes and chocolate showed sensitivity to devaluation, proving that R-O links are learned and remembered.