Chapter 10: Operant Conditioning
Fundamentals of Operant Conditioning
Definition of Operant Conditioning: Operant conditioning is a method of learning in which the future probability of a behavior is directly affected by its consequences. It involves learning that occurs when individuals seek rewards and actively avoid punishments.
Core Principle: Behaviors followed by favorable consequences become more likely to reoccur, while behaviors followed by unfavorable consequences become less likely to reoccur.
The Skinner Box (Operant Chamber)
Overview: Developed by B. F. Skinner, the operant chamber (or Skinner Box) is an isolated apparatus used to study animal behavior in a controlled environment.
Key Components of a Skinner Box:
Pellet Dispenser and Dispenser Tube: Delivers precise food rewards into the food cup.
Food Cup: Holds the delivered food pellet for consumption.
Lever (Response Mechanism): The subject (e.g., a rat) can press the lever to produce a outcome.
Signal Lights: Colored lights used to present visual cues or stimuli.
Speaker: Plays auditory signals or tones.
Electric Grid: Forms the floor of the chamber, connected to a shock generator to deliver an aversive stimulus.
Shock Generator: Controls the electrical current supplied to the grid floor.
Standard Procedure:
Behavior (Response): The rat presses the lever inside the box.
Consequence: A food pellet is automatically dispensed.
Effect: The future probability of the rat pressing the lever increases.

The Four Consequences of Operant Conditioning
Operant conditioning categorizes consequences based on two main factors: whether a stimulus is added or removed (Positive vs. Negative) and whether the behavior increases or decreases in frequency (Reinforcement vs. Punishment).

Terminology Distinction:
Reinforcement: Any consequence that increases the future frequency or probability of a behavior.
Punishment: Any consequence that decreases the future frequency or probability of a behavior.
Positive: Adding something to the environment.
Negative: Removing or subtracting something from the environment.
Positive Reinforcement
Definition: Presenting or adding a pleasant stimulus to the environment immediately following a response, which increases the future probability of that behavior.
Mechanism:
Behavior (Response): Press lever
Consequence: Rewarding stimulus presented (Food delivered)
Effect: Tendency to press lever increases
Examples:
Cat Example: A cat learns to use a new cat door, and the owner gives him a kitty treat.
Child Example: A child receives praise or treats for engaging in targeted polite behavior.
Negative Reinforcement
Definition: Removing or reducing an unpleasant (aversive) stimulus from the environment immediately following a response, which increases the future probability of that behavior.
Mechanism:
Behavior (Response): Press lever
Consequence: Aversive stimulus removed (Electrical shock turned off)
Effect: Tendency to press lever increases
Examples:
Cat Example: A cat who dislikes getting wet uses his new cat door to escape incoming rain outside.
Vehicle Example: Turning off a noisy windshield wiper switch once the rain stops.
Positive Punishment
Definition: Adding or presenting an unpleasant (aversive) stimulus to the environment immediately following a response, which decreases the future probability of that behavior.
Examples:
Talking back to a supervisor Receiving a severe verbal reprimand.
Swatting at a wasp Getting stung by the wasp.
A cat meowing constantly Being sprayed with water from a squirt bottle.
A cat scratching furniture Squirting the cat with a squirt bottle.
Negative Punishment
Definition: Removing or subtracting a pleasant or desired stimulus from the environment immediately following a response, which decreases the future probability of that behavior.
Examples:
Staying out past curfew Losing driving and car privileges.
Arguing with the boss Loss of employment/job.
Playing with food at the dinner table Loss of dessert.
Teasing a sibling Being sent to a quiet room (loss of social contact and interaction).
Cat misbehaving Moving the cat to an isolated room away from the canary and goldfish.
Behavioral Modification Techniques & Real-World Applications
Shaping
Definition: A technique in operant conditioning where successive approximations of a desired target behavior are systematically reinforced until the final target behavior is achieved.
Step-by-Step Example (Child Learning to Eat Vegetables):
Initial Baseline: The child completely refuses to eat vegetables at mealtimes.
First Approximation: The child is reinforced with praise ("YES!") simply for touching the fork.
Second Approximation: The child is now reinforced ("GOOD JOB!") only when touching the actual vegetables with the fork.
Final Target Behavior: After shaping the behavior progressively through reinforcement, the child eats the vegetables.

Practical Applications in Media and Relationships
Behavioral Modifications in Popular Culture (e.g., The Big Bang Theory):
Using chocolates as a positive reinforcer to reward specific desirable behaviors in a partner or housemate.
Using squirt bottles or proposed mild electrical shocks as punishers to suppress undesirable behaviors.
Effectiveness Comparison: Positive reinforcers (such as chocolate) generate higher compliance and co-operation compared to aggressive severe punishment (such as electrical shocks), which can provoke avoidance, anxiety, or aggression.
Inadvertent Operant Conditioning (Night Crying Example):
Panel 1: A toddler sleeps in her own crib, thinking "THIS IS GREAT!".
Panel 2: After crying during the night, the toddler gets moved into the comfortable adult bed with both parents and thinks, "I'LL HAVE TO WAKE UP CRYING IN THE MIDDLE OF THE NIGHT MORE OFTEN".
Analysis: Crying is unintentionally positively reinforced by parental attention and co-sleeping privileges, increasing night crying frequency.
Schedules of Reinforcement
Reinforcement schedules dictate the timing and pattern of delivering reinforcers following an operant behavior.
Continuous Reinforcement vs. Partial Reinforcement
Continuous Reinforcement (CRF):
Pattern: Every single correct response is immediately followed by a reinforcer (\text{Response} \n\rightarrow \text{Reinforcer}, , ).
Primary Use: Most effective during the acquisition phase when learning a brand-new behavior.
Disadvantages and Practical Limits:
Feasibility Issues: It is rarely feasible or practical to praise or reward a behavior (e.g., a child being polite) 100% of the time.
Satiation: The subject quickly becomes full or accustomed to the reward.
Rapid Extinction: Once the continuous reinforcement stops, the behavior ceases rapidly.
Partial (Intermittent) Reinforcement:
Pattern: Only some responses are followed by a reinforcer, while other responses go unreinforced (, , , ).
Primary Advantage: Produces greater resistance to extinction than continuous reinforcement schedules.
Four Types of Partial-Reinforcement Schedules
1. Fixed-Ratio Schedule (FR)
Definition: Reinforcement is delivered after a fixed, predictable number of non-reinforced responses have occurred.
Key Characteristics: High rates of response, often accompanied by a short pause immediately after reinforcement.
Examples:
Coffee Shop Loyalty Card: A "BUY 7 GET 1 FREE" stamp card where 7 distinct purchases are required to earn the 8th cup for free.
Academic Performance Rewards: Earning a bill or receiving an 'A' grade after completing a fixed set of written papers or assignments.

2. Variable-Ratio Schedule (VR)
Definition: Reinforcement is delivered after an unpredictable, variable number of non-reinforced responses.
Key Characteristics: Produces exceptionally high and steady response rates with high resistance to extinction.
Examples:
Slot Machines / Gambling: Playing a "HAYWIRE" slot machine where a win may occur after 1 pull, 15 pulls, or 3 pulls, making the exact payout schedule unpredictable.

3. Fixed-Interval Schedule (FI)
Definition: Reinforcement is delivered for the first response that occurs after a fixed, predetermined length of time has elapsed.
Key Characteristics: Response rate drops immediately following reinforcement and accelerates sharply as the time for the next expected reinforcer approaches (scalloping effect).
Examples:
Annual Retail Sales Events: Black Friday sales scheduled predictably on November 27th each year.
Scheduled Work Performance Reviews / Meetings: Staff participating enthusiastically right before or during fixed annual or monthly evaluation meetings.

4. Variable-Interval Schedule (VI)
Definition: Reinforcement is delivered for the first behavior performed after an unpredictable, varying time interval has elapsed.
Key Characteristics: Results in a steady, moderate rate of response with little to no post-reinforcement pausing.
Examples:
Checking Cell Phone Messages: Periodically checking a phone display for new incoming messages ("One new message"), where text messages arrive at completely unpredictable intervals.
Random Workplace Drug Testing: Human Resources conducting unscheduled, random drug tests at varying time intervals throughout the year.
