3.8

Foundations of Operant Conditioning

  • Operant conditioning vs. Classical conditioning: Both operant and classical conditioning are fundamental forms of associative learning, but they differ in mechanism and response type.
    • Classical conditioning forms associations between distinct stimuli (a conditioned stimulus [CS] and the unconditioned stimulus [UCS] it signals) and governs respondent behavior—involuntary, automatic responses to a stimulus (e.g., salivating first to meat powder, and subsequently to a tone).
    • Operant conditioning involves organisms associating their own voluntary actions with resulting consequences.
  • Operant Behavior: Behavior that operates on the environment to produce rewarding or punishing stimuli is defined as operant behavior.
    • Actions followed by reinforcers increase in frequency.
    • Actions followed by punishers decrease in frequency.

Skinner’s Experimental Framework and Shaping

  • B. F. Skinner (1904–19901904\text{--}1990):
    • Initially an undergraduate English major and aspiring writer, Skinner became modern behaviorism's most influential and controversial figure.
    • Skinner elaborated on the work of psychologist Edward L. Thorndike (1874–19491874\text{--}1949).
  • Thorndike’s Law of Effect:
    • Thorndike's law of effect states that rewarded behavior tends to recur (Thorndike, 18981898).
    • Experimental setup: Thorndike placed cats inside a puzzle box, using a piece of fish outside the box as an enticement to escape.
    • Results: The cats' speed and efficiency in executing the maneuvers required to escape improved over successive trials, demonstrating the law of effect.
  • Skinner’s Early Pigeon Research:
    • In 19431943, operating from a rooftop office in a Minneapolis flour mill, Skinner and his students humorously questioned whether they could teach windowsill pigeons how to bowl (Goddard, 20182018; Skinner, 19601960).
    • By shaping natural walking and pecking behaviors, they succeeded (Peterson, 20042004).
    • Skinner subsequently taught pigeons various non-natural behaviors, including walking in a figure 88, playing table tennis, and pecking at a screen target to keep a guided missile on course.
  • The Operant Chamber (Skinner Box):
    • To study operant principles systematically, Skinner designed an operant chamber (popularly called a Skinner box).
    • Features: Contains a bar (lever) that an animal presses—or a key (disk) that a bird pecks—to release a reward of food or water, alongside internal or external recording devices that tally response rates.
  • Reinforcement Definition:
    • Reinforcement is defined as any event that strengthens (increases the frequency of) a preceding response.
    • What serves as a effective reinforcer depends on the specific animal and its physiological conditions:
    • Humans: Praise, attention, or a paycheck.
    • Hungry/Thirsty Rats: Food or water.
  • Shaping Behavior:
    • Shaping is an operant procedure in which reinforcers gradually guide actions closer and closer toward a desired target behavior.
    • Method of Successive Approximations: Researchers reward responses that are progressively closer to the final target behavior while ignoring other responses.
    • Example (Conditioning a rat to press a bar):
    1. Observe existing natural behavior.
    2. Give food whenever the rat approaches the bar.
    3. Require the rat to move closer still before giving food.
    4. Require the rat to touch the bar to receive food.
    • Applied Example (MooLoo Training): 1111 cows were potty trained by reinforcing them with a treat when they entered and urinated in a specific stall designed to collect urine (Dirksen et al., 20212021). If broadly adopted, this practice could reduce environmental ammonia emissions and collect nitrogen and phosphorus for agricultural fertilizer.
    • Applied Example (Personal 5K Training): Shaping personal exercise behavior involves rewarding successive approximations—first rewarding a 15-minute15\text{-minute} walk, then a combination of walking and running that distance, then running the full distance, and finally expanding distance weekly.
  • Contextual Variability of Reinforcers:
    • Reinforcers vary across species and situations. A heat lamp is reinforcing to a cold meerkat, but not to an overheated bear; cold conditions at Sydney's Taronga Zoo alter what serves as a reinforcer compared to a hot summer day.
    • Reinforcers vary individually among humans (e.g., chocolate is reinforcing to Clarice, but Clarence prefers vanilla).

Discriminative Stimuli and Concept Formation in Animals

  • Perception Testing via Shaping:
    • Shaping allows researchers to assess what nonverbal organisms can perceive (e.g., testing if a dog discriminates red from green, or if an infant discriminates tone pitches).
    • If an organism can be conditioned to respond to one specific stimulus and not another, perceptual discrimination is proven.
  • Discriminative Stimulus Defined:
    • A discriminative stimulus is a stimulus that signals that a specific response will be reinforced (acting like a green traffic light for behavior).
  • Conceptual Capabilities in Animals:
    • Face Recognition: Pigeons reinforced for pecking only after viewing human faces learned to discriminate human faces from other images (Herrnstein & Loveland, 19641964).
    • Category Discrimination: After discrimination training, pigeons can categorize novel pictures into classes such as flowers, people, cars, and chairs (Bhatt et al., 19881988; Wasserman, 19931993).
    • Auditory Discrimination: Pigeons have been trained to discriminate between musical compositions by Bach and Stravinsky (Porter & Neuringer, 19841984).
    • Medical Imaging: Pigeons rewarded with food for correctly identifying breast tumors became as accurate as human experts at distinguishing cancerous from healthy tissue (Levenson et al., 20152015).
    • Operational Service: Animals are shaped to detect hidden explosives, sniff out illegal drugs, or locate human survivors in disaster rubble (La Londe et al., 20152015).
  • Everyday Everyday Shaping Traps:
    • Unintentional Parent-Child Reinforcement: In an interaction where Erlinda persistently nags her mother for a ride to the store until her mother yields, two operant processes occur simultaneously:
    • Erlinda’s nagging is positively reinforced because she receives a ride to the store.
    • The mother’s yielding is negatively reinforced because it terminates the aversive nagging.
    • Classroom Grade Distribution: Rewarding only student all-stars who achieve 100%100\% on spelling tests leaves hard-working students unreinforced. Operant principles suggest reinforcing students for successive approximations toward spelling challenging words correctly.

Types of Reinforcement: Positive vs. Negative

  • Definition of Reinforcement: Any consequence that strengthens or increases the frequency of a preceding behavior.
  • Positive Reinforcement:
    • Strengthens responses by presenting a desirable or pleasurable stimulus immediately after the response.
    • Examples: Petting a dog when it comes when called; paying an employee for work performed.
  • Negative Reinforcement:
    • Strengthens responses by reducing, stopping, or removing an unwanted, unpleasant, or aversive stimulus.
    • Critical Distinction: Negative reinforcement is NOT punishment. Negative reinforcement removes an aversive event to increase a behavior; punishment introduces an aversive event or removes a desirable one to decrease a behavior.
    • Examples: Taking painkillers/aspirin to remove headache pain; giving a dog a treat to silence barking; hitting a alarm clock snooze button to end an annoying sound; buckling a seatbelt to silence a warning chime; using drugs to end painful withdrawal symptoms in addiction (Baker et al., 20042004).
  • Simultaneous Co-occurrence:
    • Positive and negative reinforcement can happen at the same time. A worried student who receives a poor grade and subsequently studies harder experiences negative reinforcement through reduced anxiety and positive reinforcement through a better subsequent grade.

Primary vs. Conditioned Reinforcers and Timing

  • Primary Reinforcers:
    • Biological, unlearned reinforcers that are innately satisfying (e.g., obtaining food when hungry, escaping physical pain).
  • Conditioned (Secondary) Reinforcers:
    • Reinforcers that gain their power through learned association with primary reinforcers.
    • Experimental Example: If a light in a Skinner box turns on right before food is delivered, the rat learns to press a lever just to turn on the light.
    • Human Examples: Money, academic grades, praise, and social media "likes" (Rosenthal-von der Pütten et al., 20192019).
  • Delay of Reinforcement:
    • Animal Limitations: If an experimenter delays delivering food to a rat for more than about 30 seconds30\,\text{seconds} after a bar press, the rat will fail to learn the bar press. Instead, it will learn whatever incidental behavior (scratching, sniffing) it engaged in right before the food arrived (Austen & Sanderson, 20192019; Cunningham & Shahan, 20192019).
    • Human Educational Delay: Immediate feedback maximizes learning efficiency. Students learn academic material significantly better when given frequent quizzes with immediate feedback (Healy et al., 20172017).
    • Delayed Gratification in Humans: Unlike animals, humans can respond to delayed reinforcers (e.g., a monthly paycheck, term grades, end-of-season sports trophies).
    • The Marshmallow Study: 4-year-old4\text{-year-old} children were given a choice between eating one marshmallow immediately or waiting to receive two marshmallows later. Children who demonstrated impulse control and delayed gratification grew into adults with higher social competence and academic achievement (Mischel, 20142014; Watts et al., 20182018). Impulse control also correlates with reduced rates of adult criminal behavior (Åkerlund et al., 20162016; Logue, 1998a1998\text{a}, 1998b1998\text{b}).
    • Short-Term vs. Long-Term Trade-Offs: Immediate small rewards often dangerously override larger delayed rewards (e.g., late-night media consumption over test rest; unprotected sex over long-term safety; immediate fossil fuel comfort over global climate stability).

Reinforcement Schedules

  • Continuous Reinforcement Schedule:
    • Reinforcing the desired response every single time it occurs.
    • Characteristics: Learning/acquisition occurs rapidly, making it ideal for mastering new behaviors. However, extinction also occurs rapidly when reinforcement stops (e.g., if a reliable vending machine fails to deliver candy twice in a row, people stop inserting money).
  • Partial (Intermittent) Reinforcement Schedule:
    • Reinforcing responses only part of the time.
    • Characteristics: Acquisition is slower, but resistance to extinction is far greater than with continuous reinforcement.
    • Experimental Example: A pigeon conditioned on a gradually phased-out intermittent schedule pecked a key 150,000150{,}000 times without receiving any food reward (Skinner, 19531953).
    • Gambling Analogy: Slot machines and fishing reward individuals unpredictably, creating persistent behavioral repetition.
  • Superstitious Behaviors:
    • Accidental timing of reinforcers can produce superstitious habits. If an automatic feeder dispenses food while an animal happens to be scratching itself, that behavior is accidentally reinforced. Similarly, a gambler wearing a specific bracelet or a baseball player tapping home plate after getting a hit retains the behavior through intermittent reinforcement.
  • The Four Partial Reinforcement Schedules (Skinner, 19611961):
    • Fixed-Ratio (FR) Schedule:
    • Reinforces behavior after a set, specified number of responses.
    • Examples: Buy 1010 coffees, get 11 free; workers paid per unit produced; rats reinforced with 11 food pellet for every 3030 bar presses.
    • Response Pattern: High rate of responding with only a brief pause after receiving the reinforcer.
    • Variable-Ratio (VR) Schedule:
    • Reinforces behavior after an unpredictable, random number of responses.
    • Examples: Slot machines, fly fishing.
    • Response Pattern: Very high, consistent response rates because reinforcers increase with the number of responses.
    • Fixed-Interval (FI) Schedule:
    • Reinforces the first response after a fixed, set period of time.
    • Examples: Checking the mail near delivery time; Tuesday discount pricing.
    • Response Pattern: A "scalloped" pattern on cumulative response graphs; responding accelerates as the expected time for the reward approaches.
    • Variable-Interval (VI) Schedule:
    • Reinforces the first response after varying, unpredictable time intervals.
    • Examples: Checking a smartphone for expected messages.
    • Response Pattern: Slow, steady responding due to uncertainty about when the interval will end.
  • General Rules of Schedules:
    • Ratio schedules (linked to number of responses) yield higher response rates than interval schedules (linked to time elapsed).
    • Variable schedules (unpredictable) produce more consistent responding than fixed schedules (predictable).
  • Universality Claim: Skinner (19561956) asserted that operant reinforcement schedules function identically regardless of the species, response, or specific reinforcer ("Pigeon, rat, monkey, which is which? It doesn't matter…").

Punishment and Its Consequences

  • Definition: Punishment is any consequence that decreases the frequency of a preceding behavior.
  • Types of Punishment:
    • Positive Punishment: Administering an aversive stimulus following a behavior (e.g., spraying water on a barking dog, issuing a traffic ticket for speeding, shocking a rat).
    • Negative Punishment: Withdrawing a desirable stimulus following a behavior (e.g., revoking a misbehaving teen's driving privileges, blocking a rude social media commenter, taking away a child's toy).
  • Swiftness and Sureness vs. Severity:
    • Swift and sure punishments restrain unwanted behavior far more effectively than severe, delayed threats.
    • Example: Arizona's unusually harsh prison terms for first-time drunk drivers had little impact on drunk-driving rates, whereas Kansas City's strategy of increasing police patrols to ensure swift detection dramatically reduced crime (Darley & Alter, 20132013).
  • Major Drawbacks of Physical Punishment:
    • Meta-analysis of over 160,000160{,}000 children demonstrates that physical punishment fails to correct unwanted behavior effectively (Gershoff & Grogan-Kaylor, 20162016; APA, 2019a2019\text{a}).
    • Drawback 1: Punished behavior is suppressed, not forgotten. The temporary suppression negatively reinforces the parent's punishing behavior (child stops swearing temporarily, leading parent to believe spanking worked). Over 22 in 33 children worldwide experience regular corporal punishment (UNICEF, 20202020).
    • Drawback 2: Punishment does not teach appropriate behavior. It tells the organism what not to do, but fails to provide direction for positive alternative behaviors.
    • Drawback 3: Punishment teaches discrimination among situations. Children learn not to perform the forbidden behavior in front of the punishing agent (parents), but continue it elsewhere.
    • Drawback 4: Punishment creates fear through stimulus generalization. Children associate fear not only with the misbehavior, but also with the person administering punishment or the location (leading to school avoidance or anxiety; Gershoff et al., 20102010). Currently, 6363 countries and 3131 U.S. states ban corporal punishment in schools, with Finland reporting reduced physical abuse following child-protection legislation (Österman et al., 20142014).
    • Drawback 5: Punishment increases aggression by modeling violence as a problem-solving strategy (MacKenzie et al., 20132013; Fitton et al., 20202020). (Note: Non-experimental correlational designs cannot fully rule out genetic predispositions or pre-existing child misbehavior as confounding variables; Ferguson, 20132013; Larzelere, 20002000; Larzelere et al., 20192019).
  • Recommended Alternatives for Discipline:
    • Time-Out: Removing a child from access to positive reinforcement (e.g., parental/sibling attention; Dadds & Tully, 20192019; O'Leary et al., 19671967; Patterson et al., 19681968).
    • Positive Reframing: Reframing negative threats ("Clean your room or no dinner!") into positive incentive contingencies ("You are welcome at the table as soon as your room is cleaned"; Patterson et al., 19821982).
    • Immediate Constructive Feedback: Praising successful effort yields better long-term development than highlighting failures (Eskreis-Winkler & Fishbach, 20192019).
    • Moral Framing: Punishment enforces a morality centered on prohibition (what not to do), whereas reinforcement fosters a morality of positive obligation (Sheikh & Janoff-Bulman, 20132013).

B. F. Skinner’s Legacy and Philosophical Controversy

  • Behaviorist Stance:
    • Skinner argued that external environmental influences, rather than internal conscious thoughts, intentions, or feelings, dictate behavior.
    • Asserted that psychological behavioral science functions entirely independently of neurology (Skinner, 1938/19661938/1966).
    • Stated his own personal actions were purely "the product of my genetic endowment, my personal history, and the current setting" (Skinner, 19831983).
  • Core Controversies:
    • Critics alleged that Skinner dehumanized people, denied personal freedom, and sought total behavioral control.
    • Skinner's Rebuttal: External consequences already control human behavior randomly. Using operant principles systematically to improve education, work, and homes is more humane than relying on punitive controls.

Real-World Applications of Operant Conditioning

  • Applications in Education:
    • Skinner observed a fourth-grade math class in 19531953 and noted the inefficiency of uniform instruction pacing (Watters, 20212021).
    • Envisioned teaching machines and programmed textbooks that individualize learning, shaping knowledge in small steps with immediate reinforcement (Skinner, 19891989).
    • Modern computerized testing and interactive quizzing fulfill this vision by providing individualized pacing and immediate corrective feedback.
  • Applications in Sports:
    • Effective sports instruction shapes performance by reinforcing small baseline successes and progressively raising difficulty.
    • Putting: Starting with extremely short putts and stepping back incrementally as mastery increases.
    • Batting: Novice batters start with half-swings at oversized balls pitched from 10 feet10\,\text{feet}, gradually transitioning to standard baseballs pitched from full distance (Simek & O'Brien, 19811981, 19881988).
    • Coaching: Notre Dame basketball coach Muffet McGraw emphasized immediately catching players doing something correctly and praising them on the spot.
  • Applications in Artificial Intelligence:
    • AI developers design reinforcement algorithms that allow software agents to play complex games (chess, poker, Quake III) faster than humans, reinforcing winning actions and penalizing losing moves (Botvinick et al., 20192019; Jaderberg et al., 20192019).
  • Applications in the Workplace:
    • Reinforcement should be immediate and linked to specific, achievable performance behaviors rather than general "merit."
    • General Motors CEO Mary Barra (20152015) implemented record profit-sharing bonuses following strong worker output (Vlasic, 20152015).
    • Informal management behaviors, such as walking the workspace and delivering sincere verbal praise, effectively motivate staff.
  • Applications in Parenting:
    • Parents must avoid reinforcing bad behavior by giving in to tantrums or screaming (Wierson & Forehand, 19941994).
    • Parents can shape desirable behavior (e.g., safe teen driving; Hinnant et al., 20192019) by noticing correct actions and affirming them, while handling misbehavior with non-violent privileges removal or time-outs.
  • Step-by-Step Self-Improvement (Behavior Modification):
    1. State a realistic goal in measurable terms and announce it publicly.
    2. Decide how, when, and where to work toward the goal (implementation planning; Gollwitzer & Oettingen, 20122012; van Gelderen et al., 20182018).
    3. Monitor behavior frequency using tracking logs.
    4. Reinforce desired target behaviors with immediate rewards (Woolley & Fishbach, 20172017).
    5. Gradually reduce external rewards as internal habituation takes hold.

Biological Constraints on Operant Conditioning

  • Species-Specific Predispositions:
    • An animal's capacity for operant conditioning is bounded by its evolutionary biology (Robert Heinlein quote: "Never try to teach a pig to sing; it wastes your time and annoys the pig").
    • Hamsters: Easily conditioned with food to dig or rear up (natural food-search behaviors), but cannot easily be conditioned with food to wash their faces (Shettleworth, 19731973).
    • Pigeons: Easily learn to flap wings to avoid shock or peck to obtain food (natural escape/feeding mechanics), but struggle to learn key-pecking to avoid shock or wing-flapping to get food (Foree & LoLordo, 19731973).
  • Instinctive Drift:
    • Animal trainers Marian Breland and Keller Breland observed that trained animals frequently experience instinctive drift—reverting to biologically inherent behavioral patterns during operant tasks.
    • Example: Pigs trained to pick up large wooden dollars and deposit them in a piggy bank began dropping the coins, pushing them with their snouts, picking them up, and pushing them again, delaying their food rewards due to innate rooting instincts.

Comparing Classical and Operant Conditioning

  • Shared Characteristics: Both classical and operant conditioning are forms of associative learning. Both involve acquisition, extinction, spontaneous recovery, generalization, and discrimination.
  • Key Distinctions Summary:
    • Basic Idea:
    • Classical Conditioning: Learning associations between environmental events that the organism does not control.
    • Operant Conditioning: Learning associations between voluntary behavior and its resulting consequences.
    • Response Type:
    • Classical Conditioning: Involuntary, automatic respondent behaviors.
    • Operant Conditioning: Voluntary operant behaviors that act on the environment.
    • Acquisition Mechanism:
    • Classical Conditioning: Associating events; pairing a Neutral Stimulus (NS) with an Unconditioned Stimulus (UCS) so it becomes a Conditioned Stimulus (CS).
    • Operant Conditioning: Associating a voluntary response with a consequence (reinforcer or punisher).
    • Extinction Process:
    • Classical Conditioning: Conditioned Response (CR) decreases when the CS is repeatedly presented without the UCS.
    • Operant Conditioning: Responding decreases when reinforcement ceases.
    • Spontaneous Recovery:
    • Classical Conditioning: Reappearance, after a rest period, of a weakened CR.
    • Operant Conditioning: Reappearance, after a rest period, of an extinguished operant response.
    • Generalization:
    • Classical Conditioning: Tendency to respond to stimuli similar to the CS.
    • Operant Conditioning: Performing responses learned in one situation in other, similar situations.
    • Discrimination:
    • Classical Conditioning: Learning to distinguish between a CS and other stimuli that do not signal a UCS.
    • Operant Conditioning: Learning that specific responses will be reinforced in certain contexts, but not in others.