Reinforcement – Foundations, Superstition, Contingencies & Shaping

Administrative Announcements

  • Lecturer: Patrick Heslop (animal-lab background; honours + masters with Prof. Randolph Grace)
  • Lecture blocks: today + two next week → all on Reinforcement
  • Microphone: lapel so he can wander; signal if audio issues
  • Class-rep call: need 2 reps for PSYC384
    • Email Patrick or John; self-sign-up info will be posted; can give details at end of class
  • Labs
    • Start Monday (next week)
    • First lab slot is immediately before Monday lecture
    • Duration ≈ 50 min (listed as 1 h on Learn; will finish on time)
    • Swipe-card access; can leave early—inform demonstrator
    • Patrick preps the (non-invasively treated) rats – “little readies” – for operant work
    • Email Patrick if any sign-up problems

Recap of Previous Lecture (John)

  • Behaviour → treated as science (EAB: Experimental Analysis of Behaviour) and technology (ABA: Applied Behaviour Analysis)
  • Goal: explain behaviour via interaction with environment, not via internal/mentalistic constructs
  • Key constructs
    • Three-term contingency (ABC): Antecedent → Behaviour → Consequence
    • Environmental rules vs Organism rules (e.g.
    • Over-eating Mars bars changes internal state & environmental availability)
    • Control by context (antecedents) and control by consequences

Warm-up Discussion Question

  • Prompt to students: “What does it mean to reinforce behaviour?”
  • Common answers shared
    • "Rewarding good behaviour so habit forms"
    • Positive vs negative reinforcement = adding/removing stimulus to increase behaviour (≠ good/bad)
    • Punishment = adding/removing stimulus to decrease behaviour

Reinforcement: Definitions & Circularity

  • Danger of circular definition
    • Q: “What is a reinforcer?”
    • A: “An event that increases behaviour.”
    • Then: “What events increase behaviour?” → “Reinforcers.”
  • Need a-priori (pre-specified) list of stimuli expected to act as reinforcers
    • Contrast with post-hoc analysis

Historical Foundations of Reinforcement Theory

Thorndike’s Law of Effect (1898 → formalised 1921)
  • Puzzle-box experiments with cats
    • Tilt lever / depress treadle → door opens → food outside
    • Measured escape latency ↓ across trials → learning curve
  • Law
    • Positive Law: satisfying consequence “stamps-in” S-R connection → ↑ probability
    • Negative Law: annoying consequence weakens connection → ↓ probability
  • Contributions & issues
    • Introduces strengthening/weakening of connections
    • “Satisfying/annoying” vague; still circular
    • Highlights temporal contiguity: reinforcer must follow response quickly
Temporal Contiguity Principle
  • Learning most effective when R and Sr closely paired in time
  • Lab example: rat lever-press → milk dipper must occur within a few seconds

Superstition & Adventitious Reinforcement

Skinner’s 1948 Pigeon Study
  • 8 pigeons in operant chambers; food on Fixed-Time (FT) schedule (no response required)
  • 6/8 birds developed stereotyped behaviour patterns (e.g.
    • CCW spins, head-swaying, foot-hopping)
  • Behaviours occurred in scalloped pattern—highest rate just before pellet delivery
  • Demonstrates superstition: response reinforced by chance (no causal relation)
  • Term: Adventitious reinforcement
Human Analogue – Catania & Cutts (concurrent EXT/VI)
  • Two physical buttons
    • Left = Extinction (never reinforced)
    • Right = Variable-Interval 30 s (VI-30) → points
    • Light flashed ≈ 100 × min⁻¹ → participants had to press somewhere
  • Results
    • All participants pressed the useless left button; some > right button
    • Post-experiment rules: complex “rituals” invented (“press left twice then right”)
  • Change-Over Delay (COD) (~2–3 s) can eliminate superstitious switching by withholding Sr after a switch
Everyday Superstitions
  • Often about preventing aversive events (e.g.
    • Step on a crack → break your back
    • Throw salt over shoulder; avoid 3 cigarettes/1 match)
  • Hard to extinguish because evidence is absence of event

Detecting Causality – Killeen (1978)

  • 4 pigeons, 3 response keys
    • Centre key peck: p=0.05p=0.05 turns centre dark + illuminates both side keys
    • Background computer emits simulated “pecks” with same probability
  • Choice phase
    • Left key = “I caused it” → 4 s food (example magnitude)
    • Right key = “Computer caused it” → 2 s food
  • Findings
    • ~80%80\% correct attribution overall
    • Bias toward option delivering larger reinforcer (magnitude manipulation 4 s vs 2 s, 3.8 s vs 1.8 s, etc.)
    • Suggests causal judgments influenced by relative payoff, not just contingency detection
  • Implication: human attributions (success, failure, wealth, accidents) may reflect reinforcement-biased self-causation

Contingencies & Response Classes

  • Contingency = If–Then relation (dependency): If R → Sr
    • Skinner superstition: non-contingent Sr
    • Killeen task: Sr contingent on correct left/right report
  • Response Class: set of topographically different responses producing same outcome
Guthrie & Horton (1946) – Cat Pole-Tilt
  • Cats in box with vertical pole; any tilt opens door + food
  • Camera photographed cat posture each successful trial
  • Results
    • All cats learned quickly
    • Idiosyncratic topographies: left-paw push, bite, roll-on-pole, etc.
    • Early trials → high variability; later trials → stereotypy (response class narrowed)
  • Shows behaviour immediately preceding Sr becomes more probable next trial

Shaping by Successive Approximations

  • Powerful method to train low-probability target behaviours
  • Procedure
    • Identify broad initial response class likely to occur
    • Reinforce any instance (e.g. rat merely looks at lever)
    • Gradually raise criterion
    • Look → step toward lever → touch → press with one paw → full depress
    • Maintain tight temporal contiguity at each step
  • Widely used in service-dog training, animal acts, skill acquisition
  • Demo video: dog “Winchester” shaped to pick up duct tape & deliver to hand

Practical Lab Connections

  • Students will shape rats to lever-press for condensed milk in operant chambers
  • Keep delays < a few seconds for effective contiguity
  • Observe emergence/narrowing of response classes; note any superstitious adjunctive behaviour

Ethical / Philosophical Notes

  • Non-invasive procedures with rats; focus on humane treatment
  • Reminder: correlation ≠ causation; behavioural science strives for functional (causal) analyses while avoiding mentalistic shortcuts

Key Terms & Symbols (add to personal glossary)

  • Reinforcer (Sr) / Reinforcement
  • Punishment
  • Positive / Negative (≠ good/bad; = add/remove stimulus)
  • Three-term contingency (ABC)
  • Temporal Contiguity
  • A priori vs Post-hoc
  • Fixed-Time (FT), Variable-Interval (VI-tt)
  • Extinction (EXT)
  • Change-Over Delay (COD)
  • Adventitious Reinforcement
  • Superstition
  • Contingency
  • Response Class / Topography
  • Shaping / Successive Approximations

Numerical & Schedule References

  • p=0.05p=0.05 probability (Killeen) for stimulus change per centre peck
  • VI-30 s schedule ≈ reinforcement available on average every 30 s30\,\text{s}
  • Change-Over Delay ≈2!–3 s\approx 2!\text{–}3\,\text{s}
  • Killeen magnitude manipulations: 4 s4\,\text{s} vs 2 s2\,\text{s} food; 3.8 s3.8\,\text{s} vs 1.8 s1.8\,\text{s}, etc.

Studies Mentioned (Chronological)

  • Thorndike (1898/1921) – Puzzle-box & Law of Effect
  • Skinner (1948) – Pigeon superstition (FT schedule)
  • Guthrie & Horton – Cat pole-tilt idiosyncrasy
  • Catania & Cutts – Concurrent EXT / VI-30 & superstitious humans (+ COD)
  • Killeen (1978) – Pigeons reporting causality vs computer

Consolidated Take-Home Points

  • Law of Effect: consequences “stamp-in/out” behaviour
  • Temporal contiguity is critical; causal relation helpful but not essential
  • Adventitious reinforcement → superstitious behaviour in animals & humans
  • Explicit contingencies let us train and study behaviour systematically
  • Shaping capitalises on reinforcement to build complex actions via gradual approximations