Reinforcement – Foundations, Superstition, Contingencies & Shaping
Administrative Announcements
- Lecturer: Patrick Heslop (animal-lab background; honours + masters with Prof. Randolph Grace)
- Lecture blocks: today + two next week → all on Reinforcement
- Microphone: lapel so he can wander; signal if audio issues
- Class-rep call: need 2 reps for PSYC384
- Email Patrick or John; self-sign-up info will be posted; can give details at end of class
- Labs
- Start Monday (next week)
- First lab slot is immediately before Monday lecture
- Duration ≈ 50 min (listed as 1 h on Learn; will finish on time)
- Swipe-card access; can leave early—inform demonstrator
- Patrick preps the (non-invasively treated) rats – “little readies” – for operant work
- Email Patrick if any sign-up problems
Recap of Previous Lecture (John)
- Behaviour → treated as science (EAB: Experimental Analysis of Behaviour) and technology (ABA: Applied Behaviour Analysis)
- Goal: explain behaviour via interaction with environment, not via internal/mentalistic constructs
- Key constructs
- Three-term contingency (ABC): Antecedent → Behaviour → Consequence
- Environmental rules vs Organism rules (e.g.
- Over-eating Mars bars changes internal state & environmental availability)
- Control by context (antecedents) and control by consequences
Warm-up Discussion Question
- Prompt to students: “What does it mean to reinforce behaviour?”
- Common answers shared
- "Rewarding good behaviour so habit forms"
- Positive vs negative reinforcement = adding/removing stimulus to increase behaviour (≠ good/bad)
- Punishment = adding/removing stimulus to decrease behaviour
Reinforcement: Definitions & Circularity
- Danger of circular definition
- Q: “What is a reinforcer?”
- A: “An event that increases behaviour.”
- Then: “What events increase behaviour?” → “Reinforcers.”
- Need a-priori (pre-specified) list of stimuli expected to act as reinforcers
- Contrast with post-hoc analysis
Historical Foundations of Reinforcement Theory
- Puzzle-box experiments with cats
- Tilt lever / depress treadle → door opens → food outside
- Measured escape latency ↓ across trials → learning curve
- Law
- Positive Law: satisfying consequence “stamps-in” S-R connection → ↑ probability
- Negative Law: annoying consequence weakens connection → ↓ probability
- Contributions & issues
- Introduces strengthening/weakening of connections
- “Satisfying/annoying” vague; still circular
- Highlights temporal contiguity: reinforcer must follow response quickly
Temporal Contiguity Principle
- Learning most effective when R and Sr closely paired in time
- Lab example: rat lever-press → milk dipper must occur within a few seconds
Superstition & Adventitious Reinforcement
Skinner’s 1948 Pigeon Study
- 8 pigeons in operant chambers; food on Fixed-Time (FT) schedule (no response required)
- 6/8 birds developed stereotyped behaviour patterns (e.g.
- CCW spins, head-swaying, foot-hopping)
- Behaviours occurred in scalloped pattern—highest rate just before pellet delivery
- Demonstrates superstition: response reinforced by chance (no causal relation)
- Term: Adventitious reinforcement
Human Analogue – Catania & Cutts (concurrent EXT/VI)
- Two physical buttons
- Left = Extinction (never reinforced)
- Right = Variable-Interval 30 s (VI-30) → points
- Light flashed ≈ 100 × min⁻¹ → participants had to press somewhere
- Results
- All participants pressed the useless left button; some > right button
- Post-experiment rules: complex “rituals” invented (“press left twice then right”)
- Change-Over Delay (COD) (~2–3 s) can eliminate superstitious switching by withholding Sr after a switch
Everyday Superstitions
- Often about preventing aversive events (e.g.
- Step on a crack → break your back
- Throw salt over shoulder; avoid 3 cigarettes/1 match)
- Hard to extinguish because evidence is absence of event
Detecting Causality – Killeen (1978)
- 4 pigeons, 3 response keys
- Centre key peck: p=0.05 turns centre dark + illuminates both side keys
- Background computer emits simulated “pecks” with same probability
- Choice phase
- Left key = “I caused it” → 4 s food (example magnitude)
- Right key = “Computer caused it” → 2 s food
- Findings
- ~80% correct attribution overall
- Bias toward option delivering larger reinforcer (magnitude manipulation 4 s vs 2 s, 3.8 s vs 1.8 s, etc.)
- Suggests causal judgments influenced by relative payoff, not just contingency detection
- Implication: human attributions (success, failure, wealth, accidents) may reflect reinforcement-biased self-causation
Contingencies & Response Classes
- Contingency = If–Then relation (dependency): If R → Sr
- Skinner superstition: non-contingent Sr
- Killeen task: Sr contingent on correct left/right report
- Response Class: set of topographically different responses producing same outcome
Guthrie & Horton (1946) – Cat Pole-Tilt
- Cats in box with vertical pole; any tilt opens door + food
- Camera photographed cat posture each successful trial
- Results
- All cats learned quickly
- Idiosyncratic topographies: left-paw push, bite, roll-on-pole, etc.
- Early trials → high variability; later trials → stereotypy (response class narrowed)
- Shows behaviour immediately preceding Sr becomes more probable next trial
Shaping by Successive Approximations
- Powerful method to train low-probability target behaviours
- Procedure
- Identify broad initial response class likely to occur
- Reinforce any instance (e.g. rat merely looks at lever)
- Gradually raise criterion
- Look → step toward lever → touch → press with one paw → full depress
- Maintain tight temporal contiguity at each step
- Widely used in service-dog training, animal acts, skill acquisition
- Demo video: dog “Winchester” shaped to pick up duct tape & deliver to hand
Practical Lab Connections
- Students will shape rats to lever-press for condensed milk in operant chambers
- Keep delays < a few seconds for effective contiguity
- Observe emergence/narrowing of response classes; note any superstitious adjunctive behaviour
Ethical / Philosophical Notes
- Non-invasive procedures with rats; focus on humane treatment
- Reminder: correlation ≠ causation; behavioural science strives for functional (causal) analyses while avoiding mentalistic shortcuts
Key Terms & Symbols (add to personal glossary)
- Reinforcer (Sr) / Reinforcement
- Punishment
- Positive / Negative (≠ good/bad; = add/remove stimulus)
- Three-term contingency (ABC)
- Temporal Contiguity
- A priori vs Post-hoc
- Fixed-Time (FT), Variable-Interval (VI-t)
- Extinction (EXT)
- Change-Over Delay (COD)
- Adventitious Reinforcement
- Superstition
- Contingency
- Response Class / Topography
- Shaping / Successive Approximations
Numerical & Schedule References
- p=0.05 probability (Killeen) for stimulus change per centre peck
- VI-30 s schedule ≈ reinforcement available on average every 30s
- Change-Over Delay ≈2!–3s
- Killeen magnitude manipulations: 4s vs 2s food; 3.8s vs 1.8s, etc.
Studies Mentioned (Chronological)
- Thorndike (1898/1921) – Puzzle-box & Law of Effect
- Skinner (1948) – Pigeon superstition (FT schedule)
- Guthrie & Horton – Cat pole-tilt idiosyncrasy
- Catania & Cutts – Concurrent EXT / VI-30 & superstitious humans (+ COD)
- Killeen (1978) – Pigeons reporting causality vs computer
Consolidated Take-Home Points
- Law of Effect: consequences “stamp-in/out” behaviour
- Temporal contiguity is critical; causal relation helpful but not essential
- Adventitious reinforcement → superstitious behaviour in animals & humans
- Explicit contingencies let us train and study behaviour systematically
- Shaping capitalises on reinforcement to build complex actions via gradual approximations