Untitled
Quick recap: classical conditioning to operant conditioning
- Classical conditioning (CC) involves learning associations between stimuli and automatic responses.
- Example discussed: A student projects anxiety from a surprise quiz to the cue that precedes it.
- In that example: the surprise pop quiz is the unconditioned stimulus (US) that naturally elicits anxiety (unconditioned response, UR);
the phrase "clear your desks" becomes a conditioned stimulus (CS) after association with the quiz, and the anxiety becomes the conditioned response (CR).
- Operant conditioning (OC) involves active learning and voluntary behaviors shaped by consequences.
- Emphasizes shaping behavior through reinforcement and punishment to increase or decrease behaviors.
- Applies to humans and animals; used in education, parenting, therapy, and pet training.
Thorndike: the pioneer of operant conditioning
- Thorndike’s key contribution: systematic investigation of animal learning and trial-and-error in natural environments.
- Law of Effect: behaviors followed by satisfying consequences are strengthened; those followed by annoying/punishing consequences are weakened.
- Experimental setup: puzzle box with a cat; food outside; cat learns to perform actions that lead to escape and reward.
- Early idea: random trial-and-error exploration gradually yields coordinated, successful behaviors (instrumental learning).
- Takeaway: consequences (rewarding or punishing) shape the likelihood of future voluntary behaviors.
Skinner: operationalizing operant conditioning
- B.F. Skinner popularized operant conditioning as a theory describing how organisms learn from consequences.
- Skinner was a behaviorist: emphasized objective measurement and outwardly observable behavior; internal thoughts/emotions were not the primary focus.
- Key concept: operant conditioning (OC) is a learning process that changes the probability of a response through manipulation of consequences.
Discriminative stimulus (SD)
- A cue that signals reinforcement is available if a particular response is made.
- Examples: text message notification, green traffic light, doorbell, etc.
- In Skinner’s experiments, a light often served as the discriminative stimulus.
Reinforcement vs punishment (OC)
- Reinforcement: increases the likelihood of a response.
- Positive reinforcement: add a desirable stimulus after a response (e.g., giving a treat to reward behavior).
- Negative reinforcement: remove an aversive stimulus after a response (e.g., stopping loud noise when a desired behavior occurs).
- Punishment: decreases the likelihood of a response.
- Positive punishment: add an aversive stimulus to reduce a behavior (e.g., chores, hot sauce for misbehavior).
- Negative punishment: remove a desirable stimulus to reduce a behavior (e.g., taking away allowance).
- Core rule: reinforcement (positive or negative) always increases the likelihood of the behavior; punishment (positive or negative) aims to decrease it.
- Example narrative in lecture: giving chocolates can function as a positive reinforcement to encourage a behavior; using punishment (e.g., extra chores or removal of privileges) to curb behavior.
Primary vs secondary reinforcers
- Primary reinforcers: naturally reinforcing (biological needs) – food, water, basic needs.
- Secondary reinforcers: learned value (money, awards, frequent-flyer points) that acquire reinforcement value by association with primary reinforcers.
- Key idea: secondary reinforcers gain power through learned associations, enabling them to help obtain primary needs.
Shaping and the Skinner box
- Shaping: gradually reinforcing closer approximations to the target behavior.
- Skinner box (operant chamber): used with pigeons and rats to illustrate shaping.
- For pigeons: pecking a disc yields food; for rats: pressing a bar yields food.
- Process: start with simple behavior and reinforce successive approximations that move toward the desired behavior.
Schedules of reinforcement (partially vs continuous)
- Continuous reinforcement: reward is given after every instance of the target behavior; fastest way to learn a new behavior.
- Partial (intermittent) reinforcement: reinforcement occurs on a schedule, not every time; slower to acquire, but more resistant to extinction.
- Four main schedules of partial reinforcement:
- Fixed ratio (FR): reward after a fixed number of responses.
- Variable ratio (VR): reward after an unpredictable number of responses; high and steady response rates (e.g., gambling, lottery).
- Fixed interval (FI): reward after a fixed amount of time has passed.
- Variable interval (VI): reward after an average time interval, varying around that average.
- Real-life examples and notes:
- Example of continuous reinforcement: attendance sheet rewarded every class session.
- Gambling often operates on a variable-ratio schedule, contributing to its persistence.
- Partial reinforcement effect: extinction is slower under partial reinforcement than continuous reinforcement.
- In pigeons: when continuously reinforced, a pigeon may stop quickly after extinction; when partially reinforced, extinction takes much longer (illustrative numbers from lecture: Nc = 100 vs Np = 1000; Np ≈ 10 Nc).
- Expression: indicating greater resistance to extinction under partial reinforcement.
Extinction, punishment, and practical applications
- Extinction occurs when reinforcement is withdrawn, and the behavior gradually diminishes.
- Partial reinforcement can make behaviors more resilient to extinction, which has implications for gambling, habits, and long-term training in sports or education.
- Practical applications: coaches use reinforcement/punishment to shape athletic performance; parents use rewards and consequences to manage behavior; therapists and educators leverage these principles to reduce smoking, curb misbehavior, and encourage study habits.
Cognitive and evolutionary factors in OC
- Tolman and cognitive aspects: learning can involve cognitive maps and latent learning.
- Cognitive map: rats in mazes can form mental representations of the layout, enabling efficient navigation later, even without immediate reward.
- Late learning: learning can occur without immediate reinforcement and reveal itself later when rewards become available.
- Example analogy: humans learning routes in a familiar town and recalling them when needed.
- Tolman emphasized that learning is not purely stimulus-response; cognition mediates learning processes.
- Evolutionary influences: some behaviors resist conditioning due to instinctive drift or natural tendencies that are hard to override with new conditioning.
- Instinctive drift example: raccoon coin task – raccoons instinctively wash their food and tended to revert to natural behaviors rather than fully learning to place coins into a box; instinctual predispositions can impede conditioned responses.
Learned helplessness and its implications
- Learned helplessness (L.H.): a state in which individuals believe outcomes are uncontrollable and feel powerless to change their situation.
- Classic dog experiment: dogs conditioned to associate a sound with an unavoidable shock; when later given a chance to escape the shock by jumping a barrier, many did not attempt it because they had learned that nothing they did would change the outcome.
- Human implications: can contribute to depression, perceptions of lack of control, homelessness, or chronic stress; demonstrates how exposure to uncontrollable events can affect motivation and future behavior.
- Relevance to positive psychology and therapy: understanding L.H. informs interventions that restore perceived control and self-efficacy.
Observational learning (Bandura) and the Bobo doll study
- Observational learning (modeling) involves learning by watching others and the consequences of their actions.
- Bandura’s work showed that people learn not only from direct reinforcement but also by observing others being rewarded or punished for their behavior.
- The Bobo doll studies (and replication) demonstrated that children imitate aggression after observing an aggressive model, especially when the model is rewarded for the aggressive behavior.
- Key idea: learning can occur without direct reinforcement and can be guided by the expectation of rewards for modeled behavior.
Factors that increase imitation (observational learning)
- Imitation is more likely when the model is perceived as:
- Warm and nurturing (positive, approachable).
- Having control or power over the observer (authority figure, parent, teacher, boss).
- Similar to the observer (age, sex, interests).
- Perceived to have higher social status or to be a potential role model.
- In a context where the task is neither too easy nor too hard (moderate difficulty).
- The observer lacks confidence or is in an unfamiliar/ambiguous situation (seeks cues from others).
- Additional context: people imitate when the behavior is rewarded or when they anticipate a reward for replicating it; observation in a new or ambiguous situation increases reliance on others’ behavior.
- Elevator study anecdote: people tended to mimic others in unfamiliar situations (e.g., standing backward in an elevator) even without understanding the origin of the behavior.
Real-world relevance and ethical considerations
- Education and classrooms: using reinforcement schedules to encourage attendance, participation, and study habits.
- Parenting and sports coaching: shaping desired behaviors with a mix of reinforcement and appropriate punishment to build long-term habits.
- Therapy and behavior modification: leveraging reinforcement/punishment to reduce problematic behaviors and encourage healthier alternatives.
- Public health: applying OC and observational learning to promote smoking cessation, healthy eating, and risk-reduction behaviors.
- Ethical considerations: avoid abusive or coercive use of punishment; emphasize humane, ethical training and avoid reinforcing maladaptive behaviors.
Quick glossary of key terms
- Unconditioned stimulus (US): a stimulus that naturally elicits a response without prior learning (e.g., surprise quiz).
- Unconditioned response (UR): natural, unlearned reaction to the US (e.g., anxiety to a surprise quiz).
- Conditioned stimulus (CS): previously neutral stimulus that, after association with the US, triggers a response (e.g., cue phrase like "clear your desks").
- Conditioned response (CR): learned response to the CS (e.g., anxiety to the cue).
- Discriminative stimulus (SD): cue signaling that reinforcement is available after a behavior.
- Positive reinforcement: add a desirable stimulus to increase a behavior.
- Negative reinforcement: remove an aversive stimulus to increase a behavior.
- Positive punishment: add an aversive stimulus to decrease a behavior.
- Negative punishment: remove a desirable stimulus to decrease a behavior.
- Primary reinforcers: naturally reinforcing (food, water).
- Secondary reinforcers: learned value (money, awards).
- Shaping: reinforcing successive approximations toward a target behavior.
- Extinction (OC): when reinforcement is removed, the behavior weakens and may disappear.
- Schedules of reinforcement: FR, VR, FI, VI (see definitions above).
- Partial reinforcement extinction effect: behaviors reinforced on partial schedules take longer to extinguish than those reinforced continuously.
- Latent/late learning (Tolman): learning can occur without immediate reinforcement and reveal itself later.
- Cognitive map: mental representation of a physical environment.
- Learned helplessness: belief that outcomes are uncontrollable, leading to passivity.
- Observational learning (Bandura): learning by watching others and the consequences of their actions.
Summary of formulas and examples
- Partial reinforcement extinction resistance example: if continuous reinforcement yields Nc = 100 trials to extinction and partial reinforcement yields Np = 1000 trials, then
- Schedules overview (labels only):
- FR-n: after every n responses.
- VR-n: after an unpredictable (average) n responses.
- FI-T: after a fixed time interval T.
- VI-T: after an average time interval T.
- Illustrative note: reinforcement always increases the probability of the target response; punishment always decreases it (in the context of OC).
Connecting to broader learning science
- OC complements CC: CC emphasizes antecedents and automatic responses, while OC emphasizes consequences and voluntary actions.
- Observational learning highlights social and cognitive dimensions: we learn from others through imitation and expectancy of reward, not just direct outcomes.
- Evolution and instinct: not all learned behaviors are equally malleable; natural tendencies can oppose conditioning (instinctive drift).
- Real-world applicability: these principles are used to shape behavior across settings (schools, homes, clinics, workplaces) while considering ethical implications and individual differences.