Untitled

Quick recap: classical conditioning to operant conditioning

  • Classical conditioning (CC) involves learning associations between stimuli and automatic responses.
    • Example discussed: A student projects anxiety from a surprise quiz to the cue that precedes it.
    • In that example: the surprise pop quiz is the unconditioned stimulus (US) that naturally elicits anxiety (unconditioned response, UR);
      the phrase "clear your desks" becomes a conditioned stimulus (CS) after association with the quiz, and the anxiety becomes the conditioned response (CR).
  • Operant conditioning (OC) involves active learning and voluntary behaviors shaped by consequences.
    • Emphasizes shaping behavior through reinforcement and punishment to increase or decrease behaviors.
    • Applies to humans and animals; used in education, parenting, therapy, and pet training.

Thorndike: the pioneer of operant conditioning

  • Thorndike’s key contribution: systematic investigation of animal learning and trial-and-error in natural environments.
  • Law of Effect: behaviors followed by satisfying consequences are strengthened; those followed by annoying/punishing consequences are weakened.
  • Experimental setup: puzzle box with a cat; food outside; cat learns to perform actions that lead to escape and reward.
    • Early idea: random trial-and-error exploration gradually yields coordinated, successful behaviors (instrumental learning).
  • Takeaway: consequences (rewarding or punishing) shape the likelihood of future voluntary behaviors.

Skinner: operationalizing operant conditioning

  • B.F. Skinner popularized operant conditioning as a theory describing how organisms learn from consequences.
  • Skinner was a behaviorist: emphasized objective measurement and outwardly observable behavior; internal thoughts/emotions were not the primary focus.
  • Key concept: operant conditioning (OC) is a learning process that changes the probability of a response through manipulation of consequences.
Discriminative stimulus (SD)
  • A cue that signals reinforcement is available if a particular response is made.
  • Examples: text message notification, green traffic light, doorbell, etc.
  • In Skinner’s experiments, a light often served as the discriminative stimulus.
Reinforcement vs punishment (OC)
  • Reinforcement: increases the likelihood of a response.
    • Positive reinforcement: add a desirable stimulus after a response (e.g., giving a treat to reward behavior).
    • Negative reinforcement: remove an aversive stimulus after a response (e.g., stopping loud noise when a desired behavior occurs).
  • Punishment: decreases the likelihood of a response.
    • Positive punishment: add an aversive stimulus to reduce a behavior (e.g., chores, hot sauce for misbehavior).
    • Negative punishment: remove a desirable stimulus to reduce a behavior (e.g., taking away allowance).
  • Core rule: reinforcement (positive or negative) always increases the likelihood of the behavior; punishment (positive or negative) aims to decrease it.
  • Example narrative in lecture: giving chocolates can function as a positive reinforcement to encourage a behavior; using punishment (e.g., extra chores or removal of privileges) to curb behavior.
Primary vs secondary reinforcers
  • Primary reinforcers: naturally reinforcing (biological needs) – food, water, basic needs.
  • Secondary reinforcers: learned value (money, awards, frequent-flyer points) that acquire reinforcement value by association with primary reinforcers.
  • Key idea: secondary reinforcers gain power through learned associations, enabling them to help obtain primary needs.
Shaping and the Skinner box
  • Shaping: gradually reinforcing closer approximations to the target behavior.
  • Skinner box (operant chamber): used with pigeons and rats to illustrate shaping.
    • For pigeons: pecking a disc yields food; for rats: pressing a bar yields food.
    • Process: start with simple behavior and reinforce successive approximations that move toward the desired behavior.
Schedules of reinforcement (partially vs continuous)
  • Continuous reinforcement: reward is given after every instance of the target behavior; fastest way to learn a new behavior.
  • Partial (intermittent) reinforcement: reinforcement occurs on a schedule, not every time; slower to acquire, but more resistant to extinction.
  • Four main schedules of partial reinforcement:
    • Fixed ratio (FR): reward after a fixed number of responses.
    • Variable ratio (VR): reward after an unpredictable number of responses; high and steady response rates (e.g., gambling, lottery).
    • Fixed interval (FI): reward after a fixed amount of time has passed.
    • Variable interval (VI): reward after an average time interval, varying around that average.
  • Real-life examples and notes:
    • Example of continuous reinforcement: attendance sheet rewarded every class session.
    • Gambling often operates on a variable-ratio schedule, contributing to its persistence.
  • Partial reinforcement effect: extinction is slower under partial reinforcement than continuous reinforcement.
    • In pigeons: when continuously reinforced, a pigeon may stop quickly after extinction; when partially reinforced, extinction takes much longer (illustrative numbers from lecture: Nc = 100 vs Np = 1000; Np ≈ 10 Nc).
    • Expression: N<em>p10N</em>cN<em>p \approx 10\,N</em>c indicating greater resistance to extinction under partial reinforcement.
Extinction, punishment, and practical applications
  • Extinction occurs when reinforcement is withdrawn, and the behavior gradually diminishes.
  • Partial reinforcement can make behaviors more resilient to extinction, which has implications for gambling, habits, and long-term training in sports or education.
  • Practical applications: coaches use reinforcement/punishment to shape athletic performance; parents use rewards and consequences to manage behavior; therapists and educators leverage these principles to reduce smoking, curb misbehavior, and encourage study habits.
Cognitive and evolutionary factors in OC
  • Tolman and cognitive aspects: learning can involve cognitive maps and latent learning.
    • Cognitive map: rats in mazes can form mental representations of the layout, enabling efficient navigation later, even without immediate reward.
    • Late learning: learning can occur without immediate reinforcement and reveal itself later when rewards become available.
    • Example analogy: humans learning routes in a familiar town and recalling them when needed.
  • Tolman emphasized that learning is not purely stimulus-response; cognition mediates learning processes.
  • Evolutionary influences: some behaviors resist conditioning due to instinctive drift or natural tendencies that are hard to override with new conditioning.
  • Instinctive drift example: raccoon coin task – raccoons instinctively wash their food and tended to revert to natural behaviors rather than fully learning to place coins into a box; instinctual predispositions can impede conditioned responses.
Learned helplessness and its implications
  • Learned helplessness (L.H.): a state in which individuals believe outcomes are uncontrollable and feel powerless to change their situation.
  • Classic dog experiment: dogs conditioned to associate a sound with an unavoidable shock; when later given a chance to escape the shock by jumping a barrier, many did not attempt it because they had learned that nothing they did would change the outcome.
  • Human implications: can contribute to depression, perceptions of lack of control, homelessness, or chronic stress; demonstrates how exposure to uncontrollable events can affect motivation and future behavior.
  • Relevance to positive psychology and therapy: understanding L.H. informs interventions that restore perceived control and self-efficacy.
Observational learning (Bandura) and the Bobo doll study
  • Observational learning (modeling) involves learning by watching others and the consequences of their actions.
  • Bandura’s work showed that people learn not only from direct reinforcement but also by observing others being rewarded or punished for their behavior.
  • The Bobo doll studies (and replication) demonstrated that children imitate aggression after observing an aggressive model, especially when the model is rewarded for the aggressive behavior.
  • Key idea: learning can occur without direct reinforcement and can be guided by the expectation of rewards for modeled behavior.
Factors that increase imitation (observational learning)
  • Imitation is more likely when the model is perceived as:
    • Warm and nurturing (positive, approachable).
    • Having control or power over the observer (authority figure, parent, teacher, boss).
    • Similar to the observer (age, sex, interests).
    • Perceived to have higher social status or to be a potential role model.
    • In a context where the task is neither too easy nor too hard (moderate difficulty).
    • The observer lacks confidence or is in an unfamiliar/ambiguous situation (seeks cues from others).
  • Additional context: people imitate when the behavior is rewarded or when they anticipate a reward for replicating it; observation in a new or ambiguous situation increases reliance on others’ behavior.
  • Elevator study anecdote: people tended to mimic others in unfamiliar situations (e.g., standing backward in an elevator) even without understanding the origin of the behavior.

Real-world relevance and ethical considerations

  • Education and classrooms: using reinforcement schedules to encourage attendance, participation, and study habits.
  • Parenting and sports coaching: shaping desired behaviors with a mix of reinforcement and appropriate punishment to build long-term habits.
  • Therapy and behavior modification: leveraging reinforcement/punishment to reduce problematic behaviors and encourage healthier alternatives.
  • Public health: applying OC and observational learning to promote smoking cessation, healthy eating, and risk-reduction behaviors.
  • Ethical considerations: avoid abusive or coercive use of punishment; emphasize humane, ethical training and avoid reinforcing maladaptive behaviors.

Quick glossary of key terms

  • Unconditioned stimulus (US): a stimulus that naturally elicits a response without prior learning (e.g., surprise quiz).
  • Unconditioned response (UR): natural, unlearned reaction to the US (e.g., anxiety to a surprise quiz).
  • Conditioned stimulus (CS): previously neutral stimulus that, after association with the US, triggers a response (e.g., cue phrase like "clear your desks").
  • Conditioned response (CR): learned response to the CS (e.g., anxiety to the cue).
  • Discriminative stimulus (SD): cue signaling that reinforcement is available after a behavior.
  • Positive reinforcement: add a desirable stimulus to increase a behavior.
  • Negative reinforcement: remove an aversive stimulus to increase a behavior.
  • Positive punishment: add an aversive stimulus to decrease a behavior.
  • Negative punishment: remove a desirable stimulus to decrease a behavior.
  • Primary reinforcers: naturally reinforcing (food, water).
  • Secondary reinforcers: learned value (money, awards).
  • Shaping: reinforcing successive approximations toward a target behavior.
  • Extinction (OC): when reinforcement is removed, the behavior weakens and may disappear.
  • Schedules of reinforcement: FR, VR, FI, VI (see definitions above).
  • Partial reinforcement extinction effect: behaviors reinforced on partial schedules take longer to extinguish than those reinforced continuously.
  • Latent/late learning (Tolman): learning can occur without immediate reinforcement and reveal itself later.
  • Cognitive map: mental representation of a physical environment.
  • Learned helplessness: belief that outcomes are uncontrollable, leading to passivity.
  • Observational learning (Bandura): learning by watching others and the consequences of their actions.

Summary of formulas and examples

  • Partial reinforcement extinction resistance example: if continuous reinforcement yields Nc = 100 trials to extinction and partial reinforcement yields Np = 1000 trials, then
    • N<em>p10N</em>cN<em>p \approx 10\,N</em>c
  • Schedules overview (labels only):
    • FR-n: after every n responses.
    • VR-n: after an unpredictable (average) n responses.
    • FI-T: after a fixed time interval T.
    • VI-T: after an average time interval T.
  • Illustrative note: reinforcement always increases the probability of the target response; punishment always decreases it (in the context of OC).

Connecting to broader learning science

  • OC complements CC: CC emphasizes antecedents and automatic responses, while OC emphasizes consequences and voluntary actions.
  • Observational learning highlights social and cognitive dimensions: we learn from others through imitation and expectancy of reward, not just direct outcomes.
  • Evolution and instinct: not all learned behaviors are equally malleable; natural tendencies can oppose conditioning (instinctive drift).
  • Real-world applicability: these principles are used to shape behavior across settings (schools, homes, clinics, workplaces) while considering ethical implications and individual differences.