Comprehensive Study Guide: Choice, Preference, and Reinforcement Dynamics
Calculation of Proportions in Choice and Preference
- Formula for Relative Response Rate: To determine the proportional rate of response to a specific alternative (e.g., Key A), the formula used is: Ba/(Ba+Bb).
* Specific Example (VI 4.5 min vs VI 2.25 min):
* Response rate on Key A (Ba): 1750pecks/hour.
* Response rate on Key B (Bb): 3900pecks/hour.
* Calculation: 1750/(1750+3900)=0.31.
- Formula for Relative Reinforcement Rate: The proportional rate of reinforcement for an alternative is calculated as: Ra/(Ra+Rb).
* Specific Example:
* Reinforcement rate on Key A (Ra): 13.3reinforcers/hour.
* Reinforcement rate on Key B (Rb): 26.7reinforcers/hour.
* Calculation: 13.3/(13.3+26.7)=0.33.
- Observation: The relative response rate (0.31) closely approximates the relative reinforcement rate (0.33).
The Matching Equation and Herrnstein’s Findings
- Key Variables: Herrnstein identified the major dependent variable as the relative rate of response and the primary independent variable as the relative rate of reinforcement.
- Equation 9.1 (Proportional Response Rates): Ba+BbBa=Ra+RbRa.
- Verbal Statement: The relative rate of response matches (equals) the relative rate of reinforcement.
- Ideal vs. Actual Data: Equation 9.1 represents an ideal version of choice behavior; empirical data from subjects like Pigeon 231 approximate this matching relationship.
Extensions of the Matching Law: Time and Multiple Alternatives
- Matching Time (Equation 9.2): For continuous behaviors (e.g., talking, standing, foraging), choice is measured as time spent (T). The formula is: Ta+TbTa=Ra+RbRa.
* Significance: Proposed by Baum (2015) as a more fundamental measure of choice than discrete response counts.
- More Than Two Alternatives (Equation 9.3): Matching applies when an organism chooses among multiple sources (Ba,Bb,…,Bn): Ba+Bb+⋯+BnBa=Ra+Rb+⋯+RnRa.
- Generality Across Species: Demonstrated in pigeons (Davison & Ferguson 1978), wagtails (Houston 1986), cows (Matthews & Temple 1979), and rats (Poling 1978).
Human Communication and Social Interaction
- Conger & Killeen (1974): Humans in a group discussion on drug abuse. Relative time spent talking to specific listeners matched the relative rate of reinforcement (agreement) given by those listeners.
- Borrero et al. (2007): Discussion of juvenile delinquency; found that relative response rates were better described by the generalized matching law than relative time spent talking.
- McDowell & Caron (2010): Analyzed boys at risk for delinquency. Verbal behavior was coded as "rule-break talk" or "normative talk." Findings showed a bias toward normative talk and extreme deviations from matching as the risk for delinquency increased.
Practical Classroom Implications
- Maintenance of Behavior: Desirable (assignments) and undesirable (screaming, throwing paper) behaviors in classrooms are maintained by schedules of social reinforcement (attention, approval).
- Interval vs. Ratio Schedules in Intervention: Myerson and Hale (1984) argue that interval (VI) schedules are more successful than ratio (VR) schedules for behavior modification.
* Exclusive Preference: On concurrent ratio schedules, organisms develop exclusive preference for the higher rate alternative which can lead to intervention failure if the teacher's reinforcement rate isn't high enough.
* Comparison: A VI schedule of reinforcement for competing responses that is twice as rich as the schedule for inappropriate behavior is as effective as a VR schedule three times as rich.
Quantitative Law of Effect and Single-Operant Schedules
- Hyperbolic Curve Theory: Absolute response rate on a single schedule is a hyperbolic function of the reinforcement rate relative to the total reinforcement (scheduled + extraneous).
- Extraneous Reinforcement (Re): Unknown contingencies (scratching, sniffing, distractions) that slow the rise in response rate for the target behavior.
- Catania and Reynolds (1968): Exhaustive study of six pigeons on single VI schedules with rates ranging from 8 to 300reinforcements/hour. Statistical fits showed response rates are a hyperbolic function of reinforcement rate.
- Clinical Application (McDowell 1981): Case study of a 10-year-old boy with severe self-injurious scratching. Research identified reprimands as positive reinforcement. Matching equations accounted for >99% of the variation in scratching behavior based on reprimand rates.
Optimal Foraging, Melioration, and Choice Preference
- Maximization vs. Melioration:
* Optimal Foraging (Maximization): Organisms stabilize on a distribution that maximizes overall reinforcement.
* Melioration (Herrnstein 1982): Organisms are sensitive to momentary fluctuations; they stay on one schedule until the local rate of reinforcement drops below a second schedule.
- Preference for Choice: Animals and humans prefer alternatives that offer choice even when reinforcement rates are equal (Catania 1975).
* Developmental Data: Five out of six preschool children preferred choosing among candies (Tiger, Hanley, & Hernandez 2006).
* Brain Imaging: University students showed activity in the ventral striatum when cues signaled an upcoming choice (Leotti & Delgado 2011).
Behavioral Economics and Addiction
- Elasticity of Demand:
* Elastic: Consumption decreases significantly as price (response requirement) increases (e.g., luxury items).
* Inelastic: Consumption stays relatively stable despite price increases (e.g., groceries, addictive drugs).
- Reinforcement Substitutability:
* Substitutes: Increasing the price of one increases the consumption of another (Butter/Margarine).
* Independents: Changing the price of one has no effect on the other (Gasoline/Theater tickets).
* Complements: Increasing the price of one decreases consumption of both (Hot dogs/Buns).
- Addiction and Methadone: Methadone is considered a partial substitute for heroin, providing some reinforcing effects but typically lacks the full social context of heroin use.
- Activity Anorexia: Characterized by decreased food intake and increased wheel running in rats due to food restriction. Food and physical activity function as economic substitutes in energy-balance processes (Belke, Pierce, & Duncan 2006).
Delay Discounting of Reinforcement Value
- Devaluation: Reinforcement value decreases as the delay to receiving it increases.
- Hyperbolic Discounting Equation (9.4): Vd=1+kdA.
* Vd: Discounted value.
* A: Initial amount.
* d: Delay.
* k: Discounting rate (higher k = more impulsive).
- Populations with Higher Discounting Rates: Cigarette smokers, problem drinkers, heroin users, and pathological gamblers show steeper discounting curves compared to control groups.
- Neurobiology: Rats with lesions to the nucleus accumbens (NAc) show higher rates of discounting for large, delayed reinforcers (Bezzina et al. 2007).
Self-Control and the Ainslie–Rachlin Principle
- Principle Statement: Reinforcement value decreases hyperbolically as the delay between choice and reward increases.
- Preference Reversal: At a long delay, a larger later reward (LLR) is preferred; as the time for the smaller sooner reward (SSR) approaches, its value surpasses the LLR, leading to impulsive choice.
- Commitment Response: A behavior emitted prior to a choice point that eliminates or reduces the probability of impulsive behavior (e.g., inviting a study buddy to ensure studying occurs instead of partying).
- Pigeon Research (Green et al. 1981): Pigeons preferred 2s grain over 6s grain with short delays but reversed preference to the 6s option when an additional 18s delay was added to both.
Advanced Section: The Generalized Matching Law
- The Power Law (Equation 9.5): BbBa=k(RbRa)a.
* Bias (k): Systematic preference for one alternative caused by factors like stimulus control, effort, or history (e.g., a pigeon preferring a yellow key due to a tiny speck on it).
* Sensitivity (a): The degree to which response ratios change with reinforcement ratios.
* Undermatching (a<1): Resulting from poor discrimination; the most common outcome (averagea=0.80).
* Overmatching (a>1): Relative behavior increases faster than reinforcement; less common.
- Log-Linear Form (Equation 9.6): log(BbBa)=log(k)+a×log(RbRa).
* In a plot, the slope equals a and the intercept equals log(k).
- Preference Pulse (Davison & Baum 2000): Rapid shifts in preference following a single delivery of reinforcement, suggesting molecular dynamics underlie molar matching results.
Conditioned Reinforcement Basics
- Definition: A stimulus or event that increases or maintains an operant rate due to a history of conditioning with another reinforcer.
- Magazine Training: Deliberately pairing a feeder sound with food to establish the sound as a conditioned reinforcer.
- New-Response Method: Testing if a previously neutral stimulus (e.g., click) can condition a brand new behavior (e.g., pressing a spot on the wall).
- Clicker Training (Karen Pryor): Using a hand-held clicker followed by food. Conditioned reinforcers lose meaning if not systematically paired with backup reinforcers (Extinction).
Chain and Tandem Schedules
- Chain Schedule: Two or more simple schedules presented sequentially, each signaled by a unique discriminative stimulus (SD). Reinforcement only occurs in the final link.
* Stimuli in a chain have multiple functions: SD for the next link and Scondr for the previous behavior.
- Tandem Schedule: Sequential schedules without unique discriminative stimuli (unsignaled chain).
- Chain Types:
* Homogeneous: Topography of response is identical in each link (e.g., key pecking).
* Heterogeneous: Different responses for each link (e.g., going to a restaurant: booking, dressing, driving, eating).
- Backward Chaining: Training begins with the final link and moves toward the beginning (e.g., teaching golf by starting with short putts and working back to the tee shot).
- Observing Response (Wyckoff 1952): A topographical operant that converts a mixed schedule into a multiple schedule by producing stimuli correlated with reinforcement (SD) or extinction (SΔ).
- Good News vs. Bad News: Pigeons and humans prefer "Good News" (stimuli correlated with reinforcement) but typically do not prefer "Bad News" (stimuli correlated with extinction) unless it allows more efficient behavior (Fantino & Case 1983).
- Delay-Reduction Hypothesis: Conditioned reinforcers consist of stimuli that signal a reduction in time to positive reinforcement or an increase in time from an aversive event.
- Equation 10.1 (Concurrent-Chain Choice): RL+RRRL=(T−t2L)+(T−t2R)T−t2L.
* T: Average time to reinforcement from the start of the initial links.
* t2L,t2R: Delay in the terminal links.
Token Economies and Generalized Reinforcers
- Generalized Conditioned Reinforcer: Exchangeable for many sources of reinforcement (e.g., money, social approval). Independence from specific momentary deprivation.
- Token Schedules: Include three components:
1. Token-production schedule.
2. Exchange-production schedule.
3. Token-exchange schedule.
- Schaefer & Martin (1966): Psychiatric patients in a token economy for "apathetic" behavior. Contingent tokens for hygiene and socialization reduced return rates after discharge to 14% (vs. 28% control).