Chapter 1 Study Guide: The Quest for Causality
The Basis of Knowledge and Evidence
- Evidence as the Modern Standard: To convince others or ourselves, we must provide verifiable information. This differs from a "hunch" or innate knowledge, which lacks the objective basis required for the scientific process.
- Observing Causality: In simple cases, cause and effect are directly observable (e.g., a candle tipping over and starting a fire).
- Complexity in Social Science: Investigating events such as the 2008 U.S. Presidential election, crime reduction in the 1990s, or economic resilience during recessions involves multiple possible causes (e.g., lighting, faulty wires, arsonists), making direct causality harder to trace.
- The Role of Data: When direct observation is impossible, data allows researchers to compare outcomes. For instance, in an earthquake, researchers look for differences between collapsed and standing buildings (e.g., material, height, design, age, location).
- Caveats in Data Analysis: Observing a higher collapse rate in older buildings does not definitively prove age is the cause; other variables like design era or seismic intensity in specific neighborhoods could be the actual drivers.
- Correlation vs. Causation: The foundational principle of the text is that correlation does not imply causation. The quest of the book is to define what does imply causation, with the short answer being the identification of exogenous variation.
The Core Statistical Model
- Dependent Variable (Y): This is the outcome of interest. Its value "depends" on the changes in the independent variable.
- Independent Variable (X): This is the presumed cause of change. It is called independent because it operates outside the control of the dependent variable.
- Linear Formalization: The relationship is formalized through the equation for a line (Y=mX+b), adapted for social science as:
Yi=β0+β1Xi+ϵi
Components of the Linear Model
- The Observation (i): Subscripts represent specific people or points in a dataset (e.g., Xi or Yi is the value for person i).
- The Intercept (β0): Also called the constant. It indicates the expected value of Y when X=0. In research, this is necessary for placing the regression line but is rarely the primary focus of interest.
- The Slope Coefficient (β1): This parameter characterizes the relationship between X and Y. It indicates how much the dependent variable (Y) is expected to change for each one-unit increase in the independent variable (X).
- The Error Term (ϵi): The Greek letter epsilon represents the "wiggle room" or the difference between the actual observed value and the value predicted by the model. It captures every other factor affecting weight that is not included in the model, such as sex, height, genetics, and exercise.
Springfield Case Study: Donuts and Weight
- The Model: Weighti=β0+β1Donutsi+ϵi
- Specific Data Points (Springfield, U.S.A.):
- Homer (1): 14 donuts/week, 275pounds
- Marge (2): 0 donuts/week, 141pounds
- Lisa (3): 0 donuts/week, 70pounds
- Bart (4): 5 donuts/week, 75pounds
- Comic Book Guy (5): 20 donuts/week, 310pounds
- Mr. Burns (6): 0.75 donuts/week, 80pounds
- Smithers (7): 0.25 donuts/week, 160pounds
- Chief Wiggum (8): 16 donuts/week, 263pounds
- Principal Skinner (9): 3 donuts/week, 205pounds
- Rev. Lovejoy (10): 2 donuts/week, 185pounds
- Ned Flanders (11): 0.8 donuts/week, 170pounds
- Patty (12): 5 donuts/week, 155pounds
- Selma (13): 4 donuts/week, 145pounds
- Estimated Parameters:
- Intercept (β0≈122): The average weight for those eating zero donuts.
- Slope (β1≈9.1): Weight increases by approximately 9.1pounds for every donut eaten per week.
Fundamental Challenges: Randomness and Endogeneity
- Challenge 1: Randomness: Relationships observed in data may be the result of mere coincidence or "luck of the draw." Statistical analysis aims to distinguish results that happen by chance from those that reflect a real causal mechanism. We can never fully escape the possibility of randomness, but we can measure our confidence levels.
- Challenge 2: Endogeneity: This occurs when the independent variable (X) is correlated with factors contained in the error term (ϵ).
- Etymology: "Endo" means internal; the variable is internal to the systems captured by the error term.
- Implication: If an independent variable is endogenous, we risk wrongly attributing the effects of unmeasured variables to X. For example, if tall people eat more donuts, we might attribute the weight gain caused by height to donut consumption.
- Exogeneity: This is the opposite of endogeneity. A variable is exogenous if changes in it are unrelated to factors in the error term. This provides a "clean" view of the relationship between X and Y.
Understanding Correlation
- Definition: Two variables are correlated if they move together. Correlation ranges from −1 to 1.
- Positive Correlation: High values of X associate with high values of Y.
- Negative Correlation: High values of X associate with low values of Y.
- Zero Correlation: No linear relationship exists between the variables.
- Endogeneity via Correlation: If X is correlated with the error term (positively or negatively), endogeneity is present. If there is no correlation between X and the error term, exogeneity is present.
Detailed Case Study: Flu Shots and Mortality
- Model: Deathi=β0+β1Flushoti+ϵi
- Empirical Observation: Studies show people getting flu shots are up to 50% less likely to die.
- Endogeneity Problem: Health status is in the error term. Healthier people are more likely to seek out flu shots. Therefore, the lower death rate might result from overall health, not the shot itself.
- Evidence of Endogeneity:
- People with flu shots have a 60% lower death rate in the summer when the flu is not active.
- Historical instances of vaccine production failures or misaligned vaccine strains showed no significant change in mortality rates.
Detailed Case Study: Country Music and Suicide
- Model: Suicideratesi=β0+β1Countrymusici+ϵi
- Initial Finding: Higher suicide rates exist in metropolitan areas with more country music airtime.
- Error Term Factors: Alcohol use, drug use, gun availability, divorce rates, and poverty.
- Endogeneity Mechanism: If country music is more popular in regions with high divorce rates or high gun ownership, the correlation between music and suicide is spurious. Studies that controlled for guns and divorce found no causal relationship.
Randomized Experiments: The Gold Standard
- Function: Experiments create exogenous variation through randomization. By randomly assigning subjects to the treatment group or the control group, researchers ensure the independent variable is uncorrelated with any factors in the error term.
- The Treatment: The policy intervention or independent variable manipulated by the researcher.
- Why "Gold Standard"?: It rules out the systematic differences (e.g., athleticism, height, consciousness) between groups, allowing for a pure causal inference.
Limitations of Experiments
- Feasibility: Some experiments are too expensive or physically impossible (e.g., randomly assigning corruption or birth rates).
- Ethics: Denying potentially lifesaving treatments (like flu shots) to a control group to observe mortality is often considered unethical.
- Generalizability: Results from a specific time, place, and population may not apply elsewhere.
- Validity Types:
- Internal Validity: The confidence that the inference is unbiased within the specific experiment (i.e., X caused Y in this study).
- External Validity: The extent to which the results apply to other contexts or broader populations.
Observational Studies
- Definition: Research using data generated by non-experimental processes where the researcher does not control the variables.
- The Strategy: Since researchers cannot create exogeneity through randomization, they use statistical techniques (like multivariate regression, discussed in later chapters) to winnow down the factors in the error term and approximate exogeneity.