2.3 Notes on Correlation, Causation, Experiments, and Research Validity
Correlation and Correlation Coefficient
- Correlation means there is a relationship between two or more variables, but it does not necessarily imply cause and effect. When two variables are correlated, as one changes, the other tends to change as well, but causality is not established by correlation alone.
- Correlation coefficient (usually denoted by the letter r) measures the strength and direction of a linear relationship between two variables. It is a number from \(-1\) to \(+1\).
- The magnitude of the correlation coefficient indicates strength of the relationship:
- The closer the value is to \(|1|\, the stronger the relationship and the more predictable the changes in one variable given the other.
- The closer the value is to \(|0|\, the weaker the relationship and the less predictable the changes in one variable are.
- Example interpretations:
- A correlation coefficient of \(+0.9\) indicates a much stronger relationship than \(+0.3\).
- A correlation coefficient of \(-0.29\) indicates a weak negative relationship (e.g., the Minnesota study linking sleep duration to GPA).
- Sign of r indicates direction:
- Positive correlation (\(r > 0\)) means the variables move in the same direction (as one increases, the other increases).
- Negative correlation (\(r < 0\)) means the variables move in opposite directions (as one increases, the other decreases).
- No correlation is represented by \(r = 0\).
- Correlation has predictive value: you can use observed correlations to make predictions (e.g., admissions committee using GPA and standardized test scores). However, correlation alone cannot establish causality.
- Common misinterpretations:
- Ice cream sales and crime rise together in warm weather; temperature is a confounding variable that explains both and drives the apparent relationship.
- Correlation does not imply causation; there may be a third variable or coincidence involved.
Correlation Does Not Indicate Causation
- Correlational research reveals strength and direction of relationships but does not identify cause and effect.
- A confounding variable is an outside factor that systematically influences both variables, possibly creating a false impression of a direct relationship.
- Example: Temperature as a confounding variable for ice cream sales and crime rates.
- Another example causing misinterpretation: a cereal-health study suggesting cereal causes better health; other explanations (dietary patterns, lifestyle, socioeconomic factors) could account for observed health outcomes.
- Causality is best established through experiments, which allow manipulation of one variable and control for alternative explanations.
- Illusory correlations (false correlations) occur when people believe a relationship exists despite no reliable evidence. Classic example: the moon’s phases and human behavior. A meta-analysis of ~40 studies found no relation between the moon phase and behavior, despite perceived links (Rotton & Kelly, 1985).
- Why illusory correlations persist:
- Confirmation bias: seeking evidence that supports a hunch while ignoring disconfirming data.
- Availability heuristic: focusing on easily recalled or dramatic examples.
- Illusory correlations can contribute to prejudice and discriminatory attitudes (Fiedler, 2004).
Causality: Conducting Experiments and Using the Data
- To claim cause-and-effect, scientists conduct experiments with precise design and implementation.
- The Experimental Hypothesis: a specific, testable statement about the relationship between variables, derived from observations or prior research.
- Example: Hypothesis that the use of technology in the classroom has negative impacts on learning.
- How hypotheses are formed: from direct observation or prior research, not from limited anecdotes.
- Psychologists use experiments to determine whether changes in one variable cause changes in another.
Designing an Experiment
- Basic design involves two groups:
- Experimental group receives the manipulation (the treatment or variable being tested).
- Control group does not receive the manipulation.
- The only difference between groups should be the experimental manipulation to attribute observed differences to that manipulation.
- Operational definitions:
- Precisely describe what is being measured (e.g., how learning is measured after 45 minutes of instruction).
- Allows others to understand and replicate the study.
- Example setup: 45 minutes of algebra learning (computer program vs. in-person teacher); measure learning via a test.
- Blinding to reduce bias:
- Single-blind study: participants do not know group assignment, but researchers do.
- Double-blind study: neither participants nor researchers know group assignments.
- Experimenter bias: researchers’ expectations can influence the interpretation of ambiguous data; blinding helps control this bias.
- Placebo effect: participants’ expectations can produce perceived or actual improvements.
- In drug trials, the control group may receive a placebo (e.g., a sugar pill) so that expectations do not confound results.
- Placebo control helps ensure that differences are due to the drug, not expectancy or bias (Figure 2.16).
Independent and Dependent Variables
- Independent Variable (IV): the variable that the experimenter deliberately manipulates.
- In the classroom example, the IV is the type of learning (computer program vs. in-person instructor).
- Dependent Variable (DV): the outcome measured to assess the effect of the IV.
- In the classroom example, the DV is learning (e.g., test scores).
- Relationship: changes in the IV are expected to produce changes in the DV.
Selecting and Assigning Experimental Participants
- Participants are the individuals in the experiment.
- Representation and generalizability:
- Using college students is common but may not generalize to the broader population.
- Our example involves high school students; sample should reflect city demographics (income, race/ethnicity, religion, geography).
- Random Sampling: every member of the population has an equal chance of being selected; large samples improve representativeness.
- Population vs. sample: population refers to the whole group of interest; sample is a subset used for the study (e.g., around 200 algebra students from city schools).
- Random Assignment: after sampling, randomly assign participants to experimental or control groups to ensure equivalence, reducing preexisting differences.
- Why random assignment matters: if groups differ before the manipulation, it is unclear whether post-test differences are due to the IV or preexisting differences.
- Practical note: random assignment is often aided by software.
- Link to learning: random sampling and assignment help ensure that observed effects are attributable to the manipulation.
Issues to Consider
- Quasi-experimental designs: when researchers cannot manipulate the IV (e.g., sex) or ethically impossible to assign participants to conditions.
- These designs limit causal claims.
- Ethical constraints: some experiments cannot be conducted (e.g., creating abuse scenarios); researchers must avoid unethical manipulations.
Interpreting Experimental Findings
- After data collection, statistical analysis assesses whether observed differences are likely due to chance.
- Significance threshold: psychologists typically consider results statistically significant if there is less than a 5% chance that the observed differences would occur by chance.
- Expressed as: \(p < 0.05\).
- Strength of experiments: random sampling, random assignment, and a design that minimizes experimenter bias and participant expectancy increase confidence in causal conclusions.
- Example conclusion: if watching a violent TV program leads to more violent behavior than a nonviolent program, we can infer a causal effect of the program on behavior (under the study’s design and limitations).
Reporting Research
- The American Psychological Association (APA) publishes guidelines for writing research papers intended for peer-reviewed journals.
- Peer review: involves anonymous experts assessing the rationale, methods, and ethics; can lead to revisions or rejection to improve quality and reliability.
- Purpose of peer review:
- Quality control and facilitation of replication.
- Ensure clarity and replicability of methods and analyses.
- Replication and reliability: successful replications increase confidence in findings; repeated failures cast doubt and encourage alternative explanations or new directions.
- Replication crisis: concerns about the difficulty of replicating notable studies in psychology and other fields (Shrout & Rodgers, 2018).
- Notable real-world implications: even high-profile scientists can retract findings when replication fails (e.g., a Nobel laureate retracting a paper in 2020).
- Some view the replication crisis as an opportunity to improve scientific practices and transparency (Aschwanden, 2018).
The Vaccine-Autism Myth and Retraction of Published Studies
- Early studies claimed a link between routine childhood vaccines and autism.
- Subsequent large-scale epidemiological research found no link; several original studies were retracted due to issues such as conflicts of interest or methodological problems (Offit, 2008).
- Media coverage of debunked findings led to public hesitancy about vaccination, contributing to outbreaks (e.g., measles outbreaks in 2019; Patel et al., 2019).
- Retractions occur when data are falsified, fabricated, or severely flawed; the retraction process can be initiated by authors, collaborators, institutions, or journals.
- The vaccine-autism case illustrates how premature or flawed research can have lasting public health consequences.
Reliability and Validity
- Reliability: the consistency of a measurement instrument across time, items, or raters.
- Types include:
- Inter-rater reliability: agreement among multiple observers.
- Internal consistency: how well items on a test measure the same construct.
- Test-retest reliability: stability of measurements over time.
- Validity: the extent to which a measurement actually measures what it is intended to measure.
- Types include:
- Ecological validity: generalizability to real-world contexts.
- Construct validity: whether the instrument truly measures the intended construct.
- Face validity: whether the measure appears to assess what it is supposed to.
- Important principle: validity implies reliability, but reliability does not guarantee validity.
- Practical takeaway: researchers strive for measures that are both reliable and valid.
Everyday Connections and SAT/ACT Validity
- Predictive validity: the extent to which a test predicts future outcomes (e.g., SAT scores predicting first-year GPA).
- College Board studies suggest high predictive validity for SAT and first-year GPA (Kobrin et al., 2008).
- Debates about SAT/ACT utility:
- Potential biases that may disadvantage historically marginalized groups (Santelices & Wilson, 2010).
- Some argue predictive validity may be overstated (Rothstein, 2004).
- Debates have led many colleges to de-emphasize or relax SAT/ACT requirements (Rimer, 2008; Strauss, 2019).
Summary of Key Concepts and Their Implications
- Correlation vs. Causation: correlation reveals associations but does not prove causality; experiments are needed to establish causal relationships.
- Confounding variables: external factors that can produce spurious associations; controlling for confounds is essential in experimental design.
- Illusory correlations: false beliefs about relationships; prone to bias and prejudice; rely on systematic misinterpretation rather than data.
- Experimental design basics: random sampling, random assignment, blinding, and placebo controls help isolate causal effects and minimize bias.
- Operationalization: explicitly define how variables will be measured to enable replication and interpretation.
- Ethical considerations: some manipulations are unethical or impractical; quasi-experiments may be employed but limit causal conclusions.
- Reliability vs. validity: reliable measures yield consistent results; valid measures actually measure the intended construct and generalize appropriately.
- Replication and publication: peer review and replication are central to scientific progress; the replication crisis highlights the need for transparent methods and data sharing.
- Real-world impact: flawed studies (e.g., vaccine-autism) can have lasting public health effects; robust research and accurate communication are critical.
- Correlation coefficient range and interpretation: \(r \in \ [-1, 1]\)
- Magnitude interpretation:
- \(|r| \) close to 1 indicates a strong relationship; closer to 0 indicates a weak relationship.
- Sign of correlation:
- Positive: \(r > 0\)
- Negative: \(r < 0\)
- Statistical significance threshold:
- Example value:
- \(r = -0.29\) (weak negative correlation between sleep duration and GPA in the Minnesota study)
- Typical research benchmarks: \(<5\%) chance of observing the result if the null hypothesis were true.
- Notable cited works and data points:
- Minnesota sleep vs. GPA: \(r = -0.29\) (Lowry, Dean, & Manders, 2010)
- Cereal-health advertisements: Anderson, Hanna, Peng, & Kryscio, 2000
- Predictive validity of SAT: Kobrin, Patterson, Shaw, Mattern, & Barbuti, 2008
- SAT/ACT fairness concerns: Santelices & Wilson, 2010; Rothstein, 2004; Rimer, 2008; Strauss, 2019
- Replication and retractions: Shrout & Rodgers, 2018; Offit, 2008; Arnold retraction (2020); Aschwanden, 2018
- Measles outbreaks linked to vaccine skepticism: Patel et al., 2019
Additional Notes and Connections
- The process from observation to hypothesis to experimental testing mirrors foundational scientific principles: ask a question, formulate a testable hypothesis, design a controlled study, collect data, analyze, interpret, and report.
- Real-world relevance: understanding these concepts helps dissect media claims, advertising, and public health messaging that often rely on correlational data or selective reporting.
- Ethical practice: researchers must balance the pursuit of knowledge with participants’ well-being and societal impact, using methods that minimize harm and maximize reliability and validity of findings.