Module 0.4: Correlation and Experimentation — Comprehensive Study Notes
Correlation and Experimentation (Module 0.4) - Comprehensive Study Notes
Key terms and definitions
correlation: a measure of the extent to which two factors vary together, and thus of how well either factor predicts the other.
correlation coefficient, denoted as : a statistical index of the relationship between two variables (from to ).
variable: anything that can vary and is feasible and ethical to measure.
scatterplot: a graphed cluster of dots, each representing the values of two variables; the slope suggests the direction of the relationship, and the amount of scatter indicates the strength of the correlation (little scatter = high correlation).
perfect positive correlation:
perfect negative correlation:
correlation range reminder: correlations can range from (perfect positive) to (perfect negative); real-world data rarely show perfect correlations.
The scatterplot slope and scatter guide interpretation: upward slope = positive relation; downward slope = negative relation; more scatter = weaker relation.
The notion of correlation (abbreviated as r) is central to describing, predicting, and explaining behavioral data.
What is correlation and what does it tell us?
Correlational research is non-experimental and describes relationships between two or more variables.
Experiments are designed to establish a cause-and-effect connection.
Describing behavior via naturalistic observations and surveys often reveals that one trait or behavior tends to coincide with another; this is called a correlation.
A correlation helps us understand how strongly two variables relate and how well one predicts the other (e.g., aptitude-test scores predicting school success).
Scatterplots are useful to visualize the strength and direction of a relationship; each dot represents a pair of values for two variables.
Example context: the strength of the relationship between identical twins' personality test scores, or how well intelligence test scores predict career achievement.
Illustrative figures and interpretation
Figure 0.4-1 (scatterplots) shows a continuum from perfect positive (r = +1.00) to perfect negative (r = -1.00) correlation; in real data, r is typically less than 1 in magnitude.
A positive correlation means as one variable increases, the other tends to increase (e.g., height and weight tend to rise together).
A negative correlation means as one variable increases, the other tends to decrease (e.g., distance from head to ceiling and height; a strongly negative relation is possible but rare in real-world data).
The statement that a correlation is “negative” does not imply poor quality; it only indicates the direction of the relationship.
Example from data: a study with 2,291 Czech and Slovakian volunteers measuring fear and disgust toward 24 animals found a positive correlation between fear and disgust (r = +.72); data presented as a table (Table 0.4-1) and as a scatterplot (Figure 0.4-2) confirm the positive relation with considerable but not perfect scatter.
The example highlights how correlations reveal patterns that may not be obvious from casual observation.
Correlation as a tool for prediction and its limits
Correlation describes relationships but does not prove causation.
Directionality problem: Correlation cannot determine which variable causes the other.
Third-variable problem: A correlation could be due to a third variable that influences both variables of interest.
Example: Teen social media use correlates with teen risk of depression; this pattern may reflect a third variable (e.g., underlying mental health vulnerability) or a bidirectional influence.
When data are summarized (as in Table 0.4-1), correlations can be more apparent than in casual observation, illustrating why statistics can reveal patterns that individuals might miss.
Test Your Understanding and interpretation challenges
Table 0.4-2 (understanding correlation): asks you to categorize reports as positive or negative correlations.
The broader takeaway: correlation does not imply causation; strength and direction do not by themselves establish cause-effect.
Illusory correlations and regression toward the mean (0.4-2)
Illusory correlations: the tendency to perceive a relationship between two events when none exists, often aided by selective memory for confirming instances (e.g., believing certain dreams predict events or that a gambler’s past rolls influence future outcomes).
Illusion of control: the belief that one can influence chance events (e.g., throwing dice with different force to affect outcomes).
Regression toward the mean: extreme results tend to be followed by more average ones due to statistical tendencies, not because a special cause reappears. Examples include test scores, athletic performance, or a team's unusually poor performance improving next game.
Kahneman and Tversky (1974) highlight that regression toward the mean can mislead people into attributing extraordinary outcomes to purposeful actions when they naturally revert toward average.
Practical implication: beware of post hoc explanations for normalizing fluctuations after extreme events; correlation can mislead when interpreted as causation.
Core takeaway: Correlation reveals pattern but does not explain why the pattern exists; regression toward the mean helps explain why extreme results normalize without a causal intervention.
Correlation vs causation: core principle
The point to remember: Correlation does not equal causation.
Correlation may suggest a possible cause-effect relationship but does not prove it.
To establish causation, experimental manipulation and control of confounding variables are required.
Experimental method: isolating cause and effect (0.4-3)
What makes experimentation possible to isolate cause and effect?
Experimental design involves manipulating one or more factors (independent variables) to observe the effect on behavior or mental processes (dependent variables).
Random assignment: Participants are allocated to experimental or control groups by chance to equalize preexisting differences and reduce confounding variables.
Experimental group: receives the treatment (one version of the independent variable).
Control group: does not receive the treatment or receives a contrasting condition.
Random sampling vs random assignment:
Random sampling creates a representative sample of the population for generalizability.
Random assignment equalizes the experimental and control groups to control for confounds.
Operational definitions: Precisely define how variables are manipulated and measured to enable replication.
Confounding variables: Other factors that might influence results; random assignment helps minimize their impact.
Validity and reliability:
Validity: the extent to which the experiment tests what it is supposed to test.
Reliability (replicability): whether findings can be replicated under similar conditions.
Single-blind vs double-blind procedures:
Single-blind: participants are unaware of whether they received the treatment, reducing social desirability bias.
Double-blind: both participants and researchers are unaware of who received the treatment, reducing both placebo effects and experimenter bias.
Placebo effect: improvements due to expectations rather than the treatment itself; controlled for with placebo groups and blind procedures.
Independent and dependent variables; operational definitions; and confounding variables
Independent variable (IV): the factor that is deliberately manipulated.
Dependent variable (DV): the outcome measured.
Confounding variable: an extraneous factor that could affect the DV and confound results.
The basic experimental design requires an IV, a DV, and random assignment to conditions.
Procedure: Define the experimental and control conditions, measure the DV, and ensure procedures are replicable.
Placebo effect and the need for controls in evaluating therapies and interventions
The placebo effect demonstrates the power of expectations on perceived or actual outcomes.
To determine a treatment's true effect, researchers must control for placebo effects and other biases (e.g., demand characteristics).
Historically, some treatments were adopted because early successes were not distinguished from natural recovery or placebo responses.
Classic experimental demonstrations and modern examples
Facebook deactivation study (Allcott et al., 2020):
Design: nearly 1700 participants; random assignment to deactivate Facebook for 4 weeks vs. a control group.
Findings: those in the deactivation group spent more time socializing and watching TV, reported lower depression, greater happiness, and higher life satisfaction; post-study Facebook use decreased.
Conclusion: reduced Facebook time correlated with improved well-being in the experiment, suggesting a potential causal link under controlled conditions.
Rental housing study (Carpusor & Loges, 2006):
Independent variable: perceived ethnicity of the sender’s name in emails (e.g., Patrick McDougall, Said Al-Rahman, Tyrell Jackson).
Dependent variable: percentage of positive replies (invitations to view the apartment).
Reported responses: 89% (Patrick McDougall), 66% (Said Al-Rahman), 56% (Tyrell Jackson).
Demonstrates how manipulation of a social cue (name ethnicity) can influence responses, illustrating potential confounds and the importance of experimental control when interpreting social biases.
Self-testing vs restudy (Benassi et al. example):
Independent variable: study method (active retrieval via self-testing vs restudy by reviewing answers).
Dependent variable: final exam performance.
Result: self-testing produced higher performance (e.g., ~75% correct) than restudy ( ~51% correct) weeks later.
Demonstrates a simple manipulation showing how an experimental variable can affect learning outcomes, with replication and precise operational definitions enabling interpretation.
Practical considerations in experimental design
Validity and reliability are central to sound experiments: tests must measure what they intend to measure and produce consistent results across replications.
Ethical considerations guide what can be manipulated; all procedures must respect feasibility and ethics.
The role of randomization in controlling for unknown confounds and enabling causal inferences.
Quick recap and takeaway points
Correlation describes a relationship and predicts one variable from another but does not establish causation.
Strength of a correlation is indicated by the scatter and the magnitude of , with values closer to indicating stronger relationships.
Illusory correlations and regression toward the mean can mislead judgments about causality and the effects of interventions.
Experiments manipulate an independent variable, measure a dependent variable, and use random assignment to create comparable groups, enabling stronger causal inferences.
Placebo effects and blind procedures are critical controls in experimental research, especially in therapy and pharmacology studies.
Core formulas and definitions to remember (LaTeX)
Correlation coefficient:
Range of :
Perfect positive correlation:
Perfect negative correlation:
Random assignment vs random sampling (conceptual distinction): Random assignment = equalizes groups; random sampling = creates a representative sample for generalization.
Exam-oriented prompts and connections
Why does correlation not prove causation? Discuss directionality and third-variable problems with examples (e.g., social media use and depression).
How do illusory correlations arise, and how does regression toward the mean help explain misinterpretations of success or failure after extreme events?
In evaluating a new therapy, how do single-blind and double-blind designs help mitigate biases and placebo effects?
Describe how the Facebook deactivation experiment demonstrates the use of random assignment and control groups to infer causality.
Compare and contrast random assignment and random sampling with examples.
Note on real-world relevance
Correlation research helps forecast outcomes and identify potential relationships in fields ranging from psychology to education and public health, but it must be followed by well-designed experiments to establish causality.
Understanding bias, placebo effects, and experimental controls is essential for evaluating claims in media reports and scientific studies.
Connections to foundational principles
Statistics illuminate patterns that casual observation may miss, reinforcing the need for rigorous methodology in psychology.
The interplay between correlation and causation reflects core scientific reasoning: observation → hypothesis → testing with controlled manipulation.
Summary takeaway
Always consider directionality, potential third variables, and the difference between correlation and causation when interpreting data.
Use experimentation to isolate causal effects, leveraging random assignment, control groups, and blind procedures to minimize bias and confounding influences.