Open Science and the Replication Crisis
The Replication Crisis
Research Replicability
Scientific findings should be reliable and replicable to ensure they are not due to chance. Replication is critical for scientific progress.
A survey by Nature (Baker, 2016) found that:
- Over 70% of researchers have failed to reproduce another scientist's experiments.
- More than half have failed to reproduce their own experiments.
Replication in Psychology Research
Replication is particularly important in psychology due to the variability of human behavior.
Types of replication:
- Exact replication (direct replication): Using the same methods to see if the results are consistent.
- Conceptual replication: Testing the same idea with different methods and measures.
Neglect of Replication Before 2010s
Replication studies were rare before the 2010s because they were seen as lacking prestige and originality (Neuliep & Crandall, 1993).
Makel et al. (2012) found a replication rate of only 1.07% in high-impact psychology journals since the 1990s.
Catalyst for Change: Bem (2011)
Daryl Bem's (2011) paper on precognition, published in the Journal of Personality and Social Psychology, sparked debate about methodological and statistical practices.
- Bem D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100, 407–425. https://doi.org/10.1037/a0021524
Failed replication attempts (e.g., Ritchie et al. 2012) questioned the reliability of published findings and fueled the replication movement.
Open Science Collaboration (OSC) (2015)
A group of psychology researchers conducted 100 replication studies.
- The mean effect size (r) of the replication effects (, ) was half the magnitude of the original effects.
- Only 36% of the replications had statistically significant results compared to 97% of the original studies.
- Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://www.science.org/doi/10.1126/science.aad7243
This indicated low replication rates.
Alternative Interpretations of OSC (2015)
Gilbert, et al. (2015) argued that OSC (2015) underestimated reproducibility, citing issues like different populations and methodological infidelities.
The Many Labs project found that 54% of replications of 28 published findings showed significant evidence in the same direction, but with a substantial decline in effect sizes (Klein et al., 2018).
A review of 307 replications showed 64% with statistically significant evidence in the same direction, with effect sizes 68% as large as the original studies (Nosek et al., 2022).
Example: Elderly Priming Effect
Bargh et al. (1996) found that priming participants with elderly-related concepts caused them to walk more slowly.
Doyen et al. (2012) failed to replicate this effect.
Non-replication doesn't necessarily mean the original study was wrong.
Reasons for Non-Replication
- Chance alone (random or sampling errors)
- No replication is truly 'exact'
- Original study reported with low transparency
- Replicator errors
- Original findings are false positives (Ioannidis, 2005)
- Small sample sizes, insufficient statistical power, poor research design and data analysis
- Potentially falsified results & questionable research practices
- The replication could be a false negative.
Two Types of Errors
- False Positive: The study found a statistically significant effect when there is actually no real effect in the population.
- False Negative: The study did not find a statistically significant effect when there is actually a real effect in the population.
Research Misconduct
- Fabrication: Making up data.
- Falsification: Manipulating or omitting data.
- Plagiarism: Using others' work without acknowledgement.
- Intentionally ignoring or not acknowledging contributions by other authors
- Wrongfully passing oneself off as (co-)author
- Knowingly and willingly using incorrect methods or misinterpreting results
Diederick Stapel
A prominent psychologist, Diederick Stapel, was found to have fabricated data in dozens of studies.
Dan Ariely & Francessca Gino
Dan Ariely and Francesca Gino have both been accused of fabricating data.
Karl Popper's Hypothetico-Deductive Method
- Generate and specify hypothesis
- Design study
- Conduct study and collect data
- Analyse data and test hypothesis
- Interpret results
- Publish and/or conduct next experiment
Researcher Degrees of Freedom
Researchers make many decisions during data collection and analysis:
- When to stop collecting data
- Whether to exclude participants
- What analyses to do
- Which experimental conditions to combine
- What DV to focus on
- What control variables to include
These decisions can lead to Questionable Research Practices (QRPs) (Simmons et al., 2011).
Questionable Research Practices (QRPs)
- Running extra participants
- Excluding participants
- Running statistical tests until significant results arise (p-hacking)
- Creating hypotheses after looking at the data (HARKing)
- Selectively reporting studies that 'worked'
- Incomplete reporting of study design
- Failing to report contrary evidence
P-Hacking
P-hacking involves analyzing data in different ways until a significant result is found.
It's acceptable to run analyses in different ways, as long as it is disclosed.
Why P-Hacking is a Problem?
P-hacking violates the rules of Null Hypothesis Statistical Testing (NHST).
When there is no effect, p-values vary randomly between 0 and 1. With enough attempts, a p-value below 0.05 will eventually be found.
Ways to P-Hack
- Stop collecting data once p < .05
- Analyse many measures, but report only those with p < .05
- Collect and analyse many conditions/groups, but only report those with p < .05
- Use covariates to get p < .05
- Exclude participants to get p < .05
- Transform the data to get p < .05
HARKing - Hypothesizing After the Results are Known
Reporting a hypothesis as if it was predicted when it was not.
Prevalence of QRPs
John et al. (2012) found high rates of psychologists admitting to engaging in QRPs.
Publication Bias
Journals prefer novel studies with statistically significant results.
- File-drawer Problem: Studies with null or negative results are not published (Smaldino & McElreath, 2016).
Flawed Incentives
The "Publish or Perish" culture incentivizes researchers to prioritize publication success over methodological rigor.
What are we doing wrong?
There are a lot of false positive findings in scientific literature because it is very easy to get p < .05, even when there is actually no effect.
We can’t tell if a published positive result was p-hacked or HARKed.
Not enough transparency & Not enough verification/correction. This leads to replication crisis.
Crisis to Revolution: Open Science
Open Science aims for:
- Transparent
- Reproducible
- Accessible
- Open Data
- Open Source
- Open Notebooks
- Open Access
- Citizen Science
- Open Peer Review
- Scientific social networks
- Open educational resources
Transparency
Make your work publicly available by showing:
- Methods
- Data
- Analysis code
- All analyses
- All studies
- What you planned (pre-registration)
So we can replicate your study, reanalyze your data, see what didn’t work and tell what you predicted ahead of time.
Pre-Registration
Researchers describe their hypotheses, methods, and analyses before conducting research (Van t’Veer & Giner-Sorolla, 2016).
The elements of a Pre-registration:
- The hypotheses we plan to test
- The method we plan to use
- planned sample, design, manipulations, measures, etc.
- The analysis plan
- Contingencies and assumptions E.g., missing data, scale reliability
Where to share your pre-registration
Online pre-registration platform
- Open Science Framework (OSF, https://osf.io/)
- AsPredicted (https://aspredicted.org/)
You have the option of either making it public immediately or making it private for up to four years through an embargo.
Benefits of Pre-Registration
- Prioritizing theory and method
- Reducing questionable research practices
- Reducing reporting bias
- Distinguishing confirmatory from exploratory research
- Enhancing transparency of research process (hence credibility)
- Reducing publication bias
Scepticism
- Fear that preregistration will stifle discovery. Science isn’t just about testing hypotheses — it’s also about discovering hypotheses grounded in phenomena that are worthy of study. Aren’t we supposed to let the data guide us in our exploration? How can we make new discoveries if our studies need to be catalogued before they are run?.
Limitations of Pre-Registration
- Too restrictive and limits exploration
- More work
- Can’t stop outright cheating
- Difficult to make all decisions ahead
- Less useful for some types of research (e.g., qualitative)
- Does not guarantee quality
Transparency is Important But Not Enough
Transparency is necessary to evaluate quality, but it is not a guarantee of quality.
It should not, by itself, be taken as a marker of credibility.
What does this mean for psychology?
- Science is harder than we thought
- Confidence should be rare, doubt should be the norm
- This problem is not unique to psychology
- We’re doing something about it