1/23
Accurate Reporting of Scientific Studies
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Decline effect
Effect sizes shrink systematically in follow-up studies
P-Hacking is:
Collecting data, or conducting statistical analyses, until a non-significant result becomes significant
Driven by the pressure to produce positive findings rather than by the data or the research question
How p-hacking occurs:
Stopping data collection to early
Ending data collection the moment p < 0.05 is reached, before pre-specified sample size
Collecting data until the p-value is significant
Conducting multiple experiments; reporting only the one that worked.
Cherry-picking outcomes
Measuring many variables but only reporting those that reach significance.
Tweaking the data
Post-hoc decisions on outlier removal or data transformation to achieve significance
Doing multiple comparisons without corrections
Performing many tests without correcting for family-wise error rate
Type I errors
Untrue significant results (false positive)

Type II errors
Untrue non-significant results (false negative)

HARKing is:
When a hypothesis is invented after the data collection and results, and presented as if it was pre-formulated
Makes a chance finding look like a prior prediction
You’re pretending to have tested for something that was just a chance finding
Why is HARKing problematic?
Inflates false-positive
p = 0.05 means there is a 5% probability that your results occurred by pure random chance. If you run 20 independent tests at α = 0.05 and report the best one as a "prediction," your actual false positive rate is closer to 1 − 0.95²⁰ ≈ 64%, not 5%
Results cannot be replicated
Because results are post-hoc justified they are unlikely to replicate in a study with new data. Entire research fields can be based on non-replicable results. Researchers will spend time and money trying to replicate findings that cannot be replicated.
Invisible in publications
Readers cannot detect HARKing from the published paper alone. Peer review cannot catch HARKing without access to study logs
Publication bias
Positive findings are far more likely to be reported than null or negative results - regardless of scientific value
File-drawer effect
Direct consequence of publication bias
Null results are not written up, submitted or published
How can you detect publication bias?*
Using a funnel plot
Types of replication
Closed replication
Conceptual replication
Closed replication*
rare
Conceptual replication*
How can we improve the reporting in academia*
Pre-registration of hypotheses