Critical Evaluation and Replicability in Psychological Research
2.7 Critical Evaluation of Research Studies
To become an informed consumer of psychological research, one must adopt the principle of caveat emptor (let the buyer beware). Media reports of "University studies indicate…" should not be taken at face value without examining seven critical questions.
1. The Theoretical Framework
Logic and Flow: Does the specific hypothesis flow logically from a broader theory?
Consistency of Definitions: Are terms defined logically and consistently throughout the study?
Example: In a study on social class and intelligence (), the researcher must explain why a relationship between the two should exist and ensure those terms do not change meaning midway.
2. Sample Adequacy and Appropriateness
Representativeness: Does the sample represent the population of interest? If the goal is to generalize to the wider community, a sample consisting only of undergraduates is inappropriate.
Sample Size (): The sample must be large enough to distinguish between meaningful results and accidental chance.
Probability Example: Rolling a pair of dice six times and getting "snake eyes" twice is not enough to conclude the dice are loaded; such a result can easily happen by chance.
3. Measures and Procedures
Validity: Do the measures assess what they were intended to assess?
Controls: Were proper control groups used to rule out alternative explanations?
Confounding Variables: Did investigators control for variables like interviewer gender, which might affect participant responses?
4. Conclusiveness of Data
Result Presentation: Data should be carefully examined in the "Results" section, often found in graphs, charts, or tables.
Alternative Interpretations: The reader should ask if alternative interpretations explain the results better than the researcher's explanation. Findings may fit patterns the researcher rejected or ignored.
5. Warranted Broader Conclusions
Correlation vs. Causation: Researchers must not claim causation where only correlation exists.
Aggression Example: Finding that children who watch violent TV are more aggressive does not mean TV causes aggression. Alternatives include:
Aggressive children prefer violent TV.
Violent TV only triggers violence in children already predisposed to it.
6. Meaningfulness ("The So What? Test")
Novelty and Utility: Does the study provide new knowledge or lead to future research?
Importance: Significant studies often produce surprising findings or help choose between opposing theories ().
7. Ethics
Humanity: Are human or animal participants treated humanely?
Ends vs. Means: Does the incremental knowledge produced justify the methods used?
Regulation: The Australian Psychological Society (APS) publishes guidelines (). Universities use Ethics Committees and Institutional Review Boards (IRBs) to review proposals. Approval is required even for benign studies on memory or math.
The Crisis of Replicability
Replicability is a key attribute of empirical research; scientific legitimacy depends on getting the same results when a study is repeated. However, social sciences currently face a "crisis of replicability."
The Reproducibility Project ()
Scope: researchers attempted to replicate studies from three major psychology journals published in .
Findings:
Original studies showing statistically significant results: .
Replications showing statistically significant results: .
Perspectives on the Crisis
Premature Conclusion: argue failure to replicate may be due to different time points or populations rather than flawed original theories.
Sample Size Issues: suggest small-to-medium sample sizes can misrepresent data to suggest larger effects than exist.
Statistical Power: attribute low replicability to the recurrent use of low statistical power.
Publication Bias: Journals often prefer new, surprising results over less exciting replications. This incentivizes questionable practices like "cherry-picking" data or selective deletion of results ().
Outright Dishonesty: Researcher Diederik Stapel () admitted to fabricating data across more than publications.
Solutions for Replicability
Open Research: Making all data, results, and protocols available for transparency ().
Pre-registration: Outlining methodology and analytical strategies before data collection to prevent selective data use.
Reward Structures: argues we must shift focus from "groundbreaking" headlines to reward confirmatory work and community data sharing.
Principles of Critical Thinking
Critical thinking is the logical and rational assessment of information, examining both strengths and weaknesses (). It is supported by three key principles:
Scepticism: Always questioning assumptions and assertions, regardless of whether they are in print or from an authority figure.
Objectivity: Taking an impartial approach based on evidence and logic, setting aside personal feelings or biases.
Open-mindedness: Willingness to consider all sides and potential explanations, even if they contradict personal experience.
Fallacies in Arguments
Arguments that do not flow logically from evidence are fallacious. Four common types include:
Straw Man: Deliberately attacking a weak, decoy version of an opposing argument to make one's own look stronger.
Appeals to Popularity: Assuming an argument is true because it is widely believed (e.g., historical belief in a flat earth).
Appeals to Authority: Assuming an argument is true because a well-known or prestigious person said it.
Arguments Directed to the Person (Ad Hominem): Attacking the character or failings of the author of an alternative argument rather than the argument itself.
Popular Myths in Psychology
Misinformation is common due to mass media, self-help books, and internet hoaxes.
The Role of Education: found that while psychology students recognize myths better than the general public, the effect size is small, highlighting the need for deep processing.
Fake News and Emotion: found people are more susceptible to fake news when angry. Anxious people are more open-minded. Fact-checking information alongside fake news helps minimize belief in myths.
Common Examples of Myths
Multiple Choice Tests: The myth to "stick with your gut feeling" is refuted by evidence; you should change your answer if you have a good reason ().
Attraction: The "playing hard to get" myth is contradicted by research showing men prefer women who are open to advances.
The Spinning Ballerina (Figure ): Used to claim people are "left-brained" or "right-brained." Reality shows the brain works in an integrated fashion, though hemispheres have different specialized processes.
The Scientific Attitude and Process
The Criterion of Science
argued that the hallmark of science is the formulation of testable hypotheses that can be refuted (falsified). Experimental methods provide the strongest tests for these hypotheses.
Discovery vs. Justification
Context of Discovery: The stage where phenomena are observed and theories are built. Descriptive methods (case studies, naturalistic observation, surveys) are most useful here as they allow for unconstrained behavior.
Context of Justification: The stage where hypotheses are empirically tested. Experimental, quasi-experimental, and correlational designs are preferred. Inferential statistics help determine if findings are genuine.
Conclusion
Science is a mental map of phenomena. Optimal psychological science uses multiple methods and measures, gradually weeding out beliefs that do not withstand scientific scrutiny as technology advances.