Critical Evaluation and Replicability in Psychological Research

2.7 Critical Evaluation of Research Studies

To become an informed consumer of psychological research, one must adopt the principle of caveat emptor (let the buyer beware). Media reports of "University studies indicate…" should not be taken at face value without examining seven critical questions.

1. The Theoretical Framework
  • Logic and Flow: Does the specific hypothesis flow logically from a broader theory?

  • Consistency of Definitions: Are terms defined logically and consistently throughout the study?

  • Example: In a study on social class and intelligence (IQIQ), the researcher must explain why a relationship between the two should exist and ensure those terms do not change meaning midway.

2. Sample Adequacy and Appropriateness
  • Representativeness: Does the sample represent the population of interest? If the goal is to generalize to the wider community, a sample consisting only of undergraduates is inappropriate.

  • Sample Size (NN): The sample must be large enough to distinguish between meaningful results and accidental chance.

  • Probability Example: Rolling a pair of dice six times and getting "snake eyes" twice is not enough to conclude the dice are loaded; such a result can easily happen by chance.

3. Measures and Procedures
  • Validity: Do the measures assess what they were intended to assess?

  • Controls: Were proper control groups used to rule out alternative explanations?

  • Confounding Variables: Did investigators control for variables like interviewer gender, which might affect participant responses?

4. Conclusiveness of Data
  • Result Presentation: Data should be carefully examined in the "Results" section, often found in graphs, charts, or tables.

  • Alternative Interpretations: The reader should ask if alternative interpretations explain the results better than the researcher's explanation. Findings may fit patterns the researcher rejected or ignored.

5. Warranted Broader Conclusions
  • Correlation vs. Causation: Researchers must not claim causation where only correlation exists.

  • Aggression Example: Finding that children who watch violent TV are more aggressive does not mean TV causes aggression. Alternatives include:

    • Aggressive children prefer violent TV.

    • Violent TV only triggers violence in children already predisposed to it.

6. Meaningfulness ("The So What? Test")
  • Novelty and Utility: Does the study provide new knowledge or lead to future research?

  • Importance: Significant studies often produce surprising findings or help choose between opposing theories (Abelson,1995Abelson, 1995).

7. Ethics
  • Humanity: Are human or animal participants treated humanely?

  • Ends vs. Means: Does the incremental knowledge produced justify the methods used?

  • Regulation: The Australian Psychological Society (APS) publishes guidelines (2002,2003,20072002, 2003, 2007). Universities use Ethics Committees and Institutional Review Boards (IRBs) to review proposals. Approval is required even for benign studies on memory or math.

The Crisis of Replicability

Replicability is a key attribute of empirical research; scientific legitimacy depends on getting the same results when a study is repeated. However, social sciences currently face a "crisis of replicability."

The Reproducibility Project (Noseketal.,2015Nosek\,et\,al., 2015)
  • Scope: 270270 researchers attempted to replicate 100100 studies from three major psychology journals published in 20082008.

  • Findings:

    • Original studies showing statistically significant results: 97/10097/100.

    • Replications showing statistically significant results: 36/10036/100.

Perspectives on the Crisis
  • Premature Conclusion: StroebeandStrack(2014)Stroebe\,and\,Strack\,(2014) argue failure to replicate may be due to different time points or populations rather than flawed original theories.

  • Sample Size Issues: StangorandLemay(2016)Stangor\,and\,Lemay\,(2016) suggest small-to-medium sample sizes can misrepresent data to suggest larger effects than exist.

  • Statistical Power: SˊwiątkowskiandDompnier(2017)Świątkowski\,and\,Dompnier\,(2017) attribute low replicability to the recurrent use of low statistical power.

  • Publication Bias: Journals often prefer new, surprising results over less exciting replications. This incentivizes questionable practices like "cherry-picking" data or selective deletion of results (SocialScienceLibreTexts,2017Social\,Science\,LibreTexts, 2017).

  • Outright Dishonesty: Researcher Diederik Stapel (20112011) admitted to fabricating data across more than 3030 publications.

Solutions for Replicability
  • Open Research: Making all data, results, and protocols available for transparency (Leyser,Kingsley,&Grange,2017Leyser, Kingsley, \& Grange, 2017).

  • Pre-registration: Outlining methodology and analytical strategies before data collection to prevent selective data use.

  • Reward Structures: OttolineLeyserOttoline\,Leyser argues we must shift focus from "groundbreaking" headlines to reward confirmatory work and community data sharing.

Principles of Critical Thinking

Critical thinking is the logical and rational assessment of information, examining both strengths and weaknesses (Burton,2010Burton, 2010). It is supported by three key principles:

  1. Scepticism: Always questioning assumptions and assertions, regardless of whether they are in print or from an authority figure.

  2. Objectivity: Taking an impartial approach based on evidence and logic, setting aside personal feelings or biases.

  3. Open-mindedness: Willingness to consider all sides and potential explanations, even if they contradict personal experience.

Fallacies in Arguments

Arguments that do not flow logically from evidence are fallacious. Four common types include:

  1. Straw Man: Deliberately attacking a weak, decoy version of an opposing argument to make one's own look stronger.

  2. Appeals to Popularity: Assuming an argument is true because it is widely believed (e.g., historical belief in a flat earth).

  3. Appeals to Authority: Assuming an argument is true because a well-known or prestigious person said it.

  4. Arguments Directed to the Person (Ad Hominem): Attacking the character or failings of the author of an alternative argument rather than the argument itself.

Popular Myths in Psychology

Misinformation is common due to mass media, self-help books, and internet hoaxes.

  • The Role of Education: FurnhamandHughes(2014)Furnham\,and\,Hughes\,(2014) found that while psychology students recognize myths better than the general public, the effect size is small, highlighting the need for deep processing.

  • Fake News and Emotion: Weeks(2015)Weeks\,(2015) found people are more susceptible to fake news when angry. Anxious people are more open-minded. Fact-checking information alongside fake news helps minimize belief in myths.

Common Examples of Myths
  • Multiple Choice Tests: The myth to "stick with your gut feeling" is refuted by evidence; you should change your answer if you have a good reason (Lilienfeldetal.,2010Lilienfeld\,et\,al., 2010).

  • Attraction: The "playing hard to get" myth is contradicted by research showing men prefer women who are open to advances.

  • The Spinning Ballerina (Figure 2.72.7): Used to claim people are "left-brained" or "right-brained." Reality shows the brain works in an integrated fashion, though hemispheres have different specialized processes.

The Scientific Attitude and Process

The Criterion of Science

KarlPopper(1963)Karl\,Popper\,(1963) argued that the hallmark of science is the formulation of testable hypotheses that can be refuted (falsified). Experimental methods provide the strongest tests for these hypotheses.

Discovery vs. Justification
  • Context of Discovery: The stage where phenomena are observed and theories are built. Descriptive methods (case studies, naturalistic observation, surveys) are most useful here as they allow for unconstrained behavior.

  • Context of Justification: The stage where hypotheses are empirically tested. Experimental, quasi-experimental, and correlational designs are preferred. Inferential statistics help determine if findings are genuine.

Conclusion

Science is a mental map of phenomena. Optimal psychological science uses multiple methods and measures, gradually weeding out beliefs that do not withstand scientific scrutiny as technology advances.