Evaluating experiments reading
Extraneous Variables, Internal Validity, and Methodological Considerations
Psychologists use experiments to establish cause-and-effect relationships by manipulating an independent variable (IV) and measuring its effect on a dependent variable (DV), while attempting to control other variables that could influence the outcome.
Key consequence: if a variable is not controlled, it becomes an extraneous variable that can influence results (a confounding variable).
Example: Testing whether music affects recall of a 20-word list, but not ensuring all participants are native English speakers. If one group has more non-native speakers, this becomes a confounding variable.
Extraneous variables can also be present in materials (e.g., all words are one-syllable), which may affect recall independent of listening to music.
Internal validity: the extent to which the study tests what it claims to test. Failure to control extraneous/ confounding variables compromises internal validity, making it unclear whether observed effects are due to the IV alone.
Methodological considerations pertain to the design and procedure of the experiment and are critical evaluative points in research.
Participant biases (demand characteristics): participants form interpretations of the study’s aim and may change behavior to fit that interpretation.
More common in repeated-measures designs but can occur in independent-samples designs, observations, and interviews.
Four common types of participant biases include:
The expectancy effect (a form of compliance): participants act to help the researcher achieve what they think is desired.
Simply knowing they are in an experiment can lead to behavior they would not exhibit otherwise.
The screw-you effect: participants try to figure out the hypotheses and deliberately undermine the study’s credibility; more likely with certain sampling techniques (e.g., opportunity samples with undue pressure).
Social desirability effect: participants respond in socially acceptable ways when observed or evaluated.
- Reactivity: participants change behavior simply because they know they are being observed; can yield anxiety or overconfidence, affecting performance in interviews or clinical settings.
Page 2: Orne (1962) and demand characteristics
Orne conducted a study where participants solved addition problems with 224 numbers per page, torn up pages after each, and were told to keep working until told to stop. Despite boredom, participants continued for hours because they believed they were in an experiment.
The screw-you effect can occur when participants try to reveal or undermine researchers’ hypotheses; it is uncommon but more likely with some sampling methods (e.g., students in an opportunity sample under pressure to participate) or when researchers appear arrogant or condescending.
Participants typically try to protect their self-esteem, leading to social desirability or self-protective responses.
Relevance to practice: demand characteristics can distort results, particularly in observational and interview contexts.
Controlling demand characteristics
Use an independent samples design to avoid participants experiencing multiple conditions and guessing the goal of the study.
Debrief participants after participation to assess whether they understood what was being tested; if they can articulate the test, results may be influenced.
Deception, Order Effects, and Experimental Design
Deception is sometimes used to avoid demand characteristics, but it raises ethical concerns if it causes undue stress or harm.
Order effects: systematic changes in participants’ responses due to the order of exposure to conditions in a repeated-measures design.
Three common order effects:
Fatigue effects: tiredness or boredom across multiple conditions affects performance.
Interference effects: prior conditions influence performance in subsequent conditions (e.g., memorizing two lists of words with/without music, where words may overlap).
Practice effects: improvement due to repetition or familiarity rather than the manipulation itself.
Controlling order effects
Counterbalancing: vary the order of conditions across participants. Example: condition A (no music) first for Group 1, then condition B (with music); Group 2 receives condition B first, then A. If order effects are absent, the two groups should yield similar results.
Include a sufficiently long pause between conditions to reduce carryover effects.
Use a filler task between conditions to clear the mental palette (e.g., count backward by 3 from 100) before the next condition.
Researcher Bias, P-Hacking, and Publication/Funding Bias
Researcher biases can affect results and conclusions. Key biases include:
Confirmation bias: seeking or interpreting information to confirm a preexisting belief or hypothesis.
Example: observing playground aggression and focusing on aggressive behavior in boys while ignoring aggression in girls.
P-hacking: exploiting data in ways not originally planned to find statistically significant results; concerns arise when hypotheses are not pre-specified.
Example: after a study on music and word recall yields no significant results, examining multiple variable combinations and reporting a significant result at or beyond without a priori hypotheses.
Publication bias: tendency to publish studies with positive results more than null results.
Funding bias: potential bias when research is funded by drug companies or other corporations; meta-analyses may be biased if authors are employed by sponsors or have sponsorship.
Meta-analysis example (Ebrahim et al., 2016): antidepressant studies
Analyzed 185 meta-analyses evaluating antidepressants for depression (Jan 2007–Mar 2014).
Findings:
of meta-analyses had authors who were employees of the assessed drug manufacturer.
had sponsorship or authors related to a drug company.
Only of meta-analyses drew negative conclusions about the antidepressant.
Meta-analyses conducted by a drug-manufacturer employee were less likely to report negative statements about the drug than other meta-analyses.
Controlling Researcher Bias and Validity Concepts
Controlling researcher bias:
Decide on a hypothesis before conducting research; avoid adapting the hypothesis after seeing results. If new hypotheses arise, run a new experiment to test them.
To control confirmation bias, use researcher triangulation to improve inter-rater reliability. For example, a team observing playground aggression should converge on the same assessment across raters.
Double-blind control: standard method to reduce bias
Participants are randomly assigned to experimental vs. control groups and do not know which group they are in.
A third party knows which participants received which treatment, so the researcher analyzing data remains unaware of the assignments.
Construct validity: whether the measurement truly captures the theoretical construct of interest
Operationalization matters: the way a variable is measured should reflect the intended theoretical construct.
Example: testing attitudes toward Americans among Europeans by asking about watching US films, owning an iPhone, watching CNN, or wearing American designer clothes may not validly measure bona fide pro-American attitudes.
There are several problematic constructs in psychology (e.g., intelligence, love, aggression) due to their abstract nature and measurement difficulties.
External validity: extent to which results generalize to other situations and populations
Population (sampling) validity: the sample should be representative of the population from which it was drawn; otherwise, generalizability is limited.
Ecological validity: whether results generalize to real-world settings beyond the lab; highly controlled laboratory conditions can limit real-world applicability.
Tension between control and realism: highly artificial lab conditions may not predict what happens in normal life (e.g., lab shocks, watching a video of a car crash rather than a real car crash).
Ethics and replication:
If ethical protocols are not followed, research cannot be reliably replicated, undermining reliability of results.
Ethical considerations are integral to the interpretation and acceptability of findings.
Real-World Relevance and Examples
Illustrative considerations:
When studying music’s impact on memory, ensure that the music type, tempo, familiarity, and participants’ language and cultural background do not introduce unintended confounds.
In evaluating behavior from playground observations, ensure multiple independent raters to mitigate observer bias and confirm inter-rater reliability.
In pharmaceutical research, scrutinize funding sources and potential conflicts of interest; prefer preregistered analyses and transparent data sharing to minimize p-hacking and publication bias.
Ethics, Replication, and Practical Takeaways
Ethical guidelines are not optional; they safeguard participants and the credibility of science.
Researchers should preregister hypotheses, maintain methodological rigor, and consider multiple validity facets (construct, external, ecological).
Before drawing conclusions, consider allF potential biases, design limitations, and whether the results would hold under more naturalistic settings.
When communicating findings, transparently report methods, data analyses, and funding sources to enable replication and independent evaluation.
// Key formulas and numerical references used in the notes
Significance threshold mentioned in p-hacking example:
Reported higher-level statistics from the funding bias study:
Drug company employment in authors:
Sponsorship or related authors:
Negative conclusions about the drug:
Relative likelihood of negative statements when authored by a drug-company employee: less likely