Exhaustive Guide to Psychological Validity: Internal and External Perspectives
Defining Validity in Psychological Research
- The APA Definition of Validity: According to the APA Dictionary of Psychology (2018), validity is defined as "the degree to which empirical evidence and theoretical rationales support the adequacy and appropriateness of conclusions drawn from some form of assessment."
- Multiple Forms of Validity: Validity does not refer to a single attribute; it manifests in multiple forms depending on the research question and the specific inference being made. Prominent forms in psychology include:
- External Validity.
- Internal Validity.
- Ecological Validity.
- Statistical Conclusion Validity.
- Critical Evaluation of Findings: Research should never be accepted as true simply because it is published or reports a p-value below 0.05. Clinicians and researchers must evaluate studies to ensure methods are high-quality, trustworthy, and useful for real-world policy and practice.
External Validity and Generalization
- Definition: External validity is the extent to which research results can be generalized beyond the specific sample generated for the study.
- Difference from Reliability: Reliability focuses on whether results can be replicated under the same conditions. External validity asks: "Are these findings only unique to this particular sample?"
- Generalization Pathways: Generalization involves applying research findings to:
- A Wider Population: Moving from a small sample (e.g., 30 undergraduates) to the broader population (e.g., all university students in Australia).
- Other Studies: Comparing and building upon different research findings.
- Real-World Contexts: Applying laboratory or controlled study results to specific settings like treatment centers or clinics.
Specific Threats to External Validity
- Selection Bias: This occurs when there is a mismatch between sample characteristics and the target population. Recruitment methods, such as relying on easily accessible samples, heighten this risk.
- Volunteer Bias: Volunteers may be more curious, extroverted, or incentivized by compensation than the general population.
- Subpopulation Extrapolation: It is erroneous to assume findings from one group (e.g., women in a body image study) apply to another (e.g., men) due to differing concerns and disturbance severities.
- Cross-Species Generalization: Research on rats or monkeys may be misrepresented by media as proven in humans. The Twitter account "@justsaysinmice" highlights this issue.
- Measurement and Constructs: Concepts are narrowed into constructs and then operationalized through measurement (e.g., questionnaires or devices).
- Construct-to-Construct Generalization: Extrapolating findings from one scale (e.g., Wechsler Adult Intelligence Scale) to a different construct (e.g., emotional intelligence) is invalid.
- Single Measures/Methods: Relying on one scale (e.g., Drive for Thinness Scale for body image) or one method (e.g., self-report) may fail to capture the full scope of a phenomenon. Multiple methods (e.g., expert judging panels) provide more objective data.
- Testing Effects (Multiple Treatment Interference): Common in repeated measures studies where testing occurs over multiple time points (e.g., every 6 months for years).
- Fatigue: Participants become tired during intense or long testing phases, altering performance.
- Practice Effects: This includes habituation (familiarization with procedures) and sensitization (learning how to perform better on a specific measure, such as an IQ test).
- Experimental Effects: Participants acting unnaturally because they are in a study.
- Novelty Effect: Behavior changes because the situation feels new or exciting.
- Demand Characteristics: Participants try to guess the study aim and behave in a way they think the researcher wants (the "good participant" effect).
- Subject Effects: Includes compensatory rivalry (the control group works harder to compete) and demoralization (the control group acts sad or performs worse because they feel they are missing out on treatment).
- Experimenter Effects: Both conscious and unconscious behaviors.
- Researcher Characteristics: Age, gender, or attractiveness influencing participant responses.
- Confirmation Bias: Unconscious behaviors that increase the likelihood of data matching predictions.
Internal Validity: The Soundness of Results
- Definition: Internal validity relates to the internal structure of a study and the lack of flaws. It ensures that a single, unambiguous explanation exists for the relationship between two variables: the Independent Variable (IV) and Dependent Variable (DV).
- The Causality Question: It asks: "Can I trust that these findings are actually what they say they are, and can I draw cause-and-effect conclusions?"
- Extraneous vs. Confounding Variables:
- Extraneous Variables: Any variable in a study other than the specific variables being studied.
- Confounding Variables: Unmonitored extraneous variables that provide an alternative explanation for the results. Researchers try to identify and statistically control for these, though it is impossible to capture all of them.
Threats to Internal Validity
- Environmental Threats (History Effects): External events that occur outside the study but influence the outcome.
- Examples: Noise during testing, governmental lockdowns (COVID-19), or specific time-pressure events like the HSC exam period.
- Idiosyncratic Events: Personal events like a participant winning the lottery, creating "noise" in the data.
- Participant Variables: Individual differences such as age, wealth, or maturation.
- Selection Bias in Quasi-Experiments: When participants choose their own groups, their inherent characteristics may explain the outcome rather than the IV.
- Attrition (Dropout): Participants who leave a study may differ significantly from those who stay (e.g., those who dislike a treatment drop out, making the treatment appear more effective than it is).
- Instrumentation Threats: Changes in how an outcome is measured over time.
- Equipment: Wear and tear on devices, using different brands across sites, or calibration shifts.
- Raters: Fatigue or the replacement of a strict rater with a more lenient one in observational studies.
- Regression to the Mean: A statistical phenomenon where participants with extreme (very high or low) initial scores naturally score closer to the average on second testing. This is a major threat when specific low-scoring or high-scoring groups are selected for intervention.
Maximizing Validity and The Validity Trade-off
- The Inverse Relationship: Increasing internal validity by strictly controlling extraneous variables often reduces external validity (generalizability). Conversely, naturalistic studies have high external validity but low control.
- Phases of Research (Clinical Trial Example):
- Phase 1 (Animal Studies): High internal validity (controlled lab), zero external validity (cannot generalize to humans).
- Phase 2 (Human RCTs): High control and internal validity, moderate external validity (humans in a lab setting).
- Phase 3 (Effectiveness in Community): Maximized external validity (real-world use), minimized internal validity (less control over how patients take drugs).
Practical Methods to Reduce Risk of Bias
- Randomization: Using random allocation to ensure differences are equally dispersed across groups.
- Control Groups: Adding groups to ensure threats are equal across experimental conditions.
- Blinding:
- Participant Blinding: Using placebos or withholding the true aim until the end to reduce demand characteristics.
- Experimenter Blinding: Blinding the researcher or statistician to group allocation to reduce confirmation bias.
- Sampling Strategies:
- Random Sampling: Selecting participants randomly from the whole population.
- Stratified Random Sampling: Ensuring specific strata (subgroups), such as Indigenous Australians, are proportionally represented.
- Measurement variety: Using multiple measures or methods to capture a construct fully.
- Participant Management: Reducing contact between the principal researcher and participants; using filler tasks to mask the study purpose.
- Statistical Comparison: Comparing sample statistics to known population parameters (e.g., mean age, percentage of women) to confirm representativeness.
Questions & Discussion
- University Sample Scenario: A study compares CBT and Schema therapy for depression. The advertisement is placed in the university, seeking volunteers with no comorbidities. The therapists are clinical psychology master's students.
- External Validity Concerns: The findings may not generalize to community private practices where patients have high comorbidities and access more experienced clinicians of varying seniority levels. Volunteers may be compensated, whereas private patients pay for treatment.
- Internal Validity Strengths: The student sample and lack of comorbidities create a homogenous group, ensuring results are likely due to the therapy itself rather than other clinical factors. Uniform therapist experience level controls for the effect of clinical expertise.
- COVID-19 and Hospital Admissions: In a study comparing 2019 and 2020 hospital admissions per 10,000 people, an increase was noted. However, data from 2018 shows a year-on-year increase existing prior to the pandemic.
- Internal Validity Alert: Time might be a better explanatory factor than COVID-19 for the general trend. However, a sharp peak in admissions between March and May 2020 (post-lockdown) suggests an acute stress reaction specific to the pandemic.
- Study Location Reflection: How does selecting participants from a single postal code limit external validity?
- Response: This leads to selection bias, as a single area may have unique socioeconomic or demographic characteristics not representative of the broader population. It can be improved by multi-site recruitment or random sampling across various regions.