Within-Subjects Designs Lecture Notes

Introduction to Within-Subjects Designs

  • Definition: Within-subjects designs are research designs that involve measuring participants on a dependent variable (DV) multiple times.
  • Core Concept: Unlike between-subjects designs where different groups are compared, within-subjects designs compare the same individuals across different conditions or time points.
  • General Structure:
    • These designs provide significant advantages over between-subjects designs in terms of power and efficiency.
    • They introduce specific potential issues (e.g., order effects) that must be managed through specialized study designs and statistical analyses.
    • Researchers can combine within-subjects and between-subjects components in a single study.

Primary Types of Within-Subjects Designs

Pretest-Posttest Designs

  • Description: This involves measuring a specific outcome before (pretest) and after (posttest) a specific intervention or manipulation.
  • Purpose: Typically used to observe how the DV changes specifically due to the manipulation.
  • Clinical Relevance: Frequently applied in treatment contexts.
  • Example: Measuring depression scores before a therapeutic intervention and measuring them again after the intervention to gauge effectiveness.
  • Complexity: Considered the simplest form of within-subjects design.

Repeated Measures Designs

  • Description: These designs involve exposing participants to every level of the independent variable (IV) and measuring the outcome after each exposure.
  • Application: Used when the IV has multiple levels (>2> 2) and the researcher aims to compare the effects of each level.
  • Comparison to Pretest-Posttest: If the IV has only 22 levels, it is essentially a pretest-posttest design.
  • Example: Comparing three different therapeutic techniques—Talk therapy, Cognitive Behavioral Therapy (CBT), and Dialectical Behavior Therapy (DBT)—on depression levels. Each participant undergoes each therapy type in a pre-specified order.

Longitudinal Designs

  • Description: Similar to between-subjects longitudinal research, but focuses on comparing participants to themselves at multiple time points rather than comparing different groups.
  • Function: Allows researchers to address questions regarding the change of a DV over time.
  • Measurement: Because there are multiple points of measurement, researchers can test within-subjects hypotheses regarding individual trajectories.

Philosophy and Pragmatics

Theoretical Philosophy

  • Some argue that within-subjects designs are the only "correct" way to test most psychological hypotheses.
  • Between-Subjects Limitation: These designs only test if two groups (e.g., Treatment vs. No Treatment) are different. Researchers assume the treatment caused the difference, but change is never actually measured.
  • Within-Subjects Advantage: These designs allow for the direct measurement of change, providing a stronger basis for attributing that change to the IV.

Practical Considerations

  • Selection Criteria: You choose a within-subjects design if you are interested in demonstrating change (pre- to post-manipulation or change over time).
  • Longitudinal Correlational Research: Can be used to assess relationships between variables as they change over time.
  • Contraindications: If the research question does not involve assessing change, a between-subjects design may be more appropriate.

Benefits of Within-Subjects Designs

Statistical Power

  • Within-subjects designs are more powerful than between-subjects designs, meaning they have a higher probability of detecting an effect if one actually exists (1β1 - \beta).

Efficient Hypothesis Testing

  • Sample Efficiency: Each participant exists in every condition.
  • Sample Size Savings: The more levels of the IV, the more participants are "saved."
  • Example: To test an IV with 33 levels between-subjects, you would need 33 separate samples. In a within-subjects design, you only need 11 sample to test the same hypotheses.
  • This makes the design more efficient than even multi-group between-subjects designs.

Participants as Own Controls

  • Individuals are compared against themselves rather than a separate group of people who are merely "like" them.
  • Error Reduction: This significantly reduces the error term in statistical tests.
  • Accounting for Variance: Instead of relying on random assignment to balance out individual differences (intelligence, personality, etc.), these differences are directly accounted for because the person is the same in every condition.
  • This increase in power is distinct from the efficiency of the sample size.

Issues with Within-Subjects Designs

Independence of Data

  • Statistical Assumption: Most standard analyses assume data points are independent (each comes from a different person).
  • Violation: Within-subjects data are dependent by definition because multiple scores come from the same person (Xi1,Xi2,,XikX_{i1}, X_{i2}, \dots, X_{ik}).
  • Consequence: Traditional random assignment cannot "wash out" individual differences in the same way, and data must be handled using specific dependent-sample statistics.

Learning Effects

  • Carryover Interference: Exposing a participant to one manipulation may change their response to the next.
  • Example: If Therapy A effectively resolves depression, Therapy B will appear ineffective simply because there is no depression left to treat.
  • Testing Effect: Repeatedly using the same measurement tool (e.g., the same survey) can change participant responses or lead to a loss of reliability.
  • Measurement Error: Switching to different versions of a measure to avoid the testing effect introduces new sources of error.

Maturation and Attrition

  • Maturation Effects: Participants may change over time due to natural processes (aging, fatigue, hunger) rather than the manipulation. It is difficult to disentangle these from the treatment effect.
  • Attrition Effects: Participants may drop out of the study.
  • The "Worst Case": Losing one participant means losing data for every condition. This is particularly problematic if the manipulation itself causes the attrition (e.g., a treatment that is too difficult to complete).

Order Effects

  • These are problems caused specifically by the sequence in which materials are presented:
    • Practice Effects: Change in behavior due to familiarity with the task or measures.
    • Fatigue Effects: Boredom or tiredness over the course of the study introduces error.
    • Carryover Effects: Current responses are affected by previous manipulations.
    • Sensitization Effects: Exposure to study materials helps participants "guess" the hypothesis, which alters their natural behavior.

Solutions for Within-Subjects Issues

Oversampling

  • Definition: Collecting a larger sample than necessary for the target power level, specifically to compensate for expected attrition.
  • Calculation: Researchers look at attrition rates for similar designs, which can be between 50%50\% and 75%75\% for long studies.
  • Metric: One may need to oversample by 2×2 \times to 4×4 \times the required final NN.
  • Caveat: Oversampling can lead to excessive power. Too much power can make trivial results statistically significant even if they lack practical importance.

Control Groups

  • Used to resolve maturation and non-random attrition.
  • Method: Collect a separate control group measured over the same duration as the experimental group but without the intervention.
  • Analysis: If the control group also shows change, that change is attributed to maturation and can be statistically controlled.

Counterbalancing

  • Definition: Presenting manipulations or measures in every possible order to different participants.
  • Goal: Order effects should "wash out" across different sequences.
  • Calculation: The number of possible orders scales factorially with the number of conditions (n!n!).
    • 33 conditions: 3!=63! = 6 orders.
    • 44 conditions: 4!=244! = 24 orders.
  • Limitations:
    • As the number of conditions increases, the required sample size to cover all orders "explodes."
    • High power from oversampling might lead to finding significant differences between orders that are actually negligible.

Latin Square Designs

  • Definition: A more efficient alternative to counterbalancing that controls for the position of a measure rather than every possible sequence.
  • Efficiency: While 44 manipulations require 2424 orders for full counterbalancing, they only have 44 possible positions.
  • Structure: Each manipulation appears in each position exactly once.
  • Example Matrix for 4 Manipulations (A, B, C, D):
    • Order 1: A, B, C, D
    • Order 2: D, A, B, C
    • Order 3: C, D, A, B
    • Order 4: B, C, D, A
  • Analysis Potential: Allows researchers to check for manipulation effects, position effects (e.g., fatigue at position 44), and interactions between manipulation and position.

Analyzing Within-Subjects Data

General Approach

  • Analysts use adjusted versions of independent-sample tests.
  • These tests account for dependence by either:
    • Subtracting pairs of scores: Creating a single "difference score" (Time2Time1Time_2 - Time_1).
    • Adjusting error terms: Using previous scores as covariates (typically taught in graduate-level courses).

Paired-Samples t-tests

  • Other Names: Dependent samples t-test.
  • Use Case: Comparing exactly two sets of dependent scores (e.g., Pretest and Posttest).
  • Mechanism: Calculates the difference between scores for each participant. These difference scores are independent because each person only has one, allowing for a standard t-test against a null hypothesis of 00.

Repeated Measures ANOVA

  • Other Names: Within-subjects ANOVA.
  • Use Case: Comparing more than two sets of scores (multiple time points or multiple manipulations).
  • Utility: Tests if there is at least one significant difference between time points or levels of the IV.
  • Follow-up: Requires post-hoc tests or planned contrasts to locate the specific differences.

Questions & Discussion: Group Practice

Scenario 1: Reading Intervention in Elementary Children

Intervention vs. Standard Curriculum

  • Between-Subjects Design:

    • Participants/Assignment: Take one group of children, randomly assign half to the new intervention and half to the standard curriculum.
    • Statistical Test: Independent-samples t-test.
  • Within-Subjects Design:

    • Participants/Assignment: All children receive the standard curriculum for a period, then change to the reading intervention.
    • Statistical Test: Paired-samples t-test.
  • Pros/Cons:

    • Between: Avoids carryover/learning effects but needs more participants and has more noise from individual child differences.
    • Within: Requires fewer children and controls for individual learning speeds, but earlier exposure to standard curriculum might influence learning in the intervention phase.

Scenario 2: Multiple Interventions

Intervention A vs. Intervention B vs. Standard

  • Between-Subjects Design:

    • Participants/Assignment: Randomly assign children into three distinct groups (Group A, Group B, Group C).
    • Statistical Test: One-Way ANOVA.
  • Within-Subjects Design:

    • Participants/Assignment: Every child is exposed to the Standard curriculum, Intervention A, and Intervention B in a counterbalanced or Latin Square order.
    • Statistical Test: Repeated Measures ANOVA.
  • Pros/Cons:

    • Between: Cleaner separation of treatment effects; no fatigue from multiple reading tests.
    • Within: Highly efficient usage of a single classroom of students; direct comparison of which intervention works best for the same child. Major risk: Fatigue and practice effects on reading tests.