Within-Subjects Designs Lecture Notes
Introduction to Within-Subjects Designs
- Definition: Within-subjects designs are research designs that involve measuring participants on a dependent variable (DV) multiple times.
- Core Concept: Unlike between-subjects designs where different groups are compared, within-subjects designs compare the same individuals across different conditions or time points.
- General Structure:
- These designs provide significant advantages over between-subjects designs in terms of power and efficiency.
- They introduce specific potential issues (e.g., order effects) that must be managed through specialized study designs and statistical analyses.
- Researchers can combine within-subjects and between-subjects components in a single study.
Primary Types of Within-Subjects Designs
Pretest-Posttest Designs
- Description: This involves measuring a specific outcome before (pretest) and after (posttest) a specific intervention or manipulation.
- Purpose: Typically used to observe how the DV changes specifically due to the manipulation.
- Clinical Relevance: Frequently applied in treatment contexts.
- Example: Measuring depression scores before a therapeutic intervention and measuring them again after the intervention to gauge effectiveness.
- Complexity: Considered the simplest form of within-subjects design.
Repeated Measures Designs
- Description: These designs involve exposing participants to every level of the independent variable (IV) and measuring the outcome after each exposure.
- Application: Used when the IV has multiple levels () and the researcher aims to compare the effects of each level.
- Comparison to Pretest-Posttest: If the IV has only levels, it is essentially a pretest-posttest design.
- Example: Comparing three different therapeutic techniques—Talk therapy, Cognitive Behavioral Therapy (CBT), and Dialectical Behavior Therapy (DBT)—on depression levels. Each participant undergoes each therapy type in a pre-specified order.
Longitudinal Designs
- Description: Similar to between-subjects longitudinal research, but focuses on comparing participants to themselves at multiple time points rather than comparing different groups.
- Function: Allows researchers to address questions regarding the change of a DV over time.
- Measurement: Because there are multiple points of measurement, researchers can test within-subjects hypotheses regarding individual trajectories.
Philosophy and Pragmatics
Theoretical Philosophy
- Some argue that within-subjects designs are the only "correct" way to test most psychological hypotheses.
- Between-Subjects Limitation: These designs only test if two groups (e.g., Treatment vs. No Treatment) are different. Researchers assume the treatment caused the difference, but change is never actually measured.
- Within-Subjects Advantage: These designs allow for the direct measurement of change, providing a stronger basis for attributing that change to the IV.
Practical Considerations
- Selection Criteria: You choose a within-subjects design if you are interested in demonstrating change (pre- to post-manipulation or change over time).
- Longitudinal Correlational Research: Can be used to assess relationships between variables as they change over time.
- Contraindications: If the research question does not involve assessing change, a between-subjects design may be more appropriate.
Benefits of Within-Subjects Designs
Statistical Power
- Within-subjects designs are more powerful than between-subjects designs, meaning they have a higher probability of detecting an effect if one actually exists ().
Efficient Hypothesis Testing
- Sample Efficiency: Each participant exists in every condition.
- Sample Size Savings: The more levels of the IV, the more participants are "saved."
- Example: To test an IV with levels between-subjects, you would need separate samples. In a within-subjects design, you only need sample to test the same hypotheses.
- This makes the design more efficient than even multi-group between-subjects designs.
Participants as Own Controls
- Individuals are compared against themselves rather than a separate group of people who are merely "like" them.
- Error Reduction: This significantly reduces the error term in statistical tests.
- Accounting for Variance: Instead of relying on random assignment to balance out individual differences (intelligence, personality, etc.), these differences are directly accounted for because the person is the same in every condition.
- This increase in power is distinct from the efficiency of the sample size.
Issues with Within-Subjects Designs
Independence of Data
- Statistical Assumption: Most standard analyses assume data points are independent (each comes from a different person).
- Violation: Within-subjects data are dependent by definition because multiple scores come from the same person ().
- Consequence: Traditional random assignment cannot "wash out" individual differences in the same way, and data must be handled using specific dependent-sample statistics.
Learning Effects
- Carryover Interference: Exposing a participant to one manipulation may change their response to the next.
- Example: If Therapy A effectively resolves depression, Therapy B will appear ineffective simply because there is no depression left to treat.
- Testing Effect: Repeatedly using the same measurement tool (e.g., the same survey) can change participant responses or lead to a loss of reliability.
- Measurement Error: Switching to different versions of a measure to avoid the testing effect introduces new sources of error.
Maturation and Attrition
- Maturation Effects: Participants may change over time due to natural processes (aging, fatigue, hunger) rather than the manipulation. It is difficult to disentangle these from the treatment effect.
- Attrition Effects: Participants may drop out of the study.
- The "Worst Case": Losing one participant means losing data for every condition. This is particularly problematic if the manipulation itself causes the attrition (e.g., a treatment that is too difficult to complete).
Order Effects
- These are problems caused specifically by the sequence in which materials are presented:
- Practice Effects: Change in behavior due to familiarity with the task or measures.
- Fatigue Effects: Boredom or tiredness over the course of the study introduces error.
- Carryover Effects: Current responses are affected by previous manipulations.
- Sensitization Effects: Exposure to study materials helps participants "guess" the hypothesis, which alters their natural behavior.
Solutions for Within-Subjects Issues
Oversampling
- Definition: Collecting a larger sample than necessary for the target power level, specifically to compensate for expected attrition.
- Calculation: Researchers look at attrition rates for similar designs, which can be between and for long studies.
- Metric: One may need to oversample by to the required final .
- Caveat: Oversampling can lead to excessive power. Too much power can make trivial results statistically significant even if they lack practical importance.
Control Groups
- Used to resolve maturation and non-random attrition.
- Method: Collect a separate control group measured over the same duration as the experimental group but without the intervention.
- Analysis: If the control group also shows change, that change is attributed to maturation and can be statistically controlled.
Counterbalancing
- Definition: Presenting manipulations or measures in every possible order to different participants.
- Goal: Order effects should "wash out" across different sequences.
- Calculation: The number of possible orders scales factorially with the number of conditions ().
- conditions: orders.
- conditions: orders.
- Limitations:
- As the number of conditions increases, the required sample size to cover all orders "explodes."
- High power from oversampling might lead to finding significant differences between orders that are actually negligible.
Latin Square Designs
- Definition: A more efficient alternative to counterbalancing that controls for the position of a measure rather than every possible sequence.
- Efficiency: While manipulations require orders for full counterbalancing, they only have possible positions.
- Structure: Each manipulation appears in each position exactly once.
- Example Matrix for 4 Manipulations (A, B, C, D):
- Order 1: A, B, C, D
- Order 2: D, A, B, C
- Order 3: C, D, A, B
- Order 4: B, C, D, A
- Analysis Potential: Allows researchers to check for manipulation effects, position effects (e.g., fatigue at position ), and interactions between manipulation and position.
Analyzing Within-Subjects Data
General Approach
- Analysts use adjusted versions of independent-sample tests.
- These tests account for dependence by either:
- Subtracting pairs of scores: Creating a single "difference score" ().
- Adjusting error terms: Using previous scores as covariates (typically taught in graduate-level courses).
Paired-Samples t-tests
- Other Names: Dependent samples t-test.
- Use Case: Comparing exactly two sets of dependent scores (e.g., Pretest and Posttest).
- Mechanism: Calculates the difference between scores for each participant. These difference scores are independent because each person only has one, allowing for a standard t-test against a null hypothesis of .
Repeated Measures ANOVA
- Other Names: Within-subjects ANOVA.
- Use Case: Comparing more than two sets of scores (multiple time points or multiple manipulations).
- Utility: Tests if there is at least one significant difference between time points or levels of the IV.
- Follow-up: Requires post-hoc tests or planned contrasts to locate the specific differences.
Questions & Discussion: Group Practice
Scenario 1: Reading Intervention in Elementary Children
Intervention vs. Standard Curriculum
Between-Subjects Design:
- Participants/Assignment: Take one group of children, randomly assign half to the new intervention and half to the standard curriculum.
- Statistical Test: Independent-samples t-test.
Within-Subjects Design:
- Participants/Assignment: All children receive the standard curriculum for a period, then change to the reading intervention.
- Statistical Test: Paired-samples t-test.
Pros/Cons:
- Between: Avoids carryover/learning effects but needs more participants and has more noise from individual child differences.
- Within: Requires fewer children and controls for individual learning speeds, but earlier exposure to standard curriculum might influence learning in the intervention phase.
Scenario 2: Multiple Interventions
Intervention A vs. Intervention B vs. Standard
Between-Subjects Design:
- Participants/Assignment: Randomly assign children into three distinct groups (Group A, Group B, Group C).
- Statistical Test: One-Way ANOVA.
Within-Subjects Design:
- Participants/Assignment: Every child is exposed to the Standard curriculum, Intervention A, and Intervention B in a counterbalanced or Latin Square order.
- Statistical Test: Repeated Measures ANOVA.
Pros/Cons:
- Between: Cleaner separation of treatment effects; no fatigue from multiple reading tests.
- Within: Highly efficient usage of a single classroom of students; direct comparison of which intervention works best for the same child. Major risk: Fatigue and practice effects on reading tests.