Multi-Group Designs
Overview of Multi-Group Designs
- Multi-group designs are experimental research designs that involve more than two conditions.
- Technically, these designs consist of multiple conditions representing multiple levels of one Independent Variable (IV).
- These designs allow a researcher to compare multiple levels of a single IV simultaneously.
- While technically focused on levels of one IV, the term broadly encompasses any experiment with more than two conditions.
Philosophy and Pragmatics of Multi-Group Research
- Philosophical Foundations:
- These designs are used for the same fundamental reasons as two-group designs: to test hypotheses through experimentation.
- Additional conditions are often included because the researcher believes it is important to control for additional variables beyond what random assignment alone can account for.
- They allow the researcher to quantify the relationship between the IV and the Dependent Variables (DVs) rather than simply stating that a relationship exists.
- Pragmatic Considerations:
- Increased Efficiency: They allow for the testing of multiple hypotheses within a single sample.
- Hypothesis Multiplier: If a two-group design tests one hypothesis, adding a third group allows for the testing of 3 hypotheses in one study. A fourth group allows for 6 hypotheses.
- Confound Accountancy: It is more efficient to account for potential confounders at a single point in time. This avoids the potential confound of changes occurring over time that would exist across two separate studies.
- Implementation vs. Complexity:
- Multi-group designs are not significantly more complex to implement than two-group designs.
- However, they are more complex in terms of analysis and require recruiting a larger number of participants, which incurs extra costs.
- Recommendation: If there is no strong reason for additional conditions, the researcher should stick to a two-group design.
Efficient Testing and Comparisons
- Researchers use multi-group designs to test multiple "two-group" hypotheses—hypotheses that can be tested by comparing two specific conditions (e.g., comparing different types of therapies).
- Participant Efficiency: Testing two hypotheses in separate studies would require 4 groups of participants. Testing them in a single multi-group study requires only 3 groups.
- Calculation of Comparisons: In a study with k conditions, the number of possible comparisons is calculated as:
\frac{k(k-1)}{2}
- Note: The "reuse" of groups for multiple comparisons causes specific analytical issues that must be addressed.
Investigating Confounds and Control Conditions
- Multiple Types of Control Conditions: Multi-group designs help tease apart specific effects by using different control types.
- No Treatment Control: A "true" control condition exposed to no manipulation.
- Treatment-As-Usual Control: Checking if a new therapy is more effective than existing standards.
- Conceptual Control: A condition exposed to a "neutral" manipulation to account for nonspecific effects.
- Example: CBT and Depression:
- A researcher tests Cognitive-Behavioral Therapy (CBT) to see if it reduces depressive symptoms.
- Three conditions are utilized: CBT, No Treatment, and Nonspecific Talk Therapy.
- Talk therapy accounts for "nonspecific" treatment effects, which result from simply interacting with a therapist regardless of the technique used.
- Controlling for Multiple Factors Directly:
- Researchers compare multiple experimental conditions to a single control condition.
- Experimental conditions differ from each other on potential confounders while keeping the rest of the manipulation identical.
- Example: Music and Exercise:
- DV: Number of repetitions completed on an exercise.
- Condition 1 (True Control): No music.
- Condition 2 (Experimental): Listen to music and rate it before exercising.
- Condition 3 (Experimental): Listen to music and rate it after exercising.
- This design controls for "order effects" to ensure the act of rating the music doesn't distract the participant and confound the performance results.
Assessing Relationships Between Levels and Outcomes
- This is considered the most appropriate use of multi-group designs.
- While two-group designs only show the presence or absence of a relationship (nominal data), multiple groups at different levels can test for the direction and magnitude of the relationship.
- Level of Measurement:
- Multi-group designs improve the level of measurement for the IV.
- The IV is no longer restricted to the nominal level; it can be treated as ordinal, interval, or even ratio level data.
- Although it is hard to justify interval over ordinal, researchers often treat it as "close enough" to interval to use correlational analyses.
- Example: Therapy Duration:
- Hypothesized negative relationship: Longer therapy results in lower depressive symptoms.
- Design: No therapy (0 months), 6 months of CBT, and 1 year of CBT.
- This allows testing if 1 year is better than 6 months, and if 6 months is better than nothing.
Manipulation Design and Level Assignment
- Design Principles: Manipulations must be comparable across conditions and able to induce specific levels of the IV.
- Assigning Levels: The researcher assigns each condition to a level (e.g., 0, 6, and 12 months). This can introduce error if the levels aren't actually equidistant (e.g., the difference between 6 and 12 months may not equal the difference between 0 and 6).
- Sampling Levels: Instead of arbitrary picking, the researcher defines a range of levels and randomly samples from it.
- This is common in medication dosages but impractical in most psychological studies because it requires a huge number of conditions and a massive participant sample.
Analyzing Multi-group Data: Errors of Inference
- Multi-group data requires complex analyses because the same sample is used multiple times.
- Null-Hypothesis Significance Testing (NHST):
- If p<α, we reject the null hypothesis and assume there is a true effect.
- If p≥α, we fail to reject the null, meaning results can be explained by sampling error alone.
- Type I Error:
- Rejecting a "true" null hypothesis (a false positive).
- The significance threshold (conventional α=0.05) represents the risk of making this error.
- Type II Error:
- Failing to reject a null hypothesis that is actually false (a false negative).
- Represented by the Greek letter β. The likelihood of the error is 1−β.
- Power:
- The ability to correctly reject a false null hypothesis and detect a true effect.
- Power is based on three factors:
- Size of the "true" effect: Larger effects are easier to detect.
- Sample size: Larger samples increase power.
- Significance threshold: Higher thresholds (more lenient) make it easier to detect effects.
The Issue of Multiple Comparisons
- Statistical tests assume each test is independent (unique, non-reused samples).
- Reusing the same sample for multiple tests (as in multi-group designs) inflates the significance threshold, making it easier to wrongly reject the null.
- Pairwise Comparisons: Specific comparisons between two individual groups (e.g., a t-test between Therapy and Placebo).
- Correction Strategies:
- Familywise Error Rate: The probability of making at least one Type I error in a group (family) of tests.
- Bonferroni Correction: Divide the significance threshold by the number of tests (α/n). This is considered a conservative correction.
- False Discovery Rate (FDR): The proportion of significant results that are Type I errors.
- Benjamini-Hochberg Correction: Order p-values from smallest to largest and compare each to a modified threshold:
Threshold=total number of testsrank of p-value×FDR
Omnibus Statistical Tests
- Omnibus Test: A test that detects the presence of at least one difference between all groups simultaneously without making direct multiple comparisons.
- Analysis of Variance (ANOVA):
- Compares variance between conditions to variance within conditions.
- Used for DVs on interval or ratio scales.
- A significant ANOVA indicates that at least one condition differs from another.
- Kruskal-Wallis Test: Used as an omnibus test when the DV is ordinal.
- Chi-squared Test of Independence: Used when the DV is nominal. It checks if the number of people falling into categories differs across IV levels.
- A "goodness-of-fit" chi-squared is not used for multi-group comparisons as it only tests a single variable.
Degrees of Freedom and Follow-up Tests
- Degrees of Freedom (df): The amount of unique, non-redundant information in a data set.
- In a 3-condition ANOVA, there are 2 degrees of freedom, meaning only 2 non-redundant comparisons can be made.
- Planned Contrasts (A Priori):
- Comparisons specified before the initial test based on the study design.
- Often involves comparing sets of conditions (e.g., [Treatment + Placebo] vs. No Treatment).
- Post-hoc Tests (After the Fact):
- Conducted after a significant omnibus test.
- These are "protected" tests because the significant omnibus result reduces the likelihood of false positives.
- Fisher’s Least Significant Difference (LSD): A common post-hoc test.
- Tukey’s Honestly Significant Difference (HSD): A common post-hoc test.