Multi-Group Designs

Overview of Multi-Group Designs

  • Multi-group designs are experimental research designs that involve more than two conditions.
  • Technically, these designs consist of multiple conditions representing multiple levels of one Independent Variable (IV).
  • These designs allow a researcher to compare multiple levels of a single IV simultaneously.
  • While technically focused on levels of one IV, the term broadly encompasses any experiment with more than two conditions.

Philosophy and Pragmatics of Multi-Group Research

  • Philosophical Foundations:
    • These designs are used for the same fundamental reasons as two-group designs: to test hypotheses through experimentation.
    • Additional conditions are often included because the researcher believes it is important to control for additional variables beyond what random assignment alone can account for.
    • They allow the researcher to quantify the relationship between the IV and the Dependent Variables (DVs) rather than simply stating that a relationship exists.
  • Pragmatic Considerations:
    • Increased Efficiency: They allow for the testing of multiple hypotheses within a single sample.
    • Hypothesis Multiplier: If a two-group design tests one hypothesis, adding a third group allows for the testing of 3 hypotheses in one study. A fourth group allows for 6 hypotheses.
    • Confound Accountancy: It is more efficient to account for potential confounders at a single point in time. This avoids the potential confound of changes occurring over time that would exist across two separate studies.
  • Implementation vs. Complexity:
    • Multi-group designs are not significantly more complex to implement than two-group designs.
    • However, they are more complex in terms of analysis and require recruiting a larger number of participants, which incurs extra costs.
    • Recommendation: If there is no strong reason for additional conditions, the researcher should stick to a two-group design.

Efficient Testing and Comparisons

  • Researchers use multi-group designs to test multiple "two-group" hypotheses—hypotheses that can be tested by comparing two specific conditions (e.g., comparing different types of therapies).
  • Participant Efficiency: Testing two hypotheses in separate studies would require 4 groups of participants. Testing them in a single multi-group study requires only 3 groups.
  • Calculation of Comparisons: In a study with kk conditions, the number of possible comparisons is calculated as:    \frac{k(k-1)}{2}
  • Note: The "reuse" of groups for multiple comparisons causes specific analytical issues that must be addressed.

Investigating Confounds and Control Conditions

  • Multiple Types of Control Conditions: Multi-group designs help tease apart specific effects by using different control types.
    • No Treatment Control: A "true" control condition exposed to no manipulation.
    • Treatment-As-Usual Control: Checking if a new therapy is more effective than existing standards.
    • Conceptual Control: A condition exposed to a "neutral" manipulation to account for nonspecific effects.
  • Example: CBT and Depression:
    • A researcher tests Cognitive-Behavioral Therapy (CBT) to see if it reduces depressive symptoms.
    • Three conditions are utilized: CBT, No Treatment, and Nonspecific Talk Therapy.
    • Talk therapy accounts for "nonspecific" treatment effects, which result from simply interacting with a therapist regardless of the technique used.
  • Controlling for Multiple Factors Directly:
    • Researchers compare multiple experimental conditions to a single control condition.
    • Experimental conditions differ from each other on potential confounders while keeping the rest of the manipulation identical.
    • Example: Music and Exercise:
    • DV: Number of repetitions completed on an exercise.
    • Condition 1 (True Control): No music.
    • Condition 2 (Experimental): Listen to music and rate it before exercising.
    • Condition 3 (Experimental): Listen to music and rate it after exercising.
    • This design controls for "order effects" to ensure the act of rating the music doesn't distract the participant and confound the performance results.

Assessing Relationships Between Levels and Outcomes

  • This is considered the most appropriate use of multi-group designs.
  • While two-group designs only show the presence or absence of a relationship (nominal data), multiple groups at different levels can test for the direction and magnitude of the relationship.
  • Level of Measurement:
    • Multi-group designs improve the level of measurement for the IV.
    • The IV is no longer restricted to the nominal level; it can be treated as ordinal, interval, or even ratio level data.
    • Although it is hard to justify interval over ordinal, researchers often treat it as "close enough" to interval to use correlational analyses.
  • Example: Therapy Duration:
    • Hypothesized negative relationship: Longer therapy results in lower depressive symptoms.
    • Design: No therapy (0 months), 6 months of CBT, and 1 year of CBT.
    • This allows testing if 1 year is better than 6 months, and if 6 months is better than nothing.

Manipulation Design and Level Assignment

  • Design Principles: Manipulations must be comparable across conditions and able to induce specific levels of the IV.
  • Assigning Levels: The researcher assigns each condition to a level (e.g., 0, 6, and 12 months). This can introduce error if the levels aren't actually equidistant (e.g., the difference between 6 and 12 months may not equal the difference between 0 and 6).
  • Sampling Levels: Instead of arbitrary picking, the researcher defines a range of levels and randomly samples from it.
    • This is common in medication dosages but impractical in most psychological studies because it requires a huge number of conditions and a massive participant sample.

Analyzing Multi-group Data: Errors of Inference

  • Multi-group data requires complex analyses because the same sample is used multiple times.
  • Null-Hypothesis Significance Testing (NHST):
    • If p<αp < \alpha, we reject the null hypothesis and assume there is a true effect.
    • If p≥αp \geq \alpha, we fail to reject the null, meaning results can be explained by sampling error alone.
  • Type I Error:
    • Rejecting a "true" null hypothesis (a false positive).
    • The significance threshold (conventional α=0.05\alpha = 0.05) represents the risk of making this error.
  • Type II Error:
    • Failing to reject a null hypothesis that is actually false (a false negative).
    • Represented by the Greek letter β\beta. The likelihood of the error is 1−β1 - \beta.
  • Power:
    • The ability to correctly reject a false null hypothesis and detect a true effect.
    • Power is based on three factors:
    1. Size of the "true" effect: Larger effects are easier to detect.
    2. Sample size: Larger samples increase power.
    3. Significance threshold: Higher thresholds (more lenient) make it easier to detect effects.

The Issue of Multiple Comparisons

  • Statistical tests assume each test is independent (unique, non-reused samples).
  • Reusing the same sample for multiple tests (as in multi-group designs) inflates the significance threshold, making it easier to wrongly reject the null.
  • Pairwise Comparisons: Specific comparisons between two individual groups (e.g., a t-test between Therapy and Placebo).
  • Correction Strategies:
    • Familywise Error Rate: The probability of making at least one Type I error in a group (family) of tests.
    • Bonferroni Correction: Divide the significance threshold by the number of tests (α/n\alpha / n). This is considered a conservative correction.
    • False Discovery Rate (FDR): The proportion of significant results that are Type I errors.
    • Benjamini-Hochberg Correction: Order p-values from smallest to largest and compare each to a modified threshold:       Threshold=rank of p-valuetotal number of tests×FDRThreshold = \frac{\text{rank of p-value}}{\text{total number of tests}} \times FDR

Omnibus Statistical Tests

  • Omnibus Test: A test that detects the presence of at least one difference between all groups simultaneously without making direct multiple comparisons.
  • Analysis of Variance (ANOVA):
    • Compares variance between conditions to variance within conditions.
    • Used for DVs on interval or ratio scales.
    • A significant ANOVA indicates that at least one condition differs from another.
  • Kruskal-Wallis Test: Used as an omnibus test when the DV is ordinal.
  • Chi-squared Test of Independence: Used when the DV is nominal. It checks if the number of people falling into categories differs across IV levels.
    • A "goodness-of-fit" chi-squared is not used for multi-group comparisons as it only tests a single variable.

Degrees of Freedom and Follow-up Tests

  • Degrees of Freedom (dfdf): The amount of unique, non-redundant information in a data set.
    • In a 3-condition ANOVA, there are 2 degrees of freedom, meaning only 2 non-redundant comparisons can be made.
  • Planned Contrasts (A Priori):
    • Comparisons specified before the initial test based on the study design.
    • Often involves comparing sets of conditions (e.g., [Treatment + Placebo] vs. No Treatment).
  • Post-hoc Tests (After the Fact):
    • Conducted after a significant omnibus test.
    • These are "protected" tests because the significant omnibus result reduces the likelihood of false positives.
    • Fisher’s Least Significant Difference (LSD): A common post-hoc test.
    • Tukey’s Honestly Significant Difference (HSD): A common post-hoc test.