Introduction to Statistical Significance and Hypothesis Testing

Overview of Descriptive and Inferential Statistics

  • Statistical significance and hypothesis testing are foundational concepts in research, specifically detailed in Chapter 9 of the study materials.

  • There are two primary categories of statistics used in research: descriptive and inferential.

  • Descriptive Statistics: These are used to describe the sample data generated from a study.

    • They relate to measures such as the average (mean), standard deviation (SDSD), variance (s2s^2), and ranges.

    • They allow researchers to organize and summarize data to better describe how the sample performed.

    • Descriptive statistics provide the raw data used for further analysis but do not allow for interpretation regarding the effectiveness of an intervention in isolation.

  • Inferential Statistics: These data are used to make conclusions or "inferences" about how a general population would perform if the study were replicated or applied to them.

    • Descriptive data serve as the input for inferential statistics.

    • Inferential statistics allow researchers to determine statistical significance, effect size, practical significance, or clinical significance.

    • Common inferential tests include:

      • ANOVAs (Analysis of Variance).

      • MANOVA (Multivariate Analysis of Variance).

      • ANCOVA (Analysis of Covariance).

      • t-tests.

      • Multiple regressions.

Understanding the Concept of 'Significant' in Research

  • In everyday conversation, "significant" often means "important." In the world of research, "significant" means probably true.

  • Statistical Significance: This indicates that a found difference is real and reliable, rather than being the result of chance or error.

    • It suggests that if the study were repeated with another sample from the same population, similar results would likely be obtained.

    • It implies the difference between pre-test and post-test scores, or between a control group and an experimental group, is a "true" difference.

  • Non-Significant Findings: If scores are dispersed (i.e., people are "all over the place" with many highs, lows, and mediums), researchers cannot draw a definitive conclusion. In this case, the results lack statistical significance and provide less confidence that the Independent Variable (IVIV) caused a change.

Clinical vs. Statistical Significance

  • It is possible to have a result that is statistically significant (probably true) but not clinically or practically significant.

  • Clinical Significance/Effect Size: This refers to the magnitude of the difference.

    • A clinician might find that an intervention caused a true difference, but the size of that change is so small that it does not justify the time or expense of implementation.

    • Therefore, statistical significance measures the validity or "trueness" of a difference, while effect size measures the "bigness" or importance of that difference.

  • "Highly Significant": When a study reports results are "highly significant," it simply means there is a very high probability that the results are true; it does not inherently mean the treatment is highly important or mandatory for clinical use.

Systematic and Unsystematic Variance

  • Systematic Variance: This is a predictable change that occurs in a specific direction.

    • Desirable Systematic Variance: Change that is directly attributed to the implementation of the Independent Variable (IVIV). Researchers want to see the experimental group move systematically in the same direction (e.g., scores increasing or behaviors decreasing) while the control group remains stable.

    • Undesirable Systematic Variance (Extraneous Variables): Changes caused by factors outside the study, such as a subject receiving private speech therapy, the "history effect," or "maturation." These mask the true benefit of the IVIV. Researchers use control groups and rigorous designs to account for these threats to internal validity.

  • Unsystematic Variance (Random Variables): Influences that cause differences between pre-tests and post-tests that cannot be pinpointed or controlled.

    • Sampling Error: Problems or biases in how the sample was drawn from the population. In many fields, like speech pathology, true random selection is nearly impossible.

    • Measurement Error: No instrument or observation is perfect. Errors include scoring mistakes, failing to notice a behavior, or equipment flaws.

    • Individual Differences: Personalogical variables that make people react differently to a study (e.g., some reacting negatively while others progress rapidly).

  • Because unsystematic variance always exists, researchers can never be 100% confident in their results; this necessitates the use of probability values like p<0.05p < 0.05.

The Process of Hypothesis Testing

  • A hypothesis is a statement developed a priori (before the study begins) that predicts the relationship between the IVIV and the Dependent Variable (DVDV).

  • Hypotheses are based on previous studies and existing theories.

  • In hypothesis testing, researchers use data to decide whether to reject or accept the hypothesis.

  • Null Hypothesis (H0H_0):

    • Symbolized as HH with a subscript zero (H0H_0).

    • It states that there will be no relationship or zero change between variables.

    • It serves as the mathematical benchmark against which the study data is compared.

    • It represents "statistical non-significance."

    • Any differences found are attributed to sampling error, measurement error, or individual differences.

  • Alternative Hypothesis (HaH_a):

    • Symbolized as HH with a subscript "a" (HaH_a).

    • It represents the researcher's actual prediction (e.g., "the IVIV will cause a change in the DVDV").

    • It reflects the expected relationship based on theoretical frameworks or prior evidence.

Directionality of Hypotheses and Statistical Power

  • One-Tailed (Directional) Hypothesis: The researcher predicts the specific direction of change (e.g., "Treatment A will improve skills more than Table B").

    • These are more rigorous and provide the study with more statistical power.

  • Two-Tailed (Non-Directional) Hypothesis: The researcher predicts a difference will exist but does not specify the direction (e.g., "There will be a difference between group A and group B").

    • These are less rigorous and provide less statistical power.

Examples of Hypothesis Application

  • Example 1: Stuttering

    • HaH_a: There is a relationship between stressful situations and stuttering.

    • H0H_0: Stressful situations and stuttering are not related.

  • Example 2: Speech Intelligibility

    • HaH_a: Speech therapy will improve speech intelligibility.

    • H0H_0: Speech therapy will cause no change in speech intelligibility. (In practice, researchers expect to reject this null hypothesis based on ample evidence).

  • Example 3: AAC Devices

    • HaH_a: Using an iPad for AAC will improve a child's ability to request.

    • H0H_0: The AAC device will not improve requesting. (Researchers expect to reject this based on research showing these devices help people communicate).

Alpha Levels and Acceptable Risk

  • Level of Significance (pp value or alpha α\alpha): This is the confidence level researchers set to define the trueness of their results.

  • Common Levels:

    • 0.050.05 (95% Confident): The convention in Speech-Language Pathology for non-invasive strategies (e.g., dictionary usage).

    • 0.010.01 (99% Confident): Used for invasive procedures or high-consequence areas like swallowing (dysphagia) or Child Life, where error could lead to pneumonia or significant financial loss.

    • 0.0010.001 (99.9% Confident): The standard in medicine for surgical procedures or prescription drugs where the risk of death is a possibility.

  • The Florida Case Study: Florida adopted a reading series based on a study with an alpha of 0.050.05. Because there was still a 5% chance the study was wrong, the state invested millions in a program that ultimately failed to produce the expected results in the general population. A more rigorous alpha (0.010.01) might have prevented this financial loss.

Possible Outcomes and Errors in Hypothesis Testing

  • There are four possible outcomes when a study is concluded:

    1. Correct Decision: Accepting the H0H_0 when it is actually true (no difference exists).

    2. Correct Decision: Rejecting the H0H_0 when it is false (a difference truly exists).

    3. Type I Error (False Positive\text{False Positive}):

      • Rejecting the H0H_0 when it is actually true.

      • Claiming an intervention works when it actually does not.

      • Considered the most "grievous" error, especially in medicine.

      • Mnemonic: A capital P (for Positive) has one vertical line.

    4. Type II Error (False Negative\text{False Negative}):

      • Accepting the H0H_0 when it is actually false.

      • Claiming an intervention does not work when it actually does.

      • This prevents people from implementing effective treatments.

      • Mnemonic: A capital N (for Negative) has two vertical lines.

  • To truly know if an error occurred, one would need perfect knowledge (e.g., "asking God") or to replicate the study in the entire population.

Sample Sizes and Sensitivity

  • Sample Size (nn): The number of participants in a study.

  • Sample size is typically determined a priori using a power analysis to reach approximately 80%80\% (0.800.80) power.

  • Impact on Power:

    • Larger sample sizes generally increase power and sensitivity, making it easier to detect true differences.

    • Small sample sizes increase the risk of Type II errors (missing a true effect).

Critical Values and Decision Making

  • When running a statistic (ANOVA, t-test), a number is produced and compared against a Critical Value.

  • The range of the critical value determines whether you reject or accept the H0H_0.

  • If the statistic falls outside the critical value range, you reject the H0H_0 and celebrate finding significance.

  • This range fluctuates based on alpha levels (0.050.05 vs 0.010.01) and whether the focus is restricted (99.9%99.9\%) or broad (95%95\%).

Selection of Statistical Tests (Parametric vs. Non-Parametric)

  • Parametric Tests:

    • Preferred because they are more powerful.

    • Requirement: Data must be normally distributed (Bell curve).

    • Requirement: Measuring at the Ratio or Interval level.

    • Examples: t-test, ANOVA (FF), ANCOVA, MANCOVA.

  • Non-Parametric Tests:

    • Used when data is not normally distributed, has a small sample size, or contains many outliers (high standard deviation).

    • Used for Ordinal level data.

    • These are less sensitive and provide less confidence in the accuracy of findings.

    • Examples: Chi-square (χ2\chi^2), Wilcoxon, Mann-Whitney U.

  • Using a parametric test for ordinal data is a common mistake that reduces the reliability of research findings.