Advanced Statistical Methods: Assumptions, Power Analysis, and Hypothesis Testing

Parametric Test Assumptions and Normality

  • Core Assumptions of Parametric Statistical Tests:

    • Normal Distribution: Parametric tests rely on an underlying normal distribution. The specific entity that must be normally distributed depends on the analytical model:

    • For tt-tests and Analysis of Variance (ANOVA), the sampling distribution must be normally distributed.

    • For General Linear Models (GLMs) such as Multiple Regression, the residuals/errors of the model must be normally distributed.

    • Homogeneity of Variance: The variance of the population outcome variable must be equal across all comparison groups or levels of an independent variable.

    • Continuous Data: The dependent variable must be measured on a continuous scale, specifically an interval or ratio scale.

    • Independence of Observations: Individual data points and residuals must be independent of one another (e.g., examination scores of individual students must not be influenced by looking at another student's work).

  • The Sampling Distribution and the Central Limit Theorem:

    • Direct access to the theoretical sampling distribution is impossible in empirical research. Normality of the sampling distribution is inferred by evaluating the normality of individual variables and all linear combinations of those variables.

    • If individual variables and their linear combinations are normally distributed, the sampling distribution and regression residuals are mathematically guaranteed to be normally distributed.

    • Central Limit Theorem Principles:

    • Approximately Normal Data: If sample data are roughly normal, the sampling distribution will also be normal.

    • Non-Normal Sample Data with Small Sample Sizes (N≤40N \le 40): Severe non-normality poses a substantial threat to test validity, requiring cautious interpretation or data transformations.

    • Non-Normal Sample Data with Moderate to Large Sample Sizes (N>40N > 40): According to the Central Limit Theorem, when sample sizes reach roughly N=40N = 40 or more, the sampling distribution approaches normality regardless of the population shape. Test statistics become highly robust to non-normality.

    • Heavy-Tailed Distributions: For distributions with extremely heavy tails, significantly larger samples (e.g., N=160N = 160) may be required for the sampling distribution to reach normality.

    • Population Non-Normality Nuance: Statistical inference becomes progressively less robust as population non-normality increases. Consequently, normality should still be assessed in samples over N=40N = 40, though the practical concern diminishes as sample size grows.

  • Assessing Normality: Graphical Displays:

    • Histograms: Provide a visual depiction of the frequency distribution. A normal distribution exhibits a symmetric, bell-shaped curve centered over the mean.      

      Positively skewed vs normal histograms
    • P-P Plots (Probability-Probability Plots) and Q-Q Plots (Quantile-Quantile Plots): Compare the empirical cumulative distribution of the sample data against a theoretical normal distribution. If the data are normally distributed, the plotted points closely follow the diagonal reference line. Deviations from the diagonal line indicate skewness or kurtosis.

  

P-P Plot showing observed vs expected cumulative probability
  • Assessing Normality: Skewness and Kurtosis:

    • Skewness: Measures the asymmetry of the data distribution around its mean.

    • Positive skew indicates a tail extending toward higher positive values.

    • Negative skew indicates a tail extending toward lower negative values.

    • Kurtosis: Measures the tail-heaviness and peakedness of a distribution relative to a normal distribution (DeCarlo, 1997):

    • Leptokurtic (Positive Kurtosis): Characterized by a sharp, thin peak and heavy tails relative to the normal distribution.

    • Platykurtic (Negative Kurtosis): Characterized by a broad, flat peak ("blob") and light tails relative to the normal distribution.      

      Kurtosis shapes: Leptokurtic, Normal, Platykurtic
    • Standardized zz--score Calculation for Skewness and Kurtosis:

    • In a perfectly normal population, skewness and kurtosis statistics equal zero.

    • Skewness and kurtosis statistics are converted to standardized zz-scores by subtracting the expected population value (00 under the null hypothesis) and dividing by their respective standard errors:

      zskewness=Skewness Statistic−0SEskewnessz_{\text{skewness}} = \frac{\text{Skewness Statistic} - 0}{\text{SE}_{\text{skewness}}}

      zkurtosis=Kurtosis Statistic−0SEkurtosisz_{\text{kurtosis}} = \frac{\text{Kurtosis Statistic} - 0}{\text{SE}_{\text{kurtosis}}}

- **Worked Calculation Example (Day 1 Hygiene Data):**
  - Skewness statistic = −.004-.004
  - Standard error of skewness = .086.086

      zskewness=−.004.086=−.047z_{\text{skewness}} = \frac{-.004}{.086} = -.047

- **Interpretation Thresholds:**
  - For small-to-moderate samples, calculated zz-scores are evaluated against a continuum bounded by ±1.96\pm 1.96 (for α=.05\alpha = .05), ±2.58\pm 2.58 (for α=.01\alpha = .01), or ±3.29\pm 3.29 (for α=.001\alpha = .001).
  - Scores falling well within these boundaries (close to zero) indicate no significant deviation from normality.
  • Formal Normality Tests:

    • Kolmogorov-Smirnov Test (with Lilliefors Significance Correction) and Shapiro-Wilk Test mathematically compare the sample distribution to a normal distribution with the same mean and standard deviation.

    • Null Hypothesis (H0H_0): The sample data do not differ significantly from a normal distribution.

    • Significant Result (p<.05p < .05 or p<.001p < .001): Reject H0H_0; the data deviate significantly from a normal distribution.

    • Non-Significant Result (p>αp > \alpha): Fail to reject H0H_0; normality can be assumed.

    • Sample Size Sensitivity Limitations:

    • Formal tests are extremely sensitive to sample size.

    • In large samples (N>150N > 150), trivial, non-meaningful departures from normality produce statistically significant test results (p<.05p < .05).

    • In small samples, large, substantive departures from normality frequently fail to reach statistical significance due to low power.

    • Best Practice: Use formal tests sparingly, apply a conservative alpha threshold (α=.001\alpha = .001), and always evaluate formal test output in conjunction with visual inspection of histograms.

Homogeneity of Variance

  • Definition and Importance:

    • Homogeneity of variance assumes that the variance of the dependent variable is constant across all conditions or comparison groups being evaluated (\sigma_1^2 = \sigma_2^2 = \dots = \n\sigma_k^2).

    • Heterogeneity of variance occurs when group variances differ substantially (e.g., low variance in one group versus high variance in another).

  

Heterogeneity of Variance plot across concert venues
  • Levene's Test:

    • Evaluates the null hypothesis that population variances across groups are equal.

    • Significant Levene's Test (p<.05p < .05): Group variances are significantly different; the assumption of homogeneity of variance is violated.

    • Non-Significant Levene's Test (p>.05p > .05): Group variances do not differ significantly; the assumption is met.

    • Sample Size Bias: Similar to formal normality tests, Levene's test is highly sensitive to sample size and can flag minor, trivial variance differences as significant in large samples.

  • Hartley's Fmax⁡F_{\max} Test (Variance Ratio):

    • Serves as a complementary or alternative metric to Levene's test, particularly in large samples.

    • Calculated as the ratio of the largest group variance to the smallest group variance:

    Variance Ratio (VR)=VariancelargestVariancesmallest\text{Variance Ratio (VR)} = \frac{\text{Variance}_{\text{largest}}}{\text{Variance}_{\text{smallest}}}

  • Decision Rule: If the Variance Ratio is less than 22 (VR<2\text{VR} < 2), homogeneity of variance can be safely assumed regardless of sample size.

Data Transformations

  • Principles and Application Rules Across Designs:

    • Data transformations apply a mathematical algorithm uniformly to every score on a specific variable to reduce skewness and bring the distribution closer to normality.

    • Transforming Variables by Research Method:

    • Correlations and Multiple Regression: It is completely valid to transform only the continuous variable(s) exhibiting skewness while leaving unskewed variables in their original raw units.

    • Between-Subjects Designs (tt-tests, ANOVA): The entire variable must be transformed as a whole, applying the transformation across all comparison groups simultaneously.

    • Repeated Measures Designs (Dependent tt-tests, Repeated Measures ANOVA): All continuous variables across all measured time points or conditions must be transformed in the exact same mathematical way.

  • Mathematical Impact on Associations vs. Means:

    • Preservation of Associations: Transforming a variable alters the scale of measurement monotonically without substantially changing the relative rank-order strength of associations between variables.

    • Example: The correlation between Time 1 and Time 2 raw scores (r=.460,p=.084r = .460, p = .084) remains essentially identical when evaluating the correlation between Time 1 raw scores and the Square Root of Time 2 scores (r=.462,p=.083r = .462, p = .083).

    • Disruption of Mean Differences: Because transformations alter raw scale distances non-linearly, they directly alter raw group means and standard deviations.

    • Example: Comparing raw Time 1 (M=4.80M = 4.80) and raw Time 2 (M=5.13M = 5.13) via a paired tt-test yields a non-significant mean difference (t(14)=−0.649,p=.527t(14) = -0.649, p = .527). However, comparing raw Time 1 (M=4.80M = 4.80) against Square Root Transformed Time 2 (M=2.23M = 2.23) yields a highly significant mean difference (t(14)=6.000,p<.001t(14) = 6.000, p < .001).

    • This drastic artificial mean shift demonstrates why repeated measures variables must all undergo identical transformations prior to mean comparisons.

  • Mathematical Types of Transformations:

    • Square Root Transformation (Xi\sqrt{X_i}):

    • Applied to correct mild positive skewness.

    • Compresses the upper tail of the distribution by taking the square root of each individual score.

    • Logarithmic Transformation (log⁡10(Xi)\log_{10}(X_i) or ln⁡(Xi)\ln(X_i)):

    • Applied to correct moderate-to-severe positive skewness.

    • Exerts a stronger compressing effect on large values than the square root transformation.

    • Reciprocal / Inverse Transformation (1Xi\frac{1}{X_i}):

    • Applied to correct severe positive skewness.

    • Divides 11 by each raw score. Because this operation inherently reverses score order (large raw values become tiny fractions), the variable can be reflected prior to transformation to preserve scale direction:

      Reciprocaldirectional=1Xhighest−Xi\text{Reciprocal}_{\text{directional}} = \frac{1}{X_{\text{highest}} - X_i}

  • Caution: Analysts must re-check skewness statistics and histograms post-transformation to confirm that the mathematical operation did not overcorrect and make the skewness worse.

    • Transforming Negatively Skewed Variables:

  • Standard transformation equations (Square Root, Log, Reciprocal) are mathematically designed to compress positive skews. To transform a negatively skewed variable:

    1. Reflect (Flip) the Variable: Convert the negative skew into a positive skew by creating a new score (XflippedX_{\text{flipped}}) for each participant:

       Xflipped=(Xmaximum+1)−XiX_{\text{flipped}} = (X_{\text{maximum}} + 1) - X_i

2. **Apply Transformation:** Execute the selected standard transformation (Square Root, Log, or Reciprocal) on XflippedX_{\text{flipped}}.
3. **Reflect (Flip) Back:** Reverse the transformed variable so that higher numerical values once again correspond to higher quantities of the original construct:

       Xfinal=(Xtransformed_max+1)−XtransformedX_{\text{final}} = (X_{\text{transformed\_max}} + 1) - X_{\text{transformed}}

Consequences of Violated Assumptions

  • Sample vs. Population Inference:

    • Violating parametric test assumptions (normality, homogeneity of variance, or independence of observations) does not invalidate descriptive analysis of the immediate sample data.

    • If assumptions are violated—and assuming there are no unhandled extreme outliers or influential cases—valid, descriptive conclusions can still be drawn about the specific sample tested.

    • However, assumption violations prevent valid statistical inference and generalizability beyond the sample to the broader target population.

Statistical Power and Decision Making

  • Definition and Mathematical Formulation:

    • Statistical Power is the probability that a statistical test will correctly reject a false null hypothesis.

    • It represents the probability of detecting a statistically significant effect in a sample, given that a true effect actually exists in the population sampled.

    • Mathematically, statistical power is expressed as:

    Power=1−β\text{Power} = 1 - \beta

    where β\beta is the probability of committing a Type II error.

  • The Hypothesis Testing Decision Matrix:

  

Statistical Decision Matrix
  • Null Hypothesis True in Population (H0H_0 True, No Real Effect):

    • Decision: Reject H0H_0 (H1H_1): Type I Error (Probability = α\alpha). False Positive.

    • Decision: Fail to Reject H0H_0: Correct Decision (Probability = 1−α1 - \alpha).

  • Null Hypothesis False in Population (H0H_0 False, Real Effect Exists):

    • Decision: Reject H0H_0 (H1H_1 ): Correct Decision / Statistical Power (Probability = 1−β1 - \beta).

    • Decision: Fail to Reject H0H_0: Type II Error (Probability = β\beta). False Negative.

    • Sampling Distributions: Null vs. Alternative Population:

  • Statistical inference compares two theoretical sampling distributions along a decision axis:

    1. The sampling distribution of the null population (H0H_0), centered at an effect size of zero.

    2. The sampling distribution of the alternative population (HaH_a), centered at the true population effect size.

  

Sampling distributions of null and alternative hypothesis
  • The decision axis is established by the critical alpha threshold (α\alpha). The proportion of the alternative sampling distribution that falls beyond this decision threshold represents Statistical Power (1−β1 - \beta).

    • Factors Influencing Statistical Power:

  • Sample Size (NN):

    • As sample size increases, standard error decreases, causing the sampling distributions to narrow.

    • Narrower distributions reduce the overlap between H0H_0 and HaH_a, directly increasing statistical power.

  • Effect Size:

    • Effect size quantifies the degree to which a population phenomenon differs from zero.

    • Larger population effect sizes shift the alternative sampling distribution farther away from the null distribution along the decision axis, increasing statistical power.

  • Critical Alpha Level (α\alpha):

    • Increasing α\alpha (e.g., from .05.05 to .10.10) shifts the decision line toward the center of the null distribution, increasing the area under the alternative curve that exceeds the threshold (increasing power).

    • Crucial Caveat: Relaxing α\alpha simultaneously increases the Type I error rate. Thus, adjusting α\alpha upwards is not an acceptable strategy for increasing power in empirical research.

  

Power vs Sample Size for different correlation coefficients
  • Current State of Power in Psychological Research:

    • Consequences of Low Power: High probability of Type II errors (failing to discover true phenomena) and reduced replicability of published findings.

    • Consequences of Excessively High Power: Over-spending time, funding, and participant resources to detect trivial effects.

    • Empirical Reality: Psychological research has historically suffered from low statistical power. A large-scale empirical audit of 8,0008{,}000 published psychology articles (Stanley et al., 2018) revealed:

    • The median statistical power across the literature was only 36%36\%.

    • Only 8%8\% of published studies achieved the standard benchmark of 80%80\% power (0.800.80).

Determining Sample Size and Conducting Power Analyses

  • Three Essential Ingredients:

    • To compute required sample size via an a priori power analysis, three parameters must be specified:

    1. Target Power Level (1−β1 - \beta): Conventionally set to .80.80 (80%80\% chance of detecting a true effect). Higher targets (.90.90 or .95.95) are recommended when Type II errors carry high practical or clinical costs.

    2. Critical Alpha Level (α\alpha): Conventionally set to .05.05. More conservative thresholds (e.g., α=.005\alpha = .005; Johnson, 2013) may be selected to minimize Type I errors.

    3. Expected Population Effect Size: The anticipated magnitude of the effect in the population.

  • Methods for Estimating Effect Size:

    • Literature Review: Derived from published effect sizes in closely related studies. Warning: Published effect sizes are systematically inflated due to publication bias (the "file drawer" problem). Empirical estimates should be adjusted downward slightly.

    • Smallest Effect of Theoretical Interest: Determined in consultation with subject-matter experts as the minimum effect threshold that carries theoretical or practical significance.

    • Textbook Rules of Thumb: Used as a last resort when prior literature is absent. For example, Tabachnick and Fidell recommend a minimum sample size of N≥104+mN \ge 104 + m for multiple regression (where mm is the number of predictors).

  • G*Power Software Guidance and Navigation Rules:

    • G*Power is a free, comprehensive software package dedicated to statistical power analysis.      

      G*Power interface for correlation power analysis
    • Critical Navigation Rule: When navigating G*Power, select statistical tests using the top main menu bar (Tests →\rightarrow Correlation and Regression, Means, Proportions, etc.). Relying on bottom drop-down menus frequently leads to incorrect test selection.

  • G*Power Analysis Examples:

    • Example 1: Correlation (Bivariate Normal Model)

    • Parameters: Tail = Two, Target Power (1−β1 - \beta) = .80.80, α=.05\alpha = .05, Expected Population Correlation (ρ\rho) = .30.30.

    • Output: Required Total Sample Size N=84N = 84.

    • Example 2: Factorial ANOVA (2×22 \times 2 Design)

    • Parameters: 2×22 \times 2 design yields Number of Groups = 44. For an interaction between two 2-level IVs, numerator degrees of freedom df=(2−1)×(2−1)=1df = (2 - 1) \times (2 - 1) = 1.

    • Strategy: When testing multiple hypotheses (main effects and interaction) with varying expected effect sizes (e.g., f=.10,.60,.05f = .10, .60, .05), conduct the power calculation using the weakest expected effect size (f=.05f = .05) to guarantee adequate power across all tests.

    • Example 3: Multiple Regression I (Overall Model Power)

    • Scenario: Evaluates power to detect an overall regression model with 33 predictors explaining 25%25\% of DV variance (R2=.25R^2 = .25).

    • Effect Size Conversion: G*Power uses Cohen's f2f^2 as the input parameter for regression:

      f2=R21−R2=.251−.25=.25.75=0.3333333f^2 = \frac{R^2}{1 - R^2} = \frac{.25}{1 - .25} = \frac{.25}{.75} = 0.3333333

- *Parameters:* Effect size f2=0.3333333f^2 = 0.3333333, α=.05\alpha = .05, Power = .80.80, Number of predictors = 33
- *Output:* Required Total Sample Size N=37N = 37
  • Example 4: Multiple Regression II (Individual Predictor Power)

    • Scenario: Evaluates power for the unique contribution of a single predictor in a 33-predictor model, where the individual predictor uniquely accounts for 5%5\% of variance (Partial R2=.05R^2 = .05).

    • Effect Size Conversion: Converts Partial R2R^2 to f2f^2:

      f2=Partial R21−R2=.051−.05=0.0526316f^2 = \frac{\text{Partial } R^2}{1 - R^2} = \frac{.05}{1 - .05} = 0.0526316

- *Parameters:* Test type = Linear multiple regression: Fixed model: R2R^2 increase; Effect size f2=0.0526316f^2 = 0.0526316, α=.05\alpha = .05, Power = .80.80, Number of tested predictors = 11, Total predictors = 33
- *Output:* Required Total Sample Size N=152N = 152
  • Writing Up a Power Analysis:

    • A complete, professional power analysis write-up must document four components:

    1. The specific statistical test being performed.

    2. The estimated population effect size (including the specific metric, e.g., r,d,f2r, d, f^2) along with an explicit justification.

    3. The target statistical power level (1−β1 - \beta).

    4. The critical alpha level (α\alpha) and tail specification.

    • Template Write-Up Example:     > "A power analysis was conducted using G*Power 3 (Faul et al., 2007) to determine the required sample size for a point-biserial correlation. Based on prior research by Dio and Osborne (1985), an effect size of r=.30r = .30 is anticipated. Assuming a directional hypothesis with a critical α=.05\alpha = .05 and a target power of .80.80, the analysis indicates that a minimum of N=64N = 64 participants is required."

Integration of Key Statistical Concepts

  • Statistical Significance vs. Practical Significance:

    • Statistical Significance (p<αp < \alpha): Indicates that an observed sample effect is unlikely to have occurred purely by chance under the null hypothesis.

    • Practical Significance: Evaluates whether the magnitude of an effect is meaningful in real-world applications.

    • Core Relationship Equation:

    Statistical Significance≈Effect Size×Sample Size\text{Statistical Significance} \approx \text{Effect Size} \times \text{Sample Size}

  • Sample Size Distortion: In massive samples, tiny, practically trivial differences achieve statistical significance.

    • Illustrative Case Study: In a study of N=10,000N = 10{,}000 adults evaluating statistics enjoyment on a 10-point scale (1="can’t stand it"1 = \text{"can't stand it"}, 10="better than sex"10 = \text{"better than sex"}), males (M=3.21M = 3.21) scored significantly higher than females (M=3.16M = 3.16, p<.05p < .05). Although statistically significant, a mean difference of 0.050.05 on a 10-point scale is practically meaningless (and both means are disturbingly low).

    • Anatomy and Ingredients of Test Statistics:

  • Across all inferential statistical tests (t,F,rt, F, r), test statistics follow a universal mathematical structure:

    Test Statistic=Model EffectError=Central Tendency / Variance Explained by ModelVariability / Estimated Population Variance / Unexplained Variance\text{Test Statistic} = \frac{\text{Model Effect}}{\text{Error}} = \frac{\text{Central Tendency / Variance Explained by Model}}{\text{Variability / Estimated Population Variance / Unexplained Variance}}

  • Resampling Framework of Null Hypothesis Significance Testing:      

    Resampling framework for null hypothesis testing
    • Conceptual Resampling Procedure:

    1. Generate Null Distribution: Repeatedly sample 100100 times from a population where no effect exists (e.g., two groups drawn entirely from Scrabble players) using the study's exact sample size, calculating the test statistic each time. This creates the empirical sampling distribution under the null hypothesis.

    2. Obtain Sample Test Statistic: Conduct the actual study once using the target population (e.g., comparing Scrabble players against heavy metal concert attendees) and calculate the sample test statistic.

    3. Derive pp-value: The pp-value represents the exact proportion of times that the null sampling distribution produced a test statistic equal to or more extreme than the observed sample statistic.

      • Example: A calculated p=.02p = .02 means that in only 22 out of 100100 null resampling runs did a test statistic equal or exceed the sample statistic.

    4. Evaluate Decision Rule: Compare pp against α=.05\alpha = .05 (which splits into $2.5\% tails on both ends for a two-tailed test). If the sample test statistic falls in the extreme $2.5\% tail region, conclude that the sample was unlikely drawn from the null population and reject H0H_0

  • Confidence Intervals and Their Relation to Hypothesis Testing:      

    Confidence intervals across repeated samples
    • A $95\% Confidence Interval (CI) represents the range of values within which the true population parameter will fall across $95\% of all possible samples drawn from that population.

    • Decision Alignment with Null Hypothesis Testing:

    • Real Population Effect Exists (H0H_0 False): Individual sample $95\% CIs drawn from the alternative population will **not** span across or overlap the zero point of the null distribution.\n - **No Population Effect Exists (H_0 True):** Individual sample $95\% CIs drawn from the population will routinely span across and overlap the zero point of the null distribution.