Advanced Statistical Methods: Assumptions, Power Analysis, and Hypothesis Testing
Parametric Test Assumptions and Normality
Core Assumptions of Parametric Statistical Tests:
Normal Distribution: Parametric tests rely on an underlying normal distribution. The specific entity that must be normally distributed depends on the analytical model:
For -tests and Analysis of Variance (ANOVA), the sampling distribution must be normally distributed.
For General Linear Models (GLMs) such as Multiple Regression, the residuals/errors of the model must be normally distributed.
Homogeneity of Variance: The variance of the population outcome variable must be equal across all comparison groups or levels of an independent variable.
Continuous Data: The dependent variable must be measured on a continuous scale, specifically an interval or ratio scale.
Independence of Observations: Individual data points and residuals must be independent of one another (e.g., examination scores of individual students must not be influenced by looking at another student's work).
The Sampling Distribution and the Central Limit Theorem:
Direct access to the theoretical sampling distribution is impossible in empirical research. Normality of the sampling distribution is inferred by evaluating the normality of individual variables and all linear combinations of those variables.
If individual variables and their linear combinations are normally distributed, the sampling distribution and regression residuals are mathematically guaranteed to be normally distributed.
Central Limit Theorem Principles:
Approximately Normal Data: If sample data are roughly normal, the sampling distribution will also be normal.
Non-Normal Sample Data with Small Sample Sizes (): Severe non-normality poses a substantial threat to test validity, requiring cautious interpretation or data transformations.
Non-Normal Sample Data with Moderate to Large Sample Sizes (): According to the Central Limit Theorem, when sample sizes reach roughly or more, the sampling distribution approaches normality regardless of the population shape. Test statistics become highly robust to non-normality.
Heavy-Tailed Distributions: For distributions with extremely heavy tails, significantly larger samples (e.g., ) may be required for the sampling distribution to reach normality.
Population Non-Normality Nuance: Statistical inference becomes progressively less robust as population non-normality increases. Consequently, normality should still be assessed in samples over , though the practical concern diminishes as sample size grows.
Assessing Normality: Graphical Displays:
Histograms: Provide a visual depiction of the frequency distribution. A normal distribution exhibits a symmetric, bell-shaped curve centered over the mean.

P-P Plots (Probability-Probability Plots) and Q-Q Plots (Quantile-Quantile Plots): Compare the empirical cumulative distribution of the sample data against a theoretical normal distribution. If the data are normally distributed, the plotted points closely follow the diagonal reference line. Deviations from the diagonal line indicate skewness or kurtosis.

Assessing Normality: Skewness and Kurtosis:
Skewness: Measures the asymmetry of the data distribution around its mean.
Positive skew indicates a tail extending toward higher positive values.
Negative skew indicates a tail extending toward lower negative values.
Kurtosis: Measures the tail-heaviness and peakedness of a distribution relative to a normal distribution (DeCarlo, 1997):
Leptokurtic (Positive Kurtosis): Characterized by a sharp, thin peak and heavy tails relative to the normal distribution.
Platykurtic (Negative Kurtosis): Characterized by a broad, flat peak ("blob") and light tails relative to the normal distribution.

Standardized --score Calculation for Skewness and Kurtosis:
In a perfectly normal population, skewness and kurtosis statistics equal zero.
Skewness and kurtosis statistics are converted to standardized -scores by subtracting the expected population value ( under the null hypothesis) and dividing by their respective standard errors:
- **Worked Calculation Example (Day 1 Hygiene Data):**
- Skewness statistic =
- Standard error of skewness =
- **Interpretation Thresholds:**
- For small-to-moderate samples, calculated -scores are evaluated against a continuum bounded by (for ), (for ), or (for ).
- Scores falling well within these boundaries (close to zero) indicate no significant deviation from normality.
Formal Normality Tests:
Kolmogorov-Smirnov Test (with Lilliefors Significance Correction) and Shapiro-Wilk Test mathematically compare the sample distribution to a normal distribution with the same mean and standard deviation.
Null Hypothesis (): The sample data do not differ significantly from a normal distribution.
Significant Result ( or ): Reject ; the data deviate significantly from a normal distribution.
Non-Significant Result (): Fail to reject ; normality can be assumed.
Sample Size Sensitivity Limitations:
Formal tests are extremely sensitive to sample size.
In large samples (), trivial, non-meaningful departures from normality produce statistically significant test results ().
In small samples, large, substantive departures from normality frequently fail to reach statistical significance due to low power.
Best Practice: Use formal tests sparingly, apply a conservative alpha threshold (), and always evaluate formal test output in conjunction with visual inspection of histograms.
Homogeneity of Variance
Definition and Importance:
Homogeneity of variance assumes that the variance of the dependent variable is constant across all conditions or comparison groups being evaluated (\sigma_1^2 = \sigma_2^2 = \dots = \n\sigma_k^2).
Heterogeneity of variance occurs when group variances differ substantially (e.g., low variance in one group versus high variance in another).

Levene's Test:
Evaluates the null hypothesis that population variances across groups are equal.
Significant Levene's Test (): Group variances are significantly different; the assumption of homogeneity of variance is violated.
Non-Significant Levene's Test (): Group variances do not differ significantly; the assumption is met.
Sample Size Bias: Similar to formal normality tests, Levene's test is highly sensitive to sample size and can flag minor, trivial variance differences as significant in large samples.
Hartley's Test (Variance Ratio):
Serves as a complementary or alternative metric to Levene's test, particularly in large samples.
Calculated as the ratio of the largest group variance to the smallest group variance:
Decision Rule: If the Variance Ratio is less than (), homogeneity of variance can be safely assumed regardless of sample size.
Data Transformations
Principles and Application Rules Across Designs:
Data transformations apply a mathematical algorithm uniformly to every score on a specific variable to reduce skewness and bring the distribution closer to normality.
Transforming Variables by Research Method:
Correlations and Multiple Regression: It is completely valid to transform only the continuous variable(s) exhibiting skewness while leaving unskewed variables in their original raw units.
Between-Subjects Designs (-tests, ANOVA): The entire variable must be transformed as a whole, applying the transformation across all comparison groups simultaneously.
Repeated Measures Designs (Dependent -tests, Repeated Measures ANOVA): All continuous variables across all measured time points or conditions must be transformed in the exact same mathematical way.
Mathematical Impact on Associations vs. Means:
Preservation of Associations: Transforming a variable alters the scale of measurement monotonically without substantially changing the relative rank-order strength of associations between variables.
Example: The correlation between Time 1 and Time 2 raw scores () remains essentially identical when evaluating the correlation between Time 1 raw scores and the Square Root of Time 2 scores ().
Disruption of Mean Differences: Because transformations alter raw scale distances non-linearly, they directly alter raw group means and standard deviations.
Example: Comparing raw Time 1 () and raw Time 2 () via a paired -test yields a non-significant mean difference (). However, comparing raw Time 1 () against Square Root Transformed Time 2 () yields a highly significant mean difference ().
This drastic artificial mean shift demonstrates why repeated measures variables must all undergo identical transformations prior to mean comparisons.
Mathematical Types of Transformations:
Square Root Transformation ():
Applied to correct mild positive skewness.
Compresses the upper tail of the distribution by taking the square root of each individual score.
Logarithmic Transformation ( or ):
Applied to correct moderate-to-severe positive skewness.
Exerts a stronger compressing effect on large values than the square root transformation.
Reciprocal / Inverse Transformation ():
Applied to correct severe positive skewness.
Divides by each raw score. Because this operation inherently reverses score order (large raw values become tiny fractions), the variable can be reflected prior to transformation to preserve scale direction:
Caution: Analysts must re-check skewness statistics and histograms post-transformation to confirm that the mathematical operation did not overcorrect and make the skewness worse.
Transforming Negatively Skewed Variables:
Standard transformation equations (Square Root, Log, Reciprocal) are mathematically designed to compress positive skews. To transform a negatively skewed variable:
Reflect (Flip) the Variable: Convert the negative skew into a positive skew by creating a new score () for each participant:
2. **Apply Transformation:** Execute the selected standard transformation (Square Root, Log, or Reciprocal) on .
3. **Reflect (Flip) Back:** Reverse the transformed variable so that higher numerical values once again correspond to higher quantities of the original construct:
Consequences of Violated Assumptions
Sample vs. Population Inference:
Violating parametric test assumptions (normality, homogeneity of variance, or independence of observations) does not invalidate descriptive analysis of the immediate sample data.
If assumptions are violated—and assuming there are no unhandled extreme outliers or influential cases—valid, descriptive conclusions can still be drawn about the specific sample tested.
However, assumption violations prevent valid statistical inference and generalizability beyond the sample to the broader target population.
Statistical Power and Decision Making
Definition and Mathematical Formulation:
Statistical Power is the probability that a statistical test will correctly reject a false null hypothesis.
It represents the probability of detecting a statistically significant effect in a sample, given that a true effect actually exists in the population sampled.
Mathematically, statistical power is expressed as:
where is the probability of committing a Type II error.
The Hypothesis Testing Decision Matrix:

Null Hypothesis True in Population ( True, No Real Effect):
Decision: Reject (): Type I Error (Probability = ). False Positive.
Decision: Fail to Reject : Correct Decision (Probability = ).
Null Hypothesis False in Population ( False, Real Effect Exists):
Decision: Reject ( ): Correct Decision / Statistical Power (Probability = ).
Decision: Fail to Reject : Type II Error (Probability = ). False Negative.
Sampling Distributions: Null vs. Alternative Population:
Statistical inference compares two theoretical sampling distributions along a decision axis:
The sampling distribution of the null population (), centered at an effect size of zero.
The sampling distribution of the alternative population (), centered at the true population effect size.

The decision axis is established by the critical alpha threshold (). The proportion of the alternative sampling distribution that falls beyond this decision threshold represents Statistical Power ().
Factors Influencing Statistical Power:
Sample Size ():
As sample size increases, standard error decreases, causing the sampling distributions to narrow.
Narrower distributions reduce the overlap between and , directly increasing statistical power.
Effect Size:
Effect size quantifies the degree to which a population phenomenon differs from zero.
Larger population effect sizes shift the alternative sampling distribution farther away from the null distribution along the decision axis, increasing statistical power.
Critical Alpha Level ():
Increasing (e.g., from to ) shifts the decision line toward the center of the null distribution, increasing the area under the alternative curve that exceeds the threshold (increasing power).
Crucial Caveat: Relaxing simultaneously increases the Type I error rate. Thus, adjusting upwards is not an acceptable strategy for increasing power in empirical research.

Current State of Power in Psychological Research:
Consequences of Low Power: High probability of Type II errors (failing to discover true phenomena) and reduced replicability of published findings.
Consequences of Excessively High Power: Over-spending time, funding, and participant resources to detect trivial effects.
Empirical Reality: Psychological research has historically suffered from low statistical power. A large-scale empirical audit of published psychology articles (Stanley et al., 2018) revealed:
The median statistical power across the literature was only .
Only of published studies achieved the standard benchmark of power ().
Determining Sample Size and Conducting Power Analyses
Three Essential Ingredients:
To compute required sample size via an a priori power analysis, three parameters must be specified:
Target Power Level (): Conventionally set to ( chance of detecting a true effect). Higher targets ( or ) are recommended when Type II errors carry high practical or clinical costs.
Critical Alpha Level (): Conventionally set to . More conservative thresholds (e.g., ; Johnson, 2013) may be selected to minimize Type I errors.
Expected Population Effect Size: The anticipated magnitude of the effect in the population.
Methods for Estimating Effect Size:
Literature Review: Derived from published effect sizes in closely related studies. Warning: Published effect sizes are systematically inflated due to publication bias (the "file drawer" problem). Empirical estimates should be adjusted downward slightly.
Smallest Effect of Theoretical Interest: Determined in consultation with subject-matter experts as the minimum effect threshold that carries theoretical or practical significance.
Textbook Rules of Thumb: Used as a last resort when prior literature is absent. For example, Tabachnick and Fidell recommend a minimum sample size of for multiple regression (where is the number of predictors).
G*Power Software Guidance and Navigation Rules:
G*Power is a free, comprehensive software package dedicated to statistical power analysis.

Critical Navigation Rule: When navigating G*Power, select statistical tests using the top main menu bar (
TestsCorrelation and Regression,Means,Proportions, etc.). Relying on bottom drop-down menus frequently leads to incorrect test selection.
G*Power Analysis Examples:
Example 1: Correlation (Bivariate Normal Model)
Parameters: Tail = Two, Target Power () = , , Expected Population Correlation () = .
Output: Required Total Sample Size .
Example 2: Factorial ANOVA ( Design)
Parameters: design yields Number of Groups = . For an interaction between two 2-level IVs, numerator degrees of freedom .
Strategy: When testing multiple hypotheses (main effects and interaction) with varying expected effect sizes (e.g., ), conduct the power calculation using the weakest expected effect size () to guarantee adequate power across all tests.
Example 3: Multiple Regression I (Overall Model Power)
Scenario: Evaluates power to detect an overall regression model with predictors explaining of DV variance ().
Effect Size Conversion: G*Power uses Cohen's as the input parameter for regression:
- *Parameters:* Effect size , , Power = , Number of predictors =
- *Output:* Required Total Sample Size Example 4: Multiple Regression II (Individual Predictor Power)
Scenario: Evaluates power for the unique contribution of a single predictor in a -predictor model, where the individual predictor uniquely accounts for of variance (Partial ).
Effect Size Conversion: Converts Partial to :
- *Parameters:* Test type = Linear multiple regression: Fixed model: increase; Effect size , , Power = , Number of tested predictors = , Total predictors =
- *Output:* Required Total Sample Size Writing Up a Power Analysis:
A complete, professional power analysis write-up must document four components:
The specific statistical test being performed.
The estimated population effect size (including the specific metric, e.g., ) along with an explicit justification.
The target statistical power level ().
The critical alpha level () and tail specification.
Template Write-Up Example: > "A power analysis was conducted using G*Power 3 (Faul et al., 2007) to determine the required sample size for a point-biserial correlation. Based on prior research by Dio and Osborne (1985), an effect size of is anticipated. Assuming a directional hypothesis with a critical and a target power of , the analysis indicates that a minimum of participants is required."
Integration of Key Statistical Concepts
Statistical Significance vs. Practical Significance:
Statistical Significance (): Indicates that an observed sample effect is unlikely to have occurred purely by chance under the null hypothesis.
Practical Significance: Evaluates whether the magnitude of an effect is meaningful in real-world applications.
Core Relationship Equation:
Sample Size Distortion: In massive samples, tiny, practically trivial differences achieve statistical significance.
Illustrative Case Study: In a study of adults evaluating statistics enjoyment on a 10-point scale (, ), males () scored significantly higher than females (, ). Although statistically significant, a mean difference of on a 10-point scale is practically meaningless (and both means are disturbingly low).
Anatomy and Ingredients of Test Statistics:
Across all inferential statistical tests (), test statistics follow a universal mathematical structure:
Resampling Framework of Null Hypothesis Significance Testing:

Conceptual Resampling Procedure:
Generate Null Distribution: Repeatedly sample times from a population where no effect exists (e.g., two groups drawn entirely from Scrabble players) using the study's exact sample size, calculating the test statistic each time. This creates the empirical sampling distribution under the null hypothesis.
Obtain Sample Test Statistic: Conduct the actual study once using the target population (e.g., comparing Scrabble players against heavy metal concert attendees) and calculate the sample test statistic.
Derive -value: The -value represents the exact proportion of times that the null sampling distribution produced a test statistic equal to or more extreme than the observed sample statistic.
Example: A calculated means that in only out of null resampling runs did a test statistic equal or exceed the sample statistic.
Evaluate Decision Rule: Compare against (which splits into $2.5\% tails on both ends for a two-tailed test). If the sample test statistic falls in the extreme $2.5\% tail region, conclude that the sample was unlikely drawn from the null population and reject
Confidence Intervals and Their Relation to Hypothesis Testing:

A $95\% Confidence Interval (CI) represents the range of values within which the true population parameter will fall across $95\% of all possible samples drawn from that population.
Decision Alignment with Null Hypothesis Testing:
Real Population Effect Exists ( False): Individual sample $95\% CIs drawn from the alternative population will **not** span across or overlap the zero point of the null distribution.\n - **No Population Effect Exists (H_0 True):** Individual sample $95\% CIs drawn from the population will routinely span across and overlap the zero point of the null distribution.