Exhaustive Guide to Quantitative Hypothesis Testing, Effect Size, and Statistical Power

Foundations of Quantitative Research & Hypothesis Testing

  • Core Concepts of Variables and Investigation:

    • Effect: A general term describing the target phenomenon or relationship being investigated (determining whether an Independent Variable XX produces a change in a Dependent Variable YY).

    • Independent Variable (IV): The variable manipulated or categorized by the researcher (e.g., driver age group, pet-interaction exposure, temperature condition).

    • Dependent Variable (DV): The outcome variable measured to assess the effect of the independent variable (e.g., number of driving errors, stress level, reaction time in milliseconds).

  • Formulating Hypotheses:

    • Null Hypothesis (H0H_0): The baseline or default assumption stating that no real effect, difference, or association exists in the population beyond what occurs due to random chance or sampling variability (e.g., "Whether a driver is young or old will have no effect on driving errors").

    • Experimental / Alternative / Research Hypothesis (H1H_1): The explicit prediction that a real effect, difference, or relationship exists in the population beyond chance variability.

      • Non-directional Alternative Hypothesis: Predicts an effect exists without specifying direction (e.g., "There will be an effect of age on the number of driving errors").

      • Directional Alternative Hypothesis: Predicts the specific nature or direction of the difference or association (e.g., "Younger drivers will make fewer driving errors than older drivers").

Pre and post kitten stress estimates

Dimension 1: Direction of Data

  • Evaluating Visual and Descriptive Patterns:

    • Determining direction is the initial and simplest step in evaluating research findings.

    • Direction is established by inspecting descriptive statistics (e.g., group means MM, standard deviations SDSD) or graphical displays (e.g., bar charts, error bar graphs, scatter plots).

  • Group Differences Example (Driving Performance):

    • Evaluating driving errors across age categories yields distinct descriptive statistics:

      • Younger group: Mean driving errors M=25.45M = 25.45, standard deviation SD=4.12SD = 4.12.

      • Older group: Mean driving errors M=34.71M = 34.71, standard deviation SD=5.79SD = 5.79.

    • Reaction time measures show that younger drivers exhibit lower (faster) mean reaction times (≈280 ms\approx 280\,\text{ms}) compared to older drivers (≈340 ms\approx 340\,\text{ms}).

  • Correlational Associations Example (Exercise and Well-Being):

    • Associations between two continuous scale variables are visually inspected using scatter graphs.

    • Analyzing weekly exercise hours against well-being scores reveals a positive linear relationship (as weekly exercise hours increase, reported well-being scores also increase).

      • Low-end data point: An individual with an exercise score of 5.65 hours5.65\,\text{hours} reports a low well-being score of 1.611.61.

      • High-end data point: An individual with an exercise score of 22.08 hours22.08\,\text{hours} reports a high well-being score of 14.5414.54.

      • Empirical reference: Singh et al. (2023), British Journal of Sports Medicine, systematic review on physical activity interventions for mental health.

Exercise hours versus wellbeing scatter graph
  • Pitfalls of Eyeballing Data & Descriptive Fallacies:

    • Visual inspection or descriptive comparisons alone ("eyeballing") cannot determine whether observed differences or associations are statistically meaningful or merely random noise.

    • Common Reporting Error: Concluding definitive group superiority solely from descriptive statistics (e.g., claiming "Young people made fewer mistakes (M=25.45M = 25.45) than old people (M=31.71M = 31.71), therefore young people are better drivers") is an invalid inference without inferential testing.

    • Researchers possess an inherent confirmation bias toward overinterpreting descriptive noise in favor of their hypotheses.

    • Spurious conclusions occur when shared background variables are misidentified as causal mechanisms (e.g., observing that vodka + ice, ouzo + ice, whiskey + ice, and gin + ice all cause organ damage, and incorrectly inferring that ice is the destructive agent).

    • Methodological References: Benjamini (2016), The American Statistician; Fricker et al. (2019), The American Statistician.

Dimension 2: Null Hypothesis Significance Testing (NHST) & pp-Values

  • Foundations and Historical Evolution:

    • NHST traces back to 16th-century probability theory, with modern formalization occurring in the early 20th century.

    • Ronald Fisher (1925): Proposed using the probability (pp) of an outcome occurring under the null hypothesis to evaluate the plausibility of H0H_0. Fisher originally suggested p<0.01p < 0.01 as an appropriate threshold.

    • Jerzy Neyman & Egon Pearson (1933): Framework established that research cannot definitively "prove" the alternative hypothesis (H1H_1); rather, researchers collect empirical evidence to falsify or reject the null hypothesis (H0H_0).

  • Definition and Mathematical Properties of the pp-Value:

    • A pp-value represents the probability of obtaining test results at least as extreme as the observed results, assuming the null hypothesis (H0H_0) is true: P(data∣H0)P(\text{data} \mid H_0).

    • pp-values are bounded strictly between 00 and 11 (0≤p≤10 \le p \le 1). Absolute values of exactly 00 or 11 are mathematically impossible in empirical sampling.

      • p→0p \rightarrow 0: Indicates the observed data is extremely unlikely under pure randomness.

      • p→1p \rightarrow 1: Indicates the observed data closely matches pure randomness.

  • Decision Thresholds (α=0.05\alpha = 0.05):

    • The conventional significance threshold (α\alpha-level) is 0.050.05.

    • Statistically Significant (p<0.05p < 0.05): The observed data pattern has a less than 5%5\% probability (less than 1 in 20 chance) of occurring if the null hypothesis were true. The null hypothesis is rejected.

    • Non-Significant (p≥0.05p \ge 0.05): The observed data pattern does not deviate sufficiently from random chance expectations. The researcher fails to reject the null hypothesis.

    • Terminology Constraint: Results exceeding 0.050.05 must be reported as "non-significant," never as "insignificant."

  • Classification Practice Examples:

    • p=0.010p = 0.010 →\rightarrow Statistically Significant (p<0.05p < 0.05).

    • p=0.500p = 0.500 →\rightarrow Non-Significant (p≥0.05p \ge 0.05).

    • p=0.052p = 0.052 →\rightarrow Non-Significant (p≥0.05p \ge 0.05).

    • p=0.049p = 0.049 →\rightarrow Statistically Significant (p<0.05p < 0.05).

Null hypothesis normal distribution and significance threshold
  • Sampling Distributions and Normal Curves:

    • Under repeated sampling, dataset statistics form a theoretical normal distribution centered on the true value under the null hypothesis.

    • NHST evaluates whether the observed study sample falls within the central 95%95\% region of the null distribution (expected random variation) or within the extreme 5%5\% tails (unlikely under H0H_0).

    • A threshold of p=0.05p = 0.05 indicates that if an experiment were repeated 100 times under null conditions, a result this extreme would occur by accident only 5 times (1 in 20 iterations).

  • Cultural and Practical Rationale for 0.050.05:

    • Historical Precedent: Pre-computer statistical tables published by Fisher focused on discrete values (p=0.05p = 0.05 corresponding to ≈2 SD\approx 2\,SD and p=0.01p = 0.01 corresponding to ≈3 SD\approx 3\,SD).

    • Heuristic Balance: A 1/101/10 threshold (0.100.10) is overly lenient, while a 1/1001/100 threshold (0.010.01) is excessively stringent for subtle psychological phenomena; 1/201/20 (0.050.05) offers a practical middle ground ("One in twenty's plenty").

Methodological Limitations & Philosophical Considerations of NHST

  • Misinterpretations and "Spin" in Reporting:

    • Binary Cut-off vs. Continuum: The threshold p<0.05p < 0.05 represents a strict formal boundary. Describing non-significant results as "trending towards significance" (p=0.06p = 0.06) or "partially significant" (p=0.10p = 0.10) is statistically invalid. Non-significant results almost never trend toward significance with additional data without bias.

    • Gradation Misconception: A value of p=0.001p = 0.001 is not "more significant" than p=0.049p = 0.049; both satisfy the binary threshold for rejecting H0H_0. Exact pp-values must always be cited directly rather than compared across studies.

    • Methodological References: Wood et al. (2014), BMJ; Appelbaum et al. (2018), American Psychologist; Gandevia (2024), Retraction Watch.

  • The Asymmetry of P(data∣null)P(\text{data} \mid \text{null}):

    • NHST calculates the probability of the data given the null hypothesis, P(data∣H0)P(\text{data} \mid H_0), NOT the probability of the null hypothesis given the data, P(H0∣data)P(H_0 \mid \text{data}).

    • Low pp-values can occur for absurd or non-causal associations if the baseline theoretical model is flawed (e.g., observing a low pp-value for "unusual noise matching a T-Rex" simply because the noise is rare; Clark, 2023).

  • The Problem of Random Significance (The "Infinite Monkey" Paradigm):

    • If thousands of tests are performed on completely unrelated variables, 5%5\% will yield p<0.05p < 0.05 purely by chance.

    • Comparison of Research Scenarios:

      • Scientist 1: Conducts 2 years of rigorous theoretical research and systematic sampling, obtaining p=0.055p = 0.055 (failing to reach significance).

      • Scientist 2: Selects two random variables by a die roll, posts a social media survey, runs an arbitrary test, and obtains p=0.043p = 0.043 (statistically significant).

      • Scientist 3: Tests two variables established as unconnected by 1,000 previous studies, obtaining p=0.049p = 0.049 (statistically significant).

    • Conclusion: Without theoretical rationale and rigorous methodology, an isolated p<0.05p < 0.05 is meaningless. Statistical mechanics alone cannot replace scientific reasoning.

    • Methodological References: Greenland (2019), The American Statistician.

Dimension 3: Effect Size & Practical Significance

  • Conceptual Distinction:

    • While pp-values indicate statistical significance (whether an effect is likely real or noise), effect sizes measure practical significance or magnitude (how large, impactful, or meaningful the effect is in real-world terms).

    • Effect size is a standardized metric that quantifies the strength of a relationship or the distance between groups independently of sample size.

  • Key Advantages of Standardized Effect Sizes:

    • Continuum Measurement: Provides a scale-free continuous numerical indicator.

    • Cross-Study Comparability: Enables direct comparisons across different sample sizes, operational measures, study designs, and experimental contexts (significant pp-values cannot be compared across studies, but effect sizes can).

    • Practical Utility: Informs economic, clinical, and policy decisions (e.g., evaluating cost-benefit ratios of interventions).

    • Power Calculation: Serves as the primary input for prospective statistical power calculations.

  • Common Effect Size Statistics and Benchmarks:

    • Pearson's rr: Quantifies the linear relationship between two continuous variables.

    • Cohen's dd: Standardized difference between two group means, calculated as d=M1−M2SDpooledd = \frac{M_1 - M_2}{SD_\text{pooled}}.

      • Small Effect Size: d=0.20d = 0.20 (minimal separation; extensive overlap of distributions).

      • Medium Effect Size: d=0.50d = 0.50 (moderate separation; standard benchmark in psychological research is d≈0.40,SD=0.30d \approx 0.40, SD = 0.30).

      • Large Effect Size: d=0.80d = 0.80 (substantial separation of group distributions).

      • Very Large Effect Size: d≥2.00d \ge 2.00 (extreme separation, e.g., sex differences in human physical height).

    • Partial Eta Squared (ηp2\eta_p^2): Measures proportion of variance explained by a factor in multi-group design settings (ANOVA).

Cohen's d distribution overlaps
  • The "Huge Sample, Tiny Effect" Phenomenon:

    • Very large sample sizes drastically reduce standard error, making minuscule and clinically meaningless differences reach statistical significance (p<0.05p < 0.05).

    • Astrological Stereotype Study (Lu et al., 2020, JPSP):

      • Sample size: N=173,709N = 173,709

      • Extraversion scores compared across birth seasons: Summer birth (5.035.03) vs. Spring (5.015.01), Autumn (5.015.01), and Winter (5.005.00).

      • Statistical output: p<0.001p < 0.001 (extremely significant), but Cohen's d=0.02d = 0.02 (practically zero effect).

  • Enhancing Intuition with Natural Descriptions:

    • Abstract metrics can be translated into concrete metrics for clarity (e.g., describing a draft beer intervention as "a reduction of 29 litres, or 5%, less alcoholic beer sold").

    • Methodological References: Michal & Shah (2024), Psychological Science; Munafo (2024), The Psychologist; Richard et al. (2003), Review of General Psychology.

Dimension 4: Statistical Power & Error Management

  • Classification of Statistical Decision Errors:

    • Type I Error (α\alpha): False Positive. Occurs when a researcher rejects the null hypothesis (H0H_0) when H0H_0 is actually true (concluding an effect exists when it does not).

      • Setting α=0.05\alpha = 0.05 fixes the Type I error rate at 5%5\% for a single test.

      • Running multiple unadjusted tests inflates the overall Type I error rate ("familywise error rate").

    • Type II Error (β\beta): False Negative. Occurs when a researcher retains the null hypothesis (H0H_0) when H0H_0 is actually false (failing to detect a real existing effect).

Type I and Type II errors illustration
  • Definition and Standards of Statistical Power:

    • Statistical Power is defined as the probability of correctly rejecting a false null hypothesis (1−β1 - \beta).

    • It represents the study's ability to detect a true population effect of a given size.

    • Standard Benchmark: Psychological science adopts a power target of 1−β=0.801 - \beta = 0.80 (80%80\% probability of detecting a real effect, accepting a 20%20\% Type II error rate β=0.20\beta = 0.20).

  • Determinants of Statistical Power:

    1. Sample Size (NN): The primary controllable factor; increasing sample size increases power.

    2. Effect Size: Larger true population effects are easier to detect and require smaller samples.

    3. Alpha Threshold (α\alpha): Lowering α\alpha (e.g., from 0.050.05 to 0.010.01) reduces Type I errors but decreases power, inflating Type II errors.

  • Retrospective Power Analysis (Post-Hoc Evaluation):

    • Calculated after research completion to evaluate non-significant findings.

    • Worked Example: A driving error experiment with N=100N = 100 per condition yields p=0.07p = 0.07 and an estimated effect size d=0.20d = 0.20.\n * Consulting power tables for N = 100$ and $d = 0.20$ gives a statistical power of 0.41((41\%).\n * *Interpretation:* The study had only a 41\% chance of detecting a small effect, meaning the non-significant result cannot be confidently accepted as a true null.\n * *Methodological Warning:* Low power cannot be used for "hypothesis-laundering" (claiming an unsupported hypothesis was "nearly supported"). It serves strictly as a methodological caveat.\n\n* **Prospective Power Analysis (A Priori Sample Planning):**\n * Conducted prior to data collection to determine necessary sample size.\n * *Worked Example:* Planning a study to detect an anticipated small effect (d = 0.20)withtargetpower) with target power\ge 0.80atat\alpha = 0.05.\n * Consulting prospective power tables under d = 0.20indicatespowerreachesindicates power reaches0.88atatN = 400 per group.\n * Total required sample size = 800participants(participants (400younger,younger,400 older).\n\n# The DSS Framework & Practical Decision Matrix\n\n* **The DSS Decision Rule:**\n * Complete evaluation of any quantitative analysis requires integrating three elements:\n 1. **D**irection: Inspect descriptive statistics (M, SD) and graphs to identify structural trends.\n 2. **S**ignificance: Check whether the p−valuesatisfies-value satisfiesp < 0.05 via NHST.\n 3. **S**ize: Evaluate the standardized effect size (r, d, \eta_p^2) against conventional benchmarks.\n\n* **Comparison Matrix of Analysis Metrics:**\n\n| Metric | Primary Function | Core Reference Scale | Key Influencing Factor | Main Risk / Failure Mode |\n| :--- | :--- | :--- | :--- | :--- |\n| **p−Value∗∗∣Evaluatesstatisticalsignificanceagainst-Value** | Evaluates statistical significance againstH_0∣Binarythreshold(| Binary threshold (< 0.05vs.vs.\ge 0.05)∣Samplesize() | Sample size (N)∣Inflatedbyhuge) | Inflated by hugeN (trivial effects become significant) |\n| **Effect Size** | Measures practical magnitude under H_1∣Continuouscontinuum(| Continuous continuum (d = 0.2, 0.5, 0.8) | Underlying effect strength | Unstable in tiny samples or poor designs |\n| **Statistical Power** | Assesses reliance on non-significant H_0results∣Targetthreshold(results | Target threshold (\ge 0.80) | Sample size and target effect size | Underpowered studies cause false negatives |\n\n* **Evolution of Reporting Standards in Science:**\n * *Historical Practice (Pre-2012):* Research papers primarily cited p-values alone without effect sizes or power analyses, contributing directly to the psychology replication crisis.\n * *Modern Standard:* Journals require routine reporting of exact p-values, standardized effect sizes, and prospective power calculations for sample size justification (Wasserstein & Lazar, 2016, *The American Statistician*).\n\n# Empirical Scenarios & Results Interpretation\n\n* **Scenario 1: Pet Interaction and Stress Reduction**\n * *Design:* Comparing participant stress levels before vs. after three sessions of cat-based interaction.\n * *Statistical Output:* p = 0.002,Cohen′s, Cohen'sd = 0.50(medium),Power(medium), Power\beta = 0.68.\n * *Variables:* Independent Variable = exposure to cat interaction; Dependent Variable = stress score.\n * *Interpretation:* Statistically significant reduction in stress (p = 0.002 < 0.05).Mediumeffectsize(). Medium effect size (d = 0.50) indicates a practically meaningful intervention effect. Power is not a limiting factor because statistical significance was achieved.\n\n* **Scenario 2: Power Posing in Job Interviews**\n * *Design:* Applicants instructed to power pose vs. no pose prior to mock interviews; confidence rated by naive observers.\n * *Statistical Output:* p = 0.040,Cohen′s, Cohen'sd = 0.01(small),Power(small), Power\beta = 0.98.\n * *Variables:* Independent Variable = posing condition (pose vs. no pose); Dependent Variable = observer-rated confidence score.\n * *Interpretation:* Statistically significant difference (p = 0.040 < 0.05).However,theeffectsizeisnegligible(). However, the effect size is negligible (d = 0.01), meaning the intervention has no practical utility. High sample size drove statistical significance despite a trivial effect.\n\n* **Scenario 3: Sex Differences in Social Media Contacts**\n * *Design:* Examining weekly social media friend contacts across biological sex categories.\n * *Statistical Output:* p = 0.051,Cohen′s, Cohen'sd = 0.40(medium),Power(medium), Power\beta = 0.75.\n * *Variables:* Independent Variable = Biological Sex (Male vs. Female); Dependent Variable = weekly number of social media contacts.\n * *Interpretation:* Non-significant result (p = 0.051 \ge 0.05);failtoreject); fail to rejectH_0.Statisticalpowerisinadequate(. Statistical power is inadequate (0.75 < 0.80), introducing a distinct risk of a Type II error (false negative). Effect size magnitude is disregarded due to non-significance.\n\n* **Scenario 4: Ambient Temperature and Reaction Times**\n * *Design:* Reaction time testing completed under hot vs. cold environmental conditions (counterbalanced).\n * *Statistical Output:* p = 0.060,Cohen′s, Cohen'sd = 0.35(medium),Power(medium), Power\beta = 0.90.\n * *Variables:* Independent Variable = Ambient Temperature (Hot vs. Cold); Dependent Variable = reaction time (\text{ms}).\n * *Interpretation:* Non-significant effect (p = 0.060 \ge 0.05).Becausestatisticalpowerishigh(). Because statistical power is high (0.90 \ge 0.80), this non-significant result can be reliably accepted as a true null result rather than a false negative.\n\n* **Scenario 5: TikTok Usage in Peru vs. Lecture Absences in Staffordshire**\n * *Design:* Correlating weekly hours spent on TikTok by Peruvian teenagers with lecture absences among UK psychology students over 40 weeks.\n * *Statistical Output:* p = 0.030,Pearson′s, Pearson'sr = 0.50(largecorrelation),Power(large correlation), Power\beta = 0.80.\n * *Variables:* Independent Variable = Peruvian teen TikTok hours; Dependent Variable = UK statistics lecture absences.\n * *Interpretation:* Statistically significant (p = 0.030 < 0.05)withalargeeffectsize() with a large effect size (r = 0.50)andadequatepower() and adequate power (0.80$$). However, the finding is completely spurious and scientifically meaningless due to the absence of any theoretical mechanism or logical methodological linkage.