Exhaustive Guide to Quantitative Hypothesis Testing, Effect Size, and Statistical Power
Foundations of Quantitative Research & Hypothesis Testing
Core Concepts of Variables and Investigation:
Effect: A general term describing the target phenomenon or relationship being investigated (determining whether an Independent Variable produces a change in a Dependent Variable ).
Independent Variable (IV): The variable manipulated or categorized by the researcher (e.g., driver age group, pet-interaction exposure, temperature condition).
Dependent Variable (DV): The outcome variable measured to assess the effect of the independent variable (e.g., number of driving errors, stress level, reaction time in milliseconds).
Formulating Hypotheses:
Null Hypothesis (): The baseline or default assumption stating that no real effect, difference, or association exists in the population beyond what occurs due to random chance or sampling variability (e.g., "Whether a driver is young or old will have no effect on driving errors").
Experimental / Alternative / Research Hypothesis (): The explicit prediction that a real effect, difference, or relationship exists in the population beyond chance variability.
Non-directional Alternative Hypothesis: Predicts an effect exists without specifying direction (e.g., "There will be an effect of age on the number of driving errors").
Directional Alternative Hypothesis: Predicts the specific nature or direction of the difference or association (e.g., "Younger drivers will make fewer driving errors than older drivers").

Dimension 1: Direction of Data
Evaluating Visual and Descriptive Patterns:
Determining direction is the initial and simplest step in evaluating research findings.
Direction is established by inspecting descriptive statistics (e.g., group means , standard deviations ) or graphical displays (e.g., bar charts, error bar graphs, scatter plots).
Group Differences Example (Driving Performance):
Evaluating driving errors across age categories yields distinct descriptive statistics:
Younger group: Mean driving errors , standard deviation .
Older group: Mean driving errors , standard deviation .
Reaction time measures show that younger drivers exhibit lower (faster) mean reaction times () compared to older drivers ().
Correlational Associations Example (Exercise and Well-Being):
Associations between two continuous scale variables are visually inspected using scatter graphs.
Analyzing weekly exercise hours against well-being scores reveals a positive linear relationship (as weekly exercise hours increase, reported well-being scores also increase).
Low-end data point: An individual with an exercise score of reports a low well-being score of .
High-end data point: An individual with an exercise score of reports a high well-being score of .
Empirical reference: Singh et al. (2023), British Journal of Sports Medicine, systematic review on physical activity interventions for mental health.

Pitfalls of Eyeballing Data & Descriptive Fallacies:
Visual inspection or descriptive comparisons alone ("eyeballing") cannot determine whether observed differences or associations are statistically meaningful or merely random noise.
Common Reporting Error: Concluding definitive group superiority solely from descriptive statistics (e.g., claiming "Young people made fewer mistakes () than old people (), therefore young people are better drivers") is an invalid inference without inferential testing.
Researchers possess an inherent confirmation bias toward overinterpreting descriptive noise in favor of their hypotheses.
Spurious conclusions occur when shared background variables are misidentified as causal mechanisms (e.g., observing that vodka + ice, ouzo + ice, whiskey + ice, and gin + ice all cause organ damage, and incorrectly inferring that ice is the destructive agent).
Methodological References: Benjamini (2016), The American Statistician; Fricker et al. (2019), The American Statistician.
Dimension 2: Null Hypothesis Significance Testing (NHST) & -Values
Foundations and Historical Evolution:
NHST traces back to 16th-century probability theory, with modern formalization occurring in the early 20th century.
Ronald Fisher (1925): Proposed using the probability () of an outcome occurring under the null hypothesis to evaluate the plausibility of . Fisher originally suggested as an appropriate threshold.
Jerzy Neyman & Egon Pearson (1933): Framework established that research cannot definitively "prove" the alternative hypothesis (); rather, researchers collect empirical evidence to falsify or reject the null hypothesis ().
Definition and Mathematical Properties of the -Value:
A -value represents the probability of obtaining test results at least as extreme as the observed results, assuming the null hypothesis () is true: .
-values are bounded strictly between and (). Absolute values of exactly or are mathematically impossible in empirical sampling.
: Indicates the observed data is extremely unlikely under pure randomness.
: Indicates the observed data closely matches pure randomness.
Decision Thresholds ():
The conventional significance threshold (-level) is .
Statistically Significant (): The observed data pattern has a less than probability (less than 1 in 20 chance) of occurring if the null hypothesis were true. The null hypothesis is rejected.
Non-Significant (): The observed data pattern does not deviate sufficiently from random chance expectations. The researcher fails to reject the null hypothesis.
Terminology Constraint: Results exceeding must be reported as "non-significant," never as "insignificant."
Classification Practice Examples:
Statistically Significant ().
Non-Significant ().
Non-Significant ().
Statistically Significant ().

Sampling Distributions and Normal Curves:
Under repeated sampling, dataset statistics form a theoretical normal distribution centered on the true value under the null hypothesis.
NHST evaluates whether the observed study sample falls within the central region of the null distribution (expected random variation) or within the extreme tails (unlikely under ).
A threshold of indicates that if an experiment were repeated 100 times under null conditions, a result this extreme would occur by accident only 5 times (1 in 20 iterations).
Cultural and Practical Rationale for :
Historical Precedent: Pre-computer statistical tables published by Fisher focused on discrete values ( corresponding to and corresponding to ).
Heuristic Balance: A threshold () is overly lenient, while a threshold () is excessively stringent for subtle psychological phenomena; () offers a practical middle ground ("One in twenty's plenty").
Methodological Limitations & Philosophical Considerations of NHST
Misinterpretations and "Spin" in Reporting:
Binary Cut-off vs. Continuum: The threshold represents a strict formal boundary. Describing non-significant results as "trending towards significance" () or "partially significant" () is statistically invalid. Non-significant results almost never trend toward significance with additional data without bias.
Gradation Misconception: A value of is not "more significant" than ; both satisfy the binary threshold for rejecting . Exact -values must always be cited directly rather than compared across studies.
Methodological References: Wood et al. (2014), BMJ; Appelbaum et al. (2018), American Psychologist; Gandevia (2024), Retraction Watch.
The Asymmetry of :
NHST calculates the probability of the data given the null hypothesis, , NOT the probability of the null hypothesis given the data, .
Low -values can occur for absurd or non-causal associations if the baseline theoretical model is flawed (e.g., observing a low -value for "unusual noise matching a T-Rex" simply because the noise is rare; Clark, 2023).
The Problem of Random Significance (The "Infinite Monkey" Paradigm):
If thousands of tests are performed on completely unrelated variables, will yield purely by chance.
Comparison of Research Scenarios:
Scientist 1: Conducts 2 years of rigorous theoretical research and systematic sampling, obtaining (failing to reach significance).
Scientist 2: Selects two random variables by a die roll, posts a social media survey, runs an arbitrary test, and obtains (statistically significant).
Scientist 3: Tests two variables established as unconnected by 1,000 previous studies, obtaining (statistically significant).
Conclusion: Without theoretical rationale and rigorous methodology, an isolated is meaningless. Statistical mechanics alone cannot replace scientific reasoning.
Methodological References: Greenland (2019), The American Statistician.
Dimension 3: Effect Size & Practical Significance
Conceptual Distinction:
While -values indicate statistical significance (whether an effect is likely real or noise), effect sizes measure practical significance or magnitude (how large, impactful, or meaningful the effect is in real-world terms).
Effect size is a standardized metric that quantifies the strength of a relationship or the distance between groups independently of sample size.
Key Advantages of Standardized Effect Sizes:
Continuum Measurement: Provides a scale-free continuous numerical indicator.
Cross-Study Comparability: Enables direct comparisons across different sample sizes, operational measures, study designs, and experimental contexts (significant -values cannot be compared across studies, but effect sizes can).
Practical Utility: Informs economic, clinical, and policy decisions (e.g., evaluating cost-benefit ratios of interventions).
Power Calculation: Serves as the primary input for prospective statistical power calculations.
Common Effect Size Statistics and Benchmarks:
Pearson's : Quantifies the linear relationship between two continuous variables.
Cohen's : Standardized difference between two group means, calculated as .
Small Effect Size: (minimal separation; extensive overlap of distributions).
Medium Effect Size: (moderate separation; standard benchmark in psychological research is ).
Large Effect Size: (substantial separation of group distributions).
Very Large Effect Size: (extreme separation, e.g., sex differences in human physical height).
Partial Eta Squared (): Measures proportion of variance explained by a factor in multi-group design settings (ANOVA).

The "Huge Sample, Tiny Effect" Phenomenon:
Very large sample sizes drastically reduce standard error, making minuscule and clinically meaningless differences reach statistical significance ().
Astrological Stereotype Study (Lu et al., 2020, JPSP):
Sample size:
Extraversion scores compared across birth seasons: Summer birth () vs. Spring (), Autumn (), and Winter ().
Statistical output: (extremely significant), but Cohen's (practically zero effect).
Enhancing Intuition with Natural Descriptions:
Abstract metrics can be translated into concrete metrics for clarity (e.g., describing a draft beer intervention as "a reduction of 29 litres, or 5%, less alcoholic beer sold").
Methodological References: Michal & Shah (2024), Psychological Science; Munafo (2024), The Psychologist; Richard et al. (2003), Review of General Psychology.
Dimension 4: Statistical Power & Error Management
Classification of Statistical Decision Errors:
Type I Error (): False Positive. Occurs when a researcher rejects the null hypothesis () when is actually true (concluding an effect exists when it does not).
Setting fixes the Type I error rate at for a single test.
Running multiple unadjusted tests inflates the overall Type I error rate ("familywise error rate").
Type II Error (): False Negative. Occurs when a researcher retains the null hypothesis () when is actually false (failing to detect a real existing effect).

Definition and Standards of Statistical Power:
Statistical Power is defined as the probability of correctly rejecting a false null hypothesis ().
It represents the study's ability to detect a true population effect of a given size.
Standard Benchmark: Psychological science adopts a power target of ( probability of detecting a real effect, accepting a Type II error rate ).
Determinants of Statistical Power:
Sample Size (): The primary controllable factor; increasing sample size increases power.
Effect Size: Larger true population effects are easier to detect and require smaller samples.
Alpha Threshold (): Lowering (e.g., from to ) reduces Type I errors but decreases power, inflating Type II errors.
Retrospective Power Analysis (Post-Hoc Evaluation):
Calculated after research completion to evaluate non-significant findings.
Worked Example: A driving error experiment with per condition yields and an estimated effect size .\n * Consulting power tables for N = 100$ and $d = 0.20$ gives a statistical power of 0.4141\%).\n * *Interpretation:* The study had only a 41\% chance of detecting a small effect, meaning the non-significant result cannot be confidently accepted as a true null.\n * *Methodological Warning:* Low power cannot be used for "hypothesis-laundering" (claiming an unsupported hypothesis was "nearly supported"). It serves strictly as a methodological caveat.\n\n* **Prospective Power Analysis (A Priori Sample Planning):**\n * Conducted prior to data collection to determine necessary sample size.\n * *Worked Example:* Planning a study to detect an anticipated small effect (d = 0.20\ge 0.80\alpha = 0.05.\n * Consulting prospective power tables under d = 0.200.88N = 400 per group.\n * Total required sample size = 800400400 older).\n\n# The DSS Framework & Practical Decision Matrix\n\n* **The DSS Decision Rule:**\n * Complete evaluation of any quantitative analysis requires integrating three elements:\n 1. **D**irection: Inspect descriptive statistics (M, SD) and graphs to identify structural trends.\n 2. **S**ignificance: Check whether the pp < 0.05 via NHST.\n 3. **S**ize: Evaluate the standardized effect size (r, d, \eta_p^2) against conventional benchmarks.\n\n* **Comparison Matrix of Analysis Metrics:**\n\n| Metric | Primary Function | Core Reference Scale | Key Influencing Factor | Main Risk / Failure Mode |\n| :--- | :--- | :--- | :--- | :--- |\n| **pH_0< 0.05\ge 0.05NN (trivial effects become significant) |\n| **Effect Size** | Measures practical magnitude under H_1d = 0.2, 0.5, 0.8) | Underlying effect strength | Unstable in tiny samples or poor designs |\n| **Statistical Power** | Assesses reliance on non-significant H_0\ge 0.80) | Sample size and target effect size | Underpowered studies cause false negatives |\n\n* **Evolution of Reporting Standards in Science:**\n * *Historical Practice (Pre-2012):* Research papers primarily cited p-values alone without effect sizes or power analyses, contributing directly to the psychology replication crisis.\n * *Modern Standard:* Journals require routine reporting of exact p-values, standardized effect sizes, and prospective power calculations for sample size justification (Wasserstein & Lazar, 2016, *The American Statistician*).\n\n# Empirical Scenarios & Results Interpretation\n\n* **Scenario 1: Pet Interaction and Stress Reduction**\n * *Design:* Comparing participant stress levels before vs. after three sessions of cat-based interaction.\n * *Statistical Output:* p = 0.002d = 0.50\beta = 0.68.\n * *Variables:* Independent Variable = exposure to cat interaction; Dependent Variable = stress score.\n * *Interpretation:* Statistically significant reduction in stress (p = 0.002 < 0.05d = 0.50) indicates a practically meaningful intervention effect. Power is not a limiting factor because statistical significance was achieved.\n\n* **Scenario 2: Power Posing in Job Interviews**\n * *Design:* Applicants instructed to power pose vs. no pose prior to mock interviews; confidence rated by naive observers.\n * *Statistical Output:* p = 0.040d = 0.01\beta = 0.98.\n * *Variables:* Independent Variable = posing condition (pose vs. no pose); Dependent Variable = observer-rated confidence score.\n * *Interpretation:* Statistically significant difference (p = 0.040 < 0.05d = 0.01), meaning the intervention has no practical utility. High sample size drove statistical significance despite a trivial effect.\n\n* **Scenario 3: Sex Differences in Social Media Contacts**\n * *Design:* Examining weekly social media friend contacts across biological sex categories.\n * *Statistical Output:* p = 0.051d = 0.40\beta = 0.75.\n * *Variables:* Independent Variable = Biological Sex (Male vs. Female); Dependent Variable = weekly number of social media contacts.\n * *Interpretation:* Non-significant result (p = 0.051 \ge 0.05H_00.75 < 0.80), introducing a distinct risk of a Type II error (false negative). Effect size magnitude is disregarded due to non-significance.\n\n* **Scenario 4: Ambient Temperature and Reaction Times**\n * *Design:* Reaction time testing completed under hot vs. cold environmental conditions (counterbalanced).\n * *Statistical Output:* p = 0.060d = 0.35\beta = 0.90.\n * *Variables:* Independent Variable = Ambient Temperature (Hot vs. Cold); Dependent Variable = reaction time (\text{ms}).\n * *Interpretation:* Non-significant effect (p = 0.060 \ge 0.050.90 \ge 0.80), this non-significant result can be reliably accepted as a true null result rather than a false negative.\n\n* **Scenario 5: TikTok Usage in Peru vs. Lecture Absences in Staffordshire**\n * *Design:* Correlating weekly hours spent on TikTok by Peruvian teenagers with lecture absences among UK psychology students over 40 weeks.\n * *Statistical Output:* p = 0.030r = 0.50\beta = 0.80.\n * *Variables:* Independent Variable = Peruvian teen TikTok hours; Dependent Variable = UK statistics lecture absences.\n * *Interpretation:* Statistically significant (p = 0.030 < 0.05r = 0.500.80$$). However, the finding is completely spurious and scientifically meaningless due to the absence of any theoretical mechanism or logical methodological linkage.