Inferential Statistics Intro Continued

Fundamentals of Hypothesis Testing

  • Hypotheses as Statements of Truth: Hypotheses are defined as statements believed to be true regarding a specific phenomenon that have not yet been proven. They are structured to be testable.

  • The Null Hypothesis (H0H_0): This hypothesis states that no statistical significance or difference exists between two or more populations or groups. It represents the default position that any observed effect is due to chance.     * Specific Example (Elephant Ears): If average temperature does not affect the size of elephant ears, then species XX moving to a different climate zone will not evolve different ear sizes.

  • The Alternative Hypothesis (H1H_1 or HaH_a): This states that a phenomenon is occurring due to non-random causes and is not just a matter of chance. It is the statement we accept if the Null Hypothesis is rejected.     * Specific Example (Elephant Ears): If average temperature does affect ear size, then populations moving from warmer to cooler climates will evolve smaller ears.

  • Statistical Methodology: Scientists use specific statistical tests to determine whether to reject the Null Hypothesis. Rejecting H0H_0 allows for the acceptance of H1H_1.

Standard Deviation and Data Spread

  • Definition of Standard Deviation (SDSD): A measurement used to quantify the dispersion or variability of data points relative to their mean. It reflects how far, on average, individual values deviate from the mean.

  • Mathematical Representation (σ\sigma): The standard deviation is denoted by the Greek letter sigma (σ\sigma).

  • Calculating SDSD: The lecture defines SDSD as "the square root of the sum of the squared deviations, divided by the number of values."     * Variables involved: N=number of valuesN = \text{number of values}, μ=the mean\mu = \text{the mean}, xi=individual valuesx_i = \text{individual values}.

  • Normal Distribution Dispersion (The Empirical Rule): In a normal, symmetrical distribution:     * There is a 68.2%68.2\% chance a random sample sits within 1σ1\sigma of the mean (34.1%34.1\% above and 34.1%34.1\% below).     * There is a 95.4%95.4\% chance a random sample sits within 2σ2\sigma of the mean.     * There is a 99.7%99.7\% chance a random sample sits within 3σ3\sigma of the mean.

  • Interpretation of Spreads:     * Normal distributions can have different means (μ=900\mu = 900, 11001100, or 13001300) while maintaining the same spread.     * Normal distributions can have different standard deviations (SD=80SD = 80, 120120, or 190190), resulting in thinner or wider curves regardless of the mean.

Confidence Intervals (CICI)

  • Core Concept: The CICI is the range in which the true mean of a population lies with a specific probability, derived from a sample set.

  • Sampling Variability: Different sample sets from the same population will likely yield different means by chance. The CICI accounts for this.

  • Common Thresholds: The most used probabilities are 95%95\% and 99%99\%. In a 95%95\% CICI, if a population is sampled 9595 times, the true mean should fall within the limits 9595 times.

  • Calculating Limits:     * Lower Limit Calculation: xl=xˉ−z⋅snx_l = \bar{x} - z \cdot \frac{s}{\sqrt{n}}     * Upper Limit Calculation: xu=xˉ+z⋅snx_u = \bar{x} + z \cdot \frac{s}{\sqrt{n}}     * Components: zz-values are standard numbers from look-up tables; other values come from the specific sample set.

  • Visualization: CIsCIs are frequently added to correlation or regression plots to show the reliability of the trend (e.g., 90%90\%, 95%95\%, or 99%99\% bands).

Parametric vs. Non-Parametric Statistics

  • Parametric Tests:     * Definition: These tests have complete information about population parameters or make assumptions about the distribution (usually normal).     * Measurement Level: Metric scales, specifically Interval or Ratio scales.     * Central Tendency: Focuses on the Mean.     * Examples: t-test, ANOVA, Pearson correlation.

  • Non-Parametric Tests:     * Definition: Applied when the researcher has no knowledge of population parameters or cannot make specific assumptions about the distribution (arbitrary or non-normal).     * Measurement Level: Nominal or Ordinal scales.     * Central Tendency: Focuses on the Median.     * Examples: Mann-Whitney U test, Spearman correlation, Wilcoxon signed-rank test, Kruskal-Wallis test.

  • Variable Types in Biology: Most biological and biochemical data involve continuous variables like length, weight, pH, moles, and reaction speed.     * Independent Variable: The cause (e.g., diet).     * Dependent Variable: The effect/measurement (e.g., weight gain) that is "dependent" on the cause.

Determining Data Normality

  • Importance: Normality must be determined before selecting a statistical test; it is the "key" prerequisite.

  • Analytical/Statistical Tests for Normality: These tests evaluate the hypothesis that the data follows a normal distribution.     * Shapiro-Wilk Test: Most appropriate for small sample sizes (n<50n < 50).     * Kolmogorov-Smirnov Test: Used for larger sample sizes (n≥50n \ge 50).     * Anderson-Darling Test: An alternative normality test.     * P-value Interpretation for Normality:         * If p<0.05p < 0.05: Reject the Null Hypothesis (Data is non-normally distributed).         * If p>0.05p > 0.05: Accept the Null Hypothesis (Data is normally distributed).

  • Sample Size Influence: Normality tests are highly sensitive to sample size. A very large sample may return a p<0.05p < 0.05 (non-normal) even if it is representative of a normal population, while a small sample may return a high p-value even if it is skewed.

  • Visual Inspection Methods:     * Histograms: Plotting data density.     * Quantile-Quantile (Q-Q) Plots: A stronger method comparing theoretical quantiles of a normal distribution against observed values.         * Procedure: Sort values, divide a theoretical curve into n+1n+1 areas, calculate theoretical quantile values, and plot against actual measurements.         * Interpretation: A straight line indicates the data is normal. Deviations at the edges suggest non-normality.

The Meaning and Strength of P-values

  • Definition: The lower the P-value, the more confidence one has in rejecting the Null Hypothesis. It quantifies the strength of evidence for an experimental effect.

  • Categorization of Significance:     * p>0.1p > 0.1: No compelling evidence to reject the Null Hypothesis.     * 0.05<p<0.100.05 < p < 0.10: Mild or weak evidence to reject the Null Hypothesis.     * p<0.05p < 0.05: Moderate evidence to reject the Null Hypothesis (marked by $ * $).     * p<0.01p < 0.01: Strong evidence to reject the Null Hypothesis (marked by $ ** $).     * p<0.001p < 0.001: Very strong evidence to reject the Null Hypothesis (marked by $ *** $).     * *n.s.:* Over 0.050.05 (non-significant).

  • The 0.05 Threshold: Known as the "holy grail" of cut-offs, established by classic statisticians as a balance between strictness and leniency. Modern scientists debate using binary cut-offs and prefer describing results as statistical "power."

  • Extreme Cases: Research like the Higgs-Boson particle requires ultra-low thresholds, such as 0.00000030.0000003 (1 in 3.5 million).

  • P-hacking: The unethical practice of altering tests or data to force a result below the significane threshold (e.g., turning 0.0510.051 into 0.0490.049). This is considered scientific misconduct.

Box Plots and Data Visualisation

  • Interquartile Range (IQR): Also called the midspread or middle 50%50\%.     * Data is divided into four rank-ordered parts (quartiles).     * IQRIQR is the difference between the 75th75^{th} and 25th25^{th} percentiles.

  • Whiskers: Typically represent 1.5×IQR1.5 \times \text{IQR}.

  • Skewness in Box Plots:     * Positively Skewed: Distribution is skewed "up" or to the right; the median is closer to the bottom (Q1Q1).     * Negatively Skewed: Distribution is skewed "down" or to the left; the median is closer to the top (Q3Q3).

  • Comparison of Visuals:     * Bar Plots: Can hide data distributions and skewness.     * Violin Plots: Overlay box plots with data points to reveal the true density and outliers of the population.

Philosophical Implications in Data Analysis

  • Hypothesis as a Liability: Excessive focus on a specific hypothesis may prevent researchers from making unexpected discoveries, especially in large datasets.

  • The Gorilla Experiment: An exercise involving a dataset of BMIBMI and step counts for 17861786 people.     * Group 1: Asked to test a specific significant difference between men and women.     * Group 2: Asked to be "hypothesis-free" and draw conclusions generally.     * Outcome: If the data is manually plotted, the image of a gorilla appears in the distribution. This serves as a metaphor for "hidden" things in data that automated number-crunching might miss.

  • Conclusion: Researchers must explore data manually and visually to capture biological implications that simple stat tests might obscure.