Inferential Statistics Intro Continued
Fundamentals of Hypothesis Testing
Hypotheses as Statements of Truth: Hypotheses are defined as statements believed to be true regarding a specific phenomenon that have not yet been proven. They are structured to be testable.
The Null Hypothesis (): This hypothesis states that no statistical significance or difference exists between two or more populations or groups. It represents the default position that any observed effect is due to chance. * Specific Example (Elephant Ears): If average temperature does not affect the size of elephant ears, then species moving to a different climate zone will not evolve different ear sizes.
The Alternative Hypothesis ( or ): This states that a phenomenon is occurring due to non-random causes and is not just a matter of chance. It is the statement we accept if the Null Hypothesis is rejected. * Specific Example (Elephant Ears): If average temperature does affect ear size, then populations moving from warmer to cooler climates will evolve smaller ears.
Statistical Methodology: Scientists use specific statistical tests to determine whether to reject the Null Hypothesis. Rejecting allows for the acceptance of .
Standard Deviation and Data Spread
Definition of Standard Deviation (): A measurement used to quantify the dispersion or variability of data points relative to their mean. It reflects how far, on average, individual values deviate from the mean.
Mathematical Representation (): The standard deviation is denoted by the Greek letter sigma ().
Calculating : The lecture defines as "the square root of the sum of the squared deviations, divided by the number of values." * Variables involved: , , .
Normal Distribution Dispersion (The Empirical Rule): In a normal, symmetrical distribution: * There is a chance a random sample sits within of the mean ( above and below). * There is a chance a random sample sits within of the mean. * There is a chance a random sample sits within of the mean.
Interpretation of Spreads: * Normal distributions can have different means (, , or ) while maintaining the same spread. * Normal distributions can have different standard deviations (, , or ), resulting in thinner or wider curves regardless of the mean.
Confidence Intervals ()
Core Concept: The is the range in which the true mean of a population lies with a specific probability, derived from a sample set.
Sampling Variability: Different sample sets from the same population will likely yield different means by chance. The accounts for this.
Common Thresholds: The most used probabilities are and . In a , if a population is sampled times, the true mean should fall within the limits times.
Calculating Limits: * Lower Limit Calculation: * Upper Limit Calculation: * Components: -values are standard numbers from look-up tables; other values come from the specific sample set.
Visualization: are frequently added to correlation or regression plots to show the reliability of the trend (e.g., , , or bands).
Parametric vs. Non-Parametric Statistics
Parametric Tests: * Definition: These tests have complete information about population parameters or make assumptions about the distribution (usually normal). * Measurement Level: Metric scales, specifically Interval or Ratio scales. * Central Tendency: Focuses on the Mean. * Examples: t-test, ANOVA, Pearson correlation.
Non-Parametric Tests: * Definition: Applied when the researcher has no knowledge of population parameters or cannot make specific assumptions about the distribution (arbitrary or non-normal). * Measurement Level: Nominal or Ordinal scales. * Central Tendency: Focuses on the Median. * Examples: Mann-Whitney U test, Spearman correlation, Wilcoxon signed-rank test, Kruskal-Wallis test.
Variable Types in Biology: Most biological and biochemical data involve continuous variables like length, weight, pH, moles, and reaction speed. * Independent Variable: The cause (e.g., diet). * Dependent Variable: The effect/measurement (e.g., weight gain) that is "dependent" on the cause.
Determining Data Normality
Importance: Normality must be determined before selecting a statistical test; it is the "key" prerequisite.
Analytical/Statistical Tests for Normality: These tests evaluate the hypothesis that the data follows a normal distribution. * Shapiro-Wilk Test: Most appropriate for small sample sizes (). * Kolmogorov-Smirnov Test: Used for larger sample sizes (). * Anderson-Darling Test: An alternative normality test. * P-value Interpretation for Normality: * If : Reject the Null Hypothesis (Data is non-normally distributed). * If : Accept the Null Hypothesis (Data is normally distributed).
Sample Size Influence: Normality tests are highly sensitive to sample size. A very large sample may return a (non-normal) even if it is representative of a normal population, while a small sample may return a high p-value even if it is skewed.
Visual Inspection Methods: * Histograms: Plotting data density. * Quantile-Quantile (Q-Q) Plots: A stronger method comparing theoretical quantiles of a normal distribution against observed values. * Procedure: Sort values, divide a theoretical curve into areas, calculate theoretical quantile values, and plot against actual measurements. * Interpretation: A straight line indicates the data is normal. Deviations at the edges suggest non-normality.
The Meaning and Strength of P-values
Definition: The lower the P-value, the more confidence one has in rejecting the Null Hypothesis. It quantifies the strength of evidence for an experimental effect.
Categorization of Significance: * : No compelling evidence to reject the Null Hypothesis. * : Mild or weak evidence to reject the Null Hypothesis. * : Moderate evidence to reject the Null Hypothesis (marked by $ * $). * : Strong evidence to reject the Null Hypothesis (marked by $ ** $). * : Very strong evidence to reject the Null Hypothesis (marked by $ *** $). * *n.s.:* Over (non-significant).
The 0.05 Threshold: Known as the "holy grail" of cut-offs, established by classic statisticians as a balance between strictness and leniency. Modern scientists debate using binary cut-offs and prefer describing results as statistical "power."
Extreme Cases: Research like the Higgs-Boson particle requires ultra-low thresholds, such as (1 in 3.5 million).
P-hacking: The unethical practice of altering tests or data to force a result below the significane threshold (e.g., turning into ). This is considered scientific misconduct.
Box Plots and Data Visualisation
Interquartile Range (IQR): Also called the midspread or middle . * Data is divided into four rank-ordered parts (quartiles). * is the difference between the and percentiles.
Whiskers: Typically represent .
Skewness in Box Plots: * Positively Skewed: Distribution is skewed "up" or to the right; the median is closer to the bottom (). * Negatively Skewed: Distribution is skewed "down" or to the left; the median is closer to the top ().
Comparison of Visuals: * Bar Plots: Can hide data distributions and skewness. * Violin Plots: Overlay box plots with data points to reveal the true density and outliers of the population.
Philosophical Implications in Data Analysis
Hypothesis as a Liability: Excessive focus on a specific hypothesis may prevent researchers from making unexpected discoveries, especially in large datasets.
The Gorilla Experiment: An exercise involving a dataset of and step counts for people. * Group 1: Asked to test a specific significant difference between men and women. * Group 2: Asked to be "hypothesis-free" and draw conclusions generally. * Outcome: If the data is manually plotted, the image of a gorilla appears in the distribution. This serves as a metaphor for "hidden" things in data that automated number-crunching might miss.
Conclusion: Researchers must explore data manually and visually to capture biological implications that simple stat tests might obscure.