Lecture Notes on Conceptual Issues in Null Hypothesis Significance Testing

Assignment Clarification and Upcoming Topics

  • Assignment due on Friday.

  • Questions can be asked after the lecture or on the Ed discussion forum.

  • The course will transition from statistics to qualitative research after next week's lecture by David Moreau.

  • Frequentist statistics, including t-tests and ANOVAs, have been the focus.

  • The lecture will cover conceptual issues in null hypothesis significance testing.

  • Tomorrow's lecture will focus on descriptive and exploratory statistics like correlation and regression.

  • Next week's lecture by David Murrow covers reproducible, reliable, and transparent scientific research in psychology and Bayesian hypothesis testing.

  • Bayesian hypothesis testing provides more continuous information about the belief in the null hypothesis and is increasingly used but less common than null hypothesis significance testing.

  • Qualitative research will be covered by Shiloh, Hinayatu, and Logan, offering non-numerical research methods.

The Logic of Hypothesis Testing

  • The underlying logic of hypothesis testing involves stating a hypothesis, formulating null and alternative hypotheses, deciding on an analysis, and calculating the p-value.

  • The hypothesis and research questions pertain to the population, not just the sample data.

  • In quantitative research, the aim is to make a general claim about a population based on sample observations.

  • Qualitative research often focuses on describing the experiences of individuals in the study without generalizing.

  • The goal in quantitative research is to generalize findings beyond the sample to a broader population.

Samples and Populations

  • Data cannot be collected from the entire population.

  • Researchers randomly sample from the population.

  • The data obtained are only about the sample.

  • Statistical processes are used to make inferences from the sample to the broader population.

  • Without the need to generalize, statistics would not be necessary.

  • Statistics justify the generalization and claim about the effect in the population based on the sample.

The Problem of Inference

  • The challenge is to infer information about the population based on sample data.

  • Sample values vary from sample to sample.

  • Each study yields a slightly different answer due to the variability in samples.

  • The goal is to understand the true population value or effect size.

  • Running multiple samples leads to different answers, some far from the truth.

Sample Statistics vs. Population Statistics

  • Sample statistics are values measured in the sample (e.g., mean, standard deviation).

  • Population statistics are the corresponding values in the entire population, which are generally unknown.

  • Statistical inference helps make logical leaps from sample measurements to population estimates.

  • Estimation involves measuring things in the sample to estimate things in the population.

  • Estimates can be imprecise or biased, so it's important to quantify imprecision using confidence intervals.

P-values and Decision Making

  • P-values help make decisions about the presence of an effect in the population, considering data uncertainty.

  • They provide a way of usually making the right decision. The p-value threshold formalizes whether the difference in the sample is large and reliable enough to indicate a difference in the population.

Precision and Bias in Estimation

  • Estimates can be wrong in two ways: imprecision and bias.

  • Imprecision means the estimates are very spread out across different samples (low precision).

  • Precision means the estimates are tightly clustered together (high precision).

  • Bias refers to whether the estimates are centered on the true population value.

  • Unbiased estimates are symmetrically distributed around the true value.

  • Ideally, estimates should be both precise (low variance) and unbiased.

Law of Large Numbers

  • The law of large numbers states that a statistic calculated from a sample approaches its true value in the whole population as the sample size increases.

  • Larger samples provide better estimates, which is why larger samples are preferred.

Simulation of Sampling from a Population

  • A simulation illustrates sampling from a population (e.g., women in New Zealand) to estimate population parameters (e.g., mean height).

  • In research, the true population mean and standard deviation are typically unknown.

  • The simulation involves recruiting a sample, measuring heights, and calculating sample mean and standard deviation.

  • The true population mean is what researchers aim to estimate.

  • The sample mean and the estimate of the population mean are the same for the mean, but slightly different for standard deviation.

Sample Size and Estimation Accuracy

  • With a small sample size (e.g., five), sample statistics can be far from population values.

  • As the sample size increases, the sample statistics get closer to the true population parameters.

  • Collecting data from a large sample (e.g., 1,000 people) yields values very close to true population values.

  • The true population values are those that would be measured if the entire population were surveyed.

  • Bigger sample sizes lead to better estimates, which is a foundational principle of experimental psychology.

Confidence Intervals

  • Confidence intervals provide a range within which the true population parameter is likely to fall.

  • They quantify the uncertainty about an estimate from a sample.

  • Confidence intervals depend on sample size and data variability.

  • A 95% confidence interval means that if the experiment were repeated 100 times, the true population mean would fall within the interval in about 95 of those experiments.

  • Wider confidence intervals indicate greater uncertainty, while narrower intervals indicate more precise estimates.

Confidence Intervals and Experimental Design

  • In experimental settings, the goal is to have tight confidence intervals around the estimate.

  • Testing more people leads to tighter confidence intervals.

  • Confidence intervals are particularly useful for effect sizes, indicating the range within which the true population effect size likely falls.

  • Reporting and paying attention to confidence intervals is good practice in research.

Central Limit Theorem

  • The central limit theorem states that the distribution of sample means approaches a normal distribution as the sample size gets larger.

  • This theorem involves repeating an experiment many times and plotting the distribution of the means observed across the samples.

  • Even if the original data are not normally distributed, the distribution of sample means will be normally distributed.

Simulation of the Central Limit Theorem

  • The simulation involves measuring something (e.g., height) in a sample, taking the sample mean, and repeating this process many times.

  • Each sample mean becomes a data point in a new histogram.

  • The distribution of these sample means will be normally distributed.

Examples and Implications of the Central Limit Theorem

  • If the data are normally distributed to start with (e.g., IQ scores), the distribution of sample means will also be normal.

  • As the sample size increases, the distribution of sample means becomes narrower.

  • The central limit theorem holds true regardless of how values were originally distributed in the raw data.

  • Even with skewed data (e.g., quiz scores), the distribution of sample means converges on a normal distribution as the sample size increases.

  • The central limit theorem explains why the normal distribution is common in nature and research, as many variables are averages of multiple factors.

Assumptions in Statistical Tests

  • T-tests and ANOVAs assume normality, which simplifies the maths and allows for more efficient testing.

  • The central limit theorem justifies these assumptions of normality in statistical tests.

  • Many variables are implicitly averaging together lots of different forces. If measuring variables that are implicitly averaging together lots of different forces, a normal distribution results.

  • Because the normal distribution is common, it can be used as an assumption in lots of tests like t-tests and ANOVAS.