Central Limit Theorem for Sample Means Study Guide

Introduction to the Central Limit Theorem for Sample Means

  • The Central Limit Theorem (CLT) provides a framework for understanding what to expect when collecting data from multiple different random samples.
  • While previous studies established that a normal model represents proportions for multiple samples, the normal curve is also effective for representing the mean (μ\mu) of quantitative data across several samples.
  • A primary application discussed is evaluating the mean SAT scores for various groups of high school seniors.

Variability of Sample Means

  • When drawing random samples, the resulting means (xˉ\bar{x}) will vary because each sample surveys a different set of people.
  • Specific examples of varying sample means include:
    • 10301030
    • 10501050
    • 10351035
    • 10681068
  • Despite this variability, if certain conditions are met, the distribution of these sample means will follow a normal distribution.

Required Conditions for the Normal Model

  • To ensure that sample means are normally distributed, two primary conditions must be verified:
    • Randomization Condition: This requires that all samples collected must be randomized.
    • Large Enough Condition: The sample size (nn) must be sufficiently large. While there is no specific numerical cutoff defined, the size should be enough to allow the histogram of means to form a normal distribution.
  • If these conditions are satisfied, a histogram of the collected sample means will show a normal distribution with the actual population mean located at the center.

Characteristics of the Normal Model for Means

  • In the normal model for the mean, the actual population mean (μ\mu) is positioned at the center.
  • Most sample means will fall within three standard deviations (SDSD) above or below this central population mean.
  • The Empirical Rule (68-95-99.7 Rule):
    • 68%68\% of sample means lie between one standard deviation below and one standard deviation above the population mean (±1SD\pm 1\,SD).
    • 95%95\% of sample means lie between two standard deviations below and two standard deviations above the population mean (±2SD\pm 2\,SD).
    • 99.7%99.7\% of sample means lie between three standard deviations below and three standard deviations above the population mean (±3SD\pm 3\,SD).

Calculations for the Normal Model

  • Problems typically provide the population mean (μ\mu) and the population standard deviation (σ\sigma).
  • Model Mean: The mean of the normal model is identical to the population mean provided.
  • Model Standard Deviation Calculation: To find the standard deviation specifically for the normal model, divide the population standard deviation by the square root of the sample size.
  • Formula:Standard deviation for the normal model=σn\text{Standard deviation for the normal model} = \frac{\sigma}{\sqrt{n}}

Case Study: College Board SAT Scores

  • Population Data:
    • Mean SAT score (μ\mu): 10511051
    • Population Standard Deviation (σ\sigma): 211211
  • Sample Data:
    • Random sample size (nn): 400400 college applicants.
    • Sample mean (xˉ\bar{x}): 11001100
  • Objective: Determine if a sample mean of 11001100 is an unusually high average.
  • Condition Check: The sample was randomly collected, and a sample size of 400400 is assumed to be large enough.
  • Standard Deviation Calculation:Standard deviation for the normal model=211400=21120=10.55\text{Standard deviation for the normal model} = \frac{211}{\sqrt{400}} = \frac{211}{20} = 10.55
  • Determining the Normal Curve Ranges:
    • Mean (Center): 10511051
    • +1SD+1\,SD: 1051+10.55=1061.551051 + 10.55 = 1061.55
    • +2SD+2\,SD: 1061.55+10.55=1072.11061.55 + 10.55 = 1072.1
    • +3SD+3\,SD: 1072.1+10.55=1082.651072.1 + 10.55 = 1082.65
    • 1SD-1\,SD: 105110.55=1040.451051 - 10.55 = 1040.45
    • 2SD-2\,SD: 1040.4510.55=1029.91040.45 - 10.55 = 1029.9
    • 3SD-3\,SD: 1029.910.55=1019.351029.9 - 10.55 = 1019.35
  • Evaluation:
    • The sample mean of 11001100 lies beyond the +3SD+3\,SD marker (1082.651082.65).
    • Since 99.7%99.7\% of all sample means are within three standard deviations, a mean of 11001100 is more than three standard deviations above the mean and is therefore considered unusually high.

Comparison: Individual Data vs. Sample Means

  • Distribution of Data (Single Sample): This model looks at individual data points, such as individual SAT scores. These can range from very low to very high, resulting in a wide spread.
  • Distribution of Means (Multiple Samples): This model uses the averages calculated from large random samples. This process yields a much smaller spread of data compared to individual scores, as the means of large samples tend to be more consistent.

Summary of Key Findings

  • The normal model represents the distribution of means across several samples.
  • Critical requirements are the Randomization condition and the Large enough condition.
  • The actual population mean is centered in the normal model.
  • The standard deviation for the normal model is determined by:     Standard deviation for the normal model=Standard deviation of the populationsample size\text{Standard deviation for the normal model} = \frac{\text{Standard deviation of the population}}{\sqrt{\text{sample size}}}