Central Limit Theorem for Sample Means Study Guide
Introduction to the Central Limit Theorem for Sample Means
The Central Limit Theorem (CLT) provides a framework for understanding what to expect when collecting data from multiple different random samples.
While previous studies established that a normal model represents proportions for multiple samples, the normal curve is also effective for representing the mean (μ) of quantitative data across several samples.
A primary application discussed is evaluating the mean SAT scores for various groups of high school seniors.
Variability of Sample Means
When drawing random samples, the resulting means (xˉ) will vary because each sample surveys a different set of people.
Specific examples of varying sample means include:
1030
1050
1035
1068
Despite this variability, if certain conditions are met, the distribution of these sample means will follow a normal distribution.
Required Conditions for the Normal Model
To ensure that sample means are normally distributed, two primary conditions must be verified:
Randomization Condition: This requires that all samples collected must be randomized.
Large Enough Condition: The sample size (n) must be sufficiently large. While there is no specific numerical cutoff defined, the size should be enough to allow the histogram of means to form a normal distribution.
If these conditions are satisfied, a histogram of the collected sample means will show a normal distribution with the actual population mean located at the center.
Characteristics of the Normal Model for Means
In the normal model for the mean, the actual population mean (μ) is positioned at the center.
Most sample means will fall within three standard deviations (SD) above or below this central population mean.
The Empirical Rule (68-95-99.7 Rule):
68% of sample means lie between one standard deviation below and one standard deviation above the population mean (±1SD).
95% of sample means lie between two standard deviations below and two standard deviations above the population mean (±2SD).
99.7% of sample means lie between three standard deviations below and three standard deviations above the population mean (±3SD).
Calculations for the Normal Model
Problems typically provide the population mean (μ) and the population standard deviation (σ).
Model Mean: The mean of the normal model is identical to the population mean provided.
Model Standard Deviation Calculation: To find the standard deviation specifically for the normal model, divide the population standard deviation by the square root of the sample size.
Formula:Standard deviation for the normal model=nσ
Case Study: College Board SAT Scores
Population Data:
Mean SAT score (μ): 1051
Population Standard Deviation (σ): 211
Sample Data:
Random sample size (n): 400 college applicants.
Sample mean (xˉ): 1100
Objective: Determine if a sample mean of 1100 is an unusually high average.
Condition Check: The sample was randomly collected, and a sample size of 400 is assumed to be large enough.
Standard Deviation Calculation:Standard deviation for the normal model=400211=20211=10.55
Determining the Normal Curve Ranges:
Mean (Center): 1051
+1SD: 1051+10.55=1061.55
+2SD: 1061.55+10.55=1072.1
+3SD: 1072.1+10.55=1082.65
−1SD: 1051−10.55=1040.45
−2SD: 1040.45−10.55=1029.9
−3SD: 1029.9−10.55=1019.35
Evaluation:
The sample mean of 1100 lies beyond the +3SD marker (1082.65).
Since 99.7% of all sample means are within three standard deviations, a mean of 1100 is more than three standard deviations above the mean and is therefore considered unusually high.
Comparison: Individual Data vs. Sample Means
Distribution of Data (Single Sample): This model looks at individual data points, such as individual SAT scores. These can range from very low to very high, resulting in a wide spread.
Distribution of Means (Multiple Samples): This model uses the averages calculated from large random samples. This process yields a much smaller spread of data compared to individual scores, as the means of large samples tend to be more consistent.
Summary of Key Findings
The normal model represents the distribution of means across several samples.
Critical requirements are the Randomization condition and the Large enough condition.
The actual population mean is centered in the normal model.
The standard deviation for the normal model is determined by:
Standard deviation for the normal model=sample sizeStandard deviation of the population