Probability and Samples: The Distribution of Sample Means

Samples, Populations, and Sampling Error

  • Score Locations: The location of an individual score within a sample or population can be represented using a zz-score.
  • Research Focus: Researchers typically aim to study entire samples rather than isolated individual scores because samples provide estimates of population parameters.
  • Sampling Error Definition: Sampling error is defined as the natural discrepancy, or amount of error, between a sample statistic (such as a sample mean, MM) and its corresponding population parameter (such as the population mean, \text{\mu}).
    • Sampling error does not imply that a mistake or calculation error was made.
    • Samples possess natural variability; two distinct samples drawn from the same population are very rarely identical.

The Distribution of Sample Means

  • Sampling Variability: Selecting two separate random samples from the same population will almost always yield different sample means.
  • Definition: The distribution of sample means is defined as the collection of sample means for all possible random samples of a specific sample size (nn) that can be selected from a population.
  • Sampling Distribution: Unlike distributions of individual scores, a sampling distribution is a distribution of sample statistics (sample means).
    • The distribution of sample means forms a special population of sample means derived by extracting every possible sample of size nn from the baseline population.
  • General Characteristics:
    • Sample means tend to pile up around the population mean (μ\mu).
    • The distribution of sample means is approximately normal in shape.
    • As sample size (nn) increases, sample means cluster more tightly around the population mean (μ\mu).

Frequency Distribution Histogram for a Population of N = 4 scores

The Distribution of Sample Means for n = 2

The Central Limit Theorem

  • Applicability: The Central Limit Theorem applies to any population with a known mean (μ\mu) and standard deviation (σ\sigma).
  • Mathematical Principles:
    • As sample size (nn) approaches infinity (∞\infty), the distribution of sample means approaches a normal distribution.
    • The mean of the distribution of sample means for samples of size nn is equal to μM=μ\mu_M = \mu.
    • The standard deviation of the distribution of sample means for samples of size nn (known as standard error) is equal to:     σM=σn\sigma_M = \frac{\sigma}{\sqrt{n}}

Standard Error Formula

Properties of the Distribution of Sample Means

  • Shape Requirements for Normality: The distribution of sample means is almost perfectly normal under either of two conditions:
    • The population from which samples are selected is normally distributed.
    • The sample size (nn) in each sample is relatively large (n≥30n \ge 30).
  • Central Tendency (Expected Value of M):
    • The mean of the distribution of sample means is denoted by μM\mu_M and is identically equal to the population mean μ\mu.
    • This population parameter μM\mu_M is formally termed the expected value of MM.
  • Variability (Standard Error of M):
    • The standard deviation of the distribution of sample means is designated as the standard error of MM and written as σM\sigma_M.
    • While standard deviation measures the variability of individual scores, standard error measures the variability of sample means.
    • Standard error quantifies the average distance expected between a sample mean (MM) and the population mean (μ\mu).
    • A large standard error indicates that sample means are widely scattered.
  • Factors Influencing Standard Error:
    • Law of Large Numbers: As sample size (nn) increases, the probability that the sample mean (MM) is close to the population mean (μ\mu) increases, causing standard error to decrease.
    • Population Variance: Smaller variance or standard deviation (σ\sigma) in the population increases the probability that the sample mean (MM) is close to μ\mu, resulting in a smaller standard error.

The Relationship between Standard Error and Sample Size

z-Scores and Probability for Sample Means

  • Primary Purpose: The primary utility of the distribution of sample means is to calculate the probability of obtaining a sample with a specific sample mean (MM).
  • Procedure:
    1. Calculate the standard error σM=σn\sigma_M = \frac{\sigma}{\sqrt{n}}.
    2. Compute the zz-score for the sample mean using the formula:      z=M−μσMz = \frac{M - \mu}{\sigma_M}
    3. Consult the unit normal table using the calculated zz-score to find corresponding proportions and probabilities.
  • Interpretation of z-Scores:
    • The sign (+ or −+\text{ or }-) indicates whether the sample mean is above or below the population mean (μ\mu).
    • The numeric value specifies the distance between MM and μ\mu in units of standard error (σM\sigma_M).

z-Score Formula for Sample Means

The Distribution of Sample Means for n = 16

The Middle 80% of the Distribution of Sample Means

Standard Error and Sample Size Dynamics

  • Discrepancy and Variation: Sampling error causes discrepancy between sample means and population means. Standard error measures the extent of this variability across samples.
  • Sample Size Effect Illustrated: Assuming a population mean μ=80\mu = 80 and standard deviation σ=20\sigma = 20:
    • For sample size n=1n = 1: σM=201=20\sigma_M = \frac{20}{\sqrt{1}} = 20
    • For sample size n=4n = 4: σM=204=10\sigma_M = \frac{20}{\sqrt{4}} = 10
    • For sample size n=100n = 100: σM=20100=2\sigma_M = \frac{20}{\sqrt{100}} = 2

The Distribution of Sample Means for n = 1, n = 4, and n = 100

Application to Inferential Statistics

  • Role of Sample Data: Inferential statistics uses sample statistics to draw broad conclusions regarding target population parameters.
  • Impact of Sampling Error: Natural variability between samples and populations introduces uncertainty and error into inferential testing procedures.
  • Worked Example Structure (Treatment Study on Adult Rats):
    • Baseline Population: Weights for adult rats are normally distributed with μ=400\mu = 400 and σ=20\sigma = 20.
    • Treatment Experiment: A sample of n=25n = 25 rats is selected, treated, and measured.
    • Standard Error Calculation: σM=2025=205=4\sigma_M = \frac{20}{\sqrt{25}} = \frac{20}{5} = 4.
    • Expected Central 95%95\% Limits (z=±1.96z = \pm 1.96):
    • Lower Limit: M=400+(−1.96)(4)=392.16M = 400 + (-1.96)(4) = 392.16
    • Upper Limit: M=400+(+1.96)(4)=407.84M = 400 + (+1.96)(4) = 407.84
    • An untreated sample mean falling within 392.16392.16 and 407.84407.84 represents expected sampling variation, whereas a sample mean outside this range indicates a significant treatment effect.

The Structure of the Research Study Described in Example 7.7

The Distribution of Sample Means for Samples of n = 25 Untreated Rats

Practice Problems and Learning Checks

  • Learning Check 1:

    • Problem: A population has μ=60\mu = 60 with σ=5\sigma = 5. The distribution of sample means for samples of size n=4n = 4 selected from this population would have an expected value of _____?
    • Options: 55, 6060, 3030, 1515
    • Answer: 6060 (The expected value of MM is μM=μ=60\mu_M = \mu = 60).
    • True/False Statement 1: The shape of a distribution of sample means is always normal.
    • Answer: False (It is normal only if the population is normal or if n≥30n \ge 30).
    • True/False Statement 2: As sample size increases, the value of the standard error decreases.
    • Answer: True (Standard error is inversely proportional to n\sqrt{n}).
  • Learning Check 2:

    • Problem: A random sample of n=16n = 16 scores is obtained from a population with μ=50\mu = 50 and σ=16\sigma = 16. If the sample mean is M=58M = 58, the zz-score corresponding to the sample mean is _____?
    • Options: z=1.00z = 1.00, z=2.00z = 2.00, z=4.00z = 4.00, cannot determine
    • Calculation:       σM=σn=1616=4\sigma_M = \frac{\sigma}{\sqrt{n}} = \frac{16}{\sqrt{16}} = 4z=M−μσM=58−504=84=2.00z = \frac{M - \mu}{\sigma_M} = \frac{58 - 50}{4} = \frac{8}{4} = 2.00
    • Answer: z=2.00z = 2.00
    • True/False Statement 1: A sample mean with z=3.00z = 3.00 is a fairly typical, representative sample.
    • Answer: False (A zz-score of +3.00+3.00 represents an extreme, rare outcome far in the tail of the distribution).
    • True/False Statement 2: The mean of the sample is always equal to the population mean.
    • Answer: False (Individual sample means vary due to sampling error; only the expected value μM\mu_M, the average of all possible sample means, equals μ\mu).

Clear Your Doubts, Ask Questions Graphic