Chapter 3 (Part 1)

Chapter Three: Sampling Distributions and Inferential Statistics

The Normal Distribution

  • Introduction to the Normal Distribution
    • The concept of a normal distribution or curve was introduced in Chapters One and Two, and will be discussed in further detail in this chapter.
    • The normal curve will be heavily relied upon and extensively used throughout the text.
    • Datasets will typically be assumed to be normally distributed unless there is specific evidence to the contrary.
Importance of Normal Distributions
  • Normal distributions are fundamentally important in statistics for two primary reasons:
    1. Common Distribution Shape: Measurements of many naturally occurring phenomena, including psychological constructs like intelligence, anxiety, and mood, are commonly normally distributed.
    2. Behavior of Sample Means: If we take a sample of scores from any shaped population and calculate the mean (M) multiple times, the distribution of these sample means will approach a normal distribution as the sample size increases.
    • This observation is crucial for most statistical analyses used to test hypotheses.

Characteristics of Normal Distributions

  • For a distribution to be classified as normal, it must conform to a specific mathematical model defined by:

    • Formulation:
    • The ordinate (y) represents the height of the curve for a given x.
    • An uppercase X denotes any given score.
    • The lowercase mu (BC) represents the population mean.
    • The lowercase sigma squared (C3^2) represents the population variance.
    • The value of pi (C0) is approximately 3.1416.
    • The value of e (approximately 2.718) is the base of natural logarithms.
  • Key Points:

    1. Shape: The normal distribution is unimodal (has one peak).
    2. Symmetry: The curve is symmetrical; the right half is a mirror image of the left half.
    3. Central Tendency: The mean, median, and mode all have the same value in a normal distribution due to its symmetry and unimodality.
    4. Asymptotic Nature: The tails of the distribution never touch the horizontal axis (abscissa). Though theoretically any x value may be placed in the equation, practical limitations exist (e.g., unrealistic height measurements).
    5. Empirical Rule: In a normal distribution:
    • Approximately 68% of scores lie within one standard deviation of the mean (extuextext±extextσext{u} ext{ } ext{±} ext{ } ext{σ}).
    • Approximately 95% of scores lie within two standard deviations of the mean (extuextext±ext2extσext{u} ext{ } ext{±} ext{ } 2 ext{σ}).
    • Approximately 99.7% of scores lie within three standard deviations of the mean (extuextext±ext3extσext{u} ext{ } ext{±} ext{ } 3 ext{σ}).
  • Number of Distributions: There are an infinite number of normal distributions which can vary based on different population means and variances. All share the five listed characteristics.

Historical Background of the Normal Curve

  • The normal curve's discovery is attributed to Abraham de Moivre, dating back to a publication in 1733.
    • He was friends with notable figures such as Edmund Halley and Sir Isaac Newton and was highly regarded in intellectual circles.
  • Initial Interest: De Moivre's work stemmed from the probability of outcomes in gambling scenarios.
    • For example, determining the probability of getting between 500 and 600 heads in 1000 coin tosses results in a normal distribution as the number of tosses increases.
    • Specific methodological details from de Moivre's work remain unclear since it was customary to obscure methods.

The Central Limit Theorem (CLT)

  • The Central Limit Theorem, an essential concept in statistics, states:
    • The means of samples taken from any distribution will be normally distributed as the number of samples increases.
    • This theorem is crucial for conducting sampling distributions and hypothesis testing, to be explored in greater detail in later chapters.
    • Mathematical Foundations: Developed through contributions from various mathematicians, including:
    • Thomas Simpson: who extended the normal curve to continuous measures and emphasized the importance of minimizing error through averaging measurements.
    • Pierre Laplace: proved the central limit theorem and emphasized its broad applications beyond gambling, including astronomical observations.
    • Karl Gauss: also known as the Gaussian distribution, contributed to the methodology of predicting outcomes in various fields.
    • Significance: Gauss's contributions included the method of least squares.

Area Under the Normal Curve

  • Representation: A simple frequency distribution can be represented visually through a histogram or frequency polygon.
    • The concept of probability relates to the area under the curve, where the total area equals one.
    • Given that normal curves are symmetrical, the probability of obtaining a score above or below the mean is equal.
  • Considerations of probability will span values from 0 to 1 (where 0 indicates impossibility and 1 indicates certainty).
  • The probabilities and percentages discussed in this chapter are linked; for example:
    • Saying 50% of scores fall above the mean is equivalent to a probability of 0.5 for selecting a score above the mean.

Standard Scores (z-scores)

  • Definition: A z-score standardizes scores to enable comparisons across different distributions, regardless of the mean or variance.
    • The transformation converts original scores into a common unit, providing insights into the relative positions of scores within their distributions.
  • Relationships: A z-score indicates how many standard deviations a score is from the mean:
    • Example: Calculating z-scores for heights and weights allows for comparisons (e.g., a height z-score of +1.3 vs. a weight z-score of -0.42).
  • Applications: Z-scores allow for comparisons in various fields, such as education and psychological assessments, where different scales are used for different traits.
  • To compute a z-score:
    • For a population: z=xextuextoz = \frac{x - ext{u}}{ ext{o}}
    • For a sample: z=xextMextsz = \frac{x - ext{M}}{ ext{s}}
      where,
    • x = raw score
    • BC = population mean
    • o = population standard deviation
    • M = sample mean
    • s = sample standard deviation
  • Characteristics of z-scores:
    1. Mean of z-score distribution is 0, with a standard deviation of 1.
    2. Negative z-scores correspond to values below the mean; positive z-scores correspond to values above the mean.

3. All raw scores that equal the mean have a z-score of 0, while scores at one standard deviation from the mean are either +1 or -1, based on their position relative to the mean.