Chapter 3 (Part 1)
Chapter Three: Sampling Distributions and Inferential Statistics
The Normal Distribution
- Introduction to the Normal Distribution
- The concept of a normal distribution or curve was introduced in Chapters One and Two, and will be discussed in further detail in this chapter.
- The normal curve will be heavily relied upon and extensively used throughout the text.
- Datasets will typically be assumed to be normally distributed unless there is specific evidence to the contrary.
Importance of Normal Distributions
- Normal distributions are fundamentally important in statistics for two primary reasons:
- Common Distribution Shape: Measurements of many naturally occurring phenomena, including psychological constructs like intelligence, anxiety, and mood, are commonly normally distributed.
- Behavior of Sample Means: If we take a sample of scores from any shaped population and calculate the mean (M) multiple times, the distribution of these sample means will approach a normal distribution as the sample size increases.
- This observation is crucial for most statistical analyses used to test hypotheses.
Characteristics of Normal Distributions
For a distribution to be classified as normal, it must conform to a specific mathematical model defined by:
- Formulation:
- The ordinate (y) represents the height of the curve for a given x.
- An uppercase X denotes any given score.
- The lowercase mu (BC) represents the population mean.
- The lowercase sigma squared (C3^2) represents the population variance.
- The value of pi (C0) is approximately 3.1416.
- The value of e (approximately 2.718) is the base of natural logarithms.
Key Points:
- Shape: The normal distribution is unimodal (has one peak).
- Symmetry: The curve is symmetrical; the right half is a mirror image of the left half.
- Central Tendency: The mean, median, and mode all have the same value in a normal distribution due to its symmetry and unimodality.
- Asymptotic Nature: The tails of the distribution never touch the horizontal axis (abscissa). Though theoretically any x value may be placed in the equation, practical limitations exist (e.g., unrealistic height measurements).
- Empirical Rule: In a normal distribution:
- Approximately 68% of scores lie within one standard deviation of the mean ().
- Approximately 95% of scores lie within two standard deviations of the mean ().
- Approximately 99.7% of scores lie within three standard deviations of the mean ().
Number of Distributions: There are an infinite number of normal distributions which can vary based on different population means and variances. All share the five listed characteristics.
Historical Background of the Normal Curve
- The normal curve's discovery is attributed to Abraham de Moivre, dating back to a publication in 1733.
- He was friends with notable figures such as Edmund Halley and Sir Isaac Newton and was highly regarded in intellectual circles.
- Initial Interest: De Moivre's work stemmed from the probability of outcomes in gambling scenarios.
- For example, determining the probability of getting between 500 and 600 heads in 1000 coin tosses results in a normal distribution as the number of tosses increases.
- Specific methodological details from de Moivre's work remain unclear since it was customary to obscure methods.
The Central Limit Theorem (CLT)
- The Central Limit Theorem, an essential concept in statistics, states:
- The means of samples taken from any distribution will be normally distributed as the number of samples increases.
- This theorem is crucial for conducting sampling distributions and hypothesis testing, to be explored in greater detail in later chapters.
- Mathematical Foundations: Developed through contributions from various mathematicians, including:
- Thomas Simpson: who extended the normal curve to continuous measures and emphasized the importance of minimizing error through averaging measurements.
- Pierre Laplace: proved the central limit theorem and emphasized its broad applications beyond gambling, including astronomical observations.
- Karl Gauss: also known as the Gaussian distribution, contributed to the methodology of predicting outcomes in various fields.
- Significance: Gauss's contributions included the method of least squares.
Area Under the Normal Curve
- Representation: A simple frequency distribution can be represented visually through a histogram or frequency polygon.
- The concept of probability relates to the area under the curve, where the total area equals one.
- Given that normal curves are symmetrical, the probability of obtaining a score above or below the mean is equal.
- Considerations of probability will span values from 0 to 1 (where 0 indicates impossibility and 1 indicates certainty).
- The probabilities and percentages discussed in this chapter are linked; for example:
- Saying 50% of scores fall above the mean is equivalent to a probability of 0.5 for selecting a score above the mean.
Standard Scores (z-scores)
- Definition: A z-score standardizes scores to enable comparisons across different distributions, regardless of the mean or variance.
- The transformation converts original scores into a common unit, providing insights into the relative positions of scores within their distributions.
- Relationships: A z-score indicates how many standard deviations a score is from the mean:
- Example: Calculating z-scores for heights and weights allows for comparisons (e.g., a height z-score of +1.3 vs. a weight z-score of -0.42).
- Applications: Z-scores allow for comparisons in various fields, such as education and psychological assessments, where different scales are used for different traits.
- To compute a z-score:
- For a population:
- For a sample:
where, - x = raw score
- BC = population mean
- o = population standard deviation
- M = sample mean
- s = sample standard deviation
- Characteristics of z-scores:
- Mean of z-score distribution is 0, with a standard deviation of 1.
- Negative z-scores correspond to values below the mean; positive z-scores correspond to values above the mean.