Introduction to Hypothesis Testing and the One-Sample Z-Test

Overview of Chapter 8: The Core of Statistics

  • Importance of Chapter 8: This chapter represents the absolute "heart" of the course.
  • Building Blocks: All material covered in chapters one through seven serves as the necessary building blocks to reach this point.
  • Future Course Outlook: Every single topic and chapter following Chapter 8 will revolve around hypothesis testing. Each subsequent chapter will present different types of hypothesis tests tailored to specific situations, research designs, and study types.
  • Introductory Focus: Today’s lecture focuses on hypothesis testing in general through a specific example called the one-sample z-test.

Course Materials and Study Strategy

  • Course Packet Handouts: The lecture follows the "Chapter 8 Day 1 Lecture Handout" found on page 75 of the course packet (based on the Fall 2020 version).
  • Essential Handouts:
    • Steps for One-Sample Z-Test: Primarily useful for laboratory assignments.
    • Hypothesis Test Reminders Handout: Described as the most important handout in the class. It provides general principles that apply to all hypothesis testing topics for the remainder of the semester.
  • Study Recommendation: Because hypothesis testing is the foundation for everything moving forward, students are advised to read the textbook for Chapter 8 at least three times. This aligns with the three scheduled instruction days (Day 1, Day 2, and the Extra Day).

Review of the Distribution of Sample Means (Chapter 7)

  • Hypothesis Testing Foundation: Hypothesis testing is built directly upon the concept of the distribution of sample means developed in Chapter 7.
  • Definition: The distribution of sample means is created by taking all possible samples of size nn (where nn is the number of scores in each sample) from a population and calculating the mean (MM) for each.
  • Scale of Samples: In real-world populations (e.g., millions of students), the number of possible samples is unmanageable (millions, billions, or trillions). This makes the Central Limit Theorem essential for practical application.
  • Key Symbols:
    • Sample Mean: MM
    • Population Mean: μ\mu
    • Standard Deviation: σ\sigma

The Central Limit Theorem (CLT)

The CLT identifies three critical characteristics of the distribution of sample means:

  1. The Center (Central Tendency):

    • Known as the Expected Value of M.
    • Regardless of the parent population size, the mean of all sample means (MM) will always equal the mean of the parent population (μ\mu).
    • Formula: E(M)=μE(M) = \mu
  2. The Width (Variability):

    • Known as Standard Error (σM\sigma_M). It is essentially the standard deviation of a distribution of means.
    • Relates the spread of sample means back to the parent population standard deviation divided by the square root of the sample size.
    • Formula: σM=σn\sigma_M = \frac{\sigma}{\sqrt{n}}
  3. The Shape:

    • The distribution is considered "Normal" if at least one of two conditions is met:
      • The parent population is normal.
      • The sample size (nn) is at least 25 to 30 (a common rule of thumb).
    • Statistical accuracy increases as nn approaches infinity, though 25-30 is sufficient for most research purposes.

Chapter 7 Example: Probability Assessment

  • Scenario: Chaffey College Research Office data shows the average age of all students (μ\mu) is 26 years26\,\text{years} with a standard deviation (σ\sigma) of 6 years6\,\text{years}.
  • Research Question: What is the probability that a random sample of 3636 students (n=36n = 36) will have a mean age (MM) greater than 3030?
  • Calculation Steps:
    • Identify Distribution: Since n=36n=36, the distribution of sample means is normal.
    • Calculate Standard Error:     σM=636=66=1\sigma_M = \frac{6}{\sqrt{36}} = \frac{6}{6} = 1
    • Calculate Z-score:     z=30−261=+4.00z = \frac{30 - 26}{1} = +4.00
    • Determine Probability: Looking up a z-score of 4.004.00 in the unit normal table (column C for the tail) gives a proportion of 0.000030.00003.
    • Percent chance: 0.003%0.003\%
  • Interpretation: This result is an extreme outlier. In a world where you drew 100,000 such samples, only 3 would have a mean age above 30. If this actually happened, you would be extremely surprised; it is equivalent to winning the lottery.

Transition to Hypothesis Testing: Conditional Logic

  • The Tweak: Imagine we do not know the actual μ\mu and σ\sigma for Chaffey College. Instead, we look at Mt. San Antonio College (Mt. SAC), which reports μ=26\mu = 26 and σ=6\sigma = 6.
  • The Starting Assumption: We assume Chaffey students are identical to Mt. SAC students (μ=26\mu = 26, σ=6\sigma = 6).
  • The Result: If we draw a sample of 36 Chaffey students and get a mean of 30, the math remains the same (z=4.00z = 4.00, p=0.003%p = 0.003\%). However, the logic shifts from absolute to conditional.
  • Conditionality: The probability is only 0.003%0.003\% if our assumption about Chaffey being like Mt. SAC is true.

The Two Possible Conclusions

When a researcher observes an extremely unlikely result (low probability), they are faced with two logical choices:

  1. Stick to the Assumption: Maintain that the starting assumption (e.g., Chaffey = Mt. SAC) is true and conclude that an incredibly rare, "miraculous" event simply occurred.
  2. Reject the Assumption: Conclude that the starting assumption was likely incorrect. If the actual mean for Chaffey was higher (e.g., 28 or 29), finding a sample mean of 30 would not have been nearly as unlikely.

Real-World Metaphors for Conditional Logic

  • The Weatherman: If a meteorologist says there is a 1%1\% chance of rain and it pours, you can either believe the weatherman was right and a rare event happened, or conclude the weatherman's model (the starting assumption) was wrong.
  • The 2016 Election: Polls gave one candidate a 90%90\% chance and the other a 10%10\% chance. When the underdog won, critics either claimed the polls were fundamentally flawed (rejected the assumption) or recognized that a 10%10\% event happened (unlikely things do occur).
  • Sports Betting: When a heavy underdog wins, it represents either a failure in the odds-making logic or the occurrence of a low-probability miraculous victory.