Introduction to Hypothesis Testing and the One-Sample Z-Test
Overview of Chapter 8: The Core of Statistics
- Importance of Chapter 8: This chapter represents the absolute "heart" of the course.
- Building Blocks: All material covered in chapters one through seven serves as the necessary building blocks to reach this point.
- Future Course Outlook: Every single topic and chapter following Chapter 8 will revolve around hypothesis testing. Each subsequent chapter will present different types of hypothesis tests tailored to specific situations, research designs, and study types.
- Introductory Focus: Today’s lecture focuses on hypothesis testing in general through a specific example called the one-sample z-test.
Course Materials and Study Strategy
- Course Packet Handouts: The lecture follows the "Chapter 8 Day 1 Lecture Handout" found on page 75 of the course packet (based on the Fall 2020 version).
- Essential Handouts:
- Steps for One-Sample Z-Test: Primarily useful for laboratory assignments.
- Hypothesis Test Reminders Handout: Described as the most important handout in the class. It provides general principles that apply to all hypothesis testing topics for the remainder of the semester.
- Study Recommendation: Because hypothesis testing is the foundation for everything moving forward, students are advised to read the textbook for Chapter 8 at least three times. This aligns with the three scheduled instruction days (Day 1, Day 2, and the Extra Day).
Review of the Distribution of Sample Means (Chapter 7)
- Hypothesis Testing Foundation: Hypothesis testing is built directly upon the concept of the distribution of sample means developed in Chapter 7.
- Definition: The distribution of sample means is created by taking all possible samples of size (where is the number of scores in each sample) from a population and calculating the mean () for each.
- Scale of Samples: In real-world populations (e.g., millions of students), the number of possible samples is unmanageable (millions, billions, or trillions). This makes the Central Limit Theorem essential for practical application.
- Key Symbols:
- Sample Mean:
- Population Mean:
- Standard Deviation:
The Central Limit Theorem (CLT)
The CLT identifies three critical characteristics of the distribution of sample means:
The Center (Central Tendency):
- Known as the Expected Value of M.
- Regardless of the parent population size, the mean of all sample means () will always equal the mean of the parent population ().
- Formula:
The Width (Variability):
- Known as Standard Error (). It is essentially the standard deviation of a distribution of means.
- Relates the spread of sample means back to the parent population standard deviation divided by the square root of the sample size.
- Formula:
The Shape:
- The distribution is considered "Normal" if at least one of two conditions is met:
- The parent population is normal.
- The sample size () is at least 25 to 30 (a common rule of thumb).
- Statistical accuracy increases as approaches infinity, though 25-30 is sufficient for most research purposes.
- The distribution is considered "Normal" if at least one of two conditions is met:
Chapter 7 Example: Probability Assessment
- Scenario: Chaffey College Research Office data shows the average age of all students () is with a standard deviation () of .
- Research Question: What is the probability that a random sample of students () will have a mean age () greater than ?
- Calculation Steps:
- Identify Distribution: Since , the distribution of sample means is normal.
- Calculate Standard Error:
- Calculate Z-score:
- Determine Probability: Looking up a z-score of in the unit normal table (column C for the tail) gives a proportion of .
- Percent chance:
- Interpretation: This result is an extreme outlier. In a world where you drew 100,000 such samples, only 3 would have a mean age above 30. If this actually happened, you would be extremely surprised; it is equivalent to winning the lottery.
Transition to Hypothesis Testing: Conditional Logic
- The Tweak: Imagine we do not know the actual and for Chaffey College. Instead, we look at Mt. San Antonio College (Mt. SAC), which reports and .
- The Starting Assumption: We assume Chaffey students are identical to Mt. SAC students (, ).
- The Result: If we draw a sample of 36 Chaffey students and get a mean of 30, the math remains the same (, ). However, the logic shifts from absolute to conditional.
- Conditionality: The probability is only if our assumption about Chaffey being like Mt. SAC is true.
The Two Possible Conclusions
When a researcher observes an extremely unlikely result (low probability), they are faced with two logical choices:
- Stick to the Assumption: Maintain that the starting assumption (e.g., Chaffey = Mt. SAC) is true and conclude that an incredibly rare, "miraculous" event simply occurred.
- Reject the Assumption: Conclude that the starting assumption was likely incorrect. If the actual mean for Chaffey was higher (e.g., 28 or 29), finding a sample mean of 30 would not have been nearly as unlikely.
Real-World Metaphors for Conditional Logic
- The Weatherman: If a meteorologist says there is a chance of rain and it pours, you can either believe the weatherman was right and a rare event happened, or conclude the weatherman's model (the starting assumption) was wrong.
- The 2016 Election: Polls gave one candidate a chance and the other a chance. When the underdog won, critics either claimed the polls were fundamentally flawed (rejected the assumption) or recognized that a event happened (unlikely things do occur).
- Sports Betting: When a heavy underdog wins, it represents either a failure in the odds-making logic or the occurrence of a low-probability miraculous victory.