Hypothesis Testing Vocabulary and Logical Foundations

Introduction to Hypothesis Testing Vocabulary

  • General Concept of Hypothesis Testing: Hypothesis testing is not a single isolated concept, but rather a "big package of concepts" where every component has a specific name and role. It is similar in complexity to other statistical packages such as sampling or the distribution of sample means.

  • Learning Objectives: The primary goal is to build a foundational vocabulary by referencing specific examples—primarily analyzing a scenario comparing Chaffey College student ages to Mt. SAC student ages—and labeling the constituent parts of the hypothesis test.

The Null Hypothesis (H0H_0) and Test Values

  • Definition of Null Hypothesis (H0H_0): The null hypothesis is the starting assumption in a hypothesis test. It represents the status quo or the baseline against which data is compared.

    • The symbol for the null hypothesis is H0H_0, where the subscript is the number zero (representing "null," "nothing," or "empty"), though it is often colloquially referred to as "H-O."
  • Example Analysis (Scenario Two):

    • Context: A researcher is interested in the population of Chaffey College students (the population of interest) but lacks specific parameters like the population mean (μ\mu) or standard deviation (σ\sigma) for this group. However, parameters for a similar group, Mt. SAC students, are known (μ=26\mu = 26, σ=6\sigma = 6).
    • The Starting Assumption: Because information for Chaffey College students is unavailable, the researcher assumes that Chaffey College students are similar to Mt. SAC students.
    • Formal Statement: H0:μChaffey=26H_0: \mu_{\text{Chaffey}} = 26.
    • Symbolic Precision: We assume μChaffey=μMt. SAC=26\mu_{\text{Chaffey}} = \mu_{\text{Mt. SAC}} = 26. The subscript in the notation identifies the group to which the mean belongs (e.g., Chaffey) rather than the variable being measured (years\text{years}).
  • The Test Value: The test value is the specific number originating from the null hypothesis that is being tested. In the provided example, the test value is 2626. This is the number used in calculations (like Z-scores) to determine if the starting assumption holds up against observed data.

The P-Value and Conditional Probability

  • Formal Definition of P-Value: In hypothesis testing, the P-value is the conditional probability of observing a sample mean as extreme as, or more extreme than, the one actually observed, granted that the null hypothesis (H0H_0) is true.

  • Definition of "Extreme": In statistics, extreme means "far from the mean" or "far from the average." In the age example, if μ=26\mu = 26, a sample mean of 3030 is extreme because it is 44 years above the mean. Some tests look for extremes in one direction (e.g., greater than 3030), while others look for extremes in either direction (i.e., more than 44 years away from the average in both directions).

  • The Conditional Nature of P-Values:

    • The P-value is a "conditional" or "dependent" probability because it is entirely based on the assumption that the null hypothesis is true.
    • If the null hypothesis is incorrect (e.g., if the true μChaffey\mu_{\text{Chaffey}} is 27.527.5 instead of the assumed 2626), then the calculated Z-scores and P-values are "garbage" or meaningless.
    • Example Scenario: In the study of n=36n = 36 Chaffey students, a sample mean above 3030 resulted in a P-value of 0.000030.00003 (or 0.003%0.003\%). This probability is only valid if the assumed μ=26\mu = 26 is correct.

Logical Fallacies and Bayesian Statistics

  • Flipping the Conditional: A common mistake, even among professional researchers, is flipping the logic of the conditional probability.

    • Correct Logic: "If the null hypothesis is true, what is the chance of seeing what I saw?"
    • Incorrect Logic: "If I saw this result, what is the probability that the null hypothesis is true?"
    • Comparison to Bayesian Statistics: The latter approach (finding the probability of the hypothesis based on data) is a different branch of math known as Bayesian statistics, which has only recently become popular compared to traditional hypothesis testing.
  • The "Spot the Dog" Metaphor:

    • Statement A: "If I have a pet named Spot, what is the probability it is a dog?" (High, perhaps 50%50\%\n * Statement B: "If I have a dog, what is the probability its name is Spot?" (Very low, perhaps 2%2\%\n * Conclusion: You cannot flip the condition and the probability and expect the result to remain the same.
  • The COVID Metaphor:

    • Condition: "If I have COVID, what are the chances of not feeling well?" (High).
    • Flip: "If I do not feel well, what are the chances I have COVID?" (Low, because of other illnesses).

The Logic of Rejecting the Null Hypothesis

  • Relationship between P-value and Rejection: The smaller the P-value, the more likely a researcher is to reject the starting assumption (H0H_0).

    • Rationale: If you observe something that has a one-in-a-million chance of happening based on your assumption, it is more logical to conclude that your assumption is wrong than to believe you just witnessed a one-in-a-million miracle.
  • The Weatherman Metaphor:

    • If the weatherman claims there is a 10%10\% chance of rain and it rains, you might assume it was just luck.
    • If the weatherman claims a 1%1\% or a one-in-a-million chance and it rains, you will likely reject the starting assumption that the weatherman is reliable.
  • The New Friend Metaphor:

    • Starting Assumption (H0H_0): A new person is easy to get along with.
    • Observation: You see them screaming at someone. If you see this five times, the probability of them being "nice" but just having five bad days becomes so low that you reject the assumption that they are nice.
  • The Lottery/Scratcher Metaphor:

    • Starting Assumption (H0H_0): Your friend is competent and honest at reading a lottery ticket.
    • Scenario A: Friend says they won $10\$10. (Reasonable; keep H0H_0).
    • Scenario B: Friend says they won $10,000\$10,000. (Extremely low probability; you reject H0H_0 and check the ticket yourself).

Significance Levels (Alpha)

  • The Need for a Cutoff: Because different people have different personal thresholds for when a probability becomes "small enough," research requires a formal, binary cutoff. This is similar to the Central Limit Theorem's rule of thumb where n≥30n \geq 30 is considered "normal enough."

  • Significance Level (α\alpha): This is the Greek letter Alpha used to represent the formal cutoff for rejecting H0H_0.

  • Standard Cutoffs:

    • α=0.05\alpha = 0.05: A 5%5\% chance of an event happening by chance. This is the common standard in research.
    • α=0.01\alpha = 0.01: A 1%1\% chance of an event happening. This is used for extra-rigorous proof.
  • Decision Rule:

    • If P-value<αP\text{-value} < \alpha, Reject H0H_0.
    • If P-value≥αP\text{-value} \geq \alpha, Fail to Reject H0H_0.
  • Comparison Example:

    • If an observed P-value is 0.02280.0228:
      • Compare to α=0.05\alpha = 0.05: Since 0.0228<0.050.0228 < 0.05, we reject H0H_0.
      • Compare to α=0.01\alpha = 0.01: Since 0.02280.0228 is not less than 0.010.01, we fail to reject H0H_0.
  • Standardization in Research: The use of arbitrary numbers like 0.050.05 and 0.010.01 (possibly chosen because humans have five fingers) allows all researchers—biologists, psychologists, sociologists—to be held to the same "standard of proof." It creates common rules for the "game" of research.

Z-Obtained (ZobtZ_{\text{obt}}) vs. Z-Critical (ZcritZ_{\text{crit}})

  • Dual Z-Scores: Every hypothesis test involving proportions or means involves two different types of Z-scores:

    1. Z-Obtained (ZobtZ_{\text{obt}}): This is the Z-score calculated directly from the observed data or sample (also called ZobservedZ_{\text{observed}}). In the Chaffey example, the calculated ZobtZ_{\text{obt}} was +4+4.
    2. Z-Critical (ZcritZ_{\text{crit}}): This is the Z-score that marks the boundary of the significance level (α\alpha). It is the "critical point" or the "cutoff" point on the distribution tail.
  • Rejecting H0H_0 via Z-Scores:

    • A researcher rejects H0H_0 if the ZobtZ_{\text{obt}} falls into the "shaded tail area" past the ZcritZ_{\text{crit}}.
    • This logic always matches the P-value logic: if P<αP < \alpha, then the obtained value will inherently be in the shaded area past the critical value. They are two different ways of saying the same thing.
  • Software vs. Hand Calculation:

    • Software (e.g., SPSS): Primarily uses P-values and Alphas for interpretation.
    • Hand Calculation: Traditionally focuses on comparing ZobtZ_{\text{obt}} to ZcritZ_{\text{crit}} by shading the tails of the distribution.
  • Future Applications: While this chapter focuses on Z-scores, future chapters will apply the same "Obtained vs. Critical" logic to other tests, such as t-tests (tt), F-tests (FF), and Chi-square tests (χ2\chi^2).