Comprehensive Study Notes on Population Proportions: Confidence Intervals and Hypothesis Testing

Theoretical Foundations of Population Proportions

  • The Statistical Setting:

    • Take a Simple Random Sample (SRS) of size nn from a large population containing an unknown proportion of "successes," denoted by pp.
  • Large Sample Confidence Interval for Population Proportion:

    • An approximate level CC confidence interval for the unknown proportion pp is calculated as:         p^±zp^(1p^)n\hat{p} \pm z^* \sqrt{\frac{\hat{p}(1-\hat{p})}{n}}
    • Structure of the Interval:
      • The formula follows the standard format: Estimate±Margin of Error\text{Estimate} \pm \text{Margin of Error}.
      • The Margin of Error consists of: (Critical Value)×(Standard Error)(\text{Critical Value}) \times (\text{Standard Error}).
    • The Critical Value (zz^*): This is the level CC critical value, representing the area CC under the density curve between z-z^* and zz^* in the zz distribution with the appropriate degrees of freedom.
    • Usage Requirement: This method for a confidence interval should only be used when the hard count of the number of successes and the number of failures in the sample data are both at least 1515.

Significance Tests for Population Proportion

  • Test Statistic: To test the null hypothesis that p=p0p = p_0, the sample zz statistic (zp^z_{\hat{p}}) is used in the standard normal distribution:     zp^=p^p0p0(1p0)nz_{\hat{p}} = \frac{\hat{p}-p_0}{\sqrt{\frac{p_0(1-p_0)}{n}}}

  • Null and Alternative Hypotheses:

    • Null Hypothesis (H0H_0): Always takes the form H0:p=p0H_0: p = p_0. It establishes a set of expectations/benchmarks against which sample data is weighed.
    • Alternative Hypothesis (HaH_a): Defines what we are looking for (the suspicion).
      • One-sided (larger): Ha:p>p0H_a: p > p_0 (The true proportion of successes is greater than p0p_0; Right-tailed).
      • One-sided (smaller): Ha:p<p0H_a: p < p_0 (The true proportion of successes is less than p0p_0; Left-tailed).
      • Two-sided (different): Ha:pp0H_a: p \neq p_0 (The true proportion of successes is different from p0p_0; Two-tailed).
  • Conditions for Significance Tests:

    1. Randomness: Sample data must come from an SRS (or at least a random and representative sample).
    2. Sample Size/Large Enough: Both the hypothesized number of successes and failures must be at least 1010.
      • np010n p_0 \geq 10
      • n(1p0)10n(1 - p_0) \geq 10

Example 1: Home Field Advantage in Major League Baseball

  • Problem Context: Frequent tournaments attempt to neutralize "home field advantage." If no such advantage existed, home teams would win approximately 50%50\% of all games played.
  • Data Analysis (2013 MLB Season):
    • Total Games (nn): 24312431
    • Home Team Wins (xx): 13081308
    • Sample Proportion (p^\hat{p}): 53.81%53.81\%
  • Research Question: Does the deviation from 50%50\% represent natural sampling variability, or is it evidence of a true home field advantage in professional baseball?

Case Study 2: Binge Drinking Among College Students

  • National Benchmark: According to the National Institute on Alcohol Abuse and Alcoholism, 41%41\% of college students nationwide engage in binge drinking behavior (five or more drinks on one occasion within the past two weeks).
  • Study Details:
    • P (Population & Parameter):
      • Population: All students enrolled at the specific college.
      • Parameter (pp): The true proportion of students at this college who engage in binge drinking.
    • H (Hypotheses):
      • H0:p=0.41H_0: p = 0.41
      • Ha:p<0.41H_a: p < 0.41 (The true proportion is less than the national average).
    • A (Assumptions and Conditions):
      1. Random Samples: 462462 students selected randomly from an enrollment list.
      2. Large Enough:
        • np0=462×0.41189.4210n p_0 = 462 \times 0.41 \approx 189.42 \geq 10
        • n(1p0)=462×0.59272.5810n(1 - p_0) = 462 \times 0.59 \approx 272.58 \geq 10
    • N (Name the Test): One-proportion zz-test.
    • T (Test Statistic): z=2.59z = -2.59. This value follows the standard normal distribution.
    • O (Obtain P-value): P-value=0.005P\text{-value} = 0.005.
    • M (Make Decision):
      • Significance Level (α\alpha): 0.050.05.
      • Comparison: 0.005<0.050.005 < 0.05.
      • Decision: Reject the null hypothesis in favor of the alternative. The evidence is statistically significant.
    • S (Summary): There is statistically significant evidence at the α=0.05\alpha = 0.05 level that the proportion of students at this college who engage in binge drinking is less than 0.410.41.

Case Study 3: Changes in Smoking Behavior Since 1965

  • Context: In 1965, approximately 44%44\% of the U.S. adult population had never smoked.
  • Study Details (2010 Survey):
    • P (Population & Parameter):
      • Population: U.S. adult population in 2010.
      • Parameter (pp): The true proportion of adults who have never smoked.
      • Sample Data: n=1472n = 1472, Successes (xx) = 677677, p^0.4565\hat{p} \approx 0.4565.
    • H (Hypotheses):
      • H0:p=0.44H_0: p = 0.44 (Benchmark establish by 1965 results).
      • Ha:p0.44H_a: p \neq 0.44 (Testing for a change/difference since 1965).
    • A (Assumptions and Conditions):
      1. Random Samples: Unclear if the survey was an SRS, though it is presumably random.
      2. Large Enough:
        • np0=1472×0.44647.6810n p_0 = 1472 \times 0.44 \approx 647.68 \geq 10
        • n(1p0)=1472×0.56824.3210n(1 - p_0) = 1472 \times 0.56 \approx 824.32 \geq 10
    • N (Name the Test): One-sample, one-proportion zz-test.
    • T (Test Statistic): z=1.54z = 1.54.
    • O (Obtain P-value): P-value=0.1237P\text{-value} = 0.1237.
    • M (Make Decision):
      • Comparison: 0.1237>0.050.1237 > 0.05.
      • Decision: Fail to reject the null hypothesis. The P-valueP\text{-value} is not statistically significant.
    • S (Summary): We do not find statistically significant evidence that the proportion of U.S. adults who have never smoked has changed since 1965.

Case Study 4: Incidence of Congenital Abnormalities

  • Historical Context: In the 1980s, congenital abnormalities were believed to affect approximately 5%5\% of children.
  • Current Study:
    • P (Population & Parameter):
      • Population: Nations' children.
      • Parameter (pp): The true proportion of children with a congenital abnormality.
      • Sample Data: n=384n = 384, Successes (xx) = 4646, p^0.1197\hat{p} \approx 0.1197 (11.97%11.97\%\text{)}.
    • H (Hypotheses):
      • H0:p=0.05H_0: p = 0.05
      • Ha:p>0.05H_a: p > 0.05 (One-sided, testing for an increase since the 1980s).
    • A (Assumptions and Conditions):
      1. Random Samples: Unclear if data comes from a random sample.
      2. Large Enough:
        • np0=384×0.05=19.210n p_0 = 384 \times 0.05 = 19.2 \geq 10
        • n(1p0)=384×0.95=364.810n(1 - p_0) = 384 \times 0.95 = 364.8 \geq 10
    • N (Name the Test): One-proportion zz-test.
    • T (Test Statistic): z=6.28z = 6.28.
    • O (Obtain P-value): P-value=1.7×1010P\text{-value} = 1.7 \times 10^{-10}.
    • M (Make Decision):
      • Comparison: 1.7×1010<0.051.7 \times 10^{-10} < 0.05.
      • Decision: Reject the null hypothesis in favor of the alternative. Statistically significant evidence against the null.
    • S (Summary): We have statistically significant evidence that the incidence rate of congenital abnormalities has increased since the 1980s.