Lecture 7: Binomial Probability Distribution and Hypothesis Testing Notes

Categorical Data and Binomial Probabilities

  • Categorical data consists of variables categorized into discrete, non-numeric groups. In binary categorical data, observations fall into exactly two mutually exclusive categories:

    • Live or dead

    • Success or failure

    • Affected by disease or not affected

  • Binomial probability tools are utilized to analyze binary categorical data and test hypotheses regarding underlying population probabilities.

Mendelian Genetics Example and Sample Space

  • Single-gene inheritance traits governed by dominant (AA) and recessive (aa) alleles serve as a foundational application of binomial probability.

  • Assuming random mating and equal allele frequencies (P(A)=0.5P(A) = 0.5 and P(a)=0.5P(a) = 0.5):

    • Alleles inherited from the mother: AA or aa

    • Alleles inherited from the father: AA or aa

    • Resulting offspring genotypes: AAAA, aAaA, AaAa, and aaaa

    • Probability of recessive genotype (aaaa): P(R)=0.25P(R) = 0.25

    • Probability of dominant phenotype (AAAA, AaAa, aAaA): P(D)=0.75P(D) = 0.75

Punnett square showing Mendelian inheritance for a single gene with dominant and recessive alleles
  • When sampling 3 individuals randomly from this population, the complete sample space of outcomes for recessive (RR) and dominant (DD) phenotypes comprises 23=82^3 = 8 possibilities:

    • if the probability of one phenotype doesn’t directly relate to the probability of another, then they are considered independent

    • faster way to do this is through binomial distribution

      • RRRRRR: (0.25)×(0.25)×(0.25)=0.015625(0.25) \times (0.25) \times (0.25) = 0.015625

      • RRDRRD: (0.25)×(0.25)×(0.75)=0.046875(0.25) \times (0.25) \times (0.75) = 0.046875

      • RDRRDR: (0.25)×(0.75)×(0.25)=0.046875(0.25) \times (0.75) \times (0.25) = 0.046875

      • RDDRDD: (0.25)×(0.75)×(0.75)=0.140625(0.25) \times (0.75) \times (0.75) = 0.140625

      • DRRDRR: (0.75)×(0.25)×(0.25)=0.046875(0.75) \times (0.25) \times (0.25) = 0.046875

      • DRDDRD: (0.75)×(0.25)×(0.75)=0.140625(0.75) \times (0.25) \times (0.75) = 0.140625

      • DDRDDR: (0.75)×(0.75)×(0.25)=0.140625(0.75) \times (0.75) \times (0.25) = 0.140625

      • DDDDDD: (0.75)×(0.75)×(0.75)=0.421875(0.75) \times (0.75) \times (0.75) = 0.421875

Probability tree for sampling three individuals with recessive or dominant phenotypes

Mathematical Formulation of the Binomial Distribution

  • The binomial distribution describes the probability of obtaining a given number of "successes" (XX) from a fixed number of independent trials (nn).

  • Individual trials in a binomial process are known as Bernoulli trials, defined as random variables with exactly two possible outcomes.

  • Parameters of the binomial distribution:

    • nn: total number of independent trials

    • pp: probability of success on any single trial

    • 1−p1 - p: probability of failure on any single trial

    • XX: observed number of successful trials (X = 0, 1, 2, \ndots, n)

  • General formula for binomial probability:   P(X)=(nX)pX(1−p)n−XP(X) = \binom{n}{X} p^X (1 - p)^{n - X}

  • Combinatorial notation (nX)\binom{n}{X} (also written as nCXnCX or nCX_nC_X) represents the number of ways to choose XX successes from nn trials:   (nX)=n!X!(n−X)!\binom{n}{X} = \frac{n!}{X! (n - X)!}

  • Factorial definition and rules:

    • n!=n×(n−1)×(n−2)×⋯×3×2×1n! = n \times (n - 1) \times (n - 2) \times \dots \times 3 \times 2 \times 1

    • 0!=10! = 1

    • 1!=11! = 1

    • Example: 6!=6×5×4×3×2×1=7206! = 6 \times 5 \times 4 \times 3 \times 2 \times 1 = 720

Worked Calculation: Sampling Recessive Phenotypes

  • To calculate the probability that exactly 22 out of 33 randomly sampled individuals express a recessive phenotype (p=0.25p = 0.25):

    • Number of combinations yielding 2 recessives: (32)\binom{3}{2}

    • Probability of 2 recessive phenotypes: (0.25)2(0.25)^2

    • Probability of 1 dominant phenotype: (1−0.25)3−2=(0.75)1(1 - 0.25)^{3 - 2} = (0.75)^1

  • Step-by-step substitution and evaluation:   P(X=2)=(32)(0.25)2(1−0.25)3−2P(X = 2) = \binom{3}{2} (0.25)^2 (1 - 0.25)^{3 - 2}   P(X=2)=3!2!×1!(0.25)2(0.75)1P(X = 2) = \frac{3!}{2! \times 1!} (0.25)^2 (0.75)^1   P(X=2)=3×(0.25)2×(0.75)P(X = 2) = 3 \times (0.25)^2 \times (0.75)   P(X=2)=3×0.0625×0.75=0.140625≈0.141P(X = 2) = 3 \times 0.0625 \times 0.75 = 0.140625 \approx 0.141

Hypothesis Testing Example: Red Uniforms in Combat Sports

  • Biological Context: Male animals frequently display bright red colorations as visual signals of aggression, competitive dominance, and threat intensity.

Mandrill displaying prominent red and blue facial pigmentation as an aggressive signalOlympic wrestlers competing in red and blue uniform singlets
  • Empirical Investigation: A study evaluated whether uniform color influences outcome success in human combat sports (Hill, RA, and RA Barton 2005. "Red enhances human performance in contests." Nature 435:293).

    • Dataset: Examined round outcomes in wrestling, taekwondo, and boxing during the 2004 Olympic Games.

    • Experimental Design: Competitors were randomly assigned red or blue shirt colors.

    • Experimental Findings: In 1616 of 2020 combat rounds, the athlete assigned the red shirt won the match.

  • Formulating Hypotheses:

    • Null Hypothesis (H0H_0): Red- and blue-shirted athletes are equally likely to win, or red-shirted athletes are less likely to win (proportion of red winners ≤0.5\le 0.5).

    • Alternative Hypothesis (HAH_A): Red-shirted athletes are more likely to win (proportion of red winners >0.5> 0.5).

    • remember non directional hypothesis as opposed to directional hypotheses, which specify a direction of the expected effect. In a non-directional hypothesis, we would simply state that there is a difference in the probabilities of winning between red- and blue-shirted athletes without indicating which group is expected to be favored.

  • Sample Estimations:

    • Observed proportion of red-shirted winners: p^=1620=0.8\hat{p} = \frac{16}{20} = 0.8

    • Discrepancy: The observed sample proportion (0.80.8) deviates by 0.30.3 from the null hypothesis parameter (p=0.5p = 0.5).

Statistical Null Distribution and P-Value Calculation

  • Null Distribution Properties:

    • Under H0H_0, the expected number of red winners follows a binomial distribution with n=20n = 20 and p=0.5p = 0.5

    • Expected mean number of red winners: μ=n×p=20×0.5=10\mu = n \times p = 20 \times 0.5 = 10

  • P-Value Calculation:

    • The P-value is the probability of obtaining a sample result as extreme as or more extreme than the observed data, assuming H0H_0 is true.

    • For a one-tailed test with X=16X = 16 observed successes out of n=20n = 20:     P=P(16)+P(17)+P(18)+P(19)+P(20)P = P(16) + P(17) + P(18) + P(19) + P(20)

    • Applying the binomial formula P(X)=(20X)(0.5)X(0.5)20−XP(X) = \binom{20}{X} (0.5)^X (0.5)^{20 - X} for each outcome:     P=0.0059P = 0.0059

Binomial null distribution histogram for 20 trials with red line indicating 16 red winners
  • Statistical Decision Rule:

    • Standard significance threshold: α=0.05\alpha = 0.05

    • Result: P=0.0059P = 0.0059

    • Since P<0.05P < 0.05, the null hypothesis H0H_0 is rejected.

    • Conclusion: Red-shirted athletes were statistically significantly more likely to win their contests.