Lecture 7: Binomial Probability Distribution and Hypothesis Testing Notes
Categorical Data and Binomial Probabilities
Categorical data consists of variables categorized into discrete, non-numeric groups. In binary categorical data, observations fall into exactly two mutually exclusive categories:
Live or dead
Success or failure
Affected by disease or not affected
Binomial probability tools are utilized to analyze binary categorical data and test hypotheses regarding underlying population probabilities.
Mendelian Genetics Example and Sample Space
Single-gene inheritance traits governed by dominant () and recessive () alleles serve as a foundational application of binomial probability.
Assuming random mating and equal allele frequencies ( and ):
Alleles inherited from the mother: or
Alleles inherited from the father: or
Resulting offspring genotypes: , , , and
Probability of recessive genotype ():
Probability of dominant phenotype (, , ):

When sampling 3 individuals randomly from this population, the complete sample space of outcomes for recessive () and dominant () phenotypes comprises possibilities:
if the probability of one phenotype doesn’t directly relate to the probability of another, then they are considered independent
faster way to do this is through binomial distribution
:
:
:
:
:
:
:
:

Mathematical Formulation of the Binomial Distribution
The binomial distribution describes the probability of obtaining a given number of "successes" () from a fixed number of independent trials ().
Individual trials in a binomial process are known as Bernoulli trials, defined as random variables with exactly two possible outcomes.
Parameters of the binomial distribution:
: total number of independent trials
: probability of success on any single trial
: probability of failure on any single trial
: observed number of successful trials (X = 0, 1, 2, \ndots, n)
General formula for binomial probability:
Combinatorial notation (also written as or ) represents the number of ways to choose successes from trials:
Factorial definition and rules:
Example:
Worked Calculation: Sampling Recessive Phenotypes
To calculate the probability that exactly out of randomly sampled individuals express a recessive phenotype ():
Number of combinations yielding 2 recessives:
Probability of 2 recessive phenotypes:
Probability of 1 dominant phenotype:
Step-by-step substitution and evaluation:
Hypothesis Testing Example: Red Uniforms in Combat Sports
Biological Context: Male animals frequently display bright red colorations as visual signals of aggression, competitive dominance, and threat intensity.


Empirical Investigation: A study evaluated whether uniform color influences outcome success in human combat sports (Hill, RA, and RA Barton 2005. "Red enhances human performance in contests." Nature 435:293).
Dataset: Examined round outcomes in wrestling, taekwondo, and boxing during the 2004 Olympic Games.
Experimental Design: Competitors were randomly assigned red or blue shirt colors.
Experimental Findings: In of combat rounds, the athlete assigned the red shirt won the match.
Formulating Hypotheses:
Null Hypothesis (): Red- and blue-shirted athletes are equally likely to win, or red-shirted athletes are less likely to win (proportion of red winners ).
Alternative Hypothesis (): Red-shirted athletes are more likely to win (proportion of red winners ).
remember non directional hypothesis as opposed to directional hypotheses, which specify a direction of the expected effect. In a non-directional hypothesis, we would simply state that there is a difference in the probabilities of winning between red- and blue-shirted athletes without indicating which group is expected to be favored.
Sample Estimations:
Observed proportion of red-shirted winners:
Discrepancy: The observed sample proportion () deviates by from the null hypothesis parameter ().
Statistical Null Distribution and P-Value Calculation
Null Distribution Properties:
Under , the expected number of red winners follows a binomial distribution with and
Expected mean number of red winners:
P-Value Calculation:
The P-value is the probability of obtaining a sample result as extreme as or more extreme than the observed data, assuming is true.
For a one-tailed test with observed successes out of :
Applying the binomial formula for each outcome:

Statistical Decision Rule:
Standard significance threshold:
Result:
Since , the null hypothesis is rejected.
Conclusion: Red-shirted athletes were statistically significantly more likely to win their contests.