Probability Concepts and Simulation

Population and Sample Proportions

  • Population vs. Sample: A population consists of NN individuals, while a sample consists of nn individuals selected from the population.

Population and Sample diagram
  • Proportion Representation: Percentages are conventionally expressed in decimal form to directly reflect proportions.

Perspectives on Probability

  • Theoretical Perspective:

P(A)=number of outcomes in event Atotal number of outcomes in sample spaceP(A) = \frac{\text{number of outcomes in event } A}{\text{total number of outcomes in sample space}}

  • Frequentist Perspective:

    • Probability is the proportion of times an outcome occurs in a very long series of repetitions, represented as a value between 00 and 11.

Proportion of heads in 10 vs 500 coin tosses
  • Short-Run Misconceptions:

    • Untrained intuition incorrectly expects randomness to be predictable in the short run.

    • The "law of averages" (or "Gambler's Ruin") refers to the mistaken belief that chance outcomes must "even out" in the short run.

Simulation and Statistical Evidence

  • Simulation Definition: Using a model matching outcome counts and probabilities to imitate chance behavior and evaluate real-world results.

  • Trial Process:

    • Assign numerical labels to outcomes (e.g., 1=Joey Logano1 = \text{Joey Logano}, 2=Kevin Harvick2 = \text{Kevin Harvick}, 3=Chase Elliott3 = \text{Chase Elliott}, 4=Danica Patrick4 = \text{Danica Patrick}, 5=Jimmie Johnson5 = \text{Jimmie Johnson}).

    • Generate random integers from 11 to 55 until all outcomes appear, then record the required count.

  • Statistical Evidence Threshold:

    • An observation occurring within the 5%5\% (0.050.05) least likely outcomes of a simulation provides convincing statistical evidence.

    • A probability of 0.04990.0499 is considered convincing evidence, whereas 0.05160.0516 is not.

  • NASCAR Cereal Example:

    • Taking 2323 boxes to collect all 55 driver cards yielded a simulated probability of ≈0/50=0\approx 0/50 = 0 across 5050 trials.

    • Because this outcome is well below the 0.050.05 threshold, it provides convincing evidence that the 55 drivers' cards are not equally likely.