Distributions of Random Variables

Definition and Properties of Random Variables

  • A random variable maps a random process into a numerical outcome.

  • Notation standards designate capital letters such as XX or YY to represent the random variable itself, while lower case letters such as xx or yy denote specific realized outcomes.

  • The concept of randomness is the core defining element of these variables.

  • The law of large numbers serves as the bridging principle that allows concepts calculated from sample data to be transferred to the population level.

  • In a population study regarding 100100 enrolled students and book purchases:

    • Data reveals 20%20\% buy no books, 55%55\% buy the textbook only, and 25%25\% buy both the textbook and its study guide.

    • At the population level, the bookstore expects around 2020 students to buy 00 books, about 5555 students to buy 11 book, and approximately 2525 students to buy 22 books.

    • The total expected book sales for such a class is calculated as: 0×20+1×55+2×25=1050 \times 20 + 1 \times 55 + 2 \times 25 = 105 books.

  • It is critical to recognize sampling variability; selling slightly more or less than the exact expected value (e.g., 104104 or 106106 books) is natural and statistically normal, comparable to how flipping a coin 100100 times rarely results in exactly 5050 heads but usually stays close to that figure.

Revenue Variables and Numerical Coding

  • Analysis of bookstore revenue (YY) for a class of 100100 students:

    • The textbook costs $137\text{\$137} and the study guide costs $33\text{\$33}.

    • Expected revenue from textbook-only buyers: \text{\137} \times 55 = \text{\7535}.

    • Expected revenue from dual-book buyers: (\text{\137} + \text{\33}) \times 25 = \text{\170} \times 25 = \text{\4250}.

    • Total expected class revenue: \text{\7535} + \text{\4250} = \text{\$11785}.

  • The random variable XX represents revenue from a single student with outcomes m1=$0m_1 = \text{\$0}, m2=$137m_2 = \text{\$137}, and m3=$170m_3 = \text{\$170} occurring with probabilities 0.200.20, 0.550.55, and 0.250.25 respectively.

  • Caveat regarding Numerical Coding:

    • Non-numeric results, such as academic disciplines, are often coded numerically: Sociology (11), Psychology (22), and Pedagogy (33).

    • Coded numbers must not be treated as regular numbers; operations such as calculating the average of codes are invalid because the numerical value does not represent a magnitude (e.g., 33 is not "three times as much" as 11).

  • Variable Classification:

    • Discrete Random Variables: The set of possible outcomes is a discrete set (finite or countably infinite).

    • Continuous Random Variables: The set of outcomes consists of intervals. While transaction outcomes may appear discrete, revenue as a concept is continuous in nature.

Discrete Probability Distributions

  • A probability distribution for a discrete variable lists all possible disjoint outcomes and their associated probabilities.

  • It serves as the population-level equivalent of a sample frequency table.

  • Three Mandatory Rules for Probability Distributions:

    1. The outcomes listed must be disjoint.

    2. Each probability must be between 00 and 11.

    3. The total sum of all probabilities must be exactly 11.

  • Example: Sum of Two Dice:

    • Disjoint outcomes range from 22 to 1212.

    • Probabilities follow the rule of Laplace: P(2)=136P(2) = \frac{1}{36}, P(3)=236P(3) = \frac{2}{36}, P(4)=336P(4) = \frac{3}{36}, P(5)=436P(5) = \frac{4}{36}, P(6)=536P(6) = \frac{5}{36}, P(7)=636P(7) = \frac{6}{36}, P(8)=536P(8) = \frac{5}{36}, P(9)=436P(9) = \frac{4}{36}, P(10)=336P(10) = \frac{3}{36}, P(11)=236P(11) = \frac{2}{36}, P(12)=136P(12) = \frac{1}{36} .

  • As observations increase, sample proportions converge to these theoretical probabilities due to the law of large numbers.

  • Visual summaries are provided by bar plots where heights represent the probabilities of specific outcomes.

Cumulative Distribution and Quantile Functions

  • Cumulative Distribution Function (F(x)F(x)):

    • Defined for a random variable XX as F(x)=P(X≤x)F(x) = P(X \le x) for all real numbers xx.

    • It is the population counterpart to the empirical cumulative distribution function (F^(x)\hat{F}(x)).

  • Quantile Function (Q(p)Q(p)):

    • The inverse of the cumulative distribution function.

    • Defined as Q(p)=xQ(p) = x, where xx is the smallest outcome such that F(x)≥pF(x) \ge p.

    • This function is used to define medians, quartiles, and other specific distribution markers.

  • Calculations for the sum of two dice (XX):

    • To find F(3)F(3), calculate P(X≤3)=P(X=2)+P(X=3)=136+236=336P(X \le 3) = P(X = 2) + P(X = 3) = \frac{1}{36} + \frac{2}{36} = \frac{3}{36}.

    • To find the median Q(0.5)Q(0.5), identify the smallest xx where F(x)≥0.5F(x) \ge 0.5. Given P(X≤6)=1536P(X \le 6) = \frac{15}{36} (0.4170.417) and P(X≤7)=2136P(X \le 7) = \frac{21}{36} (0.5830.583), the result is Q(0.5)=7Q(0.5) = 7.

Expected Value and Variability

  • Expected Value (E(X)E(X) or μ\mu):

    • This is the weighted mean of all possible outcomes.

    • Formula: E(X)=∑i=1kmiP(X=mi)E(X) = \sum_{i=1}^{k} m_i P(X = m_i).

    • In physics, the expectation represents the center of gravity; the distribution balances on a triangle placed at μ\mu.

    • For the single student revenue model: E(X)=0(0.20)+137(0.55)+170(0.25)=$117.85E(X) = 0(0.20) + 137(0.55) + 170(0.25) = \text{\$117.85}.

  • Population Variance (Var(X)Var(X) or σ2\sigma^2):

    • Describes the volatility or variability of a random variable.

    • Defined as the weighted mean of squared deviations from the population mean.

    • Formula: σ2=∑j=1k(mj−μ)2P(X=mj)\sigma^2 = \sum_{j=1}^{k} (m_j - \mu)^2 P(X = m_j).

    • General formula: σ2=E[(X−E(X))2]=E(X2)−(E(X))2\sigma^2 = E[(X - E(X))^2] = E(X^2) - (E(X))^2.

  • Standard Deviation (σ\sigma):

    • The square root of the variance (σ=Var(X)\sigma = \sqrt{Var(X)}).

    • For the student revenue model: Var(X)=(0−117.85)2(0.20)+(137−117.85)2(0.55)+(170−117.85)2(0.25)=3659.3Var(X) = (0 - 117.85)^2(0.20) + (137 - 117.85)^2(0.55) + (170 - 117.85)^2(0.25) = 3659.3.

    • Standard Deviation: 3659.3=$60.49\sqrt{3659.3} = \text{\$60.49}.

The Bernoulli Distribution

  • A Bernoulli random variable studies a single trial with only two possible outcomes: "success" and "failure".

  • Success is numerically coded as 11 and failure as 00.

  • The success probability is denoted by pp, and the failure probability by q=1−pq = 1 - p.

  • Descriptive stats for Bernoulli variables:

    • Mean: μ=p\mu = p.

    • Standard Deviation: σ=p(1−p)\sigma = \sqrt{p(1 - p)}.

  • Term Utility: The label "success" does not carry a positive connotation but is a neutral mathematical label (e.g., success could be labeled as refusing to give a shock).

The Binomial Distribution

  • The Binomial distribution counts the number of successes (kk) in a fixed number (nn) of independent and identically distributed (iid) Bernoulli trials.

  • Mathematical Framework:

    • The probability of exactly kk successes in nn trials with success probability pp is: P(X=k)=(nk)pk(1−p)n−kP(X = k) = \binom{n}{k} p^k (1 - p)^{n-k}.

    • Combination formula: (nk)=n!k!(n−k)!\binom{n}{k} = \frac{n!}{k!(n - k)!}.

  • Descriptive Figures for Binomial variables:

    • Mean: μ=np\mu = np.

    • Variance: σ2=np(1−p)\sigma^2 = np(1 - p).

    • Standard Deviation: σ=np(1−p)\sigma = \sqrt{np(1 - p)}.

  • Four Required Conditions for the Binomial Model:

    1. The trials are independent.

    2. The number of trials nn is fixed.

    3. Each outcome must be classified as a success or failure.

    4. The probability of success pp remains constant across every trial.

Computational Implementation and Tips

  • Binomial probability calculations often involve factorial cancellation for efficiency.

  • Excel Functionality:

    • Individual probability P(X=k)P(X = k): \text{binom.dist}(k; n; p; \text{false}).

    • Cumulative probability P(X≤k)P(X \le k): \text{binom.dist}(k; n; p; \text{true}).

  • Rule of Thumb: Approximately 95%95\% of observations are expected to fall within 22 standard deviations of the mean.

Questions & Discussion

  • US Household Income Distributions Check

    • Three distributions were proposed for US household income ranges (0−250-25, 25−5025-50, 50−10050-100, 100+100+ in thousands).

    • (a) 0.18,0.39,0.33,0.160.18, 0.39, 0.33, 0.16: Sum is 1.061.06 (Incorrect, must be 1.001.00).

    • (b) 0.38,−0.27,0.52,0.370.38, -0.27, 0.52, 0.37: Contains a negative probability (Invalid).

    • (c) 0.28,0.27,0.29,0.160.28, 0.27, 0.29, 0.16: Correct distribution (Sum is 1.001.00, all values valid).

  • The Milgram Experiment Analysis

    • The study found that about 65%65\% of individuals administered the worst shock, meaning the probability of "success" (refusal) is p=0.35p = 0.35.

    • For a sample of 44 people, the probability exactly one refuses is: 4×(0.35)1×(0.65)3=0.384 \times (0.35)^1 \times (0.65)^3 = 0.38.

    • For a randomized sample of 4040 students, the expected number of refusers is 1414 (40×0.3540 \times 0.35) with a standard deviation of 3.023.02 (40×0.35×0.65\sqrt{40 \times 0.35 \times 0.65}).

  • Smoker and Lung Condition Scenarios

    • A random smoker has a 0.30.3 probability of a severe lung condition.

    • Check for 44 friends: Binomial model applies if friends are independent and do not share habits.

    • Probability none develop condition: (40)(0.3)0(0.7)4=0.2401\binom{4}{0} (0.3)^0 (0.7)^4 = 0.2401.

    • Probability exactly one develops condition: (41)(0.3)1(0.7)3=0.4116\binom{4}{1} (0.3)^1 (0.7)^3 = 0.4116.

    • Probability at most one develops condition: 0.2401+0.4116=0.65170.2401 + 0.4116 = 0.6517.

    • Probability at least two develop condition: 1−0.6517=0.34831 - 0.6517 = 0.3483.

  • Chemistry Book Revenue Practice

    • Costs: textbook $159\text{\$159}, supplement $41\text{\$41}. History: 25%25\% buy textbook, 60%60\% buy both.

    • (a) Percent buying nothing: 100%−25%−60%=15%100\% - 25\% - 60\% = 15\%.

    • (b) Expected Revenue (YY) per student: 0(0.15)+159(0.25)+200(0.60)=$159.750(0.15) + 159(0.25) + 200(0.60) = \text{\$159.75}.

    • (c) Variance and Volatility: Var(Y)=4800Var(Y) = 4800; σ=4800=$69.28\sigma = \sqrt{4800} = \text{\$69.28}.

  • Combinatorial Logic

    • Question: Why is (n0)=1\binom{n}{0} = 1 and (nn)=1\binom{n}{n} = 1?

    • Answer: There is only one way to have zero successes (all failures) and only one way to have all successes in nn trials.

    • Question: Ways to arrange 11 success or n−1n - 1 successes?

    • Answer: Both are equal to nn, as there are exactly nn unique slots to place a single disparate outcome (one success among failures or one failure among successes).