Distributions of Random Variables
Definition and Properties of Random Variables
A random variable maps a random process into a numerical outcome.
Notation standards designate capital letters such as or to represent the random variable itself, while lower case letters such as or denote specific realized outcomes.
The concept of randomness is the core defining element of these variables.
The law of large numbers serves as the bridging principle that allows concepts calculated from sample data to be transferred to the population level.
In a population study regarding enrolled students and book purchases:
Data reveals buy no books, buy the textbook only, and buy both the textbook and its study guide.
At the population level, the bookstore expects around students to buy books, about students to buy book, and approximately students to buy books.
The total expected book sales for such a class is calculated as: books.
It is critical to recognize sampling variability; selling slightly more or less than the exact expected value (e.g., or books) is natural and statistically normal, comparable to how flipping a coin times rarely results in exactly heads but usually stays close to that figure.
Revenue Variables and Numerical Coding
Analysis of bookstore revenue () for a class of students:
The textbook costs and the study guide costs .
Expected revenue from textbook-only buyers: \text{\137} \times 55 = \text{\7535}.
Expected revenue from dual-book buyers: (\text{\137} + \text{\33}) \times 25 = \text{\170} \times 25 = \text{\4250}.
Total expected class revenue: \text{\7535} + \text{\4250} = \text{\$11785}.
The random variable represents revenue from a single student with outcomes , , and occurring with probabilities , , and respectively.
Caveat regarding Numerical Coding:
Non-numeric results, such as academic disciplines, are often coded numerically: Sociology (), Psychology (), and Pedagogy ().
Coded numbers must not be treated as regular numbers; operations such as calculating the average of codes are invalid because the numerical value does not represent a magnitude (e.g., is not "three times as much" as ).
Variable Classification:
Discrete Random Variables: The set of possible outcomes is a discrete set (finite or countably infinite).
Continuous Random Variables: The set of outcomes consists of intervals. While transaction outcomes may appear discrete, revenue as a concept is continuous in nature.
Discrete Probability Distributions
A probability distribution for a discrete variable lists all possible disjoint outcomes and their associated probabilities.
It serves as the population-level equivalent of a sample frequency table.
Three Mandatory Rules for Probability Distributions:
The outcomes listed must be disjoint.
Each probability must be between and .
The total sum of all probabilities must be exactly .
Example: Sum of Two Dice:
Disjoint outcomes range from to .
Probabilities follow the rule of Laplace: , , , , , , , , , , .
As observations increase, sample proportions converge to these theoretical probabilities due to the law of large numbers.
Visual summaries are provided by bar plots where heights represent the probabilities of specific outcomes.
Cumulative Distribution and Quantile Functions
Cumulative Distribution Function ():
Defined for a random variable as for all real numbers .
It is the population counterpart to the empirical cumulative distribution function ().
Quantile Function ():
The inverse of the cumulative distribution function.
Defined as , where is the smallest outcome such that .
This function is used to define medians, quartiles, and other specific distribution markers.
Calculations for the sum of two dice ():
To find , calculate .
To find the median , identify the smallest where . Given () and (), the result is .
Expected Value and Variability
Expected Value ( or ):
This is the weighted mean of all possible outcomes.
Formula: .
In physics, the expectation represents the center of gravity; the distribution balances on a triangle placed at .
For the single student revenue model: .
Population Variance ( or ):
Describes the volatility or variability of a random variable.
Defined as the weighted mean of squared deviations from the population mean.
Formula: .
General formula: .
Standard Deviation ():
The square root of the variance ().
For the student revenue model: .
Standard Deviation: .
The Bernoulli Distribution
A Bernoulli random variable studies a single trial with only two possible outcomes: "success" and "failure".
Success is numerically coded as and failure as .
The success probability is denoted by , and the failure probability by .
Descriptive stats for Bernoulli variables:
Mean: .
Standard Deviation: .
Term Utility: The label "success" does not carry a positive connotation but is a neutral mathematical label (e.g., success could be labeled as refusing to give a shock).
The Binomial Distribution
The Binomial distribution counts the number of successes () in a fixed number () of independent and identically distributed (iid) Bernoulli trials.
Mathematical Framework:
The probability of exactly successes in trials with success probability is: .
Combination formula: .
Descriptive Figures for Binomial variables:
Mean: .
Variance: .
Standard Deviation: .
Four Required Conditions for the Binomial Model:
The trials are independent.
The number of trials is fixed.
Each outcome must be classified as a success or failure.
The probability of success remains constant across every trial.
Computational Implementation and Tips
Binomial probability calculations often involve factorial cancellation for efficiency.
Excel Functionality:
Individual probability :
\text{binom.dist}(k; n; p; \text{false}).Cumulative probability :
\text{binom.dist}(k; n; p; \text{true}).
Rule of Thumb: Approximately of observations are expected to fall within standard deviations of the mean.
Questions & Discussion
US Household Income Distributions Check
Three distributions were proposed for US household income ranges (, , , in thousands).
(a) : Sum is (Incorrect, must be ).
(b) : Contains a negative probability (Invalid).
(c) : Correct distribution (Sum is , all values valid).
The Milgram Experiment Analysis
The study found that about of individuals administered the worst shock, meaning the probability of "success" (refusal) is .
For a sample of people, the probability exactly one refuses is: .
For a randomized sample of students, the expected number of refusers is () with a standard deviation of ().
Smoker and Lung Condition Scenarios
A random smoker has a probability of a severe lung condition.
Check for friends: Binomial model applies if friends are independent and do not share habits.
Probability none develop condition: .
Probability exactly one develops condition: .
Probability at most one develops condition: .
Probability at least two develop condition: .
Chemistry Book Revenue Practice
Costs: textbook , supplement . History: buy textbook, buy both.
(a) Percent buying nothing: .
(b) Expected Revenue () per student: .
(c) Variance and Volatility: ; .
Combinatorial Logic
Question: Why is and ?
Answer: There is only one way to have zero successes (all failures) and only one way to have all successes in trials.
Question: Ways to arrange success or successes?
Answer: Both are equal to , as there are exactly unique slots to place a single disparate outcome (one success among failures or one failure among successes).