Comprehensive Study Notes on Discrete and Continuous Random Variables

Discrete Random Variables

  • Probability Mass Distribution (PMD): The probability mass distribution of a discrete random variable is the tabulation of the values assumed by the variable together with their corresponding probabilities.

  • Probability Mass Function (PMF):

    • In addition to tabular presentation, probabilities can be represented using a probability mass function.

    • In a probability mass function, the probability of the random variable is defined as a piecewise function.

    • The function must allow determination of the probability of any real value. Consequently, it must explicitly state that the probability is 00 for any value that the random variable cannot assume.

  • Fundamental Properties of Probability Mass Distributions and Functions:

    • For every possible value xx, the probability must satisfy:     0≤P(X=x)≤10 \le P(X = x) \le 1

    • The sum of all probabilities over the sample space must equal 1:     ∑P(X=x)=1\sum P(X = x) = 1

  • Cumulative Distribution Function (CDF):

    • The cumulative distribution function of a discrete random variable XX, denoted by F(x)F(x), represents the probability that the value of XX is at most xx:     F(x)=P(X≤x)F(x) = P(X \le x)

    • A CDF can be expressed in either tabular form or functional form.

    • In functional form, the cumulative distribution function:

    • Must allow the calculation of the cumulative probability for any value xx.

    • Starts at a cumulative probability of 00 (for values below the minimum possible value) and ends at a cumulative probability of 11 (for values at or above the maximum possible value).

  • Deriving Probability Mass Distribution from Cumulative Distribution Function:

    • Let XX be a discrete random variable assuming values in ascending order: x1,x2,…,xi−1,xi,…,xnx_1, x_2, \dots, x_{i-1}, x_i, \dots, x_n.

    • The individual probability for a specific outcome xix_i is derived from the CDF as:     P(X=xi)=F(xi)−F(xi−1)P(X = x_i) = F(x_i) - F(x_{i-1})

Expectation, Variance, and Standard Deviation of Discrete Random Variables

  • Expectation E(X)E(X):

    • The expectation (or expected value) of a discrete random variable XX, denoted by E(X)E(X) or μ\mu, represents the theoretical mean of XX:     E(X)=μ=∑xP(X=x)E(X) = \mu = \sum x P(X = x)

    • The theoretical mean can be calculated directly using classical probability without running empirical experiments.

    • In an empirical context, as the number of experimental trials increases, the sample mean obtained approaches the theoretical expectation E(X)E(X). Thus, expectation is defined as the long-run mean over a large number of trials.

  

average dice value against number of rolls
  • For example, when rolling a fair six-sided die, the theoretical expectation of the score XX is:     E(X)=(1×16)+(2×16)+(3×16)+(4×16)+(5×16)+(6×16)=3.5E(X) = (1 \times \frac{1}{6}) + (2 \times \frac{1}{6}) + (3 \times \frac{1}{6}) + (4 \times \frac{1}{6}) + (5 \times \frac{1}{6}) + (6 \times \frac{1}{6}) = 3.5

  • As demonstrated in experimental trials, as the number of die rolls approaches 10001000, the average value converges toward μ=3.5\mu = 3.5

    • Worked Example: Biased Spinner:

  • Consider a biased spinner with discrete outcome scores x∈{0,1,2,3}x \in \{0, 1, 2, 3\} and associated probability distribution:

  

biased spinner probability table
- P(X=0)=0.1 (10%)P(X = 0) = 0.1 \text{ (10\%)}
- P(X=1)=0.3 (30%)P(X = 1) = 0.3 \text{ (30\%)}
- P(X=2)=0.4 (40%)P(X = 2) = 0.4 \text{ (40\%)}
- P(X=3)=0.2 (20%)P(X = 3) = 0.2 \text{ (20\%)}
  • If the spinner is spun N=1600N = 1600 times, the expected frequency ff for each outcome is calculated as f=N×P(X=x)f = N \times P(X = x), recorded as:

  

expected frequency table
- For x=0x = 0: f=1600×0.1=160f = 1600 \times 0.1 = 160
- For x=1x = 1: f=1600×0.3=480f = 1600 \times 0.3 = 480
- For x=2x = 2: f=1600×0.4=640f = 1600 \times 0.4 = 640
- For x=3x = 3: f=1600×0.2=320f = 1600 \times 0.2 = 320
  • Derivation of the expected score E(X)E(X) for this spinner:     E(X)=∑xP(X=x)=(0×0.1)+(1×0.3)+(2×0.4)+(3×0.2)=0+0.3+0.8+0.6=1.7E(X) = \sum x P(X = x) = (0 \times 0.1) + (1 \times 0.3) + (2 \times 0.4) + (3 \times 0.2) = 0 + 0.3 + 0.8 + 0.6 = 1.7

  • Continual spinning yields a sample mean score approaching 1.71.7

    • Variance Var(X)\text{Var}(X) and Standard Deviation SD(X)\text{SD}(X):

  • Variance and standard deviation quantify the dispersion or spread of the values of a random variable around its mean.

  • Variance equation:     Var(X)=∑x2P(X=x)−[E(X)]2\text{Var}(X) = \sum x^2 P(X = x) - [E(X)]^2

  • Standard deviation equation:     SD(X)=Var(X)\text{SD}(X) = \sqrt{\text{Var}(X)}

  

square root of variance
  • Calculation for the biased spinner example:

    • First, calculate E(X2)E(X^2):       E(X2)=∑x2P(X=x)=(02×0.1)+(12×0.3)+(22×0.4)+(32×0.2)E(X^2) = \sum x^2 P(X = x) = (0^2 \times 0.1) + (1^2 \times 0.3) + (2^2 \times 0.4) + (3^2 \times 0.2)       E(X2)=0+0.3+1.6+1.8=3.7E(X^2) = 0 + 0.3 + 1.6 + 1.8 = 3.7

    • Compute Variance:       Var(X)=E(X2)−[E(X)]2=3.7−(1.7)2=3.7−2.89=0.81\text{Var}(X) = E(X^2) - [E(X)]^2 = 3.7 - (1.7)^2 = 3.7 - 2.89 = 0.81

    • Compute Standard Deviation:       SD(X)=0.81=0.9\text{SD}(X) = \sqrt{0.81} = 0.9

Integration and Area Under Curves

  • Definition of Area Under a Curve:

    • The area under a curve refers to the bounded area between the curve y=f(x)y = f(x) and the x-axis from x=ax = a to x=bx = b.

    • The region can lie entirely above the x-axis, entirely below the x-axis, or extend both above and below.

  

area under curve diagram
  • Integration Definition for Bounded Area:

    • The net signed area AA is given by the definite integral:     A=∫abf(x) dxA = \int_a^b f(x)\,dx

    • Area regions above the x-axis evaluate to positive values, while regions below the x-axis evaluate to negative values.

  • Fundamental Integration Formulas:

    • Power Rule (for n≠−1n \ne -1):     ∫xn dx=xn+1n+1+C\int x^n\,dx = \frac{x^{n+1}}{n+1} + C

    • Integration of constant 1:     ∫1 dx=x+C\int 1\,dx = x + C

    • Integration of arbitrary constant kk:     ∫k dx=kx+C\int k\,dx = kx + C

  • Properties of Definite Integrals:

    • Integral over a single point interval:     ∫aaf(x) dx=0\int_a^a f(x)\,dx = 0

    • Constant factor rule:     ∫abcf(x) dx=c∫abf(x) dx\int_a^b c f(x)\,dx = c \int_a^b f(x)\,dx

    • Sum rule:     ∫ab[f(x)+g(x)] dx=∫abf(x) dx+∫abg(x) dx\int_a^b [f(x) + g(x)]\,dx = \int_a^b f(x)\,dx + \int_a^b g(x)\,dx

    • Difference rule:     ∫ab[f(x)−g(x)] dx=∫abf(x) dx−∫abg(x) dx\int_a^b [f(x) - g(x)]\,dx = \int_a^b f(x)\,dx - \int_a^b g(x)\,dx

    • Fundamental Theorem of Calculus (Definite evaluation):     ∫abf(x) dx=[F(x)]ab=F(b)−F(a)\int_a^b f(x)\,dx = [F(x)]_a^b = F(b) - F(a)     where F(x)F(x) is an antiderivative of f(x)f(x).

Fundamentals of Continuous Random Variables and Histograms

  • Proportionality in Graphical Representations:

    • Frequency (or relative frequency) is directly proportional to the area of its corresponding bar in a histogram:     Frequency f∝Area A\text{Frequency } f \propto \text{Area } A     f=kAf = k A     where kk is a constant of proportionality measured in units per unit square.

    • Since the area of a rectangular bar with height hh and width (class width) ww is A=hwA = h w, we have:     f=khwf = k h w     h=fkwh = \frac{f}{k w}

    • If class width ww is constant across all intervals, setting k=1wk = \frac{1}{w} results in bar height equaling frequency (h=fh = f).

  • Frequency Density vs. Probability Density Height Formulations:

  

comparison of density heights
  • Using Frequency Density as Bar Height:

    • Select k=1 per unit squarek = 1\text{ per unit square}.

    • Formula for height hh:       h=fwh = \frac{f}{w}

    • Area of a bar A=hw=(fw)w=fA = h w = \left(\frac{f}{w}\right) w = f

    • Multiplying height hh by width ww yields the class frequency. Thus, hh is called frequency density.

  • Using Probability Density as Bar Height:

    • Select k=∑f per unit squarek = \sum f\text{ per unit square} (where ∑f\sum f is total sample frequency).

    • Formula for height hh:       h=fw∑fh = \frac{f}{w \sum f}

    • Area of a bar A=hw=(fw∑f)w=f∑f=ProbabilityA = h w = \left(\frac{f}{w \sum f}\right) w = \frac{f}{\sum f} = \text{Probability}

    • Multiplying height hh by width ww yields the class probability. Thus, hh is called probability density.

  

derivation of bar area and densities

Probability Density Functions (PDF)

  • Transition from Histogram to PDF:

    • When a histogram for a continuous random variable XX is constructed using probability density as the height of each bar, drawing a smooth continuous curve connecting the top midpoints of all bars defines the Probability Density Function, denoted as f(x)f(x).

  

probability density function curve over histogram
  • The area of an individual histogram bar represents the discrete class probability.

  • The area under the curve f(x)f(x) between two bounds approximates the sum of the bar areas within that range.

  • Total area of all bars equals total probability, which is 11:     Total Area=∑Probability=1\text{Total Area} = \sum \text{Probability} = 1

  • For a function restricted to interval [a,b][a, b]:     ∫abf(x) dx=1\int_a^b f(x)\,dx = 1

  • Generalized across the entire real number line (−∞,∞)(-\infty, \infty):     ∫−∞∞f(x) dx=1\int_{-\infty}^{\infty} f(x)\,dx = 1

    • Core Properties of Probability Density Functions:

  • Property 1 (Total Probability Rule):     ∫−∞∞f(x) dx=1\int_{-\infty}^{\infty} f(x)\,dx = 1

  • Property 2 (Non-negativity):     f(x)≥0∀ xf(x) \ge 0 \quad \forall \, x     The function f(x)f(x) must lie on or above the x-axis everywhere because it represents density/height.

  • Property 3 (Interval Probability):     The probability that XX lies between k1k_1 and k2k_2 is given by the integral:     P(k1≤X≤k2)=∫k1k2f(x) dxP(k_1 \le X \le k_2) = \int_{k_1}^{k_2} f(x)\,dx

  • Property 4 (Point Probability for Continuous Variables):     The probability of a continuous random variable assuming any exact single value k1k_1 is strictly 00:     P(X=k1)=∫k1k1f(x) dx=0P(X = k_1) = \int_{k_1}^{k_1} f(x)\,dx = 0

  • Property 5 (Invariance to Boundary Inclusion):     Because individual point probabilities equal zero, the inclusion or exclusion of endpoints does not alter interval probabilities:     P(k1≤X≤k2)=P(k1<X≤k2)=P(k1≤X<k2)=P(k1<X<k2)P(k_1 \le X \le k_2) = P(k_1 < X \le k_2) = P(k_1 \le X < k_2) = P(k_1 < X < k_2)

  • Validity Criterion: Any valid continuous probability density function f(x)f(x) must satisfy both Property 1 (∫−∞∞f(x) dx=1\int_{-\infty}^{\infty} f(x)\,dx = 1) and Property 2 (f(x)≥0f(x) \ge 0).

Continuous Cumulative Distribution Functions and Expectation

  • Continuous Cumulative Distribution Function (CDF):

    • Defined identically to the discrete case: F(x)F(x) is the cumulative probability that XX is at most xx:     F(x)=P(X≤x)=∫−∞xf(t) dtF(x) = P(X \le x) = \int_{-\infty}^x f(t)\,dt

    • Continuity across Piecewise Domains: The value of F(x)F(x) evaluated at boundary endpoints is identical whether calculated from the left or right piecewise sub-function, maintaining continuity across the domain boundary.

    • Finding Probabilities via CDF:     P(k1≤X≤k2)=F(k2)−F(k1)P(k_1 \le X \le k_2) = F(k_2) - F(k_1)

  • Expectation of a Continuous Random Variable:

    • The theoretical mean E(X)E(X) for a continuous random variable with density f(x)f(x) is computed as:     E(X)=∫−∞∞xf(x) dxE(X) = \int_{-\infty}^{\infty} x f(x)\,dx

  • Variance and Standard Deviation of a Continuous Random Variable:

    • The variance Var(X)\text{Var}(X) is computed as:     Var(X)=∫−∞∞x2f(x) dx−[E(X)]2\text{Var}(X) = \int_{-\infty}^{\infty} x^2 f(x)\,dx - [E(X)]^2

    • The standard deviation SD(X)\text{SD}(X) is the principal square root of variance:     SD(X)=Var(X)\text{SD}(X) = \sqrt{\text{Var}(X)}

  

standard deviation equation formula