Conditional Probability, Statistical Independence, and Simple Linear Regression

Fundamentals of Conditional Probability and Joint Distributions

  • Variable Distinction in Conditional Distributions

    • Unlike joint probability distributions, the order of variables in conditional probability distributions matters significantly.

    • In the expression f(y∣x)f(y|x), yy represents the random variable whose probability is being calculated, while xx is the conditioning variable.

  • Joint Probability versus Conditional Probability

    • Joint Probability: Measures the likelihood that an individual randomly selected from the entire population concurrently satisfies two criteria (e.g., holding a college degree and earning a low wage).

    • Conditional Probability: Measures the likelihood of an outcome within a specific subpopulation defined by the conditioning variable (e.g., among only those individuals who possess a college degree, the proportion who earn a low wage).

  • Tabular Example: Wage (yy) and Education (xx)

    • Variables and Outcomes:

    • yy (Wage): Low Wage ($10\$10), High Wage ($20\$20)

    • xx (Education): No College, College

    • Joint Probabilities and Marginals:

    • Joint probability of No College and Low Wage: f(Low Wage,No College)=0.3f(\text{Low Wage}, \text{No College}) = 0.3

    • Joint probability of No College and High Wage: f(High Wage,No College)=0.0f(\text{High Wage}, \text{No College}) = 0.0

    • Marginal probability of No College: m(No College)=0.3m(\text{No College}) = 0.3

    • Joint probability of College and Low Wage: f(Low Wage,College)=0.2f(\text{Low Wage}, \text{College}) = 0.2

    • Joint probability of College and High Wage: f(High Wage,College)=0.5f(\text{High Wage}, \text{College}) = 0.5

    • Marginal probability of College: m(College)=0.7m(\text{College}) = 0.7

  • Derivation of Conditional Probabilities f(y∣x)=f(x,y)m(x)f(y|x) = \frac{f(x, y)}{m(x)}

    • Subpopulation: No College (x=No Collegex = \text{No College}):

    • Conditional probability of Low Wage given No College:       f(Low Wage∣No College)=0.30.3=1.00f(\text{Low Wage} \mid \text{No College}) = \frac{0.3}{0.3} = 1.00

    • Conditional probability of High Wage given No College:       f(High Wage∣No College)=0.00.3=0.00f(\text{High Wage} \mid \text{No College}) = \frac{0.0}{0.3} = 0.00

    • Sum of conditional probabilities across the row: 1.00+0.00=1.001.00 + 0.00 = 1.00 (100%100\%).

    • Subpopulation: College (x=Collegex = \text{College}):

    • Conditional probability of Low Wage given College:       f(Low Wage∣College)=0.20.7≈0.29f(\text{Low Wage} \mid \text{College}) = \frac{0.2}{0.7} \approx 0.29

    • Conditional probability of High Wage given College:       f(High Wage∣College)=0.50.7≈0.71f(\text{High Wage} \mid \text{College}) = \frac{0.5}{0.7} \approx 0.71

    • Sum of conditional probabilities across the row: 0.29+0.71=1.000.29 + 0.71 = 1.00 (100%100\%).

  • Interpretation of Findings

    • No College Subpopulation: 100%100\% of individuals without a college degree earn a low wage, and 0%0\% earn a high wage.

    • College Subpopulation: 29%29\% of individuals with a college degree earn a low wage, whereas 71%71\% earn a high wage.

    • Comparison with Joint Distribution: Joint probability states that 30%30\% of the total population has no college and earns a low wage. Conditional probability demonstrates that within the group having no college education, 100%100\% earn a low wage.

Statistical Independence

  • Definition of Statistical Independence

    • Two random variables XX and YY are statistically independent if and only if their joint probability distribution equals the product of their respective marginal probability distributions:     f(x,y)=m(x)×n(y)f(x, y) = m(x) \times n(y)

    • Equivalent Conditional Form: Dividing both sides by the marginal density m(x)m(x) yields:     f(y∣x)=n(y)f(y|x) = n(y)

    • Statistical independence implies that conditioning on XX provides zero additional information about the probability distribution of YY.

  • Testing for Statistical Independence in the Wage-Education Model

    • Marginal distribution of Low Wage: n(Low Wage)=0.5n(\text{Low Wage}) = 0.5

    • Conditional distributions derived:

    • f(Low Wage∣No College)=1.00f(\text{Low Wage} \mid \text{No College}) = 1.00

    • f(Low Wage∣College)=0.29f(\text{Low Wage} \mid \text{College}) = 0.29

    • Evaluation: 1.00≠0.51.00 \neq 0.5 and 0.29≠0.50.29 \neq 0.5. Because f(y∣x)≠n(y)f(y|x) \neq n(y), wage and education are not statistically independent.

    • Rule of Distribution Comparison: If even a single conditional probability value fails to match its corresponding marginal probability value across the entire distribution, the variables are not statistically independent.

    • Theoretical scenario for independence: If education had no relationship with wage, the conditional distribution would equal the marginal distribution across all education tiers (e.g., 50%50\% low wage and 50%50\% high wage regardless of college degree).

  • Population Context

    • All values, probabilities, and distributions evaluated here refer explicitly to the full population parameters, not sample statistics.

Expected Value and Expectation Operators

  • Population Mean / Expected Value Notation

    • The notation E(Y)E(Y) represents the expected value, population mean, or average of a random variable YY.

    • The symbol EE serves as the expectation operator.

  • Formula for Expected Value of Discrete Population Variable

    • Sum of each possible value of YY weighted by its marginal probability n(y)n(y):     E(Y)=∑y×n(y)E(Y) = \sum y \times n(y)

  • Population Mean vs. Sample Mean

    • Sample Mean: Sum of observed values divided by sample size (NN); acts as an estimator for the population mean.

    • Population Mean: The true parameter derived from the actual population probability distribution weights.

  • Calculation of Overall Expected Wage E(Wage)E(\text{Wage})

    • Discrete wage values: $10\$10 (Low) and $20\$20 (High).

    • Marginal probabilities: n($10)=0.5n(\$10) = 0.5, n($20)=0.5n(\$20) = 0.5

    • Computation:     E(Wage)=10×0.5+20×0.5=5+10=$15E(\text{Wage}) = 10 \times 0.5 + 20 \times 0.5 = 5 + 10 = \$15

Conditional Expectation and Mean Independence

  • Conditional Expectation / Conditional Mean Definition

    • Notation: E(Y∣X)E(Y|X)

    • Formula: Takes each outcome yy and multiplies it by the conditional probability f(y∣x)f(y|x) rather than the marginal probability:     E(Y∣X=x)=∑y×f(y∣x)E(Y \mid X = x) = \sum y \times f(y \mid x)

  • Calculation of Conditional Expected Wage E(Wage∣Education)E(\text{Wage} \mid \text{Education})

    • Group 1: No College (x=No Collegex = \text{No College}):     E(Wage∣No College)=10×1.00+20×0.00=$10E(\text{Wage} \mid \text{No College}) = 10 \times 1.00 + 20 \times 0.00 = \$10

    • Group 2: College (x=Collegex = \text{College}):     E(Wage∣College)=10×0.29+20×0.71=2.90+14.20=$17.10E(\text{Wage} \mid \text{College}) = 10 \times 0.29 + 20 \times 0.71 = 2.90 + 14.20 = \$17.10

    • Interpretation: The overall population average wage is $15\$15. However, conditional on having no college, the average wage drops to $10\$10. Conditional on holding a college degree, the average wage increases to $17.10\$17.10

  • Definition of Mean Independence

    • Variables XX and YY are mean independent if the conditional expectation of YY given XX is identically equal to the unconditional expected value of YY:     E(Y∣X)=E(Y)E(Y \mid X) = E(Y)

    • Evaluation in example: E(\text{Wage} \mid \text{No College}) = \10 \neq \1515 and E(\text{Wage} \mid \text{College}) = \17.10 \neq \1515.

    • Conclusion: Wage and education are not mean independent.

Hierarchical Relationships Among Types of Independence

  • Logical Chain of Implication   Statistical Independence  ⟹  Mean Independence  ⟹  Uncorrelated (Cov(X,Y)=0)\text{Statistical Independence} \implies \text{Mean Independence} \implies \text{Uncorrelated } (\text{Cov}(X,Y) = 0)

  • Directionality and Non-Reversibility

    • Statistical Independence   ⟹  \implies Mean Independence: If the entire probability distributions are independent, their conditional means must be identical. However, Mean Independence does not imply Statistical Independence (two subpopulation distributions can share the exact same mean while possessing completely different variances or overall shapes).

    • Mean Independence   ⟹  \implies Uncorrelated: If the conditional expectation of YY given XX is constant, the linear correlation between XX and YY is zero. However, zero correlation does not imply Mean Independence.

  • Functional Scope

    • Correlation: Measures strictly linear relationships between variables.

    • Mean Independence: Captures both linear and non-linear relationships (e.g., quadratic, sinusoidal/sine/cosine relationships where linear correlation might yield 00, but conditional means vary non-linearly with XX).

    • Statistical Independence: The strongest condition, capturing all aspects of the full probability distribution (means, variances, higher moments).

Population Variance, Standard Deviation, and Correlation

  • Population Variance Formula and Properties

    • Definition formula:     Var(Y)=E[(Y−E(Y))2]\text{Var}(Y) = E[(Y - E(Y))^2]

    • Computational expansion:     Var(Y)=E(Y2)−(E(Y))2\text{Var}(Y) = E(Y^2) - (E(Y))^2

    • Measures population spread/dispersion around the central tendency.

    • Non-negativity Property: Because variance incorporates a squared term [Y−E(Y)]2[Y - E(Y)]^2, variance can never be negative (Var(Y)≥0\text{Var}(Y) \ge 0).

  • Standard Deviation

    • Definition: Standard deviation is the square root of the variance:     SD(Y)=Var(Y)\text{SD}(Y) = \sqrt{\text{Var}(Y)}

    • Advantage: Expressed in the exact same measurement units as the original variable (e.g., dollars \) rather than squared units (e.g., $2\$^2).

  • Conditional Variance

    • Calculated by replacing the marginal probability function with the conditional probability function inside the variance calculation.

    • Measures dispersion within a specified subpopulation.

  • Covariance and Correlation Formulas

    • Covariance:     Cov(X,Y)=E[(X−E(X))(Y−E(Y)]\text{Cov}(X, Y) = E[(X - E(X))(Y - E(Y)]

    • Correlation:     Corr(X,Y)=Cov(X,Y)SD(X)×SD(Y)\text{Corr}(X, Y) = \frac{\text{Cov}(X, Y)}{\text{SD}(X) \times \text{SD}(Y)}

Introduction to Simple Linear Regression (SLR)

  • Model Structure and Directionality

    • Simple Linear Regression analyzes directional effects of an independent explanatory variable XX on a dependent outcome variable YY (X→YX \rightarrow Y).

    • General Econometric Equation:     Y=β0+β1X+UY = \beta_0 + \beta_1 X + U

  • Model Components

    • YY: Dependent / outcome variable (e.g., Wage).

    • XX: Independent / explanatory variable (e.g., Education).

    • UU: Error term representing unobserved factors affecting YY other than XX (e.g., ability, occupation, industry).

    • β0\beta_0: Intercept parameter (a fixed, unknown constant).

    • β1\beta_1: Slope parameter (a fixed, unknown constant).

  • Core Econometric Assumptions

    • Unconditional mean of the error term is zero:     E(U)=0E(U) = 0

    • Conditional mean of the error term given XX is zero:     E(U∣X)=0E(U \mid X) = 0

    • Theoretical Connection: The assumption E(U∣X)=E(U)=0E(U \mid X) = E(U) = 0 establishes that the unobserved error term UU and explanatory variable XX are mean independent.

Questions & Discussion

  • Question: When checking if conditional expectations equal unconditional expectations for mean independence (e.g., E(Wage∣Education)E(\text{Wage} \mid \text{Education})), do you evaluate both subgroups ($10\$10 for No College and $17.10\$17.10 for College) or just one?

    • Response: Both values must be checked. For variables to be mean independent, the conditional expectation across every single subpopulation category must equal the unconditional expected value. If any single subpopulation mean differs, mean independence fails.

  • Question: Is the variance equation squared while standard deviation undoes the square?

    • Response: Yes, variance calculates average squared deviations from the mean. Standard deviation takes the square root of that result to restore units back to the original measurement scale (e.g., returning from squared dollars to dollars).

  • Question: Will students be required to memorize complex formulas for variance and correlation for exams?

    • Response: No manual calculations or formula memorization for variance and correlation will be required on exams. Software (R) will calculate variance and correlation. Exam expectations are limited to basic concepts, marginal/conditional probabilities, and simple expected value calculations.

  • Question: Is homework due as originally scheduled on the syllabus?

    • Response: No, the first homework assignment is pushed back by one day to ensure adequate coverage of linear regression material prior to assignment. A practice problem covering probability tables and expected values will be posted online.