march 5

Chapter 7: OLS with Multiple Regressors (Hypothesis Tests)

Lecture Outline

  • Hypothesis test for single coefficient in multiple regression analysis

  • Confidence interval for single coefficient in multiple regression

  • Testing hypotheses on 2 or more coefficients

  • The F-statistic

  • The overall regression F-statistic

  • Testing single restrictions involving multiple coefficients

  • Measures of fit in multiple regression model

  • SER, R2R^2 and adjusted R2R^2

  • Relation between (homoskedasticity-only) F-statistic and the R2R^2

  • Interpreting measures of fit

  • Interpreting “stars” in a table with regression output

Hypothesis Test for Single Coefficient in Multiple Regression Analysis

  • Example conducted on February 2, 2017, at 14:18:25.

    • Example command in R:

    regress test_score class_size el_pct, robust
    
    • Output Summary:

    • Number of Observations: 420

    • F-statistic: F(2,417)=223.82F(2, 417) = 223.82

    • Probability > F: 0.0000

    • R-squared Value: 0.4264

    • Root Mean Square Error (RMSE): 14.464

  • Regression Coefficients:

    • Class Size: Coef.=−1.101296Coef. = -1.101296, Std.Err.=0.4328472Std. Err. = 0.4328472, t=−2.54t = -2.54, P>∣t∣=0.011P>|t| = 0.011, [95% Conf. Interval: (−1.95213,−0.2504616)(-1.95213, -0.2504616)]

    • Percentage of English Learners: Coef.=−0.6497768Coef. = -0.6497768, Std.Err.=0.0310318Std. Err. = 0.0310318, t=−20.94t = -20.94, P>∣t∣=0.000P>|t| = 0.000, [95% Conf. Interval: (−0.710775,−0.5887786)(-0.710775, -0.5887786)]

    • Constant: Coef.=686.0322Coef. = 686.0322, Std.Err.=8.728224Std. Err. = 8.728224, t=78.60t = 78.60, P>∣t∣=0.000P>|t| = 0.000, [95% Conf. Interval: (668.8754,703.189)(668.8754, 703.189)]

  • Significant Test Question:

    • Does changing class size, while holding the percentage of English learners constant, have a statistically significant effect on test scores?

    • Significance Level: 0.05

Hypothesis Tests Steps for Single Coefficient

  • Under the four Least Squares assumptions of the multiple regression model:

    1. E(ui∣X<em>1i,X</em>2i,…,Xki)=0E(ui|X<em>{1i}, X</em>{2i}, …, X_{ki}) = 0

    2. (Y<em>i,X</em>1i,X<em>2i,…,X</em>ki)(Y<em>i, X</em>{1i}, X<em>{2i}, …, X</em>{ki}) for i=1,…,ni = 1, …, n are independent and identically distributed (i.i.d.)

    3. Large outliers are unlikely.

    4. No perfect multicollinearity.

  • The Ordinary Least Squares (OLS) estimators bjb_j for j=1,..,kj = 1, .., k are approximately normally distributed in large samples.

    • In addition, t=b<em>j−j</em>0SE(bj)∼N(0,1)t = \frac{b<em>j - j</em>0}{SE(b_j)} \sim N(0, 1).

Steps to Perform Hypothesis Tests:
  1. Null Hypothesis (H<em>0H<em>0): j=j</em>0j = j</em>0; Alternative Hypothesis (H<em>1H<em>1): j≠j</em>0j \neq j</em>0

  2. Estimate the model:

    • Y<em>i=0+1X</em>1i+…+jX<em>ji+…+kX</em>ki+u<em>iY<em>i = 0 + 1X</em>{1i} + … + jX<em>{ji} + … + kX</em>{ki} + u<em>i using OLS to obtain b</em>jb</em>j

  3. Compute the standard error of bjb_j (requires matrix algebra).

  4. Compute the t-statistic:

    • t<em>act=b</em>j−j<em>0SE(b</em>j)t<em>{act} = \frac{b</em>j - j<em>0}{SE(b</em>j)}

  5. Reject the null hypothesis if:

    • ∣tact∣>|t_{act}| > critical value

    • or if pp-value < significance level.

Example on Class Size Hypothesis Testing

  • H<em>0H<em>0: ClassSize = 0; H</em>1H</em>1: ClassSize \neq 0

    • Step 1: bClassSize=−1.10b_{ClassSize} = -1.10

    • Step 2: SE(bClassSize)=0.43SE(b_{ClassSize}) = 0.43

    • Step 3: Compute t-statistic:

    • tact=−1.10−00.43=−2.54t_{act} = \frac{-1.10 - 0}{0.43} = -2.54

    • Step 4: Reject the null hypothesis at 5% significance level if:

    • ∣−2.54∣>1.96|-2.54| > 1.96 and p−value=0.011<0.05p-value = 0.011 < 0.05.

  • Question: Do we reject H0H_0 at a 1% significance level?

Confidence Intervals for Single Coefficient in Multiple Regression

  • Robust regression output provides:

    • 95% confidence interval for ClassSize is also given in the output: [(−1.95213,−0.2504616-1.95213, -0.2504616)]

    • To calculate a 99% confidence interval for ClassSize:

    • Formula: b<em>ClassSize±2.58⋅SE(b</em>ClassSize)b<em>{ClassSize} \pm 2.58 \cdot SE(b</em>{ClassSize})

    • Resulting Interval:

    • 1.10±2.58⋅0.43=(2.21,0.01)1.10 \pm 2.58 \cdot 0.43 = (2.21, 0.01)

Hypothesis Tests on 2 or More Coefficients

  • Adding Variables Measuring Low-Income Family Background:

    • New variables:

      • meal_pct_i: measures % of students eligible for free lunch

      • calw_pct_i: measures % of students eligible for CALWORKS social assistance

    • Example regression command:

    regress test_score class_size el_pct meal_pct calw_pct, robust
    
    • Output Summary:

    • F-statistic: F(4,415)=361.68F(4, 415) = 361.68

    • Probability > F: 0.0000

    • R-squared: 0.7749

    • Root MSE: 9.0843

  • Regression Coefficients:

    • Class Size: Coef.=−1.014353Coef. = -1.014353, Std.Err.=0.2688613Std. Err. = 0.2688613, t=−3.77t = -3.77, P>∣t∣=0.000P>|t| = 0.000

    • Percentage of English Learners: Coef.=−0.1298219Coef. = -0.1298219, Std.Err.=0.0362579Std. Err. = 0.0362579, t=−3.58t = -3.58, P>∣t∣=0.000P>|t| = 0.000

    • Free Lunch: Coef.=−0.5286191Coef. = -0.5286191, Std.Err.=0.0381167Std. Err. = 0.0381167, t=−13.87t = -13.87, P>∣t∣=0.000P>|t| = 0.000

    • CALWORKS: Coef.=−0.0478537Coef. = -0.0478537, Std.Err.=0.0586541Std. Err. = 0.0586541, t=−0.82t = -0.82, P>∣t∣=0.415P>|t| = 0.415

Testing Hypotheses on Multiple Coefficients:
  1. Hypothesis for Meal Percent:

    • H<em>0H<em>0: mealpct = 0; H<em>1H<em>1: mealpct \neq 0

      • Step 1: bmealextpct=−0.5286191b_{meal ext{ pct}} = -0.5286191

      • Step 2: SE(bmealextpct)=0.038SE(b_{meal ext{ pct}}) = 0.038

      • Step 3: tmealextpct=−0.5286191−00.038=−13.87t_{meal ext{ pct}} = \frac{-0.5286191 - 0}{0.038} = -13.87 (reject null)

  2. Hypothesis for CALWORKS:

    • H<em>0H<em>0: calwpct = 0; H<em>1H<em>1: calwpct \neq 0

      • Step 1: bcalwextpct=−0.0478537b_{calw ext{ pct}} = -0.0478537

      • Step 2: SE(bcalwextpct)=0.059SE(b_{calw ext{ pct}}) = 0.059

      • Step 3: tcalwextpct=−0.0478537−00.059=−0.82t_{calw ext{ pct}} = \frac{-0.0478537 - 0}{0.059} = -0.82 (do not reject null)

Testing One Hypothesis on Two or More Coefficients

  • If testing hypothesis that both the coefficient on the % eligible for free lunch and the % eligible for CALWORKS equals zero:

    • H<em>0H<em>0: mealpct = 0 and calw_pct = 0

    • H<em>1H<em>1: mealpct \neq 0 and/or calw_pct \neq 0

  • Approach: Reject H<em>0H<em>0 if either t</em>mealextpctt</em>{meal ext{ pct}} or tcalwextpctt_{calw ext{ pct}} exceeds 1.96 (5% significance level)

  • Statistical Probability Relation:

    • If t<em>mealextpctt<em>{meal ext{ pct}} and t</em>calwextpctt</em>{calw ext{ pct}} are uncorrelated:

    Pr(t<em>mealextpct>1.96extand/ort</em>calwextpct>1.96)=1−(Pr(t<em>mealextpct≤1.96)imesPr(t</em>calwextpct≤1.96))=1−(0.95imes0.95)=0.0975>0.05Pr(t<em>{meal ext{ pct}} > 1.96 ext{ and/or } t</em>{calw ext{ pct}} > 1.96) = 1 - (Pr(t<em>{meal ext{ pct}} ≤ 1.96) imes Pr(t</em>{calw ext{ pct}} ≤ 1.96)) = 1 - (0.95 imes 0.95) = 0.0975 > 0.05

  • Testing with correlated statistics is more complicated.

F-statistic for Joint Hypothesis Testing
  • Joint hypotheses involving multiple coefficients require using the F-statistic:

    • The form of the F-statistic (with q=2q = 2 restrictions):

    F=12(t<em>12+t</em>22)(1−ρ<em>bt</em>1t2)F = \frac{1}{2} \left( t<em>1^2 + t</em>2^2 \right) \left( 1 - \rho<em>{bt</em>1t_2} \right)

    • The F-statistic considers correlation between individual t-statistics.

    • F-statistic is computed using software.

Example of F-test
  • Result summary using command for regression:

   test meal_pct calw_pct
  • Output gives:

    • F(2,415)=290.27F(2, 415) = 290.27

    • Probability > F: 0.0000 (reject H0H_0)

Testing Joint Hypothesis with q=3q = 3 Restrictions
  • Test Scenario: H<em>0H<em>0: elpct = 0 & mealpct = 0 & calwpct = 0

    • Compute the F-statistic:

    • FF-statistic output was previously F(3,415)=481.06F(3, 415) = 481.06.

  • Reject H0H_0 at 5% significance level:

    • Compare computed value against critical value from distribution (Fcritical=2.6F_{critical} = 2.6).

The Overall Regression F-statistic

  • The overall regression F-statistic tests if all slope coefficients are zero:

    • H<em>0H<em>0: β</em>1=0,β<em>2=0,…,β</em>k=0\beta</em>1 = 0, \beta<em>2 = 0, …, \beta</em>k = 0 (total of q=kq = k restrictions)

    • H1H_1: at least one slope coefficient is non-zero.

  • F-statistic computation results in:

    • Overall FF-statistic = 361.68 for the regression.

Testing Single Restrictions Involving Multiple Coefficients

  • Identify coefficients of interest, for instance, comparing mealpct with calwpct:

    • H<em>0H<em>0: mealpct = calwpct vs H</em>1H</em>1: mealpct \neq calwpct.

  1. Transform the regression model for null hypothesis simplicity.

  2. Conduct the test directly in statistical software.

Example of Testing Single Restrictions Directly
  • Command to test in R:

   test meal_pct = calw_pct
  • Output shows F-statistic:

    • F=27.32, significant with p-value 0.0000.

The R2R^2 and Adjusted R2R^2

  • R2R^2 indicates the variance explained by predictors:

    • Formula:

    R2=ESSTSS=∑(Yb−Yˉ)2∑(Y−Yˉ)2R^2 = \frac{ESS}{TSS} = \frac{\sum (Y_b - \bar{Y})^2}{\sum (Y - \bar{Y})^2}

    • $1 - \frac{SSR}{TSS}

  • Importance of R2R^2:

    • Increases with additional regressors even when explaining little.

  • Adjusted R2R^2 compensates for the number of predictors:

    • Formula:

    adjustedR2=1−(n−1n−k−1(1−R2))adjusted R^2 = 1 - \left( \frac{n-1}{n-k-1} (1 - R^2) \right)

  • Example output shows both R2R^2 and adjusted values improving model interpretations.

Interpreting Measures of Fit

  • R2R^2 indicates model prediction effectiveness.

    • Values near 1 suggest good prediction; near 0 indicates poor.

  • High R2R^2 alone is insufficient for drawing meaningful conclusions regarding variable significance, estimation validity, or complete model appropriateness.

Important Cautions
  • High R2R^2 does not imply:

    • Individual variable significance without further tests.

    • Causality implications of estimations without validity of OLS assumptions.

    • Exclusion of vital regressors leading to omitted variable bias.

Interpreting Statistical Significance in Outputs (Stars)
  • Regression output typically includes significance indicators:

    • Significance Levels: * (10%), ** (5%), *** (1%).

  • Example interpretation based on test scores and coefficients in regression output.

Conclusion

  • The chapter comprehensively covers hypothesis testing for multiple regressors in OLS regression, emphasizing statistical significance evaluation and various measures of fit to support data-driven conclusions.