Assessing Studies Based on Multiple Regression

Framework for Assessing Statistical Studies

  • Goal of Econometric Assessment: The primary objective is to evaluate the reliability and generalizability of regression results using two criteria: internal validity and external validity.

  • Internal Validity: Statistical inferences about causal effects are considered internally valid if they are accurate for the specific population being studied.

  • External Validity: Statistical inferences are externally valid if they can be generalized from the specific population and setting studied to other populations and settings.

    • Setting: Refers to the legal, policy, and physical environment, alongside other salient features of the study's context.

Threats to External Validity

Assessing external validity requires substantive knowledge and case-by-case judgment. In the context of class size studies (e.g., California), one must evaluate if results apply elsewhere based on:

  • Differences in Populations:

    • Would results from California in 2011 hold for Massachusetts in 2011 or Mexico in 2011?

  • Differences in Settings:

    • Different legal requirements (e.g., mandates for special education).

    • Distinctions in how bilingual education is treated.

  • Teacher Characteristics: Variations in training, experience, or certifications across different regions.

Threats to Internal Validity of Multiple Regression

Internal validity fails if the conditional mean independence assumption is violated: E(uiX1i,,Xki)0E(u_i | X_{1i}, \dots, X_{ki}) \neq 0. When this occurs, the Ordinary Least Squares (OLS) estimator is biased and inconsistent. The five primary threats are:

  1. Omitted Variable Bias

  2. Functional Form Misspecification

  3. Errors-in-Variables Bias

  4. Sample Selection Bias

  5. Simultaneous Causality Bias

1. Omitted Variable Bias (OVB)

  • Conditions for Bias: OVB occurs if an omitted variable is both:

    1. A determinant of the dependent variable YY.

    2. Correlated with at least one included regressor XX.

  • In Multiple Regression: One must ask if the error term remains correlated with the variable of interest even after including available control variables.

  • Solutions to OVB:

    1. Direct Inclusion: Measure the omitted causal variable and include it as a regressor.

    2. Proxy/Control Variables: Include controls that ensure conditional mean independence plausibly holds.

    3. Panel Data: Observe the same entities (individuals, districts) over multiple time periods to control for unobserved factors.

    4. Instrumental Variables (IV): Use IV regression if the omitted variable cannot be measured.

    5. Randomized Controlled Experiments: Random assignment of XX ensures it is independent of uu, forcing E(uX)=0E(u | X) = 0.

2. Functional Form Misspecification

  • Definition: This arises when the functional form of the regression is incorrect (e.g., ignoring a necessary interaction term), leading to biased causal inferences.

  • Solutions:

    • Continuous Dependent Variables: Utilize nonlinear specifications such as logarithms or interaction terms.

    • Discrete (Binary) Dependent Variables: Extend OLS methods to specialized models like Probit or Logit analysis.

3. Errors-in-Variables Bias

This bias occurs when a regressor XX is measured with error. Sources of error include data entry mistakes, recollection errors in surveys, ambiguous questions, or intentionally false responses regarding sensitive topics (e.g., income or illegal behavior).

  • Classical Measurement Error Model:

    • Defined as X~i=Xi+vi\tilde{X}_i = X_i + v_i, where viv_i is mean-zero random noise.

    • Assumes corr(Xi,vi)=0corr(X_i, v_i) = 0 and corr(ui,vi)=0corr(u_i, v_i) = 0.

    • Effect: The estimator β^1\hat{\beta}_1 is biased toward zero (attenuation bias). In the extreme case where the variable is total noise, the coefficient expectation becomes zero.

  • Best Guess Measurement Error: An alternative model where respondents provide their most accurate estimate (not detailed in this section).

  • Practical Implications:

    • Administrative data (e.g., number of teachers) tend to be accurate.

    • Survey data on sensitive topics usually involve significant measurement error.

  • Solutions:

    1. Obtain higher quality data.

    2. Develop a specific mathematical model for the measurement error process (requires specialized data cross-checking).

    3. Use Instrumental Variables regression.

4. Missing Data and Sample Selection Bias

  • Case 1: Data Missing at Random:

    • Example: A dog eats random response sheets. This is equivalent to a smaller random sample; it reduces precision (larger standard errors) but does not introduce bias.

  • Case 2: Data Missing Based on Values of X:

    • Example: Studying only school districts with STR < 20. While this limits the scope of the findings, it does not bias the OLS estimator for that specific subset.

  • Case 3: Sample Selection Bias (Missing Based on Y or u):

    • Bias occurs when the selection process influences data availability and is related to the dependent variable.

Examples of Sample Selection Bias:

  • Height of Undergraduates: Measuring height only by sampling students outside the basketball locker room leads to an upward-biased estimate of the population mean.

  • Mutual Funds: Evaluating the 10-year performance of funds available today. This ignores "failed" funds that closed during the decade, resulting in an overestimation of average returns (survivorship bias).

  • Returns to Education: Sampling only employed graduates to estimate the return on education. Because employment status is related to the error term in the wage equation, the estimate is biased.

Solutions to Sample Selection Bias:

  • Collect samples in a way that avoids selection (e.g., sampling from enrollment lists rather than specific locations).

  • For mutual funds: Use the population of funds available at the beginning of the period.

  • For education: Sample all graduates, including the unemployed.

  • Use randomized controlled experiments or construct specialized selection models.

5. Simultaneous Causality Bias

  • Definition: This occurs when XX causes YY, but YY also causes XX.

  • Mathematical Representation:

    • (a) Y_i = \beta_0 + ̢\beta_1 X_i + u_i

    • (b) Xi=γ0+γ1Yi+viX_i = \gamma_0 + \gamma_1 Y_i + v_i

  • Result: A large uiu_i leads to a large YiY_i, which in turn influences XiX_i. This makes corr(Xi,ui)0corr(X_i, u_i) \neq 0, rendering β^1\hat{\beta}_1 biased and inconsistent.

  • Real-World Examples:

    • Test Scores and Resources: Low test scores may trigger a political process that provides more resources to the district, thereby lowering the Student-Teacher Ratio (STRSTR).

    • Police and Violence: High violence levels may lead to hiring more police, making it difficult to isolate the causal effect of police on crime reduction.

  • Solutions:

    1. Randomized Controlled Experiments: Prevents feedback from YY to XX because XX is assigned randomly.

    2. Complete Modeling: Estimate both directions of causality simultaneously (difficult in practice; used in some macro models).

    3. Instrumental Variables Regression: Used to estimate the specific causal effect of XX on YY.

Internal and External Validity in Forecasting

Forecasting has different priorities than causal estimation:

  • Objective: Fit and reliability in future applications are more important than coefficient interpretation.

  • R-Squared: The value of R2R^2 matters significantly for predictive power.

  • Omitted Variables: OVB is not a concern for forecasting; you do not need a causal interpretation.

  • External Validity: This is the most critical factor, as the model built on historical data must remain valid for the near future.

Application: Assessing Class Size and Test Score Data

External Validity
  • Comparing California and Massachusetts results is necessary to determine if the findings are robust across different states.

Internal Validity Assessment for Test Score Case Study
  1. Omitted Variable Bias: Potential missing factors include student native ability, access to outside learning, and teacher quality. Regressions attempt to control for these using local demographics (IncomesIncomes, \% Free Lunch) and the fraction of English learners. The stability of the STRSTR coefficient across different specifications suggests OVB may be minimal, but judgment is still required.

  2. Wrong Functional Form: Various nonlinear forms were tested (e.g., logarithms, interactions). Because nonlinear effects appeared modest, this is likely not a major threat.

  3. Errors-in-Variables Bias: The use of administrative data makes typo-related errors unlikely. However, a "complicated" measurement error exists because district-wide STRSTR may not match the specific experience of every student taking the test. Data at the individual student level would be ideal.

  4. Sample Selection Bias: Since the sample encompasses all elementary public school districts with no missing data, selection bias is judged to be non-existent.

  5. Simultaneous Causality Bias: This would occur if funding were equalized based on test scores. During the sample period, such programs were not in place, making this bias unlikely.