1/99
A comprehensive 100-card flashcard review covering Mathematics & Statistics, OLS Foundations & Functional Form, Statistical Inference, Multivariate Regression & Specification, Panel Data, Causal Inference, Binary Dependent Variables, and Common Exam Questions for Empirical Methods in Finance.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Variance and Standard Deviation
Variance measures how dispersed X is around its mean: Var(X)=E[(X−E(X))2]. Standard deviation is: sd(X)=Var(X). A larger variance or standard deviation means observations are generally further from the mean. Standard deviation is expressed in the same units as X.
Covariance and Correlation
Covariance measures the linear co-movement of two variables: Cov(X,Y)=E[(X−E(X))(Y−E(Y))]. Cov(X,Y)>0 means X and Y tend to move together; Cov(X,Y)<0 means they tend to move in opposite directions; Cov(X,Y)=0 means no linear relationship. Correlation standardizes covariance: Corr(X,Y)=sd(X)sd(Y)Cov(X,Y). Correlation lies between −1 and 1 and does not depend on the units of measurement.
Independence vs. Zero Covariance
Independence means knowing X provides no information about Y: P(X=x,Y=y)=P(X=x)P(Y=y). If X and Y are independent, Cov(X,Y)=0. However, Cov(X,Y)=0 does not generally imply independence; zero covariance only rules out linear dependence, and a nonlinear relationship may still exist.
Conditional Probability and Dependence
Conditional probability asks how likely Y is after knowing X: P(Y=y∣X=x). If X and Y are independent, P(Y=y∣X=x)=P(Y=y), so knowing X does not change the probability of Y. If the conditional probability changes with X, the variables are dependent.
Mean vs. Median
The mean is the arithmetic average and is sensitive to extreme values. The median is the middle observation after ordering the data and is less sensitive to extreme values. Therefore, an extreme outlier can substantially change the mean while leaving the median relatively unchanged.
Regression Model and Error Term
In yi=β0+β1xi+ui, y is the dependent variable, x is the explanatory variable, and u contains all other factors affecting y that are not included in the model. β1 is the ceteris paribus effect: the change in y associated with a one-unit increase in x, holding other relevant factors constant.
Population Model vs. Estimated Regression
The population model contains the unknown true parameters: yi=β0+β1xi+ui. The estimated regression uses sample estimates: y^i=β^0+β^1xi. β1 is the unknown population effect, while β^1 is its estimate from a particular sample. Different random samples generally produce different estimates.
OLS and Residuals
OLS chooses β^0 and β^1 to minimize the sum of squared residuals: min∑u^i2, where u^i=yi−y^i. A residual is the difference between the observed and predicted value of y.
OLS Slope in Bivariate Regression
In a bivariate regression: β^1=∑(xi−xˉ)2∑(xi−xˉ)(yi−yˉ). The numerator captures how x and y move together; the denominator captures sample variation in x. Therefore, the sign of β^1 follows the sign of the sample covariance between x and y.
OLS Intercept
The OLS intercept is: β^0=yˉ−β^1xˉ. Therefore, the OLS regression line passes through (xˉ,yˉ). The intercept gives predicted y when x=0, although this interpretation is only economically meaningful when x=0 is relevant.
OLS Residual Properties
When OLS includes an intercept, ∑u^i=0 and ∑xiu^i=0. Thus, residuals have sample mean zero and are sample-uncorrelated with the included explanatory variable.
Unbiasedness
An estimator is unbiased if its expected value across repeated random samples equals the true population parameter: E(β^)=β. This does not mean every individual estimate equals β. Individual estimates can be above or below β without systematic error on average.
Consistency
An estimator is consistent if it converges to the true population parameter as the sample size becomes large: β^pβ. Unbiasedness concerns the average estimate across repeated samples; consistency concerns what happens as n→∞.
Efficiency and BLUE
BLUE means Best Linear Unbiased Estimator. Linear: the estimator is linear in y. Unbiased: E(β^)=β. Best: it has the smallest variance among linear unbiased estimators. Under the Gauss-Markov assumptions A1-A5, OLS is BLUE.
CLRM Assumptions
A1: linear in parameters.
A2: random sample from the population.
A3: variation in x / no perfect collinearity.
A4: zero conditional mean, E(ui∣Xi)=0.
A5: homoskedasticity, Var(ui∣Xi)=σ2.
Assumptions A1-A4 imply unbiased OLS; adding A5 gives efficiency and therefore BLUE.
Zero Conditional Mean
The crucial assumption for OLS unbiasedness is: E(ui∣Xi)=0. The average unobserved factors contained in u do not systematically vary with the explanatory variables. If an omitted determinant of y is correlated with an included explanatory variable, A4 is violated and OLS will generally be biased.
R-Squared
R2 measures the proportion of sample variation in y explained by the regression: R2=1−SSTSSR=TSSESS. A higher R2 means better in-sample fit. It does not imply that the model is causal, unbiased or correctly specified.
Effect of Rescaling y
If y is multiplied by a constant c, all estimated coefficients and their standard errors are also multiplied by c. Therefore, t=SE(β^)β^ is unchanged. R2, t-statistics, p-values, and statistical significance are unaffected by changing the units of y.
Effect of Rescaling x
If x is multiplied by c, its estimated coefficient and standard error are divided by c. Because both change by the same factor, the t-statistic, p-value, and statistical significance are unchanged. Rescaling changes the numerical coefficient because its units change, not the underlying economic relationship.
Standardized Coefficient
A standardized coefficient is: β^std=β^(sysx). It measures how many standard deviations predicted y changes when x increases by one standard deviation.
Level-Level, Level-Log, Log-Level & Log-Log
Level-Level: a one-unit increase in x changes y by β units.
Level-Log: a 1% increase in x changes y by approximately 100β units.
Log-Level: a one-unit increase in x changes y by approximately 100β%.
Log-Log: a 1% increase in x changes y by approximately β%; β is an elasticity.
Quadratic Model
For y=β0+β1x+β2x2+u, the effect of x is not constant. The marginal effect is: ∂x∂y=β1+2β2x. Therefore, the effect of an additional unit of x depends on the current value of x.
Turning Point
For y=β0+β1x+β2x2+u, the turning point occurs where the marginal effect equals zero: x∗=−2β2β1. If β2<0, the turning point is a maximum. If β2>0, it is a minimum.
OLS Standard Error
The standard error measures the estimated sampling uncertainty of a coefficient estimate. For the slope in the bivariate homoskedastic case: SE(β^1)=∑(xi−xˉ)2σ^2. A smaller standard error means a more precise estimate.
t-Test
A t-test tests a hypothesis about one regression coefficient: t=SE(β^)β^−β0. For H0:β=0: t=SE(β^)β^. The t-statistic measures how many standard errors the estimate is from the hypothesized value.
Statistical Significance and Decision Rule
Using the p-value: p<α⟹reject H0; p≥α⟹fail to reject H0. For a two-sided critical-value test: ∣t∣>tcritical⟹reject H0. Statistical significance means the data provide sufficient evidence against the null at the chosen significance level.
Critical-Value Decision Rule
For a two-sided test, compare ∣t∣ with the relevant critical value. If ∣t∣>tcritical, reject H0. If ∣t∣≤tcritical, fail to reject H0. For a two-sided 10% test in a large sample, the critical value is approximately 1.64.
p-Value
The p-value measures how extreme the observed test statistic would be if H0 were true. For example, p=0.03 means that, assuming H0 is true, there is a 3% probability of observing a test statistic at least as extreme as the one observed. It is not the probability that H0 is true.
One-Sided vs. Two-Sided Test
A two-sided test asks whether the coefficient differs in either direction: H1:β=β0. A one-sided test asks whether it differs in one specified direction, for example H1:β>β0 or H1:β<β0.
Confidence Interval
A confidence interval gives a range of plausible values for the population coefficient. A large-sample 95% interval is approximately: β^±1.96SE(β^). A two-sided H0:β=β0 is rejected at the corresponding level if β0 lies outside the interval.
Statistical vs. Economic Significance
Statistical significance asks whether there is sufficient statistical evidence that an effect differs from the hypothesized value. Economic significance asks whether the magnitude is economically or practically important. An effect can be statistically significant but economically small, or economically important but imprecisely estimated.
A6 - Normality
The population error u is independent of the explanatory variables and normally distributed: u∼N(0,σ2). Together with A1-A5, A6 gives a normal sampling distribution for the OLS estimator and allows exact t-inference. A6 is not required for OLS unbiasedness; it is added for statistical inference.
Endogeneity
Endogeneity occurs when an explanatory variable is correlated with the error term: Cov(x,u)=0. Then OLS is generally biased and inconsistent and cannot isolate the causal effect of x on y. Main causes include omitted variables, simultaneity/reverse causality, and measurement error in an explanatory variable.
Omitted Variable Bias
OVB occurs when an omitted variable both affects y and is correlated with an included explanatory variable. The omitted variable enters u, making the included explanatory variable correlated with u and violating zero conditional mean.
Direction of Omitted Variable Bias
Suppose the true model is y=β0+β1x+β2z+u but z is omitted. Then: E(β^1)=β1+β2δ~1, so: Bias(β^1)=β2δ~1, where δ~1 measures the relationship between z and x. Same signs for the effect z→y and Corr(x,z) imply upward bias; opposite signs imply downward bias.
Simultaneity / Reverse Causality
Simultaneity occurs when x affects y while y also affects x. This makes x correlated with the error term and violates zero conditional mean. OLS is therefore biased and inconsistent and cannot isolate the one-way causal effect of x on y.
Measurement Error in x
Suppose observed x equals true x plus classical measurement error: xobs=x+e. Classical measurement error in an explanatory variable generally biases its estimated coefficient toward zero. This is attenuation bias.
Measurement Error in y
With classical measurement error yobs=y+e, where e is unrelated to the explanatory variables, the OLS slope coefficients remain unbiased. However, measurement error adds noise, increasing uncertainty and reducing precision.
Proxy Variable
A proxy variable is an observed variable used to control for an important unobserved explanatory variable. A useful proxy contains information about the unobserved factor. Including an appropriate proxy can reduce omitted-variable bias.
Heteroskedasticity
Homoskedasticity requires Var(u∣X)=σ2. Heteroskedasticity means Var(u∣X) changes with X. If A1-A4 hold, OLS coefficients remain unbiased and consistent, but OLS is no longer BLUE and conventional standard errors can be incorrect.
Robust Standard Errors
Heteroskedasticity-robust standard errors allow valid inference when heteroskedasticity is present. They can change standard errors, t-statistics, p-values, and confidence intervals. They do not change the OLS coefficient estimates.
Breusch-Pagan Test
The Breusch-Pagan test tests for heteroskedasticity. Estimate OLS and obtain residuals; regress squared residuals on the explanatory variables; test whether the slopes are jointly zero. H0: homoskedasticity. A small p-value provides evidence of heteroskedasticity.
Irrelevant Variables and Multicollinearity
Including an irrelevant variable does not bias OLS, but can increase standard errors and reduce precision. Multicollinearity means explanatory variables are strongly correlated; it does not bias OLS but increases standard errors. Perfect multicollinearity prevents estimation.
F-Test
An F-test tests multiple coefficient restrictions jointly: F=SSRUR/(n−k−1)(SSRR−SSRUR)/q, where q is the number of restrictions. Reject H0 if F is sufficiently large or p<α. Unlike a t-test, an F-test can test several restrictions jointly.
df1=q and
Dummy Variable and Reference Group
A dummy variable takes values 0 or 1. In y=β0+β1D+u, β0 is the expected value for the reference group D=0, and β1 is the difference between D=1 and D=0.
Dummy Variable Trap
With an intercept, including a dummy for every category creates perfect multicollinearity because the dummies sum to one. One category must be omitted. The omitted category becomes the reference group, and included dummy coefficients measure differences relative to it.
Dummy Variable with Log y
In ln(y)=β0+β1D+⋯+u, the approximate percentage difference between D=1 and D=0 is 100β1%. The exact percentage difference is: 100(eβ1−1)%.
Interaction Terms
For continuous variables, y=β0+β1x1+β2x2+β3x1x2+u gives: ∂x1∂y=β1+β3x2. For a dummy × continuous interaction, y=β0+β1x+β2D+β3(Dx)+u. The slope of x is β1 for D=0 and β1+β3 for D=1; β3 is the difference in slopes.
Adjusted R-Squared
Adjusted R2 measures model fit while penalizing additional explanatory variables: Rˉ2=1−n−k−1(1−R2)(n−1). Unlike R2, adjusted R2 can decrease when a variable is added if that variable does not improve the model sufficiently.
First Differences
Starting with yit=αi+βxit+uit, first differencing gives: Δyit=βΔxit+Δuit. The time-invariant effect αi disappears. β is estimated from changes within the same unit over time.
Attenuation Bias under First Differencing
If xobs,it=xit+eit, then: Δxobs,it=Δxit+(eit−ei,t−1).
With independent measurement errors: Var(eit−ei,t−1)=2σe2.
Differencing can increase measurement-error noise relative to true variation in x, making attenuation bias more severe.
Fixed Effects
Fixed Effects controls for all time-invariant characteristics of each unit. Starting with yit=αi+βxit+uit, demeaning gives: yit−yˉi=β(xit−xˉi)+(uit−uˉi). The time-invariant αi disappears. FE estimates β using within-unit variation over time.
Time-Invariant Variables under Fixed Effects
A variable that does not change over time within a unit is eliminated by the Fixed Effects transformation. Therefore, its separate coefficient cannot be estimated with unit fixed effects.
Time Fixed Effects and Two-Way Fixed Effects
Time fixed effects control for shocks affecting all units in the same period. Two-Way Fixed Effects combine unit and time effects: yit=αi+λt+βxit+uit. αi controls permanent unit differences; λt controls common period shocks.
Fixed Effects vs. First Differences
If T=2, FE and FD produce identical slope estimates. If T>2, they need not. FE generally preserves more observations in an unbalanced panel and measurement-error bias may decrease as T grows. FD requires consecutive observations and can make attenuation bias more severe.
Random Effects
Random Effects models the unit-specific effect as random: yit=β0+β1xit+vi+uit. The crucial RE assumption is: Cov(xit,vi)=0 for all periods. Thus, the unit-specific unobserved effect must be uncorrelated with the explanatory variables.
Fixed Effects vs. Random Effects
FE allows the unit-specific effect to be correlated with explanatory variables, but cannot estimate coefficients on time-invariant variables. RE requires no such correlation, can estimate time-invariant effects, and is more efficient if its stronger assumptions hold.
Hausman Test
The Hausman test compares FE and RE estimates. H0: RE is consistent. If p<α, reject H0: evidence suggests the unit-specific effect is correlated with regressors, so FE is preferred. If p≥α, fail to reject H0.
Clustered Standard Errors
Clustered standard errors allow errors to be correlated and heteroskedastic within a cluster. Firm clustering allows serial/time correlation within a firm; time clustering allows cross-sectional correlation across firms in the same period. Without appropriate clustering, SEs may be underestimated and t-statistics overstated. Clustering changes inference, not coefficients.
Fixed Effects vs. Clustered Standard Errors
Fixed Effects removes time-invariant unobserved heterogeneity. Clustering adjusts inference for correlation of errors within clusters. They solve different problems, so using FE does not eliminate the need for clustered standard errors when within-cluster error correlation is present.
Instrumental Variables
IV estimates the causal effect of endogenous x using an instrument z. A valid instrument must satisfy relevance, Cov(z,x)=0, and exogeneity, Cov(z,u)=0. z must predict x but be unrelated to unobserved determinants of y.
Instrument Relevance and Weak Instruments
Test relevance using the first stage: xi=π0+π1zi+controls+vi. A common rule of thumb is first-stage F>10 for sufficient instrument strength. A weak instrument provides little exogenous variation in x, making IV estimates unstable, imprecise and potentially unreliable.
Instrument Exogeneity
Instrument exogeneity requires: Cov(z,u)=0. The instrument must be unrelated to unobserved determinants of y. Unlike relevance, exogeneity generally cannot be tested directly because u is unobserved; it requires a credible economic or institutional argument.
Two-Stage Least Squares (2SLS)
Stage 1: regress endogenous x on instrument z and all exogenous controls to obtain predicted x^. Stage 2: regress y on x^ and the exogenous controls. The coefficient uses variation in x generated by the instrument to estimate the causal effect.
OLS vs. IV
OLS uses all variation in x, including variation that may be correlated with u. IV uses variation generated by a valid instrument. If x is endogenous and z is valid, OLS is inconsistent while IV is consistent but generally less precise. With one x and one z: β^IV=Cov(z,x)Cov(z,y).
Difference-in-Differences: Basic Idea
DiD compares the change in the treated group with the change in the control group: DiD=(yˉT,post−yˉT,pre)−(yˉC,post−yˉC,pre). The control group's change captures the common time effect, which is subtracted from the treated group's change.
Difference-in-Differences Regression
Traditional DiD: yit=β0+β1Treati+β2Postt+β3(Treati×Postt)+uit, where β3 is the DiD treatment effect. With unit and time FE: yit=αi+λt+βDit+uit, where Dit=Treati×Postt. Unit FE absorb Treat; time FE absorb Post. β is the DiD effect, assuming parallel trends.
Parallel Trends Assumption
Without treatment, treated and control groups would have experienced the same average change in the outcome over time. This is the key identifying assumption of DiD. The post-treatment counterfactual cannot be directly verified, but similar pre-treatment trends can provide supporting evidence.
Regression Discontinuity Design
RDD exploits a treatment rule based on whether a running variable r crosses a cutoff c. Observations just below and above the cutoff are compared. The key assumption is continuity of potential outcomes at the cutoff; therefore a discontinuous jump can be attributed to treatment. RDD identifies a local effect around the cutoff.
RDD Regression
A sharp RDD can be written: yi=β0+β1Di+β2(ri−c)+β3[Di(ri−c)]+ui
with Di=1(ri≥c).
β1 is the treatment effect at the cutoff; β2 is the slope below; β2+β3 is the slope above; β3 is the slope difference. Identification requires continuity of potential outcomes at c.
Binary Dependent Variable
A binary dependent variable takes two possible values: yi∈{0,1}. For binary y: E(yi∣Xi)=P(yi=1∣Xi). Therefore, the conditional mean of y can be interpreted as a probability.
Linear Probability Model
The LPM models P(yi=1∣Xi)=β0+β1x1i+⋯+βkxki and is estimated using OLS. βj is the change in the probability that y=1 from a one-unit increase in xj, holding other variables constant. β=0.08 means an 8 percentage-point increase.
Advantages and Problems of the LPM
Advantages: easy to estimate with OLS and coefficients have a direct probability interpretation. Problems: predicted probabilities can fall below 0 or above 1; marginal effects are constant; the error is inherently heteroskedastic. Therefore, robust standard errors should be used.
Logit vs. Probit
Let z=β0+Xβ. Logit: P(y=1∣X)=1+ezez. Probit: P(y=1∣X)=Φ(z). Logit uses the logistic CDF; Probit uses the standard normal CDF. Both produce S-shaped probability curves and probabilities between 0 and 1.
Maximum Likelihood Estimation
Logit and Probit are nonlinear probability models estimated using Maximum Likelihood Estimation rather than OLS. MLE chooses parameter values that make the observed sample outcomes y=0 and y=1 as likely as possible under the model.
Logit/Probit Coefficient Interpretation
The sign of β gives the direction of the effect:
β>0 means x increases P(y=1);
β<0 means x decreases P(y=1).
But β itself is not the change in probability. For example, β=0.20 does not mean probability rises by 20 percentage points. Use marginal effects or predicted probabilities for magnitude.
Marginal Effects in Logit/Probit
A marginal effect converts a Logit/Probit coefficient into an effect on predicted probability. Because these models are nonlinear, the marginal effect generally depends on X. The Average Marginal Effect calculates the marginal effect for each observation and averages them. AME=0.06 means a 6 percentage-point increase on average.
Marginal Effect of a Dummy Variable
For a dummy D, calculate the discrete probability change: P(y=1∣D=1,X)−P(y=1∣D=0,X). For an Average Marginal Effect, calculate this probability difference for every observation and then take the average.
y^=25.4+3.20experience. Interpret experience.
A one-unit increase in experience is associated with a 3.20-unit increase in predicted salary, holding other included variables constant. State: (1) one-unit change in x, (2) β-unit change in y, and (3) ceteris paribus.
y^=30.2+8.50ln(sales). Interpret sales.
A 1% increase in sales is associated with an 1008.50=0.085 unit increase in predicted salary, holding other variables constant. Rule: in a level-log model, a 1% increase in x implies a 100β units change in y.
ln(y^)=6.20+0.035experience. Interpret experience.
A one-unit increase in experience is associated with an approximately 100(0.035)=3.5% increase in predicted salary, holding other variables constant. Rule: in a log-level model, a one-unit increase in x implies an approximately 100β% change in y.
ln(y^)=5.80+0.24ln(sales). Interpret sales.
A 1% increase in sales is associated with an approximately 0.24% increase in predicted salary, holding other variables constant. In a log-log model, β is an elasticity.
wage=12.5+2.40female+0.80education. Interpret female.
Holding education constant, the group with female=1 has predicted wage 2.40 units higher than the reference group female=0. Identify the D=1 group, the D=0 reference group, and the β-unit difference.
ln(wage)=2.30−0.15female+0.08education. Interpret female.
Approximation: 100(−0.15)=−15%, so female=1 has approximately 15% lower predicted wage than female=0, holding education constant. Exact effect: 100(e−0.15−1)≈−13.9%.
salary=30+4experience−0.10experience2. Effect at experience=10?
Do not interpret 4 alone. Compute the marginal effect: ∂experience∂salary=4−2(0.10)experience. At experience=10: 4−0.20(10)=2. One additional year increases predicted salary by approximately 2 units at experience=10.
β^=0.30, SE(β^)=0.12, p=0.012. Two-sided H0:β=0 at 5%.
Calculate t=0.120.30=2.50. Critical-value method: ∣2.50∣>1.96⟹reject H0. p-value method: 0.012<0.05⟹reject H0. Both methods give the same decision. Conclusion: β is statistically significantly different from zero at the 5% level.
β^=0.30, SE(β^)=0.20. Test H0:β≤0 vs H1:β>0 at 5%.
Calculate t=0.200.30=1.50. For a one-sided 5% test, tcritical≈1.645. Since 1.50<1.645, fail to reject H0. There is insufficient evidence that β is positive at the 5% level.
mortgage=20+0.50income; SE(β^income)=0.20. Significant at 10%? Interpret.
t=0.200.50=2.50. For a two-sided 10% test, ∣2.50∣>1.64, so income is statistically significant. A one-unit increase in income is associated with a 0.50-unit increase in predicted mortgage, holding other variables constant. Separate statistical significance from economic interpretation.
H0:β1=β2=β3=0. Output: F=4.80, p=0.003. Test at 5%.
Since 0.003<0.05, reject H0. β1, β2, and β3 are jointly statistically significant at the 5% level. Do not conclude that every coefficient is individually significant.
SSRR=500, SSRUR=400, n=100, k=4, q=2. Calculate F.
Use F=SSRUR/(n−k−1)(SSRR−SSRUR)/q. Thus F=400/(100−4−1)(500−400)/2=4.2150≈11.88. Then compare F with the appropriate critical value or use its p-value.
Regression output reports R2=0.64. Interpret it.
64% of the sample variation in the dependent variable is explained by the explanatory variables included in the regression. Do not say that 64% of y is explained, that the model is 64% correct, or that 64% of the relationship is causal.
employed=0.25+0.06education−0.03age. Interpret education.
Because this is an LPM, a one-unit increase in education is associated with a 0.06=6 percentage-point increase in the predicted probability of being employed, holding age constant. Use percentage points, not percent.
Logit model: β^education=0.40. Interpret the effect on P(employed=1).
Because 0.40>0, education has a positive effect on the predicted probability of employment. But 0.40 is not a 40 percentage-point increase. A Logit coefficient itself is not a probability change; use a marginal effect for magnitude.
Logit model reports AMEeducation=0.045. Interpret it.
A one-unit increase in education increases the predicted probability that y=1 by approximately 0.045=4.5 percentage points on average, holding other variables constant. Say percentage points, not percent.
P(y=1∣D=1)=0.62 and P(y=1∣D=0)=0.48. Interpret D.
Calculate 0.62−0.48=0.14. Changing D from 0 to 1 increases the predicted probability that y=1 by 14 percentage points.
True model includes wealth; wealth raises y and is positively correlated with income. Wealth is omitted. Bias?
Both relevant signs are positive: wealth→y is positive and Corr(income,wealth) is positive. Therefore, (+)×(+)=(+), so the estimated income coefficient is biased upward. In an OVB answer, identify both relationships before stating the bias direction.
performance=β0+β1investment+u; expected performance also raises investment. Causal β1?
No. Expected performance affects investment, creating reverse causality and Cov(investment,u)=0. Investment is endogenous, so OLS is generally biased and inconsistent. A valid IV or another credible identification strategy is needed for causal interpretation.
x is endogenous and z is proposed as an instrument. What must z satisfy?
Relevance: Cov(z,x)=0, so z must predict endogenous x. Exogeneity: Cov(z,u)=0, so z must be unrelated to unobserved determinants of y. Both conditions are required for a valid instrument.
Treated: 50→70. Control: 40→48. Calculate the DiD effect.
Treated change = 70−50=20. Control change = 48−40=8. DiD=20−8=12. The estimated treatment effect is 12 units, assuming the parallel-trends assumption holds.
yit=αi+βxit+uit. What variation identifies β?
β is identified from changes in x within the same unit over time. The unit fixed effect αi absorbs all observed and unobserved characteristics that are constant over time within that unit. β is not identified from permanent differences between units.