Empirical Methods in Finance

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/99

flashcard set

Earn XP

Description and Tags

A comprehensive 100-card flashcard review covering Mathematics & Statistics, OLS Foundations & Functional Form, Statistical Inference, Multivariate Regression & Specification, Panel Data, Causal Inference, Binary Dependent Variables, and Common Exam Questions for Empirical Methods in Finance.

Last updated 9:18 AM on 10/8/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

100 Terms

1
New cards

Variance and Standard Deviation

Variance measures how dispersed XX is around its mean: Var(X)=E[(X−E(X))2]\text{Var}(X) = E[(X - E(X))^2]. Standard deviation is: sd(X)=Var(X)\text{sd}(X) = \sqrt{\text{Var}(X)}. A larger variance or standard deviation means observations are generally further from the mean. Standard deviation is expressed in the same units as XX.

2
New cards

Covariance and Correlation

Covariance measures the linear co-movement of two variables: Cov(X,Y)=E[(X−E(X))(Y−E(Y))]\text{Cov}(X,Y) = E[(X - E(X))(Y - E(Y))]. Cov(X,Y)>0\text{Cov}(X,Y) > 0 means XX and YY tend to move together; Cov(X,Y)<0\text{Cov}(X,Y) < 0 means they tend to move in opposite directions; Cov(X,Y)=0\text{Cov}(X,Y) = 0 means no linear relationship. Correlation standardizes covariance: Corr(X,Y)=Cov(X,Y)sd(X) sd(Y)\text{Corr}(X,Y) = \frac{\text{Cov}(X,Y)}{\text{sd}(X)\,\text{sd}(Y)}. Correlation lies between −1-1 and 11 and does not depend on the units of measurement.

3
New cards

Independence vs. Zero Covariance

Independence means knowing XX provides no information about YY: P(X=x,Y=y)=P(X=x)P(Y=y)P(X=x, Y=y) = P(X=x)P(Y=y). If XX and YY are independent, Cov(X,Y)=0\text{Cov}(X,Y) = 0. However, Cov(X,Y)=0\text{Cov}(X,Y) = 0 does not generally imply independence; zero covariance only rules out linear dependence, and a nonlinear relationship may still exist.

4
New cards

Conditional Probability and Dependence

Conditional probability asks how likely YY is after knowing XX: P(Y=y∣X=x)P(Y=y \mid X=x). If XX and YY are independent, P(Y=y∣X=x)=P(Y=y)P(Y=y \mid X=x) = P(Y=y), so knowing XX does not change the probability of YY. If the conditional probability changes with XX, the variables are dependent.

5
New cards

Mean vs. Median

The mean is the arithmetic average and is sensitive to extreme values. The median is the middle observation after ordering the data and is less sensitive to extreme values. Therefore, an extreme outlier can substantially change the mean while leaving the median relatively unchanged.

6
New cards

Regression Model and Error Term

In yi=β0+β1xi+uiy_i = \beta_0 + \beta_1 x_i + u_i, yy is the dependent variable, xx is the explanatory variable, and uu contains all other factors affecting yy that are not included in the model. β1\beta_1 is the ceteris paribus effect: the change in yy associated with a one-unit increase in xx, holding other relevant factors constant.

7
New cards

Population Model vs. Estimated Regression

The population model contains the unknown true parameters: yi=β0+β1xi+uiy_i = \beta_0 + \beta_1 x_i + u_i. The estimated regression uses sample estimates: y^i=β^0+β^1xi\hat{y}_i = \hat{\beta}_0 + \hat{\beta}_1 x_i. β1\beta_1 is the unknown population effect, while β^1\hat{\beta}_1 is its estimate from a particular sample. Different random samples generally produce different estimates.

8
New cards

OLS and Residuals

OLS chooses β^0\hat{\beta}_0 and β^1\hat{\beta}_1 to minimize the sum of squared residuals: min⁡∑u^i2\min \sum \hat{u}_i^2, where u^i=yi−y^i\hat{u}_i = y_i - \hat{y}_i. A residual is the difference between the observed and predicted value of yy.

9
New cards

OLS Slope in Bivariate Regression

In a bivariate regression: β^1=∑(xi−xˉ)(yi−yˉ)∑(xi−xˉ)2\hat{\beta}_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2}. The numerator captures how xx and yy move together; the denominator captures sample variation in xx. Therefore, the sign of β^1\hat{\beta}_1 follows the sign of the sample covariance between xx and yy.

10
New cards

OLS Intercept

The OLS intercept is: β^0=yˉ−β^1xˉ\hat{\beta}_0 = \bar{y} - \hat{\beta}_1 \bar{x}. Therefore, the OLS regression line passes through (xˉ,yˉ)(\bar{x}, \bar{y}). The intercept gives predicted yy when x=0x = 0, although this interpretation is only economically meaningful when x=0x = 0 is relevant.

11
New cards

OLS Residual Properties

When OLS includes an intercept, ∑u^i=0\sum \hat{u}_i = 0 and ∑xiu^i=0\sum x_i \hat{u}_i = 0. Thus, residuals have sample mean zero and are sample-uncorrelated with the included explanatory variable.

12
New cards

Unbiasedness

An estimator is unbiased if its expected value across repeated random samples equals the true population parameter: E(β^)=βE(\hat{\beta}) = \beta. This does not mean every individual estimate equals β\beta. Individual estimates can be above or below β\beta without systematic error on average.

13
New cards

Consistency

An estimator is consistent if it converges to the true population parameter as the sample size becomes large: β^→pβ\hat{\beta} \xrightarrow{p} \beta. Unbiasedness concerns the average estimate across repeated samples; consistency concerns what happens as n→∞n \rightarrow \infty.

14
New cards

Efficiency and BLUE

BLUE means Best Linear Unbiased Estimator. Linear: the estimator is linear in yy. Unbiased: E(β^)=βE(\hat{\beta}) = \beta. Best: it has the smallest variance among linear unbiased estimators. Under the Gauss-Markov assumptions A1-A5, OLS is BLUE.

15
New cards

CLRM Assumptions

A1: linear in parameters.

A2: random sample from the population.

A3: variation in xx / no perfect collinearity.

A4: zero conditional mean, E(ui∣Xi)=0E(u_i \mid X_i) = 0.

A5: homoskedasticity, Var(ui∣Xi)=σ2\text{Var}(u_i \mid X_i) = \sigma^2.

Assumptions A1-A4 imply unbiased OLS; adding A5 gives efficiency and therefore BLUE.

16
New cards

Zero Conditional Mean

The crucial assumption for OLS unbiasedness is: E(ui∣Xi)=0E(u_i \mid X_i) = 0. The average unobserved factors contained in uu do not systematically vary with the explanatory variables. If an omitted determinant of yy is correlated with an included explanatory variable, A4 is violated and OLS will generally be biased.

17
New cards

R-Squared

R2R^2 measures the proportion of sample variation in yy explained by the regression: R2=1−SSRSST=ESSTSSR^2 = 1 - \frac{\text{SSR}}{\text{SST}} = \frac{\text{ESS}}{\text{TSS}}. A higher R2R^2 means better in-sample fit. It does not imply that the model is causal, unbiased or correctly specified.

18
New cards

Effect of Rescaling y

If yy is multiplied by a constant cc, all estimated coefficients and their standard errors are also multiplied by cc. Therefore, t=β^SE(β^)t = \frac{\hat{\beta}}{\text{SE}(\hat{\beta})} is unchanged. R2R^2, t-statistics, p-values, and statistical significance are unaffected by changing the units of yy.

19
New cards

Effect of Rescaling x

If xx is multiplied by cc, its estimated coefficient and standard error are divided by cc. Because both change by the same factor, the t-statistic, p-value, and statistical significance are unchanged. Rescaling changes the numerical coefficient because its units change, not the underlying economic relationship.

20
New cards

Standardized Coefficient

A standardized coefficient is: β^std=β^(sxsy)\hat{\beta}_{\text{std}} = \hat{\beta}\left(\frac{s_x}{s_y}\right). It measures how many standard deviations predicted yy changes when xx increases by one standard deviation.

21
New cards

Level-Level, Level-Log, Log-Level & Log-Log

Level-Level: a one-unit increase in xx changes yy by β\beta units.

Level-Log: a 1%1\% increase in xx changes yy by approximately β100\frac{\beta}{100} units.

Log-Level: a one-unit increase in xx changes yy by approximately 100β%100\beta\%.

Log-Log: a 1%1\% increase in xx changes yy by approximately β%\beta\%; β\beta is an elasticity.

22
New cards

Quadratic Model

For y=β0+β1x+β2x2+uy = \beta_0 + \beta_1 x + \beta_2 x^2 + u, the effect of xx is not constant. The marginal effect is: ∂y∂x=β1+2β2x\frac{\partial y}{\partial x} = \beta_1 + 2\beta_2 x. Therefore, the effect of an additional unit of xx depends on the current value of xx.

23
New cards

Turning Point

For y=β0+β1x+β2x2+uy = \beta_0 + \beta_1 x + \beta_2 x^2 + u, the turning point occurs where the marginal effect equals zero: x∗=−β12β2x^* = -\frac{\beta_1}{2\beta_2}. If β2<0\beta_2 < 0, the turning point is a maximum. If β2>0\beta_2 > 0, it is a minimum.

24
New cards

OLS Standard Error

The standard error measures the estimated sampling uncertainty of a coefficient estimate. For the slope in the bivariate homoskedastic case: SE(β^1)=σ^2∑(xi−xˉ)2\text{SE}(\hat{\beta}_1) = \sqrt{\frac{\hat{\sigma}^2}{\sum (x_i - \bar{x})^2}}. A smaller standard error means a more precise estimate.

25
New cards

t-Test

A t-test tests a hypothesis about one regression coefficient: t=β^−β0SE(β^)t = \frac{\hat{\beta} - \beta_0}{\text{SE}(\hat{\beta})}. For H0:β=0H_0: \beta = 0: t=β^SE(β^)t = \frac{\hat{\beta}}{\text{SE}(\hat{\beta})}. The t-statistic measures how many standard errors the estimate is from the hypothesized value.

26
New cards

Statistical Significance and Decision Rule

Using the p-value: p<α  ⟹  reject H0p < \alpha \implies \text{reject } H_0; p≥α  ⟹  fail to reject H0p \ge \alpha \implies \text{fail to reject } H_0. For a two-sided critical-value test: ∣t∣>tcritical  ⟹  reject H0|t| > t_{\text{critical}} \implies \text{reject } H_0. Statistical significance means the data provide sufficient evidence against the null at the chosen significance level.

27
New cards

Critical-Value Decision Rule

For a two-sided test, compare ∣t∣|t| with the relevant critical value. If ∣t∣>tcritical|t| > t_{\text{critical}}, reject H0H_0. If ∣t∣≤tcritical|t| \le t_{\text{critical}}, fail to reject H0H_0. For a two-sided 10%10\% test in a large sample, the critical value is approximately 1.641.64.

28
New cards

p-Value

The p-value measures how extreme the observed test statistic would be if H0H_0 were true. For example, p=0.03p = 0.03 means that, assuming H0H_0 is true, there is a 3%3\% probability of observing a test statistic at least as extreme as the one observed. It is not the probability that H0H_0 is true.

29
New cards

One-Sided vs. Two-Sided Test

A two-sided test asks whether the coefficient differs in either direction: H1:β≠β0H_1: \beta \ne \beta_0. A one-sided test asks whether it differs in one specified direction, for example H1:β>β0H_1: \beta > \beta_0 or H1:β<β0H_1: \beta < \beta_0.

30
New cards

Confidence Interval

A confidence interval gives a range of plausible values for the population coefficient. A large-sample 95%95\% interval is approximately: β^±1.96 SE(β^)\hat{\beta} \pm 1.96\,\text{SE}(\hat{\beta}). A two-sided H0:β=β0H_0: \beta = \beta_0 is rejected at the corresponding level if β0\beta_0 lies outside the interval.

31
New cards

Statistical vs. Economic Significance

Statistical significance asks whether there is sufficient statistical evidence that an effect differs from the hypothesized value. Economic significance asks whether the magnitude is economically or practically important. An effect can be statistically significant but economically small, or economically important but imprecisely estimated.

32
New cards

A6 - Normality

The population error uu is independent of the explanatory variables and normally distributed: u∼N(0,σ2)u \sim N(0, \sigma^2). Together with A1-A5, A6 gives a normal sampling distribution for the OLS estimator and allows exact t-inference. A6 is not required for OLS unbiasedness; it is added for statistical inference.

33
New cards

Endogeneity

Endogeneity occurs when an explanatory variable is correlated with the error term: Cov(x,u)≠0\text{Cov}(x,u) \ne 0. Then OLS is generally biased and inconsistent and cannot isolate the causal effect of xx on yy. Main causes include omitted variables, simultaneity/reverse causality, and measurement error in an explanatory variable.

34
New cards

Omitted Variable Bias

OVB occurs when an omitted variable both affects yy and is correlated with an included explanatory variable. The omitted variable enters uu, making the included explanatory variable correlated with uu and violating zero conditional mean.

35
New cards

Direction of Omitted Variable Bias

Suppose the true model is y=β0+β1x+β2z+uy = \beta_0 + \beta_1 x + \beta_2 z + u but zz is omitted. Then: E(β^1)=β1+β2δ~1E(\hat{\beta}_1) = \beta_1 + \beta_2 \tilde{\delta}_1, so: Bias(β^1)=β2δ~1\text{Bias}(\hat{\beta}_1) = \beta_2 \tilde{\delta}_1, where δ~1\tilde{\delta}_1 measures the relationship between zz and xx. Same signs for the effect z→yz \rightarrow y and Corr(x,z)\text{Corr}(x,z) imply upward bias; opposite signs imply downward bias.

36
New cards

Simultaneity / Reverse Causality

Simultaneity occurs when xx affects yy while yy also affects xx. This makes xx correlated with the error term and violates zero conditional mean. OLS is therefore biased and inconsistent and cannot isolate the one-way causal effect of xx on yy.

37
New cards

Measurement Error in x

Suppose observed xx equals true xx plus classical measurement error: xobs=x+ex_{\text{obs}} = x + e. Classical measurement error in an explanatory variable generally biases its estimated coefficient toward zero. This is attenuation bias.

38
New cards

Measurement Error in y

With classical measurement error yobs=y+ey_{\text{obs}} = y + e, where ee is unrelated to the explanatory variables, the OLS slope coefficients remain unbiased. However, measurement error adds noise, increasing uncertainty and reducing precision.

39
New cards

Proxy Variable

A proxy variable is an observed variable used to control for an important unobserved explanatory variable. A useful proxy contains information about the unobserved factor. Including an appropriate proxy can reduce omitted-variable bias.

40
New cards

Heteroskedasticity

Homoskedasticity requires Var(u∣X)=σ2\text{Var}(u \mid X) = \sigma^2. Heteroskedasticity means Var(u∣X)\text{Var}(u \mid X) changes with XX. If A1-A4 hold, OLS coefficients remain unbiased and consistent, but OLS is no longer BLUE and conventional standard errors can be incorrect.

41
New cards

Robust Standard Errors

Heteroskedasticity-robust standard errors allow valid inference when heteroskedasticity is present. They can change standard errors, t-statistics, p-values, and confidence intervals. They do not change the OLS coefficient estimates.

42
New cards

Breusch-Pagan Test

The Breusch-Pagan test tests for heteroskedasticity. Estimate OLS and obtain residuals; regress squared residuals on the explanatory variables; test whether the slopes are jointly zero. H0H_0: homoskedasticity. A small p-value provides evidence of heteroskedasticity.

43
New cards

Irrelevant Variables and Multicollinearity

Including an irrelevant variable does not bias OLS, but can increase standard errors and reduce precision. Multicollinearity means explanatory variables are strongly correlated; it does not bias OLS but increases standard errors. Perfect multicollinearity prevents estimation.

44
New cards

F-Test

An F-test tests multiple coefficient restrictions jointly: F=(SSRR−SSRUR)/qSSRUR/(n−k−1)F = \frac{(\text{SSR}_R - \text{SSR}_{UR})/q}{\text{SSR}_{UR}/(n - k - 1)}, where qq is the number of restrictions. Reject H0H_0 if FF is sufficiently large or p<αp < \alpha. Unlike a t-test, an F-test can test several restrictions jointly.

df1=qdf_1=q and

45
New cards

Dummy Variable and Reference Group

A dummy variable takes values 00 or 11. In y=β0+β1D+uy = \beta_0 + \beta_1 D + u, β0\beta_0 is the expected value for the reference group D=0D = 0, and β1\beta_1 is the difference between D=1D = 1 and D=0D = 0.

46
New cards

Dummy Variable Trap

With an intercept, including a dummy for every category creates perfect multicollinearity because the dummies sum to one. One category must be omitted. The omitted category becomes the reference group, and included dummy coefficients measure differences relative to it.

47
New cards

Dummy Variable with Log y

In ln⁡(y)=β0+β1D+⋯+u\ln(y) = \beta_0 + \beta_1 D + \dots + u, the approximate percentage difference between D=1D = 1 and D=0D = 0 is 100β1%100\beta_1\%. The exact percentage difference is: 100(eβ1−1)%100(e^{\beta_1} - 1)\%.

48
New cards

Interaction Terms

For continuous variables, y=β0+β1x1+β2x2+β3x1x2+uy = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + \beta_3 x_1 x_2 + u gives: ∂y∂x1=β1+β3x2\frac{\partial y}{\partial x_1} = \beta_1 + \beta_3 x_2. For a dummy ×\times continuous interaction, y=β0+β1x+β2D+β3(Dx)+uy = \beta_0 + \beta_1 x + \beta_2 D + \beta_3 (Dx) + u. The slope of xx is β1\beta_1 for D=0D = 0 and β1+β3\beta_1 + \beta_3 for D=1D = 1; β3\beta_3 is the difference in slopes.

49
New cards

Adjusted R-Squared

Adjusted R2R^2 measures model fit while penalizing additional explanatory variables: Rˉ2=1−(1−R2)(n−1)n−k−1\bar{R}^2 = 1 - \frac{(1 - R^2)(n - 1)}{n - k - 1}. Unlike R2R^2, adjusted R2R^2 can decrease when a variable is added if that variable does not improve the model sufficiently.

50
New cards

First Differences

Starting with yit=αi+βxit+uity_{it} = \alpha_i + \beta x_{it} + u_{it}, first differencing gives: Δyit=βΔxit+Δuit\Delta y_{it} = \beta \Delta x_{it} + \Delta u_{it}. The time-invariant effect αi\alpha_i disappears. β\beta is estimated from changes within the same unit over time.

51
New cards

Attenuation Bias under First Differencing

If xobs,it=xit+eitx_{\text{obs}, it} = x_{it} + e_{it}, then: Δxobs,it=Δxit+(eit−ei,t−1)\Delta x_{\text{obs}, it} = \Delta x_{it} + (e_{it} - e_{i, t-1}).

With independent measurement errors: Var(eit−ei,t−1)=2σe2\text{Var}(e_{it} - e_{i, t-1}) = 2\sigma_e^2.

Differencing can increase measurement-error noise relative to true variation in xx, making attenuation bias more severe.

52
New cards

Fixed Effects

Fixed Effects controls for all time-invariant characteristics of each unit. Starting with yit=αi+βxit+uity_{it} = \alpha_i + \beta x_{it} + u_{it}, demeaning gives: yit−yˉi=β(xit−xˉi)+(uit−uˉi)y_{it} - \bar{y}_i = \beta(x_{it} - \bar{x}_i) + (u_{it} - \bar{u}_i). The time-invariant αi\alpha_i disappears. FE estimates β\beta using within-unit variation over time.

53
New cards

Time-Invariant Variables under Fixed Effects

A variable that does not change over time within a unit is eliminated by the Fixed Effects transformation. Therefore, its separate coefficient cannot be estimated with unit fixed effects.

54
New cards

Time Fixed Effects and Two-Way Fixed Effects

Time fixed effects control for shocks affecting all units in the same period. Two-Way Fixed Effects combine unit and time effects: yit=αi+λt+βxit+uity_{it} = \alpha_i + \lambda_t + \beta x_{it} + u_{it}. αi\alpha_i controls permanent unit differences; λt\lambda_t controls common period shocks.

55
New cards

Fixed Effects vs. First Differences

If T=2T = 2, FE and FD produce identical slope estimates. If T>2T > 2, they need not. FE generally preserves more observations in an unbalanced panel and measurement-error bias may decrease as TT grows. FD requires consecutive observations and can make attenuation bias more severe.

56
New cards

Random Effects

Random Effects models the unit-specific effect as random: yit=β0+β1xit+vi+uity_{it} = \beta_0 + \beta_1 x_{it} + v_i + u_{it}. The crucial RE assumption is: Cov(xit,vi)=0\text{Cov}(x_{it}, v_i) = 0 for all periods. Thus, the unit-specific unobserved effect must be uncorrelated with the explanatory variables.

57
New cards

Fixed Effects vs. Random Effects

FE allows the unit-specific effect to be correlated with explanatory variables, but cannot estimate coefficients on time-invariant variables. RE requires no such correlation, can estimate time-invariant effects, and is more efficient if its stronger assumptions hold.

58
New cards

Hausman Test

The Hausman test compares FE and RE estimates. H0H_0: RE is consistent. If p<αp < \alpha, reject H0H_0: evidence suggests the unit-specific effect is correlated with regressors, so FE is preferred. If p≥αp \ge \alpha, fail to reject H0H_0.

59
New cards

Clustered Standard Errors

Clustered standard errors allow errors to be correlated and heteroskedastic within a cluster. Firm clustering allows serial/time correlation within a firm; time clustering allows cross-sectional correlation across firms in the same period. Without appropriate clustering, SEs may be underestimated and t-statistics overstated. Clustering changes inference, not coefficients.

60
New cards

Fixed Effects vs. Clustered Standard Errors

Fixed Effects removes time-invariant unobserved heterogeneity. Clustering adjusts inference for correlation of errors within clusters. They solve different problems, so using FE does not eliminate the need for clustered standard errors when within-cluster error correlation is present.

61
New cards

Instrumental Variables

IV estimates the causal effect of endogenous xx using an instrument zz. A valid instrument must satisfy relevance, Cov(z,x)≠0\text{Cov}(z,x) \ne 0, and exogeneity, Cov(z,u)=0\text{Cov}(z,u) = 0. zz must predict xx but be unrelated to unobserved determinants of yy.

62
New cards

Instrument Relevance and Weak Instruments

Test relevance using the first stage: xi=π0+π1zi+controls+vix_i = \pi_0 + \pi_1 z_i + \text{controls} + v_i. A common rule of thumb is first-stage F>10F > 10 for sufficient instrument strength. A weak instrument provides little exogenous variation in xx, making IV estimates unstable, imprecise and potentially unreliable.

63
New cards

Instrument Exogeneity

Instrument exogeneity requires: Cov(z,u)=0\text{Cov}(z,u) = 0. The instrument must be unrelated to unobserved determinants of yy. Unlike relevance, exogeneity generally cannot be tested directly because uu is unobserved; it requires a credible economic or institutional argument.

64
New cards

Two-Stage Least Squares (2SLS)

Stage 1: regress endogenous xx on instrument zz and all exogenous controls to obtain predicted x^\hat{x}. Stage 2: regress yy on x^\hat{x} and the exogenous controls. The coefficient uses variation in xx generated by the instrument to estimate the causal effect.

65
New cards

OLS vs. IV

OLS uses all variation in xx, including variation that may be correlated with uu. IV uses variation generated by a valid instrument. If xx is endogenous and zz is valid, OLS is inconsistent while IV is consistent but generally less precise. With one xx and one zz: β^IV=Cov(z,y)Cov(z,x)\hat{\beta}_{IV} = \frac{\text{Cov}(z,y)}{\text{Cov}(z,x)}.

66
New cards

Difference-in-Differences: Basic Idea

DiD compares the change in the treated group with the change in the control group: DiD=(yˉT,post−yˉT,pre)−(yˉC,post−yˉC,pre)\text{DiD} = (\bar{y}_{T,\text{post}} - \bar{y}_{T,\text{pre}}) - (\bar{y}_{C,\text{post}} - \bar{y}_{C,\text{pre}}). The control group's change captures the common time effect, which is subtracted from the treated group's change.

67
New cards

Difference-in-Differences Regression

Traditional DiD: yit=β0+β1Treati+β2Postt+β3(Treati×Postt)+uity_{it} = \beta_0 + \beta_1 \text{Treat}_i + \beta_2 \text{Post}_t + \beta_3 (\text{Treat}_i \times \text{Post}_t) + u_{it}, where β3\beta_3 is the DiD treatment effect. With unit and time FE: yit=αi+λt+βDit+uity_{it} = \alpha_i + \lambda_t + \beta D_{it} + u_{it}, where Dit=Treati×PosttD_{it} = \text{Treat}_i \times \text{Post}_t. Unit FE absorb Treat; time FE absorb Post. β\beta is the DiD effect, assuming parallel trends.

68
New cards

Parallel Trends Assumption

Without treatment, treated and control groups would have experienced the same average change in the outcome over time. This is the key identifying assumption of DiD. The post-treatment counterfactual cannot be directly verified, but similar pre-treatment trends can provide supporting evidence.

69
New cards

Regression Discontinuity Design

RDD exploits a treatment rule based on whether a running variable rr crosses a cutoff cc. Observations just below and above the cutoff are compared. The key assumption is continuity of potential outcomes at the cutoff; therefore a discontinuous jump can be attributed to treatment. RDD identifies a local effect around the cutoff.

70
New cards

RDD Regression

A sharp RDD can be written: yi=β0+β1Di+β2(ri−c)+β3[Di(ri−c)]+uiy_i = \beta_0 + \beta_1 D_i + \beta_2 (r_i - c) + \beta_3 [D_i(r_i - c)] + u_i

with Di=1(ri≥c)D_i = \mathbf{1}(r_i \ge c).

β1\beta_1 is the treatment effect at the cutoff; β2\beta_2 is the slope below; β2+β3\beta_2 + \beta_3 is the slope above; β3\beta_3 is the slope difference. Identification requires continuity of potential outcomes at cc.

71
New cards

Binary Dependent Variable

A binary dependent variable takes two possible values: yi∈{0,1}y_i \in \{0,1\}. For binary yy: E(yi∣Xi)=P(yi=1∣Xi)E(y_i \mid X_i) = P(y_i = 1 \mid X_i). Therefore, the conditional mean of yy can be interpreted as a probability.

72
New cards

Linear Probability Model

The LPM models P(yi=1∣Xi)=β0+β1x1i+⋯+βkxkiP(y_i = 1 \mid X_i) = \beta_0 + \beta_1 x_{1i} + \dots + \beta_k x_{ki} and is estimated using OLS. βj\beta_j is the change in the probability that y=1y = 1 from a one-unit increase in xjx_j, holding other variables constant. β=0.08\beta = 0.08 means an 88 percentage-point increase.

73
New cards

Advantages and Problems of the LPM

Advantages: easy to estimate with OLS and coefficients have a direct probability interpretation. Problems: predicted probabilities can fall below 00 or above 11; marginal effects are constant; the error is inherently heteroskedastic. Therefore, robust standard errors should be used.

74
New cards

Logit vs. Probit

Let z=β0+Xβz = \beta_0 + X\beta. Logit: P(y=1∣X)=ez1+ezP(y=1 \mid X) = \frac{e^z}{1 + e^z}. Probit: P(y=1∣X)=Φ(z)P(y=1 \mid X) = \Phi(z). Logit uses the logistic CDF; Probit uses the standard normal CDF. Both produce S-shaped probability curves and probabilities between 00 and 11.

75
New cards

Maximum Likelihood Estimation

Logit and Probit are nonlinear probability models estimated using Maximum Likelihood Estimation rather than OLS. MLE chooses parameter values that make the observed sample outcomes y=0y=0 and y=1y=1 as likely as possible under the model.

76
New cards

Logit/Probit Coefficient Interpretation

The sign of β\beta gives the direction of the effect:

β>0\beta > 0 means xx increases P(y=1)P(y=1);

β<0\beta < 0 means xx decreases P(y=1)P(y=1).

But β\beta itself is not the change in probability. For example, β=0.20\beta = 0.20 does not mean probability rises by 2020 percentage points. Use marginal effects or predicted probabilities for magnitude.

77
New cards

Marginal Effects in Logit/Probit

A marginal effect converts a Logit/Probit coefficient into an effect on predicted probability. Because these models are nonlinear, the marginal effect generally depends on XX. The Average Marginal Effect calculates the marginal effect for each observation and averages them. AME=0.06\text{AME} = 0.06 means a 66 percentage-point increase on average.

78
New cards

Marginal Effect of a Dummy Variable

For a dummy DD, calculate the discrete probability change: P(y=1∣D=1,X)−P(y=1∣D=0,X)P(y=1 \mid D=1, X) - P(y=1 \mid D=0, X). For an Average Marginal Effect, calculate this probability difference for every observation and then take the average.

79
New cards

y^=25.4+3.20 experience\hat{y} = 25.4 + 3.20\,\text{experience}. Interpret experience.

A one-unit increase in experience is associated with a 3.203.20-unit increase in predicted salary, holding other included variables constant. State: (1) one-unit change in xx, (2) β\beta-unit change in yy, and (3) ceteris paribus.

80
New cards

y^=30.2+8.50ln⁡(sales)\hat{y} = 30.2 + 8.50\ln(\text{sales}). Interpret sales.

A 1%1\% increase in sales is associated with an 8.50100=0.085\frac{8.50}{100} = 0.085 unit increase in predicted salary, holding other variables constant. Rule: in a level-log model, a 1%1\% increase in xx implies a β100\frac{\beta}{100} units change in yy.

81
New cards

ln⁡(y^)=6.20+0.035 experience\ln(\hat{y}) = 6.20 + 0.035\,\text{experience}. Interpret experience.

A one-unit increase in experience is associated with an approximately 100(0.035)=3.5%100(0.035) = 3.5\% increase in predicted salary, holding other variables constant. Rule: in a log-level model, a one-unit increase in xx implies an approximately 100β%100\beta\% change in yy.

82
New cards

ln⁡(y^)=5.80+0.24ln⁡(sales)\ln(\hat{y}) = 5.80 + 0.24\ln(\text{sales}). Interpret sales.

A 1%1\% increase in sales is associated with an approximately 0.24%0.24\% increase in predicted salary, holding other variables constant. In a log-log model, β\beta is an elasticity.

83
New cards

wage^=12.5+2.40 female+0.80 education\widehat{\text{wage}} = 12.5 + 2.40\,\text{female} + 0.80\,\text{education}. Interpret female.

Holding education constant, the group with female=1\text{female} = 1 has predicted wage 2.402.40 units higher than the reference group female=0\text{female} = 0. Identify the D=1D=1 group, the D=0D=0 reference group, and the β\beta-unit difference.

84
New cards

ln⁡(wage^)=2.30−0.15 female+0.08 education\ln(\widehat{\text{wage}}) = 2.30 - 0.15\,\text{female} + 0.08\,\text{education}. Interpret female.

Approximation: 100(−0.15)=−15%100(-0.15) = -15\%, so female=1\text{female} = 1 has approximately 15%15\% lower predicted wage than female=0\text{female} = 0, holding education constant. Exact effect: 100(e−0.15−1)≈−13.9%100(e^{-0.15} - 1) \approx -13.9\%.

85
New cards

salary=30+4 experience−0.10 experience2\text{salary} = 30 + 4\,\text{experience} - 0.10\,\text{experience}^2. Effect at experience=10\text{experience} = 10?

Do not interpret 44 alone. Compute the marginal effect: ∂salary∂experience=4−2(0.10) experience\frac{\partial \text{salary}}{\partial \text{experience}} = 4 - 2(0.10)\,\text{experience}. At experience=10\text{experience} = 10: 4−0.20(10)=24 - 0.20(10) = 2. One additional year increases predicted salary by approximately 22 units at experience=10\text{experience} = 10.

86
New cards

β^=0.30\hat{\beta} = 0.30, SE(β^)=0.12\text{SE}(\hat{\beta}) = 0.12, p=0.012p = 0.012. Two-sided H0:β=0H_0: \beta = 0 at 5%5\%.

Calculate t=0.300.12=2.50t = \frac{0.30}{0.12} = 2.50. Critical-value method: ∣2.50∣>1.96  ⟹  reject H0|2.50| > 1.96 \implies \text{reject } H_0. p-value method: 0.012<0.05  ⟹  reject H00.012 < 0.05 \implies \text{reject } H_0. Both methods give the same decision. Conclusion: β\beta is statistically significantly different from zero at the 5%5\% level.

87
New cards

β^=0.30\hat{\beta} = 0.30, SE(β^)=0.20\text{SE}(\hat{\beta}) = 0.20. Test H0:β≤0H_0: \beta \le 0 vs H1:β>0H_1: \beta > 0 at 5%5\%.

Calculate t=0.300.20=1.50t = \frac{0.30}{0.20} = 1.50. For a one-sided 5%5\% test, tcritical≈1.645t_{\text{critical}} \approx 1.645. Since 1.50<1.6451.50 < 1.645, fail to reject H0H_0. There is insufficient evidence that β\beta is positive at the 5%5\% level.

88
New cards

mortgage=20+0.50 income\text{mortgage} = 20 + 0.50\,\text{income}; SE(β^income)=0.20\text{SE}(\hat{\beta}_{\text{income}}) = 0.20. Significant at 10%10\%? Interpret.

t=0.500.20=2.50t = \frac{0.50}{0.20} = 2.50. For a two-sided 10%10\% test, ∣2.50∣>1.64|2.50| > 1.64, so income is statistically significant. A one-unit increase in income is associated with a 0.500.50-unit increase in predicted mortgage, holding other variables constant. Separate statistical significance from economic interpretation.

89
New cards

H0:β1=β2=β3=0H_0: \beta_1 = \beta_2 = \beta_3 = 0. Output: F=4.80F = 4.80, p=0.003p = 0.003. Test at 5%5\%.

Since 0.003<0.050.003 < 0.05, reject H0H_0. β1\beta_1, β2\beta_2, and β3\beta_3 are jointly statistically significant at the 5%5\% level. Do not conclude that every coefficient is individually significant.

90
New cards

SSRR=500\text{SSR}_R = 500, SSRUR=400\text{SSR}_{UR} = 400, n=100n = 100, k=4k = 4, q=2q = 2. Calculate FF.

Use F=(SSRR−SSRUR)/qSSRUR/(n−k−1)F = \frac{(\text{SSR}_R - \text{SSR}_{UR})/q}{\text{SSR}_{UR}/(n - k - 1)}. Thus F=(500−400)/2400/(100−4−1)=504.21≈11.88F = \frac{(500 - 400)/2}{400/(100 - 4 - 1)} = \frac{50}{4.21} \approx 11.88. Then compare FF with the appropriate critical value or use its p-value.

91
New cards

Regression output reports R2=0.64R^2 = 0.64. Interpret it.

64%64\% of the sample variation in the dependent variable is explained by the explanatory variables included in the regression. Do not say that 64%64\% of yy is explained, that the model is 64%64\% correct, or that 64%64\% of the relationship is causal.

92
New cards

employed=0.25+0.06 education−0.03 age\text{employed} = 0.25 + 0.06\,\text{education} - 0.03\,\text{age}. Interpret education.

Because this is an LPM, a one-unit increase in education is associated with a 0.06=60.06 = 6 percentage-point increase in the predicted probability of being employed, holding age constant. Use percentage points, not percent.

93
New cards

Logit model: β^education=0.40\hat{\beta}_{\text{education}} = 0.40. Interpret the effect on P(employed=1)P(\text{employed} = 1).

Because 0.40>00.40 > 0, education has a positive effect on the predicted probability of employment. But 0.400.40 is not a 4040 percentage-point increase. A Logit coefficient itself is not a probability change; use a marginal effect for magnitude.

94
New cards

Logit model reports AMEeducation=0.045\text{AME}_{\text{education}} = 0.045. Interpret it.

A one-unit increase in education increases the predicted probability that y=1y = 1 by approximately 0.045=4.50.045 = 4.5 percentage points on average, holding other variables constant. Say percentage points, not percent.

95
New cards

P(y=1∣D=1)=0.62P(y = 1 \mid D = 1) = 0.62 and P(y=1∣D=0)=0.48P(y = 1 \mid D = 0) = 0.48. Interpret DD.

Calculate 0.62−0.48=0.140.62 - 0.48 = 0.14. Changing DD from 00 to 11 increases the predicted probability that y=1y = 1 by 1414 percentage points.

96
New cards

True model includes wealth; wealth raises yy and is positively correlated with income. Wealth is omitted. Bias?

Both relevant signs are positive: wealth→y\text{wealth} \rightarrow y is positive and Corr(income,wealth)\text{Corr}(\text{income}, \text{wealth}) is positive. Therefore, (+)×(+)=(+)(+) \times (+) = (+), so the estimated income coefficient is biased upward. In an OVB answer, identify both relationships before stating the bias direction.

97
New cards

performance=β0+β1 investment+u\text{performance} = \beta_0 + \beta_1\,\text{investment} + u; expected performance also raises investment. Causal β1\beta_1?

No. Expected performance affects investment, creating reverse causality and Cov(investment,u)≠0\text{Cov}(\text{investment}, u) \ne 0. Investment is endogenous, so OLS is generally biased and inconsistent. A valid IV or another credible identification strategy is needed for causal interpretation.

98
New cards

xx is endogenous and zz is proposed as an instrument. What must zz satisfy?

Relevance: Cov(z,x)≠0\text{Cov}(z, x) \ne 0, so zz must predict endogenous xx. Exogeneity: Cov(z,u)=0\text{Cov}(z, u) = 0, so zz must be unrelated to unobserved determinants of yy. Both conditions are required for a valid instrument.

99
New cards

Treated: 50→7050 \rightarrow 70. Control: 40→4840 \rightarrow 48. Calculate the DiD effect.

Treated change = 70−50=2070 - 50 = 20. Control change = 48−40=848 - 40 = 8. DiD=20−8=12\text{DiD} = 20 - 8 = 12. The estimated treatment effect is 1212 units, assuming the parallel-trends assumption holds.

100
New cards

yit=αi+βxit+uity_{it} = \alpha_i + \beta x_{it} + u_{it}. What variation identifies β\beta?

β\beta is identified from changes in xx within the same unit over time. The unit fixed effect αi\alpha_i absorbs all observed and unobserved characteristics that are constant over time within that unit. β\beta is not identified from permanent differences between units.