Empirical Methods in Finance Flashcards (Weeks 1-7)

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/66

flashcard set

Earn XP

Description and Tags

A set of 114 vocabulary flashcards covering key statistical, mathematical, econometric, and causal inference terms from Weeks 1-7 of Empirical Methods in Finance.

Last updated 9:13 AM on 10/6/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

67 Terms

1
New cards

Variance and Standard Deviation

Variance measures how dispersed X is around its mean:

Var(X)=E[(X−E(X))2]Var\left(X\right)=E\left\lbrack\left(X-E\left(X\right)\right)^2\right\rbrack

Standard deviation is:

sd(x)=Var(X)sd\left(x\right)=\sqrt{Var\left(X\right)}

A larger variance or standard deviation means observations are generally further from the mean. Standard deviation is expressed in the same units as X

2
New cards

Covariance and correlation

Covariance measures the linear co-movement of two variables:

Cov(X,Y)=E[(X−E(X))(Y−E(Y))]Cov\left(X,Y\right)=E\left\lbrack\left(X-E\left(X\right)\right)\left(Y-E\left(Y\right)\right)\right\rbrack

Cov>0: they tend to move together.
Cov<0: they tend to move in opposite directions.
Cov=0: no linear relationship.

Correlation standardizes covariance:

ρ(X,Y)=Cov(X,Y)sd(X)sd(Y)\rho\left(X,Y\right)=\frac{Cov\left(X,Y\right)}{sd\left(X\right)sd\left(Y\right)}

and always lies between -1 and 1. Unlike covariance, correlation does not depend on the units of measurement.

3
New cards

Independence vs zero covariance

Independence means knowing X provides no information about Y:

P(X=x,Y=y)=P(X=x)P(Y=y)P\left(X=x,Y=y\right)=P\left(X=x\right)P\left(Y=y\right)

If X and Y are independent:

Cov(X,Y)=0Cov\left(X,Y\right)=0

But the reverse is not generally true:

Cov(X,Y)=0≠independenceCov\left(X,Y\right)=0\ne independence

because covariance only measures linear dependence. Two variables can have zero covariance but still have a nonlinear relationship

4
New cards

Conditional probability and dependence

Conditional probability asks how likely Y is after knowing X:

P(Y=y∣X=x)P\left(Y=y\left|X=x\right.\right)

If X and Y are independent:

P(Y=y∣X=x)=P(Y=y)P\left(Y=y\left|X=x\right.\right)=P\left(Y=y\right)

Knowing X does not change the probability of Y. If the conditional probability changes with X, the variables are dependent.

5
New cards

Mean vs Median

The mean is the average and is sensitive to extreme values. The median is the middle observation after ordering the data and is less sensitive to extremes.

Therefore, an extreme outlier can substantially change the mean while leaving the median unchanged

6
New cards

Regression model and error term

In

yi=α+βxi+uiy_{i}=\alpha+\beta x_{i}+u_{i}

y is the dependent variable, x the explanatory variable, and u contains all other factors affecting y that are not included in the model.

alpha is the intercept and beta measures how y changes when x increases by one unit.

7
New cards

Population model vs estimated regression

The population model contains the unknown true parameters:

yi=α+βxi+uiy_{i}=\alpha+\beta x_{i}+u_{i}

The estimated regression uses sample estimates:

yihat=αhat+βhatxiy_{i}^{hat}=\alpha^{hat}+\beta^{hat}x_{i}

betahat estimates the unknown population effect beta. Different random samples generally produce different estimates

8
New cards

OLS residuals

OLS chooses alphahat and betahat to minimize the sum of squared residuals:

∑uihat2\sum u_{i}^{hat2}

where

uihat=yi−yihatu_{i}^{hat}=y_{i}-y_{i}^{hat}

A residual is therefore the difference between the observed and predicted value of y.

9
New cards

Unbiasedness

An estimator is unbiased if its expected value across repeated random samples equals the true population parameter:

E(βhat)=βE\left(\beta^{hat}\right)=\beta

This does not mean every estimate equals beta. Individual estimates can be too high or too low, but there is no systematic error on average.

10
New cards

Efficiency and BLUE

BLUE means Best Linear Unbiased Estimator.

Linear: the estimator is linear in y.
Unbiased: E(βhat)=βE\left(\beta^{hat}\right)=\beta .
Best: it has the smallest variance among linear unbiased estimators.

Thus, “best” means most efficient within this class, not that every individual estimate is closest to the true value.

11
New cards

CLRM Assumptions

[A1] Linear in parameters.
[A2] Random sample from the population.
[A3] Variation in x.
[A4] E(ui∣xi)=0E\left(u_{i}\vert x_{i}\right)=0 .
[A5] Var(ui∣xi)=σ2Var\left(u_{i}\vert x_{i}\right)=\sigma^2.

[A1]–[A4] imply unbiased OLS.

Adding [A5] gives efficiency and therefore BLUE

12
New cards

Zero Conditional Mean

The crucial assumption for OLS unbiasedness is:

E(ui∣xi)=0E\left(u_{i}\vert x_{i}\right)=0

It means that the average unobserved factors in ui do not systematically vary with xi.

If an omitted factor affects y and is related to x, A4 is violated and OLS will generally be biased.

13
New cards

R2

R2 measures the proportion of sample variation in y explained by the regression:

R2=1−RSSTSS=ESSTSSR^2=1-\frac{RSS}{TSS}=\frac{ESS}{TSS}

A higher R2 means better in-sample fit.

It does not imply that the model is causal, unbiased, or correctly specified

14
New cards

Standard Error of OLS

The standard error measures the uncertainty of an estimated coefficient: a smaller standard error means a more precise estimate.

For the slope, precision improves when there is more variation in x, a larger sample, or less unexplained variation in y

15
New cards

Standardized coefficients

A standardized coefficient is:

Bσxσy\frac{B\sigma_{x}}{\sigma_{y}}

It measures how many standard deviations predicted y changes when x increases by one standard deviation.

16
New cards

Functional form interpretations

Dependent variable

Independent variable

Interpretation of β

Level

Level

1-unit increase in x → β-unit change in y

Level

Log

1% increase in x → β/100-unit change in y

Log

Level

1-unit increase in x → approximately 100β% change in y

Log

Log

1% increase in x → approximately β% change in y


17
New cards

OLS Standard Error formula

The standard error measures the uncertainty of an estimated coefficient. It estimates how much the coefficient estimate would vary across repeated samples.

Smaller SE → more precise estimate.

In the bivariate homoskedastic case:

Var(βhat)=σ2∑(xi−xav)2Var\left(\beta^{hat}\right)=\frac{\sigma^2}{\sum\left(x_{i}-x^{av}\right)^2}

Therefore, precision increases when there is more variation in x and decreases when there is more unexplained variation in y

18
New cards

t-Test

A t-test tests a hypothesis about one regression coefficient.

t=βhat−β0SE(βhat)t=\frac{\beta^{hat}-\beta_0}{SE\left(\beta^{hat}\right)}

For the common null hypothesis H0: beta = 0:

t=βhatSE(βhat)t=\frac{\beta^{hat}}{SE\left(\beta^{hat}\right)}

The t-statistic tells us how many standard errors the estimate is away from the value specified under H0

19
New cards

Statistical significance and decision rule

The significance level \(\alpha\) is chosen before the test, commonly 10%, 5%, or 1%.

Reject H0 when:

p<αp<\alpha

or equivalently, for a two-sided test:

|t| > tcritical

If p≥αp\ge\alpha , fail to reject H0. This means there is insufficient evidence against H0; it does not prove that H0 is true.

20
New cards

p-value

The p-value measures how extreme the observed result would be if H0 were true.

A smaller p-value → stronger evidence against H0.

Example: p=0.03 means that, assuming H0 is true, there is a 3% probability of obtaining a result at least as extreme as the observed result

21
New cards

One-sided vs Two-sided test

A two-sided test asks whether the coefficient differs from the hypothesized value in either direction.

A one-sided test asks whether the coefficient differs in one specified direction.

For the same t-statistic in the predicted direction, the one-sided p-value is half the two-sided p-value.

22
New cards

Confidence interval

A confidence interval gives a range of plausible values for the population coefficient:

βhat±tcriticalSE(βhat)\beta^{hat}\pm t_{critical}SE\left(\beta^{hat}\right)

For a large-sample 95% confidence interval:

βhat±1,96SE(βhat)\beta^{hat}\pm1,96SE\left(\beta^{hat}\right)

A two-sided H0: beta = beta0 is rejected at the corresponding significance level if beta0 lies outside the confidence interval.

23
New cards

Statistical vs economic significance

Statistical significance asks whether there is sufficient statistical evidence that an effect differs from the hypothesized value.

Economic significance asks whether the size of the estimated effect is economically meaningful.

An effect can therefore be statistically significant but economically small, or economically large but estimated too imprecisely to be statistically significant.

24
New cards

A6 Normality

The population error u is independent of the explanatory variables x and normally distributed:

u−N(0,σ2)u-N\left(0,\sigma^2\right)

A6 implies that the OLS estimator is normally distributed, which allows us to perform exact hypothesis testing using the t-distribution.

A6 is not required for OLS to be unbiased; it is added for statistical inference

25
New cards

Failure of A4 / Endogeneity

A4 requires the explanatory variables to be unrelated to the expected error term.

If an explanatory variable is correlated with u, it is endogenous and OLS is generally biased and inconsistent.

Main causes:

  • omitted variables

  • simultaneity / reverse causality

  • measurement error in an explanatory variable


26
New cards

Omitted variable bias

Omitted variable bias occurs when an omitted variable:

  1. affects y, and

  2. is correlated with an included explanatory variable.

The direction of the bias depends on:

  • the effect of the omitted variable on y;

  • its correlation with the included variable.

Same signs → positive bias.
Opposite signs → negative bias.

27
New cards

Simultaneity / reverse causality

Simultaneity occurs when x affects y while y also affects x.

This makes x correlated with the error term, violating A4. OLS is therefore biased and inconsistent and cannot isolate the one-way causal effect of x on y.

28
New cards

Measurement error

Measurement error in x → x becomes correlated with the regression error. Under classical measurement error, the estimated coefficient is typically biased toward zero. This is attenuation bias.

Measurement error in y → does not bias the OLS coefficients if it is unrelated to the explanatory variables, but increases noise and reduces precision.

29
New cards

Failure of A5 / Heteroskedasticity

A5 requires constant error variance:

Var(u∣X)=σ2Var\left(u\vert X\right)=\sigma^2

Heteroskedasticity means the variance of u changes with the explanatory variables.

If A1–A4 still hold:

  • OLS coefficients remain unbiased/consistent;

  • OLS is no longer BLUE;

  • conventional standard errors can be incorrect.


30
New cards

Robust SE

Robust standard errors allow valid inference when heteroskedasticity is present.

They change:

  • standard errors

  • t-statistics

  • p-values

  • confidence intervals

They do not change the OLS coefficient estimates.

31
New cards

Breusch-Pagan test

Tests for heteroskedasticity.

  1. Estimate OLS and obtain residuals.

  2. Regress the squared residuals on the explanatory variables.

  3. Jointly test whether the slope coefficients in this auxiliary regression equal zero.

H₀: homoskedasticity.

p < α → reject H₀ → evidence of heteroskedasticity.

32
New cards

Irrelevant variables and multicollinearity

Including an irrelevant variable does not bias OLS, but can increase standard errors and reduce precision.

Strong but imperfect multicollinearity also does not bias OLS, but increases standard errors and makes individual effects harder to estimate precisely.

33
New cards

F-test

An F-test tests multiple restrictions jointly.

The unrestricted model does not impose H₀; the restricted model does.

F=(RSSR−RSSUR)qRSSUR(n−k−1)F=\frac{\frac{\left(RSS_{R}-RSS_{UR}\right)}{q}}{\frac{RSS_{UR}}{\left(n-k-1\right)}}

where q is the number of restrictions.

Reject H₀ if F > Fcritical or p< alpha.

Unlike a t-test, which tests one coefficient, an F-test can test several coefficients jointly.

34
New cards

Dummy variables and reference group

A dummy variable takes the value 0 or 1.

Its coefficient measures the difference between the group with D = 1 and the reference group with D = 0, holding other explanatory variables constant.

With an intercept and multiple categories, omit one category as the reference group. Including every category creates perfect multicollinearity: the dummy variable trap

35
New cards

Dummy variable with log y

When y is in logs, a dummy coefficient δ represents approximately a:

100δ%\delta\%

difference between D = 1 and D = 0.

The exact percentage difference is:

100[eδ−1]%100\left\lbrack e^{\delta}-1\right\rbrack\%

36
New cards

Interaction terms

An interaction means that the effect of one explanatory variable depends on another variable.

Continuous × continuous: the effect of x₁ changes with the value of x₂.

Dummy × continuous: the interaction coefficient measures how the slope of x differs between D = 1 and D = 0.

So an interaction coefficient should not normally be interpreted as an independent direct effect

37
New cards

Unit fixed effects with dummies

Unit fixed effects can be implemented by including a dummy for each unit except one reference unit.

They control for all differences between units that are constant over time.

38
New cards

DiD basic idea

DID=(Treatedpost−Treatedpre)−(Controlpost−Controlpre)DID=\left(Treated_{post}-Treated_{pre}\right)-\left(Control_{post}-Control_{pre}\right)

DiD compares the change over time in the treated group with the change over time in the control group.

It subtracts the common time change to isolate the treatment-associated change.

39
New cards

Clustered SE

Clustered standard errors allow regression errors to be correlated within groups, such as observations from the same firm over time.

Ignoring this correlation can make conventional standard errors too small and statistical significance too strong.

Clustering changes standard errors, t-statistics, p-values and confidence intervals, but not the OLS coefficient estimates

40
New cards

Adjusted R2

Adjusted R² measures model fit while penalizing the addition of explanatory variables.

1−RSS(n−k−1)TSS(n−1)1-\frac{\frac{RSS}{\left(n-k-1\right)}}{\frac{TSS}{\left(n-1\right)}}

where:

  • n = number of observations

  • k = number of explanatory variables, excluding the intercept

Unlike R², which can never decrease when a variable is added, adjusted R² can decrease if the new variable does not improve the model sufficiently.

Therefore, adjusted R² is more useful for comparing models with different numbers of explanatory variables.

41
New cards

First Differences

First Differencing removes time-invariant unobserved effects by subtracting the previous period from the current period.

Starting with:

yit=αi+βxit+uity_{it}=\alpha_{i}+\beta x_{it}+u_{it}

we obtain:

Δyit=βΔxit+Δuit\Delta y_{it}=\beta\Delta x_{it}+\Delta u_{it}

because the time-invariant effect αi\alpha_{i} disappears.

β is therefore estimated using changes within the same unit over time.

42
New cards

Attenuation bias and FD

Measurement error in x biases the estimated coefficient toward zero. This is attenuation bias.

With First Differencing:

Δxitobs=Δxit+(vit−vi,t−1)\Delta x_{it}^{obs}=\Delta x_{it}+\left(v_{it}-v_{i,t-1}\right)

and, with independent measurement errors:

Var(vit−vi,t−1)=2σv2Var\left(v_{it}-v_{i,t-1}\right)=2\sigma_{v}^2

Therefore, differencing can increase measurement-error noise relative to the true variation in x, making attenuation bias more severe

43
New cards

Fixed Effects

Fixed Effects controls for all time-invariant characteristics of each unit, including unobserved characteristics that may be correlated with x.

Starting with:

yit=αi+βxit+uity_{it}=\alpha_{i}+\beta x_{it}+u_{it}

FE demeans the variables by subtracting each unit's time average:

yit−yiav=β(xit−xiav)+(uit−uiav)y_{it}-y_{i}^{av}=\beta\left(x_{it}-x_{i}^{av}\right)+\left(u_{it}-u_{i}^{av}\right)

The time-invariant αi\alpha_{i} disappears.

Therefore, FE estimates β from variation within the same unit over time.

44
New cards

Time Fixed Effects and Two-Way Fixed Effects

Time Fixed Effects control for shocks that affect all units in the same period, such as a financial crisis.

Two-Way Fixed Effects combine:

Unit FE → control for permanent differences between units.

Time FE → control for common shocks in each period.

Thus, Two-Way FE controls for both time-invariant unit characteristics and common time effects

45
New cards

FE vs FD

If T = 2: FE and FD produce identical slope estimates.

If T > 2: they do not necessarily produce identical estimates.

FE:

  • usually preserves more observations in an unbalanced panel;

  • measurement-error bias may decrease as T increases.

FD:

  • requires consecutive observations;

  • can make attenuation bias from measurement error more severe


46
New cards

Random Effects

Random Effects models the unit-specific effect as a random component:

yit=α+βxit+ϵi+uity_{it}=\alpha_{}+\beta x_{it}+\epsilon_{i}+u_{it}

The crucial RE assumption is:

Cov(xit,ϵi)=0Cov\left(x_{it},\epsilon_{i}\right)=0

for all periods.

Unlike FE, RE therefore requires the unit-specific unobserved effect to be uncorrelated with the explanatory variables

47
New cards

FE vs RE

FE:

  • allows the unit-specific effect to be correlated with x;

  • cannot estimate effects of time-invariant variables;

  • is more robust but less efficient.

RE:

  • requires the unit-specific effect to be uncorrelated with x;

  • can estimate time-invariant variables;

  • is more efficient if the RE assumptions hold.

Use RE only when its stronger independence assumption is credible.

48
New cards

Hausman test

The Hausman test compares FE and RE estimates to assess whether the RE assumption is appropriate.

H₀: RE is consistent.

If p<alpha → reject H₀ → evidence that the unit effect is correlated with the regressors → prefer FE.

If p>alpha → fail to reject H₀ → no evidence against RE

49
New cards

FE vs Clustered SE

FE and clustering solve different problems.

Fixed Effects remove time-invariant unobserved heterogeneity.

Clustering adjusts inference when errors are correlated within units or groups.

Using FE does not eliminate the need for clustered standard errors when within-cluster error correlation is present.

50
New cards

Clustering

Clustering allows regression errors to be correlated within the same group, such as observations of the same firm over time.

Without clustering, correlated observations are treated as providing more independent information than they actually do, which can make standard errors too small and statistical significance too strong.

Clustered standard errors:

  • allow correlation and heteroskedasticity within clusters;

  • assume independence across clusters;

  • change SEs, t-statistics, p-values and confidence intervals;

  • do not change the coefficient estimates.

Cluster at the level where errors are likely to be correlated, e.g. by firm for within-firm dependence

51
New cards

IV

IV is used to estimate the causal effect of endogenous x using an instrument z.

A valid instrument must satisfy:

Relevance:

Cov(z,x)≠0Cov\left(z,x\right)\ne0

z must predict x.

Exogeneity:

Cov(z,u)=0Cov\left(z,u\right)=0

z must be unrelated to unobserved determinants of y.

Both conditions are required for a valid instrument.

52
New cards

IV and weak instrument

Test relevance in the first-stage regression:

xi=δ0+δ1zi+controls+vix_{i}=\delta_0+\delta_1z_{i}+controls+v_{i}

The instrument is relevant when z predicts x.

Use the first-stage F-statistic:

F > 10

is the lecture's rule of thumb for sufficient instrument strength.

A weak instrument provides little useful variation in x and makes IV estimates unreliable and imprecise

53
New cards

Instrument exogeneity

Exogeneity requires:

Cov(z,u)=0Cov\left(z,u\right)=0

The instrument must be unrelated to unobserved determinants of y.

Unlike relevance, exogeneity generally cannot be tested directly because u is unobserved. It therefore requires a credible economic or institutional argument

54
New cards

2SLS

2SLS estimates an IV model in two stages.

Stage 1: regress endogenous x on instrument z and controls → obtain predicted xhat.

Stage 2: regress y on predicted xhat and the same controls.

The coefficient on xhat uses only the variation in x generated by the instrument to estimate the causal effect

55
New cards

OSL vs IV

OLS uses all variation in x, including variation that may be correlated with the error term.

IV uses only the variation in x generated by a valid instrument z.

If x is endogenous but z is relevant and exogenous:

OLS → biased/inconsistent.

IV → consistent, but generally less precise.

With one x and one instrument:

βIVhat=Cov(z,y)Cov(z,x)\beta_{IV}^{hat}=\frac{Cov\left(z,y\right)}{Cov\left(z,x\right)}


56
New cards

DiD regression

yit=β0+β1(Postt⋅Treati)+β2Treati+β3Posti+uity_{it}=\beta_0+\beta_1\left(Post_{t}\cdot Treat_{i}\right)+\beta_2Treat_{i}+\beta_3Post_{i}+u_{it}

β₁ = DiD treatment effect.

β₂ = pre-treatment difference between treated and control groups.

β₃ = time change for the control group.

The interaction \(Post\times Treat\) therefore identifies the DiD effect

57
New cards

RDD

RDD exploits a treatment rule based on a cutoff.

Observations just below and just above the cutoff are compared. A discontinuous jump in y at the cutoff identifies a local treatment effect.

Key assumption: without treatment, the expected outcome would change continuously at the cutoff.

Therefore, a jump at the cutoff can be attributed to the treatment.

58
New cards

Parallel trend assumption

Without treatment, the treated and control groups would have experienced the same average change in the outcome over time.

Formally, the lecture states the assumption for a consistent DiD estimator as:

E(u∣Post,Treat)=0E\left(u\vert Post,Treat\right)=0

Parallel trends cannot be tested directly because we cannot observe what would have happened to the treated group without treatment. Similar pre-treatment trends provide supporting evidence.

59
New cards

RDD regression

yi=β0+β1Di+β2(ri−c)+β3Di(ri−c)+uiy_{i}=\beta_0+\beta_1D_{i}+\beta_2\left(r_{i}-c\right)+\beta_3D_{i}\left(r_{i}-c\right)+u_{i}

where:

Di=1(ri≥c)D_{i}=1\left(r_{i}\ge c\right)

r = running variable
c = cutoff
D = treatment indicator

β₁ = treatment effect at the cutoff.

β₂ = slope below the cutoff.

β₂ + β₃ = slope above the cutoff.

β₃ = difference in slopes on the two sides of the cutoff.

The key RDD assumption is that potential outcomes are continuous at the cutoff, so a discontinuous jump in y at c can be attributed to treatment.

60
New cards

Binary Dependent Variable

A binary dependent variable takes only two values:

yi∈{0,1}y_{i}\in\left\lbrace0,1\right\rbrace

For binary y:

E(yi∣Xi)=P(yi=1∣Xi)E\left(y_{i}\vert X_{i}\right)=P\left(y_{i}=1\vert X_{i}\right)

Therefore, the conditional mean of y can be interpreted as a probability.

61
New cards

Linear Probability Model

The LPM models the probability of y = 1 as:

P(yi=1∣Xi)=β0+β1x1i+⋯+βkxkiP\left(y_{i}=1\vert X_{i}\right)=\beta_0+\beta_1x_{1i}+\cdots+\beta_{k}x_{ki}

and is estimated using OLS.

βj is the change in the probability that y = 1 from a one-unit increase in xj, holding other variables constant.

Example: β = 0.08 means an increase of 8 percentage points.

62
New cards

LPM advantages and problems

Advantages:
Easy to estimate with OLS and coefficients have a direct probability interpretation.

Problems:

  • predicted probabilities can be below 0 or above 1;

  • marginal effects are constant for every value of x;

  • the error term is inherently heteroskedastic.

Therefore, robust standard errors should be used.

63
New cards

Logit and Probit

Both models transform a linear index into a probability between 0 and 1:

z=β0+Xβz=\beta_0+X\beta

Logit:

P(y=1∣X)=ez1+ezP\left(y=1\vert X\right)=\frac{e^{z}}{1+e^{z}}

Probit:

P(y=1∣X)=Φ(z)P\left(y=1\vert X\right)=\Phi\left(z\right)

Logit uses the logistic CDF; Probit uses the standard normal CDF.

Both produce an S-shaped probability curve and usually give similar results

64
New cards

Maximum Likelihood Estimation

Logit and Probit are nonlinear probability models and are estimated using Maximum Likelihood Estimation rather than OLS.

MLE chooses the parameter values that make the observed sample outcomes y = 0 and y = 1 most likely under the model.

65
New cards

Logit/Probit Coefficient interpretation

The sign of β shows the direction of the effect:

β > 0 → x increases P(y=1)

β < 0 → x decreases P(y=1)

But the coefficient itself is not the change in probability.

For example, β = 0.20 does not mean that probability increases by 20 percentage points.

Use marginal effects or predicted probabilities to determine the size of the effect.

66
New cards

Marginal effects in Logit/Probit

A marginal effect converts a Logit or Probit coefficient into an effect on probability.

Because these models are nonlinear, the marginal effect of x depends on the values of the explanatory variables.

Average Marginal Effect (AME): calculate the marginal effect for every observation and take the average.

Example:

AME = 0.06 → a one-unit increase in x increases P(y=1) by 6 percentage points on average

67
New cards

Marginal Effect of a Dummy Variable

For a dummy variable, calculate the change in predicted probability when the dummy changes from 0 to 1:

P(y=1∣D=1,X)−P(y=1∣D=0,X)P\left(y=1\vert D=1,X\right)-P\left(y=1\vert D=0,X\right)

For an Average Marginal Effect, calculate this probability difference for each observation and then take the average.