1/66
A set of 114 vocabulary flashcards covering key statistical, mathematical, econometric, and causal inference terms from Weeks 1-7 of Empirical Methods in Finance.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Variance and Standard Deviation
Variance measures how dispersed X is around its mean:
Var(X)=E[(X−E(X))2]
Standard deviation is:
sd(x)=Var(X)
A larger variance or standard deviation means observations are generally further from the mean. Standard deviation is expressed in the same units as X
Covariance and correlation
Covariance measures the linear co-movement of two variables:
Cov(X,Y)=E[(X−E(X))(Y−E(Y))]
Cov>0: they tend to move together.
Cov<0: they tend to move in opposite directions.
Cov=0: no linear relationship.
Correlation standardizes covariance:
ρ(X,Y)=sd(X)sd(Y)Cov(X,Y)
and always lies between -1 and 1. Unlike covariance, correlation does not depend on the units of measurement.
Independence vs zero covariance
Independence means knowing X provides no information about Y:
P(X=x,Y=y)=P(X=x)P(Y=y)
If X and Y are independent:
Cov(X,Y)=0
But the reverse is not generally true:
Cov(X,Y)=0=independence
because covariance only measures linear dependence. Two variables can have zero covariance but still have a nonlinear relationship
Conditional probability and dependence
Conditional probability asks how likely Y is after knowing X:
P(Y=y∣X=x)
If X and Y are independent:
P(Y=y∣X=x)=P(Y=y)
Knowing X does not change the probability of Y. If the conditional probability changes with X, the variables are dependent.
Mean vs Median
The mean is the average and is sensitive to extreme values. The median is the middle observation after ordering the data and is less sensitive to extremes.
Therefore, an extreme outlier can substantially change the mean while leaving the median unchanged
Regression model and error term
In
yi=α+βxi+ui
y is the dependent variable, x the explanatory variable, and u contains all other factors affecting y that are not included in the model.
alpha is the intercept and beta measures how y changes when x increases by one unit.
Population model vs estimated regression
The population model contains the unknown true parameters:
yi=α+βxi+ui
The estimated regression uses sample estimates:
yihat=αhat+βhatxi
betahat estimates the unknown population effect beta. Different random samples generally produce different estimates
OLS residuals
OLS chooses alphahat and betahat to minimize the sum of squared residuals:
∑uihat2
where
uihat=yi−yihat
A residual is therefore the difference between the observed and predicted value of y.
Unbiasedness
An estimator is unbiased if its expected value across repeated random samples equals the true population parameter:
E(βhat)=β
This does not mean every estimate equals beta. Individual estimates can be too high or too low, but there is no systematic error on average.
Efficiency and BLUE
BLUE means Best Linear Unbiased Estimator.
Linear: the estimator is linear in y.
Unbiased: E(βhat)=β .
Best: it has the smallest variance among linear unbiased estimators.
Thus, “best” means most efficient within this class, not that every individual estimate is closest to the true value.
CLRM Assumptions
[A1] Linear in parameters.
[A2] Random sample from the population.
[A3] Variation in x.
[A4] E(ui∣xi)=0 .
[A5] Var(ui∣xi)=σ2.
[A1]–[A4] imply unbiased OLS.
Adding [A5] gives efficiency and therefore BLUE
Zero Conditional Mean
The crucial assumption for OLS unbiasedness is:
E(ui∣xi)=0
It means that the average unobserved factors in ui do not systematically vary with xi.
If an omitted factor affects y and is related to x, A4 is violated and OLS will generally be biased.
R2
R2 measures the proportion of sample variation in y explained by the regression:
R2=1−TSSRSS=TSSESS
A higher R2 means better in-sample fit.
It does not imply that the model is causal, unbiased, or correctly specified
Standard Error of OLS
The standard error measures the uncertainty of an estimated coefficient: a smaller standard error means a more precise estimate.
For the slope, precision improves when there is more variation in x, a larger sample, or less unexplained variation in y
Standardized coefficients
A standardized coefficient is:
σyBσx
It measures how many standard deviations predicted y changes when x increases by one standard deviation.
Functional form interpretations
Dependent variable | Independent variable | Interpretation of β |
|---|---|---|
Level | Level | 1-unit increase in x → β-unit change in y |
Level | Log | 1% increase in x → β/100-unit change in y |
Log | Level | 1-unit increase in x → approximately 100β% change in y |
Log | Log | 1% increase in x → approximately β% change in y |
OLS Standard Error formula
The standard error measures the uncertainty of an estimated coefficient. It estimates how much the coefficient estimate would vary across repeated samples.
Smaller SE → more precise estimate.
In the bivariate homoskedastic case:
Var(βhat)=∑(xi−xav)2σ2
Therefore, precision increases when there is more variation in x and decreases when there is more unexplained variation in y
t-Test
A t-test tests a hypothesis about one regression coefficient.
t=SE(βhat)βhat−β0
For the common null hypothesis H0: beta = 0:
t=SE(βhat)βhat
The t-statistic tells us how many standard errors the estimate is away from the value specified under H0
Statistical significance and decision rule
The significance level \(\alpha\) is chosen before the test, commonly 10%, 5%, or 1%.
Reject H0 when:
p<α
or equivalently, for a two-sided test:
|t| > tcritical
If p≥α , fail to reject H0. This means there is insufficient evidence against H0; it does not prove that H0 is true.
p-value
The p-value measures how extreme the observed result would be if H0 were true.
A smaller p-value → stronger evidence against H0.
Example: p=0.03 means that, assuming H0 is true, there is a 3% probability of obtaining a result at least as extreme as the observed result
One-sided vs Two-sided test
A two-sided test asks whether the coefficient differs from the hypothesized value in either direction.
A one-sided test asks whether the coefficient differs in one specified direction.
For the same t-statistic in the predicted direction, the one-sided p-value is half the two-sided p-value.
Confidence interval
A confidence interval gives a range of plausible values for the population coefficient:
βhat±tcriticalSE(βhat)
For a large-sample 95% confidence interval:
βhat±1,96SE(βhat)
A two-sided H0: beta = beta0 is rejected at the corresponding significance level if beta0 lies outside the confidence interval.
Statistical vs economic significance
Statistical significance asks whether there is sufficient statistical evidence that an effect differs from the hypothesized value.
Economic significance asks whether the size of the estimated effect is economically meaningful.
An effect can therefore be statistically significant but economically small, or economically large but estimated too imprecisely to be statistically significant.
A6 Normality
The population error u is independent of the explanatory variables x and normally distributed:
u−N(0,σ2)
A6 implies that the OLS estimator is normally distributed, which allows us to perform exact hypothesis testing using the t-distribution.
A6 is not required for OLS to be unbiased; it is added for statistical inference
Failure of A4 / Endogeneity
A4 requires the explanatory variables to be unrelated to the expected error term.
If an explanatory variable is correlated with u, it is endogenous and OLS is generally biased and inconsistent.
Main causes:
omitted variables
simultaneity / reverse causality
measurement error in an explanatory variable
Omitted variable bias
Omitted variable bias occurs when an omitted variable:
affects y, and
is correlated with an included explanatory variable.
The direction of the bias depends on:
the effect of the omitted variable on y;
its correlation with the included variable.
Same signs → positive bias.
Opposite signs → negative bias.
Simultaneity / reverse causality
Simultaneity occurs when x affects y while y also affects x.
This makes x correlated with the error term, violating A4. OLS is therefore biased and inconsistent and cannot isolate the one-way causal effect of x on y.
Measurement error
Measurement error in x → x becomes correlated with the regression error. Under classical measurement error, the estimated coefficient is typically biased toward zero. This is attenuation bias.
Measurement error in y → does not bias the OLS coefficients if it is unrelated to the explanatory variables, but increases noise and reduces precision.
Failure of A5 / Heteroskedasticity
A5 requires constant error variance:
Var(u∣X)=σ2
Heteroskedasticity means the variance of u changes with the explanatory variables.
If A1–A4 still hold:
OLS coefficients remain unbiased/consistent;
OLS is no longer BLUE;
conventional standard errors can be incorrect.
Robust SE
Robust standard errors allow valid inference when heteroskedasticity is present.
They change:
standard errors
t-statistics
p-values
confidence intervals
They do not change the OLS coefficient estimates.
Breusch-Pagan test
Tests for heteroskedasticity.
Estimate OLS and obtain residuals.
Regress the squared residuals on the explanatory variables.
Jointly test whether the slope coefficients in this auxiliary regression equal zero.
H₀: homoskedasticity.
p < α → reject H₀ → evidence of heteroskedasticity.
Irrelevant variables and multicollinearity
Including an irrelevant variable does not bias OLS, but can increase standard errors and reduce precision.
Strong but imperfect multicollinearity also does not bias OLS, but increases standard errors and makes individual effects harder to estimate precisely.
F-test
An F-test tests multiple restrictions jointly.
The unrestricted model does not impose H₀; the restricted model does.
F=(n−k−1)RSSURq(RSSR−RSSUR)
where q is the number of restrictions.
Reject H₀ if F > Fcritical or p< alpha.
Unlike a t-test, which tests one coefficient, an F-test can test several coefficients jointly.
Dummy variables and reference group
A dummy variable takes the value 0 or 1.
Its coefficient measures the difference between the group with D = 1 and the reference group with D = 0, holding other explanatory variables constant.
With an intercept and multiple categories, omit one category as the reference group. Including every category creates perfect multicollinearity: the dummy variable trap
Dummy variable with log y
When y is in logs, a dummy coefficient δ represents approximately a:
100δ%
difference between D = 1 and D = 0.
The exact percentage difference is:
100[eδ−1]%
Interaction terms
An interaction means that the effect of one explanatory variable depends on another variable.
Continuous × continuous: the effect of x₁ changes with the value of x₂.
Dummy × continuous: the interaction coefficient measures how the slope of x differs between D = 1 and D = 0.
So an interaction coefficient should not normally be interpreted as an independent direct effect
Unit fixed effects with dummies
Unit fixed effects can be implemented by including a dummy for each unit except one reference unit.
They control for all differences between units that are constant over time.
DiD basic idea
DID=(Treatedpost−Treatedpre)−(Controlpost−Controlpre)
DiD compares the change over time in the treated group with the change over time in the control group.
It subtracts the common time change to isolate the treatment-associated change.
Clustered SE
Clustered standard errors allow regression errors to be correlated within groups, such as observations from the same firm over time.
Ignoring this correlation can make conventional standard errors too small and statistical significance too strong.
Clustering changes standard errors, t-statistics, p-values and confidence intervals, but not the OLS coefficient estimates
Adjusted R2
Adjusted R² measures model fit while penalizing the addition of explanatory variables.
1−(n−1)TSS(n−k−1)RSS
where:
n = number of observations
k = number of explanatory variables, excluding the intercept
Unlike R², which can never decrease when a variable is added, adjusted R² can decrease if the new variable does not improve the model sufficiently.
Therefore, adjusted R² is more useful for comparing models with different numbers of explanatory variables.
First Differences
First Differencing removes time-invariant unobserved effects by subtracting the previous period from the current period.
Starting with:
yit=αi+βxit+uit
we obtain:
Δyit=βΔxit+Δuit
because the time-invariant effect αi disappears.
β is therefore estimated using changes within the same unit over time.
Attenuation bias and FD
Measurement error in x biases the estimated coefficient toward zero. This is attenuation bias.
With First Differencing:
Δxitobs=Δxit+(vit−vi,t−1)
and, with independent measurement errors:
Var(vit−vi,t−1)=2σv2
Therefore, differencing can increase measurement-error noise relative to the true variation in x, making attenuation bias more severe
Fixed Effects
Fixed Effects controls for all time-invariant characteristics of each unit, including unobserved characteristics that may be correlated with x.
Starting with:
yit=αi+βxit+uit
FE demeans the variables by subtracting each unit's time average:
yit−yiav=β(xit−xiav)+(uit−uiav)
The time-invariant αi disappears.
Therefore, FE estimates β from variation within the same unit over time.
Time Fixed Effects and Two-Way Fixed Effects
Time Fixed Effects control for shocks that affect all units in the same period, such as a financial crisis.
Two-Way Fixed Effects combine:
Unit FE → control for permanent differences between units.
Time FE → control for common shocks in each period.
Thus, Two-Way FE controls for both time-invariant unit characteristics and common time effects
FE vs FD
If T = 2: FE and FD produce identical slope estimates.
If T > 2: they do not necessarily produce identical estimates.
FE:
usually preserves more observations in an unbalanced panel;
measurement-error bias may decrease as T increases.
FD:
requires consecutive observations;
can make attenuation bias from measurement error more severe
Random Effects
Random Effects models the unit-specific effect as a random component:
yit=α+βxit+ϵi+uit
The crucial RE assumption is:
Cov(xit,ϵi)=0
for all periods.
Unlike FE, RE therefore requires the unit-specific unobserved effect to be uncorrelated with the explanatory variables
FE vs RE
FE:
allows the unit-specific effect to be correlated with x;
cannot estimate effects of time-invariant variables;
is more robust but less efficient.
RE:
requires the unit-specific effect to be uncorrelated with x;
can estimate time-invariant variables;
is more efficient if the RE assumptions hold.
Use RE only when its stronger independence assumption is credible.
Hausman test
The Hausman test compares FE and RE estimates to assess whether the RE assumption is appropriate.
H₀: RE is consistent.
If p<alpha → reject H₀ → evidence that the unit effect is correlated with the regressors → prefer FE.
If p>alpha → fail to reject H₀ → no evidence against RE
FE vs Clustered SE
FE and clustering solve different problems.
Fixed Effects remove time-invariant unobserved heterogeneity.
Clustering adjusts inference when errors are correlated within units or groups.
Using FE does not eliminate the need for clustered standard errors when within-cluster error correlation is present.
Clustering
Clustering allows regression errors to be correlated within the same group, such as observations of the same firm over time.
Without clustering, correlated observations are treated as providing more independent information than they actually do, which can make standard errors too small and statistical significance too strong.
Clustered standard errors:
allow correlation and heteroskedasticity within clusters;
assume independence across clusters;
change SEs, t-statistics, p-values and confidence intervals;
do not change the coefficient estimates.
Cluster at the level where errors are likely to be correlated, e.g. by firm for within-firm dependence
IV
IV is used to estimate the causal effect of endogenous x using an instrument z.
A valid instrument must satisfy:
Relevance:
Cov(z,x)=0
z must predict x.
Exogeneity:
Cov(z,u)=0
z must be unrelated to unobserved determinants of y.
Both conditions are required for a valid instrument.
IV and weak instrument
Test relevance in the first-stage regression:
xi=δ0+δ1zi+controls+vi
The instrument is relevant when z predicts x.
Use the first-stage F-statistic:
F > 10
is the lecture's rule of thumb for sufficient instrument strength.
A weak instrument provides little useful variation in x and makes IV estimates unreliable and imprecise
Instrument exogeneity
Exogeneity requires:
Cov(z,u)=0
The instrument must be unrelated to unobserved determinants of y.
Unlike relevance, exogeneity generally cannot be tested directly because u is unobserved. It therefore requires a credible economic or institutional argument
2SLS
2SLS estimates an IV model in two stages.
Stage 1: regress endogenous x on instrument z and controls → obtain predicted xhat.
Stage 2: regress y on predicted xhat and the same controls.
The coefficient on xhat uses only the variation in x generated by the instrument to estimate the causal effect
OSL vs IV
OLS uses all variation in x, including variation that may be correlated with the error term.
IV uses only the variation in x generated by a valid instrument z.
If x is endogenous but z is relevant and exogenous:
OLS → biased/inconsistent.
IV → consistent, but generally less precise.
With one x and one instrument:
βIVhat=Cov(z,x)Cov(z,y)
DiD regression
yit=β0+β1(Postt⋅Treati)+β2Treati+β3Posti+uit
β₁ = DiD treatment effect.
β₂ = pre-treatment difference between treated and control groups.
β₃ = time change for the control group.
The interaction \(Post\times Treat\) therefore identifies the DiD effect
RDD
RDD exploits a treatment rule based on a cutoff.
Observations just below and just above the cutoff are compared. A discontinuous jump in y at the cutoff identifies a local treatment effect.
Key assumption: without treatment, the expected outcome would change continuously at the cutoff.
Therefore, a jump at the cutoff can be attributed to the treatment.
Parallel trend assumption
Without treatment, the treated and control groups would have experienced the same average change in the outcome over time.
Formally, the lecture states the assumption for a consistent DiD estimator as:
E(u∣Post,Treat)=0
Parallel trends cannot be tested directly because we cannot observe what would have happened to the treated group without treatment. Similar pre-treatment trends provide supporting evidence.
RDD regression
yi=β0+β1Di+β2(ri−c)+β3Di(ri−c)+ui
where:
Di=1(ri≥c)
r = running variable
c = cutoff
D = treatment indicator
β₁ = treatment effect at the cutoff.
β₂ = slope below the cutoff.
β₂ + β₃ = slope above the cutoff.
β₃ = difference in slopes on the two sides of the cutoff.
The key RDD assumption is that potential outcomes are continuous at the cutoff, so a discontinuous jump in y at c can be attributed to treatment.
Binary Dependent Variable
A binary dependent variable takes only two values:
yi∈{0,1}
For binary y:
E(yi∣Xi)=P(yi=1∣Xi)
Therefore, the conditional mean of y can be interpreted as a probability.
Linear Probability Model
The LPM models the probability of y = 1 as:
P(yi=1∣Xi)=β0+β1x1i+⋯+βkxki
and is estimated using OLS.
βj is the change in the probability that y = 1 from a one-unit increase in xj, holding other variables constant.
Example: β = 0.08 means an increase of 8 percentage points.
LPM advantages and problems
Advantages:
Easy to estimate with OLS and coefficients have a direct probability interpretation.
Problems:
predicted probabilities can be below 0 or above 1;
marginal effects are constant for every value of x;
the error term is inherently heteroskedastic.
Therefore, robust standard errors should be used.
Logit and Probit
Both models transform a linear index into a probability between 0 and 1:
z=β0+Xβ
Logit:
P(y=1∣X)=1+ezez
Probit:
P(y=1∣X)=Φ(z)
Logit uses the logistic CDF; Probit uses the standard normal CDF.
Both produce an S-shaped probability curve and usually give similar results
Maximum Likelihood Estimation
Logit and Probit are nonlinear probability models and are estimated using Maximum Likelihood Estimation rather than OLS.
MLE chooses the parameter values that make the observed sample outcomes y = 0 and y = 1 most likely under the model.
Logit/Probit Coefficient interpretation
The sign of β shows the direction of the effect:
β > 0 → x increases P(y=1)
β < 0 → x decreases P(y=1)
But the coefficient itself is not the change in probability.
For example, β = 0.20 does not mean that probability increases by 20 percentage points.
Use marginal effects or predicted probabilities to determine the size of the effect.
Marginal effects in Logit/Probit
A marginal effect converts a Logit or Probit coefficient into an effect on probability.
Because these models are nonlinear, the marginal effect of x depends on the values of the explanatory variables.
Average Marginal Effect (AME): calculate the marginal effect for every observation and take the average.
Example:
AME = 0.06 → a one-unit increase in x increases P(y=1) by 6 percentage points on average
Marginal Effect of a Dummy Variable
For a dummy variable, calculate the change in predicted probability when the dummy changes from 0 to 1:
P(y=1∣D=1,X)−P(y=1∣D=0,X)
For an Average Marginal Effect, calculate this probability difference for each observation and then take the average.