1/38
Flashcards covering dummy variables, non-linear functional forms, holding factors constant, sources of variation, regression pitfalls, standard errors, hypothesis testing, and regression objectives.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Dummy Variable
A binary indicator variable taking values 1 and 0 used to represent qualitative or categorical factors such as gender, race, or type of education.
Reference Category
The baseline categorical group excluded from a regression model with an intercept, against which all included dummy variable coefficients are compared.
Interaction Effect
A model specification where an explanatory factor's effect on the outcome variable differs across categories, treating the category as a moderating factor.
Log-Linear Model
A regression model of the form ln(Y)=β0+β1X+ε, where a 1-unit increase in X is associated with a β1×100 percent change in Y.
Linear-Log Model
A regression model of the form Y=β0+β1ln(X)+ε, where a 1-percent increase in X is associated with a 100β1 unit change in Y.
Log-Log Model
A regression model of the form ln(Y)=β0+β1ln(X)+ε, where a 1-percent increase in X is associated with a β1 percent change in Y.

Quadratic Model
A non-linear regression model that adds higher powers of an explanatory variable to capture non-linear relationships, such as score^=β0+β1hours+β2hours2.

Spline Model
A regression model constructed piecewise from linear segments, using dummy variables and interaction terms to allow the slope of X to change past a specified threshold value.
Weighted Least Squares
A regression estimation method that assigns weighting schemes to observations based on survey over-sampling, corporation size, or state population rather than treating all observations equally.
Standardized Coefficient Estimate
A re-scaled regression estimate interpreted as the effect of a one standard deviation increase in X on Y, allowing comparison across variables measured on different scale units.
Good Variation
Variation in the key explanatory variable caused by factors that are not correlated with the dependent variable, other than through their impact on the key explanatory variable itself.
Bad Variation
Variation in the key explanatory variable caused by factors that could be correlated with the dependent variable beyond through the key explanatory variable.
Operative Variation
The active variation in a key explanatory variable that directly determines its estimated empirical relationship with the dependent variable in a regression model.
Held-Constant Variation
Variation in a variable that is adjusted or controlled for in a regression model, preventing it from influencing the estimated coefficient of the key explanatory variable.
Average Treatment Effect
The average change in the outcome variable if all subjects in a population were given one additional unit of treatment compared to receiving no treatment.
Mediating Factor
A variable that represents a mechanism through which the key explanatory variable affects the outcome, which should not be controlled for when estimating causal effects.
Confounding Factor
An extraneous factor that affects both the key explanatory variable and the outcome variable, generating bad variation that must be controlled for.

Pitfalls for Control Selection
Criteria governing which variables to include (such as confounders affecting both key-X and outcome) and exclude (such as mediating factors or outcomes of key-X) in a causal regression model.
Sample-Selection Bias
A systematic error occurring when subjects are non-randomly selected into a sample based on a factor related to the outcome variable.
Attrition Bias
A form of selection bias where subjects who stay in a sample over time systematically differ from those who drop out or stop responding.
Reverse Causality
A pitfall in causal regression analysis occurring when the outcome variable systematically affects the treatment variable.
Omitted-Factors Bias
Bias in a coefficient estimate resulting from omitting an unobserved factor that influences both the key explanatory variable and the outcome variable.
Self-Selection Bias
Bias arising when subjects choose or get assigned to a treatment level based on personal characteristics linked to individual benefits or costs.
Measurement Error
Coding or conceptual error in an explanatory variable, which typically attenuates coefficient estimates toward zero when the error is random.
Improper Reference Group
A modelling pitfall occurring when the omitted baseline group does not represent the correct counterfactual for evaluating the treatment effect.
Over-Weighted Groups
A pitfall occurring when treatment effects vary across subgroups, causing pooled OLS estimates to over-represent groups with greater within-group variance in the key variable.
Standard Error of the Estimate
A measure of precision for a coefficient estimate, computed as SE(β^j)=∑(Xj−Xˉj)2×(1−Rj2)σ^.

Type I Error
A false positive decision in hypothesis testing where a true null hypothesis is incorrectly rejected, analogous to convicting an innocent defendant in a criminal trial.

Type II Error
A false negative decision in hypothesis testing where a false null hypothesis is not rejected, analogous to acquitting a guilty defendant in a criminal trial.

p-value
The probability that randomness would generate a test statistic as far from its hypothetical value as observed, assuming the null hypothesis were true.
Critical Value
A cut-off threshold on a test distribution past which an observed test statistic is concluded to be too far from its hypothetical value to be caused by chance.
t-Statistic
A test statistic calculated as t=SE(β^)β^−β∗ used to evaluate hypotheses concerning an individual regression coefficient.
Two-Sided Hypothesis Test
A test evaluating H0:βi=0 against H1:βi=0, placing rejection regions in both tails of the Student's t-distribution.
One-Sided Hypothesis Test
A hypothesis test used when theory dictates a directional effect, evaluating H0:βi≤0 against H1:βi>0 (or vice versa).
Confidence Interval
An estimated range constructed as β^±tc×SE(β^) that contains the true parameter value with a specified degree of confidence.
Joint Hypothesis Test
An F-test evaluating whether a specific subset of explanatory variables collectively have a statistically significant relationship with the outcome variable.
Overall-Significance Test
An F-test reported in regression outputs testing whether all slope coefficients in a model are jointly equal to zero (H0:β1=β2=⋯=βK=0).
Practical Significance
The real-world magnitude and importance of an estimated effect, distinct from statistical significance which depends heavily on sample size.
Regression Objectives
The four main empirical goals of regression analysis: estimating causal effects, making predictions, determining predictors, and adjusting outcomes.