1/42
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
simple linear regression
Y = β0+β1X1+εi
Y
response variable
x
predictor/covariate
β0
intercept, mean when xi=0
β1
represents change in the expected value of Y for a one-unit increase in X
residual
ei= yi- yihat
desirable residual plot
random scatter around zero
null
H0
alternative
Ha
H0
β1= 0
Ha
β1 ≠ 0
p-value < 0.05
Ha: β1 ≠ 0
p-value > 0.05
H0: β1 = 0
multiple linear regression
Y = β0+β1X1+β2X2+…+βpXp+εi
when interpreting multiple regression coefficients
hold other variables constant
R²=0.72
approximately 72% of the variability response variable is explained by the predictors included in the regression model
R²
tells you how much of the variation your model explains, and the remaining percentage is left unexplained
adjusted R²
Did adding this new variable actually make the model better or did you just add another variable?
add a new predictor provides almost no useful info R² does:
increases/ stays the same
add a new predictor provides almost no useful info Adjusted R²
may decrease
if p-value < 0.05 should we keep or remove variable
keep; significant
if p-value > 0.05 should we keep or remove variable
remove; not significant
AIC
compares regression models by balancing complexity and model performance
complexity - AIC
the more variables in model the more complex
poor performance - AIC
a model that fits data poorly, considers how much error/lack of fit the model has
AIC big or small?
small
adjusted R² big or small
big
ridge-lasso
λ
controls how strong penalty is
regularization
adds a penalty to a regression model to prevent overfitting and keep coefficients from becoming too large
ridge + lasso
shrink coefficients, or perform variable selection by shrinking coefficients to 0, in order to help with multicollinearity
multicollinearity
two or more predictors are giving the model very similar information
poisson
used for count outcomes, mean ≈ variance
log link
keeps mean > 0
if over dispersion occurs
poisson model → negative binomial
IRR-incidence rate ratio
tells you how the expected rate/count changes when a predictor increases by 1 unit
how to interpret log link β
convert to IRR using exp(c) (e^c)
exp(0.08)= 1.08328
8% increase in expected count
expected count
mean
e^(-0.15) = 0.86
14% decrease
VIF-variance inflation factor
checks for multicollinearity
over dispersion occurs because
mean < variance
offset
account for different amounts of exposure or observation time across units