1/152
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Positive linear relationship
As one variable increases, the other also rises consistently; the variables move in the same direction
Negative linear relationship
As one variable increases, the other decreases; the variables move in opposite directions
Curvilinear relationship
A relationship that doesn't follow a straight line and is better described by a curve or polynomial function; a more complex relationship
Correlation
Gauges the degree to which two variables change together
Correlation coefficient (r): range and sign
Ranges from -1 to 1; positive means a positive relationship, negative means a negative relationship
What does r = 0 mean?
There is no linear relationship between the variables
How do you judge the strength of a correlation?
The closer r is to -1 or 1, the stronger the relationship
Correlation vs. causation
Correlation does not imply causation; r doesn't show that changes in one variable cause changes in another
Linear regression
A tool to model and analyse the relationship between a dependent variable and one or more independent variables
Line of best fit
The line from simple linear regression (one independent variable) fitted to the observed data
Error (for one observation)
error = y − ŷ (actual minus predicted)
Mean squared error (MSE)
Square each observation's error, then average them for the model
Training a model
Selecting the best set of values for β0, β1 (and other coefficients)
Ordinary Least Squares (OLS)
The method used to calculate the coefficients in linear regression
Root Mean Squared Error (RMSE)
The square root of the average of the squared differences between observed and predicted values; measured in the same units as the dependent variable
How do you decide if an RMSE (e.g., 3.29) is good?
Compare it with the RMSE of a baseline model that predicts the training-data mean for every case; your model should be clearly below it
Coefficient of determination (R²)
The proportion of the variance in the dependent variable that is explained by the independent variables in a regression model
Range of R² when computed on the same rows the model was fitted to
Between 0 and 1
Negative R² on held-out test rows
The model predicts the held-out rows worse than just predicting the training mean; it lost to the baseline
R² = 0
The independent variables explain none of the variability in the dependent variable
R² = 1
The independent variables perfectly explain the variability in the dependent variable
R² values in finance
R² is often far from 1, since financial outcomes are influenced by numerous unobserved factors
Interpreting R² close to 1 vs. close to 0
Close to 1: captures most of the variability; close to 0: captures little and the residuals are large
What happens to R² when you add another predictor to the same fitting sample?
R² stays the same or increases; it cannot decrease
Why doesn't a higher R² always mean an added variable is meaningful?
Because the new variable could be capturing noise rather than true signal
Overfitting
A model too tailored to its training dataset, capturing its noise and anomalies, so it is less generalizable to new data
Adjusted R²
A version of R² that accounts for the number of predictors, penalizing model complexity
Adjusted R² penalty term: what is the effect of k?
As k increases, adjusted R² is pushed down unless R² rises by enough to offset it
Relationship between adjusted R² and R²
Adjusted R² can be less than R² but never greater
Adjusted R² increases significantly after adding a variable
The variable is providing meaningful explanatory power
Adjusted R² decreases after adding a variable
The variable might not be adding meaningful information, given the complexity cost
Binary variable
A variable taking two possible values, often 0 and 1 (also called a dummy variable or indicator)
What do you often need to do with binary variables when you first receive the data?
Convert them to 0-1 values
What does a binary variable represent in regression analysis?
A distinction between two categories
Interpreting a binary (dummy) variable coefficient
The difference in the mean of the dependent variable between the two categories, among cases alike on every other predictor in the model
When is a dummy's coefficient equal to the raw difference between the two group means?
Only when the dummy is the model's sole predictor
Categorical variable
Variables with multiple categories and no natural ordering (e.g., states, colors)
How many dummy variables do you create for a categorical variable with K categories?
K-1 dummy variables
Reference category
The category left out of the dummies; the baseline that other categories are compared against
Dummy variable trap
Including all K dummies alongside the intercept, which creates perfect multicollinearity
What happens under perfect multicollinearity?
The model has no unique solution and the coefficients returned are arbitrary
Ordinal variable
Ordered categories where the order has meaning but the distance between categories is not uniform
When can an ordinal variable be treated as continuous in regression?
Sometimes, if the ordered values have a linear relationship with the dependent variable
Alternative way to handle an ordinal variable
Treat it like a categorical variable using dummy coding
Dummy variable coefficients in a model with a categorical variable are interpreted…
Relative to the reference category
How do you avoid multicollinearity with dummy variables?
Always omit one dummy; it becomes the reference category
Interaction term
Terms that model how the relationship between one independent variable and the dependent variable changes depending on the level of another independent variable
When should you include an interaction term?
When it is hypothesized that the effect of one variable depends on the category of another variable
Hypotheses tested by the p-value of a regression coefficient
H0: the regression coefficient is equal to zero. HA: the regression coefficient is not equal to zero
A low p-value (e.g., < 0.05)
You can reject the null hypothesis; the coefficient is statistically different from zero
What do you first need to calculate to get the p-value for a regression coefficient?
The t-statistic
Interpreting a regression coefficient (beta)
The average change in the dependent variable associated with a one-unit difference in the independent variable (holding the other predictors constant, in multiple regression); an association, not causation, on observational data
What does the p-value measure?
The strength of evidence against a null hypothesis; smaller means stronger evidence to reject
P-value greater than or equal to the significance level
There is insufficient evidence to reject the null hypothesis, which is not evidence that the null is true
What does 'not significant' mean?
'Not detected here', never 'shown to be zero'
Failing to reject in a Breusch-Pagan or Shapiro-Wilk test
It is not a clean bill of health; failing to reject does not show the assumption holds
Individual p-values vs. the F-test
Individual p-values test coefficients one at a time; the F-test asks whether a group of coefficients is zero together and is computed from sums of squares
P-values and model selection
They can help decide which variables to retain, but should not be the only criterion; also consider practical significance, domain knowledge, and other measures
Does a small p-value mean a big effect?
It is not a measure of effect size; the effect may not be practically significant
P-hacking
Trying multiple methods to find significant p-values; a questionable research practice that can produce non-reproducible results
F-test
A test that compares the fits of nested models and assesses multiple coefficients simultaneously
F-test hypotheses (overall significance)
H0: all slope coefficients are zero. HA: at least one coefficient is not zero
What does rejecting the F-test null tell you?
At least one predictor carries information, but not that the fit is good
A large F-statistic
The model explains a significant amount of variation, leading to rejection of the null
Confidence interval (CI)
A range of plausible values for an unknown parameter; for a coefficient, the range of values the data do not reject at a stated significance level
Meaning of a 95% confidence level
The method produces intervals that contain the true parameter in 95% of repeated samples
CI does not include 0
The coefficient is statistically significant at the matching significance level (5% for a 95% interval)
CI includes 0
We cannot reject the null that the coefficient is zero; not statistically significant at that level
Factors affecting the width of a CI
Sample size (larger gives narrower), variability in the data (more gives wider), and confidence level (higher gives wider, e.g., 99% vs. 95%)
Common CI misconception
Thinking a 95% CI means a 95% probability that this interval contains the true coefficient
Assumptions behind CI validity in linear regression
Linearity, independence, homoscedasticity, and normality
Assumption 1: Linearity
The relationship between the dependent and independent variables is linear
Assumption 2: Independence, and why it matters
Observations are independent; dependence usually doesn't bias coefficients but makes standard errors far too small, so t-stats, p-values, and CIs are overstated
Assumption 3: Homoscedasticity (constant variance)
The variance of the error terms is constant across all levels of the independent variables
Assumption 4: Normality of errors
The errors follow a normal distribution (especially important for hypothesis testing)
Assumption 5: No multicollinearity
The independent variables are not highly correlated with each other, and no predictor is an exact linear combination of the others
Binary response data
Data where the outcome variable has only two possible outcomes (e.g., success/failure, yes/no, 1/0); a binary classification problem
Why not use linear regression for a binary response?
It can produce predicted probabilities below 0 or above 1, which don't make sense
Logistic regression
A classification algorithm used to predict a binary outcome
What does logistic regression fit?
An S-shaped logistic function instead of a straight line
Interpreting a logistic regression coefficient
The change in the log odds of the outcome for a one-unit increase in that predictor, all else equal
How do you convert a logistic regression coefficient to an odds ratio?
Exponentiate it
Logistic regression assumption: Linearity
The relationship between the log odds and the predictors is linear
Logistic regression assumption: Independence
Observations are independent
Logistic regression assumption: No multicollinearity
Predictors are not perfectly correlated
Logistic regression strengths
Outputs interpretable as probabilities; easy to implement and use; very efficient to train
Logistic regression weaknesses
Makes strong assumptions; does not perform well with missing data; underperforms with multiple or non-linear decision boundaries; doesn't naturally capture complex relationships
Imbalanced data
Binary classification data where one class significantly outnumbers the other (e.g., fraud detection, rare disease prediction)
Imbalanced data difficulty: Model bias
The model becomes biased toward the majority class, often predicting it by default because it reduces overall error
Imbalanced data difficulty: Poor generalization
The model may have good overall accuracy but fail to correctly identify the minority class
Imbalanced data difficulty: Misleading metrics
Accuracy can mislead: always predicting the majority class (95% of data) would still score 95% accuracy
How should a classifier on imbalanced data be judged?
The confusion matrix, with measures such as sensitivity (recall) and precision
Sensitivity (recall)
The share of the truly positive cases that the model flags
Precision
The share of the flagged cases that turn out to be positive
SMOTE
An oversampling method that creates synthetic samples for the minority class (Synthetic Minority Over-sampling Technique)
How SMOTE creates a new sample
Take a minority sample, pick one of its k nearest minority-class neighbours, and create a new point on the line between them (original value plus a random fraction of the difference, per attribute)
Where do you apply SMOTE: training data, test data, or both?
Only the training data; the test data is left as is for a true evaluation
Principal components analysis (PCA)
A dimension reduction technique that converts variables into a set of linearly uncorrelated variables called principal components
Orthogonal transformation
The transformation PCA uses to produce principal components that are linearly uncorrelated
Main objective of PCA
Capture as much of the variability as possible with a smaller number of principal components