1/55
A complete set of vocabulary flashcards covering key statistical concepts from Simple Linear Regression, Multiple Linear Regression, Inference, Diagnostics, Indicator Variables, and Transformations.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Scatterplot
A graph showing the relationship between two quantitative variables.
Correlation
Measures direction and strength of linear relationship
R=1
positive and strong correlation
Degree of freedom for t test (SLR)
N-2
N
Sample size
K
Number of predictor variables
Test statistic for t test
b1-0 / SE ( of b1)
Test statistic for F test
MSR/MSE
Degree of freedom for t test (MLR)
N-k-1
Degree of freedom for f test
K/ n-k-1
R= 0
Weak, no Linear relationship
R= -1
Negative strong correlation
Key properties of correlation
Unit free. (Changing Units doesn’t change r) , doesn’t. depend on which variable is X or Y, only describes linear relationships
Regression equation
Least squares regression helps predict a response variable (y) using an exploratory variable (x) , finds best fit line for data by minimizing, the squared residuals
Slope
b1= r* (Sy/Sx)
What does slope represent
Average change in Y for a one unit increase in X
Sy
Standard deviation of y
Sx
Standard deviation of X
Y intercept
b0= ybar- b1*(xbar)
DOTS
An acronym for key features used to describe a scatterplot: Direction (positive, negative, none), Outliers, Trend/form (linear or nonlinear), and Strength (weak, moderate, strong).
Coefficient of Determination (R2)
A metric with a range from 0 to 1 (0 to 100 percent) that represents the percentage of variation in y explained by the regression on x.
Residual
The difference between the observed response value and predicted response value (residual=y−y^); positive when observed y is higher than predicted, and negative when lower.
Least Squares Regression
A regression method that chooses the line that makes the sum of squared residuals as small as possible.
Linear Assumption
A regression assumption requiring that the relationship looks linear and the residual plot shows random scatter.
Independence Assumption
A regression assumption satisfied when observations are collected using randomization or random selection.
Equal Variance Condition
Also known as the equal spread condition, a regression assumption requiring residuals to have consistent spread with no funnel or fan shape.
Normality Condition
A regression assumption requiring residuals to be bell-shaped or follow roughly a straight line on a normal probability plot, rather than an S-shape.
Simple Linear Regression (SLR)
A statistical technique used to predict a response variable y, model the relationship between x and y, and estimate the rate of change (slope) between them.
Fitted Line (SLR)
The regression line derived from sample statistics, represented by the equation y^=b0+b1x.
True Model (SLR)
The underlying population parameter regression model, represented by the equation y=β0+β1x+e.
Fitted line (MLR)
Y hat= b0 + b1(x) + b2(x2) ….
True model (MLR)
Y= beta0 + Beta1(x1) + Beta2 (x2) +…..+ e
t-test for Slope (SLR)
df=n−2.; determine if true slope is different from 0
Permutation / Randomization Test
A hypothesis testing method performed by rearranging data labels under the assumption that the null hypothesis H0 is true.
Bootstrap Percentile Method for Slope
A resampling method for finding possible slope values by taking the 2.5th and 97.5th percentiles of a bootstrap distribution.
Bootstrap Standard Error Method for Slope
A method for estimating plausible slope values by calculating SEboot from a bootstrap distribution and constructing the confidence interval b1±2×SEboot.
Prediction Interval (PI)
An interval giving plausible values for an individual future response y at a given x; it is much wider than the confidence interval for mean y.
Confidence Interval (CI) for Mean Y
An interval giving plausible values for the average or expected response y at a given predictor value x.
Multiple Linear Regression (MLR)
A regression technique that uses more than 1 predictor variable (x) to predict a response variable y.
MLR Slope Interpretation
The expected change in y for a 1-unit increase in predictor xi, holding all other predictor variables constant.
Overall F-test
A hypothesis test for the whole model testing H0:all slopes=0 against Ha:at least one slope=0 using the test statistic F=MSEMSR.
Individual t-test for slope (MLR)
A hypothesis test for ONE predictor variable testing H0:βi=0 against Ha:βi=0 with df=n−k−1 to determine if that predictor is useful after accounting for all other predictors.
Total Sum of Squares (SST)
The measure of total variation in response variable Y, where SST=SSR+SSE with degrees of freedom df=n−1.
Regression Sum of Squares (SSR)
The portion of total variation in Y explained by the regression model, with degrees of freedom df=k.
Sum of Squared Errors (SSE)
The portion of total variation in Y not explained by the regression model, with degrees of freedom df=n−k−1.
Mean Square Regression (MSR)
The model sum of squares divided by model degrees of freedom, computed as MSR=kSSR.
Mean Square Error (MSE)
The sum of squared errors divided by residual degrees of freedom, computed as MSE=n−k−1SSE.
F-distribution
A probability distribution used to compare the ratio of two variances, indexed by two degrees of freedom parameters (k and n−k−1).
Indicator Variable
A binary variable used to incorporate categorical variables into a regression model (1=present, 0=not present); changes the intercept
Base Category
The category in a regression model without its own indicator variable, which serves as the reference level changing the y-intercept.
Interaction Term
(quantitative×indicator) ; changes the slope
Multicollinearity
A condition in multiple regression where predictor variables are highly correlated with each other, providing redundant information to the model.
Multicollinearity Signs
t-tests produce large p-values ; F-test produces a small p-value, standard errors become very large.
Variance Inflation Factor (VIF)
A measure of multicollinearity where the lowest possible value is 1, and a value of VIF≥5 indicates a multicollinearity problem.
Adjusted R2
useful when comparing MLR models ; accounts for the number of predictors (k) rather than automatically rewarding adding more variables.
Box-Cox Transformation
A variable transformation technique for y where λ=1 results in no change, λ=0 transforms to log(y), and λ=2 transforms to y2.