Statistics: Regression Analysis and Modeling

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
full-widthPodcast
1
Card Sorting

1/55

flashcard set

Earn XP

Description and Tags

A complete set of vocabulary flashcards covering key statistical concepts from Simple Linear Regression, Multiple Linear Regression, Inference, Diagnostics, Indicator Variables, and Transformations.

Last updated 2:07 AM on 9/30/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

56 Terms

1
New cards

Scatterplot

A graph showing the relationship between two quantitative variables.

2
New cards

Correlation

Measures direction and strength of linear relationship

3
New cards

R=1

positive and strong correlation

4
New cards

Degree of freedom for t test (SLR)

N-2

5
New cards

N

Sample size

6
New cards

K

Number of predictor variables

7
New cards

Test statistic for t test

b1-0 / SE ( of b1)

8
New cards

Test statistic for F test

MSR/MSE

9
New cards

Degree of freedom for t test (MLR)

N-k-1

10
New cards

Degree of freedom for f test

K/ n-k-1

11
New cards

R= 0

Weak, no Linear relationship

12
New cards

R= -1

Negative strong correlation

13
New cards

Key properties of correlation

Unit free. (Changing Units doesn’t change r) , doesn’t. depend on which variable is X or Y, only describes linear relationships

14
New cards

Regression equation

Least squares regression helps predict a response variable (y) using an exploratory variable (x) , finds best fit line for data by minimizing, the squared residuals

15
New cards

Slope

b1= r* (Sy/Sx)

16
New cards

What does slope represent

Average change in Y for a one unit increase in X

17
New cards

Sy

Standard deviation of y

18
New cards

Sx

Standard deviation of X

19
New cards

Y intercept

b0= ybar- b1*(xbar)

20
New cards

DOTS

An acronym for key features used to describe a scatterplot: Direction (positive, negative, none), Outliers, Trend/form (linear or nonlinear), and Strength (weak, moderate, strong).

21
New cards

Coefficient of Determination (R2R^2)

A metric with a range from 00 to 11 (00 to 100100 percent) that represents the percentage of variation in yy explained by the regression on xx.

22
New cards

Residual

The difference between the observed response value and predicted response value (residual=y−y^\text{residual} = y - \hat{y}); positive when observed yy is higher than predicted, and negative when lower.

23
New cards

Least Squares Regression

A regression method that chooses the line that makes the sum of squared residuals as small as possible.

24
New cards

Linear Assumption

A regression assumption requiring that the relationship looks linear and the residual plot shows random scatter.

25
New cards

Independence Assumption

A regression assumption satisfied when observations are collected using randomization or random selection.

26
New cards

Equal Variance Condition

Also known as the equal spread condition, a regression assumption requiring residuals to have consistent spread with no funnel or fan shape.

27
New cards

Normality Condition

A regression assumption requiring residuals to be bell-shaped or follow roughly a straight line on a normal probability plot, rather than an S-shape.

28
New cards

Simple Linear Regression (SLR)

A statistical technique used to predict a response variable yy, model the relationship between xx and yy, and estimate the rate of change (slope) between them.

29
New cards

Fitted Line (SLR)

The regression line derived from sample statistics, represented by the equation y^=b0+b1x\hat{y} = b_0 + b_1 x.

30
New cards

True Model (SLR)

The underlying population parameter regression model, represented by the equation y=β0+β1x+ey = \beta_0 + \beta_1 x + e.

31
New cards

Fitted line (MLR)

Y hat= b0 + b1(x) + b2(x2) ….

32
New cards

True model (MLR)

Y= beta0 + Beta1(x1) + Beta2 (x2) +…..+ e

33
New cards

t-test for Slope (SLR)

df=n−2df = n - 2.; determine if true slope is different from 0

34
New cards

Permutation / Randomization Test

A hypothesis testing method performed by rearranging data labels under the assumption that the null hypothesis H0H_0 is true.

35
New cards

Bootstrap Percentile Method for Slope

A resampling method for finding possible slope values by taking the 2.5th2.5^{\text{th}} and 97.5th97.5^{\text{th}} percentiles of a bootstrap distribution.

36
New cards

Bootstrap Standard Error Method for Slope

A method for estimating plausible slope values by calculating SEbootSE_{\text{boot}} from a bootstrap distribution and constructing the confidence interval b1±2×SEbootb_1 \pm 2 \times SE_{\text{boot}}.

37
New cards

Prediction Interval (PI)

An interval giving plausible values for an individual future response yy at a given xx; it is much wider than the confidence interval for mean yy.

38
New cards

Confidence Interval (CI) for Mean Y

An interval giving plausible values for the average or expected response yy at a given predictor value xx.

39
New cards

Multiple Linear Regression (MLR)

A regression technique that uses more than 11 predictor variable (xx) to predict a response variable yy.

40
New cards

MLR Slope Interpretation

The expected change in yy for a 11-unit increase in predictor xix_i, holding all other predictor variables constant.

41
New cards

Overall F-test

A hypothesis test for the whole model testing H0:all slopes=0H_0: \text{all slopes} = 0 against Ha:at least one slope≠0H_a: \text{at least one slope} \ne 0 using the test statistic F=MSRMSEF = \frac{MSR}{MSE}.

42
New cards

Individual t-test for slope (MLR)

A hypothesis test for ONE predictor variable testing H0:βi=0H_0: \beta_i = 0 against Ha:βi≠0H_a: \beta_i \ne 0 with df=n−k−1df = n - k - 1 to determine if that predictor is useful after accounting for all other predictors.

43
New cards

Total Sum of Squares (SST)

The measure of total variation in response variable YY, where SST=SSR+SSESST = SSR + SSE with degrees of freedom df=n−1df = n - 1.

44
New cards

Regression Sum of Squares (SSR)

The portion of total variation in YY explained by the regression model, with degrees of freedom df=kdf = k.

45
New cards

Sum of Squared Errors (SSE)

The portion of total variation in YY not explained by the regression model, with degrees of freedom df=n−k−1df = n - k - 1.

46
New cards

Mean Square Regression (MSR)

The model sum of squares divided by model degrees of freedom, computed as MSR=SSRkMSR = \frac{SSR}{k}.

47
New cards

Mean Square Error (MSE)

The sum of squared errors divided by residual degrees of freedom, computed as MSE=SSEn−k−1MSE = \frac{SSE}{n - k - 1}.

48
New cards

F-distribution

A probability distribution used to compare the ratio of two variances, indexed by two degrees of freedom parameters (kk and n−k−1n - k - 1).

49
New cards

Indicator Variable

A binary variable used to incorporate categorical variables into a regression model (1=present1 = \text{present}, 0=not present0 = \text{not present}); changes the intercept

50
New cards

Base Category

The category in a regression model without its own indicator variable, which serves as the reference level changing the y-intercept.

51
New cards

Interaction Term

(quantitative×indicator\text{quantitative} \times \text{indicator}) ; changes the slope

52
New cards

Multicollinearity

A condition in multiple regression where predictor variables are highly correlated with each other, providing redundant information to the model.

53
New cards

Multicollinearity Signs

t-tests produce large p-values ; F-test produces a small p-value, standard errors become very large.

54
New cards

Variance Inflation Factor (VIF)

A measure of multicollinearity where the lowest possible value is 11, and a value of VIF≥5\text{VIF} \ge 5 indicates a multicollinearity problem.

55
New cards

Adjusted R2R^2

useful when comparing MLR models ; accounts for the number of predictors (kk) rather than automatically rewarding adding more variables.

56
New cards

Box-Cox Transformation

A variable transformation technique for yy where λ=1\lambda = 1 results in no change, λ=0\lambda = 0 transforms to log⁡(y)\log(y), and λ=2\lambda = 2 transforms to y2y^2.