SLR: Assessing the Model

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/36

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 12:06 PM on 10/8/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

37 Terms

1
New cards

What is a fitted value y^i\hat{y}_i?

y^i=β^0+β^1xi\hat{y}_i = \hat\beta_0 + \hat\beta_1 x_i, the estimated mean response at xix_i (a point on the fitted line).

2
New cards

What is a residual eie_i?

ei=yi−y^ie_i = y_i - \hat{y}_i (observed minus fitted), also called the crude residual. It estimates the error εi\varepsilon_i.

3
New cards

What is SSESS_E, what does it measure, and how many d.f. does it have?

SSE=∑(yi−y^i)2=∑ei2SS_E = \sum (y_i - \hat{y}_i)^2 = \sum e_i^2: variation about the fitted line; the minimum of S(β0,β1)S(\beta_0, \beta_1). d.f. n−2n - 2. Shortcut: SSE=Syy−SSRSS_E = S_{yy} - SS_R.

4
New cards

What is SSESS_E for the constant model?

SSE=SSTSS_E = SS_T, since y^i=yˉ\hat{y}_i = \bar{y}.

5
New cards

What is SSTSS_T, what does it measure, and how many d.f. does it have?

SST=∑(yi−yˉ)2SS_T = \sum (y_i - \bar{y})^2: total variation in yy about its mean. d.f. n−1n - 1.

6
New cards

What is SSRSS_R, what does it measure, and how many d.f. does it have?

SSR=∑(y^i−yˉ)2SS_R = \sum (\hat{y}_i - \bar{y})^2: variation explained by the fitted model. d.f. 11. Shortcut: SSR=β^12Sxx=β^1SxySS_R = \hat\beta_1^2 S_{xx} = \hat\beta_1 S_{xy}.

7
New cards

What is the ANOVA identity?

SST=SSR+SSESS_T = SS_R + SS_E (the cross term vanishes because ∑ei=0\sum e_i = 0 and ∑xiei=0\sum x_i e_i = 0).

8
New cards

What is the ANOVA table for simple linear regression?

Regression: d.f. 11, SSRSS_R, MSRMS_R, F=MSRMSEF = \frac{MS_R}{MS_E}. Residual: d.f. n−2n - 2, SSESS_E, MSEMS_E. Total: d.f. n−1n - 1, SSTSS_T.

9
New cards

What are degrees of freedom, and why do SSTSS_T, SSESS_E and SSRSS_R have n−1n-1, n−2n-2 and 11?

The number of independent pieces of information that go into an estimate. SSTSS_T: n−1n - 1 (one used by yˉ\bar{y}). SSESS_E: n−2n - 2 (two used by β^0,β^1\hat\beta_0, \hat\beta_1). SSRSS_R: (n−1)−(n−2)=1(n-1) - (n-2) = 1.

10
New cards

What are the mean squares MSRMS_R and MSEMS_E?

Sum of squares divided by its d.f.: MSR=SSR1MS_R = \frac{SS_R}{1}, MSE=SSEn−2MS_E = \frac{SS_E}{n-2}.

11
New cards

What is the variance ratio FF, and what does it measure?

F=MSRMSEF = \frac{MS_R}{MS_E}: variation explained by the model relative to variation due to residuals.

12
New cards

How is the F distribution defined?

If X∼χν12X \sim \chi^2_{\nu_1} and Y∼χν22Y \sim \chi^2_{\nu_2} are independent, X/ν1Y/ν2∼Fν1,ν2\frac{X/\nu_1}{Y/\nu_2} \sim F_{\nu_1, \nu_2}. Skewed; ν1\nu_1 numerator d.f., ν2\nu_2 denominator d.f.

13
New cards

How do you carry out the F test for the slope?

H0:β1=0H_0: \beta_1 = 0 vs H1:β1≠0H_1: \beta_1 \neq 0. Under H0H_0, F=MSRMSE∼Fn−21F = \frac{MS_R}{MS_E} \sim F^1_{n-2}. Reject H0H_0 at level α\alpha if Fcal>Fn−21(α)F_{cal} > F^1_{n-2}(\alpha).

14
New cards

What does rejecting H0:β1=0H_0: \beta_1 = 0 mean?

The slope is non-zero, so the full model yi=β0+β1xi+εiy_i = \beta_0 + \beta_1 x_i + \varepsilon_i is better than the constant model yi=β0+εiy_i = \beta_0 + \varepsilon_i.

15
New cards

What is E(SSE)E(SS_E)?

E(SSE)=(n−2)σ2E(SS_E) = (n - 2)\sigma^2.

16
New cards

What is the unbiased estimator of σ2\sigma^2?

S2=MSE=SSEn−2=1n−2∑(yi−y^i)2S^2 = MS_E = \frac{SS_E}{n-2} = \frac{1}{n-2}\sum (y_i - \hat{y}_i)^2. In R it appears as "Residual standard error" =MSE= \sqrt{MS_E}.

17
New cards

Is MSEMS_E the sample variance?

Not in the full model. Only in the constant model (y^i=yˉ\hat{y}_i = \bar{y}, d.f. n−1n - 1) is S2=1n−1∑(yi−yˉ)2S^2 = \frac{1}{n-1}\sum (y_i - \bar{y})^2.

18
New cards

What is the coefficient of determination R2R^2?

R2=SSRSST=1−SSESSTR^2 = \frac{SS_R}{SS_T} = 1 - \frac{SS_E}{SS_T} (× 100%): the percentage of total variation in yy explained by the fitted model.

19
New cards

What is the range of R2R^2, and what do R2=0R^2 = 0 and R2=100%R^2 = 100\% mean?

R2∈[0,100]%R^2 \in [0, 100]\%. R2=0R^2 = 0: the model explains none of the variability. R2=100%R^2 = 100\%: SSE=0SS_E = 0, all points on the line.

20
New cards

What is the main caveat when interpreting R2R^2?

It measures linear association only; a small R2R^2 does not always mean a poor relationship (e.g. it may be quadratic).

21
New cards

What is R2R^2 when all the yy values are equal?

Syy=0S_{yy} = 0, so R2=00R^2 = \frac{0}{0} is undefined; it is typically taken as 00.

22
New cards

How is R2R^2 defined for a no-intercept model, and what is the caution?

Use SST=∑yi2SS_T = \sum y_i^2, so R2=1−SSE∑yi2R^2 = 1 - \frac{SS_E}{\sum y_i^2}. It can be artificially high and is not comparable with the usual R2R^2.

23
New cards

What is the formula for adjusted R2R^2?

Radj2=1−(1−R2)n−1n−k−1R^2_{adj} = 1 - (1 - R^2)\frac{n-1}{n-k-1}, kk = number of predictors. For SLR: Radj2=1−n−1n−2(1−R2)R^2_{adj} = 1 - \frac{n-1}{n-2}(1 - R^2).

24
New cards

Why use adjusted R2R^2 instead of R2R^2?

It penalises extra parameters, so it can compare models with different numbers of predictors. R2R^2 always increases when a regressor is added; Radj2R^2_{adj} increases only if that variable's FF statistic is greater than 11.

25
New cards

What are the properties of adjusted R2R^2?

Radj2≤R2R^2_{adj} \le R^2; it can be negative (model worse than the mean); it equals 11 for a perfect fit; it is close to R2R^2 when nn is large.

26
New cards

What are the key properties of the residuals?

∑ei=0\sum e_i = 0 (so eˉ=0\bar{e} = 0), ∑xiei=0\sum x_i e_i = 0, ∑y^iei=0\sum \hat{y}_i e_i = 0, and E[ei]=0E[e_i] = 0.

27
New cards

What is the variance of the residual eie_i?

var(ei)=σ2(1−vi)\text{var}(e_i) = \sigma^2(1 - v_i) with vi=1n+(xi−xˉ)2Sxxv_i = \frac{1}{n} + \frac{(x_i - \bar{x})^2}{S_{xx}}. Not constant: it depends on ii.

28
New cards

What is the covariance of two residuals eie_i and eje_j?

cov(ei,ej)=−σ2(1n+(xi−xˉ)(xj−xˉ)Sxx)\text{cov}(e_i, e_j) = -\sigma^2\left(\frac{1}{n} + \frac{(x_i - \bar{x})(x_j - \bar{x})}{S_{xx}}\right), not 00. So residuals do not quite mimic the errors.

29
New cards

What is the difference between an error εi\varepsilon_i and a residual eie_i?

The error εi\varepsilon_i comes from the model (unobservable); the residual eie_i comes from fitting the model to the data.

30
New cards

What is the standardised residual did_i, and why use it?

di=eis2(1−vi)d_i = \frac{e_i}{\sqrt{s^2(1 - v_i)}}, vi=1n+(xi−xˉ)2Sxxv_i = \frac{1}{n} + \frac{(x_i - \bar{x})^2}{S_{xx}}. More nearly constant variance and smaller covariance than eie_i.

31
New cards

Which residual plot checks linearity?

Plot did_i against xix_i.

32
New cards

Which residual plot checks constant variance?

Plot did_i against the fitted values y^i\hat{y}_i.

33
New cards

What does an acceptable residual plot look like?

Random scatter around zero with roughly constant spread and no pattern.

34
New cards

What patterns make a residual plot unacceptable?

Fan shape (non-constant variance), curve (non-linearity), outliers, systematic pattern (missing predictor), pattern in observation order (autocorrelation), influential high-leverage point, different spread by group.

35
New cards

How do you check the normality assumption?

Normal QQ plot of the residuals: points should lie close to a straight line. Heavy tails or skewness show as departures. A formal test is better (later in the course).

36
New cards

What assumption do the F and t tests need?

Normality of the errors. If it fails, the tests may not be valid.

37
New cards