Problem Set 1 Review

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/14

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 2:13 PM on 10/6/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

15 Terms

1
New cards

prediction

  • Is y-hat close to y?

  • f-hat may be a black box


2
New cards

inference

  • How does Y change with X-sub-p?

  • f-hat must be interpretable


3
New cards

reducible error

  • How far our estimate f-hat sits from the truth f

  • A better method, better predictors, or more data can shrink this to 0


<ul><li><p>How far our estimate f-hat sits from the truth f</p></li><li><p>A better method, better predictors, or more data can shrink this to 0</p></li></ul><p></p>
4
New cards

irreducible error

  • No method, however good, gets past this floor

  • the test MSE will always stay above it


<ul><li><p>No method, however good, gets past this floor</p></li><li><p>the test MSE will always stay above it </p></li></ul><p></p>
5
New cards

parametric method

  • assume a form for f, estimate a fixed number of parameters (usually small)

  • wrong if the form is wrong → overfitting


<ul><li><p>assume a form for f, estimate a fixed number of parameters (usually small)</p></li><li><p>wrong if the form is wrong → overfitting</p></li></ul><p></p>
6
New cards

non-parametric method

  • no assumed form, estimates f directly by looking for a function that fits the observed data

  • flexible, and needs far more data (large number of observations)


7
New cards

regression

  • Y is quantitative

    • Ex: wage, price, blood pressure


8
New cards

classification

  • Y is qualitative

    • Ex: readmitted or not, labeling handwritten numbers (0-9), which customers will convert


9
New cards

training MSE

  • computed on the observations used to fit the model

  • _____ decreases as flexibility increases

  • usually smaller b/c most methods choose their parameters to minimize it

  • if the curve touches every training point, the ____ is 0

  • silver line


<ul><li><p>computed on the observations used to fit the model</p></li><li><p>_____ decreases as flexibility increases</p></li><li><p>usually smaller b/c most methods choose their parameters to minimize it</p></li><li><p>if the curve touches every training point, the ____ is 0</p></li><li><p>silver line</p></li></ul><p></p>
10
New cards

test MSE

  • computed on observations the model has never seen

  • u-shaped because of over/underfitting

  • red line


<ul><li><p>computed on observations the model has never seen</p></li><li><p>u-shaped because of over/underfitting</p></li><li><p>red line</p></li></ul><p></p>
11
New cards

degrees of freedom

  • a single number summarizing how flexible a fitted curve is

  • a smoother, more restricted curve has fewer _____ than a wiggly one


12
New cards

mean squared error

  • a metric that is used to measure how well a model’s predictions match the actual values

  • calculates the average of the squared differences between predicted and actual values


<ul><li><p>a metric that is used to measure how well a model’s predictions match the actual values </p></li><li><p>calculates the average of the squared differences between predicted and actual values</p></li></ul><p></p>
13
New cards

overfitting

  • following the noise too closely

  • low train MSE, high test MSE

  • left of the minimum, the model is too rigid to capture f

  • right of the minimum, the model is fitting the residual

  • the dashed floor is Var(residual)

  • a less flexible model would have had smaller test MSE


14
New cards

train/test split

  • [ test_size = 0.5 ] → holds out half of the 200 observations for testing

  • [ random_state = 363 ] → fixes the shuffle so that the split is identical every time the code renders


<ul><li><p>[ test_size = 0.5 ] → holds out half of the 200 observations for testing</p></li><li><p>[ random_state = 363 ] → fixes the shuffle so that the split is identical every time the code renders</p></li></ul><p></p>
15
New cards

validation set approach

  • split the data into training and testing

    • however, a different split can give a different answer

  1. fit on the training set

  2. evaluate on the test set

  3. keep the model with the smallest test error