1/21
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
bias-variance decomposition

variance
how much f-hat moves across different training sets
how much f-hat would change if estimated from a different training set
different training data gives a different f-hat
more observations shrink ____
adding an irrelevant predictor increases _____

bias
how far the average f-hat, over those training sets, sits from the truth f
the error from approximating a complicated relationship with a simpler model
a method that cannot bend enough to match f is ____
____ barely moves since it’s about the model’s form, not how much data it sees


irreducible error
neither variance nor bias touches this term
flexible methods have high variance
a flexible fit follows the training points closely, so changing even a few of those points can change the fit a lot
more flexible
low bias, high variance
less flexible
high bias, low variance
bias-variance trade-off
choosing flexibility using test error
u-shaped curve
bias falls as flexibility increases
variance increases as flexibility increases
past some point, more flexibility barely reduces bias, but raises variance a lot
moderate curve
bias and variance trade off at a middle flexibility
bias drops fast as flexibility rises, the minimum sits in the middle

near-linear
variance dominates so simple method wins
bias starts low and stays low, variance almost dominates immediately

highly non-linear
bias dominates, so flexible method wins
bias stays high until flexibility is high, variance barely rises

flexibility knobs
polynomial degree or spline degrees of freedom
K in K-nearest neighbors
number of predictors
amount of training data
K-nearest neighbors
to classify or predict at x0, average the K closest training points
small K: flexible, wiggly, high variance
large K: rigid, smooth, high bias
Bayes Error Rate
averaging that best-possible confidence over every value of X
one minus it is the lowest test error any classifier can reach

euclidian distance formula

K-nearest neighbors
to classify x-sub-0, find the K closest training points and take a vote
as K gets bigger, our flexibility goes down

curse of dimensionality
KNN works well when p is small and n is large
as p grows, the nearest neighbors stop being near
with p = 20, capturing 10% of the data needs 80% of every range