Alg for Machine Learning Final Review

0.0(0)
Studied by 0 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/237

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 12:36 PM on 7/30/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

238 Terms

1
New cards

Steps of KKN

- Choose K

- Compute distance from test point to all training points

- Sort distances

- Pick k closest

- Majority vote

2
New cards

What is K in KKN?

The number of neighbors used in voting

3
New cards

In KNN what happens if one feature has much larger values than others?

The large value dominates the distance

4
New cards

PCA

The technique in KNN is used to reduce dimensionality

5
New cards

Euclidean Distance Formula

d = √((x2 - x1)^2 + (y2 - y1)^2)

6
New cards

Manhattan Distance Formula

| x1-x2 | + | y1-y2 |

7
New cards

Hamming Distance

Counting the different binary positions

8
New cards

In the Minkowski Distance formula what value give Manhattan Distance?

1

9
New cards

In the Minkowski Distance formula what value give Euclidean Distance?

2

10
New cards

What property is not required for Minkowski Distance metric?

Linearity

11
New cards

Effect of a Small K in KKN

Overfitting

12
New cards

Effect of a large K in KKN

Underfitting

13
New cards

KKN's Main Flaw

Curse of Dimensionality

14
New cards

Curse of Dimensionality

High-dimensional data requires more samples

15
New cards

Naive Bayes

Probabilistic classifier that calculates likelihood of each class and chooses the most likely

16
New cards

Naive Bayes Steps

- Compute prior P(c)

- Compute likelihoods

- Multiply

- Choose max

17
New cards

What does Naive Bayes produce?

Multiple probabilities of classes

18
New cards

Naive Bayes Assumption

Independent Features

19
New cards

What happens if Naive Bayes' assumption is violated?

Accuracy may decrease

20
New cards

What happens if Naive Bayes encounters a feature value not seen in training?

Probability becomes zero

21
New cards

What happens if a probability becomes zero in Naive Bayes?

Entire product becomes zero

22
New cards

NB Normalized Term

P(X)

23
New cards

NB Prior Class

P(Y)

24
New cards

NB Likelihood

P(X|Y)

25
New cards

NB Posterior

P(Y|X)

26
New cards

Laplace Smoothing

Technique to handle zero probabilities in classification

27
New cards

Support Vector Machines (SVM)'s Purpose

Finding the best separating hyperplane / margin between classification classes in a model

28
New cards

SVM Decision Boundary Formula

(w^T)x + b = 0

29
New cards

Support Vectors

Data points closest to the decision boundary that determine the margin

30
New cards

What are support vectors ins SVM?

Closest points to boundary that determine the hyperplane

31
New cards

SVM Margin

The distance between the boundary and closest points

32
New cards

Small C in SVM

- Simpler Decision Boundary

- Larger Margin

- High bias, low variance

- Risks underfitting

33
New cards

Large C in SVM

- Penalizes error heavily

- Smaller Margin

- Lower bias, high variance

- Risks overfitting

34
New cards

Margin Formula

2/||w||

35
New cards

SVM Optimization Formulas

- Hard Margin

- Soft Margin

- Hinge Loss

36
New cards

Hinge Loss

Penalizes points that are inside margin or have been misclassified

37
New cards

SVM C Variable

Margin vs misclassification tradeoff

38
New cards

SVM Kernel Trick Purpose

Transforms data to higher dimensions

39
New cards

What happens when features are not scaled in SVM?

Margin becomes skewed

40
New cards

What is the expected outcome of a large Gamma(Y) In Guassian RBF?

Overfitting & Wiggly Boundary

41
New cards

What is the expected outcome of a small Gamma(Y) In Guassian RBF?

Underfitting & Smooth boundaries

42
New cards

What are the assumptions when using Gaussian Naive Bayes?

- Independent Features

- Continuous Features

- Normal Distribution

43
New cards

What are the required parameters of GNB

- mean

- variance

44
New cards

What type of algorithm is SOFTMAX Regression?

Multi-Class Classification

45
New cards

What algorithm is SOFTMAX Regression a version of?

Logistic Regression

46
New cards

What does Softmax Regression Produce?

One probability value per class

47
New cards

Range of Softmax Regression

Sum to 1

48
New cards

What does SOFTMAX produce?

- Outputs from range sum to 1

- Probabilities

49
New cards

What is the assumption when using SOFTMAX?

Classes are mutually exclusive

50
New cards

How to calculate total parameters in SOFTMAX Regression?

(features * classes) + classes

51
New cards

In softmax, what happens if one logit is much larger?

That class gets probability ≈ 1

52
New cards

What models are most sensitive to unscaled features?

- KNN

- SVM

53
New cards

Decision Tree Model

Supervised learning model that makes predictions by splitting data into branches based on feature values, forming a tree of decisions that leads to a final output

54
New cards

What is the purpose of splitting node in a Decision Tree?

To reduce impurity of child nodes

55
New cards

What is impurity in a Decision Tree's node?

Different classes within the same node

56
New cards

What are the impurity measures in a Decision Tree?

- Gini Impurity

- Entropy

57
New cards

What is the faster impurity measure to compute in Decision Trees?

Gini

58
New cards

What is the relationship between Decision Trees and Scaling?

Scaling is not required

59
New cards

Information Gain

Reduction in impurity after a split in a Decision Tree

60
New cards

Information Gain Formula

(Root Entropy) - (Weighted Entropy After Split)

61
New cards

What does increasing max_depth within reason do?

It helps the decision tree memorize training data and reduced underfitting

62
New cards

C4.5

Decision tree algorithm that extends ID3 and can handle continous attributes

63
New cards

CART

Decision tree algorithm that constructs binary trees via binary splits

64
New cards

ID3

Decision tree algorithm used for classification

65
New cards

What happens when max_depth is increased too much?

The decision tree memorizes too much training data/overtrains causing overfitting

66
New cards

What does it mean to prune a Decision Tree?

To remove unnecessary branches to improve generalization and reduce overfitting

67
New cards

min_samples_split

Parameter that controls splitting threshold and controls minimum samples

68
New cards

Entropy Increases, Information Gain __________

Decreases

69
New cards

Entropy Decrease, Information Gain __________

Increases

70
New cards

Where does max entropy occur?

p = 0.5

71
New cards

Max Entropy

1

72
New cards

What condition leads to 0 entropy?

A single-class / pure node

73
New cards

What expression represents Gini impurity for classes with probabilities p_i?

1 - Σ p_i^2

74
New cards

What do we look for to determine which entropy value is better?

The lowest value, the lower the value the higher the information gain

75
New cards

Steps of a Neural Network

- Take inputs

- Apply functions

- Produce an output

76
New cards

What Neural Network libraries are used in this course?

- Scikit-learn (MLPClassifier)

- TensorFlow

- Keras

77
New cards

Ensemble Learning

The combination of multiple models to improve overall performance by reducing error

78
New cards

Decision Tree

A single model trained on a full dataset prone to high variance/overfitting

79
New cards

Random Forest

A multi-model/ensemble of decision trees

80
New cards

Voting Classifier

Combines the class predictions from multiple models

81
New cards

Types of Voting Classifier

- Hard Voting

- Soft Voting

82
New cards

Hard Voting

Based on majority class

83
New cards

Soft Voting

Based on the nearest averaged predicted probabilities and selected the highest

84
New cards

What does Bagging/Pasting do?

They create and train multiple models independently and in parallel using subsets of data

85
New cards

Bagging

- Samples with replacement- Duplicates exist within the model- Uses Out-of-Bag Evaluation

86
New cards

Pasting

- Samples without replacement- No duplicates within the model

87
New cards

Out-of-Bag Evaluation

Takes the samples left out during the process of bagging and uses them as a validation set

88
New cards

Boosting

Models trained sequentially each focused on previous errors to reduce bias

89
New cards

Centroid

The average of all point in a cluster

90
New cards

Cluster's Job

Groups similar data points of unlabeled data & discovers structure/patterns

91
New cards

Dendrogram

A tree of clusters

92
New cards

Hierarchical Clustering

Buildings dendrogram in which each node represents a cluster of clusters

93
New cards

Agglomerative Clustering

Uses a proximity matrix and a distance formula to merge points together

94
New cards

DBSCAN

A clustering alg that groups based on density instead of distance

95
New cards

Pros of DBSCAN

- Finds arbitrary shapes

- Handles noise and outliers well

96
New cards

What are the challenges of Natural Language Processing?

- Human language is ambiguous, context-dependent, & unstructured

- Text is messy & doesn't fit tables

- Requires understanding meaning not just keywords

97
New cards

Bag of Words

Creates an occurrence matrix that tracks frequency of words, but ignores grammar and order

98
New cards

Limits of Bag of Words

- Does not preserve context

- Dominated by common words

- Understands semantics poorly

99
New cards

TF-IDF

Creates an occurrence matrix that tracks frequency of words but penalizes common words

100
New cards

Tokenization

- Splits text into words/tokens

- Removes punctuation